research/
29 pages · Updated July 26, 2026
Pages
- The World Inside Neural Networks
- Replicating Circuit Tracing for a Simple Known Mechanism
- Finding the Tree of Life in Evo 2
- Using Interpretability to Identify a Novel Class of Alzheimer's Biomarkers
- Deploying Interpretability to Production with Rakuten: SAE Probes for PII Detection
- Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers
- Understanding and Steering Llama 3 with Sparse Autoencoders
- Painting With Concepts Using Diffusion Model Latents
- Interpreting Evo 2: Arc Institute's Next-Generation Genomic Foundation Model
- Covariance-based Sequence Pooling
- Latest research - Goodfire
- Explaining 4.2 million genetic variants with state-of-the-art, interpretable predictions
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
- Towards Scalable Parameter Decomposition
- Paper Summary: Interpreting Language Model Parameters
- Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train
- Understanding Memorization via Loss Curvature
- Using Self-Correcting Search to Accelerate Materials Discovery
- Features as Rewards: Using Interpretability to Reduce Hallucinations
- Steering Along Manifolds to Control Neural Networks
- Mapping the Latent Space of Llama 3.3 70B
- Can SAEs Capture Neural Geometry?
- Reasoning Theater: Probing for Performative Chain-of-Thought
- Meandering on Manifolds: The Neural Geometry of Stories Over Time
- The Neural Geometry Series
- Interpreting Language Model Parameters
- Under the Hood of a Reasoning Model
- Verbalized Eval Awareness Inflates Measured Safety
- Discovering Undesired Rare Behaviors via Model Diff Amplification