# goodfire.ai > AI-optimized mirror of goodfire.ai containing 75 pages totalling 235,799 words of clean markdown content, structured data, and semantic HTML. Original source: https://goodfire.ai. Last updated: 2026-07-26T03:26:11.070Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [preview/index.html](/content/preview/index.html) (649 words) - [platform/index.html](/content/platform/index.html) (630 words) - [docs/index.html](/content/docs/index.html) (317 words) - [EVEE — Evo Variant Effect Explorer](/content/evee/index.html) (51 words) - [Paint with Ember — Goodfire](/content/paint/index.html) (31 words) - [Trust Center - Goodfire](/content/trust/index.html): Goodfire is a research company using interpretability to understand, learn from, and design AI systems. Our mission is to build the next generation of safe and powerful AI—not by scaling alone, but by understanding the intelligence we're building. (148 words) - [Goodfire AI](/content/site-root.html): Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. (657 words) ## Articles & Blog Posts - [Not Found](/content/join-waitlist/index.html) (11 words) - [Paint with Ember — Goodfire](/content/paint/umap-html.html) (33 words) - [Not Found](/content/papers/goodfire-ember/index.html) (11 words) - [Trust Center - Goodfire](/content/trust/controls/index.html): Goodfire is a research company using interpretability to understand, learn from, and design AI systems. Our mission is to build the next generation of safe and powerful AI—not by scaling alone, but by understanding the intelligence we're building. (633 words) - [Not Found](/content/blog/announcing-goodfire-ember/index.html) (11 words) - [Trust Center - Goodfire](/content/trust/resources/index.html): Goodfire is a research company using interpretability to understand, learn from, and design AI systems. Our mission is to build the next generation of safe and powerful AI—not by scaling alone, but by understanding the intelligence we're building. (103 words) - [static/vpd-explainer/explainer-md.html](/content/static/vpd-explainer/explainer-md.html) (854 words) - [Trust Center - Goodfire](/content/trust/subprocessors/index.html): Goodfire is a research company using interpretability to understand, learn from, and design AI systems. Our mission is to build the next generation of safe and powerful AI—not by scaling alone, but by understanding the intelligence we're building. (58 words) - [static/vpd-blog-post/post-md.html](/content/static/vpd-blog-post/post-md.html) (40,571 words) - [Not Found](/content/blog/interpreting-evo-2.html) (11 words) - [The World Inside Neural Networks](/content/research/the-world-inside-neural-networks/index.html): How neural geometry will unlock understanding and control of AI (3,024 words) - [Replicating Circuit Tracing for a Simple Known Mechanism](/content/research/replicating-circuit-tracing-for-a-simple-mechanism/index.html): Recent work by Ameisen et al. and Lindsey et al. introduced Cross-Layer Transcoders (CLTs) as a method for characterizing information processing across layers and token positions in Transformer language models. We train CLTs on GPT-2 Small and evaluate their ability to recover known computational mechanisms in a well-understood task. (5,222 words) - [Finding the Tree of Life in Evo 2](/content/research/phylogeny-manifold/index.html): In this research update, we uncover how Evo 2, a DNA foundation model, represents the “tree of life”—the phylogenetic relationships between species. We find that phylogeny is encoded geometrically in the distances along a curved manifold, one of the most complex manifold examples yet found (to our knowledge) in a foundation model. Our results support an emerging picture of feature manifolds—that they tend to have a dominant flat representation (with respect to the ambient space) plus higher curvature deviations—and point to both better ways of understanding scientific AI models and better interpretability techniques. (3,306 words) - [Using Interpretability to Identify a Novel Class of Alzheimer's Biomarkers](/content/research/interpretability-for-alzheimers-detection/index.html) (4,019 words) - [Deploying Interpretability to Production with Rakuten: SAE Probes for PII Detection](/content/research/rakuten-sae-probes-for-pii-detection/index.html) (3,030 words) - [Master Services Agreement](/content/legal/msa/index.html): Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. (4,340 words) - [Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers](/content/research/bsf-vision/index.html): We introduce Block-Sparse Featurizers (BSF), a family of methods to decompose a model’s activations into multidimensional subspaces rather than single directions. Applied to vision models, we find that BSFs find interpretable, multidimensional features which offer a more parsimonious explanation of model internals; that those features enable fine-grained steering; and that most concepts in the models are multidimensional. (2,696 words) - [Understanding and Steering Llama 3 with Sparse Autoencoders](/content/research/understanding-and-steering-llama-3/index.html) (3,284 words) - [Terms of Use](/content/legal/tos/index.html): Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. (4,651 words) - [Painting With Concepts Using Diffusion Model Latents](/content/research/painting-with-concepts/index.html) (2,889 words) - [Interpreting Evo 2: Arc Institute's Next-Generation Genomic Foundation Model](/content/research/interpreting-evo-2/index.html) (1,731 words) - [Intentionally Designing the Future of AI](/content/blog/intentional-design/index.html) (3,917 words) - [Covariance-based Sequence Pooling](/content/research/covariance-pooling/index.html): Covariance pooling: a better replacement for mean pooling that improves probing of genomic foundation model embeddings. (2,260 words) - [Latest research - Goodfire](/content/research/index.html): Our latest interpretability research to understand and intentionally design advanced AI systems. (988 words) - [You and Your Research Agent: Lessons From Using Agents for Interpretability Research](/content/blog/you-and-your-research-agent/index.html): We’ve been using AI agents to assist with interpretability research for the past several months. In this post, we’re sharing some of the lessons we’ve learned: how agents for experimentation are different than agents for software development, where they perform well, and where they still fall short. We’re also open sourcing a basic implementation of the most important tool we’ve found for enabling effective agentic experimentation — a general-purpose Jupyter server/notebook MCP package — and a suite of interpretability tasks that can be run using the tool. (3,476 words) - [Explaining 4.2 million genetic variants with state-of-the-art, interpretable predictions](/content/research/evee-explaining-genetic-variants/index.html): State-of-the-art, interpretable variant effect prediction for all 4.2 million ClinVar variants. A collaboration between Goodfire and Mayo Clinic. (2,136 words) - [On Optimism for Interpretability](/content/blog/on-optimism-for-interpretability/index.html): Why interpretable AI is achievable and essential: Eric Ho explains how mechanistic interpretability can transform opaque neural networks into understandable, debuggable systems we can trust and control. (2,489 words) - [Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training](/content/research/probe-based-data-attribution/index.html) (1,346 words) - [Towards Scalable Parameter Decomposition](/content/research/stochastic-param-decomp/index.html): The most successful methods so far, like SAEs, have focused largely on the activations that flow through a model, rather than directly inspecting the weights that transform and guide the inputs into those flows. It's a bit like trying to understand a program by only looking at its runtime variables, but never its source code. Parameter decomposition offers a way to decompose a model’s parameters—the 'source code'—into components that reveal not only what the network computes, but how it computes it. Today, we're releasing a paper on Stochastic Parameter Decomposition (SPD), which removes key barriers to the scalability of prior methods. (1,682 words) - [Paper Summary: Interpreting Language Model Parameters](/content/research/vpd-explainer/index.html) (2,006 words) - [Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train](/content/research/predictive-data-debugging/index.html): Given a preference dataset, we can accurately predict which behaviors RL will amplify or suppress before you train, trace them back to the responsible data, and reshape the dataset and/or training process to prevent undesired effects. (2,313 words) - [Pilot Agreement](/content/legal/pilot-agreement/index.html): Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. (2,332 words) - [Understanding Memorization via Loss Curvature](/content/research/understanding-memorization-via-loss-curvature/index.html) (1,897 words) - [Using Self-Correcting Search to Accelerate Materials Discovery](/content/research/self-correcting-search/index.html) (1,742 words) - [Features as Rewards: Using Interpretability to Reduce Hallucinations](/content/research/rlfr/index.html) (2,067 words) - [Understanding, Learning From, and Designing AI: Our Series B](/content/blog/our-series-b/index.html) (1,411 words) - [Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Model](/content/blog/interpretability-infra-at-frontier-scale/index.html) (2,048 words) - [Steering Along Manifolds to Control Neural Networks](/content/research/manifold-steering/index.html) (1,100 words) - [Mapping the Latent Space of Llama 3.3 70B](/content/research/mapping-latent-spaces-llama/index.html) (746 words) - [Can SAEs Capture Neural Geometry?](/content/research/can-saes-capture-neural-geometry/index.html): AKA, how to use straight lines to capture curved geometry in neural networks (1,134 words) - [Announcing Open-Source SAEs for Llama 3.3 70B and Llama 3.1 8B](/content/blog/sae-open-source-announcement/index.html) (622 words) - [Feature Steering for Reliable and Expressive AI Engineering](/content/blog/feature-steering-for-reliable-and-expressive-ai-engineering.html) (1,126 words) - [Reasoning Theater: Probing for Performative Chain-of-Thought](/content/research/reasoning-theater/index.html): Our new paper uses probes to track “performative chain-of-thought”: when models “know” their final answer but continue to generate chain-of-thought anyways. We find that its occurrence tracks with task difficulty, and show that probes can enable early exit from reasoning traces, saving up to 68% of tokens with minimal accuracy loss. (1,566 words) - [Meandering on Manifolds: The Neural Geometry of Stories Over Time](/content/research/stories-in-space/index.html): To fully understand LLM representations, we must understand how they change dynamically, over the course of a prompt or conversation. We investigate these temporal dynamics with a simple case study: how do LLMs represent human emotions while reading short stories, both geometrically (in activation space) and temporally (changing from sentence to sentence)? (1,471 words) - [Longfact++ Rollout Viewer - Features as Rewards (RLFR)](/content/demos/hallucinations-viewer/index.html): Selected rollouts of Gemma 3 12B on LongFact++ before and after RLFR to detect and reduce hallucinations. (1,212 words) - [The Neural Geometry Series](/content/research/neural-geometry/index.html) (515 words) - [Interpreting Language Model Parameters](/content/research/interpreting-lm-parameters/index.html) (49,044 words) - [Stanford Guest Lectures: AP293 (Fall 2025)](/content/blog/ap293-guest-lectures-25/index.html) (488 words) - [Announcing Our $50M Series A to Advance AI Interpretability Research](/content/blog/announcing-our-50m-series-a/index.html) (579 words) - [Announcing Goodfire’s Fellowship Program for Interpretability Research](/content/blog/fellowship-fall-25/index.html): We’re excited to announce that we’ll be bringing on several Research Fellows and Research Engineering Fellows this fall for our fellowship program. Fellows will collaborate with senior members of our technical staff, contribute to core projects, and work full time in person in our San Francisco office. For exceptional candidates, there will be an opportunity to convert to full time research positions. (591 words) - [Goodfire Silico for Life Sciences](/content/life-sciences/index.html): Uncharted discoveries, unlocked from the models you built. (338 words) - [Announcing our SOC 2 Type II Certification](/content/blog/soc-2-type-ii/index.html) (339 words) - [Build AI models the way you write software](/content/silico/index.html): Understand and debug your AI model with Silico. (305 words) - [Goodfire](/content/customer-stories/rakuten/index.html) (488 words) - [Goodfire Announces Collaboration to Advance Genomic Medicine with AI Interpretability](/content/blog/mayo-clinic-collaboration/index.html): Goodfire is excited to announce a collaboration with Mayo Clinic seeking to unlock new frontiers in genomic medicine through AI interpretability. This collaboration aims to combine Goodfire's work in interpretability of AI models with Mayo Clinic's medical expertise and investment in AI. AI interpretability is a field devoted to understanding what AI models learn and how they produce their outputs, rather than treating them as black boxes. (458 words) - [Goodfire Silico for Language Models](/content/language/index.html): Train language models with the precision of writing software. (318 words) - [Goodfire](/content/customer-stories/prima-mente/index.html) (534 words) - [Partnering with Radical AI to Advance Materials Science With Interpretability](/content/blog/radical-partnership-announcement/index.html): We're excited to announce a new partnership between Radical AI and Goodfire to fundamentally dismantle the black box of AI-driven materials discovery and design. Radical AI and Goodfire are building generative models to intelligently forge materials for a specific purpose, directly creating a material based on its desired function. Radical AI integrates proprietary AI modeling with in-house autonomous laboratory systems to create a powerful self-learning, closed-loop system for materials discovery, testing, and development. They recently announced a $55 million Series Seed+ funding round led by RTX Ventures with participation from NVentures (NVIDIA's VC arm) and other investors. Goodfire builds interpretability tools, techniques, and infrastructure to (417 words) - [Goodfire Silico for Robotics & Vision Models](/content/robotics-vision/index.html): Real understanding before real-world deployments. (250 words) - [Our Approach to Safety at Goodfire](/content/blog/our-approach-to-safety/index.html) (345 words) - [Company](/content/company/index.html): Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. (229 words) - [Careers](/content/careers/index.html): Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. (163 words) - [Dolci DPO × Llama 3.1 8B SFT  |  L20  |  20000 pair deltas  |  K=814](/content/demos/dolci-viewer/index.html): Explore the Dolci DPO dataset and the concepts that it will teach a model, surfaced via predictive data debugging. (32,457 words) - [Under the Hood of a Reasoning Model](/content/research/under-the-hood-of-a-reasoning-model/index.html) (2,425 words) - [Verbalized Eval Awareness Inflates Measured Safety](/content/research/verbalized-eval-awareness-inflates-measured-safety/index.html) (16,775 words) - [Discovering Undesired Rare Behaviors via Model Diff Amplification](/content/research/model-diff-amplification/index.html): Model diff amplification (also known as logit diff amplification, or LDA) is a simple method for efficiently identifying rare, unexpected effects of a training run on model behavior. It's useful tool for red-teaming and evaluation, monitoring training runs, detecting emergent misalignment, and detecting backdoors/sleeper agents. (2,375 words) - [Blog - Goodfire](/content/blog/index.html): Company updates, partnerships, and product demos from the Goodfire team. (240 words) ## About Pages - [Contact](/content/contact/index.html): Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. (68 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/robots.txt): Crawler directives