Mechanistic Interpretability
Key Views & Dialogues
Mechanistic Interpretability: Philosophy, Practice & Progress with Goodfire’s Daniel & Tom
- 🗓️ Date:
2025-05-29| 🎙️ Show:The Cognitive Revolution
Mechanistic interpretability now has the raw materials of a “proto-paradigm”: understandable features, linear directions, superposition, and circuits, despite lacking consensus. Existing tools already help customers, with scientific discovery and inference-time guardrails emerging as important applications while algorithms remain the softer bottleneck. Yet a scaling study found reconstruction curves bending before 99.99% or 100%, leaving memorization, higher-order geometry, or noise unresolved as Goodfire advances its $50 million independent category bet.
View Dialogue Notes & Key Takeaways
Mechanistic interpretability now has the raw materials of a “proto-paradigm,” even if the field still lacks consensus. Tom McGrath’s stack is concrete: neural networks contain understandable features; linear directions encode them, magnitude carries intensity, superposition packs far more concepts than dimensions, and features connect into circuits. These foundations do not prove that today’s tools fully recover models.
Goodfire sees algorithms as the softer current bottleneck, while insisting existing tools already create customer value. Tom says the test is whether the field could extract $1 million of information from “a million-dollar interpreter model training run”; today, it could spend the money but not productively. Daniel Balsam’s commercial framing is more immediate: an SAE is “a window into the model,” and training it is “where the work begins, not where it ends.”
Sparse autoencoders are improving, but evidence of unreconstructed “dark matter” prevents treating them as complete model audits. JumpReLU, BatchTopK, end-to-end SAEs, Matryoshka-style decompositions and other newer approaches improve or extend SAE methods, yet one scaling study suggested reconstruction curves bend before reaching 99.99% or 100%. That missing variance might be memorization, higher-order geometry or output-irrelevant noise; Tom’s stated preference is still for interpreter models to approach effectively perfect reconstruction.
Scientific discovery is a major application because capable scientific models may contain abstractions that human methods have not found. Goodfire’s Arc Institute work recovered features strongly correlated with known genomic concepts and is moving toward unsupervised searches for novel biology in models such as Evo 2. Dan’s stronger call is that even “geniuses in a data center” might prefer mechanistic interpretability: simulate physical systems on chips, extract principles from the learned computation, then better vet hypotheses before wet-lab work.
Inference-time guardrails offer a clear near-term enterprise use case. Prompt rules become “playing Whac-A-Mole,” 1,000 rules degrade task performance, frontier-model judges add another expensive call, and smaller classifiers require data and ML expertise. Goodfire instead proposes cheap monitoring of internal cognition that can trigger programmatic responses when a model appears to engage with PII or prohibited topics—while explicitly conceding that jailbreaks and other limitations remain.
Creative controls could turn latent understanding into a new interface layer, but the broader multimodal-superintelligence thesis remains contested. Paint with Ember demonstrates direct manipulation of image features, with video and music offering higher-value extensions such as asking for “a little more saxophone in the saxophone solo.” Nathan expects reasoning integrated across roughly 20 modalities, many involving scientific or natural-world domains, to produce superintelligence; Dan and Tom warn that scarce reward signals could instead yield “two models in a trench coat” without deep cross-domain understanding.
Goodfire has raised $50 million to pursue interpretability as an independent category rather than an accessory inside a scaling lab. Menlo Ventures led the round, with $1 million from Anthropic described as its first corporate investment; the host disclosed being a seed investor. The bet spans scientific discovery, guardrails and creative tooling, behind Dan’s claim that “interpretability is as big as AI itself.”
🔗 Original source & video: Mechanistic Interpretability: Philosophy, Practice & Progress with Goodfire’s Daniel & Tom