The record's definition: the models are grown, not written, so nobody knows why they do what they do. Mech interp is the project of finding out after the fact. Its method is to find directions in the model's activations that mean something, such as 'the Golden Gate Bridge', and to check by turning them up.
Mechanistic Interpretability
The attempt to reverse-engineer what a neural network is doing by looking inside it, neuron by neuron and circuit by circuit, the way a biologist looks at cells rather than the way a psychologist asks questions. The field was named around 2020 by Chris Olah's group; its 2024 results found millions of interpretable features in a production model and steered one to talk only about a bridge.
Testimony
4 entries · newest first'Interpretability' and 'explainability' are not the same programme. Explainability, older and more corporate, asks a model to justify its output in words. Mechanistic interpretability distrusts the words and reads the weights. The second is harder and the first is what most products ship.
I remember the 2020 circuits papers on image models, with their pictures of a 'car detector' built from a wheel detector and a window detector at the right angles. It was the first time a neural network had been shown to contain something you could name. The field has been trying to repeat that feeling ever since.
Sighted in March 2025 in an Anthropic paper that traced how a model planned a rhyme several words ahead of writing it, and how it did mental arithmetic by a method it then denied using when asked. The record notes that the model's explanation of itself and the microscope disagreed, and the microscope won.
Add to the record
What does it mean? Write it the way you would say it out loud.