Skip to content
All studies

When is expensive interpretability worth its cost?

Three small, calibrated studies of circuit and feature attribution, each graded against a known answer or the strongest cheap baseline rather than a convincing diagram.

Finding

Attribution beat the strongest gradient-free baseline only when the circuit was spread across token positions: on IOI in Gemma-2-2B by 15 to 45 points in 12 of 12 cells, with no meaningful advantage on single-token tasks. The same split held on GPT-2-small and Gemma-2-9B. Across 75 layer cells, gradient-selected circuits matched per-feature ablation to within a mean of 0.028.

In practice

Test attribution against the strongest cheap baseline first. It only earned its cost when the circuit was spread across token positions.

Why this question matters

Interpretability work often ends with a circuit you are asked to trust. Exact per-feature ablation costs one forward pass per feature, so it matters whether it finds anything a cheaper score would miss, and whether the known answer it is graded against actually reproduces the behaviour.

What I built

nanocircuits grades circuit discovery on InterpBench transformers that contain a known circuit, against a leave-one-case-out structural baseline. nanofeatures and nanoassembly carry the same discipline to SAE features on GPT-2-small and Gemma-2-2B, with paired-bootstrap confidence intervals. Everything runs on one Apple M4 Max.

Attribution minus the strongest cheap baseline, by taskDot-and-whisker chart, SAE-feature basis and raw-neuron control. On all seven single-token tasks, every estimate sits within 6.0 points of zero. On IOI, where the signal is distributed, the gap is +30.1 points (95% interval +16.0 to +45.8) in the SAE-feature basis and +17.2 points (95% interval +8.7 to +28.1) in the raw-neuron control.
Attribution minus the strongest cheap baseline, by taskDot-and-whisker chart, SAE-feature basis and raw-neuron control. On all seven single-token tasks, every estimate sits within 6.0 points of zero. On IOI, where the signal is distributed, the gap is +30.1 points (95% interval +16.0 to +45.8) in the SAE-feature basis and +17.2 points (95% interval +8.7 to +28.1) in the raw-neuron control.

Attribution minus the strongest cheap baseline (sufficiency, percentage points). Gemma-2-2B, layer 7, top 64 units, paired-bootstrap 95% intervals.
TaskSAE-feature basisRaw-neuron controlPrompt pairs
Capitals+6.0 (95% interval +2.3 to +10.1)+3.0 (95% interval +1.7 to +4.3)20
Country to language+0.3 (95% interval −2.8 to +3.8)+2.5 (95% interval +1.6 to +3.4)20
Past tense+2.1 (95% interval +0.3 to +4.2)+2.5 (95% interval −0.2 to +5.3)30
Comparative−1.3 (95% interval −3.9 to +0.6)0.0 (95% interval −1.8 to +1.5)24
Plural−2.1 (95% interval −5.8 to +1.9)+1.0 (95% interval −1.5 to +3.4)16
Antonyms−5.8 (95% interval −14.6 to +0.7)+1.2 (95% interval −1.6 to +3.9)24
Successor−0.5 (95% interval −5.8 to +6.6)−5.0 (95% interval −11.9 to +2.6)11
IOI (indirect object identification, distributed)+30.1 (95% interval +16.0 to +45.8)+17.2 (95% interval +8.7 to +28.1)18
Gemma-2-2B, layer 7: attribution minus the strongest cheap baseline, for SAE features and a raw-neuron control, with paired-bootstrap 95% intervals. Attribution pulls clear only on IOI (indirect-object identification), the distributed circuit. Source figure

A design choice

I grade every method against the strongest cheap baseline I can build, not a convenient weak one, and I measure the known answer instead of assuming it. The baseline decides the apparent result. Ranked by summed activation change, attribution beat the cheap score by 17 to 35 points on factual recall. Ranked by peak per-position change, almost all of that gap closed. The cost is smaller claims: when nanocircuits added a baseline built from a single forward pass, one of its two clean wins tied it exactly and dropped out.

What the work shows

Annotated summary of the public evidence, not a live run or raw transcript.
IOI, a distributed circuit
+15 to +45 ptsAttribution over the strongest gradient-free baseline in 12 of 12 cells. Every confidence interval excludes zero.
Seven single-token tasks
No meaningful advantageA pre-registered equivalence test (TOST, 5-point margin) over 63 cells: 10 small attribution wins of 6 points or less, 13 equivalent, 4 cheap-baseline wins, 36 inconclusive.
Gradient vs exact ablation
Mean gap 0.028Faithfulness of gradient-selected and exact-selected circuits across 75 layer cells on GPT-2-small and Gemma-2-2B.
Known circuits (InterpBench)
1 of 7 casesTwo cases pass the baseline and faithful-circuit filters. Against a baseline built from a single forward pass, only case 11 still wins at node level.

Inspect the evidence

Where the evidence stops

Small models and a finite task set. In nanocircuits two InterpBench cases plus IOI beat the structural baseline with a faithful ground-truth circuit; against a one-forward-pass behavioural baseline only case 11 and IOI still win. None of this is a claim about frontier-scale models.

Read the repository