When is expensive interpretability worth its cost?
Three small, calibrated studies of circuit and feature attribution, each graded against a known answer or the strongest cheap baseline rather than a convincing diagram.
Finding
Attribution beat the strongest gradient-free baseline only when the circuit was spread across token positions: on IOI in Gemma-2-2B by 15 to 45 points in 12 of 12 cells, with no meaningful advantage on single-token tasks. The same split held on GPT-2-small and Gemma-2-9B. Across 75 layer cells, gradient-selected circuits matched per-feature ablation to within a mean of 0.028.
In practice
Test attribution against the strongest cheap baseline first. It only earned its cost when the circuit was spread across token positions.
Why this question matters
Interpretability work often ends with a circuit you are asked to trust. Exact per-feature ablation costs one forward pass per feature, so it matters whether it finds anything a cheaper score would miss, and whether the known answer it is graded against actually reproduces the behaviour.
What I built
nanocircuits grades circuit discovery on InterpBench transformers that contain a known circuit, against a leave-one-case-out structural baseline. nanofeatures and nanoassembly carry the same discipline to SAE features on GPT-2-small and Gemma-2-2B, with paired-bootstrap confidence intervals. Everything runs on one Apple M4 Max.
| Task | SAE-feature basis | Raw-neuron control | Prompt pairs |
|---|---|---|---|
| Capitals | +6.0 (95% interval +2.3 to +10.1) | +3.0 (95% interval +1.7 to +4.3) | 20 |
| Country to language | +0.3 (95% interval −2.8 to +3.8) | +2.5 (95% interval +1.6 to +3.4) | 20 |
| Past tense | +2.1 (95% interval +0.3 to +4.2) | +2.5 (95% interval −0.2 to +5.3) | 30 |
| Comparative | −1.3 (95% interval −3.9 to +0.6) | 0.0 (95% interval −1.8 to +1.5) | 24 |
| Plural | −2.1 (95% interval −5.8 to +1.9) | +1.0 (95% interval −1.5 to +3.4) | 16 |
| Antonyms | −5.8 (95% interval −14.6 to +0.7) | +1.2 (95% interval −1.6 to +3.9) | 24 |
| Successor | −0.5 (95% interval −5.8 to +6.6) | −5.0 (95% interval −11.9 to +2.6) | 11 |
| IOI (indirect object identification, distributed) | +30.1 (95% interval +16.0 to +45.8) | +17.2 (95% interval +8.7 to +28.1) | 18 |
A design choice
I grade every method against the strongest cheap baseline I can build, not a convenient weak one, and I measure the known answer instead of assuming it. The baseline decides the apparent result. Ranked by summed activation change, attribution beat the cheap score by 17 to 35 points on factual recall. Ranked by peak per-position change, almost all of that gap closed. The cost is smaller claims: when nanocircuits added a baseline built from a single forward pass, one of its two clean wins tied it exactly and dropped out.
What the work shows
- IOI, a distributed circuit
- +15 to +45 ptsAttribution over the strongest gradient-free baseline in 12 of 12 cells. Every confidence interval excludes zero.
- Seven single-token tasks
- No meaningful advantageA pre-registered equivalence test (TOST, 5-point margin) over 63 cells: 10 small attribution wins of 6 points or less, 13 equivalent, 4 cheap-baseline wins, 36 inconclusive.
- Gradient vs exact ablation
- Mean gap 0.028Faithfulness of gradient-selected and exact-selected circuits across 75 layer cells on GPT-2-small and Gemma-2-2B.
- Known circuits (InterpBench)
- 1 of 7 casesTwo cases pass the baseline and faithful-circuit filters. Against a baseline built from a single forward pass, only case 11 still wins at node level.
Inspect the evidence
- Read the argument across all three
when-is-interpretability-worth-it.md: the three studies as one argument.
- Read the method and exact numbers
nanoassembly THESIS.md: the 75-cell calibration, controls and limits.
- nanocircuits on Zenodo
DOI 10.5281/zenodo.20611793. Circuit discovery graded against known circuits.
- nanofeatures on Zenodo
DOI 10.5281/zenodo.20611795. When attribution beats a free baseline on SAE features.
- nanoassembly on Zenodo
DOI 10.5281/zenodo.20611788. Multi-layer feature circuits and attribution cost.
Where the evidence stops
Small models and a finite task set. In nanocircuits two InterpBench cases plus IOI beat the structural baseline with a faithful ground-truth circuit; against a one-forward-pass behavioural baseline only case 11 and IOI still win. None of this is a claim about frontier-scale models.