# When is expensive interpretability worth its cost?

Nano interpretability: Research software, three releases. Published 9 Jun 2026. Question: Can we trust the score?
Scope: InterpBench, GPT-2-small, Gemma-2-2B and 9B · one Apple M4 Max
Page: https://jasonlovell.ai/work/nano-interpretability
Code: https://github.com/jlov7/nanoassembly

Three small, calibrated studies of circuit and feature attribution, each graded against a known answer or the strongest cheap baseline rather than a convincing diagram.

## Finding

Attribution beat the strongest gradient-free baseline only when the circuit was spread across token positions: on IOI in Gemma-2-2B by 15 to 45 points in 12 of 12 cells, with no meaningful advantage on single-token tasks. The same split held on GPT-2-small and Gemma-2-9B. Across 75 layer cells, gradient-selected circuits matched per-feature ablation to within a mean of 0.028.

## In practice

Test attribution against the strongest cheap baseline first. It only earned its cost when the circuit was spread across token positions.

## Why it matters

Interpretability work often ends with a circuit you are asked to trust. Exact per-feature ablation costs one forward pass per feature, so it matters whether it finds anything a cheaper score would miss, and whether the known answer it is graded against actually reproduces the behaviour.

## The design decision

I grade every method against the strongest cheap baseline I can build, not a convenient weak one, and I measure the known answer instead of assuming it. The baseline decides the apparent result. Ranked by summed activation change, attribution beat the cheap score by 17 to 35 points on factual recall. Ranked by peak per-position change, almost all of that gap closed. The cost is smaller claims: when nanocircuits added a baseline built from a single forward pass, one of its two clean wins tied it exactly and dropped out.

## What I built

nanocircuits grades circuit discovery on InterpBench transformers that contain a known circuit, against a leave-one-case-out structural baseline. nanofeatures and nanoassembly carry the same discipline to SAE features on GPT-2-small and Gemma-2-2B, with paired-bootstrap confidence intervals. Everything runs on one Apple M4 Max.

## What was recorded

| Measure | Value | Note |
| --- | --- | --- |
| IOI, a distributed circuit | +15 to +45 pts | Attribution over the strongest gradient-free baseline in 12 of 12 cells. Every confidence interval excludes zero. |
| Seven single-token tasks | No meaningful advantage | A pre-registered equivalence test (TOST, 5-point margin) over 63 cells: 10 small attribution wins of 6 points or less, 13 equivalent, 4 cheap-baseline wins, 36 inconclusive. |
| Gradient vs exact ablation | Mean gap 0.028 | Faithfulness of gradient-selected and exact-selected circuits across 75 layer cells on GPT-2-small and Gemma-2-2B. |
| Known circuits (InterpBench) | 1 of 7 cases | Two cases pass the baseline and faithful-circuit filters. Against a baseline built from a single forward pass, only case 11 still wins at node level. |

## Evidence

- [Read the argument across all three](https://github.com/jlov7/nanoassembly/blob/main/when-is-interpretability-worth-it.md): when-is-interpretability-worth-it.md: the three studies as one argument.
- [Read the method and exact numbers](https://github.com/jlov7/nanoassembly/blob/main/THESIS.md): nanoassembly THESIS.md: the 75-cell calibration, controls and limits.
- [nanocircuits on Zenodo](https://doi.org/10.5281/zenodo.20611793): DOI 10.5281/zenodo.20611793. Circuit discovery graded against known circuits.
- [nanofeatures on Zenodo](https://doi.org/10.5281/zenodo.20611795): DOI 10.5281/zenodo.20611795. When attribution beats a free baseline on SAE features.
- [nanoassembly on Zenodo](https://doi.org/10.5281/zenodo.20611788): DOI 10.5281/zenodo.20611788. Multi-layer feature circuits and attribution cost.

## Where the evidence stops

Small models and a finite task set. In nanocircuits two InterpBench cases plus IOI beat the structural baseline with a faithful ground-truth circuit; against a one-forward-pass behavioural baseline only case 11 and IOI still win. None of this is a claim about frontier-scale models.

---
By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me.
