# The studies, as music

Each of the ten studies is a track built from its own results, synthesised in the browser at https://jasonlovell.ai/listen. Two decks and a crossfader mix any two. A finding gets a melody; a study with no result yet gets drums and bass only. Every melodic sound stands for a number or an event in the study. Each research theme has its own groove: state and continuity is UK garage, evaluation validity is house, learning and control is afro house, and judgment and policy is deep house.

A mix can be shared as a link: https://jasonlovell.ai/listen?a=<study>&b=<study>&x=<crossfader, 0 to 1>. On that page, a browser agent can use the WebMCP tools play_studies and stop_music.

## Tracks

### MCP Continuation Replay

The action succeeded. The reply vanished. What happens next? UK garage, 132 BPM. Study: https://jasonlovell.ai/work/mcp-continuation-replay

- The stab on the one: The action. It succeeds.
- The gap where the answer should be: The reply that vanished.
- The same stab again: The retry reuses the original operation ID, so it is the same note, not a new one.
- One low note per phrase: One physical issue in the database after recovery. Never two.

### Scaffold Arena

How much of an agent’s performance comes from everything around the model? House, 124 BPM. Study: https://jasonlovell.ai/work/scaffold-arena

- Drums and bass, no melody: There is no measured result yet, so there is nothing to play. You are hearing the instrument, which is what this release built.

### RealityBridge

A simulator says it worked. Would the real service agree? UK garage, 132 BPM. Study: https://jasonlovell.ai/work/realitybridge

- Sixteen plucks a bar: The 16 new evaluation cases, one per step: the simulator on the left, pinned Gitea on the right.
- The clash on steps 4 and 9: E04 and E09, where the two responses diverged.
- Step 6: the notes agree, the bass slides away: E06. Both sides reported success, but the state underneath differed.

### ShadowSkillBench

When an agent learns a procedure, what else does it carry forward? Afro house, 122 BPM. Study: https://jasonlovell.ai/work/shadowskillbench

- Drums and bass, no melody: There is no measured result yet, so there is nothing to play. You are hearing the instrument, which is what this release built.

### Living Context

What should an agent revisit when the world changes? UK garage, 132 BPM. Study: https://jasonlovell.ai/work/living-context

- The chords: The context. When a chord changes, only the notes that changed are played again; the rest hold. That is revision tracking, with no model involved.
- The phrase that stops halfway: The learned-context extension, which came out negative or inconclusive.

### Jev Decision Lab

What changes when a model returns a typed judgment? Deep house, 118 BPM. Study: https://jasonlovell.ai/work/jev-decision-lab

- Drums and bass, no melody: There is no measured result yet, so there is nothing to play. You are hearing the instrument, which is what this release built.

### GhostTrace

Does a behavioural signal survive repeated self-distillation? Afro house, 122 BPM. Study: https://jasonlovell.ai/work/ghosttrace

- The marimba echo: Each repeat at half the level of the last: the behavioural signal decays with a measurable half-life in the toy tier.
- The dry second half: The local-LLM tier, which came out negative. No echo carries over.

### No Model Floor

Does this benchmark even need a model? House, 124 BPM. Study: https://jasonlovell.ai/work/no-model-floor

- The build: A fixed action behind a guard, rising to the benchmark ceiling.
- The drop: There isn't one. Adaptation headroom was 0 across all 328 configurations.

### CascadeShift

Are we measuring the model, or the configuration? House, 124 BPM. Study: https://jasonlovell.ai/work/cascadeshift

- The lead: It swings above and below the root, as the confidence interval spans zero. It never settles higher: no benefit is claimed.

### Nano interpretability

When is expensive interpretability worth its cost? House, 124 BPM. Study: https://jasonlovell.ai/work/nano-interpretability

- Eight notes a bar: One per task: attribution minus the strongest cheap baseline (Gemma-2-2B, layer 7). Higher means attribution won by more.
- Seven notes near the root, one leap: The seven single-token tasks show no meaningful advantage. IOI, the distributed circuit, leaps.
