The studies, as music.
Each study is a track, built from its own results. Put any two on the decks and mix them.
The sound is made in your browser as it plays; nothing is streamed. Headphones help: RealityBridge puts the simulator on the left and the real service on the right.
Deck A
A simulator says it worked. Would the real service agree?
- Sixteen plucks a bar
- The 16 new evaluation cases, one per step: the simulator on the left, pinned Gitea on the right.
- The clash on steps 4 and 9
- E04 and E09, where the two responses diverged.
- Step 6: the notes agree, the bass slides away
- E06. Both sides reported success, but the state underneath differed.
Deck B
When is expensive interpretability worth its cost?
- Eight notes a bar
- One per task: attribution minus the strongest cheap baseline (Gemma-2-2B, layer 7). Higher means attribution won by more.
- Seven notes near the root, one leap
- The seven single-token tasks show no meaningful advantage. IOI, the distributed circuit, leaps.
How a study becomes a track
- A finding gets a melody. A study with no result yet gets drums and bass only.
- Every melodic sound stands for a number or an event in the study. The key under each deck says which.
- Each research theme has its own groove: state and continuity is UK garage, evaluation validity is house, learning and control is afro house, and judgment and policy is deep house.
- The tempo follows whichever deck is louder, and a new study comes in on the next bar.
- AI agents reading the site play the percussion as they arrive. Visits are counted by agent and page only; nothing about people is recorded.