Skip to content
All studies

Does this benchmark even need a model?

Diagnostics for benchmark floors and whether an adaptive system has room to improve.

Finding

Under the stated controls, a fixed action behind a guard reached the benchmark ceiling. That left no room for a model to show it could adapt.

In practice

If a fixed baseline already reaches the ceiling, the benchmark cannot credit a model.

Why this question matters

Before crediting adaptation for a strong result, it helps to ask how far a fixed policy can get. If a simpler control reaches the ceiling, the benchmark has no room left to show an advantage.

What I built

An offline synthetic benchmark and controls for measuring the performance floor before attributing gains to adaptation.

configurations tested
328

Finite synthetic control family, 1,512 draw classes. Full result in evidence/theorem.json

Play the model on its 18 cases: break my benchmark

A design choice

I removed the model before comparing models: fixed policies in its place, scaffold left in, then rescore. The usual order compares models first and credits the best one. Running the substitution first showed that the guards already encode the answer rule. I then checked all 328 settings in the control family rather than a sample, so the zero-headroom result covers the whole family.

What the work shows

Annotated summary of the public evidence, not a live run or raw transcript.
Control coverage
328Configurations in the finite synthetic family.
Draw coverage
1,512Draw classes covered by the result artifact.
Ceiling shortfalls
0Across the integer floors in this family, with one universally dominating configuration.

Inspect the evidence

Where the evidence stops

A bounded synthetic diagnostic, not a general claim about adaptive systems or production safety.

Read the repository