Does this benchmark even need a model?
Diagnostics for benchmark floors and whether an adaptive system has room to improve.
Finding
Under the stated controls, a fixed action behind a guard reached the benchmark ceiling. That left no room for a model to show it could adapt.
In practice
If a fixed baseline already reaches the ceiling, the benchmark cannot credit a model.
Why this question matters
Before crediting adaptation for a strong result, it helps to ask how far a fixed policy can get. If a simpler control reaches the ceiling, the benchmark has no room left to show an advantage.
What I built
An offline synthetic benchmark and controls for measuring the performance floor before attributing gains to adaptation.
- configurations tested
- 328
Finite synthetic control family, 1,512 draw classes. Full result in evidence/theorem.json
A design choice
I removed the model before comparing models: fixed policies in its place, scaffold left in, then rescore. The usual order compares models first and credits the best one. Running the substitution first showed that the guards already encode the answer rule. I then checked all 328 settings in the control family rather than a sample, so the zero-headroom result covers the whole family.
What the work shows
- Control coverage
- 328Configurations in the finite synthetic family.
- Draw coverage
- 1,512Draw classes covered by the result artifact.
- Ceiling shortfalls
- 0Across the integer floors in this family, with one universally dominating configuration.
Inspect the evidence
- Inspect the result artifact
The machine-readable finite-family result and its coverage.
- Read the exhaustive tests
Configuration and draw coverage, unique universal dominance, ceiling and guard assertions.
Where the evidence stops
A bounded synthetic diagnostic, not a general claim about adaptive systems or production safety.