# Does this benchmark even need a model?

No Model Floor: Offline diagnostics. Published 21 Sep 2026. Question: Can we trust the score?
Scope: 328 configurations · 1,512 draw classes · offline
Page: https://jasonlovell.ai/work/no-model-floor
Code: https://github.com/jlov7/no-model-floor

Diagnostics for benchmark floors and whether an adaptive system has room to improve.

## Finding

Under the stated controls, a fixed action behind a guard reached the benchmark ceiling. That left no room for a model to show it could adapt.

## In practice

If a fixed baseline already reaches the ceiling, the benchmark cannot credit a model.

## Why it matters

Before crediting adaptation for a strong result, it helps to ask how far a fixed policy can get. If a simpler control reaches the ceiling, the benchmark has no room left to show an advantage.

## The design decision

I removed the model before comparing models: fixed policies in its place, scaffold left in, then rescore. The usual order compares models first and credits the best one. Running the substitution first showed that the guards already encode the answer rule. I then checked all 328 settings in the control family rather than a sample, so the zero-headroom result covers the whole family.

## What I built

An offline synthetic benchmark and controls for measuring the performance floor before attributing gains to adaptation.

## What was recorded

| Measure | Value | Note |
| --- | --- | --- |
| Control coverage | 328 | Configurations in the finite synthetic family. |
| Draw coverage | 1,512 | Draw classes covered by the result artifact. |
| Ceiling shortfalls | 0 | Across the integer floors in this family, with one universally dominating configuration. |

## Evidence

- [Inspect the result artifact](https://github.com/jlov7/no-model-floor/blob/main/evidence/theorem.json): The machine-readable finite-family result and its coverage.
- [Read the exhaustive tests](https://github.com/jlov7/no-model-floor/blob/main/tests/test_theorem.py): Configuration and draw coverage, unique universal dominance, ceiling and guard assertions.

## Where the evidence stops

A bounded synthetic diagnostic, not a general claim about adaptive systems or production safety.

---
By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me.
