Break my benchmark.
You are the model. Eighteen requests to apply a late revenue adjustment, one decision each. Then the benchmark scores you twice: as you answered, and after its guards have had their say.
These are the study's own cases, as a model saw them. Your answers are scored on the server and only the two scores are kept.
Or send in a policy instead:
What the game shows
- The guards encode the answer rule. They catch every case where acting would be wrong, so 16 of the 18 answers never reach the score.
- So a policy that always acts, and never reads a case, scores 10000. The five local models tested scored 9888 to 10000. The worst possible model still scores 8889: only 1111 basis points respond to the model at all.
- Unguarded, every one of the five models scored below refusing everything. The guards are doing the work, which is why the benchmark cannot credit a model for adapting.
- A bounded synthetic diagnostic, not a general claim about adaptive systems or production safety. The model scores are the study's exploratory run, averaged over repeated draws; a person answers once.
Read the studyThe code and evidenceHear it: a build with no drop