A simulator says it worked. Would the real service agree?
A harness that compares responses and resulting state against a pinned local reference service, and keeps the disagreements visible.
Finding
3 of 16 new cases diverged from the pinned Gitea reference: two in the response, and one where both sides reported success but the resulting state differed. Checking the response alone would have missed it.
In practice
A mock can agree on the reply and still leave a different state behind.
Why this question matters
An agent can learn to succeed in a simulator whose behaviour differs from the service it represents. A matching success response can hide that difference, so this project looks underneath it.
What I built
A bounded comparison harness for six operations against a pinned local Gitea reference, with recorded evidence and offline replay.
Recorded pilot run.
Now check what each side stored.
- responses differed: E04, E09
- state differed: E06
E06 response: two create_label calls for p2-eval-dupe, colours 111111 then 222222. The simulator and Gitea both returned success with the same name and colour each time.
16 new evaluation cases. 13 agreed; 3 diverged. Read the recorded result
E06 responses agreed: both returned success
| Case | Calls sent to both | Response | Resulting state |
|---|---|---|---|
| E01 | create_issue, create_issue, get_issue | agreed | agreed |
| E02 | edit_issue, edit_issue, get_issue | agreed | agreed |
| E03 | add_label, add_label, get_issue | agreed | agreed |
| E04 | remove_label, remove_label, get_issue | differed: Gitea succeeded, the simulator returned 404 | agreed |
| E05 | create_label | agreed | agreed |
| E06 | create_label, create_label | agreed | differed: Gitea kept 2 label rows, the simulator kept 1 |
| E07 | create_issue, get_issue | agreed | agreed |
| E08 | edit_issue, edit_issue, get_issue | agreed | agreed |
| E09 | edit_issue, add_label, get_issue | differed: Gitea succeeded, the simulator returned 404 | agreed |
| E10 | remove_label | agreed | agreed |
| E11 | create_label, add_label, add_label, get_issue | agreed | agreed |
| E12 | edit_issue, remove_label, edit_issue, get_issue | agreed | agreed |
| E13 | create_issue, add_label, get_issue | agreed | agreed |
| E14 | create_label, create_label, add_label, get_issue | agreed | agreed |
| E15 | get_issue, get_issue | agreed | agreed |
| E16 | create_label, create_label | agreed | agreed |
A design choice
The observation contract decides what counts as the same. Rows can be sorted, but duplicates are kept. Raw database IDs are not compared literally, and exact timestamps are dropped while their presence stays testable. A looser contract would agree more often and detect less, so I chose the stricter one and wrote it down. Changing it changes what the harness can see.
What the work shows
- New evaluation cases
- 3 of 16 divergedE04, E06 and E09. The other 13 agreed.
- E06 responses
- AgreedBoth paths reported success for the duplicate-label operation.
- Simulator state
- Row overwrittenThe baseline simulator replaced the duplicate label row.
- Reference state
- Distinct rowsPinned Gitea retained both. Response agreement missed this.
Inspect the evidence
- Inspect the recorded results
The frozen local E06, E04 and E09 comparisons, including the state divergence.
- Read the comparison architecture
Execution, reference readback, response comparison and the independent state oracle.
Where the evidence stops
The evidence covers the specified local operations. It does not establish platform-wide simulator equivalence or independent replication. The reference and evaluator code were AI-authored and have not had independent human review. How I build