Skip to content
All studies

A simulator says it worked. Would the real service agree?

A harness that compares responses and resulting state against a pinned local reference service, and keeps the disagreements visible.

Finding

3 of 16 new cases diverged from the pinned Gitea reference: two in the response, and one where both sides reported success but the resulting state differed. Checking the response alone would have missed it.

In practice

A mock can agree on the reply and still leave a different state behind.

Why this question matters

An agent can learn to succeed in a simulator whose behaviour differs from the service it represents. A matching success response can hide that difference, so this project looks underneath it.

What I built

A bounded comparison harness for six operations against a pinned local Gitea reference, with recorded evidence and offline replay.

Recorded pilot run.

Now check what each side stored.

  • responses differed: E04, E09
  • state differed: E06

E06 response: two create_label calls for p2-eval-dupe, colours 111111 then 222222. The simulator and Gitea both returned success with the same name and colour each time.

16 new evaluation cases. 13 agreed; 3 diverged. Read the recorded result

E06  responses agreed: both returned success

Recorded outcome of each new evaluation case
CaseCalls sent to bothResponseResulting state
E01create_issue, create_issue, get_issueagreedagreed
E02edit_issue, edit_issue, get_issueagreedagreed
E03add_label, add_label, get_issueagreedagreed
E04remove_label, remove_label, get_issuediffered: Gitea succeeded, the simulator returned 404agreed
E05create_labelagreedagreed
E06create_label, create_labelagreeddiffered: Gitea kept 2 label rows, the simulator kept 1
E07create_issue, get_issueagreedagreed
E08edit_issue, edit_issue, get_issueagreedagreed
E09edit_issue, add_label, get_issuediffered: Gitea succeeded, the simulator returned 404agreed
E10remove_labelagreedagreed
E11create_label, add_label, add_label, get_issueagreedagreed
E12edit_issue, remove_label, edit_issue, get_issueagreedagreed
E13create_issue, add_label, get_issueagreedagreed
E14create_label, create_label, add_label, get_issueagreedagreed
E15get_issue, get_issueagreedagreed
E16create_label, create_labelagreedagreed

A design choice

The observation contract decides what counts as the same. Rows can be sorted, but duplicates are kept. Raw database IDs are not compared literally, and exact timestamps are dropped while their presence stays testable. A looser contract would agree more often and detect less, so I chose the stricter one and wrote it down. Changing it changes what the harness can see.

What the work shows

Annotated summary of the public evidence, not a live run or raw transcript.
New evaluation cases
3 of 16 divergedE04, E06 and E09. The other 13 agreed.
E06 responses
AgreedBoth paths reported success for the duplicate-label operation.
Simulator state
Row overwrittenThe baseline simulator replaced the duplicate label row.
Reference state
Distinct rowsPinned Gitea retained both. Response agreement missed this.

Inspect the evidence

Where the evidence stops

The evidence covers the specified local operations. It does not establish platform-wide simulator equivalence or independent replication. The reference and evaluator code were AI-authored and have not had independent human review. How I build

Read the repository