# A simulator says it worked. Would the real service agree?

RealityBridge: Comparison harness. Published 24 Sep 2026. Question: Does the state hold up?
Scope: 16 evaluation cases · 6 operations · pinned local Gitea
Page: https://jasonlovell.ai/work/realitybridge
Code: https://github.com/jlov7/realitybridge

A harness that compares responses and resulting state against a pinned local reference service, and keeps the disagreements visible.

## Finding

3 of 16 new cases diverged from the pinned Gitea reference: two in the response, and one where both sides reported success but the resulting state differed. Checking the response alone would have missed it.

## In practice

A mock can agree on the reply and still leave a different state behind.

## Why it matters

An agent can learn to succeed in a simulator whose behaviour differs from the service it represents. A matching success response can hide that difference, so this project looks underneath it.

## The design decision

The observation contract decides what counts as the same. Rows can be sorted, but duplicates are kept. Raw database IDs are not compared literally, and exact timestamps are dropped while their presence stays testable. A looser contract would agree more often and detect less, so I chose the stricter one and wrote it down. Changing it changes what the harness can see.

## What I built

A bounded comparison harness for six operations against a pinned local Gitea reference, with recorded evidence and offline replay.

## What was recorded

| Measure | Value | Note |
| --- | --- | --- |
| New evaluation cases | 3 of 16 diverged (differs) | E04, E06 and E09. The other 13 agreed. |
| E06 responses | Agreed | Both paths reported success for the duplicate-label operation. |
| Simulator state | Row overwritten (differs) | The baseline simulator replaced the duplicate label row. |
| Reference state | Distinct rows | Pinned Gitea retained both. Response agreement missed this. |

## Evidence

- [Inspect the recorded results](https://github.com/jlov7/realitybridge/blob/main/docs/research-completion-2026-09-22/RESULTS.md): The frozen local E06, E04 and E09 comparisons, including the state divergence.
- [Read the comparison architecture](https://github.com/jlov7/realitybridge/blob/main/docs/ARCHITECTURE.md): Execution, reference readback, response comparison and the independent state oracle.

## Where the evidence stops

The evidence covers the specified local operations. It does not establish platform-wide simulator equivalence or independent replication. The reference and evaluator code were AI-authored and have not had independent human review.

---
By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me.
