# How much of an agent’s performance comes from everything around the model?

Scaffold Arena: Research workbench. Published 14 Sep 2026. Question: Can we trust the score?
Scope: Beta workbench · synthetic fixtures and mock runs
Page: https://jasonlovell.ai/work/scaffold-arena
Code: https://github.com/jlov7/scaffold-arena

A workbench and experimental protocol for studying orchestration, tools and control flow as part of the system being evaluated.

## Status

The workbench makes scaffold choices explicit before any comparison. So far it runs on local synthetic fixtures and mock runs; there is no measured result yet.

## In practice

Treat the scaffold as part of what you evaluate. Two models compared on different scaffolds is not a model comparison.

## What I built

A research workbench that makes scaffold choices explicit, with a protocol for comparing quality, reliability, cost and latency.

## Where the evidence stops

A beta research workbench, not a completed comparative live study or a validated ranking of agent scaffolds.

---
By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me.
