Skip to content
All studies

How much of an agent’s performance comes from everything around the model?

A workbench and experimental protocol for studying orchestration, tools and control flow as part of the system being evaluated.

Status

The workbench makes scaffold choices explicit before any comparison. So far it runs on local synthetic fixtures and mock runs; there is no measured result yet.

In practice

Treat the scaffold as part of what you evaluate. Two models compared on different scaffolds is not a model comparison.

Why this question matters

Change the tools or control flow around a model and you change the system being tested. A comparison needs to account for those choices before the scores mean much.

What I built

A research workbench that makes scaffold choices explicit, with a protocol for comparing quality, reliability, cost and latency.

StudyPack

Preflight

Preflight PASS

Design frozen. The run can start.

System sketch of the preflight, not a result.

A design choice

I freeze the comparison before anything runs. A StudyPack fixes the scenarios, controls, budgets and evaluation, and a preflight check returns HOLD when a control is missing. The alternative is to run first and explain the scores afterwards, which is when it gets hard to tell which change helped. The cost is friction: nothing executes until the design is declared and passes preflight.

Where the evidence stops

A beta research workbench, not a completed comparative live study or a validated ranking of agent scaffolds.

Read the repository