Skip to content
All studies

Are we measuring the model, or the configuration?

A synthetic study of configuration-aware tool use and the validity of the measurements around it.

Finding

The main comparison’s confidence interval includes zero, so the study claims no benefit and says so.

In practice

Hold configuration fixed, or measure it, before attributing a change in tool use to the model.

Why this question matters

A change in configuration can alter the tool use we observe. Before calling that a model improvement, the comparison needs to separate the two.

What I built

A research package examining configuration-aware tool use with offline measurement checks.

A design choice

An episode passes only if the goal is met and every hard constraint still holds at the end. Scoring the goal alone would credit a run that finished the task by breaking a rule. I also audited the treatment itself. For a substantial part of the shifted set, the discovery tools returned the same rules under both conditions, so the comparison tested less than its design suggested. The write-up reports that weakness next to the result.

Where the evidence stops

Synthetic case-study evidence, not a proven general performance improvement.

Read the repository