Are we measuring the model, or the configuration?
A synthetic study of configuration-aware tool use and the validity of the measurements around it.
Finding
The main comparison’s confidence interval includes zero, so the study claims no benefit and says so.
In practice
Hold configuration fixed, or measure it, before attributing a change in tool use to the model.
Why this question matters
A change in configuration can alter the tool use we observe. Before calling that a model improvement, the comparison needs to separate the two.
What I built
A research package examining configuration-aware tool use with offline measurement checks.
A design choice
An episode passes only if the goal is met and every hard constraint still holds at the end. Scoring the goal alone would credit a run that finished the task by breaking a rule. I also audited the treatment itself. For a substantial part of the shifted set, the discovery tools returned the same rules under both conditions, so the comparison tested less than its design suggested. The write-up reports that weakness next to the result.
Where the evidence stops
Synthetic case-study evidence, not a proven general performance improvement.