# Are we measuring the model, or the configuration?

CascadeShift: Synthetic case study. Published 21 Sep 2026. Question: Can we trust the score?
Scope: Synthetic case study · offline measurement checks
Page: https://jasonlovell.ai/work/cascadeshift
Code: https://github.com/jlov7/CascadeShift-Research

A synthetic study of configuration-aware tool use and the validity of the measurements around it.

## Finding

The main comparison’s confidence interval includes zero, so the study claims no benefit and says so.

## In practice

Hold configuration fixed, or measure it, before attributing a change in tool use to the model.

## What I built

A research package examining configuration-aware tool use with offline measurement checks.

## Where the evidence stops

Synthetic case-study evidence, not a proven general performance improvement.

---
By Jason Lovell. Independent work, separate from my role at PwC. I build everything here myself, end to end, with Claude Code and Codex. I write the question and the test first, read the runs, and keep the results that go against me.
