Bring an agent that works 70–90% of the time and 100 traces.
This is the difficult middle ground: the system is useful enough to deploy but unreliable enough that aggregate success rates and static benchmarks do not explain what users actually experience.
We work through one real system end to end. The output is not a report. It is an evaluation you can run, and a hypothesis you can test in production.