A public corpus of real agent failures, with forensic analyses attached.
Each case follows one fixed structure, so that failures can be compared instead of retold. Production failures are discussed today as one-off incidents inside individual teams, which makes them impossible to name, compare, or accumulate knowledge about.
Every entry answers the same fifteen questions.
Case 001 — forthcoming.
A second line of work: how production trajectories should — and should not — become repeatable evaluations. Which parts of a trace are signal, what a grader can actually measure, and when one incident generalises.
When production failures become evals, how do we prevent the failures we’ve already seen from becoming our definition of correctness?
Contribute traces