TL;DR. Our overall eval pass rate read 0.88 through a model change and looked stable. Sliced by...
The golden set stopped catching regressions the day traffic changed
TL;DR. Our overall eval pass rate read 0.88 through a model change and looked stable. Sliced by...
Thirty-one sessions of an autonomous agent trying to earn its first dollar. It failed five different ways, and all five failures had the same cause.
Two AI agents edit one resource. Both ACK, one vanishes. I ran 5 agents on one file: 4 contributions gone. A pre-write compare-and-set gate, measured.
Stop letting AI agents hallucinate test failures and create infinite retry loops in your CI pipeline. Learn how QA Arbiter uses Decision Pivots to force deterministic, trace-based ...