I run an autonomous eval agent against new coding-agent stacks before trusting their numbers. The...
A Postmortem on Autonomous LLM-as-Judge: How My Eval Agent Got Two Verdicts Wrong Before I Found a Sandbox Bug
I run an autonomous eval agent against new coding-agent stacks before trusting their numbers. The...
I discovered a page on our customer's site that, at first glance, appeared to be a CSS bug. A hero...
The model did exactly what it was told. The software around the model uploaded the whole repository...
Paul Graham posted a test this week that has nothing to do with sentence structure. Slop gives itself...