Tool-use agents fail silently when a prompt change rewires which tool gets called. 90 lines of Python, 3 judges in a ladder, runnable on a small golden set for a few dollars.
An Eval Harness for Tool-Use Agents: 90 Lines, 3 Judges, $3 Per Run
Tool-use agents fail silently when a prompt change rewires which tool gets called. 90 lines of Python, 3 judges in a ladder, runnable on a small golden set for a few dollars.
I spent the weekend with Bonsai-27B, and I think I finally understand what a "27B that runs on one...
Vision & What the Agent Does Studying for an AWS certification usually means remembering to sit...
A few years ago I started doing something embarrassing: making a list of software engineering best...