Two prompts on one eval set are paired data. The two-proportion SE said 'collect more'; McNemar said it was already decided.
Your A/B eval is paired. Your stat test probably isn't.
Two prompts on one eval set are paired data. The two-proportion SE said 'collect more'; McNemar said it was already decided.
You finish a piece of work. You ask AI to review it. It says "looks good." You publish. It wasn't...
Same agent, same faults, two ways to see the cluster. The one with the resource graph and change timeline used 76% fewer tool calls and finished in half the time.
You have built an AI agent harness. It calls tools, routes requests, and returns results. Your team...