Agentic tool-use eval on a local 35B (Q8): trap-tool avoidance is solid, but I can't tell if my failures are the model or my harness
I've been running a small agentic eval harness against a local model and I'd like a sanity check on...
I've been running a small agentic eval harness against a local model and I'd like a sanity check on...
Last month I tracked Claude Code and Codex pass rates for 95 days. The question I got most in...
Why offline agent evaluation frameworks fail in production, and what infrastructure separates agents that improve from those that plateau.
Welcome to Day 1! Python is famous for its clean, readable syntax—often described as "executable...
There is a growing sentiment in engineering circles right now that documentation is a relic of the...
The AI Skill Registry at 5,776: A Deep Dive into Reusable Modules for Code Review, Terraform, and...
I believe AI will be another service like the internet or a cell phone, and it's important to use it...
Real-Time AI Observability: Dashboards That Show Actual Database Rows Discover how TormentNexus...
MCP Protocol Deep-Dive: How Tool Discovery Actually Works Under the Hood Uncover the mechanics of...
Beyond Synchronous Hell: Why Your Multi-Agent System Needs an Event-Driven Backbone Explore how...
For the past year I've run most of my day on a system I built for exactly one user: me. Last week I...
Mexico’s Public Registry of Commerce still runs through state-level implementation. Look up the same...
Back in January, I wrote about the vibe coding hangover. The morning-after feeling of shipping fast...
You know that moment when you're reviewing a PR and you need to: Check the GitHub diff (tab...
I recently needed a short product demo for my open-source project called Agent OS. The goal was...
Over the past two years, artificial intelligence has arguably become the most talked-about topic in...
The context This talk was born from a real frustration. Marcos Ramalho and I — both...
The Closing Every series earns its final paragraph. This one earned a full article. The...
I spent last weekend comparing two ways to serve a local model: Llamafile and the more traditional...
Batch a day of agent activity into one replyable digest email and schedule the send — using a Nylas Agent Account, both the API and the CLI.
Hook When building local-first architectures, the instinct is often to tightly couple...
A paired workload protocol for validating latency, cost, correctness, and recovery before changing production models.
A practical outbound-data threat model and test procedure for coding agents with filesystem and network access.