Blog.
Writing on AI agent engineering, automation, and what it actually takes to run production AI systems.
28 posts

Why Most AI Agent Pilots Never Reach Production
79% of enterprises claim AI agent adoption. Only 11-17% actually run them in production. Here's why the gap exists, and the infrastructure patterns that close it.
Read the article
What Agents Actually Cost in Production
A single LLM call costs about four cents. An orchestrated agent task costs roughly $1.20. Here's what drives the 30x gap and how to keep it from running away.

AI Agent Failure Modes in Production: The Four That Actually Kill Systems
The four failure modes that repeat across production deployments: tool-calling errors with cascading impact, infinite retry loops burning budgets, hallucination cascades in multi-agent systems, and silent quality drift.

How to Roll Out an AI Agent Safely
Production agents never deploy all at once. The staged rollout gate — shadow, canary, expansion, stable — that Anthropic, Google Cloud, and production teams actually validate against.

AI Agent Observability: What to Track in Production
The three critical observables for production agents — tool-call trajectory, memory operations, workflow visibility — plus implementation patterns and a Langfuse vs Braintrust vs Honeycomb comparison.

AI Automation for Marketing Agencies: The 5 Systems That Actually Save Time
The 5 AI automation systems that free up agency hours without touching creative quality — brief intake, client reporting, brand voice, outreach, and status updates. Built on the tools you already use.
Have a production problem now?
You don't have to wait for a blog post. Talk to an engineer about the system you're building today.