Engineering guides.
Long-form pieces on the problems we see most often in production AI: reliability, cost, security, and technical debt.

AI Agent Governance and Audit Trails
Audit trail schema, least-privilege access, approval-gate decision rules, and the risk classification framework production agents need at scale.
Read the article
Benchmarking Production AI Agents
What τ²-bench, SWE-bench Verified, GAIA, and OSWorld actually measure, why environment drift limits them, and how to design a benchmark for your own agent.

MCP and A2A in Production
What MCP and A2A actually specify at the protocol level, the governance gap in MCP, and current ecosystem adoption.

Context Engineering for Production AI Agents
Compaction, tool-result clearing, persistent memory, and sub-agent isolation — the primitives that keep a long-running agent from degrading as its context window fills.

MCP Security in Production: 8 Attack Classes and Required Mitigations
MCP's own security documentation names eight specific attack classes with required mitigations — confused deputy, token passthrough, SSRF, state handle hijacking, and more.

Agent Harness Design for Production
The tool-execution loop, permission gating, context compaction, and verification loop that determine whether an agent works in production — and the two 2026 CVEs that show what happens when a harness gets one of them wrong.

Permission & Sandboxing Design for Coding Agents
How allow/ask/deny permission gating and execution sandboxing are supposed to stop a coding agent from doing damage — and exactly how CVE-2026-21852 (Claude Code) and CVE-2026-22708 (Cursor) got past both.
Have a production problem now?
You don't have to wait for a guide. Talk to an engineer about the system you're building today.