Glossary
Agent Harness
The software layer wrapped around a model that turns raw text generation into a working agent: the tool-execution loop, permission and approval gating, context and compaction management, and the guardrails that decide what the model's output is allowed to touch. The model generates text; the harness decides what that text can touch.
Frontier model capability has become a weaker predictor of agent performance than harness design. The same model, run through different harnesses, has been reported scoring anywhere from the mid-40s to the low-80s percent on identical SWE-bench-style tasks — the spread comes entirely from scaffolding, not the model.
Where this comes up
Context Engineering for Production AI Agents
Need more than a definition?
If this term describes a problem you have, talk to an engineer about it.