Agent Harness — Patterns & Anti-Patterns¶
Harness engineering is the practice of improving an agent by changing its harness — prompts, tools, context, limits and feedback loops — rather than the model.
Case Studies¶
LangChain: same model, better harness¶
LangChain kept the model fixed and changed only the system prompt, tools and middleware. Their coding agent on Terminal Bench 2.0 went from 52.8 to 66.5 (Top 30 → Top 5).
What they added:
| Technique | What it does |
|---|---|
| Self-verification loop | Agent checks its own result before finishing |
| Pre-completion checklist | Middleware forces a verification pass before the agent can exit |
| Local context on start | Maps the working directory and available tools at the beginning |
| Loop detection | Detects the agent editing the same thing again and again |
| "Reasoning sandwich" | More reasoning effort for planning and verification, less for routine steps |
| Trace analysis | Reads failed traces to decide what to change next |
OpenAI: an agent-first repository¶
An OpenAI team built a product with Codex writing all the code: about 1M lines and ~1,500 PRs in 5 months. The lessons are about the environment around the agent:
AGENTS.mdas a table of contents (~100 lines) that points intodocs/— design docs, specs, plans.- The repository is the system of record. Knowledge that lives only in chat or documents the agent can't read does not exist for the agent.
- Make the app legible to the agent — logs, metrics, traces and UI state are available programmatically, so the agent can check its own work.
- Enforce architecture mechanically — custom linters and structural tests; error messages explain how to fix the problem.
- "Garbage collection" — recurring agent tasks scan for drift from agreed principles and open small refactoring PRs.
- Humans steer, agents execute — engineers design the environment, specify intent and build feedback loops.
Anthropic: long-running agents¶
Two common failures of long tasks: the agent tries to do too much at once, or declares victory too early. Anthropic's harness for long-running coding agents:
- An initializer agent creates a JSON feature list (every item marked
failing), a progress log (claude-progress.txt), aninit.shscript and the first git commit. - Each following session works on one feature at a time, tests it end to end in a browser, updates the progress file and commits.
- A new session reads the progress file and git log instead of relying on a huge context.
A later setup split the work into planner → generator → evaluator, where the evaluator is a separate, skeptical agent that tests the app with Playwright. Reason: models "reliably skew positive when grading their own work." The full harness took much longer and cost more than a single run — but produced a working app where the single run did not.
Pattern: Context Engineering¶
Give the model the smallest set of high-signal tokens that lets it do the job.
| Technique | When to use |
|---|---|
| Right-altitude system prompt | Clear principles, not a list of every edge case and not vague advice |
| Just-in-time retrieval | Pass references (paths, queries, IDs) and let the agent fetch details when needed |
| Compaction | Summarise old turns when the window fills up |
| Context reset | Start a fresh context from a progress file (works better than compaction for some models and long tasks) |
| Structured notes | Agent writes notes/progress to files outside the context window |
| Sub-agents | Heavy exploration happens in a separate context; only the summary comes back |
| Tool-result clearing | Drop old, large tool outputs that are no longer needed |
Pattern: Tool Design¶
Tools are part of the prompt. Bad tools make a good model look bad.
| Do | Don't |
|---|---|
| Few, high-impact tools | Dozens of overlapping tools |
Namespaced names: asana_projects_search |
Generic names: search, get |
| Return meaningful fields (names, titles) | Return opaque IDs only |
| Paginate, filter and truncate by default | Dump thousands of rows into the context |
Offer response_format: concise / detailed |
One huge format for every case |
| Actionable errors: "Date must be YYYY-MM-DD, got 05/04" | Raw tracebacks |
| Descriptions written as for a new team member | One-word descriptions |
Improve tools with evals: run tasks, let a model read the transcripts, and fix the tools where the agent struggled.
Pattern: Verification Loops¶
Never let the agent's own "done" be the final word.
- Deterministic checks — tests, linters, type checks, schema validation
- Pre-completion checklist — middleware that runs checks before the agent can finish
- Separate evaluator — another agent (or sub-agent) verifies each claim; avoids self-preference bias
- Visual checks — screenshots or browser automation for UI work
Pattern: Guardrails and Human-in-the-Loop¶
Tool call ──► PreToolUse hook ──► allowed? ──► execute ──► PostToolUse hook
│ no
▼
blocked + reason back to the model
- Permission rules — allow / ask / deny per tool or command
- Hooks — in Claude Code, a
PreToolUsehook that exits with code2blocks the call and sends its stderr back to the model - Input / output guardrails — OpenAI Agents SDK runs them in parallel with the agent and stops the run when they trip
- Human approval — pause before risky actions (deploy, delete, payment); the run resumes after approval
Example hook that blocks dangerous shell commands (Claude Code settings.json):
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [{ "type": "command", "command": "python3 .claude/hooks/block_dangerous.py" }]
}
]
}
}
# .claude/hooks/block_dangerous.py
import json
import re
import sys
event = json.load(sys.stdin)
command = event.get("tool_input", {}).get("command", "")
if re.search(r"rm\s+-rf\s+/|git\s+push\s+--force", command):
print("Blocked: destructive command. Ask the user first.", file=sys.stderr)
sys.exit(2) # 2 = block the tool call and show the reason to the model
sys.exit(0)
Keep the Harness Up to Date¶
"Every component in a harness encodes an assumption about what the model can't do on its own." — Anthropic
When models improve, some harness parts become dead weight: - Anthropic found that context resets and sprint planning were no longer needed with newer models. - Re-test the harness after each model upgrade: remove one component at a time and measure.
Anti-Patterns¶
| Anti-pattern | Symptom | Fix |
|---|---|---|
| Doing too much at once | Half-finished features, broken builds | One feature per session, progress file |
| Premature "done" | Agent reports success, tests fail | Verification loop, pre-completion checklist |
| Self-grading | Agent always rates its own work highly | Separate evaluator agent or deterministic checks |
| Context rot | Quality drops in long sessions | Compaction, resets, sub-agents, just-in-time retrieval |
| Doom loops | Same file edited over and over | Loop detection, turn limits |
| Giant instruction file | Agent ignores important rules | Short AGENTS.md as a map, details in docs/ |
| Knowledge outside the repo | Agent repeats solved mistakes | Put decisions and specs into the repository |
| Tool sprawl | Wrong tool chosen, huge outputs | Fewer tools, better names, truncation |
| Stale harness | Extra cost and latency, no quality gain | Re-evaluate components after model upgrades |
| Undisclosed harness | Benchmark results can't be reproduced | Always record model + harness configuration |
See also¶
- AI Harness
- Agent Harness — Concepts & Components
- Eval Harness — Tools, Testing & CI
- Agentic Search & Context Engineering
- OWASP LLM Security