AI Harness¶
A harness is everything around the model that turns it into a working system. The word is used for two different things, and both matter to QA engineers:
- Agent harness — the runtime that lets a model act as an agent: the loop, tools, context, memory, permissions, sandbox.
- Evaluation harness — the infrastructure that runs evals end to end: tasks, trials, graders, metrics, reports.
"Agent = model + harness." — LangChain
The same model can score very differently depending on its harness. Changing only the harness (same model) moved LangChain's coding agent on Terminal Bench 2.0 from 52.8 → 66.5. So when you test an AI system, you are mostly testing its harness.
Sections¶
| File | Topics |
|---|---|
| Agent Harness — Concepts & Components | What a harness is, agent loop, 8 harness jobs, components, how Anthropic / OpenAI / LangChain use the term |
| Agent Harness — Patterns & Anti-Patterns | Context engineering, tool design, long-running agents, verification loops, guardrails, harness engineering case studies, failure modes |
| Eval Harness — Concepts & Metrics | Tasks, trials, graders, transcripts, pass@k vs pass^k, capability vs regression evals, building an eval suite |
| Eval Harness — Tools, Testing & CI | lm-evaluation-harness, Inspect AI, DeepEval, promptfoo, Ragas, LangSmith, Braintrust, Harbor; testing an agent harness; evals in CI/CD |
| Building a Harness with Jev | System One models, Jev question types (Noul / Choice / Score), LangChain integration, model routing and AutoMode guardrail middleware, testing probabilistic decisions |
Agent Harness vs Eval Harness¶
| Agent Harness | Evaluation Harness | |
|---|---|---|
| Purpose | Make the model do the task | Measure how well the system does the task |
| Runs | In production and in tests | In development, CI, before releases |
| Main parts | Loop, tools, context, memory, permissions, sandbox, hooks | Tasks, trials, graders, environment, metrics, reports |
| Output | Actions and results | Scores, transcripts, pass rates |
| Who owns it | Product / platform engineers | QA / eval engineers (often the same team) |
Architecture Overview¶
flowchart TD
subgraph EVAL[Evaluation Harness]
direction TB
T[Tasks × trials] --> AGENT
subgraph AGENT[Agent Harness]
direction LR
M[Model / LLM] <--> H[Loop · context · memory<br/>tools · sandbox · hooks]
end
AGENT --> O[Transcript + final state]
O --> G[Graders → aggregated scores]
end
Quick Start: Checklist for QA Engineers¶
- Name the harness — write down model, system prompt, tools, limits and sandbox. You can't compare results without it.
- Unit-test the tools — tools are plain code: test them without a model.
- Replay model calls — record real responses once, replay them in CI.
- Collect 20–50 real tasks — from bugs, support tickets and manual checks.
- Pick graders — code-based first, LLM-as-judge where needed, humans to calibrate.
- Run several trials — report pass@k and pass^k, not one lucky run.
- Read transcripts — scores tell you that it failed, transcripts tell you why.
- Gate on regressions — regression suite near 100% blocks the merge; capability suite is a trend.
See also¶
- Digital Garden: Knowledge Base
- Agentic AI Architecture
- AI Skills for Coding Agents
- DeepEval — LLM Testing Guide
- OWASP LLM Security