Eval Harness — Concepts & Metrics¶
What Is an Evaluation Harness¶
An evaluation harness is the infrastructure that runs evals end to end: it gives the agent instructions and tools, runs tasks (often in parallel), records every step, grades the results and aggregates the scores.
For a QA engineer it is the test framework of the AI world:
| Classic test automation | Eval harness |
|---|---|
| Test case | Task — input + success criteria |
| Test run | Trial — one attempt at a task |
| Assertion | Grader — checks one aspect of the result |
| Test log | Transcript / trajectory — every message and tool call |
| System state after the test | Outcome — final state of the environment |
| Test suite | Suite — a set of tasks with a shared purpose |
| Pass / fail | Score, pass rate, pass@k, pass^k |
The main difference: the system under test is non-deterministic. The same task can pass once and fail the next time, so one run proves very little.
How a Harness Runs an Eval¶
Suite
├── Task 1 ──► Trial 1 ──► agent runs in a clean environment ──► transcript + outcome
│ ├──► Trial 2 ──► ...
│ └──► Trial k ──► ...
├── Task 2 ...
▼
Graders (code / model / human) ──► scores per trial
▼
Aggregation ──► pass rate, pass@k, pass^k, cost, latency ──► report / CI gate
Every trial must start from a clean, isolated environment. Shared state between trials (leftover files, caches, a warmed-up database) creates correlated failures and fake passes.
Grader Types¶
| Grader | How it works | Strengths | Weaknesses | Use for |
|---|---|---|---|---|
| Code-based | Exact match, regex, tests, schema checks, state checks | Fast, cheap, reproducible | Brittle, can reject valid alternatives | Structured output, code, final state |
| Model-based (LLM-as-judge) | A model scores the output against a rubric | Handles open-ended answers | Non-deterministic, biased, needs calibration | Tone, helpfulness, reasoning quality |
| Human | Experts review results | Gold standard | Slow, expensive | Calibrating judges, final sign-off |
Rules of thumb: - Prefer deterministic graders where possible. - Grade the outcome, not every step — rigid step checks punish valid alternative paths. - Allow partial credit for multi-part tasks. - Known LLM-judge biases: position (prefers the first answer) and verbosity (prefers longer answers). Use pass/fail or pairwise judgements and control for length. - Calibrate every judge against human labels before you trust it.
pass@k vs pass^k¶
With a per-trial success rate p and k trials:
| Metric | Meaning | Formula | Question it answers |
|---|---|---|---|
| pass@k | At least one of k trials succeeds | 1 − (1 − p)^k |
Can the agent solve it? |
| pass^k | All k trials succeed | p^k |
Can users rely on it every time? |
Example with p = 0.8:
| k | pass@k | pass^k |
|---|---|---|
| 1 | 80% | 80% |
| 3 | 99.2% | 51.2% |
| 5 | 99.97% | 32.8% |
As k grows, pass@k goes to 100% and pass^k goes to 0%. A demo needs pass@k; production needs pass^k.
Capability vs Regression Evals¶
| Capability eval | Regression eval | |
|---|---|---|
| Question | What can the system do now? | Did we break something that worked? |
| Expected pass rate | Low at the start, grows over time | Near 100% |
| In CI | Tracked as a trend | Hard gate — blocks the merge |
| Lifecycle | Tasks graduate into the regression suite once they pass reliably | Grows with every fixed bug |
Watch for saturation: when a capability suite is close to 100%, it stops telling you anything — add harder tasks.
Building an Eval Suite: Roadmap¶
- Start early and small — 20–50 tasks taken from real failures are enough to begin.
- Convert what you already have — manual checks, bug reports and support tickets become tasks.
- Write unambiguous tasks — two experts should agree on the verdict. Keep a reference solution for each task.
- Balance the cases — include negative cases (the agent should refuse, ask, or do nothing).
- Isolate the environment — clean state per trial, pinned resources.
- Choose graders — code first, model where needed, humans to calibrate.
- Read transcripts — regularly, not only when a score drops.
- Maintain the suite — fix broken tasks, retire saturated ones, add new failures.
Pitfalls¶
| Pitfall | What happens | Prevention |
|---|---|---|
| Shared state between trials | Correlated failures, fake passes | Fresh environment per trial |
| Ambiguous task spec | Graders disagree, scores are noise | Reference solution, expert review |
| Rigid step-checking | Valid alternative solutions fail | Grade the outcome |
| Buggy grader | Wrong scores look like model problems | Test graders on known good/bad answers |
| Single trial | Random luck decides the result | Several trials, report pass@k and pass^k |
| Infrastructure noise | Scores change with CPU/RAM/timeouts | Pin resources; treat them as a variable |
| Contamination | Benchmark data leaked into training → inflated scores | Private task sets; decontamination checks |
| "Vibe-based evals" | Decisions made on a few manual chats | Logged, repeatable suites |
Infrastructure alone can move agent benchmark scores by several points — Anthropic measured a swing of about 6 points on Terminal-Bench 2.0 from resource configuration only.
Non-Determinism and Cost¶
- Run each task several times (epochs / trials) and report the spread, not just the mean.
- Set limits per sample: messages, tokens, time.
- Record and replay model responses for tests that don't need a live model.
- Use smaller models and low max-token limits in integration tests; save full runs for nightly jobs.
- Cache results in CI where the tool supports it.