Testing and evaluation · reviewed · reviewed Oct 5, 2026 · 9 min
How do you improve a coding-agent harness?
Harness engineering improves the runtime around a model: locate a failure in the trace, change the responsible context, tool, policy, state, or completion mechanism, and compare the candidate with a frozen baseline using contract tests and repeated task trials.
Choose the mechanism before changing the model
A coding agent may fail because it cannot find a file, cannot express a safe edit, loses an instruction, misreads a command result, or stops before checking the change. These failures need different repairs. “Use a stronger model” is a candidate intervention, but it does not fix a tool that drops exit codes or a runtime that accepts stale evidence.
Start with one failed task and inspect the first divergence between the required behaviour and the observed trace. Write a causal hypothesis: “The agent missed the failing assertion because the test tool returned only the last lines of output.” That suggests changing the tool response. “The agent received the assertion and ignored it” calls for a different investigation.
For a team configuring an existing agent, the available controls might be repository instructions, tool adapters, skills, trusted hooks, and permissions. A runtime maintainer can also change parsing, context assembly, persistence, and the loop. Choose a control you actually own; configuration cannot repair a defect inside an inaccessible executor.
Preserve the assertion before adding another instruction
A fictional test tool returns only the last five lines of output. The failing assertion occurs earlier and disappears from context. Adding “read errors carefully” cannot make an omitted assertion available.
Compare a candidate response contract that preserves exit status, the failing assertion, and an artifact pointer for the full log. First replay a fixed failure and prove the required evidence reaches the next context. Then measure whether real-agent trials improve without an unacceptable latency or cost increase.
This intervention is a hypothesis. A summary that extracts the wrong assertion can introduce a different failure, so retain full evidence and negative fixtures. The premature-completion example later in the article applies the same reasoning to a stop gate.
Map the failure to a small intervention
The following examples are engineering hypotheses to evaluate, not claims that a particular product implements them or that they always improve performance.
| Observed failure | Candidate change | Evidence to inspect |
|---|---|---|
| Relevant test failure is buried in large output | Return exit status, failure summary, and an artifact pointer for full output | Required assertion reaches context; omitted detail remains retrievable |
| Similar tools cause repeated invalid calls | Clarify tool purpose and arguments; remove redundant choices | Invalid-call rate and valid task outcomes |
| An edit targets stale file content | Reject stale edits and return current content for a fresh proposal | User changes survive; agent can recover within budget |
| A summary loses a repository constraint | Preserve the constraint and its source across compaction | Post-compaction trials still obey the constraint |
| A permitted action is repeatedly blocked | Narrow an overbroad policy rule using positive and negative fixtures | Legitimate action succeeds; forbidden action still produces no effect |
| Work is declared complete before checks | Bind a stop gate to current artifacts and required evidence | Premature stops fail; fully verified work can finish |
| Delegated workers overwrite each other | Isolate artifacts and define a merge owner | No lost edits; integration checks pass |
Tool ergonomics are part of the agent–computer interface: what the model can request, what the tool returns, and how errors guide the next step. A narrowly scoped search with useful surrounding context may be easier to use than a tool that dumps every log line. Keep the full evidence available outside the model's immediate working set.
Put guardrails at the boundary they control
A guardrail is a control intended to prevent or detect an unacceptable outcome. Describe its mechanism and failure policy rather than treating the word as a safety guarantee.
- Instructions and model-based screening influence or classify proposals. A screening model can trigger a blocking decision, but its judgement still has false positives and false negatives.
- Validation, authorization, and approval gate an exact operation before execution. Resolve the target, check the current principal and scope, and bind approval to the payload that will execute.
- Isolation and resource controls constrain execution: accessible files and networks, process privileges, output size, and resource consumption. They do not decide whether a business action is authorized.
For example, “never edit production configuration” in a prompt is useful guidance. An executor that rejects writes to the resolved production path enforces that boundary even if the model ignores the instruction. A sandbox limits the environment reachable by the process. These controls answer different questions and can be layered.
Specify what happens when a control is unavailable. An authorization check should not become permission when it times out; optional telemetry may have a different failure policy. Denied actions must not be retried through another tool to bypass the same restriction. The detailed contracts belong in tool permissions and agent sandboxes.
Bound the whole run, including retries and workers
A turn limit prevents one kind of runaway loop. It does not bound a long-running command, an expensive model call, or several workers consuming separate budgets. Define each limit, its owner, what counts against it, and the state reported when it is reached.
| Budget | Observable contract to verify |
|---|---|
| Turns and tool calls | Calls, retries, and child runs consume the declared allowance; resuming cannot silently reset it |
| Wall-clock time | Model calls, tool execution, waits, and retries share a deadline; cancellation also reaches subprocesses |
| Tokens and output bytes | Each request fits its context allowance; large results preserve a retrievable full artifact |
| Model and external-service cost | Reserve estimated cost before dispatch and reconcile actual usage; include failed and retried calls |
| Retry and repair attempts | One layer owns retries, with bounded attempts and backoff; ambiguous effects require reconciliation |
| Worker concurrency and shared resources | Child runs consume a shared parent allowance; parallel workers cannot each spend the entire remaining budget |
Cost estimates and provider billing can differ. Call a monetary cap approximate unless the accounting and execution boundary actually enforce a maximum charge. Context limits, response limits, and total run token allowances are also separate controls.
When a budget is exhausted, prevent further dispatch, request cancellation of in-flight work where supported, preserve partial artifacts, and report why the run stopped. Cancellation cannot undo an external effect already sent. If confirmation was lost, preserve an unknown outcome for reconciliation instead of reporting success or assuming that nothing happened.
A worked hypothesis: premature completion
Illustrative scenario; no benchmark results are claimed. An agent edits a parser and reports success. The repository contains the requested change, but the required checks never ran. The failure is premature completion, regardless of how convincing the final explanation sounds.
Compare two configurations with the same model, task, repository snapshot, tool versions, and budgets. The baseline accepts the model's stop proposal. The candidate adds a validator requiring the declared checks to pass for the final relevant artifact. On missing evidence, it returns the exact unmet requirement and permits bounded continuation. On exhausted budget, it reports incomplete work.
flowchart TB
P[Stop proposal] --> V{Evidence valid?}
V -->|Yes| C[Complete]
V -->|No| B{Budget left?}
B -->|Yes| R[Return missing evidence]
B -->|No| I[Stop incomplete]
First prove the validator's contract with scripted inputs: absent checks, failed checks, stale passing checks, and valid current evidence. Include the positive case so an overstrict guard cannot pass the suite merely by blocking every stop.
Next run repeated real-model trials from clean environments. Measure valid completion, false completion, incomplete work, extra calls, elapsed time, and total cost. A guard can eliminate false completion while making more tasks incomplete; that trade-off must remain visible. It has not necessarily improved the model's ability to fix the parser.
Freeze what is not being tested
Keep a comparison manifest containing the task and repository revisions, model identifier and sampling settings, runtime revision, instructions, enabled tools, policy, environment, grader version, budgets, and cache conditions. Reset files and session memory between independent trials. If the provider cannot pin a model snapshot, record that limitation and interleave baseline and candidate runs to reduce time-related bias.
Use failures that informed the change for development. Reserve other tasks for checking whether it generalizes. Run more than one trial per task when judging stochastic behaviour, report counts and uncertainty, and inspect regressions by task type. A scripted replay tests runtime behaviour; it cannot establish how a live model would react to changed tool output.
An ablation removes one candidate mechanism while retaining the rest of the configuration. It helps ask whether that mechanism explains the improvement. Do not remove a required safety boundary in a live environment; study it with controlled doubles or disposable fixtures.
Decide whether to keep the change
Before running the comparison, define what would justify adoption: fewer false completions, better recovery, lower cost at the same acceptable outcome rate, or a specific risk reduced without an unacceptable regression. Keep severe effects separate from average task success; extra successes cannot compensate for a newly unauthorized write.
Include failed attempts, retries, verification, and human repair when comparing cost. Preserve counterexamples as regression cases and keep the baseline available for rollback. Add mechanisms in response to evidence: a broader tool catalog, more persistent memory, or additional workers also creates new maintenance and failure surfaces.
Evaluate protection and useful work together
Keep three kinds of evidence separate. Contract tests use scripted proposals and fault injection to prove the runtime boundary. Agent evals use repeated live-model trials to measure task outcomes and trajectories. Adversarial evals challenge the system with malicious instructions and tempting forbidden actions. A pass in one suite does not imply a pass in the others.
Every protective control needs cases in both directions: a forbidden operation must be rejected, while a permitted operation remains possible. Track unauthorized effects, attack success, false refusals, legitimate task completion, recovery, incomplete outcomes, and latency and cost distributions. A system that blocks all actions can look secure while being unusable.
Add boundary cases such as the last permitted call, budget exhaustion, a policy-service timeout, changed approval arguments, and an unknown effect after cancellation. Version the cases and graders; keep failures used for tuning separate from held-out challenges. The evaluation loop explains the comparison process, and evaluation metrics explains how to interpret repeated trials.
Research to read, and what it can establish
These papers support different engineering questions. Their benchmark scores are not interchangeable release criteria for your application.
- Harness Engineering supplies a source-level architecture map; it does not measure comparative runtime performance.
- AgentDojo evaluates prompt injections through untrusted tool data, measuring both attack success and legitimate task utility. It is useful for checking whether a defense protects the agent without merely disabling useful work.
- τ-bench studies tool use with simulated users and domain policies. Its outcome checks and repeated-trial reliability are useful models for evaluating stateful workflows; its domains do not represent every coding task.
- SWE-bench evaluates patches for repository issues using tests for the fix and existing behaviour. Passing those tests is evidence about the benchmark task, not a complete audit of permissions, resource bounds, or deployment readiness.
Use research to choose a measurement method, then build cases around your own task and effect boundaries. The result should be a defensible runtime change and a record of what it improved, what it cost, and what remains uncertain.
Sources
Sources and further reading
- 01Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents — A Source-Code Study of Eleven SystemsPaul Barbaste, Tristan Darrigol, Germain Vu, and Tom Wiltberger · research · published Jul 15, 2026 · source checked Oct 5, 2026
Version 1 architecture study of dated coding-agent source snapshots; useful for subsystem vocabulary, not comparative performance claims.
- 02Writing effective tools for AI agents—using AI agentsAnthropic · guide · published Sep 11, 2025 · source checked Oct 5, 2026
First-party guidance on tool ergonomics, meaningful output, error feedback, and evaluations for improving an agent action interface.
- 03Demystifying evals for AI agentsAnthropic · guide · source checked Oct 5, 2026
A practical framework for tasks, trials, graders, transcripts, outcomes, and agent evaluation design.
- 04Excessive AgencyOWASP GenAI Security Project · standard · source checked Oct 5, 2026
A threat model organized around excessive functionality, permissions, and autonomy.
- 05A practical guide to building agentsOpenAI · guide · source checked Oct 5, 2026
Design guidance for tools, orchestration, guardrails, risk ratings, and human intervention.
- 06AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsDebenedetti et al. · research · published Jun 19, 2024 · source checked Oct 5, 2026
Primary evaluation of prompt injections in untrusted tool output, separating legitimate task utility from attacker success when comparing defenses.
- 07τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsYao, Shinn, Razavi, and Narasimhan · research · published Jun 17, 2024 · source checked Oct 5, 2026
Primary benchmark for policy-constrained tool workflows, simulated user interaction, observable database outcomes, and reliability across repeated trials.
- 08SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Jimenez et al. · research · published Oct 10, 2023 · source checked Oct 5, 2026
Primary repository-issue benchmark using fix and regression tests to evaluate generated patches; does not by itself certify safe deployment or tool authority.
