Awesome Testing

Testing and evaluation · reviewed · reviewed Oct 5, 2026 · 9 min

How do you improve a coding-agent harness?

Harness engineering improves the runtime around a model: locate a failure in the trace, change the responsible context, tool, policy, state, or completion mechanism, and compare the candidate with a frozen baseline using contract tests and repeated task trials.

Choose the mechanism before changing the model

A coding agent may fail because it cannot find a file, cannot express a safe edit, loses an instruction, misreads a command result, or stops before checking the change. These failures need different repairs. “Use a stronger model” is a candidate intervention, but it does not fix a tool that drops exit codes or a runtime that accepts stale evidence.

Start with one failed task and inspect the first divergence between the required behaviour and the observed trace. Write a causal hypothesis: “The agent missed the failing assertion because the test tool returned only the last lines of output.” That suggests changing the tool response. “The agent received the assertion and ignored it” calls for a different investigation.

For a team configuring an existing agent, the available controls might be repository instructions, tool adapters, skills, trusted hooks, and permissions. A runtime maintainer can also change parsing, context assembly, persistence, and the loop. Choose a control you actually own; configuration cannot repair a defect inside an inaccessible executor.

Preserve the assertion before adding another instruction

A fictional test tool returns only the last five lines of output. The failing assertion occurs earlier and disappears from context. Adding “read errors carefully” cannot make an omitted assertion available.

Compare a candidate response contract that preserves exit status, the failing assertion, and an artifact pointer for the full log. First replay a fixed failure and prove the required evidence reaches the next context. Then measure whether real-agent trials improve without an unacceptable latency or cost increase.

This intervention is a hypothesis. A summary that extracts the wrong assertion can introduce a different failure, so retain full evidence and negative fixtures. The premature-completion example later in the article applies the same reasoning to a stop gate.

Map the failure to a small intervention

The following examples are engineering hypotheses to evaluate, not claims that a particular product implements them or that they always improve performance.

Observed failureCandidate changeEvidence to inspect
Relevant test failure is buried in large outputReturn exit status, failure summary, and an artifact pointer for full outputRequired assertion reaches context; omitted detail remains retrievable
Similar tools cause repeated invalid callsClarify tool purpose and arguments; remove redundant choicesInvalid-call rate and valid task outcomes
An edit targets stale file contentReject stale edits and return current content for a fresh proposalUser changes survive; agent can recover within budget
A summary loses a repository constraintPreserve the constraint and its source across compactionPost-compaction trials still obey the constraint
A permitted action is repeatedly blockedNarrow an overbroad policy rule using positive and negative fixturesLegitimate action succeeds; forbidden action still produces no effect
Work is declared complete before checksBind a stop gate to current artifacts and required evidencePremature stops fail; fully verified work can finish
Delegated workers overwrite each otherIsolate artifacts and define a merge ownerNo lost edits; integration checks pass

Tool ergonomics are part of the agent–computer interface: what the model can request, what the tool returns, and how errors guide the next step. A narrowly scoped search with useful surrounding context may be easier to use than a tool that dumps every log line. Keep the full evidence available outside the model's immediate working set.

Put guardrails at the boundary they control

A guardrail is a control intended to prevent or detect an unacceptable outcome. Describe its mechanism and failure policy rather than treating the word as a safety guarantee.

  • Instructions and model-based screening influence or classify proposals. A screening model can trigger a blocking decision, but its judgement still has false positives and false negatives.
  • Validation, authorization, and approval gate an exact operation before execution. Resolve the target, check the current principal and scope, and bind approval to the payload that will execute.
  • Isolation and resource controls constrain execution: accessible files and networks, process privileges, output size, and resource consumption. They do not decide whether a business action is authorized.

For example, “never edit production configuration” in a prompt is useful guidance. An executor that rejects writes to the resolved production path enforces that boundary even if the model ignores the instruction. A sandbox limits the environment reachable by the process. These controls answer different questions and can be layered.

Specify what happens when a control is unavailable. An authorization check should not become permission when it times out; optional telemetry may have a different failure policy. Denied actions must not be retried through another tool to bypass the same restriction. The detailed contracts belong in tool permissions and agent sandboxes.

Bound the whole run, including retries and workers

A turn limit prevents one kind of runaway loop. It does not bound a long-running command, an expensive model call, or several workers consuming separate budgets. Define each limit, its owner, what counts against it, and the state reported when it is reached.

BudgetObservable contract to verify
Turns and tool callsCalls, retries, and child runs consume the declared allowance; resuming cannot silently reset it
Wall-clock timeModel calls, tool execution, waits, and retries share a deadline; cancellation also reaches subprocesses
Tokens and output bytesEach request fits its context allowance; large results preserve a retrievable full artifact
Model and external-service costReserve estimated cost before dispatch and reconcile actual usage; include failed and retried calls
Retry and repair attemptsOne layer owns retries, with bounded attempts and backoff; ambiguous effects require reconciliation
Worker concurrency and shared resourcesChild runs consume a shared parent allowance; parallel workers cannot each spend the entire remaining budget

Cost estimates and provider billing can differ. Call a monetary cap approximate unless the accounting and execution boundary actually enforce a maximum charge. Context limits, response limits, and total run token allowances are also separate controls.

When a budget is exhausted, prevent further dispatch, request cancellation of in-flight work where supported, preserve partial artifacts, and report why the run stopped. Cancellation cannot undo an external effect already sent. If confirmation was lost, preserve an unknown outcome for reconciliation instead of reporting success or assuming that nothing happened.

A worked hypothesis: premature completion

Illustrative scenario; no benchmark results are claimed. An agent edits a parser and reports success. The repository contains the requested change, but the required checks never ran. The failure is premature completion, regardless of how convincing the final explanation sounds.

Compare two configurations with the same model, task, repository snapshot, tool versions, and budgets. The baseline accepts the model's stop proposal. The candidate adds a validator requiring the declared checks to pass for the final relevant artifact. On missing evidence, it returns the exact unmet requirement and permits bounded continuation. On exhausted budget, it reports incomplete work.

flowchart TB
  P[Stop proposal] --> V{Evidence valid?}
  V -->|Yes| C[Complete]
  V -->|No| B{Budget left?}
  B -->|Yes| R[Return missing evidence]
  B -->|No| I[Stop incomplete]

First prove the validator's contract with scripted inputs: absent checks, failed checks, stale passing checks, and valid current evidence. Include the positive case so an overstrict guard cannot pass the suite merely by blocking every stop.

Next run repeated real-model trials from clean environments. Measure valid completion, false completion, incomplete work, extra calls, elapsed time, and total cost. A guard can eliminate false completion while making more tasks incomplete; that trade-off must remain visible. It has not necessarily improved the model's ability to fix the parser.

Freeze what is not being tested

Keep a comparison manifest containing the task and repository revisions, model identifier and sampling settings, runtime revision, instructions, enabled tools, policy, environment, grader version, budgets, and cache conditions. Reset files and session memory between independent trials. If the provider cannot pin a model snapshot, record that limitation and interleave baseline and candidate runs to reduce time-related bias.

Use failures that informed the change for development. Reserve other tasks for checking whether it generalizes. Run more than one trial per task when judging stochastic behaviour, report counts and uncertainty, and inspect regressions by task type. A scripted replay tests runtime behaviour; it cannot establish how a live model would react to changed tool output.

An ablation removes one candidate mechanism while retaining the rest of the configuration. It helps ask whether that mechanism explains the improvement. Do not remove a required safety boundary in a live environment; study it with controlled doubles or disposable fixtures.

Decide whether to keep the change

Before running the comparison, define what would justify adoption: fewer false completions, better recovery, lower cost at the same acceptable outcome rate, or a specific risk reduced without an unacceptable regression. Keep severe effects separate from average task success; extra successes cannot compensate for a newly unauthorized write.

Include failed attempts, retries, verification, and human repair when comparing cost. Preserve counterexamples as regression cases and keep the baseline available for rollback. Add mechanisms in response to evidence: a broader tool catalog, more persistent memory, or additional workers also creates new maintenance and failure surfaces.

Evaluate protection and useful work together

Keep three kinds of evidence separate. Contract tests use scripted proposals and fault injection to prove the runtime boundary. Agent evals use repeated live-model trials to measure task outcomes and trajectories. Adversarial evals challenge the system with malicious instructions and tempting forbidden actions. A pass in one suite does not imply a pass in the others.

Every protective control needs cases in both directions: a forbidden operation must be rejected, while a permitted operation remains possible. Track unauthorized effects, attack success, false refusals, legitimate task completion, recovery, incomplete outcomes, and latency and cost distributions. A system that blocks all actions can look secure while being unusable.

Add boundary cases such as the last permitted call, budget exhaustion, a policy-service timeout, changed approval arguments, and an unknown effect after cancellation. Version the cases and graders; keep failures used for tuning separate from held-out challenges. The evaluation loop explains the comparison process, and evaluation metrics explains how to interpret repeated trials.

Research to read, and what it can establish

These papers support different engineering questions. Their benchmark scores are not interchangeable release criteria for your application.

  • Harness Engineering supplies a source-level architecture map; it does not measure comparative runtime performance.
  • AgentDojo evaluates prompt injections through untrusted tool data, measuring both attack success and legitimate task utility. It is useful for checking whether a defense protects the agent without merely disabling useful work.
  • τ-bench studies tool use with simulated users and domain policies. Its outcome checks and repeated-trial reliability are useful models for evaluating stateful workflows; its domains do not represent every coding task.
  • SWE-bench evaluates patches for repository issues using tests for the fix and existing behaviour. Passing those tests is evidence about the benchmark task, not a complete audit of permissions, resource bounds, or deployment readiness.

Use research to choose a measurement method, then build cases around your own task and effect boundaries. The result should be a defensible runtime change and a record of what it improved, what it cost, and what remains uncertain.

Sources and further reading

  1. 01
    Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents — A Source-Code Study of Eleven SystemsPaul Barbaste, Tristan Darrigol, Germain Vu, and Tom Wiltberger · research · published Jul 15, 2026 · source checked Oct 5, 2026

    Version 1 architecture study of dated coding-agent source snapshots; useful for subsystem vocabulary, not comparative performance claims.

  2. 02
    Writing effective tools for AI agents—using AI agentsAnthropic · guide · published Sep 11, 2025 · source checked Oct 5, 2026

    First-party guidance on tool ergonomics, meaningful output, error feedback, and evaluations for improving an agent action interface.

  3. 03
    Demystifying evals for AI agentsAnthropic · guide · source checked Oct 5, 2026

    A practical framework for tasks, trials, graders, transcripts, outcomes, and agent evaluation design.

  4. 04
    Excessive AgencyOWASP GenAI Security Project · standard · source checked Oct 5, 2026

    A threat model organized around excessive functionality, permissions, and autonomy.

  5. 05
    A practical guide to building agentsOpenAI · guide · source checked Oct 5, 2026

    Design guidance for tools, orchestration, guardrails, risk ratings, and human intervention.

  6. 06
    AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsDebenedetti et al. · research · published Jun 19, 2024 · source checked Oct 5, 2026

    Primary evaluation of prompt injections in untrusted tool output, separating legitimate task utility from attacker success when comparing defenses.

  7. 07
    τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsYao, Shinn, Razavi, and Narasimhan · research · published Jun 17, 2024 · source checked Oct 5, 2026

    Primary benchmark for policy-constrained tool workflows, simulated user interaction, observable database outcomes, and reliability across repeated trials.

  8. 08
    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Jimenez et al. · research · published Oct 10, 2023 · source checked Oct 5, 2026

    Primary repository-issue benchmark using fix and regression tests to evaluate generated patches; does not by itself certify safe deployment or tool authority.