Testing and evaluation · reviewed · reviewed Oct 5, 2026 · 6 min
How do you test an agent harness?
Test the harness as an ordinary stateful system by replacing the model and tools with controlled doubles, then proving validation, authorization, effects, retries, state transitions, stopping, and trace evidence at every boundary.
Test the boundary, not the prompt
An agent harness is mostly application code. It builds context, interprets model output, validates proposed actions, checks identity and policy, invokes tools, records observations, updates state, applies budgets, and decides whether the run may continue. Those responsibilities can be tested without asking a live model to behave consistently.
Replace the model with a scripted double that returns an exact sequence of proposals. Replace external tools with adapters that record calls and can return success, rejection, timeout, malformed output, or an unknown outcome after dispatch. The test then controls every input and observes every state transition.
flowchart LR C[Scripted context and state] --> M[Model double] M --> P[Exact proposal] P --> V[Schema and semantic checks] V --> A[Identity, policy, approval] A --> T[Tool adapter] T --> O[Observation or uncertainty] O --> S[State and trace assertions]
This seam lets a test prove that unsafe proposals produce zero external effect. It also makes failures reproducible: the same proposal, identity, policy, and state should produce the same harness decision.
Make denial observable without a live model
A scripted model double proposes a refund against another tenant's order. A fake tool records every attempted dispatch. The runtime should reject before the fake tool is invoked.
given: principal tenant A; target tenant B; valid JSON proposal
when: harness evaluates the exact operation
then: denial observation + policy event + zero dispatches
Add the positive fixture too: a valid, authorized request with exact approval reaches the adapter once. A runtime that rejects everything can pass the negative case while being unusable.
These controlled fixtures prove the execution contract. Separate repeated live-model trials are still needed to learn how often an agent selects a useful proposal and obtains valid completion evidence.
Build a contract matrix
For every tool, record the tool name and version, argument schema, semantic constraints, principal, permitted resource scope, risk tier, required approval, rate and cost limits, idempotency strategy, expected observations, and log-redaction rules. Derive positive and negative tests from that matrix.
A schema-valid call can still be wrong. A path may escape the allowed directory. An account ID may belong to another tenant. A refund amount may exceed the current order balance. An approval may refer to a different operation or have expired. The harness must check current meaning and authority immediately before the effect.
Use at least these proposal classes:
- known tool with valid, authorized arguments;
- unknown tool or unsupported version;
- malformed output and extra fields;
- valid shape with invalid business meaning;
- cross-tenant, traversal, or stale-resource identifier;
- correct operation without required approval;
- replay of a previously accepted operation;
- proposal that exceeds a turn, token, time, or money budget.
Test the state machine
Model a run with explicit states such as ready, proposing, awaiting_approval, executing, observing, recovering, completed, failed, and cancelled. Then test which transitions are allowed and what evidence each transition requires.
Completion deserves its own validator. A polished final message is not evidence that the requested file exists, the database commit succeeded, or the operation stayed inside policy. The harness should enter a terminal success state only after checking the declared completion evidence. Cancellation and failure must prevent further effects.
Useful state-machine properties include:
- one accepted proposal creates at most one logical effect;
- a denied proposal creates no effect;
- a terminal run cannot execute another tool;
- an approval is bound to one exact pending operation;
- current identity and policy are checked again before execution and retry;
- every transition records enough evidence to reconstruct what happened.
Inject failures around the effect
The most important retry bug occurs when a request may have reached the external system but the response did not reach the harness. A timeout is not proof of failure. Treat the result as unknown until the harness can reconcile by operation ID or inspect the target state.
Inject failures before dispatch, during dispatch, after the external effect, while receiving the result, while persisting the observation, and while writing the next run state. Verify which failures are safe to retry, which require reconciliation, and which must stop for human review.
Also test duplicate delivery, reordered events, expired credentials, rate limits, partial tool output, oversized output, hostile content inside tool results, and loss of the process between effect and checkpoint. The trace must distinguish a tool failure from a harness failure and an unknown outcome from a confirmed no-op.
Do not ask the model to certify its harness
Harness testing is not asking a model to follow a safety prompt ten times. It is not satisfied by validating JSON against a schema, because schemas do not establish ownership, permission, freshness, or business meaning. It is not an end-to-end agent score: model capability and stochastic trajectory quality are separate concerns.
A framework's own unit tests do not prove the product policy. The application team still owns tool scope, tenant boundaries, approvals, budgets, state transitions, recovery, and terminal evidence.
Make a coding-agent stop gate falsifiable
For a code-change task, bind completion evidence to the current repository revision, the declared commands, and their actual results. A successful check from before the last edit cannot certify the final artifact. Nor can a zero exit status prove that the intended tests were discovered.
These are example fixtures for a workflow that requires tests before completion; they are not a universal rule that every task needs the same suite.
| Scripted condition | Expected runtime result |
|---|---|
| Model says “done”; required checks never ran | No successful completion; identify missing evidence |
| Required check failed but final prose claims success | Preserve the failure and allow only bounded repair |
| Checks passed, then the relevant artifact changed | Invalidate affected evidence and rerun required checks |
| Test runner exited successfully but discovered no required tests | Reject that evidence as insufficient |
| Checks passed for the final artifact and scope is satisfied | Allow completion; do not create an endless stop-hook loop |
| Repair exhausted its deadline or attempt budget | Stop with an explicit incomplete result |
Test the validator with fixed command records and artifact identities first. Then use real-model trials to measure how often the agent obtains valid evidence without excessive repair. Runtime correctness and the model's ability to reach that state are different claims. The harness-engineering guide explains how to compare a candidate change.
Verify deterministic contracts first
Use a layered suite:
- Unit-test pure context selection, validation, policy, budgets, and transition functions.
- Contract-test each tool adapter against a fake or disposable service.
- Run deterministic harness scenarios with scripted model proposals and injected tool outcomes.
- Property-test invariants such as zero effect after denial and no action after a terminal state.
- Run a small number of real-model integration trials to verify the boundary is connected correctly.
- Re-run dangerous and previously failing traces as permanent regression fixtures.
Assert observable contracts rather than private implementation details. The strongest evidence is a trace that shows the proposal, the exact validation and policy decision, whether dispatch occurred, the observed outcome, the resulting state, and why the run stopped.
Sources
Sources and further reading
- 01A practical guide to building agentsOpenAI · guide · source checked Oct 5, 2026
Design guidance for tools, orchestration, guardrails, risk ratings, and human intervention.
- 02Excessive AgencyOWASP GenAI Security Project · standard · source checked Oct 5, 2026
A threat model organized around excessive functionality, permissions, and autonomy.
- 03Artificial Intelligence Risk Management Framework 1.0NIST · standard · published Jan 26, 2023 · source checked Oct 6, 2026
A system-lifecycle framework for mapping context, measuring trustworthiness, and managing AI risk.
- 04Demystifying evals for AI agentsAnthropic · guide · source checked Oct 5, 2026
A practical framework for tasks, trials, graders, transcripts, outcomes, and agent evaluation design.
