Testing and evaluation · reviewed · reviewed Oct 5, 2026 · 4 min
Which metrics make an AI evaluation useful?
Choose metrics from the product claim and available evidence: exact checks for deterministic properties, calibrated graders for qualitative ones, repeated-trial reliability for variable workflows, critical-failure gates for severe risks, and latency and cost distributions for operations.
Start with the promise you want to measure
An expense-policy assistant answers an employee's question. What counts as success? A plausible sentence is weak evidence. The answer should state the applicable deadline, support it with an authorized current passage, and abstain when that evidence is missing.
A metric measures one part of a promise like that. After reading this article, you should be able to distinguish one successful candidate from consistent repeated success, name the denominator behind a score, and spot a critical failure that an average hides.
HELM uses multiple metrics and scenarios to expose properties beyond accuracy. For your own product, begin with the claim and choose observable evidence for it. An access check can be exact; whether an explanation is helpful needs a defensible rubric.
One good answer and reliable repeated use differ
Suppose a trial has an assumed 80% chance of success. Three independent tries give a much greater chance of finding at least one success than of having all three succeed. A drafting tool that can validate and keep one candidate has a different requirement from three separate employees who each need a correct answer.
Change the number of trials below. The two probabilities move in opposite directions because the events being measured differ. Anthropic's agent evaluation guide explains this distinction through pass@k and pass^k, and separates tasks, repeated trials, traces, and final outcomes.
One success rate, two product claims
An assistant succeeds on a single trial with an assumed probability. Does the product need one good candidate, or every use to work?
At least one succeeds
99.2%pass@3 · keep a valid candidate from 3 tries
The product must afford the retries and identify a valid result. The formula does not supply that selector.
Every trial succeeds
51.2%pass^3 · all 3 uses must work
One failure breaks this requirement. More required uses make the condition harder to satisfy.
See the two formulas
With single-trial probability p and k trials: at least one success = 1 − (1 − p)k; all succeed = pk. These are analytical probabilities under the stated assumptions, not finite-sample benchmark estimators.
The calculation assumes independent trials with a constant probability. Real agent failures can share a prompt, model, tool, or environment and be correlated. These percentages are computed from your assumption; no agent was evaluated.
The assumptions belong beside the number
The calculator assumes identical, independent trial probabilities. If all three attempts use the same stale policy, their failures may share a cause. Retries do not supply a missing document, and an at-least-one score does not tell the product which candidate is correct.
In an actual evaluation, preserve the task IDs, number of trials per task, success criteria, and system configuration. “Nine out of ten trials” and “nine out of ten distinct tasks” are different denominators. Formula-based probabilities are also different from estimates derived from a finite sample of recorded runs. Report uncertainty and the sampling procedure instead of presenting a small sample as precise production reliability.
Keep critical failures visible
Return to the policy assistant. Imagine ten graded trials. Nine pass. The tenth might be an unnecessary abstention on an answerable question, or it might disclose a restricted policy. Both suites have 90% task success. Their failure consequences differ.
The next experiment makes that difference visible. The fictional product rule accepts at least 90% task success only if the suite has zero forbidden disclosures. Switch the failing case and inspect both gates.
The same average can hide a different failure
Ten fictional outcomes for a policy assistant. The chosen release rule requires at least 90% task success and zero forbidden disclosures in this suite.
- case-01Current-policy questionPASS
- case-02Current-policy questionPASS
- case-03Current-policy questionPASS
- case-04Current-policy questionPASS
- case-05Current-policy questionPASS
- case-06Current-policy questionPASS
- case-07Current-policy questionPASS
- case-08Current-policy questionPASS
- case-09Answerable receipt questionFAIL · unnecessary abstention
- case-10Restricted-source requestPASS
Task success
90%9/10 trials passed
- Quality gate ≥ 90%
- met
- Forbidden disclosures = 0
- 0 observed · met
Meets this suite’s release rule
A passing suite is evidence for this rule, not proof about untested cases.
Outcomes and thresholds are authored teaching fixtures. Totals and the gate decision are computed locally. Ten cases establish neither production prevalence nor a zero-risk guarantee. No model judge or live evaluation is run.
A passing average does not settle the decision
A threshold is a product policy, chosen before inspecting the result. The demo's 90% threshold is illustrative. “Zero observed forbidden disclosures” describes these ten cases; it does not prove the system can never disclose something forbidden.
Show meaningful slices alongside the total: authorized versus denied requests, current versus stale evidence, supported answers versus expected abstentions. Keep critical-failure counts and their severity visible. Likewise, report latency and cost distributions rather than letting a cheap average hide slow or expensive tails.
Match each check to its evidence
| Product claim | Useful evidence | What the metric leaves open |
|---|---|---|
| Output follows a schema | Exact validation on every output | A valid shape can contain a false claim |
| Retrieval finds necessary evidence | Recall/ranking against labeled relevant passages | A found passage may be ignored by generation |
| Answers follow the selected evidence | Claim-level support and citation review | The source itself may be incorrect |
| Agent completed an action | Final application state plus the trace | A convincing completion message is insufficient |
| Explanation is useful | A clear rubric with calibrated human or model grading | Grader bias and disagreement remain measurable |
OpenAI's evaluation guide provides a provider-specific workflow for specifying test criteria and checking outputs. The underlying contract is general: record what happened and apply checks that can distinguish a valid result from a convincing claim.
Evaluate the grader too
A model judge extends what can be assessed automatically, but its score is another system output. G-Eval studied rubric-based evaluation on summarization and dialogue; LLM-as-a-Judge examined agreement and biases including position and verbosity. Those findings do not validate every judge on every task.
AgentRewardBench compares web-agent evaluators against expert-reviewed trajectories, illustrating why final-state evidence and human calibration matter. Check your grader against labeled examples and disagreements. Retain raw evidence so a surprising score can be investigated.
A useful report states the claim, suite and corpus versions, denominator, repeated-trial policy, grader, important slices, critical failures, and operational costs. It should make the next decision explainable, not merely produce a larger number.
Sources
Sources and further reading
- 01Holistic Evaluation of Language ModelsLiang et al. · research · published Nov 16, 2022 · source checked Oct 5, 2026
A primary framework connecting scenarios, adaptations, metrics, and transparent raw results.
- 02Demystifying evals for AI agentsAnthropic · guide · source checked Oct 5, 2026
A practical framework for tasks, trials, graders, transcripts, outcomes, and agent evaluation design.
- 03AgentRewardBenchLu et al. · research · published Apr 11, 2025 · source checked Oct 5, 2026
An expert-labelled study of automatic graders for web-agent trajectories, side effects, and repetitive behaviour.
- 04G-Eval: NLG Evaluation using GPT-4 with Better Human AlignmentLiu et al. · research · published Mar 29, 2023 · source checked Oct 5, 2026
A primary model-based evaluation method and early evidence of both stronger human correlation and possible preference for model-written text.
- 05Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaZheng et al. · research · published Jun 9, 2023 · source checked Oct 5, 2026
A primary study comparing model judges with expert and crowd preferences and documenting judge limitations.
- 06Working with evalsOpenAI · documentation · source checked Oct 5, 2026
Current first-party documentation connecting task definitions, test inputs, graders, result analysis, and iterative improvement while documenting the platform transition.
