Testing and evaluation · reviewed · reviewed Oct 6, 2026 · 7 min
How should an AI system handle uncertainty?
Treat uncertainty as a measured property of a specific claim and operating condition: combine calibrated task evidence, source and tool checks, repeated behaviour, and risk-aware policy, then let the harness verify, clarify, escalate, or abstain when evidence is insufficient.
Separate the signals
A language model predicts a distribution over the next token given its context. A high probability for one token means the model strongly preferred that continuation at that position. It does not directly mean that the completed answer is factually correct, that cited evidence supports it, or that an external operation will succeed.
Several things called “confidence” need separate names:
- token probability: preference among possible next token IDs;
- sequence score: an aggregate over token probabilities, strongly affected by wording and length;
- semantic consistency: whether repeated samples converge on the same meaning;
- self-report: text such as “I am 90% confident,” generated through the same language mechanism;
- grader score: a separate instrument applying a rubric to visible evidence;
- empirical calibration: how often a defined prediction was correct among comparable cases assigned a given score;
- system confidence: an application decision combining evidence, dependencies, policy, and cost of error.
Fluency is not on the list. A polished answer can be unsupported, while a hesitant style can accompany a correct one. The harness should not infer reliability from tone.
Ask for the fact the policy does not contain
A fictional assistant retrieves an employee reimbursement policy, but the user asks about a contractor. The passage is clear; its scope does not answer this case.
| Missing or conflicting condition | Useful next step |
|---|---|
| Employment type unknown | Ask which type applies |
| Contractor policy absent | Retrieve it or escalate to its owner |
| Two current sources disagree | Preserve the conflict and seek adjudication |
| Submitted write lost its response | Reconcile the operation before another effect |
“I'm 90% confident” supplies none of the missing evidence. State the concrete gap and a next step. These are authored policy decisions, not calibrated percentages or a claim that a model recognizes every unsupported answer.
Calibration is a population property
Suppose a classifier emits 0.7 for many comparable cases. It is calibrated at that level if roughly 70% of those cases are correct under the declared label and data distribution. Calibration does not say which individual case is correct, and it does not transfer automatically to a new language, user group, task, model revision, or traffic shift.
flowchart LR
S[Raw signal] --> C[Calibration on held-out labels]
C --> P[Estimated risk for a declared slice]
P --> D{Policy threshold}
D -->|low risk| A[Answer or act]
D -->|unclear| V[Retrieve, verify, or ask]
D -->|high risk| H[Abstain or escalate]
For fixed-label tasks, reliability diagrams compare assigned score bands with observed correctness. Proper scoring rules such as log loss or the Brier score reward useful probabilities and penalize confident errors. A single expected-calibration-error number can hide small severe slices, so keep the underlying counts and plots.
Choose a threshold without hiding its denominator
The experiment contains thirty authored cases: ten each with scores 0.60, 0.80, and 0.95. In the matched population, 6, 8, and 9 cases in those respective groups are correct. In the shifted population, the scores stay the same while the correct counts become 3, 4, and 5.
These scores represent an illustrative instrument's estimate of a defined correctness event, not token probabilities or an LLM's self-report. The labels are independent authored outcomes. The scenarios provide no evidence about a real model's calibration.
Raise the threshold and inspect accepted errors, coverage, and accuracy among accepted cases. Then switch population while keeping the same threshold. The reliability plot always uses all thirty cases, so filtering accepted cases cannot make the plotted population look better.
Explore the mechanism
Trade coverage against errors on a population
| Score | Cases | Correct | Observed rate |
|---|---|---|---|
| 0.60 | 10 | 6 | 60.0% |
| 0.80 | 10 | 8 | 80.0% |
| 0.95 | 10 | 9 | 90.0% |
- Accepted cases
- 30 / 30
- Coverage
- 100.0%
- Accepted errors
- 7
- Accuracy among accepted
- 76.7%
0 cases abstained. The threshold changes which cases are accepted; it does not recalibrate their scores.
Inspect all cases and population scores
All-population Brier score: 0.164. Fixed-bin ECE: 0.017. Neither depends on the acceptance threshold.
| Case | Score | Correct | Accepted |
|---|---|---|---|
| 1 | 0.60 | Yes | Yes |
| 2 | 0.60 | Yes | Yes |
| 3 | 0.60 | Yes | Yes |
| 4 | 0.60 | Yes | Yes |
| 5 | 0.60 | Yes | Yes |
| 6 | 0.60 | Yes | Yes |
| 7 | 0.60 | No | Yes |
| 8 | 0.60 | No | Yes |
| 9 | 0.60 | No | Yes |
| 10 | 0.60 | No | Yes |
| 11 | 0.80 | Yes | Yes |
| 12 | 0.80 | Yes | Yes |
| 13 | 0.80 | Yes | Yes |
| 14 | 0.80 | Yes | Yes |
| 15 | 0.80 | Yes | Yes |
| 16 | 0.80 | Yes | Yes |
| 17 | 0.80 | Yes | Yes |
| 18 | 0.80 | Yes | Yes |
| 19 | 0.80 | No | Yes |
| 20 | 0.80 | No | Yes |
| 21 | 0.95 | Yes | Yes |
| 22 | 0.95 | Yes | Yes |
| 23 | 0.95 | Yes | Yes |
| 24 | 0.95 | Yes | Yes |
| 25 | 0.95 | Yes | Yes |
| 26 | 0.95 | Yes | Yes |
| 27 | 0.95 | Yes | Yes |
| 28 | 0.95 | Yes | Yes |
| 29 | 0.95 | Yes | Yes |
| 30 | 0.95 | No | Yes |
Thirty cases, scores, and correctness labels are authored teaching data. Rates, reliability bins, Brier score, and fixed-bin calibration error are computed locally. No model, benchmark, or calibrator is fitted. Population changes preserve the selected threshold so the same policy can be compared; these scenarios do not measure real distribution drift.
At threshold 0.95, the matched fixture accepts ten cases and gets nine right. The shifted fixture accepts the same score group but gets only five right. A high score's interpretation depends on the population and the quality of the scoring instrument.
At threshold 1, no case qualifies. Accuracy among accepted cases is then undefined: the denominator is zero. Reporting 100% would disguise a policy that answered nothing. Coverage and accepted errors must remain visible beside conditional accuracy.
The threshold chooses which cases to accept; it does not recalibrate their scores. The optional inspector calculates a Brier score and a count-weighted calibration gap for the three declared score groups on the whole fixture. Those population quantities remain fixed when only the threshold changes.
In a real system, estimate and validate calibration using appropriately separated labelled data and relevant slices. Thirty invented examples are enough to reveal the arithmetic and its boundaries, not to certify a deployment.
Agreement still needs independent evidence
Free-form generation is harder. Many token sequences express the same answer, and longer sequences naturally accumulate lower probability. Sampling several answers and grouping them by meaning can reveal semantic disagreement; research on semantic uncertainty shows this can detect some confabulations. It still does not detect every error—several samples can agree on the same false claim.
Locate what is uncertain
An application may face different uncertainties:
- the request is ambiguous or missing constraints;
- relevant knowledge is absent, stale, or outside the model;
- retrieved sources conflict or do not support the claim;
- a tool result is missing, malformed, or from an unknown state;
- the model produces several incompatible interpretations;
- the request differs from the calibrated evaluation population;
- a grader is unavailable or disagrees with deterministic evidence;
- the real-world effect is irreversible or has an unknown outcome.
These conditions call for different responses. Ask a clarifying question for ambiguity. Retrieve current evidence for missing knowledge. Re-run an idempotent read when a tool result is unavailable. Require approval for consequential effects. Escalate a high-risk domain decision. Preserve “unknown outcome” after an ambiguous write rather than retrying blindly.
Abstention belongs to the system contract
Abstention means deliberately not making a claim or effect under specified conditions. It can be a refusal, a request for missing information, a limited answer with named uncertainty, a handoff to a person, or a safe no-op. Define which form applies and what the user can do next.
Thresholds should reflect consequences. A slightly uncertain movie recommendation and an uncertain medication dose cannot share one policy. Compare at least three rates by slice: correct actions, harmful or incorrect actions, and abstentions or escalations. Increasing abstention can raise accuracy among answered cases while making the product useless; reducing it can improve completion while increasing severe false actions.
Evaluation metrics influence behaviour. A benchmark that scores only exact correct answers and treats abstention exactly like a wrong answer rewards guessing whenever there is any chance of success. Score confident errors, useful abstentions, and unnecessary refusals in a way that matches the actual decision cost.
Do not delegate the final decision to a prompt saying “answer only if confident.” The model can contribute a signal, but the harness owns evidence checks, calibrated thresholds, permissions, routing, and the allowed terminal states.
Communicate without false precision
Show users what matters: the evidence used, its date and scope, unresolved conflict, assumptions, actions taken, and a clear next step. A naked “87% confidence” suggests a universal frequency unless it was calibrated for that exact kind of claim and population.
Prefer statements such as “The retrieved policy does not cover contractors, so I need the employment type” or “The write timed out after submission; its outcome is unknown, so I will check the transaction before retrying.” These expose the missing evidence and safe response rather than imitating a feeling.
Confidence is not permission
Uncertainty is not one scalar stored inside a language model. Entropy is not the probability that a sentence is false. Repeated agreement is not independent verification. A model saying “I don't know” is not automatically calibrated, and a citation is not evidence until the cited material is retrieved and shown to support the claim.
Calibration is also not permanent certification. It depends on labels, scoring rules, slices, thresholds, and the system version. Change the model, prompt, tools, context construction, or population and the evidence must be checked again.
Evaluate abstention as a decision policy
Build labelled cases spanning easy, hard, ambiguous, unanswerable, conflicting, stale, out-of-distribution, and adversarial inputs. Include missing tools, malformed results, insufficient permissions, and unknown write outcomes. Define acceptable answer, verification, clarification, abstention, and escalation states before running the system.
For numeric scores, calculate reliability by score band and important slice; report sample counts, Brier or log loss, false-confidence events, and coverage-versus-risk curves. Choose thresholds on protected evidence, then evaluate once on a separate held-out set. Recheck after material version or traffic changes.
Probe superficial invariance. Paraphrase questions, alter tone, reorder irrelevant details, translate equivalent cases, and repeat generations. Confidence should not change merely because an answer is longer or more assertive. Insert false citations and confident model-generated rationales to confirm that grounding checks remain independent.
Finally, test the policy outcomes. Every high-impact confident error deserves review. So does every unnecessary refusal in an essential workflow. Simulate unavailable uncertainty services and malformed scores; the application should fall back to a declared safe state rather than treating missing confidence as permission to proceed.
Sources
Sources and further reading
- 01On Calibration of Modern Neural NetworksGuo et al. · research · published Jun 14, 2017 · source checked Oct 6, 2026
A primary empirical study defining confidence calibration and evaluating reliability diagrams, expected calibration error, and post-hoc temperature scaling.
- 02Detecting hallucinations in large language models using semantic entropyFarquhar et al. · research · published Jun 19, 2024 · source checked Oct 6, 2026
A primary study that groups sampled generations by meaning and evaluates semantic uncertainty for detecting a subset of unsupported generations.
- 03Evaluating large language models for accuracy incentivizes hallucinationsKalai et al. · research · published Apr 22, 2026 · source checked Oct 6, 2026
A primary analysis of why accuracy-only evaluation can reward guessing and why evaluation should distinguish errors from appropriate abstention.
- 04Artificial Intelligence Risk Management Framework 1.0NIST · standard · published Jan 26, 2023 · source checked Oct 6, 2026
A system-lifecycle framework for mapping context, measuring trustworthiness, and managing AI risk.
