Awesome Testing

Models · reviewed · reviewed Oct 5, 2026 · 4 min

How do reasoning models use extra computation?

An LLM system can spend extra inference work on intermediate tokens, candidate answers, revision, or search, then use a checking and selection rule to choose an answer. Learned parameters normally stay fixed; better results depend on the task, proposals, verifier, and budget.

More work before the answer

A language model still generates tokens from its current context. Additional intermediate tokens create further forward computations and can supply context for later decisions. A system can also generate several candidate answers, revise a proposal, or search a space of possible solutions before selecting what to return.

These are different uses of test-time compute: work spent while answering the current request. A longer response is only one possible use. Repeated attempts, checks, and discarded candidates can consume work without appearing in the final answer.

Training taught the model and any learned verifier their behaviour earlier. During an ordinary inference run, those parameters remain fixed. The current run changes its context, candidate set, and search state; it does not learn a new checkpoint just because it spends longer on a problem.

Propose, check, select

Separate the responsibilities before judging the result:

flowchart LR
  Q[Question and budget] --> P[Generate a candidate]
  P --> V[Check the declared requirements]
  V --> D{Accept?}
  D -->|yes| A[Selected answer]
  D -->|no, work remains| P
  D -->|no, budget exhausted| U[Unresolved result]

A proposer supplies possibilities. A verifier measures whether a possibility meets a criterion. A selection rule chooses among available evidence, and a stop rule bounds further work. These roles can live inside a model system, a surrounding application, or a combination. The diagram is a teaching workflow, not the hidden architecture of every reasoning model.

Checking an exact arithmetic contract is different from asking a learned judge whether an explanation sounds convincing. A good candidate is useful only if it can be found and recognized.

Spend a budget on one small puzzle

Use the numbers 4, 7, and 9 exactly once, with addition or subtraction, to reach 20. A result of 20 is necessary, but it is not the whole contract: 9 + 9 + 2 reaches 20 using the wrong inputs.

The experiment exposes a fixed sequence of authored candidate expressions. Increase the candidate allowance and compare a verifier that checks the complete contract with one that checks only the numeric result. Arithmetic, number usage, selection, and inspected-candidate counts are calculated locally.

Explore the mechanism

Spend work on proposals and checks

Reach 20 using 4, 7, 9 exactly once, with + or −.

Inspected candidates
1
Unused allowance
0
  1. 9 + 7 − 4 = 12
    Numbers match: yes; result matches: no.
    Verifier: reject.
Selected answer: None; allowance exhausted

Complete task contract: unresolved

Four authored expressions in a fixed order. Arithmetic, complete-contract validity, selection, and stop counts are computed. No LLM, learned judge, sampling, or real timing runs; this is not a hidden reasoning trace.

With the complete verifier, two candidates leave the task unresolved; the third supplies a valid solution. The result-only verifier accepts the second candidate even though it violates the input constraint. Additional work helped only when the selection rule recognized the right property.

The sequence is deliberately repeatable. It does not model the probability that an LLM discovers these expressions, and the counters are not tokens, seconds, FLOPs, or benchmark scores.

There is more than one way to spend compute

Generate several paths and aggregate. Self-consistency research samples different solution paths and combines their final answers. Agreement can help under the studied conditions, but several paths can share the same mistake. A vote is not independent factual evidence.

Revise a candidate. A later request or continuation can use an earlier proposal and feedback. Useful feedback can expose a missing constraint; a vague request to “think again” can also preserve or introduce an error.

Search with a verifier. Outcome checks assess a final answer; process checks assess intermediate steps. A learned process verifier can guide exploration, but it remains an imperfect measurement system. Deterministic checks are preferable where the requirement is exact.

These approaches have different costs and failure modes. Research on test-time scaling found that the useful allocation depends on the task's difficulty and the model's capabilities. It does not establish a universal budget that improves every request.

More computation cannot supply missing truth

A system can spend its entire budget elaborating a false premise, repeatedly proposing an unavailable solution, or optimizing a weak verifier. More tokens do not provide a current policy, a missing image region, a revoked permission, or an authoritative tool result.

Match the response to the gap. Retrieve missing evidence, use a calculator for exact arithmetic, ask about an ambiguous constraint, or stop with an unresolved result. A model's narrated explanation is another output to check; it cannot certify its own correctness.

For a practical comparison, freeze the task and acceptance rule, vary the compute policy, and report accepted outcomes beside total work and severe failures. A system that obtains more correct answers by multiplying cost may be useful, but that trade-off should remain visible.

What to record when comparing reasoning budgets

Record the model and verifier versions, prompt and evidence, candidate or revision policy, decoding settings, maximum work, stop reason, accepted answer, and the checks that support it. Distinguish generated candidates from candidates actually checked or selected.

Keep training changes separate from inference-budget changes. Repeat real-model trials when estimating reliability; use fixed candidate fixtures when proving selection and stopping contracts. Do not silently remove unresolved cases from the denominator.

Sources and further reading

  1. 01
    Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model ParametersSnell et al. · research · published Aug 6, 2024 · source checked Oct 5, 2026

    Primary study of proposer/verifier policies, revision, search, and task-dependent allocation of inference work; its results are not universal reasoning-budget guarantees.

  2. 02
    Self-Consistency Improves Chain of Thought Reasoning in Language ModelsWang et al. · research · published Mar 21, 2022 · source checked Oct 5, 2026

    Primary method for sampling diverse solution paths and aggregating final answers, distinct from learned verification and from proof that agreement establishes truth.

  3. 03
    Let's Verify Step by StepLightman et al. · research · published May 31, 2023 · source checked Oct 5, 2026

    Primary comparison of outcome and process supervision for learned math verifiers, with explicit limits on transfer beyond the studied domain.