Awesome Testing

Models · reviewed · reviewed Oct 5, 2026 · 4 min

What is context engineering?

Context engineering is the design of the complete, temporary input available to a model: instructions, examples, history, retrieved evidence, tool definitions, state, and observations. The goal is sufficient, relevant, trustworthy context—not the largest possible prompt.

Build the input for this step

A coding assistant is asked to fix a CSV export. The failing example is small: “Paris, France” must remain one quoted cell, rather than becoming two columns. The assistant needs the task, the API constraint, the failing example, the relevant code, and an appropriate way to check the repair.

It probably does not need a conversation about dashboard colors or the complete output of unrelated CI jobs. Context engineering is the decision about what the model should see for this call. After reading this article, you should be able to identify a missing input fact, distinguish it from a capacity problem, and explain who constructs the next call.

Anthropic's context engineering guide describes curation as an ongoing part of an agent loop. Prompt wording matters, but the complete input also includes selected history, files, retrieved evidence, tool definitions, and observations from earlier actions.

Enough room is not enough information

A context window sets a capacity limit. It does not certify relevance, authority, or task coverage. Our CSV request can fit comfortably while omitting the test that explains the bug. Adding thousands of unrelated tokens can consume room without supplying that missing example.

Use the existing selection experiment below to compare sufficient, incomplete, and overloaded inputs. The right-hand packet shows the authored content behind each selection. Remove repository rules to lose the API constraint, or the failing test to lose the concrete regression case.

Curate one model call

The task is to repair a CSV quoting failure. Choose what the model sees, then inspect the exact information you have removed or added.

2,750/4,096 toy tokens · bounded working set

Required input facts are present.

Assembled input packet

  1. Current goal

    Fix CSV export so a comma inside a field stays inside one quoted cell.

  2. Repository rules

    Keep the public API unchanged. Reuse the existing download helper.

  3. Failing test

    Input: ["Paris, France"]. Expected: one quoted CSV cell. Actual: two cells.

  4. Target files

    reports/export.ts joins raw fields with commas. export.test.ts contains the regression case.

  5. Relevant tools

    read_file, edit_file, and a focused export test command.

The counts and quality labels are fixed teaching data. The packet shows authored excerpts, not their tokenized full text. Coverage checks use this task’s declared requirements; they do not predict model success. No model is called. The 4,096-token teaching budget covers input only; real requests also reserve space for output.

Read the result as an input check

The “missing evidence” preset is smaller than the sufficient packet. It still lacks the failing example. The overloaded packet includes every required item, but exceeds the teaching budget. These conditions require different repairs: retrieve the missing evidence in the first case; select or compress material in the second.

“Required input facts are present” means only that the declared checklist is satisfied. It does not establish that a model will find the correct patch. Model behavior must be observed separately. Lost in the Middle found that evidence position affected performance in its studied long-context tasks. RULER tested several kinds of long-context work and distinguished advertised capacity from effective task performance. Neither result gives this local token counter a model-quality score.

The harness owns the next packet

The application or harness selects and labels the input before calling the model. The model proposes a response or tool call. If a tool runs, the harness receives its result and decides how that observation enters the following call.

For the CSV repair, a focused test failure should replace uncertainty with a specific observation. The next packet should preserve the current goal and constraints, include the relevant result, and omit material no longer needed. Persistent files and application state survive outside this temporary packet; the model does not automatically receive all of them.

flowchart LR
  G[Goal + constraints] --> C[Harness builds context]
  E[Selected evidence + tools] --> C
  C --> M[Model proposes next action]
  M --> T[Harness checks and runs tool]
  T --> O[Result + updated state]
  O --> C

Select, label, and preserve what matters

Select the smallest useful evidence set for the current action. Label where each fact came from and which instruction has authority. Compress repeated history when it helps, while retaining decisions, unresolved questions, failures, and exact references that a summary could lose.

A repository file or retrieved page can contain text that looks like an instruction. Its appearance in context does not give it the authority to change the user's goal or permissions. Keep trusted constraints separate from untrusted evidence and enforce consequential limits in the application.

Persist important state outside the prompt, then retrieve the relevant part when needed. A summary is a lossy view, not the canonical record. Splitting work across agents can isolate useful contexts, but introduces its own coordination and shared-budget obligations.

Diagnose the missing resource

Before adding more tokens, inspect the actual input sent to the model. Was the important fact selected? Was it current and authorized? Was its source clear? Did a tool result survive into the next call? Did compression erase a constraint? Measure tokens with the actual tokenizer and reserve output space according to the chosen API; official model guidance is provider-specific, not a universal budget.

Then change one part and evaluate the task again. Context engineering succeeds when the next action has the evidence and constraints it needs—not when the packet is as large as possible.

Sources and further reading

  1. 01
    Effective context engineering for AI agentsAnthropic · guide · published Sep 29, 2025 · source checked Oct 5, 2026

    A first-party systems view of selecting high-signal context, just-in-time retrieval, compact tools, compaction, persistent notes, and context isolation.

  2. 02
    Model guidanceOpenAI · documentation · source checked Oct 5, 2026

    Current first-party guidance for outcome-focused prompts, explicit constraints, approval boundaries, tool descriptions, success criteria, and evaluation against representative tasks.

  3. 03
    Lost in the Middle: How Language Models Use Long ContextsLiu et al. · research · published Jul 6, 2023 · source checked Oct 5, 2026

    A primary evaluation showing that access to a long input does not guarantee robust use of information at every position.

  4. 04
    RULER: What's the Real Context Size of Your Long-Context Language Models?Hsieh et al. · research · published Apr 9, 2024 · source checked Oct 5, 2026

    A configurable benchmark extending simple retrieval tests with multi-hop tracing and aggregation across long contexts.