Production systems · reviewed · reviewed Oct 5, 2026 · 4 min
What is retrieval-augmented generation?
Retrieval-augmented generation selects authorized evidence from an external corpus at request time, places the most useful passages into model context, and generates or checks an answer against that evidence.
Give the model evidence it did not learn
An employee asks: “How many days do I have to submit an expense claim?” The company's current policy says 30 days. Last year's policy said 14. A model's fluent recollection cannot tell us which policy this employee may access or which version applies today.
Retrieval-augmented generation, or RAG, brings external evidence into the current model call. After reading this article, you should be able to follow that evidence from a document to a claim—and explain what changes when the document disappears, becomes obsolete, or is no longer authorized.
The original RAG paper combined a trained generator with a retrieved document index. Today's RAG applications use many variations of that pattern. At request time, retrieving a passage normally changes the model's input, while its learned parameters stay fixed. Updating a corpus is distinct from retraining a model.

An illustrative metaphor: the archive stays outside the model; selected fragments enter one reading tray. The image does not show a literal model architecture.
From a question to an evidence packet
Before a question arrives, the application parses documents, splits them into passages, and indexes them for search. Each passage needs provenance: its document, version, location, and access rules. Splitting the deadline from its exception could make an apparently relevant fragment misleading.
At request time, search finds candidates. The application filters for access and freshness, ranks relevance, and assembles a bounded context packet. The model then receives the question, instructions, and selected evidence. Generating an answer and checking its claims are further stages; finding a document does not complete them.
Try the local example below. Remove the current policy first. Then allow archived versions. Finally, deny policy access or switch to an answer that invents a deadline. Watch the selected packet and the support verdict separately.
Follow the evidence, then inspect the claim
One question, four fictional passages. Change the corpus or access rules and watch what actually reaches the answer stage.
How many days do I have to submit an expense claim?
- RetrieveApplication · match + filter
- Assemble contextApplication · keep provenance
- Answer + checkTemplate · inspect support
Candidate passages
- Expense policypolicy-2026 · v2026
Submit an expense claim within 30 days of purchase.
selected · 4/4 keyword matches - Archived expense policypolicy-2025 · v2025
Submit an expense claim within 14 days of purchase.
obsolete version - Receipt FAQreceipt-faq · v2026
Every expense claim needs an itemized receipt.
selected · 2/4 keyword matches - Office opening hoursoffice-note · v2026
The office is open from nine until five.
no matching terms
What the answer stage receives
Submit an expense claim within 30 days of purchase.
policy-2026 · v2026
Every expense claim needs an itemized receipt.
receipt-faq · v2026
Submit your expense claim within 30 days.
Attached citation: policy-2026Supported by the current policyWhat this demo computes
The search counts exact matches for expense, claim, submit, and days; ties keep corpus order. Availability, access, and version rules filter candidates before the top two enter context. Claim checking compares the deadline with structured facts attached to these authored passages. Real retrieval and semantic support checks are more complex.
All documents, access rules, and answer templates are teaching fixtures. No LLM or external search is called. A supported claim can still be wrong if its source is wrong.
A citation answers only one of the questions
With the current policy present, the authored answer can cite a passage stating 30 days. With only the archive available, 14 days is supported by that old passage but fails the requirement for a current answer. With no deadline evidence, the grounded template abstains. The invented seven-day answer remains unsupported even when it has a citation.
Those are different failures: missing evidence, obsolete evidence, and a claim that does not follow from evidence. A real model might also ignore a correct passage. Even a faithful answer can be false if the source itself is wrong. Keep provenance, freshness, and claim support visible rather than collapsing them into “the system found something.”
Retrieved documents are also data, not higher-authority instructions. An embedded command cannot grant access or authorize a tool effect. Access checks must happen before restricted material reaches model input; deleting it from the final answer cannot undo that exposure. These are application controls within the broader lifecycle approach described by NIST's AI Risk Management Framework.
Search is a design choice
The demo's exact keyword matching is deliberately small. Lexical retrieval is useful for identifiers and exact phrases. Dense retrieval compares learned vector representations and can find related wording. Hybrid retrieval combines signals; reranking revisits a smaller candidate set with a more expensive relevance method. BEIR evaluates retrieval across diverse tasks and shows why a method should be tested in its intended setting.
There is no requirement to use a vector database. A direct database lookup may be the strongest route to an exact deadline. Use RAG when selecting and interpreting external evidence is useful; choose a simpler authoritative lookup when it answers the product question better.
Inspect the first stage that lost the contract
Save the question, corpus version, authorized evidence, selected passages, generated claims, and citations. If the correct passage was never selected, improve retrieval or ingestion. If it was selected but omitted from context, inspect assembly. If the answer misstates it, inspect generation and claim checking.
RAGAs separates retrieval focus, faithful evidence use, and answer quality. Automated scores still need validation against the cases that matter to your product. Include revoked access, stale versions, missing documents, conflicting passages, and expected abstentions. The useful result explains where evidence was lost and whether the final claim remained justified.
Sources
Sources and further reading
- 01Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksLewis et al. · research · published May 22, 2020 · source checked Oct 5, 2026
Primary RAG formulation combining parametric generation with retrieved non-parametric evidence.
- 02BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval ModelsThakur et al. · research · published Apr 17, 2021 · source checked Oct 6, 2026
A heterogeneous retrieval benchmark demonstrating why retriever quality and out-of-domain behaviour need explicit evaluation.
- 03RAGAS: Automated Evaluation of Retrieval Augmented GenerationEs et al. · research · published Mar 1, 2024 · source checked Oct 5, 2026
A component-level evaluation framework separating context relevance, answer faithfulness, and answer quality.
- 04Artificial Intelligence Risk Management Framework 1.0NIST · standard · published Jan 26, 2023 · source checked Oct 6, 2026
A system-lifecycle framework for mapping context, measuring trustworthiness, and managing AI risk.
