TATECHATLAS
◎ English
Artificial intelligence

What RAG changes—and what it does not

How retrieval connects a language model to documents, and why a citation is not a guarantee.

On this page

RAG retrieves relevant material before a model writes an answer. It can make documents available at answer time without retraining the model. The quality still depends on what was found and how faithfully the answer uses it.

An example: answering a question about company policy

A colleague asks how many days they have to submit an expense claim. A useful RAG system finds the relevant policy, its current version and the paragraph about the deadline. The model then explains that paragraph and links to its source.

If retrieval finds only a policy for another country or an old version, the answer should say the evidence is insufficient. A confident sentence with a citation to the wrong policy is still a bad answer. This example shows why document scope matters as much as text similarity.

Follow the evidence path

A typical flow is: question, retrieval, selected passages, answer. Keep document titles and locations with each passage. If useful evidence is absent, return that limitation instead of filling the gap with a plausible claim.

Build the smallest useful retrieval flow

Start with a small set of trusted documents. Extract readable text, retain headings and split it into passages that preserve the rule and its exceptions. Store the document name, version, section and access restrictions with each passage. Retrieve a few relevant passages, then ask for an answer supported by them.

For a small collection, ordinary keyword search may be enough to establish whether the idea helps. Vector search can help with different wording; hybrid search combines lexical and vector results. Neither requires you to begin with a large infrastructure or a universal chunk size.

Diagnose retrieval separately

First check whether the relevant passage was retrieved. Then check whether it supports the answer. Short chunks can lose context; oversized chunks can dilute the relevant point. Evaluate chunking with representative questions, not a universal size.

Separate three different failure modes

When an answer is wrong, first inspect the retrieved passages. If the right information never appeared, investigate indexing, filters and retrieval. If it appeared but lost an exception, investigate chunk boundaries or selection. If the evidence was complete and the model contradicted it, investigate answer generation.

Prepare a modest evaluation set: a clear question, a rephrased question, a question requiring an exception and one with no answer in the documents. Record the expected passage and whether the answer should be refused. That makes changes comparable instead of judging a demo by one impressive response.

Keep permissions and costs under control

Filter documents by the user’s permissions before their text reaches the model. A warning in the prompt is not an access-control mechanism. Treat instructions inside retrieved documents as document content, not as authority to change the system’s rules.

Limit retrieved passages and output length, cache where appropriate and update the index when documents change. Measure both answer usefulness and spending. Before adding reranking or another model call, identify the specific failure it is intended to fix; extra stages also add latency and operating cost.

Things to check

  • Can the system find the needed passage?
  • Does the cited passage actually support the claim?
  • Are document versions and access rules respected?

Retrieval reduces some knowledge gaps but does not guarantee factual correctness. A source may itself be wrong or out of date.

Sources

  1. Retrieval-Augmented Generation, Lewis et al. ↗
  2. Microsoft Learn: RAG overview ↗
  3. Microsoft Learn: document chunking ↗
  4. Microsoft Learn: vector search ↗
Back to top ↑