Retrieval-Augmented Generation (RAG): Best Practices
Retrieval-augmented generation (RAG) connects a language model to an external knowledge source so answers can use evidence that is current, private, or too specific to be reliably recalled from model weights. For developers and platform teams building document assistants, support search, and internal knowledge tools, RAG is a common architecture, but adding a vector database alone does not make a useful system. This explainer follows the full path from source documents through retrieval and answer generation to evaluation.
What Is Retrieval-Augmented Generation?
RAG retrieves relevant passages for a user query and supplies them to a language model as context for generating an answer. The source collection may be manuals, product records, policies, code, or other material that changes independently of model training. The model still generates the response; retrieval provides evidence rather than changing the model’s weights.
The basic sequence has two phases. During indexing, a system reads source material, divides it into passages, attaches metadata, creates embeddings or searchable text indexes, and stores the results. During querying, it searches the index, selects and arranges useful passages, and sends them with the question to a model. The Microsoft Learn overview of RAG in Azure AI Search describes this retrieval-and-generation flow and distinguishes a search-backed RAG application from simpler prompt-only approaches.
The Problem RAG Solves
A language model’s built-in knowledge may be stale, incomplete, or unrelated to an organization’s private data. Asking it to answer from memory can produce plausible but unsupported details. Fine-tuning can change a model’s behavior, but it is not a reliable replacement for searching a frequently updated corpus: model updates are costly, and a generated answer does not automatically identify the source behind a claim.
RAG makes the source collection separately searchable and lets the application present evidence alongside the question. That helps with freshness, traceability, and domain-specific material. It does not guarantee correctness. A retriever can return the wrong passage, a prompt can omit an important qualification, and a model can still misread or ignore relevant evidence. Treat RAG as a system for improving grounded answers, not a switch that eliminates hallucinations.
How a RAG System Works
A practical pipeline usually includes these stages:
- Ingest and normalize: Read supported formats, remove boilerplate, preserve headings, and retain a stable source identifier and revision.
- Split and index: Divide long sources into passages that retain enough context to be meaningful. Store text, metadata, and one or more search representations.
- Retrieve candidates: Convert the query into a search request and find passages that may answer it. Apply access-control and scope filters as part of this operation.
- Rerank and assemble context: Score a manageable candidate set more carefully, remove duplicates, and fit the strongest evidence into the model’s context window.
- Generate and return evidence: Ask the model to answer from the supplied material, then show citations that resolve to the underlying sources.
- Measure and update: Evaluate retrieval and answers separately; re-index changed or deleted material and monitor production behavior.
The best retrieval method depends on how people ask questions and what the corpus contains:
| Retrieval method | Useful when | Main limitation |
|---|---|---|
| Keyword or sparse search | Queries contain exact names, error codes, or phrases | Can miss relevant text expressed with different wording |
| Dense vector search | Queries describe a concept without matching the source vocabulary | Similarity is not proof of relevance or factual support |
| Hybrid search | Users mix exact terms with natural-language questions | Requires combining result lists and tuning additional components |
| Reranking | Initial search finds useful candidates but their order is weak | Adds latency and cost; cannot recover passages that were never retrieved |
For hybrid systems, combining vector and sparse retrieval explains candidate fusion and how to test its effect. The LlamaIndex RAG documentation also breaks the application into indexing, retrieval, and response-synthesis stages, which is useful when diagnosing which component is responsible for a poor answer.
Components and Design Choices
Source preparation and chunking. Extract text without losing tables, headings, or references that make a passage intelligible. Split by document structure first, then use a token limit suitable for the embedding model and downstream prompt. Very large chunks dilute the relevant detail; very small chunks can separate a claim from its definitions or exceptions. Overlap can preserve continuity across boundaries, but excessive overlap duplicates evidence and consumes storage and context.
Metadata and permissions. Store fields such as source ID, title, section, revision, and access scope with each passage. Apply tenant and user permissions before content reaches the model, not after retrieval. Re-index changed sources and remove obsolete passages so that a correct answer cannot be grounded in a revoked or stale document.
Retrieval and context assembly. Use the same embedding model and compatible normalization for document and query vectors. Add keyword search when exact identifiers matter, and consider reranking only if evaluation shows it improves the results enough to justify extra work. Retrieve more candidates than will fit in the prompt, then deduplicate and choose the passages with the strongest evidence. Keep source labels so the application can render citations that point to an actual document location.
Prompt and answer behavior. Tell the model to use only supplied evidence for factual claims, distinguish missing evidence from a negative answer, and state when it cannot answer. Ask for references in a format the application can validate against retrieved source IDs. A prompt cannot repair poor retrieval, but explicit rules reduce the chance that the model silently fills gaps with unsupported details. See prompt engineering for language models for more on structuring model instructions.
Evaluation. Keep a representative set of questions, relevant passages, and expected answer properties. Measure retrieval independently from generation: recall at a chosen rank checks whether useful evidence was found, while answer-level checks examine relevance and support. The Ragas metric documentation describes measures such as context precision, context recall, faithfulness, and answer relevancy. No single score captures correctness for every corpus, so review failures as well as averages. The NIST Generative AI Profile provides a risk-management framework for identifying and measuring generative AI risks; it is useful context for deciding which failures need human review or stronger controls.
Real-World Use Cases
An internal support assistant can retrieve approved procedures and cite the exact section used in its response. A product documentation search can combine semantic matches with exact version numbers or error codes. A legal or policy tool can filter sources by jurisdiction, date, and user permissions before generating a summary. In each case, RAG is most useful when answers need to reflect a bounded, inspectable source collection.
For high-impact decisions, retrieval should support a human reviewer rather than make the decision by itself. A retrieved passage may be outdated, incomplete, or outside the user’s authority. Record which sources informed an answer and make the limits of the system visible.
Getting Started: Build a Small Retrieval Prototype
The following Python example builds an in-memory semantic index from three short passages. It demonstrates embedding, nearest-neighbor search, and source labels; it is not a production ingestion or authorization system.
Create an environment and install the dependencies:
Windows PowerShell
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install sentence-transformers faiss-cpu
macOS or Linux
python3 -m venv .venv
source .venv/bin/activate
python -m pip install sentence-transformers faiss-cpu
Check that the active environment can import both packages:
python -c "import faiss, sentence_transformers; print('RAG dependencies are available')"
Save this as rag_retrieval.py and run it with python rag_retrieval.py:
import faiss
import numpy as np
from sentence_transformers import SentenceTransformer
chunks = [
{
"source": "handbook/retention#policy",
"text": "Customer support logs are retained for 30 days unless a legal hold applies.",
},
{
"source": "handbook/access#requests",
"text": "Access requests must be approved by the resource owner before permissions change.",
},
{
"source": "runbook/backups#schedule",
"text": "Database backups run every six hours and are checked by a daily restore test.",
},
]
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
texts = [chunk["text"] for chunk in chunks]
vectors = model.encode(
texts,
normalize_embeddings=True,
convert_to_numpy=True,
).astype("float32")
index = faiss.IndexFlatIP(vectors.shape[1])
index.add(vectors)
question = "How long are customer support logs kept?"
query_vector = model.encode(
[question],
normalize_embeddings=True,
convert_to_numpy=True,
).astype("float32")
scores, positions = index.search(query_vector, k=min(2, len(chunks)))
for score, position in zip(scores[0], positions[0]):
if position >= 0:
chunk = chunks[position]
print(f"{score:.3f} {chunk['source']}: {chunk['text']}")
The example normalizes vectors and uses inner product so ranking corresponds to cosine similarity. Check that the retention passage appears first, then inspect the second result: similarity rankings can include plausible but irrelevant text. A real system needs document parsing, token-aware chunking, persistent storage, source versioning, permission filters, and deletion handling. Add hybrid search or reranking only after a test set shows where semantic retrieval falls short. Feed retrieved passages and their source IDs to a model, then verify that returned citations refer to those IDs.
If imports fail, confirm which interpreter is active with python -c "import sys; print(sys.executable)", then check the installed packages with python -m pip show sentence-transformers faiss-cpu. Using python -m pip ties installation and diagnostics to the same interpreter that runs the example.
In production, log latency and retrieval outcomes without exposing private prompt content unnecessarily. Track indexing failures, stale sources, empty retrievals, and changes in answer quality. OpenTelemetry for LLM production observability covers tracing model operations and related measurements; model-side caching and inference costs are discussed in LLM inference optimization and KV caching.
Common Misconceptions
- “A vector database is a RAG system.” It stores and retrieves data, but ingestion, permissions, prompt construction, generation, citations, and evaluation remain application responsibilities.
- “More retrieved passages always improve the answer.” Low-quality or duplicated context can distract the model and use tokens that could have held better evidence. Evaluate the number and ranking of passages.
- “A citation proves the answer is correct.” A citation can point to a real source that does not support the associated claim. Validate source-to-claim support and show users enough context to check it.
- “RAG prevents hallucinations.” It gives the model external evidence, but retrieval can fail and generation can still go beyond that evidence. Test abstention and unsupported-answer behavior.
- “One chunk size works for every corpus.” Useful boundaries depend on document structure, query style, embedding model, and context budget. Compare alternatives on representative questions.
Related Articles
- Hybrid Search for RAG: Combining Vector and Sparse Retrieval explains ways to retrieve both semantic matches and exact terms.
- Vector Databases for RAG: How Retrieval Works covers embeddings, indexes, and passage retrieval.
- Prompt Engineering for LLMs explores how to guide a model’s use of retrieved evidence.
- OpenTelemetry for LLM Production Observability explains tracing and production measurements.
Changelog
- Initial publication.
Last updated: September 29

