A support engineer asks the bot: "What's the refund window for enterprise contracts?" The bot answers correctly in four seconds. No fine-tuning. No retraining. The model had never seen that policy document during pre-training. What just happened, mechanically?
That question is what retrieval-augmented generation (RAG) answers - and understanding the actual pipeline, not just the name, changes how you build and debug AI features. This post walks through each step with real numbers, then lands on one non-obvious problem most introductions skip entirely.
The five-step pipeline behind RAG retrieval
RAG connects large language models to external knowledge bases through a five-stage pipeline: ingestion, embedding, retrieval, augmentation, and generation - enabling accurate, domain-specific answers without retraining the model. Here is what each step actually does.
1. Ingestion and chunking. Your documents - PDFs, Notion pages, Confluence docs - get parsed into plain text and then sliced into chunks. Documents are split into chunks and encoded into dense vector embeddings, which are stored in a vector database that supports similarity search. Chunk size is not arbitrary. For most production RAG systems, recursive chunking with 400-800 token chunks and 20% overlap provides the best balance of performance and efficiency. Go too small and a chunk loses enough context to be useless; go too large and you drag in neighboring noise. If you retrieve 8 chunks averaging 700 tokens each, you're at 5,600 tokens before the model has generated anything.
2. Embedding. Raw data - documents or videos - are first converted into chunks, and each chunk is converted into a high-dimensional embedding vector using an embedding model and stored in a vector database. An embedding is just a list of numbers - typically 768 to 3,072 floating-point values - that encodes what a piece of text means, not just what words it contains. Each sentence becomes a vector - a point in high-dimensional space - where semantic closeness means directional similarity. Instead of matching exact words, embeddings capture meaning.
3. Retrieval. When the user queries the system, the query is converted to an embedding vector using the same embedding model and passed to the retriever. The vector database searches for embeddings close to the query embedding using some distance metric, and returns the relevant data chunks.
When millions of vectors exist, finding the nearest ones efficiently demands specialized data structures such as FAISS indexes. These indexes allow approximate nearest-neighbor searches with only a small trade-off in precision.
4. Augmentation. The retrieved data chunks and the user query are combined into a single prompt and passed to the LLM. The model never "learned" the content of your documents - it is reading them fresh, right now, in its context window. By merging two knowledge streams - the fixed, general knowledge embedded in the LLM and the flexible, domain-specific information augmented on demand - the system aligns the model with both established and emerging information.
5. Generation. The model produces a grounded answer, citing the material it was just handed. By grounding generation in retrieved evidence, RAG reduces hallucinations and improves factual accuracy without requiring model retraining.
Why the chunk boundary is where most RAG systems break
The piece of the pipeline that gets least attention is the one that matters most: how you split documents before indexing.
The document chunking process plays a central role in the performance of RAG pipelines. Incoherent document splits and inappropriate chunk sizes hinder retrieval efficiency and contextual accuracy. The classic failure mode is a chunk that contains the answer but lacks the context that makes the answer mean anything. Imagine splitting a legal policy document right before the sentence that names the effective date - the chunk with the rule no longer knows when the rule applies.
Anthropic's Contextual Retrieval technique addresses this directly. Contextual Retrieval prepends a short, chunk-specific explanation to each document chunk before it is embedded and indexed, so retrieval keeps the surrounding context that naive chunking strips away.
The improvement from doing this is measurable. Contextual Embeddings reduced the top-20-chunk retrieval failure rate by 35% (5.7% → 3.7%).
Combining Contextual Embeddings and Contextual BM25 reduced the top-20-chunk retrieval failure rate by 49% (5.7% → 2.9%).
Reranked Contextual Embedding and Contextual BM25 reduced the top-20-chunk retrieval failure rate by 67% (5.7% → 1.9%).
That last number - 67% fewer retrieval failures - comes from layering three techniques: context-enriched embeddings, hybrid keyword search, and a reranker that re-scores candidates before selecting the final chunks.
The problem nobody mentions: position bias inside the context window
Here is the non-obvious part. RAG can retrieve the right chunk and still produce a wrong answer - because of where that chunk ends up in the prompt.
The performance of LLMs often degrades when crucial information is in the middle of a long context, a "lost-in-the-middle" phenomenon that mirrors the primacy and recency effects in human memory.
A 2024 study by researchers at MIT and Google Cloud AI showed that this blind spot stems from a U-shaped attention bias: LLMs consistently favor the start and end of input sequences, neglecting the middle even when it contains the most relevant content.
Placement of key information in different positions of a prompt can significantly affect accuracy. So if your retriever returns five chunks and the correct one lands in position three of five, the model may effectively ignore it - not because retrieval failed, but because of how attention weights distribute across a long prompt.
The practical implication: production fixes include two-stage retrieval (broad recall plus cross-encoder reranking), hybrid search (semantic plus BM25), and strategic ordering - placing top evidence at the start and end of the prompt. Keep only the most relevant 3-5 documents in the prompt.
This is why a reranker is not just a nice-to-have. It controls which chunk lands where, and that order changes what the model actually attends to. A teammate like Beagle, answering questions against a Notion or Confluence knowledge base, has to solve exactly this problem: retrieve the right fragment, rerank it to the top, and put it where attention is highest.
Hybrid search: why vectors alone are not enough
Pure vector search retrieves by semantic similarity. That is powerful, but it misses exact-match cases - product names, version numbers, alphanumeric IDs - where the user's words need to match the document's words precisely.
Full-text search uses techniques like TF-IDF and BM25 to search documents by matching the keywords in a query against a database of documents.
Vector search finds and ranks documents based on cosine similarity or other distance measures between the query vector and document vectors, capturing deeper semantic meanings.
The best production systems do both simultaneously and then fuse the results.
| Signal | Strength | Weakness | Best for |
|---|---|---|---|
| Vector / semantic | Captures meaning, handles paraphrase | Misses exact product names or codes | "What's the refund policy?" |
| BM25 / keyword | Exact term match, fast | No semantic generalization | "What does clause 4.2.1 say?" |
| Hybrid (fused) | Both | More complex to tune | Most real-world knowledge bases |
| Reranker on top | Corrects ordering post-retrieval | Adds latency (~50-200 ms) | Any production system past MVP |
A practical pattern is to retrieve 50-200 candidates cheaply, then rerank down to 5-12 for context. The two-stage architecture keeps latency tight while letting you cast a wide recall net.
How RAG retrieval works: common questions
What is retrieval-augmented generation in plain English?
RAG is a method for giving an LLM access to documents it was never trained on. At query time, it fetches the most relevant fragments from a knowledge base, injects them into the prompt, and lets the model answer from that fresh context - without retraining. The model reads; it does not recall.
What chunk size should I use for a RAG system?
Start at 400-600 tokens with 10-15% overlap - this fits most prose and keeps your prompt budget predictable. FAQ and support content works well at 200-400 tokens; technical docs at 400-800 tokens; legal and regulatory text at 800-1,200 tokens. Measure retrieval recall on a test set before assuming a size works.
Why does RAG fail even when the right document is in the database?
Two common reasons: the chunk boundary stripped the context needed to make the answer meaningful, or the correct chunk landed in the middle of a long prompt where attention weights are lowest. Both are fixable - contextual chunking addresses the first; reranking and strategic prompt ordering address the second.
Is RAG better than fine-tuning?
They solve different problems. RAG gives a model access to documents that change frequently without retraining. Fine-tuning adjusts the model's behavior and style on a fixed dataset. Most production systems use RAG for live knowledge and fine-tuning for task adaptation - not one or the other.
What is a reranker and why does it matter?
A reranker is a second-stage model (typically a cross-encoder) that re-scores retrieved candidates by reading both the query and each chunk together, rather than comparing vectors separately. Adding a reranking stage pushes the total retrieval failure reduction to about 67%, taking the failure rate from 5.7% down to 1.9%. The three techniques are complementary: contextualization improves what gets indexed, hybrid search broadens recall, and reranking improves the final ordering.