Your Notion runbook gets split into 43 pieces before any AI reads a word of it. That split - how big, where it breaks, how much it overlaps with the next piece - turns out to matter more than most teams expect when they're debugging why their AI assistant keeps returning the wrong paragraph.
Retrieval-augmented generation (RAG) is the mechanism behind almost every AI assistant that reads your internal docs: your Slack bot that surfaces tickets, your support agent that pulls from a knowledge base, a teammate like Beagle answering a question about last quarter's incident timeline. The architecture is not complicated, but the details at each step have compounding effects. Here is exactly what happens.
The five mechanical steps of a RAG pipeline
RAG operates through four core stages: ingestion (loading your data into a source like a vector database), retrieval (fetching relevant data based on a query), augmentation (combining the retrieved data and the query into a prompt), and generation (the model producing output from the augmented prompt). In practice, retrieval itself splits into two sub-steps - embedding and search - so there are really five things to understand.
Step 1 - Chunking. Your raw document gets sliced into segments before anything else happens. Chunk size is the single most impactful hyperparameter in a RAG pipeline. NVIDIA's research found that factoid queries perform best at 256-512 tokens, while multi-hop analytical queries benefit from 512-1,024 tokens. Get it wrong and you pay: getting this wrong by one bracket degrades context precision by 15-30%.
The failure modes run in both directions. Chunks that are too small lose context - in the FloTorch 2026 benchmark, semantic chunking produced fragments averaging 43 tokens, and those fragments scored only 54% accuracy on end-to-end questions.
A January 2026 systematic analysis identified a "context cliff" around 2,500 tokens where response quality drops.
The practical default: 256 to 512 tokens with 10 to 25% overlap. Microsoft Azure recommends 512 tokens with 25% overlap (128 tokens) as a starting point, using BERT tokens rather than character counts.
Step 2 - Embedding. Each chunk is fed through an embedding model, which converts it into a high-dimensional vector - a list of floating-point numbers that encodes the chunk's meaning geometrically. Vector databases store documents as embeddings in a high-dimensional space, allowing for fast and accurate retrieval based on semantic similarity. The same embedding model must be used at query time, which is a common deployment trap: swap the embedding model after indexing and every stored vector becomes nonsense.
Step 3 - Indexing. The corpus is divided into chunks, converted into vector representations using an embedding model, and those representations are stored in a vector database for later retrieval. The index structure determines query speed. Approximate nearest-neighbor algorithms like IVFPQ trade a tiny bit of recall for orders-of-magnitude faster search at scale - fine for most teams, but worth knowing when you're debugging recall problems.
Step 4 - Query embedding and search. When a question arrives, the same embedding model converts it into a vector. That question vector queries the vector database, identifying the top N chunks most similar to it - typically via shortest cosine distance - and those N chunks are extracted as context for the LLM.
Step 5 - Augmentation and generation. In hybrid search, results from both dense and sparse indexes are combined, de-duplicated, and reranked by a unified relevance score. The most relevant matches are then used to construct an augmented prompt - both the search results and the user's query - to send to the LLM.
Using the augmented prompt, the LLM now has access to the most pertinent grounding facts from your vector database, reducing the likelihood of hallucination.
What actually goes into the augmented prompt
This is where most explanations hand-wave. The prompt the LLM receives is not your original question. It is a structured document assembled from three parts: the system prompt (who the agent is, what it can do), the retrieved chunks (verbatim text from your vector store, with metadata), and the user's query.
A minimal version looks roughly like this: