Searching one million 768-dimensional vectors with a flat scan takes roughly 2-3 seconds on modern hardware. That is why every production AI knowledge tool you use today runs a different algorithm entirely-and understanding that one detail unlocks how retrieval-augmented generation (RAG) actually works from end to end.
RAG is now the default pattern for grounding LLMs in real documents. It is how internal knowledge bots find your runbooks, how legal tools cite statutes, how support agents pull from product docs. But most explanations stop at "the model retrieves relevant chunks." That is barely the beginning. Here is the full picture.
The pipeline from document to answer
RAG connects large language models to external knowledge bases through a five-stage pipeline-ingestion, embedding, retrieval, augmentation, and generation-enabling accurate, domain-specific answers without retraining the model. Here is what each stage actually does.
Ingestion. Raw documents-PDFs, Notion pages, Confluence wikis, Slack exports-get loaded and cleaned. Data may come from PDFs, websites, databases, CRM systems, ERP systems, internal wikis, or product documentation.
Chunking. Chunking breaks large documents into smaller segments before generating embeddings.
Chunk size directly affects output quality: chunks that are too small lose surrounding context, while chunks that are too large dilute the specific passage most relevant to the user question.
Common strategies include fixed-size splitting by token count and sentence-boundary splitting with overlapping borders to reduce the risk of key context falling at a boundary.
Overlap helps preserve context across chunk boundaries.
Embedding. Embedding transforms each chunk into a dense vector-a numerical representation that captures the semantic meaning of the text. These vectors, generated by pre-trained language models, are stored in a vector database.
The embedding model vectorizes the data in a multidimensional mathematical space, arranging data points by similarity-points judged to be closer in relevance are placed closer together.
Retrieval. When a user submits a query, the RAG system applies the same embedding model to convert the user input into a vector representation and queries the vector store, executing a similarity search that returns the top-k most relevant document chunks.
The k value-how many chunks to retrieve-trades off retrieval coverage against context window consumption and must be tuned for the target LLM.
Generation. The retrieved documents are concatenated as context with the original input prompt and fed to the text generator, which produces the final output.
Why vector search is the hard part
This is where most explanations wave their hands. The retrieval step has a real engineering problem: how do you find the closest matching vectors in a corpus of millions without scanning every single one?
Searching one million 768-dimensional vectors with a flat scan takes roughly 2-3 seconds on modern hardware. At 10 million items it takes 20-30 seconds. At the latencies required by real search systems-results in under 50 milliseconds-exact search is not viable past a few hundred thousand items.
The answer is approximate nearest-neighbor search, and specifically an algorithm called HNSW (Hierarchical Navigable Small World).
A vector index using HNSW organizes embeddings to enable approximate nearest-neighbor search at scale, reducing retrieval from a linear scan of all embeddings to a sub-millisecond lookup.
To answer a query, HNSW starts at the topmost layer and greedily walks toward the query vector, descending one layer at a time until it reaches the base layer in O(log|idx|) expected steps.
The word "approximate" sounds worrying. It is not. For the vast majority of production retrieval workloads, recall@10 of 0.95 or above is indistinguishable from exact search in downstream quality metrics. You are trading a fraction of a percent of theoretical accuracy for two orders of magnitude of speed.
Align chunk size with context windows, use task-specific sentence transformers, and default to HNSW with metadata filtering for sub-100 ms retrieval at 95%+ recall. That is the practical production target.
A concrete example: a support team's runbook
Say you have 400 internal runbooks in Notion. Someone types into Slack: "What's the rollback procedure for the payments service?"
Here is what fires:
- The query gets embedded into a vector-let's say 1,536 dimensions if you're using OpenAI's
text-embedding-3-large. - The system runs HNSW search against the indexed chunks of those 400 runbooks and returns the top 5 by cosine similarity.
- Those 5 chunks-perhaps 300 tokens each-get prepended to the query as context.
- The LLM reads the augmented prompt and generates an answer that cites the specific steps from the runbook.
The whole retrieve-plus-generate round trip in a well-tuned system runs in under two seconds. Without RAG, the model would either hallucinate a procedure or admit it does not know.
RAG vs. fine-tuning: where each one belongs
Teams often ask whether they should fine-tune instead. The short answer: RAG is used primarily to inject new knowledge into a model, while fine-tuning is best for changing behavior, tone, or task structure.
Fine-tuning is not, despite how it is often marketed, a great mechanism for teaching a model facts-especially facts that change. Every time your runbooks are updated, you'd need another fine-tuning run. With RAG, you update the document in Notion and re-index.
| Dimension | RAG | Fine-tuning |
|---|---|---|
| Knowledge stays current | Yes-update the corpus | No-requires a new training run |
| Citable sources | Yes-chunks link back to origin | No-knowledge is in weights |
| Cost to update | Low (re-index) | High (GPU hours) |
| Behavior/tone control | Weak | Strong |
| Hallucination risk | Lower with good corpus | Higher for factual recall |
A hybrid approach combining both RAG and fine-tuning typically outperforms either method alone-fine-tuning handles behavioral consistency while RAG keeps responses factually current from live knowledge bases. Most serious enterprise deployments end up here: RAG for the facts, a lightly tuned model for the voice.
Where RAG actually fails
One genuine non-obvious problem: a 2024 PubMed study on RAG chatbots for cancer information found hallucination rates up to 35% when using general web search as the retrieval corpus, versus 6% when using a curated, domain-specific knowledge base. The architecture is identical in both cases. The corpus is what changed.
The embedding corpus quality matters as much as the embedding model. A RAG system over a messy, contradictory, or outdated knowledge base will confidently retrieve the wrong thing and hand the model plausible-sounding garbage to summarize.
The other failure mode is positional. When a RAG system retrieves multiple chunks and passes them all to the LLM as context, the model does not weight them equally. This is sometimes called the "lost in the middle" problem-chunks near the start and end of the context window get more attention than those in the center. Reranking with a cross-encoder before handing off to the LLM is the standard fix: it reorders the top-k chunks by relevance before they enter the context.
RAG requires continuous updates to maintain data relevance; production systems need automated pipelines that detect updated source documents and trigger re-embedding on a scheduled or event-driven basis. That re-indexing cost is invisible in demos and real in production.
RAG retrieval: common questions
What is retrieval-augmented generation?
RAG is an architecture that gives an LLM access to an external knowledge base at query time, without changing the model's weights. The user's question gets converted into a vector, the most similar document chunks are retrieved, and those chunks are injected into the prompt before the model generates an answer.
How does RAG reduce hallucinations?
By grounding the model in retrieved text rather than its training memory alone. RAG, which forces models to ground answers in external documents, can reduce hallucinations by 40-71% in many scenarios. But it is not a guarantee-hallucination rates depend heavily on whether the retrieved chunks are accurate and relevant in the first place.
What chunk size should I use for RAG?
There is no universal answer, but 256-512 tokens with a 10-20% overlap is a common starting point for prose documents. When chunks are too large, data points can become too general and fail to correspond to user queries. When chunks are too small, they can lose semantic coherency. Test against your actual queries and measure retrieval recall.
When should I use RAG instead of fine-tuning?
Choose RAG when you need up-to-date, verifiable knowledge-it connects a language model to your own document base and retrieves relevant passages for each query so answers stay current and can cite their source. Choose fine-tuning when you need specific behavior-a consistent tone, format, or specialized reasoning pattern.
Do I need a dedicated vector database?
Not necessarily for small corpora. Postgres with the pgvector extension handles millions of vectors adequately. Dedicated databases like Pinecone, Qdrant, or Weaviate add operational tooling-filtering, namespacing, automatic re-indexing-that matters once you have multiple collections or high query volume.