How RAG Actually Works: Chunks, Vectors, and Retrieval

Retrieval-augmented generation is the architecture behind most serious AI knowledge tools. Here is exactly how the pipeline works-from chunking to vector search to the final answer.

Cover art for How RAG Actually Works: Chunks, Vectors, and Retrieval

Searching one million 768-dimensional vectors with a flat scan takes roughly 2-3 seconds on modern hardware. That is why every production AI knowledge tool you use today runs a different algorithm entirely-and understanding that one detail unlocks how retrieval-augmented generation (RAG) actually works from end to end.

RAG is now the default pattern for grounding LLMs in real documents. It is how internal knowledge bots find your runbooks, how legal tools cite statutes, how support agents pull from product docs. But most explanations stop at "the model retrieves relevant chunks." That is barely the beginning. Here is the full picture.

The pipeline from document to answer

RAG connects large language models to external knowledge bases through a five-stage pipeline-ingestion, embedding, retrieval, augmentation, and generation-enabling accurate, domain-specific answers without retraining the model. Here is what each stage actually does.

Ingestion. Raw documents-PDFs, Notion pages, Confluence wikis, Slack exports-get loaded and cleaned. Data may come from PDFs, websites, databases, CRM systems, ERP systems, internal wikis, or product documentation.

Chunking. Chunking breaks large documents into smaller segments before generating embeddings.

Chunk size directly affects output quality: chunks that are too small lose surrounding context, while chunks that are too large dilute the specific passage most relevant to the user question.

Common strategies include fixed-size splitting by token count and sentence-boundary splitting with overlapping borders to reduce the risk of key context falling at a boundary.

Overlap helps preserve context across chunk boundaries.

Embedding. Embedding transforms each chunk into a dense vector-a numerical representation that captures the semantic meaning of the text. These vectors, generated by pre-trained language models, are stored in a vector database.

The embedding model vectorizes the data in a multidimensional mathematical space, arranging data points by similarity-points judged to be closer in relevance are placed closer together.

Retrieval. When a user submits a query, the RAG system applies the same embedding model to convert the user input into a vector representation and queries the vector store, executing a similarity search that returns the top-k most relevant document chunks.

The k value-how many chunks to retrieve-trades off retrieval coverage against context window consumption and must be tuned for the target LLM.

Generation. The retrieved documents are concatenated as context with the original input prompt and fed to the text generator, which produces the final output.

Why vector search is the hard part

This is where most explanations wave their hands. The retrieval step has a real engineering problem: how do you find the closest matching vectors in a corpus of millions without scanning every single one?

Searching one million 768-dimensional vectors with a flat scan takes roughly 2-3 seconds on modern hardware. At 10 million items it takes 20-30 seconds. At the latencies required by real search systems-results in under 50 milliseconds-exact search is not viable past a few hundred thousand items.

The answer is approximate nearest-neighbor search, and specifically an algorithm called HNSW (Hierarchical Navigable Small World).

A vector index using HNSW organizes embeddings to enable approximate nearest-neighbor search at scale, reducing retrieval from a linear scan of all embeddings to a sub-millisecond lookup.

To answer a query, HNSW starts at the topmost layer and greedily walks toward the query vector, descending one layer at a time until it reaches the base layer in O(log|idx|) expected steps.

The word "approximate" sounds worrying. It is not. For the vast majority of production retrieval workloads, recall@10 of 0.95 or above is indistinguishable from exact search in downstream quality metrics. You are trading a fraction of a percent of theoretical accuracy for two orders of magnitude of speed.

Align chunk size with context windows, use task-specific sentence transformers, and default to HNSW with metadata filtering for sub-100 ms retrieval at 95%+ recall. That is the practical production target.

~2-3 secexact scan of 1M vectorsnot viable in production
sub-50 msHNSW approximate retrievalthe production standard
35% → 6%hallucination rate dropcurated corpus vs. general web in one RAG study

A concrete example: a support team's runbook

Say you have 400 internal runbooks in Notion. Someone types into Slack: "What's the rollback procedure for the payments service?"

Here is what fires:

  1. The query gets embedded into a vector-let's say 1,536 dimensions if you're using OpenAI's text-embedding-3-large.
  2. The system runs HNSW search against the indexed chunks of those 400 runbooks and returns the top 5 by cosine similarity.
  3. Those 5 chunks-perhaps 300 tokens each-get prepended to the query as context.
  4. The LLM reads the augmented prompt and generates an answer that cites the specific steps from the runbook.

The whole retrieve-plus-generate round trip in a well-tuned system runs in under two seconds. Without RAG, the model would either hallucinate a procedure or admit it does not know.

Beagle in action#platform-eng, 11:02am
The ask
'anyone know the rollback steps for payments?'
Beagle drafts
embeds the query, retrieves top 3 runbook chunks from the connected Notion workspace, drafts a reply with the procedure and a direct link to the source doc
You approve
you approve; the answer posts with full provenance, no copy-paste required
Do this in your workspace →

RAG vs. fine-tuning: where each one belongs

Teams often ask whether they should fine-tune instead. The short answer: RAG is used primarily to inject new knowledge into a model, while fine-tuning is best for changing behavior, tone, or task structure.

Fine-tuning is not, despite how it is often marketed, a great mechanism for teaching a model facts-especially facts that change. Every time your runbooks are updated, you'd need another fine-tuning run. With RAG, you update the document in Notion and re-index.

Dimension RAG Fine-tuning
Knowledge stays current Yes-update the corpus No-requires a new training run
Citable sources Yes-chunks link back to origin No-knowledge is in weights
Cost to update Low (re-index) High (GPU hours)
Behavior/tone control Weak Strong
Hallucination risk Lower with good corpus Higher for factual recall

A hybrid approach combining both RAG and fine-tuning typically outperforms either method alone-fine-tuning handles behavioral consistency while RAG keeps responses factually current from live knowledge bases. Most serious enterprise deployments end up here: RAG for the facts, a lightly tuned model for the voice.

Where RAG actually fails

One genuine non-obvious problem: a 2024 PubMed study on RAG chatbots for cancer information found hallucination rates up to 35% when using general web search as the retrieval corpus, versus 6% when using a curated, domain-specific knowledge base. The architecture is identical in both cases. The corpus is what changed.

The embedding corpus quality matters as much as the embedding model. A RAG system over a messy, contradictory, or outdated knowledge base will confidently retrieve the wrong thing and hand the model plausible-sounding garbage to summarize.

The other failure mode is positional. When a RAG system retrieves multiple chunks and passes them all to the LLM as context, the model does not weight them equally. This is sometimes called the "lost in the middle" problem-chunks near the start and end of the context window get more attention than those in the center. Reranking with a cross-encoder before handing off to the LLM is the standard fix: it reorders the top-k chunks by relevance before they enter the context.

RAG requires continuous updates to maintain data relevance; production systems need automated pipelines that detect updated source documents and trigger re-embedding on a scheduled or event-driven basis. That re-indexing cost is invisible in demos and real in production.

Answering 'where is the runbook?' in Slack
Without Beagle
someone posts the question; three people DM different links; the correct one gets found 8 minutes later after a thread of 'try this one'
With Beagle
a RAG-backed teammate retrieves the right chunk, posts the exact procedure with a source link, and the thread closes in 20 seconds

RAG retrieval: common questions

What is retrieval-augmented generation?

RAG is an architecture that gives an LLM access to an external knowledge base at query time, without changing the model's weights. The user's question gets converted into a vector, the most similar document chunks are retrieved, and those chunks are injected into the prompt before the model generates an answer.

How does RAG reduce hallucinations?

By grounding the model in retrieved text rather than its training memory alone. RAG, which forces models to ground answers in external documents, can reduce hallucinations by 40-71% in many scenarios. But it is not a guarantee-hallucination rates depend heavily on whether the retrieved chunks are accurate and relevant in the first place.

What chunk size should I use for RAG?

There is no universal answer, but 256-512 tokens with a 10-20% overlap is a common starting point for prose documents. When chunks are too large, data points can become too general and fail to correspond to user queries. When chunks are too small, they can lose semantic coherency. Test against your actual queries and measure retrieval recall.

When should I use RAG instead of fine-tuning?

Choose RAG when you need up-to-date, verifiable knowledge-it connects a language model to your own document base and retrieves relevant passages for each query so answers stay current and can cite their source. Choose fine-tuning when you need specific behavior-a consistent tone, format, or specialized reasoning pattern.

Do I need a dedicated vector database?

Not necessarily for small corpora. Postgres with the pgvector extension handles millions of vectors adequately. Dedicated databases like Pinecone, Qdrant, or Weaviate add operational tooling-filtering, namespacing, automatic re-indexing-that matters once you have multiple collections or high query volume.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle