How Does RAG Actually Work? The Pipeline, Step by Step

Retrieval-augmented generation is behind most AI Q&A tools, but the actual pipeline is rarely explained plainly. Here's how RAG works under the hood, from chunking to vector search to the answer.

Cover art for How Does RAG Actually Work? The Pipeline, Step by Step

Ask an AI assistant where last quarter's refund policy lives and, if it answers correctly, a six-step mechanical process just ran faster than you blinked. That process is retrieval-augmented generation - RAG - and it powers everything from internal knowledge bots to customer-facing support tools. Most explanations skip from "it retrieves relevant docs" straight to "the model answers." The middle part is where the interesting failures live.

Here is what actually happens.

The offline phase: turning documents into searchable geometry

Before any question is asked, you have to build the index. A RAG pipeline starts by loading data, breaking text into overlapping chunks, creating embeddings, and uploading them to a vector database. None of that happens at query time. It happens once - or whenever your documents change.

Chunking is the first decision, and it is underappreciated. The quality of a RAG system lives or dies by how you chunk documents - it sits at the critical junction between raw data and the vector store.

Get it wrong and you will either retrieve fragments too small to be useful, or chunks so large they bury the relevant information in noise.

The research on chunk size is messier than most blog posts admit. Query type affects optimal chunk size: factoid questions perform best with 256-512 tokens; analytical queries need 1,024 or more. A clinical decision-support study found that adaptive chunking aligned to logical topic boundaries hit 87% accuracy versus 13% for fixed-size baselines. But a separate benchmark found a fixed size of 200 words could match or even outperform semantic chunking on real-world datasets. The honest answer: start with a standard recursive splitter, measure recall on representative queries, then tune.

A good starting baseline is a chunk size of 512 tokens and a chunk overlap of 50-100 tokens. The overlap stops a sentence from being split across two chunks with neither chunk carrying its full meaning.

Embedding converts each chunk into a vector - a list of floating-point numbers that encodes meaning as position in high-dimensional space. RAG works by embedding text chunks into a vector space where similar chunks are close to each other and can be found using fast nearest-neighbor search.

The numbers here are concrete. OpenAI's text-embedding-3-small and text-embedding-3-large produce 1,536- and 3,072-dimensional vectors respectively and are a common default for RAG.

Cohere's embed-english-v3.0 outputs 1,024-dimensional embeddings.

OpenAI's model uses Matryoshka Representation Learning, so embeddings can be shortened by truncating from the end - and its API accepts up to 8,192 tokens of input, while Cohere's limit is 512 tokens per chunk.

OpenAI's text-embedding-3-small costs $0.02 per million tokens; text-embedding-3-large costs $0.13 per million tokens. Embedding a 300-page book with the small model costs under $3.

Once embeddings exist, they get written to a vector database - Pinecone, Chroma, Weaviate, pgvector, or a dozen others - indexed for fast approximate nearest-neighbor lookup.

The online phase: what happens when someone asks a question

A user types a question. Here is what fires:

  1. Query embedding - the same embedding model that processed your documents now encodes the query into the same vector space. The AI sends the query to the embedding model, which converts it into a numeric vector.

  2. Vector search - the embedding model compares these numeric values to vectors in the index and retrieves the related data when it finds matches. This is not keyword matching. Vector search focuses on semantic relationships, not surface-level word overlap.

  3. Reranking (optional but important) - modern RAG pipelines implement a two-stage retrieval approach: the first stage uses vector similarity to retrieve a larger candidate set, typically 20-100 results, prioritizing recall over precision.

The second stage applies a reranking model that evaluates each document's relevance with greater precision - cross-encoder models like BERT-based rerankers jointly encode query and document, capturing semantic relationships that bi-encoder embeddings miss.

  1. Prompt augmentation - the top-ranked chunks get inserted into the prompt alongside the original question. You create an augmented prompt with both the search results and the user's query to send to the LLM. A typical template is just: "Using the CONTEXT provided, answer the QUESTION. If the CONTEXT doesn't contain the answer, say you don't know."

  2. Generation - the LLM combines the retrieved content and its own response to the query into a final answer, potentially citing sources the embedding model found.

Beagle in action#support, 2:07pm
The ask
'what's our current SLA for P2 tickets?'
Beagle drafts
converts the question to a query embedding, retrieves the top-3 chunks from the linked Confluence SLA doc, assembles an augmented prompt
You approve
you approve a reply that cites the exact clause and page - not a hallucination from training data, but text pulled seconds ago from the live doc
Do this in your workspace →

The "lost in the middle" problem - and why more chunks isn't better

This is the part that most RAG tutorials omit entirely. You can retrieve the right document, pass it into the prompt, and still get a wrong answer.

Large language models now support context windows extending to millions of tokens, yet research reveals a critical limitation: these models struggle to effectively use information located in the middle of long contexts - a phenomenon known as the "lost in the middle" problem.

LLM performance on multi-document question answering follows a U-shaped function of information position: accuracy is highest when relevant information appears at the beginning or end of the input context and degrades by more than 30% when relevant information is positioned in the middle. This finding replicated across six model families including GPT-3.5-Turbo, GPT-4, Claude 1.3, and Cohere Command.

The architectural reason: the root cause lies in RoPE's long-term decay property, which reduces dot-product similarity between distant token pairs and systematically decreases attention on mid-context information. Softmax normalisation amplifies this by concentrating attention on the highest-scoring tokens.

In practical terms: your system retrieved five perfect chunks. The answer was in chunk #3. The LLM read chunk #1 carefully, skimmed chunks #2 through #4, and paid close attention to chunk #5. It is not carelessness - it is architecture.

If RAG retrieves too many chunks or orders them poorly, the correct chunk can still land in a weak middle position. Reranking helps by surfacing the most relevant chunk to position #1 in the prompt. But the real fix is keeping your context lean: fewer, better chunks beat a prompt stuffed with tangentially related text.

30%+accuracy dropwhen the right chunk lands in the middle of the prompt
87% vs 13%adaptive vs fixed chunkingon a clinical RAG benchmark
$0.02per million tokensto embed with OpenAI text-embedding-3-small

When hybrid search beats pure vector similarity

Pure vector search can miss exact matches. If a user asks for a specific product SKU or a regulation number, semantic similarity does not help - the phrase has to appear verbatim. This is where hybrid search earns its keep.

In hybrid search, you query both a dense (vector) index and a sparse (keyword) index, combine and de-duplicate the results, and use a reranking model to produce a unified relevance score.

Retrieval type Strength Weakness
Dense vector Finds semantically similar text Misses exact strings, codes, IDs
Sparse (BM25 / keyword) Nails exact-match lookups Blind to paraphrase and synonyms
Hybrid Covers both More infrastructure, needs a reranker

A good rule of thumb: if your users will ever paste in a ticket number or a clause reference, hybrid is the right default. For open-ended conversational Q&A, pure vector retrieval is fine and cheaper.

Answering "what does our refund policy say about digital goods?"
Without Beagle
the model draws on training data, produces a plausible-sounding but possibly outdated or hallucinated answer with no source
With Beagle
RAG retrieves the current policy doc chunk, augments the prompt, and the model answers from that text - with a clause link the team can check

A teammate like Beagle, which pulls answers from linked docs in Slack, is running a version of this pipeline on every query - embed the question, retrieve from the connected knowledge base, augment, generate, then hold the draft for a human to approve before it posts. The draft-and-approve step is the part the pipeline itself cannot provide: a check that the retrieved chunk actually answers the question being asked, not just the question the model assumed was being asked.

How does RAG work: common questions

What is retrieval-augmented generation in plain terms?

RAG is a two-phase system. First, it converts your documents into numeric vectors and stores them in a searchable index. When a question arrives, it converts the question into the same vector format, finds the closest document chunks, inserts them into the model's prompt, and generates an answer grounded in that retrieved text rather than the model's training data.

Why do RAG systems give wrong answers even when the document exists?

Three common reasons: the relevant chunk was not retrieved (a recall failure, usually from chunking or embedding choices); the chunk was retrieved but landed in the middle of a long prompt where the model underweights it; or the retrieved chunk was adjacent to the answer but did not contain it. Each failure mode has a different fix.

What chunk size should I use for RAG?

Factoid queries perform best with 256-512 tokens; analytical queries need 1,024 or more.

A reasonable starting baseline is 512 tokens with 50-100 tokens of overlap. Measure recall on your actual query distribution before committing to a size - benchmarks on other datasets do not reliably transfer.

What is the difference between vector search and keyword search in RAG?

Vector search encodes meaning as geometry and finds semantically similar text even when different words are used. Keyword search finds exact string matches. Hybrid search combines both and uses a reranker to merge the results - useful when users include precise codes, names, or identifiers that pure vector search can miss.

Does a larger context window make RAG unnecessary?

No. Long-context models can miss evidence buried mid-window, while vanilla RAG can add noise through weak retrieval and chunking. Large context windows reduce the need to chunk aggressively, but they do not remove the positional attention bias - and stuffing a million-token context with every document is slower and more expensive than retrieving the five chunks that actually matter.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle