LLM Context Windows: What Actually Happens to Your Tokens

An LLM context window isn't just a reading limit - it's the boundary of what a model can reason over, and the physics behind it have real cost and accuracy consequences. Here's how it works.

Cover art for LLM Context Windows: What Actually Happens to Your Tokens

A single request to Llama 3.1 70B at 128,000 tokens needs roughly 43 gigabytes of memory just for the attention cache - before the model weights even load. That number isn't a vendor complaint. It's the physics of how a context window works, and understanding it changes how you use these models.

Here's the plain version of what's happening inside every LLM inference call you send.

What an LLM context window actually is

The size of the context window determines the maximum number of tokens the model can consider at once. That sounds simple. The more interesting part is what "consider" means computationally.

Before text enters a context window, it converts into tokens through tokenization. Most modern LLMs use Byte-Pair Encoding, which breaks text into subword units - roughly one token per four characters, or about three-quarters of a word. So a 100,000-token context holds around 75,000 words. One well-annotated codebase, or about half a Harry Potter novel.

Context windows aren't just for conversations. They also store system prompts, attached documents, and source code. A few long documents can quickly fill up even a large context window.

Context windows have expanded by roughly two orders of magnitude since the original transformer architecture, from a few thousand tokens in early models to 1-2M tokens today.

Model Context window Practical equivalent
GPT-4o (128K) 128,000 tokens ~96,000 words / full novel
Claude Opus 4.6 (200K) 200,000 tokens ~150,000 words / thick manual
Gemini 2.5 Pro (1M) 1,000,000 tokens ~750,000 words / mid-size codebase
Llama 3.1 (128K) 128,000 tokens same as GPT-4o; open weights

How self-attention reads every token against every other token

Self-attention allows the model to weigh the relevance of every token to every other token in the entire input simultaneously. When processing the word "bank," the model can look at the whole sentence at once - "river" nearby makes it more likely to mean riverbank; "money" makes it more likely to mean financial institution.

That parallel awareness is the thing that makes transformers powerful. It is also what makes long contexts expensive.

The tradeoff: attention is computationally expensive at O(n²) - every time you double the context length, you quadruple the compute required for attention. A 4x cost jump per doubling is why early models were capped at 2,000 tokens and why every vendor now runs specialized infrastructure to serve 1M-token requests.

The KV cache is how inference stays tractable. A decoder-only transformer generates one token at a time. To produce the next token, the attention mechanism needs the key and value vectors for every previous token. The naive approach would recompute those for the entire sequence on every step - quadratic work that gets slower as text grows. The KV cache avoids this: when the model processes a token, it computes that token's key and value vectors once and stores them in GPU memory.

The stored vectors are why the memory bill runs high. A single Llama 3.1 70B request at 128K context uses approximately 42.9 GB for the KV cache at BF16 precision - calculated as 2 × 80 layers × 8 KV heads × 128 head dimensions × 131,072 tokens × 2 bytes.

With FP8 KV quantization this drops to roughly 21.5 GB.

KV memory scales linearly: a model needing 1.25 GiB at 4K tokens reaches 40 GiB at 128K and exceeds 300 GiB per request at 1M context. On a hosted API this becomes a dollar figure on your invoice. On hardware you own, it becomes a hard physical limit.

43 GBKV cache for one 128K requestLlama 3.1 70B at BF16
4xcompute growth per 2x contextdue to O(n²) attention
300 GBKV memory at 1M contextexceeds a single H100 card

The lost-in-the-middle problem - why bigger isn't always better

Here is the non-obvious consequence of how attention works. A model with a 1M-token context window does not read every part of that window equally.

The "lost-in-the-middle" effect is well-documented: LLMs perform significantly worse when relevant information sits in the middle of their context rather than at the beginning or end. Liu et al. (2024) measured a 30%+ accuracy drop on multi-document question answering when the answer document moved from position 1 to position 10 in a 20-document context.

The architectural root cause lies in the RoPE long-term decay property: reduced dot-product similarity between distant token pairs systematically decreases attention weight on mid-context information. Softmax normalisation amplifies this by concentrating attention on the highest-scoring tokens, reinforcing primacy and recency advantages.

The result is a U-shaped attention curve - strong at the start and end of a context, weak in the middle - that mirrors something from 1960s memory research. This phenomenon is strikingly similar to serial position effects in human memory, where people preferentially recall items from the beginning (primacy) and end (recency) of a list with higher accuracy.

Chroma Research tested 18 frontier models and found accuracy drops of 20-50% from 10K to 100K tokens. Adding full conversation history (~113K tokens) can drop accuracy by 30% compared to a focused 300-token version.

The practical implication: ordering matters more than size. Put the most important context at the top or bottom of your prompt. A shorter, better-curated input often beats a larger one padded with loosely relevant text.

Beagle in action#legal-ops, 11:02am
The ask
'can you check the indemnity clause against our standard policy?'
Beagle drafts
reads the pinned policy doc (placed at the top of context) and the pasted contract section (placed at the bottom), drafts a comparison with the relevant clause numbers
You approve
a teammate reviews and approves; the answer cites both sources and took 18 seconds
Do this in your workspace →

What the token count in a model card doesn't tell you

The headline context size is the ceiling, not the floor of what the model reliably uses.

RULER demonstrates that the effective context length of models is often far below their advertised maximum, with task-dependent degradation.

Unlike larger context windows, which suffer accuracy degradation beyond 32K tokens due to the "lost-in-the-middle" effect, small context windows maintain consistent attention distribution. In practice, most long-context models show sharp performance drops past 32K tokens.

What you should actually benchmark:

  • Needle position sensitivity - put the key fact at position 10%, 50%, and 90% in your context and compare answers
  • Distractor density - add loosely related docs and measure how often the model cites the wrong one
  • Token cost at your P90 request - not the maximum allowed, but what your actual workload sends

Don't trust the spec sheet: benchmark your actual use case at your target context length.

And on cost: pick the context length you actually use, not the maximum the model allows. Most coding and document tasks live comfortably under 32K. Setting a 200K window "just in case" reserves VRAM - and API budget - you'll never fill.

Passing context into a long thread
Without Beagle
the full Slack thread (300+ messages, ~60K tokens) is dumped into the prompt; the model answers well on the first and last few messages, misses a key decision buried in message 150
With Beagle
Beagle retrieves the three most relevant messages by content, puts them at the top of a focused prompt; the answer is accurate and the token bill is 8x lower

LLM context window: common questions

What is a context window in an LLM?

A context window is the total number of tokens an LLM can process in one inference call - including the system prompt, user messages, retrieved documents, and the model's output. Everything outside the window is invisible to the model. Tokens are roughly three-quarters of a word, so a 128K window holds about 96,000 words.

Why does a larger context window cost more?

Because attention requires every token to compare itself against every other token in the window. The compute cost scales at O(n²) - double the context, quadruple the attention work. The KV cache stores intermediate results to avoid recomputation, but that cache grows linearly with context length and sits on GPU memory, which has a hard size limit.

What is the lost-in-the-middle problem?

It is the documented tendency of LLMs to attend more strongly to information at the start and end of a context, and less to information in the middle. Liu et al. (2024) found more than a 30% accuracy drop when a key document moved from first to tenth position in a 20-document prompt. The effect holds across every major model family tested.

How much GPU memory does a long context actually use?

A lot. A single request to Llama 3.1 70B at 128K tokens requires roughly 43 GB of KV cache memory at standard BF16 precision - before model weights. At 1M tokens that figure exceeds 300 GB, which is why million-token inference runs on multi-GPU infrastructure, not a single card.

Does a bigger context window replace RAG?

Not cleanly. Large windows reduce the need for retrieval in some workflows - you can load an entire codebase or policy library directly. But the lost-in-the-middle effect means a model may still miss facts buried in the middle of that window. For anything requiring precise retrieval across large corpora, a retrieval layer that selects and orders context deliberately will outperform blindly stuffing the full window.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle