A single request to Llama 3.1 70B at 128,000 tokens needs roughly 43 gigabytes of memory just for the attention cache - before the model weights even load. That number isn't a vendor complaint. It's the physics of how a context window works, and understanding it changes how you use these models.
Here's the plain version of what's happening inside every LLM inference call you send.
What an LLM context window actually is
The size of the context window determines the maximum number of tokens the model can consider at once. That sounds simple. The more interesting part is what "consider" means computationally.
Before text enters a context window, it converts into tokens through tokenization. Most modern LLMs use Byte-Pair Encoding, which breaks text into subword units - roughly one token per four characters, or about three-quarters of a word. So a 100,000-token context holds around 75,000 words. One well-annotated codebase, or about half a Harry Potter novel.
Context windows aren't just for conversations. They also store system prompts, attached documents, and source code. A few long documents can quickly fill up even a large context window.
Context windows have expanded by roughly two orders of magnitude since the original transformer architecture, from a few thousand tokens in early models to 1-2M tokens today.
| Model | Context window | Practical equivalent |
|---|---|---|
| GPT-4o (128K) | 128,000 tokens | ~96,000 words / full novel |
| Claude Opus 4.6 (200K) | 200,000 tokens | ~150,000 words / thick manual |
| Gemini 2.5 Pro (1M) | 1,000,000 tokens | ~750,000 words / mid-size codebase |
| Llama 3.1 (128K) | 128,000 tokens | same as GPT-4o; open weights |
How self-attention reads every token against every other token
Self-attention allows the model to weigh the relevance of every token to every other token in the entire input simultaneously. When processing the word "bank," the model can look at the whole sentence at once - "river" nearby makes it more likely to mean riverbank; "money" makes it more likely to mean financial institution.
That parallel awareness is the thing that makes transformers powerful. It is also what makes long contexts expensive.
The tradeoff: attention is computationally expensive at O(n²) - every time you double the context length, you quadruple the compute required for attention. A 4x cost jump per doubling is why early models were capped at 2,000 tokens and why every vendor now runs specialized infrastructure to serve 1M-token requests.
The KV cache is how inference stays tractable. A decoder-only transformer generates one token at a time. To produce the next token, the attention mechanism needs the key and value vectors for every previous token. The naive approach would recompute those for the entire sequence on every step - quadratic work that gets slower as text grows. The KV cache avoids this: when the model processes a token, it computes that token's key and value vectors once and stores them in GPU memory.
The stored vectors are why the memory bill runs high. A single Llama 3.1 70B request at 128K context uses approximately 42.9 GB for the KV cache at BF16 precision - calculated as 2 × 80 layers × 8 KV heads × 128 head dimensions × 131,072 tokens × 2 bytes.
With FP8 KV quantization this drops to roughly 21.5 GB.
KV memory scales linearly: a model needing 1.25 GiB at 4K tokens reaches 40 GiB at 128K and exceeds 300 GiB per request at 1M context. On a hosted API this becomes a dollar figure on your invoice. On hardware you own, it becomes a hard physical limit.
The lost-in-the-middle problem - why bigger isn't always better
Here is the non-obvious consequence of how attention works. A model with a 1M-token context window does not read every part of that window equally.
The "lost-in-the-middle" effect is well-documented: LLMs perform significantly worse when relevant information sits in the middle of their context rather than at the beginning or end. Liu et al. (2024) measured a 30%+ accuracy drop on multi-document question answering when the answer document moved from position 1 to position 10 in a 20-document context.
The architectural root cause lies in the RoPE long-term decay property: reduced dot-product similarity between distant token pairs systematically decreases attention weight on mid-context information. Softmax normalisation amplifies this by concentrating attention on the highest-scoring tokens, reinforcing primacy and recency advantages.
The result is a U-shaped attention curve - strong at the start and end of a context, weak in the middle - that mirrors something from 1960s memory research. This phenomenon is strikingly similar to serial position effects in human memory, where people preferentially recall items from the beginning (primacy) and end (recency) of a list with higher accuracy.
Chroma Research tested 18 frontier models and found accuracy drops of 20-50% from 10K to 100K tokens. Adding full conversation history (~113K tokens) can drop accuracy by 30% compared to a focused 300-token version.
The practical implication: ordering matters more than size. Put the most important context at the top or bottom of your prompt. A shorter, better-curated input often beats a larger one padded with loosely relevant text.
What the token count in a model card doesn't tell you
The headline context size is the ceiling, not the floor of what the model reliably uses.
RULER demonstrates that the effective context length of models is often far below their advertised maximum, with task-dependent degradation.
Unlike larger context windows, which suffer accuracy degradation beyond 32K tokens due to the "lost-in-the-middle" effect, small context windows maintain consistent attention distribution. In practice, most long-context models show sharp performance drops past 32K tokens.
What you should actually benchmark:
- Needle position sensitivity - put the key fact at position 10%, 50%, and 90% in your context and compare answers
- Distractor density - add loosely related docs and measure how often the model cites the wrong one
- Token cost at your P90 request - not the maximum allowed, but what your actual workload sends
Don't trust the spec sheet: benchmark your actual use case at your target context length.
And on cost: pick the context length you actually use, not the maximum the model allows. Most coding and document tasks live comfortably under 32K. Setting a 200K window "just in case" reserves VRAM - and API budget - you'll never fill.
LLM context window: common questions
What is a context window in an LLM?
A context window is the total number of tokens an LLM can process in one inference call - including the system prompt, user messages, retrieved documents, and the model's output. Everything outside the window is invisible to the model. Tokens are roughly three-quarters of a word, so a 128K window holds about 96,000 words.
Why does a larger context window cost more?
Because attention requires every token to compare itself against every other token in the window. The compute cost scales at O(n²) - double the context, quadruple the attention work. The KV cache stores intermediate results to avoid recomputation, but that cache grows linearly with context length and sits on GPU memory, which has a hard size limit.
What is the lost-in-the-middle problem?
It is the documented tendency of LLMs to attend more strongly to information at the start and end of a context, and less to information in the middle. Liu et al. (2024) found more than a 30% accuracy drop when a key document moved from first to tenth position in a 20-document prompt. The effect holds across every major model family tested.
How much GPU memory does a long context actually use?
A lot. A single request to Llama 3.1 70B at 128K tokens requires roughly 43 GB of KV cache memory at standard BF16 precision - before model weights. At 1M tokens that figure exceeds 300 GB, which is why million-token inference runs on multi-GPU infrastructure, not a single card.
Does a bigger context window replace RAG?
Not cleanly. Large windows reduce the need for retrieval in some workflows - you can load an entire codebase or policy library directly. But the lost-in-the-middle effect means a model may still miss facts buried in the middle of that window. For anything requiring precise retrieval across large corpora, a retrieval layer that selects and orders context deliberately will outperform blindly stuffing the full window.