Gemini 2.5 Pro accepts 2 million tokens. Llama 4 Scout claims 10 million. And yet a 2024 Stanford and University of Washington study found that model performance can drop by more than 30% when the critical information shifts from the edges of the context to the middle. Bigger windows, same blind spot.
That gap - between the number on the spec sheet and what the model actually uses - is the thing worth understanding. The LLM context window is the most consequential dial in modern AI systems, and it's also the most commonly misread one.
What the context window actually contains
The context window is the maximum amount of text - or other tokenized input - available to the model at one time when generating output. That's the clean definition. The messier reality is that the window is a shared budget, and everything competes for space inside it.
An LLM context window is a temporary working space. It contains the current prompt, instructions, conversation history, retrieved information, tool results, and the model's output. That last part matters more than it sounds: the tokens the model writes back also consume the window. The context window covers both input and output tokens combined. A model with a 128K context window can handle 128,000 tokens total - if your input uses 100K tokens, the model only has 28K tokens left for its response.
It is measured in tokens, which are units produced by the model's tokenizer rather than words or characters.
In English, one token roughly equals three-quarters of a word. So a 200,000-token window holds somewhere around 150,000 words - about 500 dense pages of text, all held simultaneously in the model's working view.
LLMs are stateless: they do not inherently remember past interactions. The context window is the model's working memory, so it determines how much of the earlier conversation or task the model can recall. There is no persistent awareness sitting in the background. Every inference call re-reads the full context from scratch.
How window sizes got so large, so fast
Context windows have expanded by roughly two orders of magnitude since the original transformer architecture, from a few thousand tokens in early models to 1-2M tokens today. The milestone timeline shows just how compressed that expansion was.
Claude 2 made the 100K window broadly available in July 2023. GPT-4 Turbo became the first mainstream 128K model in November 2023. Claude 2.1 doubled that to 200K the same month. Then in February 2024, Gemini 1.5 Pro, initially in private preview, became the first one-million-token-class window.
Today the 1M token threshold is essentially the baseline for flagship models.
| Model | Context window | Max output |
|---|---|---|
| GPT-5 | 1,000,000 tokens | 32,768 tokens |
| Gemini 2.5 Pro | 1,000,000 tokens | 65,536 tokens |
| Claude Opus 4 | 200,000 tokens | 32,000 tokens |
| Llama 4 Scout | 10,000,000 tokens | 16,384 tokens |
| DeepSeek V3 | 128,000 tokens | 8,192 tokens |
Output limits are much smaller than input limits. Even models with 1M input contexts typically cap output at 8K-65K tokens. The context window is asymmetric - you can feed the model a lot, but it will not write a novel-length response in one go.
The engineering question is: what made 1M-token windows physically possible?
Every decoder-only transformer model generates text autoregressively. Each new token attends to every previous token in the context. Without caching, this requires recomputing the key and value matrices for every prior token on every single forward pass. The compute cost is quadratic in context length: double the context, quadruple the computation.
That quadratic scaling is the wall. The KV cache is what gets you past it.
The KV cache is the optimization that avoids recomputing prior tokens. When the model processes a token, it computes that token's key and value vectors once and stores them in GPU memory. Every subsequent step reuses them. Generation becomes linear in the number of new tokens instead of quadratic in the sequence length.
The trade-off is memory. For a Llama-3 70B-class model, the per-token KV cache cost works out to approximately 0.3125 MB per token. Multiply that by a 128K context and you're looking at roughly 40 GB just for the cache - the same order as the model weights themselves. This means even if you have enough VRAM to store the KV cache for a 128K context window, you might still hit an out-of-memory error during the prefill stage because the attention matrix becomes too large to fit in memory.
Techniques like FlashAttention-2 and FlashAttention-3 mitigate this by using tiling and recomputation to avoid materializing the full attention matrix in VRAM. They don't change the math; they change the order of operations so the GPU never has to hold the whole thing at once.
The problem nobody mentions in the spec sheet
Here is the non-obvious part: bigger context windows don't fix the model's ability to reason across all of them.
The 'lost-in-the-middle' problem occurs when LLMs prioritize the beginning and end of a context window over critical information buried in the middle.
Research from Stanford and the University of Washington demonstrates that LLMs exhibit a U-shaped performance curve when processing long contexts. Models achieve highest accuracy when relevant information appears at the beginning or end of the input context, but performance degrades significantly when critical information is positioned in the middle.
Results showed that performance can degrade by more than 30% when relevant information shifts from the start or end positions to the middle of the context window. And this is not a 2023-era problem that has since been patched. Chroma's 2025 context rot report tested 18 LLMs, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 models. The report found that newer models still do not use context uniformly, and performance grows less reliable as input length grows.
The U-curve exists even in models with context windows of 4K, 16K, and 32K tokens. Research from 2025 confirmed it persists in models with 128K+ windows. Bigger windows mean more middle, which means more room for information to get lost.
The gap between the advertised window and the effective one is measurable. Even a 70B model trained to 128K context (Llama 3.1) was found to only leverage about 64K effectively.
Empirical tests show that even models advertising 1M-token input rarely sustain high-accuracy reasoning across more than half of it. In real use, effective context often caps around 30-60% of the stated window before recall decay sets in.
Why? The recency effect directly aligns with the shape of short-term memory demand in the training data, while the primacy effect is induced by uniform long-term memory demand and is influenced by the model's autoregressive properties and attention sinks. In other words, the model learned to pay attention to beginnings and ends because that is where useful information lived in training data. It is a feature masquerading as a bug.
What teams actually do about it
The cleanest strategy is to not fill the window. A larger context window allows the model to handle more complex and lengthy prompts, but more context isn't automatically better. As token count grows, accuracy and recall degrade - a phenomenon known as context rot. This makes curating what's in context just as important as how much space is available.
Signal density matters more than filling the window. Remove irrelevant or obsolete material while preserving the structure and explanation needed to interpret important evidence.
When you're building agents that run long enough to hit the limit, compaction is the current mitigation of choice. Anthropic's server-side compaction automatically summarizes earlier parts of the conversation so it can continue past the context window limit. It is available in beta for Claude 4.6 and later models.
Long-running conversations and agentic tasks often hit the context window. Context compaction automatically summarizes and replaces older context when the conversation approaches a configurable threshold, letting Claude perform longer tasks without hitting limits.
The cost of compaction is fidelity. Claude just can't remember what you were chatting about. The coding conventions established, the specific phrasing asked for - all summarised into oblivion. A teammate like Beagle, working inside Slack threads, avoids this by keeping each interaction short and tightly scoped - never accumulating a sprawling context in the first place.
The architecture-level fix is sub-agent isolation. In Anthropic's production multi-agent research system, a lead agent spawns sub-agents that each get a self-contained task, an output format, and a fresh, isolated context window. The heavy, noisy search context stays inside the sub-agent; only the distilled result returns to the lead, protecting the lead's attention budget.
The retrieval approach maps directly to the lost-in-the-middle research: put the relevant material where attention is strongest - at the start of the context - rather than burying it somewhere the model statistically ignores.
LLM context window explained: common questions
What is a context window in an LLM?
A context window is the total number of tokens - input and output combined - that a model can process in one inference call. It includes the system prompt, conversation history, retrieved documents, tool results, and the model's own reply. Nothing outside the window exists to the model.
Why does the context window matter for AI agents?
Agents accumulate context fast: tool calls, results, reasoning traces, and conversation history all stack up inside the same window. Once the window fills, the model either starts losing earlier context (truncation), gets compacted into a lossy summary, or fails outright. Window management is the core engineering challenge of long-running agents.
What is the lost-in-the-middle problem?
It is the empirical finding that LLMs attend more reliably to tokens at the beginning and end of their context than to tokens in the middle. Performance on multi-document question answering can drop over 30% when the relevant passage is placed in the middle rather than the edges. Larger windows don't eliminate this - they add more middle.
What is a KV cache and how does it relate to the context window?
The KV cache stores the key and value vectors for every token already processed, so the model doesn't recompute them on each generation step. Without it, compute cost would scale quadratically with context length - doubling the context would quadruple the work. The trade-off is memory: a 70B-class model uses roughly 0.3 MB of VRAM per token in the cache.
How much of a 1M-token context window can a model actually use?
In practice, effective reasoning typically holds across 30-60% of the stated window before recall accuracy begins to drop. A 1M-token window doesn't mean the model reasons equally well across all 1M positions. The edges are reliable; the deep middle is not. RAG and sub-agent isolation are the practical workarounds.