The largest advertised context window in mid-2026 is 10 million tokens, but effective performance typically degrades well before that limit - most models deliver reliable quality at 60-70% of their stated maximum. That gap is not a footnote. It changes how you architect anything that puts long documents, conversation history, or retrieved chunks into a prompt.
Here is what is actually happening under the hood, and why throwing more tokens at a problem is often the wrong answer.
What the context window actually is
The context window determines how much text the model can "see" at once - your prompt, any system instructions, retrieved documents, conversation history, and the model's own response all have to fit inside it. That is the whole budget. Nothing persists between calls; every time you send a message to a language model, the entire conversation history is sent with it - the model does not remember previous messages, it receives them all, in order, as a single input sequence.
One token represents about four characters or three-quarters of a word, though this varies by language and the specific tokenizer. A 1M-token window holds roughly 750,000 words - or about three average-sized novels. That sounds spacious. The problem is not the advertised ceiling. It is the engineering reality underneath it.
The computational cost of attention scales quadratically with sequence length - doubling the context window roughly quadruples the computation for the attention layers. This is why three constraints create context window limits: O(n²) complexity in self-attention, KV cache memory growth, and GPU memory bandwidth.
The KV cache is worth understanding specifically. A decoder-only transformer generates one token at a time. To produce the next token, the attention mechanism needs the key and value vectors for every token that came before it. The naive approach would recompute those keys and values for the entire sequence on every single step - quadratic work that gets slower as the text grows. The KV cache is the optimization that avoids it. But it has a cost: at 128K tokens, a 7B model's KV cache (~64 GB) exceeds the capacity of an A100 GPU. On a hosted API, that cost shows up in your invoice. Filling the same 1M-token window costs $0.14 on DeepSeek V4 Flash and $10.00 on Claude Fable 5.
The lost-in-the-middle problem: attention is not uniform
This is the non-obvious part - the thing most teams miss when they first hit a long-context failure.
The lost-in-the-middle effect is a demonstrated phenomenon where LLMs perform significantly worse when relevant information is placed in the middle of the input context rather than at the beginning or end. It produces a U-shaped attention curve: the model attends well to what comes first (system prompt, task framing) and what comes last (the current query), and poorly to everything sandwiched between.
Liu et al. (2024) measured a 30%+ accuracy drop on multi-document question answering when the answer document moved from position 1 to position 10 in a 20-document context.
This finding replicated across six model families - GPT-3.5-Turbo, GPT-4, Claude 1.3, LongChat-13B, MPT-30B, and Cohere Command.
The architectural cause: the root cause lies in the RoPE long-term decay property - reduced dot-product similarity between distant token pairs systematically decreases attention weight on mid-context information. Softmax normalisation amplifies this by concentrating attention on the highest-scoring tokens, reinforcing primacy and recency advantages.
A key fact placed at position 60,000 in a 100,000-token context has a 20-30% lower chance of being correctly used compared to the same fact placed in the first 5,000 tokens. This effect has improved in newer models - Claude Opus 4.6 and GPT-5.5 handle it better than their predecessors - but has not been eliminated.
A second mechanism compounds this: transformer attention is quadratic, so 100K tokens means 10 billion pairwise relationships.
As context length increases, each individual token receives proportionally less of the model's attention budget. A relevant 500-token passage competes for attention with every other token. In a 10K-token context, it gets roughly 5% of the attention. In a 1M-token context, it gets 0.05%.
The gap between advertised and effective context
A model's advertised window is its capacity. Effective context is the length over which quality holds. Every long-context benchmark ever published shows a gap between the two.
Among proprietary flagships, GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8, and Claude Sonnet 5 have converged on 1 million tokens, so the practical differentiator between them is now reasoning quality, multimodality, and price rather than raw window size.
| Model tier | Advertised window | Approx. effective quality threshold | Notes |
|---|---|---|---|
| Llama 4 Scout | 10M tokens | No published benchmark shows quality holding near this | Best for archival breadth work |
| GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro | 1M tokens | ~600-700K before degradation | Converged flagship tier |
| DeepSeek V3, Mistral Large | 128K tokens | ~80-90K | Open-weight practical ceiling |
| GPT-4o class | 128K tokens | Quality drops sharply past ~32K | Well-documented in benchmarks |
The gap between advertised and effective context length is substantial: models typically break 30-40% before their claimed limit - a 200K model becomes unreliable around 130K tokens.
Beyond the accuracy issues, two hard limits keep enormous context windows from being a universal answer: latency and cost. The time to first token for a multi-million-token prompt can be measured in tens of seconds to minutes, even on modern hardware.
What most teams find in production: if you are regularly hitting context limits on Claude or GPT-4o, the problem is almost always conversation management, not context window size.
What to do instead of filling the window
The non-obvious insight from all of this: position matters as much as presence. Getting the right context into the window is only half the job. Putting it in the right place is the other half.
Front-load critical facts. System-level instructions and the most important constraints belong at the top of your prompt, not buried after retrieved documents.
End with the question. The model attends strongly to the most recent tokens. Put the actual query last, not mid-prompt.
Summarize aggressively before inserting. A 500-token summary of a 10K document often outperforms stuffing the whole document in context.
Prefer retrieval over stuffing. A 100K-line repo fills a 1M window before you add the system prompt, conversation history, or tool outputs, and the model still needs room to write its answer. In practice, anything above roughly half the window forces a choice: retrieve only the relevant slice or compress the history you carry forward.
Compact long conversations. Long-running applications need to periodically summarize and restart context. Your users won't notice if you do it well.
Don't benchmark with needles. Despite achieving perfect results in needle-in-haystack tests, almost all models fail to maintain performance in other RULER tasks as context length increases. Passing basic needle-in-haystack tests doesn't guarantee true long-context understanding capabilities. Test with your own documents.
LLM context windows: common questions
How does a context window actually work?
Transformer models use the self-attention mechanism, which calculates the relationships between different parts of the input. The self-attention mechanism computes vectors of weights, where each weight represents how important one token is to another. All tokens in your prompt - instructions, history, retrieved chunks, tool outputs - are processed together as a single sequence, up to the model's token limit.
Why does information in the middle of a long context get ignored?
The beginning of a context window contains system instructions and task framing that become strong anchors. The end sits closest to the current user request. The middle has neither advantage - it is farther from task framing and the final query, and competes with more nearby tokens and distractors. The result is a measurable 30%+ accuracy drop for facts positioned mid-context.
Is a 1M-token context window actually useful?
For tasks where you need to search across a very large codebase or document set, large context windows genuinely help. For typical development tasks - debugging, code review, architecture discussion - the effective range of 200K is more than sufficient when used well. Match the window to the task; bigger is not automatically better, and it is always more expensive.
What is the KV cache and why does it matter?
The KV cache stores the computed key and value vectors for every token already processed, so the model does not recompute them on each generation step. KV cache memory grows approximately linearly with the maximum context length and key architectural parameters. On your own hardware that means VRAM pressure; on a hosted API it means higher per-call cost at long contexts, even when output is short.
How do I know what context length is safe to use?
Don't trust the spec sheet: benchmark your actual use case at your target context length. Use your own documents, not synthetic retrieval tasks. Monitor context length in production and set up alerts when conversations exceed the effective context threshold - usually 30-50% of the advertised maximum.