You paste a 40-page contract into Claude and ask it to flag the indemnification clause. The model returns the right answer. Then you paste a 200-page contract and ask the same question - and it misses the clause, even though the document technically fits inside the advertised 1M-token window. This is not a bug report. It is how context windows work.
The number on the spec sheet is a hard ceiling, not a quality guarantee. Understanding the difference between the two changes how you build with these models.
What the token limit actually measures
A context window is the maximum number of tokens a model can process in a single forward pass. It defines the maximum tokens in one request - including both input tokens (your prompt) and output tokens (the model's reply). Everything goes into that shared budget: every time you send a message, the entire conversation history is sent with it. The model does not remember previous messages - it receives them all, in order, as a single input sequence. Tool use results, if the model used function calling, are concatenated in there too.
Before text enters a context window, it converts into tokens through a process called tokenization. Most modern LLMs use Byte-Pair Encoding, which breaks text into subword units. As a rough approximation, one token represents about four characters or three-quarters of a word.
A standard 250-word page of English prose is approximately 330 tokens. So a 200-page contract at roughly 250 words per page is around 66,000 tokens - well inside a 1M window on paper, but that is only part of the story.
Context windows have expanded by roughly two orders of magnitude since the original transformer architecture, from a few thousand tokens in early models to 1-2M tokens today. The current frontier: Claude Opus and Sonnet ship 1M tokens, matching GPT-4.1's 1M and sitting under Gemini 2.5 Pro (1M, with 2M coming) and Llama 4 Scout's claimed 10M.
The spec-sheet race, though, has outrun what the architecture can actually do with all those tokens.
Why filling the window is not the same as using it
Transformer models use the self-attention mechanism. It calculates the relationships between different parts of the input - for example, how relevant the word at the beginning of a sentence is to the word at the end. The mechanism computes vectors of weights, where each weight represents how important one token is to another.
The computational problem is that the attention mechanism requires quadratic O(n²) scaling in computational requirements as the input sequence length grows. Double the context, quadruple the work. At 1M tokens, generating every token requires 1,765 seconds with over 96% of latency spent on attention operations without the optimization that makes long context practical at all: the KV cache.
A decoder-only transformer generates one token at a time. To produce the next token, the attention mechanism needs the key and value vectors for every token that came before it. The naive approach would recompute those keys and values for the entire sequence on every single step - quadratic work that gets slower as the text grows. The KV cache is the optimization that avoids it. When the model processes a token, it computes that token's key and value vectors once and stores them in GPU memory.
That storage is expensive. A 7B model with 32 layers, 32 heads, and head dimension 128 in FP16 costs approximately 0.5 MB per token. At 128K context, that is still 64 GB - most of a single GPU's VRAM budget before you account for model weights or activation memory.
A single Llama 3.1 405B user at 128K tokens burns 66 GB of GPU memory before the model even responds.
On a hosted API you do not pay in VRAM - you pay in dollars. If your agent session runs 50+ tool calls, the history alone can exceed 150K tokens, billed again on every subsequent call. Long-context costs come from this compounding re-send, not from the one-off large request.
The attention blind spot in the middle
Here is the part most developers miss, and it explains the 200-page contract failure from the opening.
A 2024 study by researchers at MIT and Google Cloud AI showed that this blind spot stems from a U-shaped attention bias: LLMs consistently favor the start and end of input sequences, neglecting the middle even when it contains the most relevant content. The original 2023 Stanford and UNC paper established the pattern: LLM performance on multi-document question answering and key-value retrieval follows a U-shaped function of information position - accuracy is highest when relevant information appears at the beginning or end of the context and degrades by more than 30% when relevant information is positioned in the middle.
This finding replicated across six model families: GPT-3.5-Turbo, GPT-4, Claude 1.3, LongChat-13B, MPT-30B, and Cohere Command.
The architectural root cause lies in the RoPE long-term decay property: reduced dot-product similarity between distant token pairs systematically decreases attention weight on mid-context information. Softmax normalisation amplifies this by concentrating attention on the highest-scoring tokens, reinforcing primacy and recency advantages.
Vendors have pushed back on this with needle-in-a-haystack benchmarks: GPT-4.1 advertises 100% recall of needles anywhere inside a 1M-token prompt, Claude 3 reports 99%+ on a strengthened Needle-in-a-Haystack evaluation, and Gemini 1.5 Pro reaches 99.7% recall across up to 1M tokens. The catch is that NVIDIA's RULER puts usable context at 50-65% of advertised. The NoLiMa benchmark, which forces the model to infer a link rather than string-match, found that 11 of 13 models advertising at least 128K context dropped below 50% of their short-context baseline by 32K tokens, and GPT-4o fell from 99.3% to 69.7%.
In other words: needle tests measure retrieval of one planted fact. Reasoning over information spread through a long document is a harder, different problem - and that is the problem most real workflows actually present.
What this means when you build
Three practical rules that follow from the mechanics:
- Put the most critical content first or last. If you cannot control where information lands in a long context, structure your prompt so the instructions and most relevant facts bookend the long middle material.
- Treat the advertised window as a maximum, not a target. Unlike larger context windows which suffer accuracy degradation beyond 32K tokens due to the "lost-in-the-middle" effect, small focused contexts maintain consistent attention distribution. In practice, most long-context models show sharp performance drops past 32K tokens. For most tasks - debugging, Q&A, document review - a tight, well-curated context outperforms a bloated one.
- Monitor token spend in multi-turn agent sessions. If you are regularly hitting context limits, the problem is almost always conversation management, not context window size. Summarize or compact earlier turns before the window fills, not after.
A common shortcut teams take: paste an entire Notion doc or Slack thread export into a prompt and call it "context." The model sees it all, but the middle of that paste is the part attention weights least. A teammate like Beagle retrieves only the relevant sections before sending anything to the model - a small change that turns a 60K-token bloated prompt into a sharp 4K one.
How LLM context windows work: common questions
What is a context window in an LLM?
A context window is the maximum number of tokens - roughly three-quarters of a word each - that a model can process in a single request. It includes your system prompt, conversation history, any tool outputs, and the model's response. Everything draws from the same shared budget, and nothing outside the window influences the model's output.
Does a bigger context window mean better performance?
Not automatically. A 1M-token context window does not give you 5x the effective context of a 200K window because attention quality degrades as the sequence grows. Larger windows help for tasks that genuinely require scanning a full codebase or document set. For typical focused tasks, a well-structured short context consistently outperforms a large one stuffed with loosely relevant material.
What is the "lost in the middle" problem?
It is a documented attention bias where LLMs perform significantly worse on information positioned in the middle of a long prompt. Accuracy drops significantly for information near the center of the context window - a pattern strikingly similar to serial position effects in human memory, where people recall items from the beginning and end of a list with higher accuracy, producing a U-shaped curve. Mitigation: put critical facts at the top or bottom of your prompt, and use retrieval to avoid stuffing the middle in the first place.
What is a KV cache and why does it matter?
Transformer-based LLMs generate text autoregressively, one token at a time. The KV cache eliminates redundant work by storing the key and value vectors for previously processed tokens the first time they are computed, so only the new token's vectors need to be calculated at each step. The tradeoff is that this cache is not free: it lives in the same GPU memory pool as the model's weights and grows for the entire duration of a request. On hosted APIs, that growth shows up directly in your token bill on long multi-turn sessions.
How should I decide how much context to pass to a model?
Pass the minimum that answers the question. A 500-token summary of a 10K document often outperforms stuffing the whole document in context. For multi-document tasks, retrieve the relevant sections first rather than concatenating everything. Set a practical budget - say, 30K tokens - as your default ceiling, and expand it only for tasks verified to need it.