Filling a 1M-token context window costs $0.14 on one model and $10.00 on another - a 71x price gap for the same number of tokens. That single fact tells you context windows are a real engineering constraint, not just a spec sheet row. But the cost gap is only the start of what is worth understanding.
What a context window actually contains
The context window is the total amount of text an LLM can process in a single request, measured in tokens. It includes everything the model reads - your prompt, previous messages, system instructions - and everything it writes back. Nothing about that list is passive. Each item competes for the same fixed budget.
As a rough rule of thumb in English, 1,000 tokens is roughly 750 words. So a model with a 1 million token context window can hold around 750,000 words, or about 1,500 pages of text, in view at one time. That sounds enormous until you remember how fast it fills. An agentic workflow that attaches a system prompt, pulls a few docs via RAG, and carries a multi-turn conversation history can burn through 50K tokens before a single line of useful output appears.
The window is a shared budget: maximum output tokens (typically 64K-128K on most models) and, on reasoning models, thinking tokens all draw from the same limit as your input. That is the first thing most people miss: the output is not in addition to the window - it comes out of it.
Research documented a four-times increase in average prompt length - from 1.5K to 6K tokens - between 2024 and 2025, driven by agentic workflows. Agents stuff in tool outputs, memory summaries, prior reasoning steps. Contexts that were comfortable at 8K a year ago now routinely hit six figures.
Why bigger windows are slow and expensive
The reason context size is hard to scale is architectural. Transformer-based LLMs use self-attention to let every token "attend to" every other token in the sequence. The computational cost scales quadratically with sequence length - doubling the context window roughly quadruples the computation for the attention layers. This is why larger context windows are expensive.
Put it in concrete numbers: 1,000 tokens means roughly 1 million attention operations. 10,000 tokens means 100 million. 200,000 tokens means 40 billion. The growth is not linear - it is quadratic.
Ring Attention is one key enabler of the 1M+ token era. Standard self-attention needs the entire key-value cache to sit in one GPU's high-bandwidth memory. For a 10 million token sequence, that cache would run into terabytes - far beyond a single H100's 80GB. Ring Attention distributes the sequence across many GPUs arranged in a logical ring.
Model providers have invested heavily in architectural optimizations - sparse attention, ring attention, efficient KV-cache management - to make million-token contexts feasible, but the fundamental cost relationship remains.
Compute requirements for self-attention push Time to First Token from a few seconds to over a minute at 1M tokens. That latency matters as much as the API bill. A 20-second wait before the first word appears is a hard limit for anything user-facing.
The number on the spec sheet is not what you actually get
Here is where the gap between marketing and mechanics is widest. Most models market a context window range, but effective context often falls far below the advertised maximum.
Research found that all models fell short of their advertised maximum context window by more than 99% in some cases.
The underlying mechanism has a name: the lost-in-the-middle effect. Lost-in-the-middle is the tendency of LLMs to use information at the beginning and end of a context window more reliably than information placed in the middle. The model may "see" the right passage or policy, but if it is buried mid-window, it may not carry enough weight in the final answer.
Research demonstrated that LLM performance on multi-document question answering and key-value retrieval follows a U-shaped function of information position: accuracy is highest when relevant information appears at the beginning or end of the input context and degrades by over 30% when relevant information is positioned in the middle. This finding replicated across six model families.
The architectural root cause lies in the RoPE long-term decay property: reduced dot-product similarity between distant token pairs systematically decreases attention weight on mid-context information. Softmax normalisation amplifies this by concentrating attention on the highest-scoring tokens, reinforcing primacy and recency advantages.
Chroma's research tested 18 frontier models and found that every single one gets worse as input length increases. Not most - all.
Being able to accept 10 million tokens does not guarantee reasoning over them. Independent evaluations of Llama 4 Scout indicate that its reliable reasoning degrades substantially beyond the 128K-256K range, even though it physically accepts far more.
Where context windows stand right now - and what it means in practice
As of September 2026, Grok 4.20's 2M-token window is the largest in production, though GPT-5.6 (1.05M), Claude Opus 5/Sonnet 5 (1M), and Gemini 3.1 Pro (1M) lead among current flagships.
| Model | Window (advertised) | Output cap | Reliable range |
|---|---|---|---|
| Grok 4.20 | 2M tokens | - | Not published |
| GPT-5.6 | 1.05M tokens | ~128K | - |
| Claude Opus 5 / Sonnet 5 | 1M tokens | 300K (beta) | Scores well to 1M |
| Gemini 3.1 Pro | 1M tokens | 65K | - |
| Llama 4 Scout (open) | 10M tokens | - | ~128-256K reliable |
There is a distinction worth internalising: the advertised context window is the maximum a model accepts, while the cost-efficient context window is the range where pricing remains economical for a given workload. For Gemini's tiered pricing, a 300K-token prompt costs up to twice as much per input token as a 100K-token prompt.
The practical upshot is three rules for anyone building on top of these APIs:
Put the most important content first or last. The middle of a long context is the place your model is most likely to ignore. If a critical instruction is buried in paragraph 12 of a pasted document, move it to the top of the system prompt.
Budget each slot explicitly. Every LLM request has four token consumers - system prompt, user input, RAG context, and reserved output - and each needs an explicit budget before you write a line of application code.
Use RAG instead of brute-force stuffing. Simply pasting everything into a huge prompt is rarely the best strategy. Retrieving only the relevant chunks keeps the signal-to-noise ratio high and keeps the critical content near the edges where the model actually attends.
A teammate like Beagle, which answers questions in Slack by pulling from linked sources, works within this constraint by design - it retrieves a targeted document excerpt rather than dumping an entire knowledge base into the prompt.
How LLM context windows work: common questions
What is an LLM context window, in plain language?
The context window is the model's working memory for a single request. It holds your system prompt, conversation history, any documents you attach, and the model's own response - all at once. When the total token count exceeds the window, the model either errors out or silently drops the oldest content. Think of it as a whiteboard that gets erased from one end when it runs out of space.
Why does context window size affect cost so much?
Because attention computation scales quadratically. Doubling your context from 64K to 128K tokens roughly quadruples the computation inside the model's attention layers. Providers absorb this with architectural tricks like Flash Attention and Ring Attention, but the cost still rises - and shows up as both higher per-token prices and longer latency before you see the first output word.
Does a bigger context window mean better answers?
Not automatically. Research across 18 frontier models found that every one gets worse as input grows. The lost-in-the-middle effect means content buried in the middle of a long context receives far less attention than content at the edges - with accuracy dropping over 30% in documented benchmarks. A smaller, well-curated context often outperforms a large, noisy one.
What is the difference between advertised and effective context length?
The advertised limit is the maximum tokens a model will accept without throwing an error. The effective limit is how far into that window the model can reliably use information for reasoning and retrieval. Research found models failing to use information reliably well below their advertised ceilings - in some cases degrading at just a few thousand tokens. The effective limit varies by task type, not just token count.
How should I decide how much context to send?
Put critical instructions in the system prompt (always near the top). Retrieve only relevant chunks via RAG rather than pasting entire documents. Reserve explicit token budgets for each component - system prompt, user message, retrieved context, and output. And watch your cost tier: some providers charge up to 2x per token once you cross a threshold like 128K, so the cheapest call is often the one that stays narrow.