Context Window Sizes in 2026: GPT-6 Astra, Claude Sonnet 5, Gemini 3.1 Pro, Llama 4, Kimi K3 — and What It Costs to Fill Them
Context window comparison for GPT-6 Astra, Claude Opus 5, Gemini 3.1 Pro, Llama 4 Scout, Kimi K3, Grok 4.6 and more, with the cost of one full-window request.
"What's the context window of model X?" is one of the most repeated questions in LLM engineering, and the answer changes every few months. This post collects the current numbers for the major models in one table, then does the part the spec sheets skip: what a request that actually fills the window costs, and why the advertised number is never the number you can use.
Window sizes and prices below are from the Token Inspector catalog, last verified against provider documentation on 2026-09-15. Prices are USD per million input tokens.
Context Windows by Model
| Model | Provider | Context window (tokens) | Input $/M |
|---|---|---|---|
| GPT-6 Astra | OpenAI | 1,050,000 | $10.00 |
| GPT-5.6 Sol | OpenAI | 1,050,000 | $4.00 |
| GPT-5.6 Terra | OpenAI | 1,050,000 | $2.00 |
| GPT-5.6 Luna | OpenAI | 1,050,000 | $0.20 |
| o3 | OpenAI | 200,000 | $2.00 |
| GPT-4o | OpenAI | 128,000 | $2.50 |
| Claude Fable 5.1 | Anthropic | 1,000,000 | $10.00 |
| Claude Opus 5 | Anthropic | 1,000,000 | $5.00 |
| Claude Sonnet 5 | Anthropic | 1,000,000 | $2.00 |
| Claude Haiku 4.5 | Anthropic | 200,000 | $1.00 |
| Gemini 3.1 Pro | 1,000,000 | $2.00 | |
| Gemini 3.6 Flash | 1,000,000 | $0.75 | |
| Kimi K3 | Moonshot AI | 1,048,576 | $3.00 |
| Kimi K2.6 | Moonshot AI | 262,144 | $0.95 |
| Llama 4 Maverick | Meta (via API providers) | 1,048,576 | $0.27 |
| Llama 4 Scout | Meta (via API providers) | 1,048,576 | $0.18 |
| DeepSeek-V4 Pro | DeepSeek | 1,000,000 | $0.66 |
| Qwen3.8-Max | Alibaba | 1,000,000 | $2.00 |
| GLM-5.2 | Zhipu AI | 1,000,000 | $1.40 |
| Grok 4.6 | xAI | 500,000 | $2.00 |
| Mistral Large 3 | Mistral | 262,144 | $0.50 |
| Phi-4 | Microsoft | 16,000 | $0.07 |
Two things stand out. First, the million-token window is now the default for flagship and mid-tier models from every major provider; the 128K–200K models are legacy or small-tier. Second, window size no longer tracks price: Llama 4 Scout and Claude Fable 5.1 have the same ~1M window and a 55x difference in input price.
What One Full-Window Request Costs
Window size only matters if you can afford to use it. Filling the entire window once, input tokens only:
| Model | Tokens | Cost to fill once |
|---|---|---|
| GPT-6 Astra | 1,050,000 | $10.50 |
| Claude Fable 5.1 | 1,000,000 | $10.00 |
| Claude Opus 5 | 1,000,000 | $5.00 |
| Kimi K3 | 1,048,576 | $3.15 |
| GPT-5.6 Terra | 1,050,000 | $2.10 |
| Claude Sonnet 5 | 1,000,000 | $2.00 |
| Gemini 3.1 Pro | 1,000,000 | $2.00 |
| Grok 4.6 | 500,000 | $1.00 |
| Gemini 3.6 Flash | 1,000,000 | $0.75 |
| DeepSeek-V4 Pro | 1,000,000 | $0.66 |
| GPT-4o | 128,000 | $0.32 |
| GPT-5.6 Luna | 1,050,000 | $0.21 |
| Llama 4 Scout | 1,048,576 | $0.19 |
A flagship request at full window is a $5–$10 line item per call. An agent loop that re-sends a near-full context every turn pays that every turn, which is why prompt caching and context trimming matter more than the headline window. Output tokens are billed at a higher rate and aren't included above.
Why the Advertised Window Isn't the Usable Window
The number on the spec sheet is a ceiling for input plus output plus everything else the request carries. In practice you lose room to:
- The system prompt and tool schemas. A production agent with 20 tools can spend thousands of tokens on definitions before the user says a word.
- Reserved output. The window is shared. A request that fills the window with input leaves no room for the answer, and reasoning models also spend tokens on thinking.
- Tokenizer differences. The same text is a different number of tokens in each model family, so "1M tokens" is not "1M tokens of your text" across vendors. We covered the mechanics in Why the Same Prompt Has Different Token Counts on GPT-4o, Claude, and Gemini. Newer tokenizers can also change counts within a single provider's lineup, so re-measure when you upgrade models.
- Retrieval quality. A model that accepts a million tokens doesn't necessarily use the middle of them equally well. Long-context performance varies by model and task, so test on your own documents instead of assuming the window is uniformly usable.
A Rough Conversion Table
As an English-text rule of thumb, one token is about 0.75 words, so:
| Window | Approximate English words |
|---|---|
| 16K (Phi-4) | ~12,000 |
| 128K (GPT-4o) | ~96,000 |
| 200K (o3, Claude Haiku 4.5) | ~150,000 |
| 262K (Mistral Large 3, Kimi K2.6) | ~196,000 |
| 1M (most flagships) | ~750,000 |
Code, JSON, and non-English text tokenize less efficiently and fit fewer words. Measure your real input instead of estimating.
How to Pick
- Whole-codebase or whole-contract analysis in one shot: pick a 1M-window model, and check the full-window cost above first. Gemini 3.6 Flash, DeepSeek-V4 Pro, and GPT-5.6 Luna make this cheap enough to do routinely.
- Chat and agents with long sessions: window size matters less than how fast the context fills. Predict when your context window will fill and plan trimming or summarization before you hit the ceiling.
- RAG: a bigger window doesn't remove the need for good chunking. Retrieving less, better, is cheaper and often more accurate than stuffing everything in. See token vs sentence vs paragraph chunking.
- Small, fixed prompts: a 128K–262K model is plenty and usually cheaper. Don't pay for a window you won't fill.
Check Your Own Prompt
Paste your actual prompt into Token Inspector to see its token count and projected cost across every model in the table, using each model's tokenizer family. That tells you how much of each window you really use, and which model fits your workload at your volume.
Follow Trango Compute on LinkedIn
We post updates on new tools, context engineering patterns, and LLM cost research.