Models / Context windows

LLM context windows, ranked

Every generally-available model, ranked by how many tokens it can hold at once. Sort by context to find the biggest window, or by price to find the cheapest large-context model.

The short answer

The largest context window among the GA LLMs we track is Llama 4 Scout (17B-16E Instruct) (Meta) at 10M tokens. Next is GPT-5.6 Sol at 1.1M. 32 models now offer a context window of 1M tokens or more; the cheapest of those is Qwen-Flash at $0.138 per 1M tokens blended.

Last verified: 19 Aug 2026 · sourced from official provider documentation

The ranking

78 models, largest context first

Ranked by context window. Click a header to re-sort by price, or filter by provider.

Model Status Context Input $/M Output $/M Blended $/M Cutoff
Llama 4 Scout (17B-16E Instruct)open
Meta · text, image
GA10M$0.1$0.3$0.152024-08
GPT-5.6 Sol
OpenAI · text, image
GA1.1M$5$30$11.252026-02-16
GPT-5.6 Terra
OpenAI · text, image
GA1.1M$2$12$4.502026-02-16
GPT-5.6 Luna
OpenAI · text, image
GA1.1M$0.2$1.20$0.452026-02-16
GPT-5.5
OpenAI · text, image
GA1.1M$5$30$11.252025-12-01
GPT-5.5 Pro
OpenAI · text, image
GA1.1M$30$180$67.502025-12-01
GPT-5.4
OpenAI · text, image
GA1.1M$2.50$15$5.632025-08-31
GPT-5.4 Pro
OpenAI · text, image
GA1.1M$30$180$67.502025-08-31
Gemini 3.7 Flash
Google · text, image, video, audio, pdf
GA1.0M$0.75$3.75$1.50
Gemini 3.6 Flash
Google · text, image, video, audio, pdf
GA1.0M$0.75$3.75$1.50
Gemini 3.5 Flash
Google · text, image, video, audio, pdf
GA1.0M$1.50$9$3.382025-01
Gemini 3.5 Flash-Lite
Google · text, image, video, audio, pdf
GA1.0M$0.3$2.50$0.85
Gemini 3.1 Flash-Lite
Google · text, image, video, audio, pdf
GA1.0M$0.25$1.50$0.5632025-01
Claude Fable 5
Anthropic · text, image, vision
GA1M$10$50$20
Claude Opus 4.8
Anthropic · text, image, vision
GA1M$5$25$102026-01
Claude Opus 4.7
Anthropic · text, image, vision
GA1M$5$25$102026-01
Claude Opus 4.6
Anthropic · text, image, vision
GA1M$5$25$102025-05
Claude Sonnet 5
Anthropic · text, image, vision
GA1M$3$15$62026-01
Claude Sonnet 4.6
Anthropic · text, image, vision
GA1M$3$15$62025-08
Grok 4.3
xAI · text-in, image-in, text-out
GA1M$1.25$2.50$1.56
Grok 4.20 (0309) Reasoning
xAI · text-in, image-in, text-out
GA1M$1.25$2.50$1.56
Grok 4.20 (0309) Non-Reasoning
xAI · text-in, image-in, text-out
GA1M$1.25$2.50$1.56
Llama 4 Maverick (17B-128E Instruct)open
Meta · text, image
GA1M$0.15$0.6$0.2622024-08
Amazon Nova 2 Lite
Amazon · text, image, video
GA1M$0.3$2.50$0.852025-10
Qwen3.7-Max
Alibaba · text, image, video
GA1M$2.50$7.50$3.75
Qwen3.7-Plus
Alibaba · text, image, video
GA1M$0.4$1.60$0.7
Qwen3.6-Flash
Alibaba · text, image, video
GA1M$0.25$1.50$0.563
Qwen3.7-Plus (snapshot 2026-05-26)
Alibaba · text, image, video
GA1M$0.4$1.60$0.7
Qwen3.6-Plus
Alibaba · text, image, video
GA1M$0.5$3$1.13
Qwen3.5-Flash
Alibaba · text, image
GA1M$0.1$0.4$0.175
Qwen-Plus (Qwen3-series)
Alibaba · text
GA1M$0.4$1.20$0.6
Qwen-Flash
Alibaba · text
GA1M$0.05$0.4$0.138
Grok 4.5
xAI · text-in, image-in, text-out
GA500K$2$6$3
GPT-5.6 Cyber
OpenAI · text, image
GA400K$12.50$75$28.132026-02-16
GPT-5.4 mini
OpenAI · text, image
GA400K$0.75$4.50$1.692025-08-31
GPT-5.4 nano
OpenAI · text, image
GA400K$0.2$1.25$0.4632025-08-31
GPT-5.3-Codex
OpenAI · text, image
GA400K$1.75$14$4.812025-08-31
Amazon Nova Lite
Amazon · text, image, video
GA300K$0.06$0.24$0.1052024-10
Amazon Nova Pro
Amazon · text, image, video
GA300K$0.8$3.20$1.402024-10
Qwen3-Max
Alibaba · text
GA262K$1.20$6$2.402025-06
Qwen3.5-Plus
Alibaba · text, image
GA262K$0.4$2.40$0.9
Kimi K2.7 Codeopen
Moonshot · text
GA262K$0.95$4$1.71
Kimi K2.7 Code HighSpeedopen
Moonshot · text
GA262K$1.90$8$3.42
Kimi K2.6open
Moonshot · text, image, video
GA262K$0.95$4$1.71
Kimi K2.5open
Moonshot · text, image, video
GA262K$0.6$3$1.20
Grok Build 0.1
xAI · text-in, image-in, text-out
GA256K$1$2$1.25
Mistral Large 3open
Mistral · text, vision
GA256K$0.5$1.50$0.75
Mistral Small 4open
Mistral · text, vision
GA256K$0.15$0.6$0.262
Ministral 3 8Bopen
Mistral · text, vision
GA256K$0.15$0.15$0.15
Command Aopen
Cohere · text
GA256K
Command A Reasoningopen
Cohere · text
GA256K
Claude Opus 4.5
Anthropic · text, image, vision
GA200K$5$25$102025-05
Claude Sonnet 4.5
Anthropic · text, image, vision
GA200K$3$15$62025-01
Claude Haiku 4.5
Anthropic · text, image, vision
GA200K$1$5$22025-02
Sonar Pro
Perplexity · text, image
GA200K$3$15$6
Moonshot v1 128K
Moonshot · text
GA131K$2$5$2.75
Llama 3.3 70B Instructopen
Meta · text
GA128K$0.1$0.32$0.1552023-12
Llama 3.3 8B Instruct (Llama API)open
Meta · text
GA128K
Llama 3.1 405B Instructopen
Meta · text
GA128K2023-12
Llama 3.1 70B Instructopen
Meta · text
GA128K2023-12
Llama 3.1 8B Instructopen
Meta · text
GA128K$0.02$0.03$0.0222023-12
Llama 3.2 90B Vision Instructopen
Meta · text, image
GA128K2023-12
Llama 3.2 11B Vision Instructopen
Meta · text, image
GA128K2023-12
Llama 3.2 3B Instructopen
Meta · text
GA128K2023-12
Llama 3.2 1B Instructopen
Meta · text
GA128K2023-12
Codestral (v25.08)open
Mistral · text, code
GA128K$0.3$0.9$0.45
Command A Plusopen
Cohere · text, image
GA128K
Command A Visionopen
Cohere · text, image
GA128K
Command R7Bopen
Cohere · text
GA128K$0.037$0.15$0.066
Command R+ (08-2024)open
Cohere · text
GA128K$2.50$10$4.38
Command R (08-2024)open
Cohere · text
GA128K$0.15$0.6$0.262
Amazon Nova Micro
Amazon · text
GA128K$0.035$0.14$0.0612024-10
Sonar
Perplexity · text, image
GA128K$1$1$1
Sonar Reasoning Pro
Perplexity · text
GA128K$2$8$3.50
Qwen-Max (Qwen2.5-Max)
Alibaba · text
GA33K$1.60$6.40$2.80
Moonshot v1 32K
Moonshot · text
GA33K$1$3$1.50
Moonshot v1 8K
Moonshot · text
GA8K$0.2$2$0.65
Command A Translateopen
Cohere · text
GA8K

Blended = 0.75 × input + 0.25 × output $/M tokens (a fair single-number cost proxy). Click any header to sort.

FAQ

LLM context windows

Which LLM has the largest context window?

Llama 4 Scout (17B-16E Instruct) (Meta) has the largest context window of the generally-available models we track, at 10M tokens. GPT-5.6 Sol is next at 1.1M tokens. Context windows are read from official provider documentation and refreshed daily.

What is a context window?

A context window is the maximum number of tokens — roughly ¾ of a word each — that a model can consider at once, spanning your prompt, any retrieved documents or code, the conversation history and the model’s own reply. A larger window lets you feed in more source material (whole codebases, long PDFs, long chats) in a single request without chunking.

Does a bigger context window always mean better answers?

No. Effective recall often degrades well before the stated limit — a model with a 1M-token window may still miss details buried in the middle of a very long prompt. Bigger windows also cost more, because you pay per input token. Match the window to what you actually need to fit, then verify recall on your own long-context task.

How many models offer a 1M+ token context window?

32 of the generally-available models we track offer a context window of at least 1M tokens, and the cheapest of them is Qwen-Flash (Alibaba) at a blended $0.138 per 1M tokens. Sort the table by context to see them ranked, or by blended price to find the cheapest large-context option.