AI Model Comparison: Context Window & Price
Compare the frontier models used in coding agents by context window, max output, and per-token price. Every row is checked against the vendor's own docs; sort by any column.
| GPT-5.6 (Sol) | OpenAI | 1,049k | 128k | $5 | $30 |
| Gemini 3.6 Flash | 1,049k | 66k | $1.5 | $7.5 | |
| Claude Opus 5 | Anthropic | 1,000k | 128k | $5 | $25 |
| Claude Sonnet 5 | Anthropic | 1,000k | 128k | $2 | $10 |
| GLM-5.2 | Z.ai | 1,000k | 128k | $1.4 | $4.4 |
| DeepSeek V4 (flash) | DeepSeek | 1,000k | 384k | $0.14 | $0.28 |
| Claude Haiku 4.5 | Anthropic | 200k | 64k | $1 | $5 |
| GLM-4.6 | Z.ai | 200k | 128k | $0.6 | $2.2 |
Context and max output are token counts; prices are per-1M-token API list rates. Last checked 2026-08-11. Models whose specs are not confirmable at a primary source are left off rather than shown unverified.
Common questions
- Which AI model has the largest context window?
- Among the models here, the largest context windows are around 1 million tokens (Claude Opus and Sonnet at 1,000,000; the GPT and Gemini flagships at 1,048,576). The context window is the total tokens a model can consider at once, input plus output.
- What is the difference between context window and max output?
- The context window is the total tokens the model can take in for one request; max output is the largest response it can generate. A 1M-token window with 128k max output means you can send a very large prompt but the reply is capped at 128k tokens.
- How current are these numbers?
- Each row is checked against the vendor's own documentation on the date shown, and a freshness test fails the build if a row goes stale. Models whose specs could not be confirmed at a primary source are left off rather than shown unverified.