Context Token Budget Planner
Plan what fits in a model's context window. Add the pieces of your prompt, reserve room for the output, and see whether the total fits or overflows. All counts are estimates and stay in your browser.
Budget
Input tokens used
9,200
+ reserved output
128,000
Context window
1,000,000
Estimated 14% of the window used. Fits, with 862,800 tokens to spare.
Per-item breakdown
| Item | Est. tokens | Share of window |
|---|---|---|
| System prompt | 1,200 | 0.1% |
| Retrieved context | 8,000 | 0.8% |
| Reserved for output | 128,000 | 12.8% |
Token counts are an estimate based on roughly 4 characters per token, not a real tokenizer. The actual count differs by model, so treat the fit or overflow as a rough guide and leave headroom.
Common questions
- What does the planner check?
- It adds up the estimated tokens for every context item plus the tokens you reserve for the output, then compares that total to the selected model's context window. It shows the percentage used and whether the plan fits or overflows.
- How are the token counts estimated?
- When you paste text, the planner uses a heuristic of roughly 4 characters per token (characters divided by 4, rounded up). It does not run a real tokenizer, so the counts are estimates and the actual figures differ by model.
- Why reserve tokens for output?
- The context window is shared by input and output, so tokens spent on a long response are not available for your prompt. Reserving output room, seeded from the model's max output, keeps the plan honest. Leave extra headroom because the estimate is rough.