Quick presets
Input tokens
0
~0 char-based
Output tokens
0
0 estimated
Total tokens
0
0 combined
Cheapest single call
<$0.0001
DeepSeek R1
| Model | Input cost | Output cost | Single call | Monthly | Context |
|---|---|---|---|---|---|
GPT-4o OpenAI | <$0.0001 $2.5/M tok | <$0.0001 $10/M tok | <$0.0001 | <$0.0001 100/day × 30 | ✓ 128K |
GPT-4.1 OpenAI | <$0.0001 $2/M tok | <$0.0001 $8/M tok | <$0.0001 | <$0.0001 100/day × 30 | ✓ 1M |
Claude Opus 4 Anthropic | <$0.0001 $15/M tok | <$0.0001 $75/M tok | <$0.0001 | <$0.0001 100/day × 30 | ✓ 200K |
Claude Sonnet 4 Anthropic | <$0.0001 $3/M tok | <$0.0001 $15/M tok | <$0.0001 | <$0.0001 100/day × 30 | ✓ 200K |
Gemini 2.5 Pro Google | <$0.0001 $1.25/M tok | <$0.0001 $10/M tok | <$0.0001 | <$0.0001 100/day × 30 | ✓ 1M |
Llama 4 Scout Meta / Together AI | <$0.0001 $0.18/M tok | <$0.0001 $0.59/M tok | <$0.0001 | <$0.0001 100/day × 30 | ✓ 131K |
DeepSeek R1cheapestpriciest DeepSeek | <$0.0001 $0.55/M tok | <$0.0001 $2.19/M tok | <$0.0001 | <$0.0001 100/day × 30 | ✓ 128K |
Pricing reference (May 2026)
Prices are list prices as of May 2026. Llama 4 pricing via Together AI. Actual costs may vary with batch discounts, cached tokens, or on-premise hosting.
Tokenization is the fundamental unit of cost for every large language model API. GPT-4o, Claude, Gemini, and competitors all price by the million tokens — but the ratio of input to output pricing and the absolute price per million tokens vary by more than 50x across the market. This tool lets you estimate cost before you write a single line of code: paste text, set an output length expectation, and see side-by-side costs for seven major models including budget options like DeepSeek R1 and Llama 4.
Token estimation uses a blended heuristic: the well-known 4-characters-per-token approximation (accurate for English prose) combined with a word-count approach (1.33 tokens per word, accounting for punctuation and subword splits). The two estimates are blended 60/40 in favour of the character-based method, which is more reliable for code and mixed content. Cost is computed as (tokens / 1,000,000) × price_per_million, separately for input and output, then summed. The monthly projector multiplies single-call cost by requests/day × 30.
A developer planning a RAG system pastes a sample document chunk and sees that GPT-4.1 costs 4x less than Claude Opus 4 for the same retrieval query.
A product team budgeting a translation feature enters 1,400 input tokens and 1,400 output tokens, sets 500 requests/day, and compares monthly costs across models before choosing Gemini 2.5 Pro.
An indie hacker building a code-review bot uses the Code Review 500 Lines preset and discovers DeepSeek R1 cuts monthly costs from $180 to $12 vs. GPT-4o at 1,000 requests/day.
A solutions architect demonstrating AI ROI to a client copies a 2,000-word specification document and shows the exact per-call and monthly costs to justify model selection.
Paste your text in the input area or switch to manual token entry
Adjust the output ratio slider to match your expected response length
View the per-call and monthly cost breakdown across all 7 models
Use the monthly projector to enter requests per day and see scaled costs
Click a preset scenario to load a common use-case token profile instantly
We packaged the .claude config that runs this site — 26 specialist agents, 14 workflow skills, and 6 rule files from 31 real production incidents. From $9.
Get the Claude Code Production Pack — $29