Loading…
Loading…
Price an agent fleet per run, day and month on every major model at once
GPT-6.1 Sol prices USD per 1M tokens
One-fifth of Astra token prices. API model id gpt-6.1-sol.
A run is one task: triage a ticket, review a PR, answer a lead.
Every tool call round-trip re-sends the context, so each one is a billed call.
System prompt + tools + history + retrieved context, as sent on each call.
Include reasoning tokens if the model bills them as output.
Share of input billed at the cached-read price. Stable system prompts cache well.
Always-on agents run 30. Business-hours agents run about 22.
5 agents × 20 runs/day on GPT-6.1 Sol
$464/ month
47% of the bill is output tokens. At 60% cache hits, each step costs $0.013.
5 of 5 models cost less than the $12,000 of human time your agents save each month.
Quality differs between these models. The cheapest bill only wins if the model finishes the run in the same number of steps.
Prices are vendor list prices and aggregator listings checked on 8 Oct 2026; they change without notice. Excluded: tool and API call fees, cache-write fees, long-context surcharges (Astra above 272K input tokens, Haiku 5.5 above 100K), batch discounts and taxes. INR uses ₹85 per $1. Counting tokens first? Try the AI token estimator, trim the system prompt with the prompt bloat analyzer, or price one-off calls in the AI model cost calculator.
Run your projects on Claude Code?
We packaged the .claude config that runs this site — 26 specialist agents, 14 workflow skills, and 6 rule files from 31 real production incidents. From $9.
Get the Claude Code Production Pack — $29Always-on agents changed the unit of AI spend. A chat costs one call; an agent that triages tickets all day makes thousands, and each one re-sends its context. This simulator turns a workload (agents, runs, steps, tokens and cache hits) into a bill per step, run, day, month and year, then prices the same workload on GPT-6 Astra, GPT-6.1 Sol, Claude Haiku 5.5, Mistral Large 4 and Gemini 3.5 Flash. A break-even panel sets that bill against the human hours the agents replace.
Every number starts from one model call, which the tool calls a step. Input is billed at a blended rate: the cache-hit share of input tokens at the cached-read price and the rest at the standard input price. Output tokens are billed at the output price. Prices are per 1M tokens, so a step costs (input tokens × blended price + output tokens × output price) ÷ 1,000,000.
A run is steps per run × step cost. Daily cost multiplies by runs per agent per day and by the number of agents; monthly cost multiplies by active days, and yearly cost is twelve months. Tokens per month and runs per month come from the same multiplication, so you can sanity-check them against your provider dashboard.
The comparison chart reruns the identical workload on every preset with whatever prices you have edited, flags the lowest monthly bill as Best value and shows every other model as a multiple of it. The break-even panel multiplies hours saved per agent per day by agents, days and the hourly cost you enter. It reports the net monthly difference, the value returned per dollar or rupee, the hourly cost at which the agents only break even, and the minutes each agent must save per day to cover its bill. INR figures use the site’s fixed USD-to-INR rate, printed under the tool.
An engineering lead deciding whether a 20-agent support triage fleet runs on GPT-6.1 Sol or Claude Haiku 5.5 before asking finance for a monthly budget line.
A solo founder pricing an always-on sales research agent in rupees and checking it still pays back at ₹1,500 an hour of their own time.
A platform team measuring what a better prompt cache is worth by moving the cache-hit share from 20% to 70% on the same workload.
A consultant building a client proposal that puts the agent bill next to the human hours it replaces, with a share card for the deck.
Scope note: Prices are vendor list prices and aggregator listings checked on 8 Oct 2026 and change without notice; Mistral Large 4 also shows a lower preview rate. The simulator excludes tool and API call fees, cache-write fees, long-context surcharges (GPT-6 Astra above 272K input tokens, Claude Haiku 5.5 above 100K), batch discounts and taxes. Token counts are your estimates, and models differ in how many steps a task takes.
Pick a model preset, or choose Custom and type the per-1M-token prices from your provider’s pricing page.
Set the workload: number of agents, runs per agent per day, model calls per run, and input and output tokens per call.
Set the cache-hit share (the part of each prompt billed at the cached price) and how many days a month the agents run.
Read the cost per step, run, day, month and year, then check the comparison chart for the cheapest model on the same workload.
Enter the hourly cost of the person doing this work and the hours each agent saves to see net savings and the break-even point, then share the card.
About the Always-On AI Agent Cost Simulator
Multiply agents × runs per agent per day × model calls per run × active days to get calls per month. Each call costs its input tokens at the input price plus its output tokens at the output price, with the cached share of input billed at the cached-read rate; divide per-1M prices by 1,000,000 first. Example: 36,000 calls a month at 8,000 input tokens (60% cached) and 600 output tokens cost $463.68 on GPT-6.1 Sol at $2 / $10 / $0.10 per 1M tokens.
An agent re-sends its context on every step. A run with 12 tool calls bills the system prompt, tool definitions and growing history 12 times, so input token volume balloons. Prompt caching is the main lever: Claude Haiku 5.5 bills cached reads at $0.01 per 1M tokens against $0.10 uncached, and GPT-6.1 Sol at $0.10 against $2.00.
On list prices checked 8 Oct 2026, Claude Haiku 5.5 is the cheapest preset for prompts up to 100K tokens ($0.10 in / $0.50 out per 1M). On the default workload it costs about $24 a month against $271 for Mistral Large 4, $393 for Gemini 3.5 Flash, $464 for GPT-6.1 Sol and $2,405 for GPT-6 Astra. A lower token price only wins if the model finishes the task in a similar number of steps.
It is the share of each call’s input tokens that the provider serves from its prompt cache and bills at the cached-read price. Stable prefixes such as system prompts and tool schemas cache well; fresh tool output and new user messages do not. If you do not know your rate, read the cached-token field in your API usage logs, or start at 0% for a worst case. Cache-write fees are not included.
No. It prices model tokens only. Web search, code execution, browser sessions, vector database reads and per-call tool fees are billed separately by most providers, as are long-context surcharges (GPT-6 Astra above 272K input tokens, Claude Haiku 5.5 above 100K). Add a margin for those, or fold them into the Custom preset as an effective per-token price.
Score your system prompt for bloat and get a cut list
Open →DeveloperBuilder/critic agent prompts with brakes built in
Open →DeveloperFormat, validate & diff JSON — runs entirely in browser
Open →DeveloperTest regex live — railroad diagrams + plain English explained
Open →