What AI actually costs

Every number below is the vendor's own published rate, per million tokens, in US dollars. The same job can cost several times less on a different model — that is not a discount we invented, it is the same work run on a different engine.

For reference only · prices change without notice · your quote is what counts Every figure below is taken from the vendor’s own official pricing page (linked per row), but vendors change their rates at any time and price differently by region and by context tier. This table is not a quote and is not a commitment of any kind — the final quote we give you is what governs.

Note

Model pricing keeps moving — new releases, price cuts, currency swings. Every row is the vendor's own rate as checked on 2026-08-10, not a fixed discount we made up. Before quoting a client, confirm against the vendor’s live pricing page. Three rows were re-read against the vendor’s own page on 2026-09-05: Claude Sonnet 5 (its launch introductory rate became the standard price and the scheduled increase was cancelled, so the figure is unchanged), Qwen3-Max and Qwen3-Plus (corrected to the Singapore/International tiered rates). Every other row still carries the 2026-08-10 figure.

US frontier models
Model Input /1M Cached /1M Output /1M Source
Claude Opus 5Anthropic $5.00$0.50$25.00 anthropic ↗
Claude Sonnet 5Anthropic · the launch introductory rate is now the standard price; the increase scheduled for 2026-09-01 was cancelled $2.00$0.20$10.00 anthropic ↗
Claude Haiku 4.5Anthropic $1.00$0.10$5.00 anthropic ↗
GPT-5.6 SolOpenAI · flagship tier $5.00$0.50$30.00 openai ↗
GPT-5.6 TerraOpenAI · standard tier $2.00$0.20$12.00 openai ↗
Gemini 3.1 ProGoogle · preview, ≤200k context (above that: $4.00 / $0.40 / $18.00) $2.00$0.20$12.00 google ↗
Chinese models
Model Input /1M Cached /1M Output /1M Source
DeepSeek V4 ProDeepSeek · vendor flags a price increase coming soon $0.435$0.0036$0.87 deepseek ↗
DeepSeek V4 FlashDeepSeek · vendor flags a price increase coming soon $0.14$0.0028$0.28 deepseek ↗
Kimi K3Moonshot · flagship tier, tax not included $3.00$0.30$15.00 kimi ↗
GLM-5.2Zhipu · flagship tier $1.40$0.26$4.40 zhipu ↗
GLM-4.7Zhipu $0.60$0.11$2.20 zhipu ↗
Qwen3-MaxAlibaba Cloud · Singapore/International · ≤32k context (32k-128k: $2.40 / $12.00; 128k-256k: $3.00 / $15.00) $1.20≈10% of input$6.00 alibaba ↗
Qwen3-PlusAlibaba Cloud · Singapore/International · ≤256k context · limited-time 20% off with no published end date (list $0.40 / $1.60, discounted on both input and output; 256k-1M lists at $1.20 / $4.80) $0.32≈10% of input$1.28 alibaba ↗

What that means for your monthly bill

Pick the shape of your team, or type your own numbers. A developer running an AI coding agent through a working day is the heaviest realistic case — that is what these presets assume. The maths uses the list prices above; the cached-input column is a separate saving and is left out of it.

millions
millions
Run the same month on… Per month Per year vs. Claude Opus 5

This figure is an estimate worked out from the published list prices in the table above. It is for reference only and is not a quote. Vendors change their rates at any time, and your real usage, context lengths and cache hit rate will all move the result — the final quote is what governs.

Trying it out

There is no contract to sign first and no minimum top-up. Put in $1 and start — run your own work through it, look at the output and the bill, then decide whether to move more across. Unused balance does not expire.

Where this breaks down

Not every Chinese model is cheap — Kimi K3 lists at $3.00 / $15.00, which is dearer than Claude Sonnet 5 at $2.00 / $10.00. The saving comes from picking the right model for the job, not from the vendor's flag.

Longer context can jump to a higher tier — Gemini splits at 200k, Qwen at 32k/256k/1M; the rows above show the lowest tier. Anthropic and OpenAI charge one flat rate across the full context.

Prices move in both directions — Claude Sonnet 5 was scheduled to rise on 2026-09-01, and instead the vendor made the introductory rate permanent and cancelled the increase; DeepSeek's own page flags an increase coming. A stale table can therefore quote too high as easily as too low. After the checked date, defer to the vendor's live page.

Capability is not identical — depending on the task (coding, reasoning, image, video) expect roughly 80–90% of the frontier US models. For a great deal of real work that is the whole job; for the hardest 10% it is not, and we will say so rather than sell you the wrong thing. That 80–90% is our own estimate, not a published benchmark — unlike every price above, there is no vendor page to link for it.

Caching cuts both sides — unevenly — a cached input token costs $0.50 on Claude Opus 5 and $0.0036 on DeepSeek V4 Pro. On cache-heavy agent workloads the gap widens rather than narrows. The calculator above leaves that column out, so what it shows is the conservative figure.

All figures are each vendor's own published rate, checked 2026-08-10 — the Claude Sonnet 5, Qwen3-Max and Qwen3-Plus rows re-read 2026-09-05, as noted above — linked per row. Currency, regional pricing, and enterprise contracts can differ from these — this is the public standard API rate, not what we quote clients.

Already paying for one of these? Tell us your usage and we'll work out what the same job costs on a Chinese model instead.

See the letter we send →