Pick a task your organisation actually does, set the volume, and see what it costs on each vendor's comparable model at published API prices, plus when a per user seat is cheaper than paying per use. Nothing you enter leaves your browser.
Prices reviewed 5 October 2026
At published list prices reviewed 5 October 2026, the cheapest fast model in this comparison is GPT-6 Luna from OpenAI at $0.10 input and $0.50 output per million tokens, and the cheapest most capable model is Mistral Medium 3.5 at $1.50 input and $7.50 output. The most expensive most capable models are GPT-6 Astra and Claude Fable 5.1, at $10 input and $50 output. Price per token is not cost per completed task, so test the cheapest few on your own work before choosing.
Each task carries an indicative size in tokens, which you can change below. Models are compared like for like within a tier.
By default each vendor's model for the selected tier is used. Pick another to compare. The token adjustment lets you reflect measured differences in how each vendor counts tokens for your text.
List prices are the starting point. What you actually pay depends on the models you route to, caching, volume commitments and the terms in the contract. Tell us the work you want AI to do and the volumes, and we will tell you which models fit, what it should really cost, and how to keep it predictable. We have no model of our own to sell.
Prefer email? Write to hello@c4cgroup.co.uk.
For each vendor, the tool takes the task's input and output tokens, multiplies by the number of model calls, and prices them at that model's published rate per million tokens. Cached input is priced at the vendor's published cached rate, and batch processing at each model's published batch discount: 50 percent for most models, 20 percent for some of xAI's, and none where the vendor says there is none or publishes no rate. Hidden reasoning is added as output, because that is how it is billed. The result is converted to pounds at the rate you set.
The comparison is like for like by tier. For each vendor, "most capable" is its most capable generally available model, "balanced" its recommended production model, and "fast" its cheapest general purpose model, using each vendor's own positioning. The ranking comes purely from the prices below. A cheaper model is only cheaper if it does the job: what matters is cost per completed task, not price per token, so test quality on your own work before choosing.
Two things it cannot know. Tokenisers differ between vendors, so the same text produces different token counts, and Anthropic, for example, says its newer Claude models produce approximately 30 percent more tokens than its earlier ones for the same text. And the same models bought through Microsoft Foundry, Amazon Bedrock or Google Cloud can carry different prices, including a premium for regional processing. Our comparison of Foundry, Bedrock and Vertex AI covers that.
What 1,000 support ticket answers cost on each vendor's model in each tier, at list prices, assuming 6,000 input and 400 output tokens per answer and no caching or batch discount. US dollars, as the vendors bill.
| Fast tier | Per 1,000 tasks |
|---|---|
| GPT-6 LunaOpenAI | $0.80 |
| Mistral Small 4Mistral | $1.14 |
| Gemini 3.5 Flash LiteGoogle (Gemini) | $2.80 |
| Claude Haiku 4.5Anthropic (Claude) | $8.00 |
| Grok 4.3xAI (Grok) | $8.50 |
| Balanced tier | Per 1,000 tasks |
|---|---|
| Mistral Large 3Mistral | $3.60 |
| Gemini 3.8 FlashGoogle (Gemini) | $6.00 |
| Grok 4.3xAI (Grok) | $8.50 |
| GPT-6.1 SolOpenAI | $16.00 |
| Claude Sonnet 5.5Anthropic (Claude) | $16.00 |
| Most capable tier | Per 1,000 tasks |
|---|---|
| Mistral Medium 3.5Mistral | $12.00 |
| Grok 4.7xAI (Grok) | $14.40 |
| Gemini 3.1 Pro (preview)Google (Gemini) | $16.80 |
| GPT-6 AstraOpenAI | $80.00 |
| Claude Fable 5.1Anthropic (Claude) | $80.00 |
US dollars per million tokens, from each vendor's pricing page, reviewed 5 October 2026. xAI has no cheaper general purpose model than Grok 4.3, so it is used for both the balanced and fast tiers.
| Model | Input | Cached input | Output | Batch discount |
|---|---|---|---|---|
| OpenAI | ||||
| GPT-6 AstraMost capable | $10 | $1 | $50 | 50% |
| GPT-6.1 SolBalanced | $2 | $0.10 | $10 | 50% |
| GPT-6 LunaFastest and cheapest | $0.10 | $0.010 | $0.50 | 50% |
| Anthropic (Claude) | ||||
| Claude Fable 5.1Most capable generally available | $10 | $0.25 | $50 | 50% |
| Claude Opus 5.5Complex reasoning | $4 | $0.20 | $20 | 50% |
| Claude Sonnet 5.5Balanced | $2 | $0.20 | $10 | 50% |
| Claude Haiku 4.5Fastest | $1 | $0.10 | $5 | 50% |
| Google (Gemini) | ||||
| Gemini 3.1 Pro (preview)Most capable; preview; prompts up to 200k tokens | $2 | $0.20 | $12 | 50% |
| Gemini 3.8 FlashIntroductory price to 31 December 2026, then 1.50 and 7.50 | $0.75 | $0.075 | $3.75 | 50% |
| Gemini 3.5 FlashBalanced | $1.50 | $0.15 | $9 | not published |
| Gemini 3.5 Flash LiteFastest | $0.30 | not published | $2.50 | 50% |
| xAI (Grok) | ||||
| Grok 4.7Most capable; no batch discount; prompts under 200k tokens | $2 | $0.50 | $6 | none |
| Grok 4.3Cheapest general purpose; used for balanced and fast; 20% batch discount | $1.25 | $0.20 | $2.50 | 20% |
| Grok Build 0.1Positioned as faster and cheaper; xAI does not state its intended use | $1 | $0.20 | $2 | not published |
| Mistral | ||||
| Mistral Medium 3.5Described by Mistral as its most powerful | $1.50 | not published | $7.50 | 50% |
| Mistral Large 3Balanced | $0.50 | not published | $1.50 | 50% |
| Mistral Small 4Fastest | $0.15 | not published | $0.60 | 50% |
Seat prices used for the seat comparison: ChatGPT Business $20 per user per month (published, annual billing); Claude Team $20 per user per month (published, annual billing); Microsoft Copilot $30 per user per month (widely reported list price, add on to microsoft 365).
At list prices reviewed 5 October 2026, the cheapest fast model in this comparison is GPT-6 Luna at $0.10 input and $0.50 output per million tokens, and the cheapest most capable model is Mistral Medium 3.5 at $1.50 input and $7.50 output. The cheapest per token is not always cheapest per completed task, because a model that needs more attempts or more checking costs more overall.
Sometimes. What matters is cost per completed task, not price per token. A more capable model that finishes complex work in fewer steps, with fewer retries and less human correction, can cost less overall, while for simple, high volume work a fast model is usually enough. Measure cost per completed task in a pilot before deciding.
At list prices reviewed 5 October 2026, Claude Haiku 4.5 costs $1 input and $5 output, Claude Sonnet 5.5 $2 and $10 and Claude Fable 5.1 $10 and $50, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.
At list prices reviewed 5 October 2026, GPT-6 Luna costs $0.10 input and $0.50 output, GPT-6.1 Sol $2 and $10 and GPT-6 Astra $10 and $50, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.
At list prices reviewed 5 October 2026, Gemini 3.5 Flash Lite costs $0.30 input and $2.50 output, Gemini 3.8 Flash $0.75 and $3.75 (introductory price to 31 December 2026, then 1.50 and 7.50) and Gemini 3.1 Pro (preview) $2 and $12, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.
At list prices reviewed 5 October 2026, Grok 4.3 costs $1.25 input and $2.50 output and Grok 4.7 $2 and $6, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.
At list prices reviewed 5 October 2026, Mistral Small 4 costs $0.15 input and $0.60 output, Mistral Large 3 $0.50 and $1.50 and Mistral Medium 3.5 $1.50 and $7.50, all per million tokens. The vendor publishes no specific cached input rates for these models. Use the calculator above to see what a specific task costs on each.
It depends on the task and the tier of model. At published API list prices, the cheapest option changes between fast, balanced and most capable models, and with how much of the work is input versus output. Compare like for like on your own task, then test quality, because a cheaper model that needs a second attempt costs more.
Usually fractions of a penny to a few pence per task at API prices. A short classification on a fast model costs a tiny fraction of a penny, while a multi step agent task on a most capable model can cost far more. Volume is what turns small numbers into large bills, so multiply by tasks per month.
For light, occasional users, paying per use through the API is often cheaper. For heavy users, a seat at around 20 dollars per user per month usually wins, and it includes the chat interface and tools. The calculator shows the break even number of tasks per person for each model.
Each vendor uses its own tokeniser, so the same text becomes a different number of tokens. Anthropic, for example, says its newer Claude models produce approximately 30 percent more tokens for the same text than its earlier ones. Measure token counts on your own text for each model you shortlist.
Use the smallest model that does the job well, cache repeated context, use batch processing for work that is not urgent, and route each request to the right model. Most vendors publish around 50 percent off for batch work, though xAI offers 20 percent on some models and none on others, and large discounts for cached input. An AI gateway can apply caching and routing across vendors.