Tool · Enterprise AI

LLM API Cost Calculator: What Your Tasks Cost on OpenAI, Claude, Gemini, Grok and Mistral

Pick a task your organisation actually does, set the volume, and see what it costs on each vendor's comparable model at published API prices, plus when a per user seat is cheaper than paying per use. Nothing you enter leaves your browser.

Prices reviewed 5 October 2026

The short answer

At published list prices reviewed 5 October 2026, the cheapest fast model in this comparison is GPT-6 Luna from OpenAI at $0.10 input and $0.50 output per million tokens, and the cheapest most capable model is Mistral Medium 3.5 at $1.50 input and $7.50 output. The most expensive most capable models are GPT-6 Astra and Claude Fable 5.1, at $10 input and $50 output. Price per token is not cost per completed task, so test the cheapest few on your own work before choosing.

Your task

Each task carries an indicative size in tokens, which you can change below. Models are compared like for like within a tier.

Set automatically for each task. Change it to compare another tier.
Used for the seat comparison. Set to 0 to hide it.
Adjust the task size
Reasoning models bill their thinking as output. Measure it on your own tasks.
Discounts and currency
Repeated system prompts and documents. Models with no published cached rate get no discount.
Vendors bill in dollars. Indicative rate.
Choose specific models

By default each vendor's model for the selected tier is used. Pick another to compare. The token adjustment lets you reflect measured differences in how each vendor counts tokens for your text.

Turn the list price into a real number

List prices are the starting point. What you actually pay depends on the models you route to, caching, volume commitments and the terms in the contract. Tell us the work you want AI to do and the volumes, and we will tell you which models fit, what it should really cost, and how to keep it predictable. We have no model of our own to sell.

Prefer email? Write to hello@c4cgroup.co.uk.

How the calculation works

For each vendor, the tool takes the task's input and output tokens, multiplies by the number of model calls, and prices them at that model's published rate per million tokens. Cached input is priced at the vendor's published cached rate, and batch processing at each model's published batch discount: 50 percent for most models, 20 percent for some of xAI's, and none where the vendor says there is none or publishes no rate. Hidden reasoning is added as output, because that is how it is billed. The result is converted to pounds at the rate you set.

The comparison is like for like by tier. For each vendor, "most capable" is its most capable generally available model, "balanced" its recommended production model, and "fast" its cheapest general purpose model, using each vendor's own positioning. The ranking comes purely from the prices below. A cheaper model is only cheaper if it does the job: what matters is cost per completed task, not price per token, so test quality on your own work before choosing.

Two things it cannot know. Tokenisers differ between vendors, so the same text produces different token counts, and Anthropic, for example, says its newer Claude models produce approximately 30 percent more tokens than its earlier ones for the same text. And the same models bought through Microsoft Foundry, Amazon Bedrock or Google Cloud can carry different prices, including a premium for regional processing. Our comparison of Foundry, Bedrock and Vertex AI covers that.

LLM API pricing compared by tier

What 1,000 support ticket answers cost on each vendor's model in each tier, at list prices, assuming 6,000 input and 400 output tokens per answer and no caching or batch discount. US dollars, as the vendors bill.

Fast tierPer 1,000 tasks
GPT-6 LunaOpenAI$0.80
Mistral Small 4Mistral$1.14
Gemini 3.5 Flash LiteGoogle (Gemini)$2.80
Claude Haiku 4.5Anthropic (Claude)$8.00
Grok 4.3xAI (Grok)$8.50
Balanced tierPer 1,000 tasks
Mistral Large 3Mistral$3.60
Gemini 3.8 FlashGoogle (Gemini)$6.00
Grok 4.3xAI (Grok)$8.50
GPT-6.1 SolOpenAI$16.00
Claude Sonnet 5.5Anthropic (Claude)$16.00
Most capable tierPer 1,000 tasks
Mistral Medium 3.5Mistral$12.00
Grok 4.7xAI (Grok)$14.40
Gemini 3.1 Pro (preview)Google (Gemini)$16.80
GPT-6 AstraOpenAI$80.00
Claude Fable 5.1Anthropic (Claude)$80.00

Published prices used

US dollars per million tokens, from each vendor's pricing page, reviewed 5 October 2026. xAI has no cheaper general purpose model than Grok 4.3, so it is used for both the balanced and fast tiers.

ModelInputCached inputOutputBatch discount
OpenAI
GPT-6 AstraMost capable$10$1$5050%
GPT-6.1 SolBalanced$2$0.10$1050%
GPT-6 LunaFastest and cheapest$0.10$0.010$0.5050%
Anthropic (Claude)
Claude Fable 5.1Most capable generally available$10$0.25$5050%
Claude Opus 5.5Complex reasoning$4$0.20$2050%
Claude Sonnet 5.5Balanced$2$0.20$1050%
Claude Haiku 4.5Fastest$1$0.10$550%
Google (Gemini)
Gemini 3.1 Pro (preview)Most capable; preview; prompts up to 200k tokens$2$0.20$1250%
Gemini 3.8 FlashIntroductory price to 31 December 2026, then 1.50 and 7.50$0.75$0.075$3.7550%
Gemini 3.5 FlashBalanced$1.50$0.15$9not published
Gemini 3.5 Flash LiteFastest$0.30not published$2.5050%
xAI (Grok)
Grok 4.7Most capable; no batch discount; prompts under 200k tokens$2$0.50$6none
Grok 4.3Cheapest general purpose; used for balanced and fast; 20% batch discount$1.25$0.20$2.5020%
Grok Build 0.1Positioned as faster and cheaper; xAI does not state its intended use$1$0.20$2not published
Mistral
Mistral Medium 3.5Described by Mistral as its most powerful$1.50not published$7.5050%
Mistral Large 3Balanced$0.50not published$1.5050%
Mistral Small 4Fastest$0.15not published$0.6050%

Seat prices used for the seat comparison: ChatGPT Business $20 per user per month (published, annual billing); Claude Team $20 per user per month (published, annual billing); Microsoft Copilot $30 per user per month (widely reported list price, add on to microsoft 365).

Frequently asked questions

What is the cheapest LLM API?

At list prices reviewed 5 October 2026, the cheapest fast model in this comparison is GPT-6 Luna at $0.10 input and $0.50 output per million tokens, and the cheapest most capable model is Mistral Medium 3.5 at $1.50 input and $7.50 output. The cheapest per token is not always cheapest per completed task, because a model that needs more attempts or more checking costs more overall.

Is a more expensive AI model worth paying for?

Sometimes. What matters is cost per completed task, not price per token. A more capable model that finishes complex work in fewer steps, with fewer retries and less human correction, can cost less overall, while for simple, high volume work a fast model is usually enough. Measure cost per completed task in a pilot before deciding.

How much does the Claude API cost?

At list prices reviewed 5 October 2026, Claude Haiku 4.5 costs $1 input and $5 output, Claude Sonnet 5.5 $2 and $10 and Claude Fable 5.1 $10 and $50, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.

How much does the OpenAI API cost?

At list prices reviewed 5 October 2026, GPT-6 Luna costs $0.10 input and $0.50 output, GPT-6.1 Sol $2 and $10 and GPT-6 Astra $10 and $50, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.

How much does the Gemini API cost?

At list prices reviewed 5 October 2026, Gemini 3.5 Flash Lite costs $0.30 input and $2.50 output, Gemini 3.8 Flash $0.75 and $3.75 (introductory price to 31 December 2026, then 1.50 and 7.50) and Gemini 3.1 Pro (preview) $2 and $12, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.

How much does the Grok API cost?

At list prices reviewed 5 October 2026, Grok 4.3 costs $1.25 input and $2.50 output and Grok 4.7 $2 and $6, all per million tokens. Cached input is cheaper on models with a published cached rate. Use the calculator above to see what a specific task costs on each.

How much does the Mistral API cost?

At list prices reviewed 5 October 2026, Mistral Small 4 costs $0.15 input and $0.60 output, Mistral Large 3 $0.50 and $1.50 and Mistral Medium 3.5 $1.50 and $7.50, all per million tokens. The vendor publishes no specific cached input rates for these models. Use the calculator above to see what a specific task costs on each.

Which is cheaper, ChatGPT, Claude, Gemini or Grok?

It depends on the task and the tier of model. At published API list prices, the cheapest option changes between fast, balanced and most capable models, and with how much of the work is input versus output. Compare like for like on your own task, then test quality, because a cheaper model that needs a second attempt costs more.

How much does an AI task cost?

Usually fractions of a penny to a few pence per task at API prices. A short classification on a fast model costs a tiny fraction of a penny, while a multi step agent task on a most capable model can cost far more. Volume is what turns small numbers into large bills, so multiply by tasks per month.

Is a ChatGPT or Claude seat cheaper than the API?

For light, occasional users, paying per use through the API is often cheaper. For heavy users, a seat at around 20 dollars per user per month usually wins, and it includes the chat interface and tools. The calculator shows the break even number of tasks per person for each model.

Why do token counts differ between models?

Each vendor uses its own tokeniser, so the same text becomes a different number of tokens. Anthropic, for example, says its newer Claude models produce approximately 30 percent more tokens for the same text than its earlier ones. Measure token counts on your own text for each model you shortlist.

How can we cut AI model costs?

Use the smallest model that does the job well, cache repeated context, use batch processing for work that is not urgent, and route each request to the right model. Most vendors publish around 50 percent off for batch work, though xAI offers 20 percent on some models and none on others, and large discounts for cached input. An AI gateway can apply caching and routing across vendors.