Tool · AI Infrastructure Advisory

AI Inference Cost Calculator

At what volume does owning your inference infrastructure beat paying per token? Enter your usage, pick a model class, and see the API cost against the owned cost over 36 months, including the crossover point if there is one. This tool is deliberately even handed: if the numbers say stay on the API, it says so. No sign up, nothing leaves your browser.

Assumptions last reviewed August 2026

Your inference workload

The calculator runs entirely in your browser, so none of these figures are sent anywhere.

Input plus output combined. Your API invoice usually states it.
Selecting a class sets indicative prices and capacity below. Overwrite them with your own figures.
Be honest. Utilisation is the number that decides this comparison, and most owned AI estates run far below their theoretical capacity.
Adjust assumptions

Defaults are indicative estimates in pounds, set by the model class above and editable here. For a real picture, use the blended rate from your API invoice and a hardware quote from your own sizing.

Size it properly before you decide

This tool compares one owned configuration against one API rate, which is enough to see the shape of the decision but not enough to make it. A sizing conversation goes further: we model your actual models, concurrency and growth, price a validated configuration through our Dell, HPE and distribution relationships, and compare it honestly against staying on the API. If the API is still the cheaper home for your workload, that is exactly what we will tell you. The infrastructure conversation is only worth having when the numbers are real.

Prefer email? Reach us directly at hello@c4cgroup.co.uk.

How the calculation works

The API side is simple: your monthly token volume multiplied by a blended price per million tokens, grown at your chosen rate, accumulated over 36 months.

The owned side is where honesty matters. We amortise the hardware over 36 months, add power, cooling and support as an annual percentage of the hardware cost, and add a staffing allowance, because someone has to run it and that cost is real even when it is a fraction of an existing team. Capacity is the node's theoretical monthly throughput scaled by your utilisation estimate, and if your volume outgrows one node, the model adds nodes and their costs as needed.

Utilisation deserves its own sentence: it is the single number that decides most of these comparisons. A node running at 30 percent utilisation costs three times as much per token as the same node at 90 percent, and most owned AI estates run nearer the former. The API's great advantage is that you pay for exactly what you use; owned infrastructure only wins once volume is high enough, and steady enough, to keep the hardware genuinely busy.

Two things this simple model deliberately leaves out, both of which favour ownership in specific cases rather than in general: data constraints that price the API at infinity for regulated workloads (see our Private AI Pod), and negotiated API discounts at committed volume, which cut the other way. The sizing conversation prices both.

Default assumptions

The defaults below are indicative estimates, last reviewed in August 2026, and every one of them is editable in the tool. They exist to make the comparison honest in shape, not to price your deployment: real API rates vary by provider and commitment, and real hardware costs come from a sizing, not a web page.

Model classBlended API, £ per million tokensNode hardwareNode capacity, M tokens / month
Smallaround £0.80around £60,000around 20,000
Mediumaround £4.00around £250,000around 6,000
Largearound £18.00around £600,000around 1,500

Plus in every case: staff and operations at £40,000 per year and power, cooling and support at 15 percent of hardware cost per year, both editable. Capacity figures assume served open weight models with standard optimisations, before your utilisation estimate is applied.

Please read this as a shape, not a quote. Every figure here is an editable estimate, and the honest version of this comparison needs your models, your concurrency and a real hardware sizing. The tool exists to tell you whether the question is worth asking properly. If your result says the API wins, believe it, that is not a sales funnel talking. If it says ownership wins, verify it with a sizing before spending anything.

Frequently asked questions

When is owned infrastructure cheaper than the API?

When volume is high, steady and genuinely utilises the hardware. As a shape: owned cost is fixed per month while API cost scales with usage, so there is a volume above which ownership wins. At typical enterprise utilisation that volume is higher than most organisations expect, which is why the honest first answer for low and unpredictable usage is stay on the API.

What does it cost to run an LLM on premises?

Three components: the hardware amortised over its life, running costs of roughly 10 to 20 percent of hardware per year for power, cooling and support, and a real staffing allowance. Divided by the tokens you actually serve, not the tokens the hardware could serve, that gives your true cost per million tokens, and utilisation is what moves it most. Performance per watt is the quiet second lever: two configurations at the same sticker price can differ sharply in sustained power draw, and over a three year life that difference is real money the headline number hides.

What utilisation should I assume for AI hardware?

Lower than you would like. Internal workloads follow office hours, demand is spiky, and capacity is bought ahead of growth, so sustained utilisation of 30 percent is common and 60 percent is good. Assuming 90 percent in a business case flatters ownership; sensitivity check the decision at half whatever you assumed.

Should we move from the API to our own infrastructure?

Move when at least one of three things is true: your volume is high and steady enough that the owned cost per token is clearly lower, your data cannot leave your estate, or committed API spend is approaching what a sized deployment would cost. Move for none of those reasons and you have bought depreciation and an ops burden to serve a workload the API was already handling.

Do data residency requirements change the calculation?

Completely. If regulation, contracts or policy mean prompts and documents cannot go to an external API, the comparison is no longer about price and the question becomes how to run private inference well. That is a design problem rather than a calculator problem, and it is exactly what our Private AI Pod validated design exists for.

Can we mix the API and our own infrastructure?

Yes, and mature estates usually do: steady, high volume or sensitive workloads served on owned infrastructure, spiky and experimental workloads on the API, with routing between them. The mix shifts as usage matures, which is why the decision should be revisited on real numbers rather than made once and defended forever.