C4C AI Pods · Tier 1

Inference Pod

The tier most enterprises actually need. A validated reference architecture for serving open weight models, internal chatbots, document Q&A and retrieval over internal data, sized from your workload and configured to order. If your AI estate answers questions rather than trains models, this is the design for it.

Designed for

The workloads enterprises actually run

Serving open weight models

Internal chatbots and assistants on open models, served to your own users at enterprise concurrency.

Retrieval over internal data

Document Q&A and retrieval augmented generation over your own estate, where the value is your data, not the model size.

Document and call processing

Summarisation, extraction, classification and transcription pipelines running as steady production workloads.

Established machine learning

The recommendation, scoring and forecasting workloads you already run, consolidated onto shared accelerated infrastructure.

Not designed for

Stated plainly, as every validated design should

Foundation model training. Different discipline, different economics, and almost never the workload an enterprise really has. If someone is sizing you for training, ask them to name the model you will train and the budget line that funds the run.

Regular fine tuning at scale. Adapting open models on your own data as a planned, recurring workload belongs on the Fine Tuning Pod, which carries the accelerator and storage capacity that work genuinely needs.

Regulated sovereign workloads. Where data residency and regulatory control are the design constraint, the Private AI Pod is the right tier: the same inference discipline with the security and residency wrap built in rather than bolted on.

Logical architecture

The stack, as layers

Inference Pod logical architectureLayered block diagram. Users and applications connect to a serving and orchestration layer, which runs on an accelerator and compute layer, backed by a flash storage tier for models and vector indexes and a capacity storage tier for the document estate, all connected by an enterprise Ethernet network.Users and applicationsServing and orchestration layerAccelerator and compute layerFlash tiermodels and vector indexesCapacity tierdocument estateEnterprise Ethernet network  connecting layer

What is in the box

Component classes, not SKUs

Compute class

Enterprise rack servers from our Dell and HPE lines, sized for the serving and retrieval processes around the accelerators, not just the accelerators themselves.

Accelerator class

Inference class GPUs, selected by model size, quantisation and concurrency. Serving models rarely needs the accelerator tier that training does, and the sizing proves it either way.

Storage class

Fast flash for models and vector indexes, capacity tiers for the document estate feeding retrieval. Throughput matters more than headline capacity.

Networking and software layer

Standard enterprise Ethernet, a supported serving and orchestration stack for open weight models, and monitoring that reports utilisation honestly.

No part numbers and no prices on this page, deliberately. Published SKUs go stale, and the right bill of materials depends on your sizing. The specifics are produced in the sizing engagement, configured to order through our Dell, HPE and distribution relationships.

Sizing signals

The Inference Pod fits when the workload sounds like this:

  • You are serving internal users, not training models.
  • The models you plan to serve sit in roughly the 7 to 40 billion parameter range, open weight models run with retrieval.
  • Your workload is questions over your own data: documents, policies, tickets, calls.
  • You can name the models you plan to run, or want help choosing them.
  • Your API bill or your data constraints are telling you to bring inference in house.

Lifecycle

Every pod is wrapped by the IDEAL framework.

Sizing, procurement, deployment, adoption, then refresh and expansion managed proactively rather than at the vendor's renewal cadence. The pod is the starting configuration; the lifecycle is what keeps it right sized as the workload grows.

Questions

Frequently asked questions

What hardware do I need to run an LLM on premises?

It depends on three numbers: the size of the model you will serve, the quantisation you will accept, and the concurrent users you must support. Those determine accelerator memory and count, which determine everything else. A mid sized open weight model serving a few hundred internal users needs far less than vendor sizing conversations tend to assume, which is exactly why we size from the workload first.

Do I need H100 class GPUs for inference?

Usually not. Top tier accelerators are built for training economics; most enterprise inference is served well by inference class GPUs at a fraction of the cost, especially with quantised open weight models. There are genuine exceptions, very large models at high concurrency among them, and the sizing conversation identifies them with numbers rather than instinct.

What does the Inference Pod cost?

No published price, deliberately: the honest number depends on the sizing, and a padded number helps nobody. The sizing conversation produces a validated bill of materials configured to order through our Dell, HPE and distribution relationships, with expansion priced as an option you exercise when growth is real.

Can the Inference Pod fine tune models?

Light adaptation, sometimes. Periodic, planned fine tuning on your own data is what the Fine Tuning Pod exists for, with the additional accelerator and storage capacity that work actually needs. If fine tuning is on your roadmap, say so in the sizing conversation and we will tell you honestly which tier the roadmap points at.

What is the Inference Pod not designed for?

Foundation model training, large scale fine tuning, and regulated workloads that need a sovereign wrap. Training is a different discipline with different economics. Regular fine tuning belongs on the Fine Tuning Pod, and regulated estates in financial services, healthcare and critical infrastructure belong on the Private AI Pod, where security and data residency are part of the design rather than an addition.

Size it before you buy it

Bring the workload: models, users, data. We will bring the numbers, and a straight answer if the API is still the cheaper home for it.

Book a sizing conversation