CloudPe
Glossary

AI compute

CloudPe Team •
AI compute

What is AI compute?

AI compute is the processing power, mainly GPUs, used to train and run AI models. Every step in an AI model’s life, from learning patterns in data to answering a user’s prompt in real time, consumes compute. How much compute a task needs, and what kind, is one of the biggest factors in how much AI actually costs to build and run.

AI compute isn’t one single thing. It splits into two very different phases with very different demands: training and inference.

Training compute vs inference compute

Training computeInference compute
What it doesTeaches the model by adjusting its parameters over huge datasetsRuns the already-trained model to generate real outputs
When it happensOnce (or periodically, for updates), before deploymentEvery single time a user sends a request, for the life of the product
Cost typeA large, fixed, upfront investmentAn ongoing, continuous cost that scales with usage
Hardware priorityMassive parallelism and memory bandwidth across many GPUs at onceLow latency and efficient, real-time response per request

Training is a one-time (or occasional) heavy lift. Inference is the meter that keeps running for as long as people use the model.

Why this split matters

Training compute is a capital expense. You spend it once to produce a working model, and that model’s weights can then be reused indefinitely, across any number of deployments, without retraining.

Inference compute is closer to an operating expense. It starts the moment a product launches and never stops, since it runs on every single user request. This is why a model that was relatively cheap to train can still become very expensive to run in production; if it’s used by millions of people, every one of those requests draws compute.

Which one costs more?

It depends on stage and scale. Early in a model’s life, training dominates the cost. Once it’s deployed and being used at scale, inference costs can add up faster and for much longer than the original training run did.

Working out the crossover point

The honest answer is that there’s a crossover point: a moment when everything you’ve spent on inference so far passes what you spent on training. Where it lands depends entirely on how much the model gets used, so it’s worth seeing how the arithmetic works.

For example: Take a hypothetical 70-billion-parameter model trained on 2 trillion tokens. All numbers below are illustrative and rounded; swap in your own.

Training side

A rough estimate for training compute is 6 × parameters × training tokens:

6 × 70B × 2T = 8.4 × 10²³ FLOPs

A modern data-centre GPU delivers somewhere near 400 teraFLOPs of useful work per second on this kind of job, well below its headline peak, because no cluster runs at 100% efficiency. That gives:

  • ~580,000 GPU-hours
  • ~$1.45M, at $2.50 per GPU-hour

Spent once.

Inference side

Say serving costs $1.00 per million tokens, and an average request uses 1,500 tokens in and out. That’s $0.0015 per request. Now, the same model, two products:

Consumer appInternal tool
Requests per month200 million88,000 (200 staff, ~20/day)
Inference cost per month$300,000$132
Time to pass $1.45M training cost~5 months~900 years

Same model. Same training bill. Completely different economics.

That’s the part the training/inference split alone doesn’t tell you. Training cost is fixed the moment the run finishes. Inference cost is a function of how popular you are which is why a successful product can find its compute bill inverting within a year, while an internal one never gets close.

How training and inference use hardware differently

Training Inference use difeerent hardware

Let’s look at how training and inference use hardware:

Training clusters: Built for extreme parallelism, many GPUs working in sync on the same job, connected through high-speed interconnects so they can share data constantly without becoming a bottleneck.

Inference clusters: Built for responsiveness and efficiency per request, prioritizing low latency and often running on hardware optimized for cost-effective, real-time responses rather than raw peak throughput.

This is also why the same GPU isn’t always the best choice for both. Some hardware is optimized for training’s raw parallel throughput; other setups are tuned specifically for fast, efficient inference at scale.

What determines how much compute you need

Below are the specs that determine how much compute you need

  • Model size: Larger models, more parameters, generally need more compute for both training and inference.
  • Dataset size (for training): More training data typically means more compute to process it.
  • Precision: Lower-precision formats (like FP8 instead of FP32) let hardware do more useful work per second, directly affecting how much FLOPs capacity you actually need.
  • Usage volume (for inference): More users and more requests directly multiply your ongoing inference compute needs.

Where AI compute comes from

Compute is priced and sold in GPU-hours: one GPU running for one hour. Every figure above, and nearly every quote you’ll be given, reduces to that unit.

You can buy those hours four ways:

  • On-demand: pay by the hour, cancel anytime. Highest rate, no commitment.
  • Reserved (or committed): you commit to a term, usually a year or more, for a lower rate.
  • Spot (or preemptible): spare capacity sold cheap, reclaimed at short notice when someone else needs it.
  • Owned or colocated: you buy the hardware. Capital expense up front, plus power, cooling, and depreciation for as long as you keep it.

These aren’t interchangeable, and the reason maps straight back to the training/inference split.

Training is a bounded, interruptible batch job. It has an end date, nobody’s waiting on it in real time, and good training code checkpoints regularly. So if a spot instance is reclaimed, you resume from the last checkpoint. That tolerance for interruption is exactly what makes cheap capacity usable.

Inference is unbounded and latency-sensitive. It runs for the life of the product, and every request has a user attached waiting for a response. You can’t have that capacity pulled mid-request, which rules out spot for anything user-facing and pushes you toward reserved or owned capacity.

So the split that shapes your hardware also shapes your procurement: training buys cheap, interruptible bursts; inference buys steady, guaranteed capacity.

Where AI compute fits in your infrastructure

AI compute is delivered through GPU-backed infrastructure, often described under the broader term AI cloud, and measured in terms of raw processing capability (FLOPs) alongside memory capacity (VRAM), since both together determine what a given piece of hardware can actually handle.