CloudPe
Glossary

FLOPs

CloudPe Team •
FLOPs

What is FLOPs?

FLOPs stands for Floating-Point Operations Per Second. It measures how many mathematical calculations, specifically operations on decimal (floating-point) numbers, a processor can perform in one second. It’s one of the main numbers used to describe how powerful a CPU or GPU is, especially for AI and scientific computing.

A higher FLOPs number generally means a processor can do more math faster, which matters directly for how quickly an AI model trains or generates a response.

But the number on a spec sheet is only half the story. To understand why FLOPs became the industry’s yardstick and why it can also mislead you, it helps to know where it came from.

Where the metric came from

Before FLOPs, computers were rated in MIPS (Millions of Instructions Per Second). 

The problem: MIPS was easy to game. 

A processor could rack up millions of trivial instructions without doing any meaningful scientific calculation.

Seymour Cray changed that. 

In 1964, his CDC 6600 became the first machine widely recognized as a supercomputer, running at roughly 3 megaFLOPS, three million floating-point operations per second. 

It was designed for the kind of decimal-heavy math that nuclear simulation, weather forecasting, and aerospace design actually needed, which is why FLOPs, not instruction count, became the metric that stuck.

Twelve years later, Cray’s own follow-up, the Cray-1 (1976), pushed that to 160 megaFLOPS. By processing entire arrays of numbers at once instead of one value at a time. This “vector processing” approach cemented FLOPs as the definitive way to rank supercomputers for the next several decades.

From there, the scale climbed roughly a thousandfold every decade or so:

MilestoneYearSystemPeak performance
MegaFLOPS (10⁶)1964CDC 6600~3 MFLOPS
MegaFLOPS (10⁶)1976Cray-1~160 MFLOPS
GigaFLOPS (10⁹)1985Cray-2~1.9 GFLOPS
TeraFLOPS (10¹²)1997ASCI Red~1.3 TFLOPS
PetaFLOPS (10¹⁵)2008IBM Roadrunner~1.0 PFLOPS
ExaFLOPS (10¹⁸)2022Frontier~1.1 EFLOPS

Six decades, six orders of magnitude from a machine that could do three million operations a second to one that does over a billion times more.

How FLOPs is measured today

Because modern processors handle enormous volumes of calculations, FLOPs is expressed with a scale prefix, the same one used in the historical table above:

TermValue
MFLOPSMillion FLOPs
GFLOPSBillion FLOPs
TFLOPSTrillion FLOPs
PFLOPSQuadrillion FLOPs

Modern GPUs are typically measured in TFLOPS. NVIDIA’s H100, for example, delivers up to 989 TFLOPS at FP8 precision; nearly a quadrillion floating-point operations every second at that precision level.

FLOPs and precision: why the number changes

FLOPS and Precision

Here is a key detail: a single GPU doesn’t have just one FLOPs number. It has several, depending on the numerical precision used for the calculation.

Lower precision means fewer decimal digits tracked per number. This means the hardware can push through more operations per second, at some cost to numerical accuracy. 

That’s why NVIDIA lists multiple FLOPs figures for the same chip. The H100 delivers roughly 989 TFLOPS at FP8, but only about 60 TFLOPS at FP64 (double precision), over 15 times lower, on the same GPU.

This isn’t a new idea; it’s actually a reversal of one. For fifty years, computing pushed toward more precision from 16-bit up to 64-bit for scientific accuracy. Deep learning flipped that trajectory:

  • 2012 (AlexNet era): Neural networks trained using standard FP32 math on standard graphics cards.
  • 2017 (NVIDIA Volta): Tensor Cores arrived, hardware built specifically for fast, mixed-precision matrix math using FP16.
  • 2022–present (NVIDIA Hopper): FP8 “Transformer Engines” trade even more bit-depth for throughput.

Neural networks tolerate this trade-off well because they act like statistical filters; they’re robust to a bit of numerical noise. AI training and inference use lower-precision formats like FP16, BF16, or FP8 deliberately, because the accuracy loss is small and the speed gain is large.

Theoretical FLOPs vs. real-world performance

The FLOPs number on a spec sheet is a theoretical peak, the maximum possible under ideal conditions. Real-world training rarely hits it.

The underlying reason is a counter-intuitive shift in modern hardware. Performing a floating-point calculation is now essentially free; moving the data to the processor to do that calculation is what’s expensive.

  • A single FP32 multiply-accumulate operation costs roughly 1–3 picojoules of energy.
  • Fetching 32 bits from DRAM costs roughly 1,000–2,000 picojoules, several hundred times more.

To hit its advertised peak, a GPU needs to perform hundreds of math operations for every byte of data it pulls from memory. Fall short of that ratio, and the chip sits partially idle, bottlenecked by memory bandwidth rather than compute, regardless of how many TFLOPS are printed on the box.

This gap between advertised and actual throughput has a name: Model FLOPs Utilization (MFU).

MFU = Actual observed FLOPs/sec/Theoretical peak FLOPs/sec

In real-world LLM training and inference, MFU typically lands between 30% and 55%, driven down by:

  • Memory bandwidth limits moving weights from HBM to on-chip memory often takes longer than the math itself.
  • Communication overhead in multi-GPU clusters, interconnects like NVLink or InfiniBand stall compute while syncing data between chips.
  • Non-math operations like LayerNorm or Softmax are memory-bound, not compute-bound, so they don’t benefit from a chip’s raw FLOPs at all.

This is also why low-precision formats won out beyond just raw speed. Thus, cutting a number from FP32 to FP8 doesn’t just make the math faster; it cuts memory traffic by roughly 75%, which is often the actual bottleneck standing between a GPU and its peak FLOPs.

FLOPs vs. VRAM: different bottlenecks

FLOPs measures raw compute speed. VRAM measures how much data the GPU can hold close at hand. A workload can be limited by either one.

  • Training a model that fits comfortably in memory is usually compute-bound; FLOPs matters most.
  • Running inference on a very large model is often memory-bound; VRAM capacity and bandwidth matter more than raw FLOPs.

The NVIDIA H100 and H200 make this concrete. 

Both chips share the same Hopper Tensor Cores and the identical peak FLOPs across every precision level. The only difference is memory: the H200 carries 141 GB of HBM3e at 4.8 TB/s, versus the H100’s 80 GB at 3.35 TB/s bandwidth about 43% more bandwidth, zero more compute. 

For memory-bound work like the decode phase of LLM inference, that difference alone lets the H200 push meaningfully more tokens per second, despite having no FLOPs advantage on paper.

Where FLOPs fits when choosing a GPU

FLOPs is one part of a larger picture that includes VRAM, memory bandwidth, and the specific precision your workload actually uses. A GPU with the highest advertised FLOPs number isn’t automatically the best choice if your workload is limited by memory rather than raw compute as the H100/H200 comparison shows, two chips can share identical FLOPs and still perform very differently in practice.