A Kubernetes platform built for AI workloads needs GPU node pools that support fractional GPU sharing, attribute-based scheduling, and topology-aware placement for distributed training, instead of just a pool of GPU nodes bolted onto a standard cluster.
An average GPU utilization across production Kubernetes clusters is just 5%, according to Cast AI’s 2026 State of Kubernetes Optimization Report, largely because default scheduling treats every GPU as a single, indivisible unit, regardless of how much of it a workload actually needs.
For context, CPU utilization across the same clusters averages 8%, and memory averages 20%. GPUs are the most expensive resource in the cluster by far, and they’re also the least utilized, the opposite of where the waste should be concentrated.
Here’s what separates a GPU node pool that works from one that quietly wastes most of its capacity.
Why default Kubernetes GPU scheduling wastes so much capacity
Default Kubernetes GPU scheduling wastes capacity because it treats every GPU as a single, indivisible unit, regardless of its size or how much of it a workload actually needs. When something asks Kubernetes for a GPU, Kubernetes just finds any node with a free one and hands over the whole card, whether the workload needs it all or just a small slice.
This creates what’s sometimes called a GPU waste tax. For instance, picture ten small AI services, each one only needing about 15% of a GPU’s power. Under default scheduling, each of those ten services still reserves an entire physical GPU for itself, since Kubernetes has no way to hand out partial access.
Ten GPUs are reserved, even though the actual combined demand is only roughly 1.5. That gap, repeated across every workload in a cluster, accounts for most of the 5% average utilization figure.
Deploy VMs and GPU within minutes with CloudPe.
Try for freeWhat fractional GPU sharing solves
Fractional GPU sharing solves the waste problem by allowing multiple workloads to use parts of the same physical GPU rather than each claiming a whole one. NVIDIA supports this through three distinct approaches, and picking the right one matters for both performance and monitoring.
Time-slicing and MPS for shared, lower-isolation workloads
Time-slicing splits GPU access across pods in rotating turns, suited to lower-criticality workloads that can tolerate some contention. NVIDIA’s Multi-Process Service (MPS) goes a step further, enabling higher throughput for trusted workloads that share a GPU simultaneously rather than taking turns. Neither provides hardware-level isolation between workloads.
MIG for hardware-isolated GPU slices
Multi-Instance GPU (MIG) partitions a physical GPU into genuinely isolated hardware slices, each with dedicated compute and memory. This only works on A100s, H100s, and newer architectures, but it’s the only option that gives proper per-instance monitoring.
Time-sliced workloads only report node-level metrics, since monitoring tools like DCGM Exporter can’t attribute usage to individual containers without MIG’s hardware separation.
How Dynamic Resource Allocation changes GPU scheduling
Dynamic Resource Allocation (DRA) changes GPU scheduling by allowing pods to request GPUs based on attributes such as architecture and memory size, rather than a simple integer count. A claim requesting an Ampere-architecture GPU with 20GB of memory is matched to any hardware that meets that requirement, rather than requiring someone to manually pick an instance type.
This has real practical payoffs. Cast AI documented a demo running seven replicas of a workload, each with a dedicated 1g.5gb MIG partition, all sharing a single A100, generating output simultaneously with full hardware isolation between them, without adding a single additional GPU node.
Serving a 7-billion-parameter model on a full H100 wastes most of that GPU’s capacity, but requesting a smaller MIG profile through DRA lets multiple models share the same physical card, reducing the effective per-model GPU cost.
Autoscalers built around DRA can take this further, reading a pod’s resource claim, simulating which available instance type would satisfy it, and provisioning the cheapest option automatically, without anyone manually selecting a GPU tier.
If your team is currently provisioning full GPUs per workload by default, it’s worth checking whether a Kubernetes platform with proper GPU node pool support could recover a meaningful share of that unused capacity without adding hardware.
Why distributed training needs topology-aware scheduling
Distributed training needs topology-aware scheduling because raw GPU count doesn’t guarantee the placement a training job actually needs. A cluster with 40 GPUs available can still fail to satisfy a request for 8 GPUs on a single node, since that request depends on NVLink bandwidth within a single physical node, not just on aggregate GPU availability across the cluster.
Schedulers like Volcano and Apache YuniKorn address this with topology-aware placement policies that account for NVLink and PCIe layout, though they require accurate topology information in node labels to work correctly.
Above the scheduler layer, Kueue manages team-level resource quotas through ClusterQueue and LocalQueue resources, holding jobs that exceed a team’s allocation before they ever reach the scheduling pool.
Deploy VMs and GPU within minutes with CloudPe.
Try for freeWhat separates a safe multi-tenant GPU platform from a risky one
A safe multi-tenant GPU platform cleanly separates teams so that one workload can’t degrade another’s performance, which matters more for GPUs than for almost any other cluster resource, given how expensive shared capacity is.
Namespaces and RBAC form the basic isolation boundary, with ResourceQuotas setting hard GPU limits, physical or MIG-sliced, per team or project.
PriorityClasses add a second layer of protection. Production inference endpoints typically get elevated priority with preemption enabled, so an unthrottled research job can’t quietly consume enough GPU capacity to degrade latency on a system serving real traffic.
Spot-aware scheduling adds cost efficiency on top. Tools like the KAI Scheduler support automatic failover, rescheduling workloads to on-demand nodes the moment a spot instance gets reclaimed, without manual intervention.
For teams running mixed research and production workloads on the same cluster, this kind of isolation and priority-aware GPU scheduling is what prevents a shared platform from becoming a shared point of failure.
What to watch when monitoring GPU node pools
Monitoring GPU node pools requires looking past utilization percentage alone, since a healthy-looking number can hide a real bottleneck.
Illustrative example: a quantized 7B model running on a 16GB GPU can show utilization figures in the 60-70% range that look perfectly healthy, while memory is nearly maxed out, leaving little headroom for batching.
Once concurrent requests climb past a handful, latency spikes purely from memory pressure, a problem the utilization metric never surfaces on its own. Moving to a GPU with more memory directly resolves this class of issues.
Conclusion
A Kubernetes platform is only as good for AI workloads as its GPU scheduling underneath it. Fractional sharing through MIG or time-slicing, attribute-based allocation through DRA, and topology-aware placement for distributed training are what separate a platform running at 5% GPU utilization from one that genuinely gets value from expensive hardware.
If your current setup treats every GPU request the same way it treats a CPU core, that gap is probably showing up on your bill whether or not it’s showing up in your dashboards.
None of this requires rebuilding a cluster from scratch. Most of it comes down to which scheduling layer and sharing strategy a platform supports out of the box, which is exactly the kind of detail that’s easy to overlook until a GPU bill arrives that doesn’t match the workload it’s paying for.
If you’re evaluating infrastructure for GPU-heavy Kubernetes workloads, it’s worth considering how a managed Kubernetes platform with dedicated GPU node pools handles these specifics before assuming that any GPU-labeled node pool will do the job equally well.
Deploy VMs and GPU within minutes with CloudPe.
Try for freeFrequently Asked Questions
What is MIG in Kubernetes GPU scheduling?
MIG (Multi-Instance GPU) partitions a physical GPU into hardware-isolated slices, each with dedicated compute and memory, available on A100, H100, and newer NVIDIA architectures.
Does Dynamic Resource Allocation replace the NVIDIA device plugin?
DRA represents the newer, Kubernetes-native approach to device allocation, allowing workloads to request GPUs by attribute rather than by integer count. The device plugin path still exists, but DRA is increasingly the direction most GPU scheduling tooling is moving toward for structured requirements.
Why is average GPU utilization on Kubernetes so low?
Mainly because default scheduling reserves entire GPUs per pod regardless of actual usage, and Cast AI’s 2026 report puts the resulting average utilization across production clusters at just 5%.
Can Kubernetes GPU node pools use spot instances safely?
Yes, with the right scheduler. Tools like the KAI Scheduler support spot-aware scheduling, automatically rescheduling workloads to on-demand capacity if a spot instance gets reclaimed mid-job.
What’s the difference between time-slicing and MIG?
Time-slicing rotates GPU access between pods without hardware isolation and only reports node-level metrics. MIG creates genuinely isolated hardware slices with proper per-instance monitoring, but only on supported GPU generations.