Generic server rack
– Getty Images

Kubernetes clusters are running far below capacity, and the gap is widening, according to Cast AI’s "2026 State of Kubernetes Optimization Report."

The vendor analyzed tens-of-thousands of clusters across Amazon Web Services (AWS), Microsoft Azure, and Google Cloud, finding average central processing unit (CPU) utilization at 8%, which is down from 10% a year earlier. Memory utilization fell from 23% to 20%. Graphic processing unit (GPU) utilization across AI and machine learning (ML) workloads averaged 5%, highlighting a growing mismatch between infrastructure spend and actual usage.

Cast AI sells optimization software, which is worth noting for context. But the scale and consistency of the data points to a broader pattern rather than isolated inefficiencies.

CPU over-provisioning rose from 40% to 69% year on year. In practice, this means teams are paying for capacity their workloads do not use. The report attributes this to a combination of defensive engineering practices and the way Kubernetes resource management works in production environments.

Engineers typically over-allocate CPU and memory to avoid performance issues or crashes. Those settings are often baked into templates and reused across services, even as workloads change. Autoscalers then provision infrastructure based on those inflated requests, locking in excess capacity at the cluster level. Over time, this creates a system where inefficiency compounds rather than corrects itself.

The impact becomes more pronounced with GPUs. Cloud GPU pricing is already high, and AWS increased H200 Capacity Block prices by 15% in January 2026, the report said, marking a rare upward shift in compute pricing. Running at 5% utilization against that cost base raises questions about how effectively AI infrastructure is being used in enterprise environments.

Spot instances, which could help offset some of that cost, have seen limited adoption for GPU workloads. Fewer than 2% of GPUs ran on Spot through 2025, largely due to limited availability for higher-end hardware. The report notes early signs of improved availability for lower-end GPUs such as Nvidia T4 in some AWS regions, but it remains unclear whether similar stability will emerge for newer, high-performance GPUs such as H100 and H200.

Not all findings reinforce the idea that more capacity improves reliability. In one example, a cluster recorded up to 50 out-of-memory (OOM) kills per measurement interval despite heavy over-provisioning. After automated rightsizing reduced CPU allocation by half, OOM events dropped to near zero. The result suggests that excessive resource allocation can introduce instability rather than prevent it.

On the infrastructure side, Arm-based CPUs are gaining traction. Arm nodes grew 3.5-times faster than x86 between mid-2024 and the end of 2025, and now account for around 9% of deployments in the dataset. For stateless, containerized workloads, the price-performance advantage of processors such as AWS Graviton is driving broader adoption.

Overall, the report suggests Kubernetes efficiency does not improve automatically as environments scale. Organizations that have reduced the utilization gap are treating resource configuration as an ongoing process, rather than something set at deployment and revisited only occasionally.