Velocity Stream LogoVelocity Stream Logo
Back to Insights
AI/ML
FinOps
Kubernetes

Stop Burning Money on Idle GPUs: How We Fixed 5% Utilization on EKS

How we replaced the default Kubernetes scheduler with Volcano and Karpenter to increase AI workload GPU utilization from 5% to 75% for a GenAI startup.

7 min read
Futuristic isometric server rack showing optimized GPU compute streams

The Problem: The $50,000/Month Idle Fleet

In 2026, compute is the new oil, and NVIDIA H100s are the refineries. But for a Series-A GenAI Video Startup we recently audited, those refineries were mostly sitting dormant. They were spending upwards of $50,000 a month on AWS p5.48xlarge instances, yet their internal metrics showed an average GPU utilization of just 5.2%.

This wasn't a code issue; it was a Kubernetes scheduling issue.

The "Gang Scheduling" Failure

Standard Kubernetes schedulers process pods one-by-one. When launching a distributed training job that requires 32 GPUs to start simultaneously, the default scheduler would place 16 pods, realize the cluster was full, and leave those 16 GPUs completely idle while waiting for someone to manually scale the cluster. This is known as "resource fragmentation."

The Fix: Ripping Out the Default Scheduler

To fix this financial bleed, we re-architected their EKS cluster to handle the specialized demands of AI/ML workloads. Here is exactly what we did.

1. Implementing the Volcano Batch Scheduler

We bypassed the default kube-scheduler and installed Volcano, a batch scheduling engine built specifically for high-performance computing (HPC) and AI workloads.

Volcano supports gang scheduling natively. If a training job requires 32 GPUs, Volcano ensures that either all 32 pods are scheduled simultaneously, or none of them are. This prevents partial deployments from locking up expensive GPUs.

apiVersion: batch.volcano.sh/v1alpha1
kind: Job
metadata:
  name: distributed-video-training
spec:
  minAvailable: 32 # Gang scheduling requirement
  schedulerName: volcano
  tasks:
    - replicas: 32
      name: gpu-worker
      template:
        spec:
          containers:
            - name: training-container
              image: genai-model:v2
              resources:
                limits:
                  nvidia.com/gpu: 1

2. Aggressive Spot Scaling with Karpenter

With Volcano ensuring jobs didn't hang, we needed a way to provision nodes instantly. We implemented Karpenter to replace the sluggish Cluster Autoscaler. We configured Karpenter to heavily utilize AWS Spot Instances for ephemeral training jobs, falling back to On-Demand only when Spot capacity was unavailable in the region.

3. Dynamic Resource Allocation (DRA) for Inference

For their smaller inference workloads (which only needed 20% of a GPU's VRAM), allocating a full GPU per pod was incredibly wasteful. We leveraged Kubernetes Dynamic Resource Allocation (DRA) and NVIDIA's Time-Slicing features to allow up to 4 inference pods to share a single physical GPU securely, quadrupling their serving capacity per node.

Before Velocity Stream

  • 5.2% Average GPU Utilization
  • Hanging distributed training jobs
  • $50k/mo on idle P-series instances
  • 1 GPU = 1 Inference Pod

After Implementation

  • 75.4% Average GPU Utilization
  • Guaranteed Gang Scheduling
  • 60% reduction in AWS Compute Bill
  • 1 GPU = 4 Inference Pods (via DRA)

The ROI of AI FinOps

By treating Kubernetes as a high-performance batch computing engine rather than a simple web-server orchestrator, we turned a massive financial liability into a competitive advantage. The client now trains models faster, serves inference cheaper, and has extended their runway by months.

If your cloud bill is growing faster than your revenue, you likely have an orchestration problem, not a pricing problem.

Bleeding money on idle cloud infrastructure?

Our senior engineers specialize in AI FinOps, Kubernetes optimization, and cloud architecture for high-growth startups.

Chat with an Engineer