The conversation usually starts in a budget review. The GPU line has tripled, the CFO wants to know why, and procurement's answer is to renegotiate the rate. They'll get eight percent. Meanwhile, the clusters those GPUs live in are running at utilization numbers nobody wants to put on a slide.
That's the uncomfortable truth about AI infrastructure economics: the money isn't lost at the pricing table. It's lost in the scheduler. A discounted GPU that sits idle 70% of the time is still one of the most expensive things in your cloud bill.
Utilization is the real unit price
Enterprises buy GPUs the way they bought VMs in 2015: one workload, one box, statically allocated, sized for peak. But AI workloads don't behave like VMs. Inference traffic is bursty. Training is spiky and schedulable. Experimentation is constant and small. Give each of those a dedicated GPU and you've built a fleet of expensive space heaters.
The platform mechanics that fix this are well understood — they're just rarely implemented together. NVIDIA's MIG partitions a single GPU into isolated slices, so seven small inference services don't each need their own card. Time-slicing lets bursty, latency-tolerant workloads share what's left. Karpenter-style autoscaling consolidates nodes so the cluster shrinks when the queue does. None of this is exotic. All of it is architecture, and none of it comes from procurement.
The build-vs-rent line moves with volume
The same discipline applies one layer up. Per-token API pricing is the right answer at low volume — no platform to run, no capacity to manage. But the crossover point where self-hosted serving becomes cheaper arrives earlier than most teams expect, and when it arrives, the deciding factor isn't the model. It's whether your platform can actually run vLLM on shared GPU capacity with autoscaling that works. Organizations without that platform stay on per-token pricing years past the crossover, because the alternative is a platform build nobody scoped.
The static fleet
- One workload per GPU, sized for peak
- Utilization measured never, or in a crisis
- Capacity requests settled by seniority
- Per-token API bills growing unbounded
- Savings sought at the pricing table
The scheduled platform
- MIG slices & time-slicing matched to workload profiles
- Utilization on a dashboard, reviewed like spend
- Capacity governed by namespace quotas & priorities
- Self-hosted serving where volume justifies it
- Savings engineered in the scheduler
Where to start
Not with a reservation purchase. Start by measuring: per-workload GPU utilization, queue wait times, and the fully-loaded cost per thousand inferences for your top services. In my experience those three numbers are enough to find 30–50% of spend that's recoverable through scheduling alone — before any pricing conversation happens.
Then treat the fix as a platform project with an owner, not a collection of tuning tickets. Fractionalization, autoscaling, quotas, and showback are one coherent system. Built together, they compound. Built piecemeal, they fight.