Perspectives · AI infrastructure · Cost

Your GPU bill is an architecture decision, not a procurement problem.

The conversation usually starts in a budget review. The GPU line has tripled, the CFO wants to know why, and procurement's answer is to renegotiate the rate. They'll get eight percent. Meanwhile, the clusters those GPUs live in are running at utilization numbers nobody wants to put on a slide.

That's the uncomfortable truth about AI infrastructure economics: the money isn't lost at the pricing table. It's lost in the scheduler. A discounted GPU that sits idle 70% of the time is still one of the most expensive things in your cloud bill.

The most expensive GPU in your fleet is the one doing nothing — and most fleets are full of them.

Utilization is the real unit price

Enterprises buy GPUs the way they bought VMs in 2015: one workload, one box, statically allocated, sized for peak. But AI workloads don't behave like VMs. Inference traffic is bursty. Training is spiky and schedulable. Experimentation is constant and small. Give each of those a dedicated GPU and you've built a fleet of expensive space heaters.

The platform mechanics that fix this are well understood — they're just rarely implemented together. NVIDIA's MIG partitions a single GPU into isolated slices, so seven small inference services don't each need their own card. Time-slicing lets bursty, latency-tolerant workloads share what's left. Karpenter-style autoscaling consolidates nodes so the cluster shrinks when the queue does. None of this is exotic. All of it is architecture, and none of it comes from procurement.

The build-vs-rent line moves with volume

The same discipline applies one layer up. Per-token API pricing is the right answer at low volume — no platform to run, no capacity to manage. But the crossover point where self-hosted serving becomes cheaper arrives earlier than most teams expect, and when it arrives, the deciding factor isn't the model. It's whether your platform can actually run vLLM on shared GPU capacity with autoscaling that works. Organizations without that platform stay on per-token pricing years past the crossover, because the alternative is a platform build nobody scoped.

The static fleet

  • One workload per GPU, sized for peak
  • Utilization measured never, or in a crisis
  • Capacity requests settled by seniority
  • Per-token API bills growing unbounded
  • Savings sought at the pricing table

The scheduled platform

  • MIG slices & time-slicing matched to workload profiles
  • Utilization on a dashboard, reviewed like spend
  • Capacity governed by namespace quotas & priorities
  • Self-hosted serving where volume justifies it
  • Savings engineered in the scheduler

Where to start

Not with a reservation purchase. Start by measuring: per-workload GPU utilization, queue wait times, and the fully-loaded cost per thousand inferences for your top services. In my experience those three numbers are enough to find 30–50% of spend that's recoverable through scheduling alone — before any pricing conversation happens.

Then treat the fix as a platform project with an owner, not a collection of tuning tickets. Fractionalization, autoscaling, quotas, and showback are one coherent system. Built together, they compound. Built piecemeal, they fight.

This is the work I do.

Bounded proofs on real data, agent platforms your organization owns, and the operating-model design that makes them stick — delivered end to end, corp-to-corp through Mazo Cloud Group LLC.

Start an engagement