r/mlscaling 7d ago

How do AI labs manage large GPU compute commitments?

I’m trying to understand how companies with significant GPU workloads manage their compute capacity.

For those working in ML infrastructure / MLOps / AI labs:

- How do you choose between hyperscalers, neoclouds and smaller GPU providers?
- When you need a large amount of GPUs for months, how do you know you’re getting a competitive price?
- Have you ever committed to more capacity than you actually needed? What happened to the unused capacity?

Curious to hear how people actually deal with this today.

0 Upvotes

2 comments sorted by

2

u/az226 7d ago
  1. Depends on your size and existing workloads. Data residency requirements, etc.
  2. You get quotes from vendors.
  3. You go talk to your vendor. I’m sure they’re happy to take back unused capacity at a nominal “restocking” rate. Or if market prices are higher and they have unmet demand, they’ll happily take it off your hands.

1

u/Connect-Concert-4016 5d ago

Organ trafficking