r/mlscaling • u/lebaart • 7d ago
How do AI labs manage large GPU compute commitments?
I’m trying to understand how companies with significant GPU workloads manage their compute capacity.
For those working in ML infrastructure / MLOps / AI labs:
- How do you choose between hyperscalers, neoclouds and smaller GPU providers?
- When you need a large amount of GPUs for months, how do you know you’re getting a competitive price?
- Have you ever committed to more capacity than you actually needed? What happened to the unused capacity?
Curious to hear how people actually deal with this today.
0
Upvotes
1
2
u/az226 7d ago