H100 GPU self-service orchestration portal
Turned a manually-coordinated H100 cluster into a Kueue self-service portal. 2× utilization, 18 months zero downtime.
Why it was hard.
Expensive H100 cards were shared across research teams, but every allocation went through a person. Getting a GPU took about a week, and with no return process, idle cards stayed locked away from other teams. As a medical-data environment, isolation and audit demands were high.
Constraints
- Research in progress — zero downtime was a given
- Medical data — isolation, access control, audit logs required
- Each team had different workloads (training, inference, batch)
What we did.
- Measure & diagnoseQuantify the bottleneck from utilization and wait times
- Scheduling designQueues, quotas, priority and preemption with Kueue
- Self-service portalResearchers request and release on their own
- Operational handoverUtilization dashboards, guardrails, docs
Outcome.
Allocation dropped from a week to same-day, and the same cards now run more experiments. Utilization doubled and idle resources fell. The portal has run with zero downtime for 18 months.
Stack.
More work from this service.
- ↳Cost & performance turnaround on an undocumented legacy systemFinance·fintechReverse-engineered a system whose owners had left and whose docs were gone, then stopped the cost leaks and bottlenecks.→
- ↳Root-cause analysis of a DDoS·breach and a rebuilt defense systemE-commerce·gamingPinned down the root cause of DDoS·hacking attempts and rebuilt the defense, from detection through response.→
Got a similar challenge? Let's talk it through, case in hand.
In 30 minutes we'll pin down what matches and what differs.