Major university hospital · healthcare

H100 GPU self-service orchestration portal

Turned a manually-coordinated H100 cluster into a Kueue self-service portal. 2× utilization, 18 months zero downtime.

2.1×
GPU utilization
34%
idle resources
18mo
zero downtime
THE CHALLENGE

Why it was hard.

Expensive H100 cards were shared across research teams, but every allocation went through a person. Getting a GPU took about a week, and with no return process, idle cards stayed locked away from other teams. As a medical-data environment, isolation and audit demands were high.

Constraints

  • Research in progress — zero downtime was a given
  • Medical data — isolation, access control, audit logs required
  • Each team had different workloads (training, inference, batch)
OUR APPROACH

What we did.

  1. Measure & diagnoseQuantify the bottleneck from utilization and wait times
  2. Scheduling designQueues, quotas, priority and preemption with Kueue
  3. Self-service portalResearchers request and release on their own
  4. Operational handoverUtilization dashboards, guardrails, docs
OUTCOME

Outcome.

Allocation dropped from a week to same-day, and the same cards now run more experiments. Utilization doubled and idle resources fell. The portal has run with zero downtime for 18 months.

STACK

Stack.

EKSKueueKarpenterNVIDIA H100GrafanaTerraform

Got a similar challenge? Let's talk it through, case in hand.

In 30 minutes we'll pin down what matches and what differs.

Already trusted by teams across finance · healthcare · media · public
Request a technical review