Reservations open
GB300
NVIDIA NVL72
- GPUs
- 72x Blackwell Ultra per rack
- Memory
- 288 GB HBM3e per GPU
- Fabric
- NVLink 5 / Quantum-X800
Managed training and inference workloads with serverless or dedicated compute. From 64 GPUs to multi-thousand-GPU islands, configured around your workload.
64 to 10k+
GPUs
1 mo+
Terms
24 / 7
Support
Current NVIDIA training platforms from our own racks and a vetted provider network, with status updated as capacity comes online.
Rack-scale / NVL72 systems
Reservations open
NVIDIA NVL72
Available now
NVIDIA NVL72
Node-scale / HGX systems
NVIDIA HGX
Available now
NVIDIA HGX
Available now
NVIDIA HGX
Available now
Need another node count, storage ratio, or generation? Share the workload and we will configure the cluster around it.

NVIDIA is an investor and our closest engineering partner. In a market where frontier GPUs are allocated before they ship, that partnership is why the fleet above reads available, and it carries through how every cluster is designed, built, and supported.
Each new platform generation is allocated to us as it ships, so reservations are fulfilled on silicon timelines, not the resale market.
Every cluster is built and burned in to NVIDIA reference designs and validated before handover.
When an issue needs NVIDIA's attention, it reaches NVIDIA engineering through the partnership, not a vendor queue.
Before this platform ran an external workload, it trained Reka's multimodal frontier models end to end. Link flaps, straggler nodes, checkpoint stalls, and silent data corruption were found and fixed on our runs, not yours.
10,000+
GPUs in one training fabric
99.2%
measured training goodput
24 / 7
operator coverage
The machine is the floor. What you are buying is everything that keeps it busy.
No hypervisor between your job and the silicon. You get the whole machine, and every FLOP you pay for.
Full-bisection, rail-optimized InfiniBand built for collective operations at cluster scale.
High-throughput storage sits beside the compute so checkpoints do not gate step time.
Slurm or Kubernetes configured, images loaded, and a reference job completed before handoff.
Hot spares, uptime commitments, and reporting on the hours your workload actually trained.
Around-the-clock coverage from engineers who operate frontier training clusters themselves.
We size compute, fabric, storage, and regions around your workload.
Lock the capacity against a firm go-live date, from one month to multi-year.
We configure your scheduler, images, networking, access, and observability.
Start on a burned-in cluster with root access and live engineering support.
Bare metal. Your jobs run directly on the machines, with root access and the full fabric available to your workload.
Standard HGX blocks can be ready in days. Custom scheduler, image, storage, and networking layouts typically take a few weeks after signature.
One month to multi-year. Reserve for a single training run or lock a longer allocation with a defined go-live date.
Compute, power, cooling, storage, fabric, and support are included in the GPU-hour rate, with no ingress or egress fees.
Hot spares remain in the same fabric, failed nodes are swapped quickly, and downtime is credited.
Reka Cloud does, 24/7. The same engineering discipline used to train frontier multimodal models is applied to your cluster.
Share the workload and timeline. We will return with concrete capacity, topology, and commercial options.
Reserve capacity