Sapience runs two NVIDIA Blackwell platforms: the HGX B300, built for training-scale clusters, and the RTX PRO 6000 Blackwell Server Edition, built for inference, agentic workloads and visualisation. Both come as 8-GPU Supermicro servers in dedicated clusters of 8, 24 or 32 servers. The differences are memory, interconnect, terms and availability.
Side by side
| HGX B300 | RTX PRO 6000 | |
|---|---|---|
| GPU memory | 288 GB HBM3e per GPU · 2.3 TB per node | 96 GB GDDR7 per GPU · 768 GB per server |
| Compute (FP4) | 120 PFLOPS per node | 32 PFLOPS per server |
| GPU-to-GPU | 1.8 TB/s NVLink | PCIe Gen 5.0 |
| Node fabric | 1.6 Tb/s, ConnectX-8 ready | 400G |
| Terms | Multi-year reservation | From 3 months; longer terms on request |
| Availability | Reservations open · deliveries from December 2026 | Fully reserved through 2026 · 2027 reservations open |
When HGX B300 is the right node
- Training large models. Mixed-precision training with Adam needs roughly 16 bytes per parameter before activations. The B300's three times larger memory per GPU and NVLink between GPUs reduce sharding and communication overhead.
- Full-parameter fine-tunes of large open-weight models.
- Serving the largest models, or serving at the highest throughput: 2.3 TB of GPU memory per node.
When RTX PRO 6000 is the right node
- Inference for most models. A 70B-parameter model at FP8 (about 70 GB of weights) fits on one 96 GB GPU, with room for the KV cache.
- Adapter fine-tuning (LoRA, QLoRA), where memory is dominated by the frozen base model.
- Agentic AI, multi-application pipelines and visualisation, the workloads NVIDIA designed the Server Edition for.
- Shorter projects, with terms from three months.
A quick decision rule
If the job is pre-training or a full fine-tune of a model in the tens of billions of parameters or more, or if a single model needs more than 96 GB per GPU at your chosen precision, start from HGX B300. Otherwise start from RTX PRO 6000, and move up when the profile of the run says so. The inference sizing guide has the memory arithmetic.
Specifications as published by NVIDIA and Sapience AI for the configurations deployed. Memory figures are rules of thumb; profile your own workload before committing.
Questions about your workload?
Describe the model, the data and the timeline. You get a recommendation and a quote, not a sales deck.