LLM and foundation-model training
288 GB per GPU and 1.8 TB/s NVLink keep large models and their training state on the GPUs; the 1.6 Tb/s fabric scales across up to 32 servers.
LLM training →Each node is an 8-GPU NVIDIA HGX B300 system built by Supermicro. Nodes are connected on a 1.6 Tb/s fabric and delivered as bare metal: your operating system, your drivers, your orchestration.
| Specification | NVIDIA HGX B300 node |
|---|---|
| System | 8-GPU HGX B300 system, built by Supermicro |
| GPUs per node | 8 × B300 SXM (Blackwell Ultra) |
| GPU memory | 288 GB HBM3e per GPU · 2.3 TB per node |
| Compute (FP4) | 120 PFLOPS per node |
| GPU interconnect | 1.8 TB/s NVLink, GPU to GPU |
| Node fabric | 1.6 Tb/s, ConnectX-8 ready |
| Power | 3,000 W redundant, Titanium-grade power supplies |
| Storage | Encrypted NVMe on every node (AES-256 at rest, TLS 1.3 in transit) |
| Tenancy | Single-tenant: dedicated compute, storage and network |
Specifications as published by NVIDIA and Sapience AI for the configuration deployed. System memory per server, storage and networking options are confirmed in your quote.
HGX B300 is sold by the cluster, not by the GPU. Aggregate figures below are the per-node specification multiplied by the node count; delivered performance depends on the workload.
| Cluster | GPUs | HBM3e memory | FP4 compute |
|---|---|---|---|
| 8 servers | 64 × B300 | 18.4 TB | 960 PFLOPS |
| 24 servers | 192 × B300 | 55.3 TB | 2,880 PFLOPS |
| 32 servers | 256 × B300 | 73.7 TB | 3,840 PFLOPS |
HGX B300 is the node for work that has to keep a large model, its optimiser state and its activations close to the GPUs, across several servers.
288 GB per GPU and 1.8 TB/s NVLink keep large models and their training state on the GPUs; the 1.6 Tb/s fabric scales across up to 32 servers.
LLM training →Full-parameter fine-tuning of large open-weight models, where weights, gradients and optimiser state together run to terabytes.
Fine-tuning →Models too large for a 96 GB GPU, served with the highest throughput from 2.3 TB of GPU memory per server.
Inference →How the two Blackwell nodes compare on memory, interconnect, terms and availability.
Read the comparison →Anything else: sales@sapienceai.eu or +421 233 329 562.
One cluster of 8 servers (64 GPUs) on a multi-year term. Clusters come in 8-, 24- and 32-server sizes and are sold whole.
Reservations are open now. First deliveries are scheduled from December 2026; the delivery date for your cluster is confirmed in the contract.
Per GPU-hour on a reserved term, quoted in writing. Rates are in US dollars and exclude VAT; the quote confirms rate, term, delivery date and availability.
In a colocation data centre in Bratislava, Slovakia, on systems operated by Sapience with 24/7 monitoring by our own engineers. Customer data is processed and stored in the EU.
Nodes are delivered as bare metal today, so you can install the scheduler and stack you already use. Managed Kubernetes and SLURM are in development on the same infrastructure.
Tell us the cluster size, term and start month. The quote confirms rate, availability and the delivery date in writing.