Inference at scale
Serving open-source or proprietary models from dedicated European servers, with a predictable cost per GPU-hour and no shared tenancy.
Inference →Each server carries eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs in a 4U Supermicro chassis optimised for air cooling. Servers are linked on a 400G fabric for multi-node work and delivered as bare metal.
| Specification | NVIDIA RTX PRO 6000 server |
|---|---|
| System | 8-GPU server, built by Supermicro |
| GPUs per server | 8 × RTX PRO 6000 Blackwell Server Edition |
| GPU memory | 96 GB GDDR7 per GPU · 768 GB per server · 14.3 TB/s aggregate |
| Compute (FP4) | 32 PFLOPS per server |
| GPU interconnect | PCIe Gen 5.0 |
| Node fabric | 400G high-speed fabric for multi-node scaling |
| Form factor | 4U rackmount, optimised for air cooling |
| Storage | Encrypted NVMe on every node (AES-256 at rest, TLS 1.3 in transit) |
| Tenancy | Single-tenant: dedicated compute, storage and network |
Specifications as published by NVIDIA and Sapience AI for the configuration deployed. System memory per server is confirmed in your quote.
RTX PRO 6000 capacity is reserved by the cluster. Aggregate figures are the per-server specification multiplied by the server count.
| Cluster | GPUs | GPU memory | FP4 compute |
|---|---|---|---|
| 8 servers | 64 × RTX PRO 6000 | 6.1 TB | 256 PFLOPS |
| 24 servers | 192 × RTX PRO 6000 | 18.4 TB | 768 PFLOPS |
| 32 servers | 256 × RTX PRO 6000 | 24.6 TB | 1,024 PFLOPS |
As a rule of thumb, model weights take about one byte per parameter at FP8 and two at FP16/BF16. A 70-billion-parameter model at FP8 (about 70 GB of weights) fits on a single RTX PRO 6000 with room for the KV cache; larger models are sharded across the 768 GB of a server. The inference sizing guide works through the numbers.
RTX PRO 6000 Blackwell Server Edition is NVIDIA's platform for agentic AI, multi-application workflows and professional visualisation, and a cost-effective node for serving models.
Serving open-source or proprietary models from dedicated European servers, with a predictable cost per GPU-hour and no shared tenancy.
Inference →Adapter and parameter-efficient fine-tunes of models that fit within 96 GB per GPU or 768 GB per server, on terms from three months.
Fine-tuning →Agent workflows, multi-application pipelines, simulation and professional graphics that mix AI with rendering.
All solutions →How the two Blackwell nodes compare on memory, interconnect, terms and availability.
Read the comparison →Anything else: sales@sapienceai.eu or +421 233 329 562.
One cluster of 8 servers (64 GPUs). Terms start at three months, with longer terms on request.
The current capacity is fully reserved through 2026. Reservations for 2027 are open; tell us your start month and we confirm availability in the quote.
Per GPU-hour on a reserved term, quoted in writing. Rates are in US dollars and exclude VAT.
For moderate multi-node jobs, yes, on the 400G fabric. GPUs within a server communicate over PCIe Gen 5.0 rather than NVLink, so large-scale training is better served by HGX B300.
In a colocation data centre in Bratislava, Slovakia, on systems operated by Sapience. Customer data is processed and stored in the EU.
Pick the cluster size, the term and the start month; the quote confirms rate and availability in writing.