GPU cloud for LLM inference in the EUServe models from dedicated European GPUs.

Run open-source or proprietary models on reserved NVIDIA Blackwell clusters in Bratislava: no shared tenancy, a predictable cost per GPU-hour, and requests and outputs that stay in the EU.

Most models
RTX PRO 6000
96 GB per GPU · 768 GB per server
The largest models
HGX B300
288 GB per GPU · 2.3 TB per server
Inference sizing

Which models fit on which GPU

For serving, the first question is whether the model's weights fit in GPU memory with room left for the KV cache. Weights take roughly 2 bytes per parameter at FP16/BF16, 1 at FP8 and 0.5 at FP4.

Model sizeFP16 / BF16FP8FP4
8B parameters≈ 16 GB≈ 8 GB≈ 4 GB
70B parameters≈ 140 GB≈ 70 GB≈ 35 GB
405B parameters≈ 810 GB≈ 405 GB≈ 203 GB

Against the hardware:

  • RTX PRO 6000: 96 GB per GPU, 768 GB per server. A 70B model at FP8 fits on one GPU with KV-cache headroom; a 405B model at FP8 fits across one server.
  • HGX B300: 288 GB per GPU, 2.3 TB per server, NVLink between GPUs. A 405B model at FP8 fits on two GPUs; the largest models and the highest throughput belong here.

Weights only. The KV cache grows with context length and concurrent requests and needs its own headroom; we size it with you from your traffic profile.

Predictable cost, no noisy neighbours

Inference capacity on Sapience is reserved for a term and priced per GPU-hour, so the cost of serving is known in advance. Every cluster is single-tenant, with dedicated compute, storage and network, so latency is not affected by other customers. See pricing for terms.

Inference inside the EU

Prompts and outputs often contain personal or confidential data. Serving from Bratislava, on systems operated by a European company under EU law, keeps that traffic and its logs in the EU. The security page sets out the controls and the status of ISO 27001 and SOC 2 Type II.

Questions

Inference questions.

Anything else: sales@sapienceai.eu or +421 233 329 562.

Can I run a 70B model on one GPU?

At FP8, yes: about 70 GB of weights fit in the 96 GB of an RTX PRO 6000, leaving headroom for the KV cache. At FP16 the same model needs two GPUs.

Which inference server can I use?

Any you choose. Nodes are bare metal, so you install your own serving stack and models, open-source or proprietary.

Is there per-token or per-hour pay-as-you-go pricing?

No. Capacity is reserved for a term and the rate is expressed per GPU-hour, quoted in writing.

When can inference capacity start?

RTX PRO 6000 is fully reserved through 2026, with 2027 reservations open. HGX B300 deliveries start from December 2026.

Size an inference cluster.

Send the models, context lengths and expected traffic. We recommend a node and quote it in writing.

Get a quote Talk to our team