Choose the node by method and model size
Memory use in fine-tuning depends mostly on the method. Parameter-efficient methods such as LoRA train small adapter matrices while the base model stays frozen, so memory is dominated by the base weights. QLoRA goes further by holding the frozen base model in 4-bit precision. Full-parameter fine-tuning updates every weight and needs gradients and optimiser state for all of them.
| Method | Memory driver | Typical fit |
|---|---|---|
| LoRA / adapters | Frozen base weights (≈2 B/param at BF16) plus small adapters | RTX PRO 6000: models up to tens of billions of parameters per GPU, larger ones across a 768 GB server |
| QLoRA | Base weights at 4-bit (≈0.5 B/param) plus adapters | RTX PRO 6000: larger models on fewer GPUs |
| Full-parameter | Weights, gradients and optimiser state (≈16 B/param with Adam) | HGX B300: 288 GB per GPU, NVLink, sharded across the cluster |
Rules of thumb before activations; sequence length and batch size add to them. We size the cluster with your team.
Terms that match the project
Fine-tuning projects are often shorter than pre-training programmes. RTX PRO 6000 clusters can be reserved from three months, with longer terms on request and 2027 start dates open. HGX B300 is reserved on multi-year terms, with deliveries from December 2026. Either way the rate is expressed per GPU-hour, so you can compare it. See pricing.
Your data stays yours
Fine-tuning data is usually proprietary: support tickets, contracts, clinical notes, code. On Sapience, clusters are single-tenant, storage is encrypted (AES-256 at rest, TLS 1.3 in transit) and customer data is processed and stored in the EU. Details on the security page.
Fine-tuning questions.
Anything else: sales@sapienceai.eu or +421 233 329 562.
Which node should I choose for LoRA fine-tuning?
Usually RTX PRO 6000: 96 GB per GPU and 768 GB per server hold the frozen base weights of most open-weight models, on terms from three months.
When do I need HGX B300?
For full-parameter fine-tuning of large models, where weights, gradients and optimiser state run to terabytes and fast GPU-to-GPU links matter.
Can I fine-tune proprietary models?
Yes. You run open-source or proprietary models on bare metal you control; Sapience operates the infrastructure, not your models.
What is the minimum commitment?
One cluster of 8 servers (64 GPUs). RTX PRO 6000 terms start at three months; HGX B300 is reserved on multi-year terms.
Size a fine-tuning cluster.
Tell us the base model, the method and the data volume; we recommend a node and quote it in writing.