AI inference dedicated servers
GPU passthrough for serious AI inference workloads. RTX PRO 4500 (32 GB VRAM) or RTX PRO 6000 (96 GB VRAM). Full CUDA, no MIG slicing.
Best-fit dedicated plans
These are the closest stock plans in our catalogue for AI inference server workloads. Fully customizable — port, storage, RAM, DDoS profile.
GPU VPS · RTX PRO 4500
16 vCPU · 128 GB DDR5 · 1 TB NVMe Gen5 · RTX PRO 4500 (32 GB VRAM) full passthrough
GPU VPS · RTX PRO 6000
32 vCPU · 256 GB DDR5 · 2 TB NVMe Gen5 · RTX PRO 6000 (96 GB VRAM) full passthrough
2× / 4× / 8× GPU clusters
Custom-quoted per build — multiple Blackwell / L40S / L4 GPUs, NVLink where supported
About AI inference server
Inference workloads want three things from a GPU host: enough VRAM to hold the model, full PCIe bandwidth to feed it, and zero contention from noisy neighbours. Our dedicated GPU tier gives you all three — no MIG slicing, no shared GPU with 8 other tenants.
RTX PRO 4500 Blackwell (32 GB VRAM) comfortably serves quantized 13-14B models (Llama 3, Mistral, Qwen at INT4/INT8) or full-precision 7-8B models. Typical throughput: 40-80 tokens/s single stream, 200-400 tokens/s batched. RTX PRO 6000 (96 GB VRAM) scales to quantized 70B or full 34B — 15-40 tokens/s single, 100-250 batched.
All builds ship with Ubuntu 22.04 + CUDA 12.4, cuDNN, TensorRT, PyTorch/JAX/vLLM pre-installed. Docker + NVIDIA Container Toolkit ready. IPMI 2.0 for out-of-band recovery. 1 TB NVMe Gen5 storage for model weights + hot cache.
For multi-GPU workloads (2× / 4× / 8× GPU per box, NVLink where supported) — custom-quoted per build. Multi-tenant training jobs, LoRA fine-tuning, LLM serving at scale — send the workload, we'll spec the box.
Frequently asked questions
Which model size fits in 32 GB VRAM?
Full FP16: 7-8B (Llama 3 8B, Mistral 7B). INT8: 13-14B. INT4: 30-34B. With CPU offload of some layers you can push 70B at INT4 but throughput drops significantly.
Is it worth going for RTX PRO 6000 vs multiple 4500s?
Depends. Single-model serving at low latency → prefer single 6000 (96 GB). Multi-tenant serving with many small models → multiple 4500s scale better economically.
Do you offer H100 or H200?
Not currently in stock. If your workload requires H100, we can source and rack them — custom-quoted, 2-3 weeks lead time.