all systems operational
Get started
RTX PRO 4500 / 6000 Blackwell · dedicated

AI inference dedicated servers

GPU passthrough for serious AI inference workloads. RTX PRO 4500 (32 GB VRAM) or RTX PRO 6000 (96 GB VRAM). Full CUDA, no MIG slicing.

32 / 96 GB
VRAM per GPU
PCIe 5.0 x16
full-speed passthrough
CUDA 12 +
pre-installed
from $550
/mo (RTX PRO 4500)

About AI inference server

Inference workloads want three things from a GPU host: enough VRAM to hold the model, full PCIe bandwidth to feed it, and zero contention from noisy neighbours. Our dedicated GPU tier gives you all three — no MIG slicing, no shared GPU with 8 other tenants.

RTX PRO 4500 Blackwell (32 GB VRAM) comfortably serves quantized 13-14B models (Llama 3, Mistral, Qwen at INT4/INT8) or full-precision 7-8B models. Typical throughput: 40-80 tokens/s single stream, 200-400 tokens/s batched. RTX PRO 6000 (96 GB VRAM) scales to quantized 70B or full 34B — 15-40 tokens/s single, 100-250 batched.

All builds ship with Ubuntu 22.04 + CUDA 12.4, cuDNN, TensorRT, PyTorch/JAX/vLLM pre-installed. Docker + NVIDIA Container Toolkit ready. IPMI 2.0 for out-of-band recovery. 1 TB NVMe Gen5 storage for model weights + hot cache.

For multi-GPU workloads (2× / 4× / 8× GPU per box, NVLink where supported) — custom-quoted per build. Multi-tenant training jobs, LoRA fine-tuning, LLM serving at scale — send the workload, we'll spec the box.

/faq

Frequently asked questions

Which model size fits in 32 GB VRAM?

Full FP16: 7-8B (Llama 3 8B, Mistral 7B). INT8: 13-14B. INT4: 30-34B. With CPU offload of some layers you can push 70B at INT4 but throughput drops significantly.

Is it worth going for RTX PRO 6000 vs multiple 4500s?

Depends. Single-model serving at low latency → prefer single 6000 (96 GB). Multi-tenant serving with many small models → multiple 4500s scale better economically.

Do you offer H100 or H200?

Not currently in stock. If your workload requires H100, we can source and rack them — custom-quoted, 2-3 weeks lead time.

Chat with us@hostfory