VPS Arena

Run DeepSeek-R1 Distill Qwen 32B on a VPS.

September 2026 · 68 of 280 plans fit · re-ranked daily

The strongest distill that still fits 32 GB; o1-mini-class on benchmarks at launch. On a CPU-only VPS it needs about 23.4 GB of RAM: 19.9 GB of Q4_K_M weights, 2 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 68 of the 280 plans in our index fit; the cheapest comfortable pick is Contabo's Cloud VPS 12 · 12 vCPU · 48 GB at $29.05/mo, streaming an estimated 1.2–2.4 tok/s.

RAM needed · CPU inference

23.4 GB

weights · Q4_K_M
19.9 GB
KV cache · 8K context
2 GB
OS + runtime headroom
1.5 GB

Other quants: Q8_0 weights 34.8 GB (near-lossless, about half of F16); F16 65.5 GB. Comfortable = 20% headroom over the total.

Model card

Size
33B parameters
Context
128K tokens
Kind
reasoning
Released
2025-01
License
MIT

Facts fetched from Hugging Face on 3 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 561,497 downloads.

ollama run deepseek-r1:32bmodel card ↗DeepSeek

Best VPS plans for DeepSeek-R1 Distill Qwen 32B

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    netcupRS 4000 G12 · 12 dedicated cores · 32 GB

    12 vCPU dedicated · 32 GB RAM · 1024 GB NVME

    runs comfortably~1.8–3.6 tok/s$38.99/mo
  2. 2
    ContaboCloud VPS 12 · 12 vCPU · 48 GB

    12 vCPU shared · 48 GB RAM · 400 GB SSD

    runs comfortably~1.2–2.4 tok/s$29.05/mo
  3. 3
    netcupVPS 4000 G12 · 12 vCore · 32 GB

    12 vCPU shared · 32 GB RAM · 1024 GB NVME

    runs comfortably~1.2–2.4 tok/s$31.66/mo
  4. 4
    HetznerCX53 · 16 vCPU · 32 GB

    16 vCPU shared · 32 GB RAM · 320 GB NVME

    runs comfortably~1.2–2.4 tok/s$34.27/mo
  5. 5
    ContaboCloud VPS 16 · 16 vCPU · 64 GB

    16 vCPU shared · 64 GB RAM · 500 GB SSD

    runs comfortably~1.2–2.4 tok/s$43/mo
  6. 6
    netcupRS 8000 G12 · 16 dedicated cores · 64 GB

    16 vCPU dedicated · 64 GB RAM · 2048 GB NVME

    runs comfortably~1.8–3.6 tok/s$69.70/mo

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbours are — treat these as order-of-magnitude.

Or buy hardware · 8 reference machines fit

all machines →
  • Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~4–6.2 tok/s · tight$1,299= 45 mo of VPS
  • Framework Desktop (Ryzen AI Max+ 395, 64 GB)64 GB · 256 GB/s · ~6.8–10 tok/s$1,959= 67 mo of VPS
  • Mac mini (M5 Pro, 64 GB)64 GB · 307 GB/s · ~8.1–12 tok/s$2,299= 79 mo of VPS
  • Mac Studio (M5 Max, 36 GB)36 GB · 460 GB/s · ~12–19 tok/s$2,499= 86 mo of VPS
  • Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~6.8–10 tok/s$3,449= 119 mo of VPS
  • NVIDIA DGX Spark (128 GB)128 GB · 273 GB/s · ~7.2–11 tok/s$4,699*= 10+ yrs of VPS

Buy · per month

$37.45

$36.08 hardware + $1.36 power

Rent · per month

$29.05

Contabo Cloud VPS 12 · 12 vCPU · 48 GB

Break-even

The Mac mini (M6, 32 GB) pays for itself after 47 months of replacing the VPS — and it streams an estimated 4–6.2 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does DeepSeek-R1 Distill Qwen 32B need?
About 23.4 GB for CPU inference at Q4_K_M: 19.9 GB of weights, 2 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (256 KB per token for this model).
What is the cheapest VPS that can run DeepSeek-R1 Distill Qwen 32B?
Contabo Cloud VPS 12 · 12 vCPU · 48 GB (48 GB RAM, 12 vCPU) at $29.05/mo excl. VAT runs it comfortably as of 5 Sept 2026. Contabo Cloud VPS 8 · 8 vCPU · 24 GB at $16.27/mo is a tight fit.
How fast will DeepSeek-R1 Distill Qwen 32B run on a VPS without a GPU?
Roughly 1.8–3.6 tok/s on the top pick. Token generation is bound by memory bandwidth — each token streams the full weights once — so more and dedicated vCPUs help; a GPU or Apple-silicon machine is 10–50× faster.
Same lineupDeepSeek-R1-0528 Qwen3 8BDeepSeek-R1 Distill Qwen 14BDeepSeek-R1 Distill Llama 70BSame size classQwen3 32BQwen2.5-Coder 32BGemma 3 27BMistral Small 3.2 24BVPS for AI agents (API-hosted models) →