Run Qwen3 30B-A3B (MoE) on a VPS.
September 2026 · 68 of 280 plans fit · re-ranked daily
The CPU darling: 30B of knowledge, 3B active per token, so it streams like a small model on a 24–32 GB VPS. On a CPU-only VPS it needs about 20.9 GB of RAM: 18.6 GB of Q4_K_M weights, 819 MB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 68 of the 280 plans in our index fit; the cheapest comfortable pick is Contabo's Cloud VPS 12 · 12 vCPU · 48 GB at $29.05/mo, streaming an estimated 6–12 tok/s.
RAM needed · CPU inference
20.9 GB
- weights · Q4_K_M
- 18.6 GB
- KV cache · 8K context
- 819 MB
- OS + runtime headroom
- 1.5 GB
Other quants: Q8_0 weights 32.5 GB (near-lossless, about half of F16); F16 61.1 GB. Comfortable = 20% headroom over the total.
Model card
- Size
- 31B total · 3.3B active
- Context
- 40K tokens
- Kind
- chat
- Released
- 2025-04
- License
- Apache-2.0
Facts fetched from Hugging Face on 3 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 2,375,624 downloads.
Best VPS plans for Qwen3 30B-A3B (MoE)
ranked by estimated tokens/s per dollar, comfortable fits first
- 1runs comfortably~9–18 tok/s$38.99/mo
- 2runs comfortably~6–12 tok/s$29.05/mo
- 3runs comfortably~6–12 tok/s$31.66/mo
- 4runs comfortably~6–12 tok/s$34.27/mo
- 5runs comfortably~6–12 tok/s$43/mo
- 6runs comfortably~9–18 tok/s$69.70/mo
Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated, ×0.5 for MoE routing overhead), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbours are — treat these as order-of-magnitude.
Or buy hardware · 8 reference machines fit
all machines →- Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~20–31 tok/s$1,299= 45 mo of VPS
- Framework Desktop (Ryzen AI Max+ 395, 64 GB)64 GB · 256 GB/s · ~34–52 tok/s$1,959= 67 mo of VPS
- Mac mini (M5 Pro, 64 GB)64 GB · 307 GB/s · ~40–62 tok/s$2,299= 79 mo of VPS
- Mac Studio (M5 Max, 36 GB)36 GB · 460 GB/s · ~60–93 tok/s$2,499= 86 mo of VPS
- Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~34–52 tok/s$3,449= 119 mo of VPS
- NVIDIA DGX Spark (128 GB)128 GB · 273 GB/s · ~36–55 tok/s$4,699*= 10+ yrs of VPS
Buy · per month
$37.45
$36.08 hardware + $1.36 power
Rent · per month
$29.05
Contabo Cloud VPS 12 · 12 vCPU · 48 GB
Break-even
The Mac mini (M6, 32 GB) pays for itself after 47 months of replacing the VPS — and it streams an estimated 20–31 tok/s against the VPS's CPU-only pace.
* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.
Frequently asked
- How much RAM does Qwen3 30B-A3B (MoE) need?
- About 20.9 GB for CPU inference at Q4_K_M: 18.6 GB of weights, 819 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (96 KB per token for this model).
- What is the cheapest VPS that can run Qwen3 30B-A3B (MoE)?
- Contabo Cloud VPS 12 · 12 vCPU · 48 GB (48 GB RAM, 12 vCPU) at $29.05/mo excl. VAT runs it comfortably as of 5 Sept 2026. Contabo Cloud VPS 8 · 8 vCPU · 24 GB at $16.27/mo is a tight fit.
- How fast will Qwen3 30B-A3B (MoE) run on a VPS without a GPU?
- Roughly 9–18 tok/s on the top pick. Token generation is bound by memory bandwidth — each token streams the active-expert weights once — so more and dedicated vCPUs help; a GPU or Apple-silicon machine is 10–50× faster.