Which VPS can run this model?
Start your search from the model, not the server. For 30 open-weight models we compute the RAM a CPU-only VPS needs at Q4_K_M, pick the cheapest plans that fit from all 280 tracked offers, estimate tokens per second from memory bandwidth, and price the local-hardware alternative — September 2026 data, refreshed daily.
Up to 4B
Qwen · chat
Qwen3 0.6B
0.75B parameters · 32K context
2.9 GB
RAM needed
from $5.35/mo
~11–21 tok/s · 242 plans fit
Gemma · chat
Gemma 3 1B
1B parameters · 32K context
2.5 GB
RAM needed
from $5.35/mo
~7.9–16 tok/s · 242 plans fit
Meta Llama · chat
Llama 3.2 1B
1.24B parameters · 128K context
2.6 GB
RAM needed
from $5.35/mo
~6.4–13 tok/s · 242 plans fit
Qwen · chat
Qwen3 1.7B
2.03B parameters · 32K context
3.7 GB
RAM needed
from $6.39/mo
~7.8–16 tok/s · 242 plans fit
Meta Llama · chat
Llama 3.2 3B
3.21B parameters · 128K context
4.4 GB
RAM needed
from $6.39/mo
~4.9–9.9 tok/s · 207 plans fit
Phi · chat
Phi-4-mini 3.8B
3.84B parameters · 128K context
5 GB
RAM needed
from $6.39/mo
~4.1–8.2 tok/s · 207 plans fit
Qwen · chat
Qwen3 4B
4.02B parameters · 40K context
5.1 GB
RAM needed
from $6.39/mo
~3.9–7.9 tok/s · 207 plans fit
Gemma · chat · vision
Gemma 3 4B
4.3B parameters · 128K context
5.1 GB
RAM needed
from $6.39/mo
~3.7–7.4 tok/s · 207 plans fit
7–9B
Qwen · code
Qwen2.5-Coder 7B
7.62B parameters · 128K context
6.6 GB
RAM needed
from $6.39/mo
~2.1–4.2 tok/s · 207 plans fit
Mistral · chat
Ministral 8B
8.02B parameters · 32K context
7.5 GB
RAM needed
from $8.72/mo
~3–5.9 tok/s · 207 plans fit
Meta Llama · chat
Llama 3.1 8B
8.03B parameters · 128K context
7.4 GB
RAM needed
from $8.72/mo
~3–5.9 tok/s · 207 plans fit
Qwen · chat
Qwen3 8B
8.19B parameters · 40K context
7.6 GB
RAM needed
from $8.72/mo
~2.9–5.8 tok/s · 207 plans fit
DeepSeek R1 · reasoning
DeepSeek-R1-0528 Qwen3 8B
8.19B parameters · 128K context
7.6 GB
RAM needed
from $8.72/mo
~2.9–5.8 tok/s · 207 plans fit
12–15B
Gemma · chat · vision
Gemma 3 12B
12B parameters · 128K context
11.8 GB
RAM needed
from $16.27/mo
~2.6–5.2 tok/s · 148 plans fit
Mistral · chat
Mistral Nemo 12B
12B parameters · 1M context
10.3 GB
RAM needed
from $16.27/mo
~2.6–5.2 tok/s · 148 plans fit
Phi · reasoning
Phi-4 14B
15B parameters · 16K context
12.2 GB
RAM needed
from $16.27/mo
~2.2–4.3 tok/s · 138 plans fit
Qwen · chat
Qwen3 14B
15B parameters · 40K context
11.8 GB
RAM needed
from $16.27/mo
~2.1–4.3 tok/s · 148 plans fit
DeepSeek R1 · reasoning
DeepSeek-R1 Distill Qwen 14B
15B parameters · 128K context
12 GB
RAM needed
from $16.27/mo
~2.1–4.3 tok/s · 148 plans fit
24–32B
Mistral · code
Devstral Small 24B
24B parameters · 128K context
17.1 GB
RAM needed
from $16.27/mo
~1.3–2.7 tok/s · 68 plans fit
Mistral · chat · vision
Mistral Small 3.2 24B
24B parameters · 128K context
17.1 GB
RAM needed
from $16.27/mo
~1.3–2.6 tok/s · 68 plans fit
Gemma · chat · vision
Gemma 3 27B
27B parameters · 128K context
22 GB
RAM needed
from $29.05/mo
~1.4–2.9 tok/s · 68 plans fit
Qwen · chat
Qwen3 32B
33B parameters · 40K context
23.3 GB
RAM needed
from $29.05/mo
~1.2–2.4 tok/s · 68 plans fit
Qwen · code
Qwen2.5-Coder 32B
33B parameters · 128K context
23.4 GB
RAM needed
from $29.05/mo
~1.2–2.4 tok/s · 68 plans fit
DeepSeek R1 · reasoning
DeepSeek-R1 Distill Qwen 32B
33B parameters · 128K context
23.4 GB
RAM needed
from $29.05/mo
~1.2–2.4 tok/s · 68 plans fit
70B+
Mixture-of-experts
gpt-oss · reasoning
gpt-oss-20b (MoE)
21B total · 3.6B active · 128K context
14 GB
RAM needed
from $16.27/mo
~4.4–8.7 tok/s · 138 plans fit
Qwen · chat
Qwen3 30B-A3B (MoE)
31B total · 3.3B active · 40K context
20.9 GB
RAM needed
from $29.05/mo
~6–12 tok/s · 68 plans fit
Meta Llama · chat · vision
Llama 4 Scout (109B MoE)
109B total · 17B active · 10.24M context
68.4 GB
RAM needed
from $115.06/mo
~1.2–2.3 tok/s · 2 plans fit
gpt-oss · reasoning
gpt-oss-120b (MoE)
117B total · 5.1B active · 128K context
65.5 GB
RAM needed
from $115.06/mo
~4.2–8.5 tok/s · 2 plans fit
Embeddings & speech
Embeddings & speech · embedding
nomic-embed-text v1.5
0.14B parameters · 2K context
2.1 GB
RAM needed
from $5.35/mo
242 plans fit
Embeddings & speech · embedding
BGE-M3
0.568B parameters · 8K context
2.6 GB
RAM needed
from $5.35/mo
242 plans fit
Embeddings & speech · speech
Whisper large-v3
1.54B parameters · — context
4.6 GB
RAM needed
from $6.39/mo
207 plans fit
How these numbers are computed
RAM needed
Weights at Q4_K_M (4.85 bits per weight × parameters) + KV cache for a 8,192-token context (from each model's layer/head geometry) + 1.5 GB for the OS and runtime. A plan is “comfortable” at 20% headroom, “tight” at exactly enough.
Tokens per second
CPU generation is memory-bandwidth-bound: every token streams the active weights once. We assume 4 GB/s per shared vCPU (max 40) and 6 GB/s per dedicated vCPU (max 60), divide by active weight bytes, and show a ±20–40% band. Treat it as an order of magnitude, not a benchmark.
Catalog
Model facts are re-fetched weekly from Hugging Face — parameter counts and licenses from the model API, exact GGUF file sizes from the quant repo, KV-cache geometry from the GGUF header (last refresh 3 Sept 2026). Only the one-line descriptions are hand-written. Prompt processing, batching and GPU offload are out of scope — this is the honest CPU-only floor.