Run Qwen3 0.6B on a VPS.
September 2026 · 242 of 280 plans fit · re-ranked daily
Sub-gigabyte model for rerankers, intent detection and edge devices. On a CPU-only VPS it needs about 2.9 GB of RAM: 512 MB of Q4_K_M weights, 922 MB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 242 of the 280 plans in our index fit; the cheapest comfortable pick is OVHcloud's VPS-1 2027 · 2 vCPU / 4 GB at $5.35/mo, streaming an estimated 11–21 tok/s.
RAM needed · CPU inference
2.9 GB
- weights · Q4_K_M
- 512 MB
- KV cache · 8K context
- 922 MB
- OS + runtime headroom
- 1.5 GB
Other quants: Q8_0 weights 819 MB (near-lossless, about half of F16); F16 1.5 GB. Comfortable = 20% headroom over the total.
Model card
- Size
- 0.75B parameters
- Context
- 32K tokens
- Kind
- chat
- Released
- 2025-04
- License
- Apache-2.0
Facts fetched from Hugging Face on 3 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 22,741,013 downloads.
Best VPS plans for Qwen3 0.6B
ranked by estimated tokens/s per dollar, comfortable fits first
- 1runs comfortably~32–63 tok/s$8.72/mo
- 2runs comfortably~21–42 tok/s$6.39/mo
- 3runs comfortably~63–127 tok/s$20.93/mo
- 4runs comfortably~42–85 tok/s$16.27/mo
- 5runs comfortably~32–63 tok/s$12.49/mo
- 6runs comfortably~42–85 tok/s$18.58/mo
Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbours are — treat these as order-of-magnitude.
Or buy hardware · 14 reference machines fit
all machines →- Raspberry Pi 5 (16 GB)16 GB · 17 GB/s · ~9.8–15 tok/s$305= 57 mo of VPS
- GeForce RTX 5060 Ti 16 GB (card only)16 GB · 448 GB/s · ~100+ tok/s$805*= 10+ yrs of VPS
- Mac mini (M6, 16 GB)16 GB · 153 GB/s · ~100+ tok/s$899= 10+ yrs of VPS
- GeForce RTX 5070 Ti 16 GB (card only)16 GB · 896 GB/s · ~100+ tok/s$1,100*= 10+ yrs of VPS
- Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~100+ tok/s$1,299= 10+ yrs of VPS
- GeForce RTX 5080 16 GB (card only)16 GB · 960 GB/s · ~100+ tok/s$1,500*= 10+ yrs of VPS
Buy · per month
$8.78
$8.47 hardware + $0.31 power
Rent · per month
$5.35
OVHcloud VPS-1 2027 · 2 vCPU / 4 GB
Break-even
The Raspberry Pi 5 (16 GB) pays for itself after 61 months of replacing the VPS — and it streams an estimated 9.8–15 tok/s against the VPS's CPU-only pace.
* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.
Frequently asked
- How much RAM does Qwen3 0.6B need?
- About 2.9 GB for CPU inference at Q4_K_M: 512 MB of weights, 922 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (112 KB per token for this model).
- What is the cheapest VPS that can run Qwen3 0.6B?
- OVHcloud VPS-1 2027 · 2 vCPU / 4 GB (4 GB RAM, 2 vCPU) at $5.35/mo excl. VAT runs it comfortably as of 5 Sept 2026.
- How fast will Qwen3 0.6B run on a VPS without a GPU?
- Roughly 32–63 tok/s on the top pick. Token generation is bound by memory bandwidth — each token streams the full weights once — so more and dedicated vCPUs help; a GPU or Apple-silicon machine is 10–50× faster.