Run Mistral Nemo 12B on a VPS.
September 2026 · 148 of 280 plans fit · re-ranked daily
Apache-licensed 12B built with NVIDIA; a reliable creative-writing and chat workhorse. On a CPU-only VPS it needs about 10.3 GB of RAM: 7.5 GB of Q4_K_M weights, 1.3 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 148 of the 280 plans in our index fit; the cheapest comfortable pick is Contabo's Cloud VPS 8 · 8 vCPU · 24 GB at $16.27/mo, streaming an estimated 2.6–5.2 tok/s.
RAM needed · CPU inference
10.3 GB
- weights · Q4_K_M
- 7.5 GB
- KV cache · 8K context
- 1.3 GB
- OS + runtime headroom
- 1.5 GB
Other quants: Q8_0 weights 13 GB (near-lossless, about half of F16); F16 24.5 GB. Comfortable = 20% headroom over the total.
Model card
- Size
- 12B parameters
- Context
- 1M tokens
- Kind
- chat
- Released
- 2024-07
- License
- Apache-2.0
Facts fetched from Hugging Face on 3 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 360,703 downloads.
Best VPS plans for Mistral Nemo 12B
ranked by estimated tokens/s per dollar, comfortable fits first
- 1runs comfortably~3.9–7.8 tok/s$20.93/mo
- 2runs comfortably~2.6–5.2 tok/s$16.27/mo
- 3runs comfortably~2.6–5.2 tok/s$18.58/mo
- 4runs comfortably~2.6–5.2 tok/s$18.80/mo
- 5runs comfortably~4.8–9.7 tok/s$38.99/mo
- 6runs comfortably~3.2–6.5 tok/s$29.05/mo
Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbours are — treat these as order-of-magnitude.
Or buy hardware · 14 reference machines fit
all machines →- Raspberry Pi 5 (16 GB)16 GB · 17 GB/s · ~0.6–0.9 tok/s$305= 19 mo of VPS
- GeForce RTX 5060 Ti 16 GB (card only)16 GB · 448 GB/s · ~32–49 tok/s$805*= 50 mo of VPS
- Mac mini (M6, 16 GB)16 GB · 153 GB/s · ~11–17 tok/s$899= 55 mo of VPS
- GeForce RTX 5070 Ti 16 GB (card only)16 GB · 896 GB/s · ~63–97 tok/s$1,100*= 68 mo of VPS
- Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~11–17 tok/s$1,299= 80 mo of VPS
- GeForce RTX 5080 16 GB (card only)16 GB · 960 GB/s · ~68–104 tok/s$1,500*= 92 mo of VPS
Buy · per month
$8.78
$8.47 hardware + $0.31 power
Rent · per month
$16.27
Contabo Cloud VPS 8 · 8 vCPU · 24 GB
Break-even
The Raspberry Pi 5 (16 GB) pays for itself after 19 months of replacing the VPS — and it streams an estimated 0.6–0.9 tok/s against the VPS's CPU-only pace.
* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.
Frequently asked
- How much RAM does Mistral Nemo 12B need?
- About 10.3 GB for CPU inference at Q4_K_M: 7.5 GB of weights, 1.3 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (160 KB per token for this model).
- What is the cheapest VPS that can run Mistral Nemo 12B?
- Contabo Cloud VPS 8 · 8 vCPU · 24 GB (24 GB RAM, 8 vCPU) at $16.27/mo excl. VAT runs it comfortably as of 5 Sept 2026. Contabo Cloud VPS 6 · 6 vCPU · 12 GB at $8.72/mo is a tight fit.
- How fast will Mistral Nemo 12B run on a VPS without a GPU?
- Roughly 3.9–7.8 tok/s on the top pick. Token generation is bound by memory bandwidth — each token streams the full weights once — so more and dedicated vCPUs help; a GPU or Apple-silicon machine is 10–50× faster.