VPS Arena

Run Mistral Nemo 12B on a VPS.

September 2026 · 148 of 280 plans fit · re-ranked daily

Apache-licensed 12B built with NVIDIA; a reliable creative-writing and chat workhorse. On a CPU-only VPS it needs about 10.3 GB of RAM: 7.5 GB of Q4_K_M weights, 1.3 GB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 148 of the 280 plans in our index fit; the cheapest comfortable pick is Contabo's Cloud VPS 8 · 8 vCPU · 24 GB at $16.27/mo, streaming an estimated 2.6–5.2 tok/s.

RAM needed · CPU inference

10.3 GB

weights · Q4_K_M
7.5 GB
KV cache · 8K context
1.3 GB
OS + runtime headroom
1.5 GB

Other quants: Q8_0 weights 13 GB (near-lossless, about half of F16); F16 24.5 GB. Comfortable = 20% headroom over the total.

Model card

Size
12B parameters
Context
1M tokens
Kind
chat
Released
2024-07
License
Apache-2.0

Facts fetched from Hugging Face on 3 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 360,703 downloads.

ollama run mistral-nemo:12bmodel card ↗Mistral AI

Best VPS plans for Mistral Nemo 12B

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    netcupRS 2000 G12 · 8 dedicated cores · 16 GB

    8 vCPU dedicated · 16 GB RAM · 512 GB NVME

    runs comfortably~3.9–7.8 tok/s$20.93/mo
  2. 2
    ContaboCloud VPS 8 · 8 vCPU · 24 GB

    8 vCPU shared · 24 GB RAM · 300 GB SSD

    runs comfortably~2.6–5.2 tok/s$16.27/mo
  3. 3
    HetznerCX43 · 8 vCPU · 16 GB

    8 vCPU shared · 16 GB RAM · 160 GB NVME

    runs comfortably~2.6–5.2 tok/s$18.58/mo
  4. 4
    netcupVPS 2000 G12 · 8 vCore · 16 GB

    8 vCPU shared · 16 GB RAM · 512 GB NVME

    runs comfortably~2.6–5.2 tok/s$18.80/mo
  5. 5
    netcupRS 4000 G12 · 12 dedicated cores · 32 GB

    12 vCPU dedicated · 32 GB RAM · 1024 GB NVME

    runs comfortably~4.8–9.7 tok/s$38.99/mo
  6. 6
    ContaboCloud VPS 12 · 12 vCPU · 48 GB

    12 vCPU shared · 48 GB RAM · 400 GB SSD

    runs comfortably~3.2–6.5 tok/s$29.05/mo

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbours are — treat these as order-of-magnitude.

Or buy hardware · 14 reference machines fit

all machines →
  • Raspberry Pi 5 (16 GB)16 GB · 17 GB/s · ~0.6–0.9 tok/s$305= 19 mo of VPS
  • GeForce RTX 5060 Ti 16 GB (card only)16 GB · 448 GB/s · ~32–49 tok/s$805*= 50 mo of VPS
  • Mac mini (M6, 16 GB)16 GB · 153 GB/s · ~11–17 tok/s$899= 55 mo of VPS
  • GeForce RTX 5070 Ti 16 GB (card only)16 GB · 896 GB/s · ~63–97 tok/s$1,100*= 68 mo of VPS
  • Mac mini (M6, 32 GB)32 GB · 153 GB/s · ~11–17 tok/s$1,299= 80 mo of VPS
  • GeForce RTX 5080 16 GB (card only)16 GB · 960 GB/s · ~68–104 tok/s$1,500*= 92 mo of VPS

Buy · per month

$8.78

$8.47 hardware + $0.31 power

Rent · per month

$16.27

Contabo Cloud VPS 8 · 8 vCPU · 24 GB

Break-even

The Raspberry Pi 5 (16 GB) pays for itself after 19 months of replacing the VPS — and it streams an estimated 0.6–0.9 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does Mistral Nemo 12B need?
About 10.3 GB for CPU inference at Q4_K_M: 7.5 GB of weights, 1.3 GB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (160 KB per token for this model).
What is the cheapest VPS that can run Mistral Nemo 12B?
Contabo Cloud VPS 8 · 8 vCPU · 24 GB (24 GB RAM, 8 vCPU) at $16.27/mo excl. VAT runs it comfortably as of 5 Sept 2026. Contabo Cloud VPS 6 · 6 vCPU · 12 GB at $8.72/mo is a tight fit.
How fast will Mistral Nemo 12B run on a VPS without a GPU?
Roughly 3.9–7.8 tok/s on the top pick. Token generation is bound by memory bandwidth — each token streams the full weights once — so more and dedicated vCPUs help; a GPU or Apple-silicon machine is 10–50× faster.
Same lineupMinistral 8BMistral Small 3.2 24BDevstral Small 24BSame size classQwen3 14BGemma 3 12BPhi-4 14BDeepSeek-R1 Distill Qwen 14BVPS for AI agents (API-hosted models) →