VPS Arena

Run gpt-oss-120b (MoE) on a VPS.

September 2026 · 2 of 280 plans fit · re-ranked daily

o4-mini-class reasoning in ~65 GB of MXFP4 weights with only 5B active — the case for a 96–128 GB dedicated box. On a CPU-only VPS it needs about 65.5 GB of RAM: 63.4 GB of weights in the model's native format, 614 MB of KV cache for an 8,192-token context and 1.5 GB for the OS and runtime. 2 of the 280 plans in our index fit; the cheapest comfortable pick is Contabo's Cloud VPS Plus 18 · 18 vCPU · 96 GB at $115.06/mo, streaming an estimated 4.2–8.5 tok/s.

RAM needed · CPU inference

65.5 GB

weights · native format
63.4 GB
KV cache · 8K context
614 MB
OS + runtime headroom
1.5 GB

Model card

Size
117B total · 5.1B active
Context
128K tokens
Kind
reasoning
Released
2025-08
License
Apache-2.0

Facts fetched from Hugging Face on 3 Sept 2026: exact GGUF file sizes, KV geometry from the GGUF header · 5,376,994 downloads.

ollama run gpt-oss:120bmodel card ↗OpenAI

Best VPS plans for gpt-oss-120b (MoE)

ranked by estimated tokens/s per dollar, comfortable fits first

  1. 1
    ContaboCloud VPS Plus 18 · 18 vCPU · 96 GB

    18 vCPU shared · 96 GB RAM · 900 GB NVME

    runs comfortably~4.2–8.5 tok/s$115.06/mo
  2. 2
    LinodeLinode 90GB

    4 vCPU dedicated · 90 GB RAM · 90 GB SSD

    runs comfortably~2.5–5.1 tok/s$240/mo

Speed = effective memory bandwidth ÷ active weight bytes (4 GB/s per shared vCPU, 6 per dedicated, ×0.5 for MoE routing overhead), shown as a band. Real numbers depend on the host CPU generation, AVX-512/AMX support and how noisy the neighbours are — treat these as order-of-magnitude.

Or buy hardware · 3 reference machines fit

all machines →
  • Framework Desktop (Ryzen AI Max+ 395, 128 GB)128 GB · 256 GB/s · ~24–36 tok/s$3,449= 30 mo of VPS
  • NVIDIA DGX Spark (128 GB)128 GB · 273 GB/s · ~25–39 tok/s$4,699*= 41 mo of VPS
  • Mac Studio (M5 Ultra, 96 GB)96 GB · 1200 GB/s · ~100+ tok/s · tight$5,499= 48 mo of VPS

Buy · per month

$100.47

$95.81 hardware + $4.67 power

Rent · per month

$115.06

Contabo Cloud VPS Plus 18 · 18 vCPU · 96 GB

Break-even

The Framework Desktop (Ryzen AI Max+ 395, 128 GB) pays for itself after 31 months of replacing the VPS — and it streams an estimated 24–36 tok/s against the VPS's CPU-only pace.

* approximate: August 2026 US retail median rather than list price. Local speed = peak bandwidth × 0.7 efficiency (0.35 on CPU-only boards) ÷ active weight bytes; unified-memory machines are assumed to give models 75% of their RAM. GPU cards need a host PC that is not included in the price.

Frequently asked

How much RAM does gpt-oss-120b (MoE) need?
About 65.5 GB for CPU inference at Q4_K_M: 63.4 GB of weights, 614 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (72 KB per token for this model).
What is the cheapest VPS that can run gpt-oss-120b (MoE)?
Contabo Cloud VPS Plus 18 · 18 vCPU · 96 GB (96 GB RAM, 18 vCPU) at $115.06/mo excl. VAT runs it comfortably as of 5 Sept 2026.
How fast will gpt-oss-120b (MoE) run on a VPS without a GPU?
Roughly 4.2–8.5 tok/s on the top pick. Token generation is bound by memory bandwidth — each token streams the active-expert weights once — so more and dedicated vCPUs help; a GPU or Apple-silicon machine is 10–50× faster.
Same lineupgpt-oss-20b (MoE)Same size classLlama 4 Scout (109B MoE)Qwen3 30B-A3B (MoE)VPS for AI agents (API-hosted models) →