VPS Arena

Run Whisper large-v3 on a VPS.

September 2026 · 207 of 280 plans fit · re-ranked daily

Speech-to-text via whisper.cpp; ~3 GB in F16, CPU-friendly for batch transcription. On a CPU-only VPS it needs about 4.6 GB of RAM: 3.1 GB of weights in the model's native format and 1.5 GB for the OS and runtime. 207 of the 280 plans in our index fit; the cheapest comfortable pick is Contabo's Cloud VPS 4 · 4 vCPU · 8 GB at $6.39/mo.

RAM needed · CPU inference

4.6 GB

weights · native format
3.1 GB
OS + runtime headroom
1.5 GB

Model card

Size
1.54B parameters
Context
tokens
Kind
speech
Released
2023-11
License
Apache-2.0

Facts fetched from Hugging Face on 3 Sept 2026: sizes from bits-per-weight · 4,972,119 downloads.

Best VPS plans for Whisper large-v3

ranked by price among plans that fit

  1. 1
    ContaboCloud VPS 4 · 4 vCPU · 8 GB

    4 vCPU shared · 8 GB RAM · 100 GB SSD

    runs comfortably$6.39/mo
  2. 2
    ContaboCloud VPS 6 · 6 vCPU · 12 GB

    6 vCPU shared · 12 GB RAM · 200 GB SSD

    runs comfortably$8.72/mo
  3. 3
    HetznerCX33 · 4 vCPU · 8 GB

    4 vCPU shared · 8 GB RAM · 80 GB NVME

    runs comfortably$9.87/mo
  4. 4
    OVHcloudVPS-2 2027 · 4 vCPU / 8 GB

    4 vCPU shared · 8 GB RAM · 75 GB NVME

    runs comfortably$10/mo
  5. 5
    OVHcloudVPS-2 LZ 2027 · 4 vCPU / 8 GB

    4 vCPU shared · 8 GB RAM · 75 GB NVME

    runs comfortably$10/mo
  6. 6
    netcupVPS 1000 G12 · 4 vCore · 8 GB

    4 vCPU shared · 8 GB RAM · 256 GB NVME

    runs comfortably$10.12/mo

Frequently asked

How much RAM does Whisper large-v3 need?
About 4.6 GB for CPU inference at Q4_K_M: 3.1 GB of weights, 0 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (0 KB per token for this model).
What is the cheapest VPS that can run Whisper large-v3?
Contabo Cloud VPS 4 · 4 vCPU · 8 GB (8 GB RAM, 4 vCPU) at $6.39/mo excl. VAT runs it comfortably as of 5 Sept 2026.
How fast will Whisper large-v3 run on a VPS without a GPU?
Speech models are small and CPU-friendly; throughput scales with vCPUs rather than RAM.
Same lineupnomic-embed-text v1.5BGE-M3VPS for AI agents (API-hosted models) →