VPS Arena

Run BGE-M3 on a VPS.

September 2026 · 242 of 280 plans fit · re-ranked daily

Multilingual dense + sparse embeddings for hybrid search. On a CPU-only VPS it needs about 2.6 GB of RAM: 1.1 GB of weights in the model's native format and 1.5 GB for the OS and runtime. 242 of the 280 plans in our index fit; the cheapest comfortable pick is OVHcloud's VPS-1 2027 · 2 vCPU / 4 GB at $5.35/mo.

RAM needed · CPU inference

2.6 GB

weights · native format
1.1 GB
OS + runtime headroom
1.5 GB

Model card

Size
0.568B parameters
Context
8K tokens
Kind
embedding
Released
2024-01
License
MIT

Facts fetched from Hugging Face on 3 Sept 2026: sizes from bits-per-weight · 37,079,940 downloads.

ollama run bge-m3model card ↗various

Best VPS plans for BGE-M3

ranked by price among plans that fit

  1. 1
    OVHcloudVPS-1 2027 · 2 vCPU / 4 GB

    2 vCPU shared · 4 GB RAM · 40 GB NVME

    runs comfortably$5.35/mo
  2. 2
    netcupVPS 500 G12 · 2 vCore · 4 GB

    2 vCPU shared · 4 GB RAM · 128 GB NVME

    runs comfortably$5.78/mo
  3. 3
    HetznerCX23 · 2 vCPU · 4 GB

    2 vCPU shared · 4 GB RAM · 40 GB NVME

    runs comfortably$6.38/mo
  4. 4
    ContaboCloud VPS 4 · 4 vCPU · 8 GB

    4 vCPU shared · 8 GB RAM · 100 GB SSD

    runs comfortably$6.39/mo
  5. 5
    HetznerCAX11 · 2 vCPU · 4 GB

    2 vCPU shared · 4 GB RAM · 40 GB NVME

    runs comfortably$6.96/mo
  6. 6
    ContaboCloud VPS 6 · 6 vCPU · 12 GB

    6 vCPU shared · 12 GB RAM · 200 GB SSD

    runs comfortably$8.72/mo

Frequently asked

How much RAM does BGE-M3 need?
About 2.6 GB for CPU inference at Q4_K_M: 1.1 GB of weights, 0 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (0 KB per token for this model).
What is the cheapest VPS that can run BGE-M3?
OVHcloud VPS-1 2027 · 2 vCPU / 4 GB (4 GB RAM, 2 vCPU) at $5.35/mo excl. VAT runs it comfortably as of 5 Sept 2026.
How fast will BGE-M3 run on a VPS without a GPU?
Embedding models are small and CPU-friendly; throughput scales with vCPUs rather than RAM.
Same lineupnomic-embed-text v1.5Whisper large-v3VPS for AI agents (API-hosted models) →