Run BGE-M3 on a VPS.
September 2026 · 242 of 280 plans fit · re-ranked daily
Multilingual dense + sparse embeddings for hybrid search. On a CPU-only VPS it needs about 2.6 GB of RAM: 1.1 GB of weights in the model's native format and 1.5 GB for the OS and runtime. 242 of the 280 plans in our index fit; the cheapest comfortable pick is OVHcloud's VPS-1 2027 · 2 vCPU / 4 GB at $5.35/mo.
RAM needed · CPU inference
2.6 GB
- weights · native format
- 1.1 GB
- OS + runtime headroom
- 1.5 GB
Model card
- Size
- 0.568B parameters
- Context
- 8K tokens
- Kind
- embedding
- Released
- 2024-01
- License
- MIT
Facts fetched from Hugging Face on 3 Sept 2026: sizes from bits-per-weight · 37,079,940 downloads.
Best VPS plans for BGE-M3
ranked by price among plans that fit
- 1runs comfortably$5.35/mo
- 2runs comfortably$5.78/mo
- 3runs comfortably$6.38/mo
- 4runs comfortably$6.39/mo
- 5runs comfortably$6.96/mo
- 6runs comfortably$8.72/mo
Frequently asked
- How much RAM does BGE-M3 need?
- About 2.6 GB for CPU inference at Q4_K_M: 1.1 GB of weights, 0 MB of KV cache at 8,192 tokens of context, and 1.5 GB of headroom. Longer contexts need more KV cache (0 KB per token for this model).
- What is the cheapest VPS that can run BGE-M3?
- OVHcloud VPS-1 2027 · 2 vCPU / 4 GB (4 GB RAM, 2 vCPU) at $5.35/mo excl. VAT runs it comfortably as of 5 Sept 2026.
- How fast will BGE-M3 run on a VPS without a GPU?
- Embedding models are small and CPU-friendly; throughput scales with vCPUs rather than RAM.