mlx-community / bge-m3-mlx-8bit

huggingface.co
Total runs: 1.9K
24-hour runs: -11
7-day runs: 53
30-day runs: -488
Model's Last Updated: March 16 2026
feature-extraction

Introduction of bge-m3-mlx-8bit

Model Details of bge-m3-mlx-8bit

BGE-M3 MLX (8-bit Quantized)

This is the BAAI/bge-m3 model converted to MLX format with 8-bit quantization for Apple Silicon.

Model Description

BGE-M3 is a versatile embedding model capable of:

  • Dense retrieval
  • Sparse retrieval
  • Multi-vector (ColBERT) retrieval

This 8-bit quantized version offers the best quality among quantized variants.

Model Details
Property Value
Architecture XLM-RoBERTa
Precision 8-bit (affine quantization)
Embedding Dimension 1024
Max Sequence Length 8192
Model Size ~592 MB
Quantization Group Size 64
Languages 100+ languages
Size Comparison
Version Size Compression
FP16 1.1 GB -
8-bit 592 MB 46%
6-bit 457 MB 58%
4-bit 321 MB 71%
Usage
With MLX
from mlx_embeddings.utils import load_model, load_tokenizer
import mlx.core as mx

model_path = "mlx-community/bge-m3-mlx-8bit"

# Load model and tokenizer
model = load_model(model_path)
tokenizer = load_tokenizer(model_path)

# Generate embeddings
text = "Hello, world!"
tokens = tokenizer.encode(text)
input_ids = mx.array([tokens])
output = model(input_ids)
embedding = output.last_hidden_state.mean(axis=1)  # Mean pooling

print(f"Embedding shape: {embedding.shape}")  # (1, 1024)
With oMLX
curl http://127.0.0.1:8000/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model": "bge-m3-mlx-8bit", "input": "Your text here"}'
Quantization Details
  • Method : Affine quantization
  • Bits per weight : 8
  • Group size : 64
  • Source : Converted from MLX FP16 version
Recommended Use Cases
  • Quality-critical applications
  • Production deployments where quality is prioritized
  • Best choice when memory allows ~600MB
License

MIT license (inherited from BAAI/bge-m3)

Citation
@article{bge_m3,
  title={BGE M3-Embedding: Accurate, Efficient and Versatile Text Embedding},
  author={Chen, Jianlv and Xiao, Shitao and Zhang, Peitian and Luo, Kun and Zhang, Zheng},
  journal={arXiv preprint arXiv:2402.03216},
  year={2024}
}
Disclaimer

This is an unofficial MLX conversion of the BAAI/bge-m3 model. For the original model, see BAAI/bge-m3 .

Runs of mlx-community bge-m3-mlx-8bit on huggingface.co

1.9K
Total runs
-11
24-hour runs
-9
3-day runs
53
7-day runs
-488
30-day runs

More Information About bge-m3-mlx-8bit huggingface.co Model

More bge-m3-mlx-8bit license Visit here:

https://choosealicense.com/licenses/mit

bge-m3-mlx-8bit huggingface.co

bge-m3-mlx-8bit huggingface.co is an AI model on huggingface.co that provides bge-m3-mlx-8bit's model effect (), which can be used instantly with this mlx-community bge-m3-mlx-8bit model. huggingface.co supports a free trial of the bge-m3-mlx-8bit model, and also provides paid use of the bge-m3-mlx-8bit. Support call bge-m3-mlx-8bit model through api, including Node.js, Python, http.

mlx-community bge-m3-mlx-8bit online free

bge-m3-mlx-8bit huggingface.co is an online trial and call api platform, which integrates bge-m3-mlx-8bit's modeling effects, including api services, and provides a free online trial of bge-m3-mlx-8bit, you can try bge-m3-mlx-8bit online for free by clicking the link below.

mlx-community bge-m3-mlx-8bit online free url in huggingface.co:

https://huggingface.co/mlx-community/bge-m3-mlx-8bit

bge-m3-mlx-8bit install

bge-m3-mlx-8bit is an open source model from GitHub that offers a free installation service, and any user can find bge-m3-mlx-8bit on GitHub to install. At the same time, huggingface.co provides the effect of bge-m3-mlx-8bit install, users can directly use bge-m3-mlx-8bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

bge-m3-mlx-8bit install url in huggingface.co:

https://huggingface.co/mlx-community/bge-m3-mlx-8bit

Url of bge-m3-mlx-8bit

Provider of bge-m3-mlx-8bit huggingface.co

mlx-community
ORGANIZATIONS

Other API from mlx-community