8-bit MLX-compatible quantization for Apple Silicon.
MLX port of
openbmb/VoxCPM2
— a 2B-parameter
multilingual diffusion-autoregressive TTS model with
48 kHz
studio-quality output,
voice cloning, and instruction-driven voice design.
Bundle size
: 2.95 GB
Capabilities
30 languages
including English, Chinese, Indonesian, Japanese, Korean
48 kHz output
Zero-shot synthesis
— generate speech from text alone
Voice cloning
— clone a target speaker from a single reference clip
Voice design
— natural-language style control (e.g.
"young female voice, warm and gentle"
)
Ultimate cloning
— reference audio + transcript for prosody-preserving cloning
Streaming generation
— patch-level decoding for low-latency synthesis
Quantization
Format
: MLX
QuantizedLinear
, 8 bits per element, group size 64,
per-group scales and biases stored as float16.
What is quantized
: All
Linear
layers inside the LM backbones
(
base_lm
,
residual_lm
), the DiT estimator decoder,
feat_encoder.encoder
, and the top-level projection heads
(
enc_to_lm_proj
,
lm_to_dit_proj
,
res_to_dit_proj
,
fusion_concat_proj
,
stop_proj
,
stop_head
,
fsq_layer.*
, the
time/delta-time MLPs).
What stays bfloat16
: All
audio_vae.*
weights, RMSNorm /
LayerNorm gain tensors, RoPE lookup tables, Snake
alpha
,
embedding tables, and 1-D parameters.
Round-trip fidelity vs bf16
: mean relative L2 error
0.53 %
, worst-layer relative L2
0.78 %
(
stop_head
).
40 % smaller than bf16 with negligible quality impact in practice.
import VoxCPM2TTS
let model =tryawaitVoxCPM2TTSModel.fromPretrained(
modelId: "aufklarer/VoxCPM2-MLX-int8"
)
let audio =tryawait model.generate(text: "Hello from VoxCPM2.", language: "english")
Apache 2.0 — inherited from the upstream openbmb/VoxCPM2 model.
Responsible use
Voice cloning capability is included. Users are responsible for obtaining consent
for any voice that is cloned and for not using the model to impersonate individuals
without their permission, generate disinformation, or commit fraud.
Runs of aufklarer VoxCPM2-MLX-int8 on huggingface.co
266
Total runs
-1
24-hour runs
-9
3-day runs
73
7-day runs
-57
30-day runs
More Information About VoxCPM2-MLX-int8 huggingface.co Model
VoxCPM2-MLX-int8 huggingface.co is an AI model on huggingface.co that provides VoxCPM2-MLX-int8's model effect (), which can be used instantly with this aufklarer VoxCPM2-MLX-int8 model. huggingface.co supports a free trial of the VoxCPM2-MLX-int8 model, and also provides paid use of the VoxCPM2-MLX-int8. Support call VoxCPM2-MLX-int8 model through api, including Node.js, Python, http.
VoxCPM2-MLX-int8 huggingface.co is an online trial and call api platform, which integrates VoxCPM2-MLX-int8's modeling effects, including api services, and provides a free online trial of VoxCPM2-MLX-int8, you can try VoxCPM2-MLX-int8 online for free by clicking the link below.
aufklarer VoxCPM2-MLX-int8 online free url in huggingface.co:
VoxCPM2-MLX-int8 is an open source model from GitHub that offers a free installation service, and any user can find VoxCPM2-MLX-int8 on GitHub to install. At the same time, huggingface.co provides the effect of VoxCPM2-MLX-int8 install, users can directly use VoxCPM2-MLX-int8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.