CosyVoice3 (Mandarin) — CoreML Models for FluidAudio
CoreML conversions of CosyVoice3's four inference stages, frozen to the exact
shapes the
FluidAudio
Swift
package's
CosyVoice3TtsManager
loads at runtime. Targets Apple Silicon
(M-series) with the Neural Engine for LLM + HiFT, CPU for Flow.
A default voice ships in
voices/
so the repo is self-contained. Additional
voices (as they're extracted) live in the companion repo
FluidInference/cosyvoice3-voices-zh
.
Shipping configuration (frozen)
Each model is shipped in two formats:
.mlpackage
(source, portable) and
.mlmodelc
(pre-compiled for macOS 14 / iOS 17 + Apple Silicon). Swift can
load either;
.mlmodelc
skips the one-time compile step on first use
(~20-30 s for Flow without it).
Single-step AR decode, 768-slot KV cache, 24 layers × 2 KV heads × 64 dim
fp16
Flow-N250-fp16
CPU + GPU
Speech-token → mel (80-bin, 24 kHz), N_total=250
fp16 (pure CPU overflows fused LayerNorm → NaN; ANE refuses to compile; GPU path uses fp32 accumulators internally and is stable)
HiFT-T500-fp16
CPU + ANE
Mel → 24 kHz PCM, T=500 frames
fp16
Total disk footprint (
.mlmodelc
+
.mlpackage
+ runtime tables): ~6.6 GB on
disk. If you only need one format, delete the other after download.
Runtime tables
embeddings/
embeddings-runtime-fp32.safetensors
— 542 MB. Qwen2
model.embed_tokens.weight
at
runtime
(post-
.float()
) dtype. Required for bit-exact parity with
the Python reference — shipping raw
.pt
weights introduces ~4.7e-4 error
through the HuggingFace dtype round-trip. Swift mmaps this file.
aishell3-zh-SSB*.safetensors
— 10 AISHELL-3 speakers bootstrapped via
verify/bootstrap_aishell3_voices.py
(5 female + 5 male, north + south
accents). See
aishell3-bootstrap.json
for per-voice provenance.
Each
.safetensors
ships with a
.json
prompt-text sidecar and follows the
schema documented in the companion
cosyvoice3-voices-zh
repo.
CosyVoice3-0.5B-coreml huggingface.co is an AI model on huggingface.co that provides CosyVoice3-0.5B-coreml's model effect (), which can be used instantly with this FluidInference CosyVoice3-0.5B-coreml model. huggingface.co supports a free trial of the CosyVoice3-0.5B-coreml model, and also provides paid use of the CosyVoice3-0.5B-coreml. Support call CosyVoice3-0.5B-coreml model through api, including Node.js, Python, http.
CosyVoice3-0.5B-coreml huggingface.co is an online trial and call api platform, which integrates CosyVoice3-0.5B-coreml's modeling effects, including api services, and provides a free online trial of CosyVoice3-0.5B-coreml, you can try CosyVoice3-0.5B-coreml online for free by clicking the link below.
FluidInference CosyVoice3-0.5B-coreml online free url in huggingface.co:
CosyVoice3-0.5B-coreml is an open source model from GitHub that offers a free installation service, and any user can find CosyVoice3-0.5B-coreml on GitHub to install. At the same time, huggingface.co provides the effect of CosyVoice3-0.5B-coreml install, users can directly use CosyVoice3-0.5B-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
CosyVoice3-0.5B-coreml install url in huggingface.co: