On‑device multilingual TTS model converted to Core ML for Apple platforms.
This is a hand‑port of
Supertone Supertonic‑3 v1.7.3
from ONNX → PyTorch → Core ML, suitable for FluidAudio's TTS pipeline on
macOS/iOS. 31 languages, 44.1 kHz output, flow‑matching diffusion with
classifier‑free guidance (8 denoising steps).
text_encoder
— token embeddings → contextual text features
[B, 256, T]
.
duration_predictor
— predicts utterance duration from text features.
vector_estimator
— flow‑matching diffusion in latent space
(8 steps, classifier‑free guidance via batch‑2 duplication, ConvNeXt + cross‑attention to text + style attention).
End‑to‑end on M2: ≈ 0.74 s to synthesize 6.32 s of audio for a single English
sentence (RTFx ≈ 8.5×), 8 denoising steps. Output verified against
FluidAudio Parakeet TDT ASR.
Note on
vector_estimator
: 100 % of its ops are ANE‑eligible after
the float‑mask + precompute refactor, but Apple's ANECCompile currently
returns opaque error 11 on this graph and silently falls back to CPU/GPU.
See
coreml/trials.md
in the conversion repo for the full investigation.
Files
Both
.mlpackage
(Core ML source bundle, includes weights + spec) and the
precompiled
.mlmodelc
(ready for direct
MLModel(contentsOf:)
load) are
shipped — use
.mlmodelc
to skip the on‑device compile step on first load.
TextEncoder.mlpackage
/
TextEncoder.mlmodelc
— fixed
T=128
text input.
DurationPredictor.mlpackage
/
DurationPredictor.mlmodelc
— fixed
T=128
text input.
VectorEstimator.mlpackage
/
VectorEstimator.mlmodelc
—
latent.L
and
text.T
as RangeDim(17..512), FP16 weights (122 MB).
For the curious / for sanity checking, this repo ships a small self‑contained
script
infer.py
that loads all four modules directly via
coremltools
and
writes a 44.1 kHz WAV. No external repo clone required.
# 1. Download the repo (e.g. via huggingface_hub or `git lfs clone`).
git lfs clone https://huggingface.co/FluidInference/supertonic-3-coreml
cd supertonic-3-coreml
# 2. Install the 3 deps (macOS, Python 3.11+ recommended).
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# 3. Synthesize.
python infer.py "Hello, world." --voice-style voice_styles/M1.json -o hello.wav
python infer.py "Bonjour le monde." --lang fr --voice-style voice_styles/M1.json -o fr.wav
# Use the int8-quantized VectorEstimator (62 MB instead of 122 MB).
python infer.py "Hello, int8 build." --vector-estimator VectorEstimator_int8.mlpackage -o int8.wav
# Optional: pick a compute unit explicitly.
python infer.py "Test" --compute-units CPU_AND_NE -o ne.wav
The Python script loads
.mlpackage
(which is what
coremltools
accepts);
the
.mlmodelc
bundles are for direct Swift / Objective‑C use
(
MLModel(contentsOf:)
) where they skip the on‑device compile step.
Production (Swift / FluidAudio)
For production use, the FluidAudio Swift framework handles model loading,
text frontend, batching, chunking, and the diffusion / vocoder loop.
Swift (FluidAudio)
import AVFoundation
import FluidAudio
Task {
// Download and load Supertonic-3 models (first run only)let models =tryawaitSupertonic3Models.downloadAndLoad()
// Initialize the TTS managerlet tts =Supertonic3Manager(config: .default)
tryawait tts.initialize(models: models)
// Synthesize speech for some text with a voice stylelet style =tryVoiceStyle.load(path: "voice_styles/M1.json")
let audio =tryawait tts.synthesize(text: "Hello, world.", style: style)
// audio.samples is 44.1 kHz Float32 PCM in [-1, 1]tryAudioWriter.writeWav(audio.samples, sampleRate: 44_100, to: "hello.wav")
tts.cleanup()
}
44.1 kHz output is high quality but heavier than 16/22.05 kHz TTS — plan
for the bandwidth and storage cost.
vector_estimator
currently runs on CPU + GPU instead of ANE due to an
Apple‑side ANE compiler limitation (see
Performance
).
Text frontend currently uses fixed
T=128
token windows; longer text
must be segmented by the caller.
License
OpenRAIL‑M (inherited from upstream
Supertone/supertonic-3
).
The Core ML conversion tooling and FluidAudio integration are MIT‑licensed.
See the
FluidAudio repository
for details and usage guidance.
Runs of FluidInference supertonic-3-coreml on huggingface.co
179
Total runs
0
24-hour runs
-1
3-day runs
42
7-day runs
50
30-day runs
More Information About supertonic-3-coreml huggingface.co Model
supertonic-3-coreml huggingface.co is an AI model on huggingface.co that provides supertonic-3-coreml's model effect (), which can be used instantly with this FluidInference supertonic-3-coreml model. huggingface.co supports a free trial of the supertonic-3-coreml model, and also provides paid use of the supertonic-3-coreml. Support call supertonic-3-coreml model through api, including Node.js, Python, http.
supertonic-3-coreml huggingface.co is an online trial and call api platform, which integrates supertonic-3-coreml's modeling effects, including api services, and provides a free online trial of supertonic-3-coreml, you can try supertonic-3-coreml online for free by clicking the link below.
FluidInference supertonic-3-coreml online free url in huggingface.co:
supertonic-3-coreml is an open source model from GitHub that offers a free installation service, and any user can find supertonic-3-coreml on GitHub to install. At the same time, huggingface.co provides the effect of supertonic-3-coreml install, users can directly use supertonic-3-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
supertonic-3-coreml install url in huggingface.co: