MLX port of
Stability AI Stable Audio 3 (optimized)
.
Latent-diffusion text-to-audio with
mask-based inpainting
and continuation, with the DiT denoiser quantized to
INT8
for Apple Silicon.
What's in this bundle
Component
Format
Notes
DiT (Medium, 1.4B)
INT8
Diffusion Transformer denoiser, group-size 64
SAME-L encoder
FP32
Audio → latents (codec is precision-sensitive — differential attention cancels in FP16)
SAME-L decoder
FP32
Latents → 44.1 kHz stereo waveform
T5Gemma text encoder
FP16
Prompt conditioning
Codec stays FP32 because the SAME differential attention catastrophically cancels in FP16 (per Stability's own MLX runtime). T5Gemma stays FP16 — it's small relative to the DiT and quantization gives no speed-up on the short prompt encode pass.
Files
File
Size
Format
dit_medium/model.safetensors
1 GB
int8
same_l_encoder/model.safetensors
2 GB
fp32
same_l_decoder/model.safetensors
2 GB
fp32
t5gemma/model.safetensors
541 MB
fp16
Capabilities
Text-to-audio
generation (music + SFX depending on DiT specialisation)
Inpainting / region editing
via masked latent diffusion
Audio continuation
from a short prompt clip
Variable-length
generation up to several minutes
The
DiT-Small-Music-*
variant is music-specialised;
DiT-Small-SFX-*
is sound-effects specialised;
DiT-Medium-*
is the higher-quality general model.
Usage
This bundle is the quantized weights only — inference uses Stability AI's official pure-MLX runtime at
stable-audio-3/optimized/mlx
. At load time, each
(base.weight, base.scales, base.biases)
triplet is dequantized via
mlx.core.dequantize
back to FP16; codec and T5Gemma load as-is.
from huggingface_hub import snapshot_download
import mlx.core as mx
bundle = snapshot_download("aufklarer/Stable-Audio-3-DiT-Medium-MLX-8bit")
defload_component(comp_dir):
w = dict(mx.load(f"{comp_dir}/model.safetensors"))
bases = {k[:-7] for k in w if k.endswith(".scales")
iff"{k[:-7]}.weight"in w andf"{k[:-7]}.biases"in w}
out = {}
for k, v in w.items():
if k.endswith((".scales", ".biases")) and k.rsplit(".", 1)[0] in bases:
continueif k.endswith(".weight") and k[:-7] in bases:
base = k[:-7]
out[k] = mx.dequantize(w[f"{base}.weight"], w[f"{base}.scales"],
w[f"{base}.biases"], group_size=64, bits=8)
else:
out[k] = v
return out
Plug the rehydrated dict into the matching model class from
stable-audio-3/optimized/mlx/models/defs/
.
Stability AI Community License
— free for non-commercial research and for commercial use up to the revenue threshold defined by Stability AI; see the
license text
. T5Gemma component additionally inherits the
Gemma Terms of Use
.
Runs of aufklarer Stable-Audio-3-DiT-Medium-MLX-8bit on huggingface.co
62
Total runs
-7
24-hour runs
-12
3-day runs
-18
7-day runs
-18
30-day runs
More Information About Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co Model
More Stable-Audio-3-DiT-Medium-MLX-8bit license Visit here:
Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co is an AI model on huggingface.co that provides Stable-Audio-3-DiT-Medium-MLX-8bit's model effect (), which can be used instantly with this aufklarer Stable-Audio-3-DiT-Medium-MLX-8bit model. huggingface.co supports a free trial of the Stable-Audio-3-DiT-Medium-MLX-8bit model, and also provides paid use of the Stable-Audio-3-DiT-Medium-MLX-8bit. Support call Stable-Audio-3-DiT-Medium-MLX-8bit model through api, including Node.js, Python, http.
Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co is an online trial and call api platform, which integrates Stable-Audio-3-DiT-Medium-MLX-8bit's modeling effects, including api services, and provides a free online trial of Stable-Audio-3-DiT-Medium-MLX-8bit, you can try Stable-Audio-3-DiT-Medium-MLX-8bit online for free by clicking the link below.
aufklarer Stable-Audio-3-DiT-Medium-MLX-8bit online free url in huggingface.co:
Stable-Audio-3-DiT-Medium-MLX-8bit is an open source model from GitHub that offers a free installation service, and any user can find Stable-Audio-3-DiT-Medium-MLX-8bit on GitHub to install. At the same time, huggingface.co provides the effect of Stable-Audio-3-DiT-Medium-MLX-8bit install, users can directly use Stable-Audio-3-DiT-Medium-MLX-8bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Stable-Audio-3-DiT-Medium-MLX-8bit install url in huggingface.co: