aufklarer / Stable-Audio-3-DiT-Medium-MLX-8bit

huggingface.co
Total runs: 62
24-hour runs: -7
7-day runs: -18
30-day runs: -18
Model's Last Updated: May 28 2026
text-to-audio

Introduction of Stable-Audio-3-DiT-Medium-MLX-8bit

Model Details of Stable-Audio-3-DiT-Medium-MLX-8bit

Stable-Audio-3-DiT-Medium-MLX-8bit

MLX port of Stability AI Stable Audio 3 (optimized) . Latent-diffusion text-to-audio with mask-based inpainting and continuation, with the DiT denoiser quantized to INT8 for Apple Silicon.

What's in this bundle
Component Format Notes
DiT (Medium, 1.4B) INT8 Diffusion Transformer denoiser, group-size 64
SAME-L encoder FP32 Audio → latents (codec is precision-sensitive — differential attention cancels in FP16)
SAME-L decoder FP32 Latents → 44.1 kHz stereo waveform
T5Gemma text encoder FP16 Prompt conditioning

Codec stays FP32 because the SAME differential attention catastrophically cancels in FP16 (per Stability's own MLX runtime). T5Gemma stays FP16 — it's small relative to the DiT and quantization gives no speed-up on the short prompt encode pass.

Files
File Size Format
dit_medium/model.safetensors 1 GB int8
same_l_encoder/model.safetensors 2 GB fp32
same_l_decoder/model.safetensors 2 GB fp32
t5gemma/model.safetensors 541 MB fp16
Capabilities
  • Text-to-audio generation (music + SFX depending on DiT specialisation)
  • Inpainting / region editing via masked latent diffusion
  • Audio continuation from a short prompt clip
  • Variable-length generation up to several minutes

The DiT-Small-Music-* variant is music-specialised; DiT-Small-SFX-* is sound-effects specialised; DiT-Medium-* is the higher-quality general model.

Usage

This bundle is the quantized weights only — inference uses Stability AI's official pure-MLX runtime at stable-audio-3/optimized/mlx . At load time, each (base.weight, base.scales, base.biases) triplet is dequantized via mlx.core.dequantize back to FP16; codec and T5Gemma load as-is.

from huggingface_hub import snapshot_download
import mlx.core as mx

bundle = snapshot_download("aufklarer/Stable-Audio-3-DiT-Medium-MLX-8bit")

def load_component(comp_dir):
    w = dict(mx.load(f"{comp_dir}/model.safetensors"))
    bases = {k[:-7] for k in w if k.endswith(".scales")
              if f"{k[:-7]}.weight" in w and f"{k[:-7]}.biases" in w}
    out = {}
    for k, v in w.items():
        if k.endswith((".scales", ".biases")) and k.rsplit(".", 1)[0] in bases:
            continue
        if k.endswith(".weight") and k[:-7] in bases:
            base = k[:-7]
            out[k] = mx.dequantize(w[f"{base}.weight"], w[f"{base}.scales"],
                                   w[f"{base}.biases"], group_size=64, bits=8)
        else:
            out[k] = v
    return out

Plug the rehydrated dict into the matching model class from stable-audio-3/optimized/mlx/models/defs/ .

Source
License

Stability AI Community License — free for non-commercial research and for commercial use up to the revenue threshold defined by Stability AI; see the license text . T5Gemma component additionally inherits the Gemma Terms of Use .

Runs of aufklarer Stable-Audio-3-DiT-Medium-MLX-8bit on huggingface.co

62
Total runs
-7
24-hour runs
-12
3-day runs
-18
7-day runs
-18
30-day runs

More Information About Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co Model

More Stable-Audio-3-DiT-Medium-MLX-8bit license Visit here:

https://choosealicense.com/licenses/stability-ai-community-license

Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co

Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co is an AI model on huggingface.co that provides Stable-Audio-3-DiT-Medium-MLX-8bit's model effect (), which can be used instantly with this aufklarer Stable-Audio-3-DiT-Medium-MLX-8bit model. huggingface.co supports a free trial of the Stable-Audio-3-DiT-Medium-MLX-8bit model, and also provides paid use of the Stable-Audio-3-DiT-Medium-MLX-8bit. Support call Stable-Audio-3-DiT-Medium-MLX-8bit model through api, including Node.js, Python, http.

Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co Url

https://huggingface.co/aufklarer/Stable-Audio-3-DiT-Medium-MLX-8bit

aufklarer Stable-Audio-3-DiT-Medium-MLX-8bit online free

Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co is an online trial and call api platform, which integrates Stable-Audio-3-DiT-Medium-MLX-8bit's modeling effects, including api services, and provides a free online trial of Stable-Audio-3-DiT-Medium-MLX-8bit, you can try Stable-Audio-3-DiT-Medium-MLX-8bit online for free by clicking the link below.

aufklarer Stable-Audio-3-DiT-Medium-MLX-8bit online free url in huggingface.co:

https://huggingface.co/aufklarer/Stable-Audio-3-DiT-Medium-MLX-8bit

Stable-Audio-3-DiT-Medium-MLX-8bit install

Stable-Audio-3-DiT-Medium-MLX-8bit is an open source model from GitHub that offers a free installation service, and any user can find Stable-Audio-3-DiT-Medium-MLX-8bit on GitHub to install. At the same time, huggingface.co provides the effect of Stable-Audio-3-DiT-Medium-MLX-8bit install, users can directly use Stable-Audio-3-DiT-Medium-MLX-8bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Stable-Audio-3-DiT-Medium-MLX-8bit install url in huggingface.co:

https://huggingface.co/aufklarer/Stable-Audio-3-DiT-Medium-MLX-8bit

Url of Stable-Audio-3-DiT-Medium-MLX-8bit

Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co Url

Provider of Stable-Audio-3-DiT-Medium-MLX-8bit huggingface.co

aufklarer
ORGANIZATIONS

Other API from aufklarer

huggingface.co

Total runs: 2.5K
Run Growth: 2.3K
Growth Rate: 94.66%
Updated:September 16 2025
huggingface.co

Total runs: 151
Run Growth: 93
Growth Rate: 61.59%
Updated:October 15 2025