ACE-Step / acestep-v15-xl-sft-diffusers

huggingface.co
Total runs: 4.1K
24-hour runs: -18
7-day runs: 326
30-day runs: 2.0K
Model's Last Updated: June 03 2026
text-to-audio

Introduction of acestep-v15-xl-sft-diffusers

Model Details of acestep-v15-xl-sft-diffusers

ACE-Step v1.5 XL SFT Diffusers

Diffusers-format checkpoint of ACE-Step v1.5 XL SFT - the supervised fine-tuned 5B-parameter flow-matching DiT for text-to-music generation ( hidden_size=2560 , 32 layers, 32 heads; encoder_hidden_size=2048 on the condition encoder).

This repository is the official Diffusers-format version of the ACE-Step v1.5 XL SFT checkpoint. It can be loaded directly with AceStepPipeline , which is available in huggingface/diffusers .

Weights are produced by scripts/convert_ace_step_to_diffusers.py from the upstream release and packaged in the standard Diffusers pipeline layout ( model_index.json + one subdirectory per module), so the full pipeline can be loaded in a single from_pretrained call.

Usage

Install Diffusers from source until the next package release includes AceStepPipeline .

pip install git+https://github.com/huggingface/diffusers.git
import torch
import soundfile as sf
from diffusers import AceStepPipeline

pipe = AceStepPipeline.from_pretrained(
    "ACE-Step/acestep-v15-xl-sft-diffusers",
    torch_dtype=torch.bfloat16,
)
pipe = pipe.to("cuda")

# Long-form audio: enable VAE tiling to keep decode memory bounded.
pipe.vae.enable_tiling()

output = pipe(
    prompt="An upbeat synthwave track with driving drums and a catchy lead",
    lyrics="[Verse]\nNeon lights are calling me\n[Chorus]\nRide the wave tonight",
    audio_duration=30.0,
    num_inference_steps=8,
    guidance_scale=7.0,
    shift=3.0,
    generator=torch.Generator(device="cuda").manual_seed(42),
)

audio = output.audios[0]  # (channels, samples), 48 kHz
sf.write("acestep-xl-sft.wav", audio.T.cpu().float().numpy(), pipe.sample_rate)

Unlike the turbo checkpoint, XL SFT is not guidance-distilled. The pipeline uses ACE-Step's APG guidance path when guidance_scale > 1.0 ; guidance_scale=7.0 and shift=3.0 are the recommended defaults. You can increase num_inference_steps for slower, higher-quality sampling.

For batched prompts with padding and FlashAttention, use the variable-length backend:

pipe.transformer.set_attention_backend("flash_varlen")
pipe.condition_encoder.set_attention_backend("flash_varlen")

For single-prompt generation, the regular flash backend is also suitable.

Repository layout
├── model_index.json
├── transformer/        # AceStepTransformer1DModel (DiT, 5B params, bf16)
├── condition_encoder/  # AceStepConditionEncoder (with baked-in silence_latent)
├── audio_tokenizer/    # AceStepAudioTokenizer
├── audio_token_detokenizer/ # AceStepAudioTokenDetokenizer
├── vae/                # AutoencoderOobleck (48 kHz stereo)
├── text_encoder/       # Qwen3-Embedding-0.6B
├── tokenizer/          # Qwen3 tokenizer
├── scheduler/          # FlowMatchEulerDiscreteScheduler config
└── silence_latent.pt   # Raw reference (kept for debugging; not needed at runtime)
License
  • ACE-Step weights: MIT (same as upstream )
  • text_encoder/ (Qwen3-Embedding-0.6B): Apache 2.0 - redistributed per Qwen's license
Citation
@misc{gong2026acestep,
  title = {ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation},
  author = {Junmin Gong, Yulin Song, Wenxiao Zhao, Sen Wang, Shengyuan Xu, Jing Guo},
  howpublished = {\url{https://github.com/ace-step/ACE-Step-1.5}},
  year = {2026},
  note = {GitHub repository}
}

Runs of ACE-Step acestep-v15-xl-sft-diffusers on huggingface.co

4.1K
Total runs
-18
24-hour runs
177
3-day runs
326
7-day runs
2.0K
30-day runs

More Information About acestep-v15-xl-sft-diffusers huggingface.co Model

More acestep-v15-xl-sft-diffusers license Visit here:

https://choosealicense.com/licenses/mit

acestep-v15-xl-sft-diffusers huggingface.co

acestep-v15-xl-sft-diffusers huggingface.co is an AI model on huggingface.co that provides acestep-v15-xl-sft-diffusers's model effect (), which can be used instantly with this ACE-Step acestep-v15-xl-sft-diffusers model. huggingface.co supports a free trial of the acestep-v15-xl-sft-diffusers model, and also provides paid use of the acestep-v15-xl-sft-diffusers. Support call acestep-v15-xl-sft-diffusers model through api, including Node.js, Python, http.

acestep-v15-xl-sft-diffusers huggingface.co Url

https://huggingface.co/ACE-Step/acestep-v15-xl-sft-diffusers

ACE-Step acestep-v15-xl-sft-diffusers online free

acestep-v15-xl-sft-diffusers huggingface.co is an online trial and call api platform, which integrates acestep-v15-xl-sft-diffusers's modeling effects, including api services, and provides a free online trial of acestep-v15-xl-sft-diffusers, you can try acestep-v15-xl-sft-diffusers online for free by clicking the link below.

ACE-Step acestep-v15-xl-sft-diffusers online free url in huggingface.co:

https://huggingface.co/ACE-Step/acestep-v15-xl-sft-diffusers

acestep-v15-xl-sft-diffusers install

acestep-v15-xl-sft-diffusers is an open source model from GitHub that offers a free installation service, and any user can find acestep-v15-xl-sft-diffusers on GitHub to install. At the same time, huggingface.co provides the effect of acestep-v15-xl-sft-diffusers install, users can directly use acestep-v15-xl-sft-diffusers installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

acestep-v15-xl-sft-diffusers install url in huggingface.co:

https://huggingface.co/ACE-Step/acestep-v15-xl-sft-diffusers

Url of acestep-v15-xl-sft-diffusers

acestep-v15-xl-sft-diffusers huggingface.co Url

Provider of acestep-v15-xl-sft-diffusers huggingface.co

ACE-Step
ORGANIZATIONS

Other API from ACE-Step

huggingface.co

Total runs: 60.8K
Run Growth: 7.9K
Growth Rate: 13.02%
Updated:February 03 2026