stepfun-ai / Step-Audio-EditX

huggingface.co
Total runs: 3.9K
24-hour runs: 86
7-day runs: -134
30-day runs: -6.8K
Model's Last Updated: February 14 2026
text-to-speech

Introduction of Step-Audio-EditX

Model Details of Step-Audio-EditX

Step-Audio-EditX

✨ Demo Page | 🌟 GitHub | 📑 Paper

Check our open-source repository https://github.com/stepfun-ai/Step-Audio-EditX for more details!

We are open-sourcing Step-Audio-EditX , a powerful LLM-based audio model specialized in expressive and iterative audio editing . It excels at editing emotion , speaking style , and paralinguistics , and also features robust zero-shot text-to-speech (TTS) capabilities.

Features
  • Zero-Shot TTS

    • Excellent zero-shot TTS cloning for Mandarin, English, Sichuanese, and Cantonese.
    • To use a dialect, just add a [Sichuanese] or [Cantonese] tag before your text.
  • Emotion and Speaking Style Editing

    • Remarkably effective iterative control over emotions and styles, supporting dozens of options for editing.
      • Emotion Editing : [ Angry , Happy , Sad , Excited , Fearful , Surprised , Disgusted , etc. ]
      • Speaking Style Editing: [ Act_coy , Older , Child , Whisper , Serious , Generous , Exaggerated , etc.]
      • Editing with more emotion and more speaking styles is on the way. Get Ready! 🚀
  • Paralinguistic Editing :

    • Precise control over 10 types of paralinguistic features for more natural, human-like, and expressive synthetic audio.
    • Supporting Tags:
      • [ Breathing , Laughter , Suprise-oh , Confirmation-en , Uhm , Suprise-ah , Suprise-wa , Sigh , Question-ei , Dissatisfaction-hnn ]

For more examples, see demo page .

Model Usage
📜 Requirements

The following table shows the requirements for running Step-Audio-EditX model:

Model Setting
(sample frequency)
GPU Optimal Memory
Step-Audio-EditX 41.6Hz 32GB
  • An NVIDIA GPU with CUDA support is required.
    • The model is tested on a single L40S GPU.
  • Tested operating system: Linux
🔧 Dependencies and Installation
git clone https://github.com/stepfun-ai/Step-Audio-EditX.git
conda create -n stepaudioedit python=3.10
conda activate stepaudioedit

cd Step-Audio-EditX
pip install -r requirements.txt

git lfs install
git clone https://huggingface.co/stepfun-ai/Step-Audio-Tokenizer
git clone https://huggingface.co/stepfun-ai/Step-Audio-EditX

After downloading the models, where_you_download_dir should have the following structure:

where_you_download_dir
├── Step-Audio-Tokenizer
├── Step-Audio-EditX
Run with Docker

You can set up the environment required for running Step-Audio using the provided Dockerfile.

# build docker
docker build . -t step-audio-editx

# run docker
docker run --rm --gpus all \
    -v /your/code/path:/app \
    -v /your/model/path:/model \
    -p 7860:7860 \
    step-audio-editx
Launch Web Demo

Start a local server for online inference. Assume you have one GPU with at least 32GB memory available and have already downloaded all the models.

# Step-Audio-EditX demo
python app.py --model-path where_you_download_dir --model-source local
Local Inference Demo

For optimal performance, keep audio under 30 seconds per inference.

# zero-shot cloning
python3 tts_infer.py \
    --model-path where_you_download_dir \
    --output-dir ./output \
    --prompt-text "your prompt text"\
    --prompt-audio your_prompt_audio_path \
    --generated-text "your target text" \
    --edit-type "clone"

# edit
python3 tts_infer.py \
    --model-path where_you_download_dir \
    --output-dir ./output \
    --prompt-text "your promt text" \
    --prompt-audio your_prompt_audio_path \
    --generated-text "" \ # for para-linguistic editing, you need to specify the generatedd text
    --edit-type "emotion" \
    --edit-info "sad" \
    --n-edit-iter 2
Citation
@misc{yan2025stepaudioeditxtechnicalreport,
      title={Step-Audio-EditX Technical Report}, 
      author={Chao Yan and Boyong Wu and Peng Yang and Pengfei Tan and Guoqiang Hu and Yuxin Zhang and Xiangyu and Zhang and Fei Tian and Xuerui Yang and Xiangyu Zhang and Daxin Jiang and Gang Yu},
      year={2025},
      eprint={2511.03601},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2511.03601}, 
}

Runs of stepfun-ai Step-Audio-EditX on huggingface.co

3.9K
Total runs
86
24-hour runs
15
3-day runs
-134
7-day runs
-6.8K
30-day runs

More Information About Step-Audio-EditX huggingface.co Model

More Step-Audio-EditX license Visit here:

https://choosealicense.com/licenses/apache-2.0

Step-Audio-EditX huggingface.co

Step-Audio-EditX huggingface.co is an AI model on huggingface.co that provides Step-Audio-EditX's model effect (), which can be used instantly with this stepfun-ai Step-Audio-EditX model. huggingface.co supports a free trial of the Step-Audio-EditX model, and also provides paid use of the Step-Audio-EditX. Support call Step-Audio-EditX model through api, including Node.js, Python, http.

Step-Audio-EditX huggingface.co Url

https://huggingface.co/stepfun-ai/Step-Audio-EditX

stepfun-ai Step-Audio-EditX online free

Step-Audio-EditX huggingface.co is an online trial and call api platform, which integrates Step-Audio-EditX's modeling effects, including api services, and provides a free online trial of Step-Audio-EditX, you can try Step-Audio-EditX online for free by clicking the link below.

stepfun-ai Step-Audio-EditX online free url in huggingface.co:

https://huggingface.co/stepfun-ai/Step-Audio-EditX

Step-Audio-EditX install

Step-Audio-EditX is an open source model from GitHub that offers a free installation service, and any user can find Step-Audio-EditX on GitHub to install. At the same time, huggingface.co provides the effect of Step-Audio-EditX install, users can directly use Step-Audio-EditX installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Step-Audio-EditX install url in huggingface.co:

https://huggingface.co/stepfun-ai/Step-Audio-EditX

Url of Step-Audio-EditX

Step-Audio-EditX huggingface.co Url

Provider of Step-Audio-EditX huggingface.co

stepfun-ai
ORGANIZATIONS

Other API from stepfun-ai

huggingface.co

Total runs: 600.2K
Run Growth: -100.8K
Growth Rate: -16.79%
Updated:February 04 2025
huggingface.co

Total runs: 41.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 29 2026
huggingface.co

Total runs: 4.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 02 2025
huggingface.co

Total runs: 83
Run Growth: -209
Growth Rate: -251.81%
Updated:January 14 2026