The
first ONNX export of the standard
htdemucs
(non-FT) model
on
the Hugging Face Hub. Runs in
onnxruntime
on CPU out of the box, and
on CoreML / CUDA / DirectML with a one-line provider change.
No PyTorch required at inference.
This repo is the single-file companion to
StemSplitio/htdemucs-ft-onnx
.
You get all 4 stems out of one 316 MB
.onnx
file (
htdemucs.onnx
),
or 166 MB if you grab the fp16weights variant. The FT bag is higher
quality; this single model is ~30% faster and uses 1 session instead of 4.
infer.py
— pure-numpy reference inference (~200 lines, no torch).
requirements.txt
— three small packages, no PyTorch.
Quality
The official
htdemucs
model is the precursor to
htdemucs_ft
— same
architecture, single set of weights instead of 4 specialist sub-models.
On MUSDB18-HQ:
Metric
htdemucs
(this)
htdemucs_ft
(4-bag)
Median vocals SDR
~8.8 dB
9.19 dB
Median drums SDR
~9.5 dB
10.11 dB
Total model size
316 MB
1.26 GB
Sessions to load
1
4
Speed vs the bag
~1.4× faster
baseline
Parity vs PyTorch fp32 (random input, 7.8 s segment):
htdemucs.onnx
max abs diff:
6.62 × 10⁻⁴
htdemucs_fp16weights.onnx
max abs diff (vs fp32 weights):
4.6 × 10⁻⁵
Both well within the 1e-3 publish threshold.
Performance
Single 7.8 s segment, Apple M4 Pro CPU:
Variant
RAM
Latency
RTF
htdemucs.onnx
(fp32)
~1.1 GB
~1.6 s
0.20
htdemucs_fp16weights.onnx
~1.1 GB
~1.6 s
0.20
For comparison:
htdemucs_ft
(4-session bag)
~4.0 GB
~6.4 s
0.49
CUDA / DirectML / CoreML EPs are typically ≥ 5× faster on real GPUs.
Quick start
Python
import soundfile as sf
import infer
audio, sr = sf.read("your-song.mp3", dtype="float32", always_2d=True)
stems = infer.separate(audio.T, sr,
model_path=infer.DEFAULT_MODEL,
providers=["CPUExecutionProvider"])
for stem, arr in stems.items():
sf.write(f"{stem}.wav", arr.T, sr)
Stereo, 44.1 kHz, 7.8 s segment. Values in [-1, 1].
Output
stems
(1, 4, 2, 343980)
float32
Stems in order
[drums, bass, other, vocals]
. All 4 are real predictions (unlike the FT specialists).
For longer audio, chunk with overlap-add — see
infer.py::separate
for
a working 60-line implementation.
Tooling —
demucs-onnx
Python package
This model can be run (and re-exported from PyTorch) via the open-source
demucs-onnx
Python package
on PyPI. It auto-downloads from this repo on first use, so you don't
have to clone or wrangle file paths.
The export pipeline lives in the open-source
demucs-onnx
package at
demucs_onnx/export/
.
It applies four patches to make
torch.onnx.export
work on htdemucs:
Complex-typed
torch.stft
outputs →
Conv1d
with sin/cos kernels.
model.segment
fractions.Fraction
→ plain
float
.
random.randrange
in transformer pos-embedding → hardcoded
shift=0
.
aten::_native_multi_head_attention
(no ONNX symbolic) → drop-in
nn.MultiheadAttention.forward
built from
Linear
/
bmm
/
softmax
.
These are the four blockers every previous community attempt at "demucs
onnx" stalled on. See the
README of the demucs-onnx package
for the full write-up with code references.
Related work
Sibling ONNX repos from the same export pipeline:
Repo
Format
Stems
Use when
htdemucs-onnx
(this)
Single file
4
Faster startup, fewer sessions, ~30% lower latency than the FT bag.
Don't want to bundle a 316 MB model in your app, manage a GPU pool, or
write overlap-add chunking? Use the
StemSplit API
instead — same model under the hood, hosted for you, with credits and a
dashboard.
This repo is
MIT-licensed
, matching the original HT-Demucs.
@inproceedings{rouard2023hybrid,
title = {Hybrid Transformers for Music Source Separation},
author = {Rouard, Simon and Massa, Francisco and D{\'e}fossez, Alexandre},
booktitle = {ICASSP},
year = {2023}
}
htdemucs-onnx huggingface.co is an AI model on huggingface.co that provides htdemucs-onnx's model effect (), which can be used instantly with this adowu htdemucs-onnx model. huggingface.co supports a free trial of the htdemucs-onnx model, and also provides paid use of the htdemucs-onnx. Support call htdemucs-onnx model through api, including Node.js, Python, http.
htdemucs-onnx huggingface.co is an online trial and call api platform, which integrates htdemucs-onnx's modeling effects, including api services, and provides a free online trial of htdemucs-onnx, you can try htdemucs-onnx online for free by clicking the link below.
adowu htdemucs-onnx online free url in huggingface.co:
htdemucs-onnx is an open source model from GitHub that offers a free installation service, and any user can find htdemucs-onnx on GitHub to install. At the same time, huggingface.co provides the effect of htdemucs-onnx install, users can directly use htdemucs-onnx installed effect in huggingface.co for debugging and trial. It also supports api for free installation.