GGUF weights for native inference with
audio.cpp
. RE-USE restores degraded
speech while preserving its input sample rate. It uses a convolutional
encoder/decoder and 30 bidirectional time-frequency Mamba blocks.
Upstream
The checkpoint and reference implementation are pinned to
nvidia/RE-USE, revision 022e920d727347a64d6c21fbf0628f2a5f37ad78
.
Each GGUF includes the model weights, configuration, and audio.cpp model spec.
No separate encoder or tokenizer is required. Packages are converted directly
from the original safetensors, not from another GGUF.
Packages
File
Size
Storage
Recommendation
reuse-f32.gguf
37.80 MiB
F32
Default; fastest of the tested C++ formats.
Only F32 is provided. F16 and Q8 were rejected because their small disk savings
did not improve speed or peak VRAM in testing, while introducing waveform drift
and additional packaging complexity.
F32 was checked on CUDA and Vulkan. Metal, and HIP were not validated in
this comparison.
Controlled Parity With Python
These are
separate accuracy runs
, not the performance runs below. Both
implementations used original F32 weights with TF32 disabled. Comparisons cover
the final waveform, with no time alignment or normalization applied.
Input
Rate / channels
C++ vs Python cosine
Relative RMSE
Official
noisy_audio/mic_test2.wav
, 5.10 s
44.1 kHz / mono
0.99999846
0.001756
sample_16k.wav
, 14.07 s
16 kHz / mono
0.999999995
0.0000972
c.wav
, 7.53 s
24 kHz / stereo
0.99964691
0.026572
The last two inputs are in audio.cpp's
audio fixtures
.
The outputs are not bit-identical to Python; the stereo example has noticeably
larger numerical drift than the mono examples. These checks do not establish
identical output on every recording or backend.
Ordinary CUDA Performance
Remeasured on 2026-09-27 after the scoped SSM-convolution optimization.
Original F32 checkpoint for Python. Normal inference defaults,
no TF32
overrides or precision overrides
. Logging enabled for C++.
Five sequential requests in one loaded session. Time is the median of the
last three; model loading and WAV file I/O are excluded. C++ includes its host
STFT/ISTFT, and Python includes its GPU STFT/ISTFT.
Peak VRAM is process memory sampled with
nvidia-smi
every 50 ms across
loading and all five requests, not PyTorch allocator-only memory.
RTF is inference time divided by input duration; lower is faster.
5.10-second official microphone clip
Inference time
RTF
Peak VRAM
Python F32
273.58 ms
0.0536
5,604 MiB
C++ F32
277.47 ms
0.0544
3,012 MiB
C++ F32 has similar latency to Python while using less peak VRAM in this
measurement. These results do not establish the same trade-off for every input
or backend.
Long Recordings
The 327.6-second
qwen3_tts_longform_asr_input.wav
fixture uses
10-second
chunks with 1-second overlap in both implementations
. Same measurement
method as above. This is chunked processing, not full-recording parity.
Chunking changes bidirectional and normalization context and can change output
quality. It is not equivalent to full-recording inference. Input and output
audio still occupy host memory proportional to recording length.
RE-USE-GGUF huggingface.co is an AI model on huggingface.co that provides RE-USE-GGUF's model effect (), which can be used instantly with this audio-cpp RE-USE-GGUF model. huggingface.co supports a free trial of the RE-USE-GGUF model, and also provides paid use of the RE-USE-GGUF. Support call RE-USE-GGUF model through api, including Node.js, Python, http.
RE-USE-GGUF huggingface.co is an online trial and call api platform, which integrates RE-USE-GGUF's modeling effects, including api services, and provides a free online trial of RE-USE-GGUF, you can try RE-USE-GGUF online for free by clicking the link below.
audio-cpp RE-USE-GGUF online free url in huggingface.co:
RE-USE-GGUF is an open source model from GitHub that offers a free installation service, and any user can find RE-USE-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of RE-USE-GGUF install, users can directly use RE-USE-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.