GGUF weights for native inference with
audio.cpp
. SAM Audio separates a sound
described by text, a masked image or video, or positive and negative time spans
from an input recording.
Upstream
These packages use the original Meta checkpoints, not the separate
-tv
variants.
Each GGUF contains SAM weights, T5 text encoder, tokenizer, configuration, and
audio.cpp model spec. No separate text encoder is needed. F32 packages retain
the original weights, checked byte-for-byte against the source tensors.
BF16 and Q8_0 packages are also available for each variant; both were converted
directly from the source safetensors. Biases, normalization vectors, and
scale/shift tables remain F32. Q8_0 denotes mixed quantized storage, not that
every tensor is quantized.
Outputs are
target.wav
and
residual.wav
, mono 48 kHz. Optional temporal
prompts use
--request-option 'anchors=[["+",0.5,2.0],["-",3.0,4.0]]'
.
The runtime also accepts a masked static image with
--request-option reference_image_path=masked.png
, or a masked video with
--request-option reference_video_path=masked.mp4
. Video decoding requires
FFmpeg/libav runtime libraries. Use one visual option at a time and keep the
video and input audio time origins aligned. Automatic span prediction and
candidate ranking are not included.
See the
model usage guide
for all request options. Code tested on CUDA/Vulkan/Metal.
All parity, performance, and VRAM results below were measured with watermarking enabled to match the official Python implementation, which adds a watermark to its output. The final audio.cpp version will return audio WITHOUT the added watermark, so these results do not describe the final implementation.
Controlled Parity With Python
Parity was tested separately from performance on CUDA with F32 weights and
TF32 disabled in both implementations (
NVIDIA_TF32_OVERRIDE=0
). Python used
seed 42; C++ replayed its exact diffusion noise and watermark bits. Both used
the same 14.975-second office recording,
man speaking
prompt, and 16 midpoint
steps. C++ computed its own encoder features, diffusion, and decoded audio.
Each row compares the final C++ waveforms directly against official Python,
including bounded-memory mode. Higher cosine similarity is better; 1 is exact
directional agreement, not a claim of byte-identical waveforms.
Variant
C++ mode
Target cosine
Residual cosine
Small F32
Normal
0.99999945
0.99999864
Small F32
Bounded memory
0.99999968
0.99999896
Base F32
Normal
0.99999956
0.99999558
Base F32
Bounded memory
0.99999950
0.99999753
Large F32
Normal
0.99999980
0.99999888
Large F32
Bounded memory
0.99999984
0.99999936
All rows exceeded 0.9999 for both outputs. Repeating each controlled C++ run
within the same session produced byte-identical outputs.
Ordinary CUDA Performance With Python
Measured on an NVIDIA RTX 5090 (32 GB), using a 14.975-second recording from the
official office example
and the text prompt
man speaking
.
Both implementations use the original weights, 16 midpoint steps, one
candidate, and no visual or span prompts. Optional ranking and span-prediction
models are not loaded in Python.
Ordinary inference: no fixed or replayed random noise, no TF32 override, and
no forced precision/autocast changes. C++ uses
--seed -1
; Python does not
set a seed. Python uses PyTorch 2.11.0+cu128; C++ uses a debug build and 8 threads.
Warm inference time is the median of the last three of five sequential
requests in one session. It excludes model loading and file I/O. RTF is
inference time divided by input duration; lower is faster.
Peak VRAM is process memory sampled with
nvidia-smi
every 50 ms across
loading and all five requests, using the same method for Python and C++.
Variant
C++ warm time
C++ RTF
C++ peak VRAM
Python warm time
Python RTF
Python peak VRAM
Small F32
684 ms
0.0456
8,500 MiB
746 ms
0.0498
10,972 MiB
Base F32
918 ms
0.0613
11,010 MiB
1,105 ms
0.0738
13,414 MiB
Large F32
1,542 ms
0.1030
17,826 MiB
2,103 ms
0.1405
20,568 MiB
These are text-conditioned, single-recording measurements, not a benchmark of
visual prompting, automatic span prediction, or candidate ranking.
Outputs from these independent-randomness performance runs are not used for
parity comparisons.
BF16 and Q8 Compared With C++ F32
The following are a separate C++-only comparison on the same RTX 5090 and
14.975-second input. Ordinary performance uses the five-request methodology
above,
--seed -1
, and no TF32 or precision overrides. F32 was remeasured for
this comparison; no Python model was run.
Variant
Storage
Warm time
RTF
Peak VRAM
Small
F32
700 ms
0.04673
8,500 MiB
Small
BF16
674 ms
0.04501
7,366 MiB
Small
Q8_0
673 ms
0.04493
6,834 MiB
Base
F32
960 ms
0.06408
11,010 MiB
Base
BF16
785 ms
0.05241
8,644 MiB
Base
Q8_0
794 ms
0.05302
7,524 MiB
Large
F32
1,566 ms
0.10457
17,826 MiB
Large
BF16
1,146 ms
0.07652
12,092 MiB
Large
Q8_0
1,051 ms
0.07018
9,378 MiB
Output drift was measured in
separate controlled C++ runs
, using seed 42
and
NVIDIA_TF32_OVERRIDE=0
for every dtype. Each output is compared against
that variant's C++ F32 waveform, not Python. These runs supply no performance
or VRAM figures.
Variant
Storage
Target cosine vs F32
Residual cosine vs F32
Small
BF16
0.999940
0.999763
Small
Q8_0
0.999476
0.999095
Base
BF16
0.999918
0.999788
Base
Q8_0
0.999623
0.999069
Large
BF16
0.999743
0.999712
Large
Q8_0
0.999590
0.999273
BF16 reduced peak VRAM by 13-32%; Q8_0 reduced it by 20-47% in this workload.
Drift is expected. These short text-conditioned checks do not establish equal
quality for every recording or validate reduced-precision visual prompting
and bounded-memory long-form execution. F32 remains available as the reference.
Bounded-Memory Mode
Enable the experimental session option to reduce long-recording GPU workspace:
--session-option sam_audio.memory_bounded=true
The option defaults to
false
, and its name is provisional. The existing Small
GGUF predates this option; add
--model-spec-override model_specs/sam_audio.json
from a checkout that includes bounded-memory support. The Base and Large GGUFs
include the option in their embedded specs.
The mode retains full-recording diffusion and attention. Codec convolutions
use tiles with receptive-field context, and watermark LSTM state carries
between tiles. It does not separate independent audio chunks and crossfade
them. Intermediate sequences reside in host RAM, which still grows with input
length. Transfers and overlap computation can make it slower on short inputs;
small numerical differences are possible. The configured model context limit
remains 400 seconds.
The following CUDA tests used a
180-second input
made by repeating the same
example to 180 seconds, with the same prompt and inference settings as above.
Each C++ result is the first inference in a fresh session, including graph
setup but excluding model loading and file I/O. Peak VRAM includes loading
and inference. These are not warm medians.
Variant
C++ bounded time
C++ RTF
C++ peak VRAM
Official Python, same input
Small F32
8.768 s
0.0487
6,550 MiB
CUDA out of memory
Base F32
11.109 s
0.0617
9,212 MiB
CUDA out of memory
Large F32
17.038 s
0.0947
16,244 MiB
CUDA out of memory
All three C++ runs produced the complete target and residual recordings.
Python ran out of memory in its audio encoder on the same 32 GB GPU, using
its default precision and allocator settings. This demonstrates completion
and memory usage for this workload,
not 180-second parity against Python
:
Python did not produce a full-length reference. Other hardware or Python
memory-saving configurations may behave differently.
License
SAM Audio weights are distributed under the
SAM License
, dated
November 19, 2025. The included T5 encoder is from
Google T5 Base
and retains its
Apache-2.0 license
. Conversion does not replace the upstream license terms.
Runs of audio-cpp SAM-Audio-GGUF on huggingface.co
4.1K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About SAM-Audio-GGUF huggingface.co Model
SAM-Audio-GGUF huggingface.co is an AI model on huggingface.co that provides SAM-Audio-GGUF's model effect (), which can be used instantly with this audio-cpp SAM-Audio-GGUF model. huggingface.co supports a free trial of the SAM-Audio-GGUF model, and also provides paid use of the SAM-Audio-GGUF. Support call SAM-Audio-GGUF model through api, including Node.js, Python, http.
SAM-Audio-GGUF huggingface.co is an online trial and call api platform, which integrates SAM-Audio-GGUF's modeling effects, including api services, and provides a free online trial of SAM-Audio-GGUF, you can try SAM-Audio-GGUF online for free by clicking the link below.
audio-cpp SAM-Audio-GGUF online free url in huggingface.co:
SAM-Audio-GGUF is an open source model from GitHub that offers a free installation service, and any user can find SAM-Audio-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of SAM-Audio-GGUF install, users can directly use SAM-Audio-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.