GGUF conversion of
m-a-p/SheetSage2
for
audio.cpp
. SheetSage2 transcribes music into melody, chords, beats, key, structure, and editable scores. This is a converted checkpoint, not a new model or an official upstream release.
The packaged file uses the audio.cpp SheetSage2 runtime, not the upstream Transformers loader.
Weights and Size
File
Storage
Size
sheetsage2-orig.gguf
Unquantized FP32, merged model
2,708,224,512 bytes (2.71 GB / 2.52 GiB)
The upstream
model.safetensors
is only 228,738,564 bytes (228.74 MB): it contains SheetSage2 adapters and task-specific weights and relies on the separately downloaded MERT-v2-FullSong backbone. The GGUF includes that backbone with the LoRA updates already merged, plus the task weights and embedded configuration. No separate backbone checkpoint is needed at inference time.
The merged GGUF stores 677,020,953 FP32 values across 1,039 tensors. The larger size comes from packaging the complete merged model, not GGUF container overhead or accidental duplication. Both the upstream SheetSage2 safetensors and this GGUF use FP32; this is not an FP16-to-FP32 size increase.
Conversion Validation
The tensor audit passed all 1,039 tensors: 96 matched the official FP32 LoRA merge byte-for-byte, and 943 matched the source tensors unchanged. There were no missing or extra tensors, dtype or shape errors, or byte mismatches.
Two complete songs (69.84 s and 161.53 s) matched Python token sequences and ABC output exactly with FP32 weights and CUDA TF32 enabled in both implementations. Event timestamps differed only by floating-point rounding (at most 2.85e-14 s). This is matched-math-policy validation, not a guarantee of identical output across backends or with Python TF32 disabled.
Q8 Warning
Q8 is not safe for preserving transcription parity. Use
sheetsage2-orig.gguf
.
On the tested 161.53-second song, Q8_0 first changed a chord at 30.48 seconds. Compared with orig, it produced 538 instead of 542 events and 387 instead of 392 notes. At shared subbeats, 17 chord fields and 32 melody fields differed, and the ABC output changed. This is more than a one-off timestamp drift. It is evidence from one song, not a general accuracy benchmark; Q8 is not included in this package.
License
The upstream weights are licensed under
Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
. This conversion retains that license; conversion does not grant commercial-use rights. Attribute the original SheetSage2 and MERT2 creators and source repositories, and identify the GGUF conversion and LoRA merge as modifications.
SheetSage2-GGUF huggingface.co is an AI model on huggingface.co that provides SheetSage2-GGUF's model effect (), which can be used instantly with this audio-cpp SheetSage2-GGUF model. huggingface.co supports a free trial of the SheetSage2-GGUF model, and also provides paid use of the SheetSage2-GGUF. Support call SheetSage2-GGUF model through api, including Node.js, Python, http.
SheetSage2-GGUF huggingface.co is an online trial and call api platform, which integrates SheetSage2-GGUF's modeling effects, including api services, and provides a free online trial of SheetSage2-GGUF, you can try SheetSage2-GGUF online for free by clicking the link below.
audio-cpp SheetSage2-GGUF online free url in huggingface.co:
SheetSage2-GGUF is an open source model from GitHub that offers a free installation service, and any user can find SheetSage2-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of SheetSage2-GGUF install, users can directly use SheetSage2-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.