cocktailpeanut / muscriptor-large

huggingface.co
Total runs: 336
24-hour runs: 0
7-day runs: -11
30-day runs: 31
Model's Last Updated: July 11 2026

Introduction of muscriptor-large

Model Details of muscriptor-large

MuScriptor — large (≈1.3B)

MuScriptor is an open-weight model for general-purpose, multi-instrument automatic music transcription (AMT) : it converts a music recording (any genre, multiple simultaneous instruments) into a stream of notes played. This repository hosts the large variant (≈1.3B parameters) — the flagship, best-quality checkpoint .

For a smaller footprint use muscriptor-medium (≈300M, good trade-off) or muscriptor-small (≈100M, fastest).

Table of contents
Quickstart

Install the muscriptor package (it uses huggingface_hub to fetch weights automatically):

pip install git+https://github.com/muscriptor/muscriptor.git
# TODO (PyPI release forthcoming: pip install muscriptor)
Python
from pathlib import Path
from muscriptor import TranscriptionModel

# "large" resolves to hf://MuScriptor/muscriptor-large and downloads on first use.
model = TranscriptionModel.load_model("large")

# Get a MIDI file directly:
Path("out.mid").write_bytes(model.transcribe_to_midi("audio.wav"))

# Or stream note events as they are transcribed:
for event in model.transcribe("audio.wav"):
    print(event)  # NoteStartEvent / NoteEndEvent / ProgressEvent

load_model accepts a size keyword ( "small" / "medium" / "large" ), a local .safetensors path, or an hf:// / https:// URL. Weights loaded by size keyword (or any hf:// URL) are cached in the standard Hugging Face cache ( ~/.cache/huggingface/hub , configurable via HF_HOME ); weights fetched from a plain http(s):// URL are cached under ~/.cache/muscriptor/ . Input audio can be WAV or any format libsndfile reads (mp3, flac, ogg, m4a, …); it is resampled to 16 kHz mono internally.

CLI
muscriptor transcribe --model large audio.wav -o out.mid
Model description

MuScriptor performs transcription by autoregressively predicting a MIDI-like token sequence given the mel-spectrogram of a short audio segment, following the sequence-to-sequence AMT paradigm (cf. MT3). It deliberately avoids complex architectural tweaks in favor of a simple, decoder-only Transformer.

  • Architecture: decoder-only Transformer (this variant: dim=1536 , num_heads=24 , num_layers=48 ).
  • Input: raw waveform (16 kHz, mono) of a 5-second segment → mel-spectrogram (STFT n_fft=2048 , hop 160 → 100 Hz frame rate, 512 mel bins). The spectrogram is projected to the model dimension and used as a prefix condition.
  • Output tokenization: MT3-like note events; the 128 MIDI programs are mapped to 36 instrument subgroups using the MT3_FULL_PLUS taxonomy. Decoding is greedy (argmax) by default, with optional classifier-free guidance (CFG).
  • Inference: audio is processed in 5-second chunks; note events are emitted in temporal order. Optional instrument conditioning stabilizes predictions across chunk boundaries and lets you restrict/customize the transcription (see below).

Note on the representation: the tokenizer recovers onset/offset timing, pitch, and instrument, but not velocity . It also cannot represent two notes of the same pitch and instrument sounding at the same time. Drums are onset-only.

Model variants
Repo Params dim heads layers Notes
muscriptor-small ≈100M 768 12 14 smallest / fastest
muscriptor-medium ≈300M 1024 16 24 good trade-off
muscriptor-large ≈1.3B 1536 24 48 this model · best quality

All variants share the same input pipeline, tokenizer, and training recipe; they differ only in latent dimension, attention heads, and depth.

Intended uses & limitations

Intended uses

  • General-purpose transcription of real, multi-instrument music across genres (classical → heavy metal) into MIDI.
  • A building block for music information retrieval (chord/key recognition), musicological analysis, generative-modeling data pipelines, and tools for musicians.

Out of scope / use with care

  • Not a substitute for a hand-annotated score; expect errors, especially on dense mixes, unusual timbres, and heavily processed audio.
  • Velocity/dynamics are not produced (see note above).
  • Onset/offset precision is lower for some styles (e.g. choral music), and exact offsets are inherently harder than onsets.

Limitations & biases

  • Training data skews toward pop and Western classical music, and the instrument distribution is long-tailed (piano/guitar/bass/drums are most frequent). Rare instruments and underrepresented genres may be transcribed less reliably.
  • The fixed MT3_FULL_PLUS 36-group instrument taxonomy limits instrument granularity.
  • Simultaneous same-pitch/same-instrument notes cannot be represented by the tokenizer.
Instrument conditioning

The model can be told which instrument groups are present in the track. Supplying the correct set improves quantitative scores and produces more coherent instrument assignments across segments.

from muscriptor.tokenizer.mt3 import MT3_FULL_PLUS_GROUP_NAMES

# `instrument_group` is a space-separated string of MT3_FULL_PLUS group IDs.
# Convert readable group names to IDs:
names = ["acoustic_piano", "acoustic_guitar", "acoustic_bass"]
instrument_group = " ".join(str(MT3_FULL_PLUS_GROUP_NAMES[n]) for n in names)  # -> "0 4 7"

# Only expect piano, acoustic guitar and bass in this track:
model.transcribe_to_midi("audio.wav", instrument_group=instrument_group)
muscriptor transcribe --model large --instruments "acoustic_piano,acoustic_guitar,acoustic_bass" audio.wav -o out.mid
muscriptor list-instruments   # show all available group names
Evaluation

Metrics are instrument-agnostic F1 scores computed with mir_eval .

Headline results on D_Test

D_Test is the authors' held-out test set of 372 multi-instrument tracks. Results below are for this 1.3B model with the full training pipeline ( D_Synth + D_Real + D_RL ), CFG = 2:

Model Onset F1 Frame F1 Offset F1 Drums F1 Multi F1
YourMT3+ (baseline) 32.5 45.5 17.8 41.4 21.9
MuScriptor 1.3B 60.4 72.4 48.6 49.6 47.8
Model-size comparison

F1 ↑ on D_Test from the paper's scaling study (models trained on D_Real only, CFG = 2). Note these ablation numbers omit synthetic pre-training and RL, so they are lower than the full-pipeline results above:

Variant Params Onset Frame Offset Drums Multi
muscriptor-small 100M 51.2 67.2 38.7 41.5 38.2
muscriptor-medium 300M 52.4 68.0 40.3 42.0 39.7
muscriptor-large 1.3B 53.2 68.7 41.0 42.5 40.5
Citation
@inproceedings{muscriptor2026,
  title     = {MuScriptor: An Open Model for Multi-Instrument Music Transcription},
  author    = {Rouard, Simon and Krause, Michael and Roebel, Axel and
               Simon-Gabriel, Carl-Johann and D{\'e}fossez, Alexandre},
  year      = {2026},
  note      = {Kyutai, Mirelo AI, IRCAM}
}
License

Code released under the MIT License . Weights released under CC-BY-NC.

Runs of cocktailpeanut muscriptor-large on huggingface.co

336
Total runs
0
24-hour runs
2
3-day runs
-11
7-day runs
31
30-day runs

More Information About muscriptor-large huggingface.co Model

More muscriptor-large license Visit here:

https://choosealicense.com/licenses/cc-by-nc-4.0

muscriptor-large huggingface.co

muscriptor-large huggingface.co is an AI model on huggingface.co that provides muscriptor-large's model effect (), which can be used instantly with this cocktailpeanut muscriptor-large model. huggingface.co supports a free trial of the muscriptor-large model, and also provides paid use of the muscriptor-large. Support call muscriptor-large model through api, including Node.js, Python, http.

cocktailpeanut muscriptor-large online free

muscriptor-large huggingface.co is an online trial and call api platform, which integrates muscriptor-large's modeling effects, including api services, and provides a free online trial of muscriptor-large, you can try muscriptor-large online for free by clicking the link below.

cocktailpeanut muscriptor-large online free url in huggingface.co:

https://huggingface.co/cocktailpeanut/muscriptor-large

muscriptor-large install

muscriptor-large is an open source model from GitHub that offers a free installation service, and any user can find muscriptor-large on GitHub to install. At the same time, huggingface.co provides the effect of muscriptor-large install, users can directly use muscriptor-large installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

muscriptor-large install url in huggingface.co:

https://huggingface.co/cocktailpeanut/muscriptor-large

Url of muscriptor-large

Provider of muscriptor-large huggingface.co

cocktailpeanut
ORGANIZATIONS

Other API from cocktailpeanut

huggingface.co

Total runs: 7.9K
Run Growth: 5.1K
Growth Rate: 64.28%
Updated:April 12 2025
huggingface.co

Total runs: 2.6K
Run Growth: -287
Growth Rate: -10.94%
Updated:June 05 2025
huggingface.co

Total runs: 978
Run Growth: -22
Growth Rate: -2.25%
Updated:August 02 2024
huggingface.co

Total runs: 200
Run Growth: 156
Growth Rate: 78.00%
Updated:September 19 2024
huggingface.co

Total runs: 181
Run Growth: 131
Growth Rate: 72.38%
Updated:October 06 2024
huggingface.co

Total runs: 166
Run Growth: 124
Growth Rate: 74.70%
Updated:September 22 2024
huggingface.co

Total runs: 135
Run Growth: 108
Growth Rate: 80.00%
Updated:September 17 2024
huggingface.co

Total runs: 118
Run Growth: 94
Growth Rate: 79.66%
Updated:September 19 2024
huggingface.co

Total runs: 93
Run Growth: 53
Growth Rate: 56.99%
Updated:September 29 2024
huggingface.co

Total runs: 91
Run Growth: 69
Growth Rate: 75.82%
Updated:September 22 2024
huggingface.co

Total runs: 74
Run Growth: -32
Growth Rate: -43.24%
Updated:October 22 2024
huggingface.co

Total runs: 45
Run Growth: 33
Growth Rate: 73.33%
Updated:October 31 2024
huggingface.co

Total runs: 14
Run Growth: -1
Growth Rate: -7.14%
Updated:December 05 2024
huggingface.co

Total runs: 5
Run Growth: 0
Growth Rate: 0.00%
Updated:September 26 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:February 04 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 25 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:June 13 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 22 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:October 21 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 15 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 03 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 13 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:September 14 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:September 30 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:September 09 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 07 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 13 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 06 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:February 29 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 31 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:September 01 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 10 2025