MOSS-Transcribe-Diarize 0.9B is an end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
Given an audio or video file, the model generates a compact speaker-aware transcript in one pass, including timestamps and anonymous speaker labels such as
[S01]
,
[S02]
, and beyond.
News
2026-07-09: Released MOSS-Transcribe-Diarize 0.9B.
MOSS-Transcribe-Diarize 0.9B turns real-world long-form audio into structured, speaker-aware transcripts in one pass. Instead of stitching together separate ASR and diarization systems, it jointly performs speech transcription and speaker diarization, producing time-aligned text with consistent speaker labels.
The model is built for meetings, calls, podcasts, interviews, lectures, videos, and other long or messy multi-speaker recordings. It can also emit acoustic event annotations, giving downstream systems a richer view of what happened, who spoke, and when.
Core capabilities:
Long-form transcription
: Converts long audio or video recordings into timestamped text.
Speaker-aware diarization
: Assigns anonymous speaker labels such as
[S01]
and
[S02]
without a separate diarization pipeline.
WhisperFeatureExtractor
, 16 kHz, 80 mel bins, 30 s chunks
Audio-text bridge
4x temporal merge + MLP adaptor
Fusion
Audio features replace
<|audio_pad|>
embeddings via
masked_scatter
Output format
Compact
[start][Sxx]text[end]
transcript with speaker tags such as
[S01]
This Hugging Face repository includes the custom Transformers remote code required to load the model with
trust_remote_code=True
.
Evaluation
We evaluate MOSS-Transcribe-Diarize using three objective metrics: Character Error Rate (CER), concatenated minimum-permutation Character Error Rate (cpCER), and Delta-cp. Lower is better for all metrics. A dash (
-
) indicates that the result is unavailable.
Model
AISHELL‑4
Alimeeting
Podcast
Movies
CER↓
cpCER↓
Δcp↓
CER↓
cpCER↓
Δcp↓
CER↓
cpCER↓
Δcp↓
CER↓
cpCER↓
Δcp↓
Doubao
18.18
27.86
9.68
25.25
37.57
12.31
7.93
10.54
2.61
9.94
30.88
20.94
ElevenLabs
19.58
37.95
18.36
25.70
36.69
10.99
8.50
11.34
2.85
11.49
17.85
6.37
GPT-4o
-
-
-
-
-
-
-
-
-
14.37
23.67
9.31
Gemini 2.5 Pro
42.70
53.42
10.72
27.43
41.64
14.21
7.38
10.23
2.85
15.46
24.15
8.69
Gemini 3 Pro
22.75
27.43
4.68
26.75
32.84
6.09
-
-
-
8.62
14.73
6.11
VIBEVOICE ASR
21.40
24.99
3.59
27.40
29.33
1.93
27.94
48.30
20.36
14.59
42.54
27.94
MOSS Transcribe Diarize 0.9B
14.84
15.83
0.99
24.86
22.17
-2.69
5.97
7.37
1.40
6.36
12.76
6.40
MOSS Transcribe Diarize Pro
13.78
14.02
0.24
18.22
13.94
-4.27
4.46
6.97
2.51
5.86
11.78
5.92
Quickstart
Serve with SGLang Omni
The recommended way to serve MOSS-Transcribe-Diarize is
SGLang Omni
through the OpenAI-compatible
/v1/audio/transcriptions
endpoint. Install
sglang-omni
by following the
Installation guide
, then download the model:
Max generated tokens; raise for long audio (e.g.
65536
)
prompt
string
unset
Optional instruction override; omit to use the built-in transcribe+diarize prompt
verbose_json
parses the model markup into OpenAI-style
segments
with
start
,
end
, and speaker-prefixed
text
(for example
[S01]...
).
json
/
text
return the full transcript string without segment parsing.
The GitHub package provides helper utilities such as audio/video loading, transcription message construction, transcript parsing, CLI inference, and the subtitle web app. The model weights and remote-code model files are loaded from this Hugging Face repository.
Python Usage
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
from moss_transcribe_diarize import parse_transcript
from moss_transcribe_diarize.inference_utils import (
build_transcription_messages,
generate_transcription,
resolve_device,
)
model_id = "OpenMOSS-Team/MOSS-Transcribe-Diarize"
audio_path = "audio.wav"
device = resolve_device("auto")
dtype = torch.bfloat16 if device.type == "cuda"else torch.float32
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype="auto",
).to(dtype=dtype).to(device).eval()
processor = AutoProcessor.from_pretrained(
model_id,
trust_remote_code=True,
)
messages = build_transcription_messages(audio_path)
result = generate_transcription(
model,
processor,
messages,
max_new_tokens=2048,
do_sample=False,
device=device,
dtype=dtype,
)
print(result["text"])
for segment in parse_transcript(result["text"]):
print(segment.start, segment.end, segment.speaker, segment.text)
The message flow follows the common Qwen multimodal pattern:
processor.apply_chat_template(messages, tokenize=False)
renders text with audio placeholders.
The helper utilities load audio waveforms from the same messages.
processor(text=text, audio=audios)
computes Whisper input features and expands audio placeholders.
model.generate(...)
produces timestamped transcription and diarization text.
Custom Prompt and Hotwords
The default prompt is optimized for timestamped transcription and speaker diarization:
MOSS-Transcribe-Diarize also supports vLLM serving through the OpenAI-compatible transcription API. Use a pinned vLLM nightly build that includes the MOSS-Transcribe-Diarize model registration. Choose one of the following commands: for CUDA 12 environments, use
cu129
; for CUDA 13 environments, use
cu130
.
Open
http://127.0.0.1:7860
, upload an audio/video file, review the parsed subtitle segments, then download JSON/SRT/ASS or burn an MP4 if
ffmpeg
and
ffprobe
are available on
PATH
.
[0.48][S01]Welcome everyone[1.66][12.26][S02]The new transcription pipeline is ready for evaluation[13.81][14.36][S01]Great, include the diarization results in the report[18.76]
In this format:
start_time
and
end_time
are timestamps in seconds.
[S01]
,
[S02]
, and similar labels are anonymous model-generated speaker labels.
Speaker labels are relative labels within the input audio and should not be interpreted as real speaker identities.
MOSS-Transcribe-Diarize huggingface.co is an AI model on huggingface.co that provides MOSS-Transcribe-Diarize's model effect (), which can be used instantly with this OpenMOSS-Team MOSS-Transcribe-Diarize model. huggingface.co supports a free trial of the MOSS-Transcribe-Diarize model, and also provides paid use of the MOSS-Transcribe-Diarize. Support call MOSS-Transcribe-Diarize model through api, including Node.js, Python, http.
MOSS-Transcribe-Diarize huggingface.co is an online trial and call api platform, which integrates MOSS-Transcribe-Diarize's modeling effects, including api services, and provides a free online trial of MOSS-Transcribe-Diarize, you can try MOSS-Transcribe-Diarize online for free by clicking the link below.
OpenMOSS-Team MOSS-Transcribe-Diarize online free url in huggingface.co:
MOSS-Transcribe-Diarize is an open source model from GitHub that offers a free installation service, and any user can find MOSS-Transcribe-Diarize on GitHub to install. At the same time, huggingface.co provides the effect of MOSS-Transcribe-Diarize install, users can directly use MOSS-Transcribe-Diarize installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
MOSS-Transcribe-Diarize install url in huggingface.co: