Most speech-to-text systems never actually decide whether to write down what
was
said
or what was
meant
. They inherit that choice from their training
data and apply it inconsistently. CrisperWhisper 2.0 makes it an explicit,
controllable choice. One recording, two transcripts:
Verbatim
, exactly what was said, in one consistent format:
[um] so we we need to, to reschedule the th- thursday meeting to [uh] march third at nine thirty [laughter]
Intended
, the clean version the speaker meant, with numbers, dates,
and emails formatted the way you'd write them:
So we need to reschedule the Thursday meeting to March 3 at 9:30.
On top of that:
Word-level timings.
Around 30 ms mean boundary error on read speech
and 41 ms on conversational speech, the most precise word timing of any
system we benchmarked, on both.
Verbatimize.
Upgrade transcripts you already have: given audio plus a
trusted clean transcript, the model reproduces your content word-for-word
and inserts only the disfluencies and vocal events actually present in the
audio (rare-word recall jumps from 6.8% to 96.1% vs. re-transcribing).
This turns the world's abundant clean corpora into verbatim ones, ready
for TTS data, clinical speech analysis, and dataset construction.
Multilingual.
Verbatim and intended modes work across most languages
Whisper supports. CrisperWhisper 2.0 tops the
Nyra Verbatim Speech Benchmark
leaderboard for disfluency F1 across ten languages, ahead of every
closed-source alternative we tested.
Seamless longform.
Audio of any length, transcribed without the usual
chunk-boundary artifacts: each window continues from the words already
transcribed (
conditional continuation
), so there are no duplicated or
dropped words at the seams and no fragile timestamp-token bookkeeping.
Production inference.
A CTranslate2 runtime with speculative decoding
and built-in mitigation of Whisper's looping-hallucination failure mode.
Performance
The
Nyra Verbatim Speech Benchmark
scores fillers, repetitions, cut-offs, and vocal sounds as separate, typed
metrics. Its headline number is
disfluency F1
: how reliably a system
writes down the disfluencies that were actually spoken, without inventing
ones that weren't. Averaged over ten languages:
#
System
Disfluency F1
1
CrisperWhisper 2.0 Pro
93.5
2
CrisperWhisper 2.0
87.8
3
ElevenLabs Scribe v2
79.2
4
Microsoft MAI-Transcribe-1.5
77.5
5
CrisperWhisper 1.0*
64.8
6
Inworld STT
59.5
7
xAI Grok Speech-to-Text
42.8
8
Deepgram Nova-3
37.8
9
Fish Audio ASR
35.0
10
AssemblyAI Universal-3 Pro
30.5
* CrisperWhisper 1.0 is English/German-only; its average covers those
two languages. English and German use human-labeled evaluation sets; the
other eight languages use synthetic verbatim sets. Per-language breakdowns
and how the metric is computed are in the
benchmark post
.
Word-timing accuracy
Mean absolute word-boundary error on read speech (TIMIT), lower is better:
#
System
Boundary error
1
CrisperWhisper 2.0
29.6 ms
2
xAI Grok Speech-to-Text
37.1 ms
3
CTC-seg
49.3 ms
4
ElevenLabs Scribe v2
51.3 ms
5
NeMo-FA
60.0 ms
6
Deepgram Nova-3
63.3 ms
7
WhisperX
64.8 ms
8
Cartesia Ink-Whisper
69.4 ms
9
Canary
85.5 ms
Scored on exactly the words each system gets right. How the timings
are extracted from supervised cross-attention, plus results on
conversational speech, are in the
aligner post
.
Install
# NVIDIA GPU (Linux): fastest, includes speculative decoding.# An NVIDIA driver is all you need; CUDA libraries arrive via pip.
pip install "crisperwhisper[ct2]"# Pure PyTorch: runs anywhere torch does (macOS, Windows, CPU)
pip install "crisperwhisper[transformers]"
Quickstart
from crisperwhisper import CrisperWhisperModel
model = CrisperWhisperModel() # nyralabs/CrisperWhisper2.0_large# or pick a size: CrisperWhisperModel("turbo") # turbo / medium / small# Verbatim transcription (default): every filler, repetition, stutter,# false start, and vocal event
result = model.transcribe("meeting.wav", language="en")
print(result.text)
# Intended: the clean, readable version
clean = model.transcribe("meeting.wav", language="en", mode="intended")
# Word-level timestamps
result = model.transcribe("meeting.wav", language="en", word_timestamps=True)
for w in result.words:
print(f"{w.start:6.2f}-{w.end:6.2f}{w.word}")
# Verbatimize: upgrade an existing clean transcript with the# disfluencies that are actually in the audio
result = model.verbatimize("clip.wav", "I think we should ship it Friday.")
Audio longer than 30 seconds is handled automatically (see
longform
below). The first load of a model
downloads it from HuggingFace and, on the
ct2
backend, converts it once
into a local cache.
Models
Shorthand
HuggingFace ID
Notes
"large"
(default)
nyralabs/CrisperWhisper2.0_large
Best open quality
"turbo"
nyralabs/CrisperWhisper2.0_turbo
Fastest, with some quality degradation; recommended as the speculative draft
"medium"
nyralabs/CrisperWhisper2.0_medium
Near-large quality; best tradeoff between size and quality
Pro
: our best models, with improved performance, hotword boosting, trained on additional proprietary data
The standard models are released under a
non-commercial research license
and are available for commercial licensing. The
Pro
models are available
under commercial license only. For both,
get in touch
.
Faster inference: speculative decoding (ct2)
A small draft model proposes tokens and the main model verifies them. Same
output, 1.3 to 1.4x faster:
model = CrisperWhisperModel("large", draft_model="turbo")
result = model.transcribe("meeting.wav", language="en",
speculative_decoding=True)
What else is in the box
Everything below works out of the box and is covered in depth in
DOCS.md
:
Option
What it does
mode="verbatim" / "intended"
Choose what-was-said vs. what-was-meant per call
word_timestamps=True
Per-word start/end times from supervised cross-attention alignment
hotwords=[...]
Bias recognition toward names and rare terms (
Pro models only
)
model.transcribe_dual(...)
Verbatim
and
intended in one pass (ct2)
model.verbatimize(audio, transcript)
Insert real disfluencies into a trusted clean transcript
model.forced_align(audio, text)
Timings for a transcript you already have
Longform
Audio >30s transcribed seamlessly via conditional continuation, with no chunk-boundary duplicates, drops, or stitching
Hallucination mitigation
On by default: detects and suppresses Whisper's looping-repetition failure mode during decoding
DOCS.md
covers the full API: backends and their trade-offs,
every
transcribe()
option, dual-mode transcription, forced alignment,
longform strategies, speculative-K tuning, hallucination-repair thresholds,
quantization, model conversion, and the result object.
License
The model weights are released under the
Nyra Health Non-Commercial Research License
:
free for research and other non-commercial use; any commercial use requires
a commercial license. The Pro models are available under commercial license
only. For commercial licensing of either,
contact Nyra
.
Runs of nyralabs CrisperWhisper2.0_medium on huggingface.co
4.5K
Total runs
0
24-hour runs
258
3-day runs
1.1K
7-day runs
1.1K
30-day runs
More Information About CrisperWhisper2.0_medium huggingface.co Model
CrisperWhisper2.0_medium huggingface.co is an AI model on huggingface.co that provides CrisperWhisper2.0_medium's model effect (), which can be used instantly with this nyralabs CrisperWhisper2.0_medium model. huggingface.co supports a free trial of the CrisperWhisper2.0_medium model, and also provides paid use of the CrisperWhisper2.0_medium. Support call CrisperWhisper2.0_medium model through api, including Node.js, Python, http.
CrisperWhisper2.0_medium huggingface.co is an online trial and call api platform, which integrates CrisperWhisper2.0_medium's modeling effects, including api services, and provides a free online trial of CrisperWhisper2.0_medium, you can try CrisperWhisper2.0_medium online for free by clicking the link below.
nyralabs CrisperWhisper2.0_medium online free url in huggingface.co:
CrisperWhisper2.0_medium is an open source model from GitHub that offers a free installation service, and any user can find CrisperWhisper2.0_medium on GitHub to install. At the same time, huggingface.co provides the effect of CrisperWhisper2.0_medium install, users can directly use CrisperWhisper2.0_medium installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
CrisperWhisper2.0_medium install url in huggingface.co: