nyralabs / CrisperWhisper2.0_medium

huggingface.co
Total runs: 4.5K
24-hour runs: 0
7-day runs: 1.1K
30-day runs: 1.1K
Model's Last Updated: August 03 2026
automatic-speech-recognition

Introduction of CrisperWhisper2.0_medium

Model Details of CrisperWhisper2.0_medium

CrisperWhisper 2.0

PyPI

The most accurate verbatim speech recognition you can run in production: controllable, multilingual, and timed to the word.

Release post · Paper · Full documentation · Models · Benchmark · Benchmark repo

Try it now

Most speech-to-text systems never actually decide whether to write down what was said or what was meant . They inherit that choice from their training data and apply it inconsistently. CrisperWhisper 2.0 makes it an explicit, controllable choice. One recording, two transcripts:

Verbatim , exactly what was said, in one consistent format: [um] so we we need to, to reschedule the th- thursday meeting to [uh] march third at nine thirty [laughter]

Intended , the clean version the speaker meant, with numbers, dates, and emails formatted the way you'd write them: So we need to reschedule the Thursday meeting to March 3 at 9:30.

On top of that:

  • Word-level timings. Around 30 ms mean boundary error on read speech and 41 ms on conversational speech, the most precise word timing of any system we benchmarked, on both.
  • Verbatimize. Upgrade transcripts you already have: given audio plus a trusted clean transcript, the model reproduces your content word-for-word and inserts only the disfluencies and vocal events actually present in the audio (rare-word recall jumps from 6.8% to 96.1% vs. re-transcribing). This turns the world's abundant clean corpora into verbatim ones, ready for TTS data, clinical speech analysis, and dataset construction.
  • Multilingual. Verbatim and intended modes work across most languages Whisper supports. CrisperWhisper 2.0 tops the Nyra Verbatim Speech Benchmark leaderboard for disfluency F1 across ten languages, ahead of every closed-source alternative we tested.
  • Seamless longform. Audio of any length, transcribed without the usual chunk-boundary artifacts: each window continues from the words already transcribed ( conditional continuation ), so there are no duplicated or dropped words at the seams and no fragile timestamp-token bookkeeping.
  • Production inference. A CTranslate2 runtime with speculative decoding and built-in mitigation of Whisper's looping-hallucination failure mode.
Performance

The Nyra Verbatim Speech Benchmark scores fillers, repetitions, cut-offs, and vocal sounds as separate, typed metrics. Its headline number is disfluency F1 : how reliably a system writes down the disfluencies that were actually spoken, without inventing ones that weren't. Averaged over ten languages:

# System Disfluency F1
1 CrisperWhisper 2.0 Pro 93.5
2 CrisperWhisper 2.0 87.8
3 ElevenLabs Scribe v2 79.2
4 Microsoft MAI-Transcribe-1.5 77.5
5 CrisperWhisper 1.0* 64.8
6 Inworld STT 59.5
7 xAI Grok Speech-to-Text 42.8
8 Deepgram Nova-3 37.8
9 Fish Audio ASR 35.0
10 AssemblyAI Universal-3 Pro 30.5

* CrisperWhisper 1.0 is English/German-only; its average covers those two languages. English and German use human-labeled evaluation sets; the other eight languages use synthetic verbatim sets. Per-language breakdowns and how the metric is computed are in the benchmark post .

Word-timing accuracy

Mean absolute word-boundary error on read speech (TIMIT), lower is better:

# System Boundary error
1 CrisperWhisper 2.0 29.6 ms
2 xAI Grok Speech-to-Text 37.1 ms
3 CTC-seg 49.3 ms
4 ElevenLabs Scribe v2 51.3 ms
5 NeMo-FA 60.0 ms
6 Deepgram Nova-3 63.3 ms
7 WhisperX 64.8 ms
8 Cartesia Ink-Whisper 69.4 ms
9 Canary 85.5 ms

Scored on exactly the words each system gets right. How the timings are extracted from supervised cross-attention, plus results on conversational speech, are in the aligner post .

Install
# NVIDIA GPU (Linux): fastest, includes speculative decoding.
# An NVIDIA driver is all you need; CUDA libraries arrive via pip.
pip install "crisperwhisper[ct2]"

# Pure PyTorch: runs anywhere torch does (macOS, Windows, CPU)
pip install "crisperwhisper[transformers]"
Quickstart
from crisperwhisper import CrisperWhisperModel

model = CrisperWhisperModel()          # nyralabs/CrisperWhisper2.0_large
# or pick a size: CrisperWhisperModel("turbo")  # turbo / medium / small

# Verbatim transcription (default): every filler, repetition, stutter,
# false start, and vocal event
result = model.transcribe("meeting.wav", language="en")
print(result.text)

# Intended: the clean, readable version
clean = model.transcribe("meeting.wav", language="en", mode="intended")

# Word-level timestamps
result = model.transcribe("meeting.wav", language="en", word_timestamps=True)
for w in result.words:
    print(f"{w.start:6.2f}-{w.end:6.2f}  {w.word}")

# Verbatimize: upgrade an existing clean transcript with the
# disfluencies that are actually in the audio
result = model.verbatimize("clip.wav", "I think we should ship it Friday.")

Audio longer than 30 seconds is handled automatically (see longform below). The first load of a model downloads it from HuggingFace and, on the ct2 backend, converts it once into a local cache.

Models
Shorthand HuggingFace ID Notes
"large" (default) nyralabs/CrisperWhisper2.0_large Best open quality
"turbo" nyralabs/CrisperWhisper2.0_turbo Fastest, with some quality degradation; recommended as the speculative draft
"medium" nyralabs/CrisperWhisper2.0_medium Near-large quality; best tradeoff between size and quality
"small" nyralabs/CrisperWhisper2.0_small Smallest
"large_pro" / "turbo_pro" / "medium_pro" / "small_pro" nyralabs/CrisperWhisper2.0_<size>_pro Pro : our best models, with improved performance, hotword boosting, trained on additional proprietary data

The standard models are released under a non-commercial research license and are available for commercial licensing. The Pro models are available under commercial license only. For both, get in touch .

Faster inference: speculative decoding (ct2)

A small draft model proposes tokens and the main model verifies them. Same output, 1.3 to 1.4x faster:

model = CrisperWhisperModel("large", draft_model="turbo")
result = model.transcribe("meeting.wav", language="en",
                          speculative_decoding=True)
What else is in the box

Everything below works out of the box and is covered in depth in DOCS.md :

Option What it does
mode="verbatim" / "intended" Choose what-was-said vs. what-was-meant per call
word_timestamps=True Per-word start/end times from supervised cross-attention alignment
hotwords=[...] Bias recognition toward names and rare terms ( Pro models only )
model.transcribe_dual(...) Verbatim and intended in one pass (ct2)
model.verbatimize(audio, transcript) Insert real disfluencies into a trusted clean transcript
model.forced_align(audio, text) Timings for a transcript you already have
Longform Audio >30s transcribed seamlessly via conditional continuation, with no chunk-boundary duplicates, drops, or stitching
Hallucination mitigation On by default: detects and suppresses Whisper's looping-repetition failure mode during decoding
compute_type="float16" / "int8_float16" Quantization
How it works

Each mechanism has a deep-dive post:

Documentation

DOCS.md covers the full API: backends and their trade-offs, every transcribe() option, dual-mode transcription, forced alignment, longform strategies, speculative-K tuning, hallucination-repair thresholds, quantization, model conversion, and the result object.

License

The model weights are released under the Nyra Health Non-Commercial Research License : free for research and other non-commercial use; any commercial use requires a commercial license. The Pro models are available under commercial license only. For commercial licensing of either, contact Nyra .

Runs of nyralabs CrisperWhisper2.0_medium on huggingface.co

4.5K
Total runs
0
24-hour runs
258
3-day runs
1.1K
7-day runs
1.1K
30-day runs

More Information About CrisperWhisper2.0_medium huggingface.co Model

More CrisperWhisper2.0_medium license Visit here:

https://choosealicense.com/licenses/nyra-health-non-commercial-research

CrisperWhisper2.0_medium huggingface.co

CrisperWhisper2.0_medium huggingface.co is an AI model on huggingface.co that provides CrisperWhisper2.0_medium's model effect (), which can be used instantly with this nyralabs CrisperWhisper2.0_medium model. huggingface.co supports a free trial of the CrisperWhisper2.0_medium model, and also provides paid use of the CrisperWhisper2.0_medium. Support call CrisperWhisper2.0_medium model through api, including Node.js, Python, http.

CrisperWhisper2.0_medium huggingface.co Url

https://huggingface.co/nyralabs/CrisperWhisper2.0_medium

nyralabs CrisperWhisper2.0_medium online free

CrisperWhisper2.0_medium huggingface.co is an online trial and call api platform, which integrates CrisperWhisper2.0_medium's modeling effects, including api services, and provides a free online trial of CrisperWhisper2.0_medium, you can try CrisperWhisper2.0_medium online for free by clicking the link below.

nyralabs CrisperWhisper2.0_medium online free url in huggingface.co:

https://huggingface.co/nyralabs/CrisperWhisper2.0_medium

CrisperWhisper2.0_medium install

CrisperWhisper2.0_medium is an open source model from GitHub that offers a free installation service, and any user can find CrisperWhisper2.0_medium on GitHub to install. At the same time, huggingface.co provides the effect of CrisperWhisper2.0_medium install, users can directly use CrisperWhisper2.0_medium installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

CrisperWhisper2.0_medium install url in huggingface.co:

https://huggingface.co/nyralabs/CrisperWhisper2.0_medium

Url of CrisperWhisper2.0_medium

CrisperWhisper2.0_medium huggingface.co Url

Provider of CrisperWhisper2.0_medium huggingface.co

nyralabs
ORGANIZATIONS

Other API from nyralabs