VibeVoice-ASR (7B)
, the original BF16 weights as Microsoft published them,
with its
DFlash 2
drafter bundled
in
drafter/
:
one download, and
vibevoice.c
decodes with speculative decoding -- the drafter proposes 8 tokens in one
pass, the model checks them in one pass and keeps the ones it agrees with.
The check is exact
: every checked row is computed with the arithmetic of
the model's own decode step, so the transcript is byte-for-byte the one
without the drafter.
(With a drafter vibevoice.c attends with flashinfer, which checks rows
together; the transcript is then
--attn flashinfer
's.)
Use
Needs vibevoice.c with DFlash 2 support: branch
dflash2
(
PR #48
), in the next
release. A model directory's
drafter/
is used without asking:
vibevoice.c runs BF16 weights as dense FP16 (about 19 GB of VRAM for a 7B,
drafter included);
--quant int4
quantizes them at load, and the drafter
works with that too.
BF16 weights run as dense FP16. Plain = the same model with
--draft none --attn flashinfer
(with a drafter vibevoice.c attends with flashinfer); against the default without a drafter (fa2): 2.79x on the 2-minute file, 2.51x on the 32-minute file.
--draft-check exact
, the check width measured as it runs.
Inside
The model: the files of
microsoft/VibeVoice-ASR
at revision
d0c9efdb
, unchanged (17.37 GB) -- its card has the model, the evaluation and the license.
tokenizer.json
,
tokenizer_config.json
,
vocab.json
,
merges.txt
,
special_tokens_map.json
,
preprocessor_config.json
: from
Ar4ikov/VibeVoice-ASR-AWQ-W4A16-ASYM
at revision
22b44c85
-- the upstream repository leaves them out (they are Qwen2.5's tokenizer and the audio preprocessor's settings), and vibevoice.c needs them next to the weights.
drafter/
:
Ar4ikov/VibeVoice-ASR-DFlash2-Drafter
at revision
795e2a5d
(1.66 GB): 5 Qwen3-style layers reading the model's
layers 1/7/13/19/25, a candidate selector, a 32768-id draft vocabulary;
BF16, held as INT4 by default at load (
--draft-quant f16
keeps FP16). Its card has the architecture and the training.
License
MIT, like VibeVoice.
Runs of Ar4ikov VibeVoice-ASR-DFlash2 on huggingface.co
24
Total runs
0
24-hour runs
1
3-day runs
22
7-day runs
22
30-day runs
More Information About VibeVoice-ASR-DFlash2 huggingface.co Model
VibeVoice-ASR-DFlash2 huggingface.co is an AI model on huggingface.co that provides VibeVoice-ASR-DFlash2's model effect (), which can be used instantly with this Ar4ikov VibeVoice-ASR-DFlash2 model. huggingface.co supports a free trial of the VibeVoice-ASR-DFlash2 model, and also provides paid use of the VibeVoice-ASR-DFlash2. Support call VibeVoice-ASR-DFlash2 model through api, including Node.js, Python, http.
VibeVoice-ASR-DFlash2 huggingface.co is an online trial and call api platform, which integrates VibeVoice-ASR-DFlash2's modeling effects, including api services, and provides a free online trial of VibeVoice-ASR-DFlash2, you can try VibeVoice-ASR-DFlash2 online for free by clicking the link below.
Ar4ikov VibeVoice-ASR-DFlash2 online free url in huggingface.co:
VibeVoice-ASR-DFlash2 is an open source model from GitHub that offers a free installation service, and any user can find VibeVoice-ASR-DFlash2 on GitHub to install. At the same time, huggingface.co provides the effect of VibeVoice-ASR-DFlash2 install, users can directly use VibeVoice-ASR-DFlash2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
VibeVoice-ASR-DFlash2 install url in huggingface.co: