Training
: Self-supervised pre-training on LibriSpeech 960h, fine-tuned with CTC loss
Language
: English only
License
: Apache 2.0
WER
: 1.89% (LibriSpeech test-clean), 4.07% (test-other)
Usage with CrispASR
# Uses the wav2vec2 backend (auto-detected from GGUF architecture)
crispasr --backend wav2vec2 -m data2vec-audio-base-960h-q4_k.gguf -f audio.wav
Architecture Notes
Data2Vec Audio differs from standard wav2vec2 in three ways handled by the converter:
5-layer positional convolution
(vs 1 for wav2vec2), each with Conv1d + LayerNorm(no affine) + GELU
Global encoder LayerNorm BEFORE transformer layers
(vs after for wav2vec2)
POST-norm encoder
despite using LayerNorm in CNN (wav2vec2-large uses pre-norm)
All three are auto-detected from the HuggingFace model config and stored as GGUF metadata flags.
Files
File
Size
JFK Transcription
data2vec-audio-base-960h-f16.gguf
196 MB
perfect
data2vec-audio-base-960h-q4_k.gguf
79 MB
perfect
data2vec-audio-base-960h-q8_0.gguf
120 MB
perfect
Accuracy
Tested on JFK inaugural address (11s):
AND SO A MY FELLOW AMERICANS ASK NOT WHAT YOUR COUNTRY CAN DO FOR YOU
ASK WHAT YOU CAN DO FOR YOUR COUNTRY
Identical to the Python HuggingFace reference output. All quantized variants produce the same transcription.
Citation
@inproceedings{baevski2022data2vec,
title={data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language},
author={Baevski, Alexei and Hsu, Wei-Ning and Xu, Qiantong and Babu, Arun and Gu, Jiatao and Auli, Michael},
booktitle={ICML},
year={2022}
}
Runs of cstr data2vec-audio-960h-GGUF on huggingface.co
936
Total runs
34
24-hour runs
46
3-day runs
96
7-day runs
205
30-day runs
More Information About data2vec-audio-960h-GGUF huggingface.co Model
data2vec-audio-960h-GGUF huggingface.co is an AI model on huggingface.co that provides data2vec-audio-960h-GGUF's model effect (), which can be used instantly with this cstr data2vec-audio-960h-GGUF model. huggingface.co supports a free trial of the data2vec-audio-960h-GGUF model, and also provides paid use of the data2vec-audio-960h-GGUF. Support call data2vec-audio-960h-GGUF model through api, including Node.js, Python, http.
data2vec-audio-960h-GGUF huggingface.co is an online trial and call api platform, which integrates data2vec-audio-960h-GGUF's modeling effects, including api services, and provides a free online trial of data2vec-audio-960h-GGUF, you can try data2vec-audio-960h-GGUF online for free by clicking the link below.
cstr data2vec-audio-960h-GGUF online free url in huggingface.co:
data2vec-audio-960h-GGUF is an open source model from GitHub that offers a free installation service, and any user can find data2vec-audio-960h-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of data2vec-audio-960h-GGUF install, users can directly use data2vec-audio-960h-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
data2vec-audio-960h-GGUF install url in huggingface.co: