Audio8-ASR-0.1B
is a compact autoregressive ASR model whose language-model
component has only 0.1B parameters. It supports multilingual speech recognition
for languages including Chinese, English, French, German, Japanese, Korean, and
Cantonese. We position it as one of the smallest usable performance ASR models
in the LLM era.
This base repository provides the Hugging Face Transformers checkpoint. We also
provide deployment-focused releases:
The ONNX Runtime release is designed for edge-device deployment and can run
with roughly 1.1 GB peak memory footprint, depending on device, runtime
configuration, and workload.
The iOS release is designed for local iPhone transcription with roughly 200 MB
peak runtime memory footprint, depending on device, iOS version, and workload.
The Open ASR results use the seven current public splits from
hf-audio/open-asr-leaderboard
at dataset revision
b6bdcd0beb34f8975dc659796176d88f43aff502
. They were
measured with the standalone Transformers package on standardized H200 Hugging
Face Jobs using BF16, eager attention, greedy decoding,
max_new_tokens=256
,
and the documented 30-second audio cap. Per-split batch sizes were 1152, 1024,
1408, 1024, 1024, 2048, and 628. Raw manifests are stored in
hf://buckets/AutoArk-AI/audio8-asr-open-asr-results
, and the corresponding
machine-readable results are provided in
.eval_results/open_asr_leaderboard.yaml
.
The internal canonical WenetSpeech results come from the reproducibility-checked
teacher0p6B-step3000
export with batch size 128. Its effective model tensors
are byte-identical to this standalone release; the release only removes a
redundant tied LM-head tensor and packages the same weights for standalone use.
Chinese results are reported as character error rate. AISHELL is intentionally
excluded from this table.
Model Overview
Task:
automatic speech recognition
Checkpoint format:
safetensors
Sampling rate:
16 kHz
Decoder:
8-layer Qwen-style causal LM
Audio front end:
Qwen3-ASR audio encoder plus MLP adapter/projector
The root
config.json
is intentionally kept in this repository so Hugging Face
can recognize the model package and count downloads through normal model-file
queries.
Hotwords are applied at decode time by nudging logits for tokenizer paths that
match the requested words. This does not modify model weights and does not
inject the hotwords into the prompt.
The default examples target short-form ASR and truncate audio at 30 seconds.
Hotword boosting can help with near-miss terms but can also over-bias decoding
when boost values are too high.
Some Transformers/tokenizers versions emit a Qwen tokenizer regex warning. The
staged tokenizer config is kept in the loadable form used by this package; pass
explicit tokenizer regex flags only after testing your local Transformers version.
Acknowledgements
The audio encoder backbone is based on
Qwen3-ASR-0.6B
, with the audio
adapter and projector trained as part of Audio8-ASR. The language-model backbone
is based on
Ref-Pretrain-Qwen-104M
.
Runs of Edge0 Audio8-ASR-0.1B on huggingface.co
30.0K
Total runs
786
24-hour runs
1.6K
3-day runs
5.5K
7-day runs
23.6K
30-day runs
More Information About Audio8-ASR-0.1B huggingface.co Model
Audio8-ASR-0.1B huggingface.co is an AI model on huggingface.co that provides Audio8-ASR-0.1B's model effect (), which can be used instantly with this Edge0 Audio8-ASR-0.1B model. huggingface.co supports a free trial of the Audio8-ASR-0.1B model, and also provides paid use of the Audio8-ASR-0.1B. Support call Audio8-ASR-0.1B model through api, including Node.js, Python, http.
Audio8-ASR-0.1B huggingface.co is an online trial and call api platform, which integrates Audio8-ASR-0.1B's modeling effects, including api services, and provides a free online trial of Audio8-ASR-0.1B, you can try Audio8-ASR-0.1B online for free by clicking the link below.
Edge0 Audio8-ASR-0.1B online free url in huggingface.co:
Audio8-ASR-0.1B is an open source model from GitHub that offers a free installation service, and any user can find Audio8-ASR-0.1B on GitHub to install. At the same time, huggingface.co provides the effect of Audio8-ASR-0.1B install, users can directly use Audio8-ASR-0.1B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.