akisviete / azuolas-qwen-lt

huggingface.co
Total runs: 383
24-hour runs: 95
7-day runs: 383
30-day runs: 383
Model's Last Updated: September 29 2026
automatic-speech-recognition

Introduction of azuolas-qwen-lt

Model Details of azuolas-qwen-lt

Ąžuolas v2 — Qwen3-ASR-1.7B Lithuanian (LIEPA-3, full SFT)

The second generation of Ąžuolas , the Lithuanian ASR fine-tune series. v1 ( akisviete/azuolas-whisper-lt ) was whisper-large-v3 ; v2 replaces the backbone with Qwen/Qwen3-ASR-1.7B-hf and keeps everything that made v1 what it was: a full fine-tune on the complete LIEPA-3 Lithuanian corpus — read, spontaneous and phonetic, ~ 9,806 hours .

Qwen3-ASR's technical report lists 52 supported languages, and Lithuanian is not among them ( resolve_language() raises on lt ). v2 is therefore not a refinement of an existing capability — it is a capability installed by fine-tuning, and the training recipe reflects that (most importantly: a learning rate 5× higher than the official SFT default, which is the single largest measured lever — see Training ).

The two generations split by job. Under a matched decode (beam 5, one version-matched stack) v1 is the more accurate model on the two large clean sets (8.93 vs 9.74 FLEURS, 5.09 vs 5.35 CV26). v2 rejects non-speech where v1 hallucinates over it , runs ~30 % faster under PyTorch at every decode setting tested, is 4.5× smaller, and is the integrated Subtitle Edit / CrispASR deployment path. One operational rule: run this model with beam search, not greedy — beam 5 is the decode every matched-decode comparison in this card uses and the one the CrispASR deployment path runs. The measured v2 vs v1 has the full picture.

Output is lowercase and unpunctuated , following the LIEPA-3 label convention, and uses the realised-speech convention — it writes turim , not normative turime , because that is what the corpus labels do.

Results

Evaluated on the author's eval_compare.py harness: normalize_lt() on both sides, jiwer.wer / jiwer.cer . Both sets are out of domain : neither FLEURS nor Common Voice is part of LIEPA-3, and neither shares its recording conditions.

FLEURS LT — 985 clips, 2.96 h, 17,092 reference words Common Voice 26.0 LT — 5,877 clips, 8.50 h, 39,653 reference words (release cv-corpus-26.0-2026-06-12 )

All numbers are beam 5 on one fresh version-matched stack (uv venv, torch 2.14.0+cu130, transformers 5.17.0, beam 5 wherever the architecture supports it):

model (decode) FLEURS WER / CER CV26 WER / CER
Ąžuolas v1, large-v3 (beam 5) 8.93 / 5.84 5.09 / 1.38
Ąžuolas v2 1.7B — this model (beam 5) 9.74 / 6.79 5.35 / 1.70
kmynas-parakeet-lt-v4, TDT 0.6B (NeMo beam 5) 10.56 / 6.18 6.56 / 1.42
paprika v3, distilled (beam 5) 11.60 / 6.41 8.08 / 1.72

Notes on the table. Harness validation: the same 2026-09-25 head-to-head also ran both Ąžuolas generations on their canonical (greedy) decode, and both reproduced their previously published figures to the hundredth (9.12 / 4.97 for v1, 9.74 / 5.11 for v2) — that is what makes the matched beam-5 comparison above trustworthy. Cross-harness deltas below ~0.5 points are noise. kmynas-parakeet-lt-v4 (a 3,087 h LIEPA-3 transducer fine-tune) is greedy-only in the transformers path ( ParakeetForTDT supports GREEDY_SEARCH only); its row above is the NeMo path (nemo 3.0.0), where transducer beam 5 changes almost nothing versus greedy (10.56 / 6.60) — for this TDT, beam search ≈ greedy — while the optimized NeMo decode is the speed leader by far (on the 57-minute hard-chunk long-form run: 2.0 s wall, 6,568 words, zero degeneration). It also emits punctuation, casing and standardised endings; normalize_lt strips punctuation but not the daryt→daryti standardisation, so its WER carries a small convention tax. paprika v3 is a distilled Whisper (32-layer large-v3 encoder + 4-layer decoder, 1.6 GB) — the raw-speed option.

Reading: v1 leads the clean sets (−0.81 FLEURS / −0.26 CV26 vs next-best); kmynas-v4 is the speed leader by far ; this model sits second on both clean sets — ahead of kmynas-v4 (10.56 / 6.56) and paprika v3 (11.60 / 8.08) on both. Context from prior harnesses, not in the table: stock openai/whisper-large-v3 24.37 FLEURS / 30.07 CV26, kmynas-v3 (v4's predecessor) 17.73 / 22.86.

The same ranking survives a completely different inference stack. Re-run through CrispASR (q8_0 GGUF, Vulkan, RTX 3080 Ti) over the same full sets:

model (q8_0, CrispASR) FLEURS WER / CER CV26 WER / CER speed
Ąžuolas v1 9.07 / 5.87 4.52 / 1.06 11.4× RT
Ąžuolas v2 1.7B (this model) 9.79 / 6.24 5.03 / 1.28 13.7× RT

q8_0 quantization is effectively lossless for this model: 9.79/5.03 (q8_0) against 9.74/5.11 (PyTorch bf16), inside the ~0.2 noise floor. The q8_0 GGUF (2.5 GB) is the artifact that runs in Subtitle Edit — see Deployment as GGUF .

Long-form — a 57-minute public event, unscripted. Vidas Mačiulis's book launch at Ąžuolyno biblioteka (57.1 min, audience, overlapping speech, applause). No human reference exists for this recording, so no WER is reported; what is measurable is coverage, degeneration and speed:

model (CrispASR, --vad ) sec ×RT words coverage distinct-5-gram loop
Ąžuolas v2 1.7B (this model) 292.3 11.7 6,596 99.8 % 1.000 none
Ąžuolas v1 299.6 11.4 6,631 99.8 % 1.000 none

No degeneration, and the two models' outputs agree within a few percent. Qwen3-ASR's audio encoder handles long input at inference regardless of training clip length — a model trained almost entirely on ~5 s raw clips (an ablation, not a deployment candidate) ran the same 57-minute file with identical segment counts and zero looping.

Non-speech rejection — the axis where this model wins. Trained with a non-speech pool injected at a realised 1.40 %, it scored 100 % rejection, 0 characters on digital silence, faint hiss, room tone, ESC-50 and DEMAND (the last two never used in training).

Same-protocol head-to-head (20 s pure silence + 20 s white noise at −40 dB, PyTorch, both models):

input v1 (whisper) v2 (this model)
pure silence 22 chars 0 chars
−40 dB noise 414 chars of fluent hallucination 0 chars

v1 also invents up to 947 characters over faint hiss on its own card's 60 s protocol, and its no-speech gate does not fire — the production VAD exists to suppress exactly this. If your audio contains silence, noise or music, this is the model that stays quiet.

Intended use

Lithuanian transcription of read and spontaneous speech: interviews, public events, parliamentary and broadcast audio, phone recordings. 16 kHz mono input (the processor resamples for you).

If you need punctuation, casing or speaker labels, those are separate models.

How to run

Requires transformers >= 5.13.0 (Qwen3-ASR is natively supported from that version):

pip install "transformers>=5.13.0"
Short clips
from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL = "akisviete/azuolas-qwen3-asr-1.7b-lt"
processor = AutoProcessor.from_pretrained(MODEL)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL, dtype=torch.bfloat16, device_map="cuda")

inputs = processor.apply_transcription_request(
    audio="clip.wav",
    language="Lithuanian",          # force the language; auto-detect also works
).to(model.device, model.dtype)

out_ids = model.generate(**inputs, max_new_tokens=256, num_beams=5)
gen_ids = out_ids[:, inputs["input_ids"].shape[1]:]

# raw output is  language Lithuanian<asr_text>…
print(processor.decode(gen_ids, return_format="transcription_only")[0])

return_format="parsed" gives a {"language": …, "transcription": …} dict if you also want the detected language.

Use beam search, not greedy. This model is an encoder-LLM and is sensitive to the decode setting in a way Whisper models are not — on in-distribution sets the difference is small, but far out of distribution the LLM decoder can collapse on rare words and proper names under greedy. num_beams=5 is the setting every matched-decode comparison in this card uses and the one the CrispASR deployment path runs.

Long recordings — cut at pauses, not on a fixed stride

Decode in windows of at most ~22–26 s , cut at the midpoints of pauses found by Silero VAD , or use an engine that does this for you (CrispASR with --vad -vm silero ). The training windows average 25.44 s, so this is the regime the model was built for.

Fixed-stride chunking is measured to under-transcribe: on 180 s of Lithuanian speech, 30 s/28 s-stride chunks reached 82.4 % agreement with the reference against 94.9 % for pause-cut ~22 s windows — it does not truncate, it silently omits runs of 10–16 words from the middle.

Deployment as GGUF (CrispASR / Subtitle Edit)

A q8_0 GGUF of these weights (2.5 GB) is the deployed artifact for Subtitle Edit's Crisp ASR Qwen3 engine (CrispASR's convert-qwen3-asr-to-gguf.py , llama.cpp-style tensor names). Verified end to end: 180 s of audio in 49 s (3.7× RT, q8_0, Vulkan, RTX 3080 Ti), word-level timestamps via Qwen3-ForcedAligner, diacritics intact. Always pass --vad (CrispASR's qwen3 path runs beam search + VAD, which is the production decode this model wants — see Results ). A GGUF of this exact checkpoint can be published alongside the HF weights if requested.

VRAM

Measured reference points:

path weights in memory what was measured
PyTorch bf16, beam 5, clips ≈3.4 GB (1.7 B params × 2 bytes) the beam-5 evaluation ran at batch 12 with ~19 GB free on a 96 GB card time-sliced with a 78 GB vLLM tenant; greedy batch 32 / beam batch 16 OOM'd in the same headroom
PyTorch bf16, whole file ≈3.4 GB + audio-feature prefill prefill of 57 min of audio OOM'd with ~19 GB free — long audio must be VAD-chunked (see above)
GGUF q8_0 (CrispASR) 2.5 GB runs on a 12 GB RTX 3080 Ti with the forced aligner resident; the f16 GGUF (4.4–4.7 GB) does not fit 12 GB once KV cache + aligner (~5 GB overhead) are added
Training (for reference) bf16 peak 77.93 GB on one 96 GB GPU

Rule of thumb derived from those points: the GGUF path wants ~8 GB minimum / 12 GB comfortable ; the PyTorch path wants 16 GB for comfortable beam-5 clip inference. The minimum figures are derived, not measured — everything in the table is measured.

Limitations
  • Decode-sensitive: use beam search, not greedy. On in-distribution sets greedy and beam 5 agree within ~0.5 points, but as an LLM decoder this model is decode-sensitive in a way Whisper models are not (±0.2), and greedy can collapse on rare words far out of distribution. num_beams=5 is the production decode — this card's numbers and the CrispASR path.
  • Lowercase, unpunctuated, realised-speech output by design.
  • No dialect slice. LIEPA-3's dial subset (~100 h) was excluded: 99.95 % of its transcripts carry phonetic dialect annotation rather than standard orthography.
  • Terser under hard chunking, and the loop grows with the window. Under hard 30 s greedy chunks v2 drops ~28 % of v1's words (conversational fillers; LIEPA-3 clean-text bias), and the greedy loop near the 55-min mark of the 57-minute file grows from ~340 words at 30 s windows to ~780 at 60 s. Use the VAD-cut path (≤ ~26 s) for long audio.
  • 16 kHz mono. Read and spontaneous speech dominate the mix.
  • Every WER figure is a single run.
Training

Full SFT of Qwen/Qwen3-ASR-1.7B-hf (no LoRA — see below), 1 epoch on the full 9,806 h corpus, ~16 h wall time on one 96 GB GPU.

base Qwen/Qwen3-ASR-1.7B-hf (full fine-tune, all weights)
data LIEPA-3, 9,806 h — read 5,230 h (53.3 %) + spontaneous 4,359 h (44.5 %) + phonetic 217 h (2.2 %), Whisper-prep fused ~30 s windows (mean 25.44 s), streamed at the corpus's natural subset proportions
non-speech pool 19,755 clips / 42.57 h (synthetic, MUSAN, AudioSet — all CC BY 4.0 or generated), injected at nonspeech_rate 0.0142 , realised 1.40 % — this is what buys the 100 % rejection
optimiser AdamW, β (0.9, 0.98 ), ε 1e-9 , weight decay 0.1 , max_grad_norm 1.0
lr 1e-4 , linear decay to zero, warmup ratio 0.02 (~220 steps)
batch 16 × 8 accumulation = effective 128
steps 10,995
precision bf16, FlashAttention 2, no gradient checkpointing
checkpoint bf16
The learning rate

The official Qwen3-ASR SFT recipe specifies lr 2e-5 . It is ~5× too low for this task:

lr FLEURS CV26
2e-5 (official) 12.95 9.83
1e-4 (this model) 9.74 5.11
2e-4 10.97 5.96

The 2e-5 convention assumes the model already has the capability and you are nudging it. Lithuanian is absent from the base model's 52 languages, so the weights must travel much further than a domain nudge can in ~1,100 steps — and max_grad_norm 1.0 was clipping the under-warmed 2e-5 run most of the time, compounding the undershoot. The curve peaks at 1e-4; 2e-4 is a real regression (+1.23 FLEURS), so this is a peak, not a plateau.

What else was measured, and found inert or harmful
  • LoRA loses to full SFT at every rank tested at full scale (r32: 13.87 / 10.59; r8: 18.10 / 15.97). Rank capacity dominates once optimizer steps are adequate. The LoRA line is closed.
  • Encoder LR multiplier: swept 0.5 → 1.0 → 1.5 → 2.0; monotonically worse in both directions from 1.0. The audio tower already trains at the full base LR; pushing it past the LM destabilises the joint optimisation.
  • Batch: effective 128 beats 64 at full corpus (0.53 FLEURS / 0.57 CV26) despite half the optimizer steps — the 1,000 h proxy's "batch is free" result does not survive full scale.
  • Schedule: cosine is +0.39 FLEURS worse than linear at the peak LR.
  • Prepped corpus beats the raw LIEPA-3 parquet outright (9.74 vs 10.33 FLEURS, 1.62× faster training) and is the only variant that keeps silence rejection at 100 % — the raw-clip run collapsed the realised non-speech injection rate 4.7× because the pool capped, regressing rejection to 96.7 %.
  • Training loss is not a validity signal for this model class. Two separate incidents: a label-smoothing path bug produced 100 % WER while reporting a lower loss, and the raw-clip run reached half the training loss while scoring worse on both OOD sets.
Comparison — v2 vs v1

Same corpus, same series, different backbone — run side by side on one PyTorch harness on 2026-09-25, including the full five-model beam-5 leaderboard ( Results ) and a worst-case long-form stress test (hard 30 s chunks, greedy).

On accuracy, v1 is the better ASR model on the two large clean sets , under a matched decode (beam 5, one fresh version-matched stack for both): 8.93 vs 9.74 FLEURS, 5.09 vs 5.35 CV26.

v2's defensible advantages are the other four axes, all measured:

  • Silence robustness out of the box. 0 characters on pure silence and −40 dB noise; v1 hallucinates confidently (414 characters) and needs the production VAD to suppress it. That is the difference between a model you gate and one you don't have to.
  • Speed under PyTorch — ~30 % faster at every decode setting tested (greedy and beam 5, on both sets). In the deployed engines the ranking inverts (CrispASR 13.7× vs 11.4× RT) — the PyTorch gap is a framework/engine artifact, not a model property.
  • 4.5× smaller deployment footprint — which is what makes the GGUF/CrispASR path possible at all, and 13.7× realtime on a 12 GB consumer GPU.
  • It is an LLM backbone — promptable (hotwords via prompt ), DPO-tunable, usable for downstream reasoning over the transcript. v1 is a seq2seq model that does exactly one job.

Rule of thumb. For verbatim reference transcription of clean Lithuanian speech — e.g. the TTS data pipeline — v1 keeps the edge on the big clean sets. For the interactive / Subtitle Edit / CrispASR deployment, where latency, silence robustness and footprint matter, v2 is the right choice — run it with beam search, not greedy.

Against the field, on the matched beam-5 leaderboard: second on both clean sets — behind v1 (8.93 / 5.09 vs 9.74 / 5.35) — and ahead of kmynas-v4 (10.56 / 6.56) and paprika v3 (11.60 / 8.08) on both. It is the only one of the field that ships a working non-speech guarantee, and the only encoder-LLM.

Sources

Base model

Models compared

Data

Tooling

Licence

Weights CC BY 4.0 , inheriting the LIEPA-3 corpus attribution chain (Vilnius University / raštija.lt). Base model Qwen/Qwen3-ASR-1.7B-hf is Apache-2.0. The non-speech pool is all CC BY 4.0 or generated.

Contact

Questions and suggestions: [email protected] .

Runs of akisviete azuolas-qwen-lt on huggingface.co

383
Total runs
95
24-hour runs
217
3-day runs
383
7-day runs
383
30-day runs

More Information About azuolas-qwen-lt huggingface.co Model

More azuolas-qwen-lt license Visit here:

https://choosealicense.com/licenses/cc-by-4.0

azuolas-qwen-lt huggingface.co

azuolas-qwen-lt huggingface.co is an AI model on huggingface.co that provides azuolas-qwen-lt's model effect (), which can be used instantly with this akisviete azuolas-qwen-lt model. huggingface.co supports a free trial of the azuolas-qwen-lt model, and also provides paid use of the azuolas-qwen-lt. Support call azuolas-qwen-lt model through api, including Node.js, Python, http.

azuolas-qwen-lt huggingface.co Url

https://huggingface.co/akisviete/azuolas-qwen-lt

akisviete azuolas-qwen-lt online free

azuolas-qwen-lt huggingface.co is an online trial and call api platform, which integrates azuolas-qwen-lt's modeling effects, including api services, and provides a free online trial of azuolas-qwen-lt, you can try azuolas-qwen-lt online for free by clicking the link below.

akisviete azuolas-qwen-lt online free url in huggingface.co:

https://huggingface.co/akisviete/azuolas-qwen-lt

azuolas-qwen-lt install

azuolas-qwen-lt is an open source model from GitHub that offers a free installation service, and any user can find azuolas-qwen-lt on GitHub to install. At the same time, huggingface.co provides the effect of azuolas-qwen-lt install, users can directly use azuolas-qwen-lt installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

azuolas-qwen-lt install url in huggingface.co:

https://huggingface.co/akisviete/azuolas-qwen-lt

Url of azuolas-qwen-lt

azuolas-qwen-lt huggingface.co Url

Provider of azuolas-qwen-lt huggingface.co

akisviete
ORGANIZATIONS

Other API from akisviete