Introduction of fastconformer-quran-coreml-streaming
Model Details of fastconformer-quran-coreml-streaming
FastConformer-Quran — Streaming (CoreML / Apple Neural Engine)
Real-time, on-device
streaming
Quranic recitation ASR for iOS & macOS. Cache-aware
FastConformer-Hybrid (CTC), fp16, runs on the
Apple Neural Engine
at a few milliseconds per chunk —
built for live recitation tracking (word highlighting, follow-along, real-time feedback).
This is the
streaming
member of the FastConformer-Quran family. For maximum-accuracy full-utterance
transcription see the
offline CoreML repo
;
for the source model / ONNX /
.nemo
, see
Muno459/fastconformer-quran
.
Riwayah:
Hafs only — not a general Arabic ASR.
Output:
Arabic with full tashkīl (diacritics).
Architecture:
cache-aware FastConformer-Hybrid, CTC head,
att_context_size = [70, 13]
(~1.04 s lookahead), fixed chunk of 112 mel frames (1120 ms).
11 / 11
Al-Fātiḥah + Al-Ikhlās ayāt correct (incl. the Basmala),
0 NaN
ANE residency
~99%
— 1094 ops on ANE / 9 on CPU (no silent GPU/CPU fallback)
Latency
5–8 ms per 1120 ms chunk
(real-time, large margin)
The chunked, limited-context attention bounds fp16 accumulation by design, so CTC margins stay safely
positive in fp16 on the ANE.
Scope:
the on-device check above is on 11 clean EveryAyah test ayāt — strong evidence the design
holds. A broad multi-reciter WER sweep on-device is future work; a few frames sit on a thin positive
margin (~0.01–0.46 nats), so a very noisy input could surface an edge case.
Accuracy (held-out WER / CER %)
Evaluated on a
leakage-free
held-out set (EveryAyah reciters never used in training + a held-out QUL
reciter + real phone-recorded recitation), CTC greedy, alef-insensitive in parentheses:
Test set
Streaming WER
CER
EveryAyah (held-out reciters, clean studio)
6.3 (6.0)
2.2
QUL — Al-Nufais (held-out reciter, clean)
11.6 (11.2)
6.7
Real phone recitation (tlog)
19.6 (14.3)
7.1
All
9.8 (8.6)
4.0
Streaming trades some accuracy for low-latency, state-carrying inference. If you need the lowest WER and
latency isn't critical, the offline variant scores ~3% WER on the same clips.
All shapes are concrete (no dynamic axes, no
length
input), so the Neural Engine pre-compiles one
kernel and runs it without fallback. Feed each chunk, carry the three
*_next
cache tensors into the
next call.
80-channel log-mel, identical to NeMo
FilterbankFeatures
:
16 kHz, mono
window 25 ms (400 samples), Hann · hop 10 ms (160 samples) · 512-pt FFT
80 mel bins (Slaney), power spectrum,
log(mel + 1e-5)
pre-emphasis 0.97, then per-feature mean/var normalization
Python reference:
tajweed/aligner.py
.
~200 lines in Swift with
Accelerate
for the FFT.
Decoding
Argmax
logprobs
per frame → token IDs.
CTC collapse: drop blanks (id
1024
) and dedupe consecutive identical IDs.
SentencePiece-decode (
tokenizer.model
) → Arabic text. Append across chunks for a rolling transcript.
Pronunciation head (optional)
pronunciation-head.mlpackage
is trained on features pooled from
this streaming encoder
(so its
input distribution matches what the model emits on-device). Inputs: pooled
encoder_output
per token
(512-d) + token ID →
prob_correct
(P the token was pronounced correctly). All ops are ANE-friendly;
sigmoid-bounded, no fp16 concerns.
Precision note
fp16 throughout (no int8/int4 — the ANE is natively fp16). The
RelPositionalEncoding
xscale multiply
(
×√512
) can exceed the fp16 max on large activations, so it is computed then
saturated to ±65 504
(
(x·√512).clamp(±65504)
) — exactly the ANE's own behaviour, so it's a no-op on-device yet prevents
inf→NaN if an op is ever evicted off-ANE. Baked into the graph as a single
clip
op.
License
Apache 2.0 — same as NVIDIA FastConformer-Hybrid and the upstream FastConformer-Quran model.
Citation
@misc{fastconformer_quran_coreml_streaming_2026,
title = {FastConformer-Quran (Streaming, CoreML): on-device Quranic ASR for Apple Neural Engine},
year = {2026},
url = {https://huggingface.co/Muno459/fastconformer-quran-coreml-streaming}
}
Benchmark
Leakage-free held-out WER vs nvidia / whisper / seamless / mms / omniASR / Tarteel:
Quranic ASR Leaderboard
.
Runs of Muno459 fastconformer-quran-coreml-streaming on huggingface.co
25
Total runs
0
24-hour runs
0
3-day runs
5
7-day runs
7
30-day runs
More Information About fastconformer-quran-coreml-streaming huggingface.co Model
More fastconformer-quran-coreml-streaming license Visit here:
fastconformer-quran-coreml-streaming huggingface.co is an AI model on huggingface.co that provides fastconformer-quran-coreml-streaming's model effect (), which can be used instantly with this Muno459 fastconformer-quran-coreml-streaming model. huggingface.co supports a free trial of the fastconformer-quran-coreml-streaming model, and also provides paid use of the fastconformer-quran-coreml-streaming. Support call fastconformer-quran-coreml-streaming model through api, including Node.js, Python, http.
fastconformer-quran-coreml-streaming huggingface.co is an online trial and call api platform, which integrates fastconformer-quran-coreml-streaming's modeling effects, including api services, and provides a free online trial of fastconformer-quran-coreml-streaming, you can try fastconformer-quran-coreml-streaming online for free by clicking the link below.
Muno459 fastconformer-quran-coreml-streaming online free url in huggingface.co:
fastconformer-quran-coreml-streaming is an open source model from GitHub that offers a free installation service, and any user can find fastconformer-quran-coreml-streaming on GitHub to install. At the same time, huggingface.co provides the effect of fastconformer-quran-coreml-streaming install, users can directly use fastconformer-quran-coreml-streaming installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
fastconformer-quran-coreml-streaming install url in huggingface.co: