Parrotlet-a 2.5 Pro is a purpose-built automatic speech recognition (ASR) model for medical speech in Indian healthcare settings. It transcribes
Indian English, Hindi, Marathi, Kannada and Telugu
, including the heavily code-mixed speech typical of real consultations (English drug names and clinical terms embedded in Indic speech).
The model combines a
Whisper large-v3 encoder
with a
MedGemma 4B decoder
through a lean projector layer, and was further tuned with GRPO on medical conversation data. Weights are stored in bfloat16.
Benchmark results
Re-benchmarked against nine production ASR systems on medical evaluation sets in five Indian languages, plus AI4Bharat's IndicVoices as an out-of-domain check. Every medical sample carries a SQUIM objective STOI-based audio-quality rating (binned into Excellent / Good / Bad / Poor) and medical-entity annotations.
Metric — Semantic WER
: word error rate after medical-unit, number, transliteration and orthography normalization, so a model is not penalised for spelling a drug name in a different script. Lower is better.
Semantic WER — medical conversations
Model
Indian English
Hindi
Marathi
Kannada
Telugu
Parrotlet-a 2.5 Pro
(Eka.care)
9.15
18.16
31.61
29.24
24.77
Gemini 3.1 Pro (Google)
12.09
19.08
39.26
36.38
31.80
Gemini 3.6 Flash (Google)
12.50
21.96
44.44
43.43
37.81
GPT Transcribe (OpenAI)
13.00
26.71
51.98
58.92
42.56
Saaras v3 (Sarvam)
14.59
21.64
40.75
33.33
29.79
Scribe V2 (ElevenLabs)
16.23
21.91
38.14
39.61
33.40
Whisper Large v3 (OpenAI)
17.46
52.26
101.42
99.59
96.22
Prisma v2.5 (Gnani)
26.12
35.90
52.08
51.99
53.82
IndicConformer 600M (AI4Bharat)
—
37.49
56.02
58.37
42.09
Best in every language on the medical conversation sets.
Medical entity accuracy
Share of medical keyword tokens recovered — drug names, doses, conditions, procedures. This is the number that decides whether a transcript is clinically usable. Higher is better.
Model
Indian English
Hindi
Marathi
Kannada
Telugu
Parrotlet-a 2.5 Pro
(Eka.care)
95.05
89.10
82.46
86.01
88.99
Gemini 3.1 Pro (Google)
94.94
88.08
57.43
70.38
71.55
Gemini 3.6 Flash (Google)
94.00
82.88
52.75
63.94
64.08
GPT Transcribe (OpenAI)
90.96
73.58
42.19
43.69
54.44
Saaras v3 (Sarvam)
87.23
76.71
46.11
62.87
63.18
Scribe V2 (ElevenLabs)
86.16
77.64
69.18
75.29
79.42
Whisper Large v3 (OpenAI)
83.81
47.42
17.37
8.50
19.95
Prisma v2.5 (Gnani)
72.26
50.93
32.31
42.37
40.63
IndicConformer 600M (AI4Bharat)
—
50.36
27.37
35.73
44.76
Semantic WER — IndicVoices (out-of-domain)
General-domain read and spontaneous speech, outside the medical domain the model is tuned for. Parrotlet-a 2.5 Pro leads on Kannada and sits within a point or two of the best models on Telugu and Marathi — specialisation has not cost general coverage.
Model
Hindi
Marathi
Kannada
Telugu
Saaras v3 (Sarvam)
10.54
13.16
25.53
20.57
Prisma v2.5 (Gnani)
11.65
14.20
25.16
19.20
IndicConformer 600M (AI4Bharat)
10.76
12.58
26.25
20.69
Parrotlet-a 2.5 Pro
(Eka.care)
12.45
13.27
24.71
20.01
Gemini 3.1 Pro (Google)
12.51
18.51
33.72
24.51
Scribe V2 (ElevenLabs)
12.21
19.58
38.66
28.08
Gemini 3.6 Flash (Google)
15.99
24.04
41.72
30.69
GPT Transcribe (OpenAI)
14.24
23.27
47.12
28.96
Whisper Large v3 (OpenAI)
25.52
77.91
88.79
105.53
Audio-quality robustness
On the medical sets split by audio quality (Excellent → Poor, estimated from the audio itself using SQUIM objective STOI, binned at 0.5 / 0.65 / 0.85), Parrotlet-a 2.5 Pro keeps the flattest degradation profile and is the best model in nearly every quality bucket in every language — e.g. Indian English semantic WER of 6.49 on Excellent audio and 25.14 on Poor audio, versus 8.59 → 43.86 for the next-best model.
Audio (wav, mp3) must be passed at
16 kHz
— load with
librosa.load(..., sr=16000)
as above. Other sample rates are not
currently handled by
transcribe()
.
The model handles short-form audio up to
30 seconds; longer clips are
silently truncated
— chunk longer recordings yourself before
transcribing.
Weights are bfloat16; the model runs on GPU (CUDA) or CPU.
The combined weights are therefore distributed under the
Health AI
Developer Foundations Terms of Use
— by downloading or using this model
you agree to those terms, including their use restrictions and the
requirement that further derivatives carry the same terms.
Citation
If you use this model, please cite:
@software{parrotlet_a_2_5_pro,
author = {{Eka Care}},
title = {Parrotlet-a 2.5 Pro: medical speech recognition for Indian languages},
year = {2026},
url = {https://huggingface.co/ekacare/parrotlet-a-2.5-pro}
}
Runs of ekacare parrotlet-a-2.5-pro on huggingface.co
1.2K
Total runs
-1
24-hour runs
3
3-day runs
-26
7-day runs
888
30-day runs
More Information About parrotlet-a-2.5-pro huggingface.co Model
parrotlet-a-2.5-pro huggingface.co is an AI model on huggingface.co that provides parrotlet-a-2.5-pro's model effect (), which can be used instantly with this ekacare parrotlet-a-2.5-pro model. huggingface.co supports a free trial of the parrotlet-a-2.5-pro model, and also provides paid use of the parrotlet-a-2.5-pro. Support call parrotlet-a-2.5-pro model through api, including Node.js, Python, http.
parrotlet-a-2.5-pro huggingface.co is an online trial and call api platform, which integrates parrotlet-a-2.5-pro's modeling effects, including api services, and provides a free online trial of parrotlet-a-2.5-pro, you can try parrotlet-a-2.5-pro online for free by clicking the link below.
ekacare parrotlet-a-2.5-pro online free url in huggingface.co:
parrotlet-a-2.5-pro is an open source model from GitHub that offers a free installation service, and any user can find parrotlet-a-2.5-pro on GitHub to install. At the same time, huggingface.co provides the effect of parrotlet-a-2.5-pro install, users can directly use parrotlet-a-2.5-pro installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
parrotlet-a-2.5-pro install url in huggingface.co: