ArmBench-ASR: A Benchmark for Armenian ASR

Community Article
Published August 20, 2026

Automatic speech recognition for Armenian is improving, but comparing models remains difficult. Public evaluations often cover a single speech style, use inconsistent text processing, or lack standardized evaluation procedures. ArmBench-ASR provides a shared evaluation across five Armenian speech datasets, with strict and normalized scores reported.

This initial v0.1 release evaluates almost 30 models, both open-weight and closed. The evaluation suite consists of 10,113 clips totaling approximately 20.7 hours of audio.

Top result: Gemini 2.5 Pro ranks first by strict combined WER, with 14.31% WER and 6.42% normalized WER.

Armenian remains underrepresented in speech technology. A model that performs well on a standard read-speech test set may still struggle with the variability of real-world speech. A single aggregate score cannot expose these differences.

ArmBench-ASR evaluates the models on the same fixed collection of five datasets:

  • Common Voice 26: public, crowdsourced read speech.
  • FLEURS: public read speech. We use the Armenian test split and apply a small set of manual transcription corrections.
  • Poems: expressive literary speech with background music (not open-sourced).
  • Movies: conversational dialogue from produced media, sometimes with background noise (not open-sourced).
  • Infocom: Armenian news text recited by a speaker (not open-sourced).

In addition to applying a small set of manual corrections to FLEURS, we manually reviewed the references in our three collected datasets — Poems, Movies, and Infocom — and corrected transcription issues identified during the review.

Together, these datasets provide a broader picture of Armenian ASR quality than any one test set can provide.

Treemaps showing the ArmBench ASR dataset composition by sample count and audio duration

Methodology

All models are evaluated against the same references and fixed dataset splits. The same deterministic pre-processing and normalization functions are applied symmetrically to references and predictions.

We publish two complementary scoring views:

  • WER / CER use stricter pre-processed text and remain sensitive to punctuation and capitalization differences.
  • WER-norm / CER-norm apply additional normalization on top of preprocessing, including lowercasing and punctuation removal.

The pre-processing pipeline applies Unicode NFKC, standardizes punctuation and quote variants, canonicalizes apostrophes, Armenian comma marks, hyphens and dashes, and makes spacing and initial capitalization consistent. The normalized pipeline additionally lowercases text and removes punctuation while retaining Armenian, English and Russian letters plus digits.

Default parameters were used for all models. The Gemini runs used temperature 0 and thinking was disabled or constrained to a minimal setting.

Scores are computed per dataset and over the combined evaluation set. Combined metrics are calculated from pooled edit counts rather than by averaging the five dataset percentages.

Main Findings

  • Performance is strongly domain-dependent. Rankings and error rates change across read speech, poetry, movies, and narrated news. Reporting only one public test set hides important weaknesses.
  • Movie audio is the most difficult domain in this release. Every evaluated model has its highest strict WER on Movies. The median Movies WER is 61.87%, compared with 17.29% on Common Voice 26 and 17.02% on FLEURS. The possible reasons behind this is the background noise.
  • Normalization materially changes the apparent error rate. The gap between strict and normalized scores shows that punctuation, capitalization, spacing, and orthographic variation account for a meaningful share of measured errors. The reason behind this could be that the punctuation sometimes gets subjective, especially in the Poems.
  • No single dataset tells the full story. Strong combined performance does not guarantee the best result on every domain, which makes per-dataset reporting essential.
  • Closed systems currently lead on aggregate accuracy. The 8 lowest combined WER scores belong to closed systems. The strongest open model is NVIDIA's Armenian FastConformer at number 9 wit a combined WER of 20.21%.

The ArmBench-ASR interactive leaderboard allows users to switch between the combined view and individual datasets, compare WER and CER variants, and filter systems by availability.

Top Models and Analysis

Ranking by strict combined WER, Google’s gemini-2.5-pro leads at 14.31%, followed by HiSpeech’s model-23012026, a conversational model developed in Armenia, at 16.81%. Google’s gemini-2.5-flash ranks 3rd at 17.55%, followed closely by Talk2Edit, ElevenLabs Scribe V2, and HiSpeech’s formal model-01052025, all at around 18% WER.

Strict and normalized WER should not be directly compared, as they can produce different rankings. A model may perform better on strict WER but slightly worse on normalized WER.

Top 10 Models, WER/CER

Gemini models perform best on the Infocom dataset: 4 of the 5 tested models rank at the top of the leaderboard with ≤14% WER. One possible explanation is that Infocom includes news-style narration, a widely available type of speech data that is likely well represented in large-scale training corpora.

As mentioned earlier, NVIDIA’s Armenian FastConformer is the top open-source model and also leads on MCV with a 7.70% WER. This is the only dataset where an open model ranks first. The tbb-asr model follows closely at 8.15% WER, with both Armenian models outperforming gemini-2.5-pro at 9.69%.

For the Poems dataset, both HiSpeech models outperform Gemini on standard WER (~23% vs. ~25%). However, Gemini leads on normalized WER at 4%, compared with 4.62% and 4.85% for HiSpeech. HiSpeech’s newer conversational model also outperforms its older formal model on nearly every dataset.

Average Dataset Scores on Top-10 Models

Limitations

  • The benchmark mainly covers Eastern Armenian. It does not evaluate Western Armenian or regional dialects, so no claims about those varieties should be inferred from these results.
  • Building a representative dialect dataset is particularly difficult because many Armenian dialects do not have a single accepted written form, making consistent reference transcription and scoring challenging.
  • Three domain-focused datasets are private, which limits independent reproduction of those portions of the benchmark.
  • The benchmark measures transcription accuracy, not punctuation quality, speaker diarization, timestamps, streaming behavior, robustness to long-form audio, or downstream task performance.
  • The evaluations were conducted in late July and early August 2026. Hosted models and APIs may change over time, including without a publicly versioned checkpoint.

Benchmark results should be interpreted in the context of this evaluation setup. Model performance can depend on prompting, API configuration, speech domain, and product-specific optimizations, so these rankings should not be treated as a universal measure of model quality or suitability for every use case.

Caveats

  • Lower is better for every WER and CER metric.
  • Strict and normalized scores answer different questions and should be compared within the same scoring view.
  • Corrected FLEURS references: We made slight fixes to the Armenian FLEURS references and published the corrected rows in the Metric-AI/fleurs-corrections dataset.
  • Manually reviewed references: We manually reviewed the references in the collected Poems, Movies, and Infocom datasets to ensure data quality.
  • Poems may contain background music, and Movies may contain background noise or other produced-media audio.
  • Infocom is news content recited by one speaker.
  • Closed APIs may apply undocumented preprocessing or model routing, and their results may not be exactly reproducible later.

What Comes Next

ArmBench-ASR is intended to grow beyond the current five-dataset transcription benchmark. Future releases can broaden both the speech represented in the data and the tasks used to evaluate speech systems.

  • Code-switching: dedicated data containing natural switching between Armenian and languages commonly used alongside it, including English and Russian.
  • More kinds of dialogue: spontaneous conversations, interviews, customer-support calls, meetings, podcasts, and multi-speaker discussions with interruptions and overlapping speech.
  • Specialized topics: domain-focused speech from medicine, law, finance, education, technology, and public administration, where terminology and proper names create different recognition challenges.
  • Speaker diarization: evaluating not only what was said, but also who spoke when in multi-speaker and overlapping recordings.
  • Additional transcription tasks: comparing verbatim transcription with clean or fixed transcription, including punctuation restoration, capitalization, disfluency handling, and correction of spoken-form text into a readable final transcript.
  • Timestamp and segmentation quality: measuring word or utterance timing, long-form segmentation, and the handling of pauses and speaker changes.

These additions would make the benchmark more representative of production speech systems, where useful output requires more than minimizing WER on short, single-speaker clips.

Community

Sign up or log in to comment