New speech recognition model
አማርኛ

Amharic speech to text, now in Gaston

Our own speech-to-text model for Amharic, fine-tuned on about 1,060 hours of Amharic speech. It turns interviews, broadcasts and everyday conversation into searchable text in Ethiopic script, with timestamps for every word.

30.7%
word error rate on conversational speech
19.5%
word error rate on read speech
~1,060 h
of Amharic speech in training
~9×
faster than real time

Word error rate counts the words that differ from a careful human transcript. Lower is better.

01 · Benchmark

The lowest word error rate among openly licensed Amharic models

We measured every model on Horn-ASR: 774 clips of real, unscripted interviews across seven dialect groups. It is the hardest test set we have, and it was never used for training.

Our lead over hohe-asr is not a fluke: with 95% confidence, our error rate is 0.6 to 1.9 percentage points lower.

Horn-ASR, conversational speech
Lower is better
Gaston Amharic Gaston
30.7%
snapwre/hohe-asr-amharic CC-BY-4.0
32.0%
badrex/Ethio-ASR-amharic CC-BY-4.0
47.0%*
facebook/mms-1b-all CC-BY-NC
58.0%*

* Measured earlier on whole clips rather than short pieces. That method does not change results for these model types, but treat these numbers as indicative.

02 · Versus OpenAI

About half the word errors of OpenAI's best transcription model

We ran OpenAI's gpt-transcribe on exactly the same clips with the same scoring. On conversational, read and descriptive Amharic alike, our model makes about half as many word errors.

The gap holds in every dialect group, for example Gojjam 33.7% vs 62.3% and Wollo 34.0% vs 62.9%.

Word error rate on three public test sets
Lower is better
Gaston Amharic OpenAI gpt-transcribe
Horn-ASR · Conversational speech
30.7%
56.7%
Dataset.ET · Read sentences
19.5%
46.6%
WAXAL · Spontaneous picture descriptions
22.7%
47.1%

Tested in October 2026 on 1,774 public test clips (about 6.5 hours). 95% confidence intervals for the gap: 25 to 27 points (Horn-ASR), 26 to 29 (Dataset.ET), 23 to 26 (WAXAL). gpt-transcribe was prompted that the audio is Amharic. WAXAL scores here are without the repetition filter, for a like-for-like comparison. OpenAI models change over time.

03 · Accuracy

Strong on read speech, strong on real conversation

Horn-ASR
Spontaneous interviews, 7 dialect groups
30.7%
WER
12.9%
CER
WER CER
Gaston Amharic 30.7 12.9
Our previous model 32.7 13.9
hohe-asr 32.0 12.5
Dataset.ET
Read sentences
19.5%
WER
7.8%
CER
WER CER
Gaston Amharic 19.5 7.8
Our previous model 20.6 8.4
hohe-asr 22.2 8.3
WAXAL
Spontaneous picture descriptions
21.7%
WER
8.9%
CER
WER CER
Gaston Amharic 21.7 8.9
Our previous model 22.1 8.9
hohe-asr 22.5 8.7

With the repetition filter Gaston uses in production. Without it: 22.7% WER. On this set we and hohe-asr are tied within noise.

Where we are not ahead yet
On character error rate, hohe-asr is still slightly better on conversational audio (12.5% vs 12.9%) and on WAXAL. Our lead is on whole words, which is what readers notice.

WER: word error rate. CER: character error rate. Lower is better for both.

04 · Dialects

Built for how Amharic is actually spoken

Amharic sounds different in Gojjam than in Addis Ababa. The model beats the best other open model on 6 of 7 dialect groups, and on Addis Ababa it is within half a point.

The biggest gains over our previous model are in Wollo, Gojjam and Shewa. Gojjam dialect recordings were added to training in this version.

Gaston Amharic Our previous model hohe-asr
Horn-ASR word error rate by dialect group
Lower is better
Addis Ababa 158 clips
27.6
29.1
27.1●
Gojjam 128 clips
33.7●
36.1
34.8
Gonder 63 clips
30.2●
31.7
32.2
Shewa 56 clips
37.2●
39.4
39.2
Wollo 111 clips
34.0●
36.6
37.2
Non-native (L2) speakers 163 clips
27.6●
29.8
29.5
Unknown 95 clips
30.2●
31.6
31.1

● lowest error rate in the group. Each dialect group has only 56 to 163 clips, so differences under about one point are within noise.

05 · Training

What went into it

We started from OpenAI Whisper large-v3 and fine-tuned it on a large, varied collection of Amharic speech: read sentences, spontaneous descriptions, dialect recordings over phone-quality audio and speech from YouTube.

1.55 B
parameters, based on OpenAI Whisper large-v3
~1,060 h
of unique Amharic training audio
~809 h
transcribed by people
~276 h
labelled by a teacher model and cross-checked by a second model
1,300+
speakers, counting only the sources that report speaker numbers
~146
GPU-hours of training on NVIDIA RTX 3090 cards
Dialects in training
Addis Ababa Gojjam Gonder Shewa Wollo

Plus broad material from many other sources and speakers.

Trained for messy audio
Background noise Music Room echo Phone-line quality Speed changes Volume changes

About 9 hours of synthetic silence and noise teach it to output nothing instead of inventing text.

06 · In Gaston

What you get

The model runs inside the Gaston transcription service, so every Amharic recording becomes a transcript you can search, translate, bookmark and export.

ሀ ለ ሐ
Ethiopic script
Amharic audio becomes Amharic text in Ge'ez script (fidel), not a romanized approximation.
Every word timestamped
Word-level alignment lets you click any word in the transcript and jump straight to that moment.
Silence stays silent
Trained to output nothing on silence and noise. No repetition loops on any of the 774 conversational test clips, and the service adds its own repetition filter.
Fast
A 7.3-minute recording is transcribed in about 48 seconds on a single GPU, roughly 9 times faster than real time.
Numbers are written as words
Spoken numbers appear as Amharic words, not digits, because the training transcripts follow that convention.
Known weak spots
  • •English loanwords and brand names
  • •Occasionally merging two words into one
  • •Mixing old and new spellings of a few letters (ጕ/ጉ, ኵ/ኩ)
07 · Method

How we measured

  • Exactly the way the production service runs it: audio cut into pieces of at most 20 seconds, decoded with beam search 5.
  • Spelling variants of letters that sound the same are not counted as errors, for example: ሀ/ሐ/ኀ, ሰ/ሠ
  • Horn-ASR was never used for training, but it helped us pick the best training checkpoint, which makes its number very slightly optimistic. The other two test sets were not used for any decision.
Test set Size used Notes
Horn-ASR
LesanAI · CC-BY-SA-4.0
774 clips up to 30 s (~2.5 h) Spontaneous interviews, never used for training
Dataset.ET
snapwre/amharic-speech · CC-BY-4.0
500-clip random sample Speakers separate from its training split
WAXAL
Google WaxalNLP · CC-BY-SA-4.0
500-clip random sample Speakers separate from its training split

Credits

This model would not exist without openly licensed Amharic speech data. Thank you to everyone who built and shared it.

  • Based on OpenAI Whisper large-v3 (MIT)
  • Amharic Linguistic Datasets © 2026 Leyu (gheero), CC BY 4.0 - recordings were segmented and filtered
  • Google WaxalNLP (WAXAL), CC-BY-SA-4.0
  • snapwre (Dataset.ET, amharic-speech), CC-BY-4.0
  • Google FLEURS, CC-BY-4.0
  • Mozilla Common Voice, CC0
  • Ejigu, Y. A. (2024), BDU-speech
  • ESPnet YODAS2, CC-BY-3.0
  • snapwre/hohe-asr-amharic, CC-BY-4.0 (teacher model, not part of the product)
  • MUSAN corpus (background noise)
  • LesanAI Horn-ASR, CC-BY-SA-4.0 (test only)
ሰላም

Try it on your own Amharic recordings

Upload a file or paste a YouTube link, and Gaston does the rest.

Create an account