Amharic speech to text, now in Gaston
Our own speech-to-text model for Amharic, fine-tuned on about 1,060 hours of Amharic speech. It turns interviews, broadcasts and everyday conversation into searchable text in Ethiopic script, with timestamps for every word.
Word error rate counts the words that differ from a careful human transcript. Lower is better.
The lowest word error rate among openly licensed Amharic models
We measured every model on Horn-ASR: 774 clips of real, unscripted interviews across seven dialect groups. It is the hardest test set we have, and it was never used for training.
Our lead over hohe-asr is not a fluke: with 95% confidence, our error rate is 0.6 to 1.9 percentage points lower.
* Measured earlier on whole clips rather than short pieces. That method does not change results for these model types, but treat these numbers as indicative.
About half the word errors of OpenAI's best transcription model
We ran OpenAI's gpt-transcribe on exactly the same clips with the same scoring. On conversational, read and descriptive Amharic alike, our model makes about half as many word errors.
The gap holds in every dialect group, for example Gojjam 33.7% vs 62.3% and Wollo 34.0% vs 62.9%.
Tested in October 2026 on 1,774 public test clips (about 6.5 hours). 95% confidence intervals for the gap: 25 to 27 points (Horn-ASR), 26 to 29 (Dataset.ET), 23 to 26 (WAXAL). gpt-transcribe was prompted that the audio is Amharic. WAXAL scores here are without the repetition filter, for a like-for-like comparison. OpenAI models change over time.
Strong on read speech, strong on real conversation
| WER | CER | |
|---|---|---|
| Gaston Amharic | 30.7 | 12.9 |
| Our previous model | 32.7 | 13.9 |
| hohe-asr | 32.0 | 12.5 |
| WER | CER | |
|---|---|---|
| Gaston Amharic | 19.5 | 7.8 |
| Our previous model | 20.6 | 8.4 |
| hohe-asr | 22.2 | 8.3 |
| WER | CER | |
|---|---|---|
| Gaston Amharic | 21.7 | 8.9 |
| Our previous model | 22.1 | 8.9 |
| hohe-asr | 22.5 | 8.7 |
With the repetition filter Gaston uses in production. Without it: 22.7% WER. On this set we and hohe-asr are tied within noise.
WER: word error rate. CER: character error rate. Lower is better for both.
Built for how Amharic is actually spoken
Amharic sounds different in Gojjam than in Addis Ababa. The model beats the best other open model on 6 of 7 dialect groups, and on Addis Ababa it is within half a point.
The biggest gains over our previous model are in Wollo, Gojjam and Shewa. Gojjam dialect recordings were added to training in this version.
● lowest error rate in the group. Each dialect group has only 56 to 163 clips, so differences under about one point are within noise.
What went into it
We started from OpenAI Whisper large-v3 and fine-tuned it on a large, varied collection of Amharic speech: read sentences, spontaneous descriptions, dialect recordings over phone-quality audio and speech from YouTube.
Plus broad material from many other sources and speakers.
About 9 hours of synthetic silence and noise teach it to output nothing instead of inventing text.
What you get
The model runs inside the Gaston transcription service, so every Amharic recording becomes a transcript you can search, translate, bookmark and export.
- •English loanwords and brand names
- •Occasionally merging two words into one
- •Mixing old and new spellings of a few letters (ጕ/ጉ, ኵ/ኩ)
How we measured
- Exactly the way the production service runs it: audio cut into pieces of at most 20 seconds, decoded with beam search 5.
- Spelling variants of letters that sound the same are not counted as errors, for example: ሀ/ሐ/ኀ, ሰ/ሠ
- Horn-ASR was never used for training, but it helped us pick the best training checkpoint, which makes its number very slightly optimistic. The other two test sets were not used for any decision.
| Test set | Size used | Notes |
|---|---|---|
Horn-ASR LesanAI · CC-BY-SA-4.0 |
774 clips up to 30 s (~2.5 h) | Spontaneous interviews, never used for training |
Dataset.ET snapwre/amharic-speech · CC-BY-4.0 |
500-clip random sample | Speakers separate from its training split |
WAXAL Google WaxalNLP · CC-BY-SA-4.0 |
500-clip random sample | Speakers separate from its training split |
Credits
This model would not exist without openly licensed Amharic speech data. Thank you to everyone who built and shared it.
- Based on OpenAI Whisper large-v3 (MIT)
- Amharic Linguistic Datasets © 2026 Leyu (gheero), CC BY 4.0 - recordings were segmented and filtered
- Google WaxalNLP (WAXAL), CC-BY-SA-4.0
- snapwre (Dataset.ET, amharic-speech), CC-BY-4.0
- Google FLEURS, CC-BY-4.0
- Mozilla Common Voice, CC0
- Ejigu, Y. A. (2024), BDU-speech
- ESPnet YODAS2, CC-BY-3.0
- snapwre/hohe-asr-amharic, CC-BY-4.0 (teacher model, not part of the product)
- MUSAN corpus (background noise)
- LesanAI Horn-ASR, CC-BY-SA-4.0 (test only)
Try it on your own Amharic recordings
Upload a file or paste a YouTube link, and Gaston does the rest.
Create an account