Other people's Ethiopian language datasets
Speech, text and translation data published by other researchers and organisations. None of it is ours. It is gathered here so that anyone starting work on an Ethiopian language has one place to look instead of searching and hoping.
These are third-party datasets. Read this first.
None of this is our data, and we have not checked its quality. Every entry belongs to the researcher or organisation that published it. We have not listened to it, measured it, or verified that it contains what its authors say it contains. This is a directory to make things findable, nothing more. Judge each one yourself, and credit its authors, not us.
We host none of it. Every link goes to the publisher, and the licence shown is the one they declared, copied rather than inferred. Where a publisher declared no licence we say so, because that means you have no permission to use it until you ask them.
Links and licences were last confirmed by hand on 29 August 2026. Things move. If you find something out of date or missing, tell us and we will fix it.
5 of the 28 entries below declare no licence at all. They are marked, and you should treat them as all rights reserved.
Speech
12 entries
ALFFA Amharic
MITALFFA project, via hadamard-2 · 10K–100K rows
The classic academic Amharic read-speech corpus for ASR, converted to the Hugging Face datasets format.
Long-standing reference corpus; several published Amharic ASR models are trained on it, so it is a reasonable baseline to compare against.
Amharic Speech Dataset 110HRS
None declaredKYAGABA · 10K–100K rows
One of the larger community-contributed Amharic speech collections.
The name claims 110 hours; the dataset card declares no licence and no hour count. Without a licence you have no permission to use it, so contact the author before building anything on it.
Common Voice 17 (Amharic and Tigrinya)
CC0 1.0Mozilla, via fsicoli
Mozilla's crowd-sourced read-speech corpus, with per-clip age, gender, accent and up/down vote metadata.
This is an unofficial community conversion, not Mozilla's own upload. The per-clip metadata schema is close to ours and is a useful reference point.
Amharic ASR Dataset
Apache 2.0Ephrem · Under 1K rows
A small community-contributed Amharic ASR set.
Tagged under 1,000 rows, so treat it as a sample rather than a training corpus.
Amharic Speech (protocol demonstration)
CC BY-NC 4.0Silencio Network · Under 1K rows
A small speaker-metadata-rich set published to demonstrate a collection protocol rather than to train models.
Non-commercial licence, so it cannot be used to train a commercial model. Its card openly documents gender and regional imbalance, which is a good template for how any corpus should state its own limits.
WaxalNLP
CC BY-SA 4.0 / CC BY 4.0Google's large multilingual African speech corpus, covering Amharic among more than forty languages.
By far the most downloaded dataset on this page. Widely used as the Amharic ASR training baseline, and the set most published WER figures are measured against. Dual-licensed; check which licence applies to the portion you use.
Leyu Amharic dialect corpora
CC BY 4.0Leyu · 1K–10K rows each
Five separate Amharic speech corpora, one per regional dialect: Addis Ababa, Shewa, Gojjam, Gonder and Wello.
Partitioned by dialect rather than pooled, which is unusual and useful if you care about regional coverage or want to measure how a model generalises across dialects. Openly licensed. Leyu works in the same space we do.
Ethiopian Languages Speech Dataset
CC BY 4.0Benji-fish · 1K–10K rows
Speech covering four Ethiopian languages in one release, openly licensed.
One of the few sets that spans several Ethiopian languages rather than Amharic alone.
Sagalee
CC BY-NC 4.0turiabu · 10K–100K rows
A read-speech corpus for Afaan Oromo ASR.
One of the largest open Oromo speech sets. Non-commercial licence, so it cannot be used to train a commercial model. Several re-uploads exist under other accounts; this is the one to cite.
Ethiopian Speech collection
None declaredBadr al-Absi (badrex) · 10K–100K rows
Speech data across five Ethiopian languages, alongside per-language sets and the Ethio-ASR models trained on them.
The data behind the Ethio-ASR models. No licence is declared on the dataset, so ask before building on it, even though the models themselves are openly licensed.
Tigrinya Speech Data
CC BY 4.0Professor · 10K–100K rows
An openly licensed Tigrinya speech corpus.
Tigrinya is badly under-served; this is among the larger open sets. The same publisher has Wolaytta and Sidama equivalents.
Amharic TTS Benchmark
Other (see dataset card)Addis AI · Under 1K rows
An evaluation set for Amharic text-to-speech systems.
A benchmark rather than training data, published on the Addis AI organisation account. Addis AI also ran the independent seven-system ASR comparison against our own test set.
Model
6 entries
SpeechBrain DVoice Amharic ASR
Apache 2.0SpeechBrain
A wav2vec2 + CTC Amharic ASR system trained on the ALFFA corpus.
A model rather than a dataset. Listed because it is one of the few ready-to-run Amharic ASR baselines.
Ethio-ASR
See individual modelsBadr al-Absi (badrex)
A family of wav2vec2-BERT ASR models covering five Ethiopian languages, with both per-language and multilingual variants.
The widest open ASR coverage of Ethiopian languages we are aware of. We benchmark against these.
Amharic TTS (SpeechT5)
MITAddisuSeteye
A SpeechT5 model fine-tuned for Amharic speech synthesis.
Shook (Amharic ASR experiments)
None declaredBiniyam Daniel (b1n1yam)
A large set of Amharic ASR and language-model checkpoints, mostly Whisper fine-tunes, published in several sizes.
A personal account, not an Addis AI product release, though the author is a member of that organisation. Most cards are the auto-generated training template, so any reported WER is against the author's own evaluation set and is not comparable with figures measured on a shared benchmark. No licence is declared on the weights.
EthioLLM
MITEthioNLP
A family of multilingual language models covering several Ethiopian languages, released with an accompanying paper.
YehaTranslate
See model cardHasab AI
A translation model for Amharic from Hasab AI.
Hasab AI's only public release on the Hub at the time of checking.
Text
8 entries
Amharic Sentences Corpus
Apache 2.0a3xrfgb · 1M–10M rows
A large Amharic sentence corpus extracted from public Telegram channel exports.
The extraction tooling is open source, so the corpus can be reproduced and extended.
Amharic News Text Classification
CC BY 4.0Israel Abebe Azime · 10K–100K rows
A labelled Amharic news corpus with topic categories (Azime & Mohammed, 2021).
Widely cited. Also mirrored on GitHub.
Amharic News Category Classification
CC BY 4.0rasyosef · 10K–100K rows
Amharic news articles labelled by category, also usable for summarisation.
Amharic Instruction Dataset
None declaredEthioNLP · 100K–1M rows
Instruction-tuning data for Amharic, produced alongside the EthioLLM model family.
The largest open Amharic instruction set we know of. No licence is declared, so ask the authors before using it.
EthioSenti, EthioHate and EthioPOS
MIT (EthioSenti); varies by setEthioNLP
Labelled corpora for sentiment, hate-speech detection and part-of-speech tagging across several Ethiopian languages.
EthioNLP is the most consistently active academic group in this space, and also maintains the resource survey listed below.
Amharic Combined Corpus
None declaredBiniyam Daniel (b1n1yam) · 10M–100M rows
A very large aggregated Amharic text corpus, assembled for language-model pretraining.
The largest Amharic text collection on this page by an order of magnitude. No licence declared, and being an aggregate it likely carries mixed upstream terms; check before use.
Amharic Wikipedia
Apache 2.0Addis AI · 10K–100K rows
Amharic Wikipedia prepared for language-model training.
Tigrinya SQuAD
CC BY-SA 4.0fgaim · 10K–100K rows
A question-answering dataset for Tigrinya in SQuAD format.
The same publisher has Tigrinya abusive-language and OCR datasets.
Translation
1 entry
Directory
1 entry
Know one we are missing?
Everything here was opened and checked by hand, and the licence on each row is copied from the publisher rather than guessed. Send us anything that belongs here, including your own work. We list datasets that compete directly with ours: a researcher who cannot find data is a worse outcome than a researcher who picks someone else's.