Third-party datasets

Other people's Ethiopian language datasets

Speech, text and translation data published by other researchers and organisations. None of it is ours. It is gathered here so that anyone starting work on an Ethiopian language has one place to look instead of searching and hoping.

These are third-party datasets. Read this first.

None of this is our data, and we have not checked its quality. Every entry belongs to the researcher or organisation that published it. We have not listened to it, measured it, or verified that it contains what its authors say it contains. This is a directory to make things findable, nothing more. Judge each one yourself, and credit its authors, not us.

We host none of it. Every link goes to the publisher, and the licence shown is the one they declared, copied rather than inferred. Where a publisher declared no licence we say so, because that means you have no permission to use it until you ask them.

Links and licences were last confirmed by hand on 29 August 2026. Things move. If you find something out of date or missing, tell us and we will fix it.

5 of the 28 entries below declare no licence at all. They are marked, and you should treat them as all rights reserved.

Speech

12 entries

ALFFA Amharic

MIT

ALFFA project, via hadamard-2 · 10K–100K rows

The classic academic Amharic read-speech corpus for ASR, converted to the Hugging Face datasets format.

Long-standing reference corpus; several published Amharic ASR models are trained on it, so it is a reasonable baseline to compare against.

Amharic

Amharic Speech Dataset 110HRS

None declared

KYAGABA · 10K–100K rows

One of the larger community-contributed Amharic speech collections.

The name claims 110 hours; the dataset card declares no licence and no hour count. Without a licence you have no permission to use it, so contact the author before building anything on it.

Amharic

Common Voice 17 (Amharic and Tigrinya)

CC0 1.0

Mozilla, via fsicoli

Mozilla's crowd-sourced read-speech corpus, with per-clip age, gender, accent and up/down vote metadata.

This is an unofficial community conversion, not Mozilla's own upload. The per-clip metadata schema is close to ours and is a useful reference point.

AmharicTigrinya100+ others

Amharic ASR Dataset

Apache 2.0

Ephrem · Under 1K rows

A small community-contributed Amharic ASR set.

Tagged under 1,000 rows, so treat it as a sample rather than a training corpus.

Amharic

Amharic Speech (protocol demonstration)

CC BY-NC 4.0

Silencio Network · Under 1K rows

A small speaker-metadata-rich set published to demonstrate a collection protocol rather than to train models.

Non-commercial licence, so it cannot be used to train a commercial model. Its card openly documents gender and regional imbalance, which is a good template for how any corpus should state its own limits.

Amharic

WaxalNLP

CC BY-SA 4.0 / CC BY 4.0

Google

Google's large multilingual African speech corpus, covering Amharic among more than forty languages.

By far the most downloaded dataset on this page. Widely used as the Amharic ASR training baseline, and the set most published WER figures are measured against. Dual-licensed; check which licence applies to the portion you use.

Amharic40+ African languages

Leyu Amharic dialect corpora

CC BY 4.0

Leyu · 1K–10K rows each

Five separate Amharic speech corpora, one per regional dialect: Addis Ababa, Shewa, Gojjam, Gonder and Wello.

Partitioned by dialect rather than pooled, which is unusual and useful if you care about regional coverage or want to measure how a model generalises across dialects. Openly licensed. Leyu works in the same space we do.

Amharic

Ethiopian Languages Speech Dataset

CC BY 4.0

Benji-fish · 1K–10K rows

Speech covering four Ethiopian languages in one release, openly licensed.

One of the few sets that spans several Ethiopian languages rather than Amharic alone.

AmharicOromoSidamaTigrinya

Sagalee

CC BY-NC 4.0

turiabu · 10K–100K rows

A read-speech corpus for Afaan Oromo ASR.

One of the largest open Oromo speech sets. Non-commercial licence, so it cannot be used to train a commercial model. Several re-uploads exist under other accounts; this is the one to cite.

Oromo

Ethiopian Speech collection

None declared

Badr al-Absi (badrex) · 10K–100K rows

Speech data across five Ethiopian languages, alongside per-language sets and the Ethio-ASR models trained on them.

The data behind the Ethio-ASR models. No licence is declared on the dataset, so ask before building on it, even though the models themselves are openly licensed.

AmharicOromoTigrinyaSidamaWolaytta

Tigrinya Speech Data

CC BY 4.0

Professor · 10K–100K rows

An openly licensed Tigrinya speech corpus.

Tigrinya is badly under-served; this is among the larger open sets. The same publisher has Wolaytta and Sidama equivalents.

Tigrinya

Amharic TTS Benchmark

Other (see dataset card)

Addis AI · Under 1K rows

An evaluation set for Amharic text-to-speech systems.

A benchmark rather than training data, published on the Addis AI organisation account. Addis AI also ran the independent seven-system ASR comparison against our own test set.

Amharic

Model

6 entries

SpeechBrain DVoice Amharic ASR

Apache 2.0

SpeechBrain

A wav2vec2 + CTC Amharic ASR system trained on the ALFFA corpus.

A model rather than a dataset. Listed because it is one of the few ready-to-run Amharic ASR baselines.

Amharic

Ethio-ASR

See individual models

Badr al-Absi (badrex)

A family of wav2vec2-BERT ASR models covering five Ethiopian languages, with both per-language and multilingual variants.

The widest open ASR coverage of Ethiopian languages we are aware of. We benchmark against these.

AmharicOromoTigrinyaSidaamaWolaytta

Amharic TTS (SpeechT5)

MIT

AddisuSeteye

A SpeechT5 model fine-tuned for Amharic speech synthesis.

Amharic

Shook (Amharic ASR experiments)

None declared

Biniyam Daniel (b1n1yam)

A large set of Amharic ASR and language-model checkpoints, mostly Whisper fine-tunes, published in several sizes.

A personal account, not an Addis AI product release, though the author is a member of that organisation. Most cards are the auto-generated training template, so any reported WER is against the author's own evaluation set and is not comparable with figures measured on a shared benchmark. No licence is declared on the weights.

Amharic

EthioLLM

MIT

EthioNLP

A family of multilingual language models covering several Ethiopian languages, released with an accompanying paper.

AmharicOromoTigrinyaSomaliand others

YehaTranslate

See model card

Hasab AI

A translation model for Amharic from Hasab AI.

Hasab AI's only public release on the Hub at the time of checking.

Amharic

Text

8 entries

Amharic Sentences Corpus

Apache 2.0

a3xrfgb · 1M–10M rows

A large Amharic sentence corpus extracted from public Telegram channel exports.

The extraction tooling is open source, so the corpus can be reproduced and extended.

Amharic

Amharic News Text Classification

CC BY 4.0

Israel Abebe Azime · 10K–100K rows

A labelled Amharic news corpus with topic categories (Azime & Mohammed, 2021).

Widely cited. Also mirrored on GitHub.

Amharic

Amharic News Category Classification

CC BY 4.0

rasyosef · 10K–100K rows

Amharic news articles labelled by category, also usable for summarisation.

Amharic

Amharic Instruction Dataset

None declared

EthioNLP · 100K–1M rows

Instruction-tuning data for Amharic, produced alongside the EthioLLM model family.

The largest open Amharic instruction set we know of. No licence is declared, so ask the authors before using it.

Amharic

EthioSenti, EthioHate and EthioPOS

MIT (EthioSenti); varies by set

EthioNLP

Labelled corpora for sentiment, hate-speech detection and part-of-speech tagging across several Ethiopian languages.

EthioNLP is the most consistently active academic group in this space, and also maintains the resource survey listed below.

AmharicOromoTigrinyaand others

Amharic Combined Corpus

None declared

Biniyam Daniel (b1n1yam) · 10M–100M rows

A very large aggregated Amharic text corpus, assembled for language-model pretraining.

The largest Amharic text collection on this page by an order of magnitude. No licence declared, and being an aggregate it likely carries mixed upstream terms; check before use.

Amharic

Amharic Wikipedia

Apache 2.0

Addis AI · 10K–100K rows

Amharic Wikipedia prepared for language-model training.

AmharicEnglish

Tigrinya SQuAD

CC BY-SA 4.0

fgaim · 10K–100K rows

A question-answering dataset for Tigrinya in SQuAD format.

The same publisher has Tigrinya abusive-language and OCR datasets.

Tigrinya

Translation

1 entry

HornMT

See repository

Asmelash Teka Hadgu and contributors

A machine translation benchmark covering six languages of the Horn of Africa.

One of very few resources covering several Ethiopian languages together rather than Amharic alone.

AmharicOromoTigrinyaSomaliAfarEnglish

Directory

1 entry

Ethiopian Language Survey

See repository

EthioNLP

A periodically updated survey and resource list of publicly available NLP resources for Ethiopian languages.

Broader than this page for text and NLP resources. If you are looking for something we do not list, start here.

AmharicOromoTigrinyaWolaytta

Know one we are missing?

Everything here was opened and checked by hand, and the licence on each row is copied from the publisher rather than guessed. Send us anything that belongs here, including your own work. We list datasets that compete directly with ours: a researcher who cannot find data is a worse outcome than a researcher who picks someone else's.

Reviewed by a person before it is listed.