Dataset.ET
  • Mission
  • Languages
  • How It Works
  • Community
  • Try Hohe
  • Dataset
  • Directory
  • Writing
  • Contact
Contribute
Writing

Notes from the work

What we are building, what we find when we look closely at the data, and what we get wrong along the way.

Release21 September 20268 min

Hohe ASR: Amharic speech to Ge'ez letters, running on an ordinary processor

16.1% word error on speakers it has never heard, no GPU needed to run it, and an honest account of the 47% it gets wrong when two people talk at once. You can try it on Telegram.

Research6 September 20269 min

We tested every open Amharic speech model, then made the best one better

The smaller model beat the bigger one, two fair metrics disagreed about the winner, and a 538 MB file fixes three words in every hundred. You can talk to the result.

Research4 September 20268 min

When two speech models agree, are they right? We measured it, and our first answer was wrong

Benchmarking two open Amharic ASR models on our test set, and how a single truncation bug produced a result that looked like success.

Release27 August 20267 min

We released 22.7 hours of Amharic speech, and Addis AI benchmarked it in days

Our first open dataset, what we found when we screened it, and why one in seven recordings humans approved were not good enough to ship.

Dataset.ET — The future of AI speaks Ethiopian

Open infrastructure for Ethiopian language AI. Dataset.ET is an open dataset initiative building speech, text, and translation datasets for 80+ Ethiopian languages.

Member ofNVIDIA Inception

Project

  • Mission
  • Languages
  • How It Works

Community

  • Contribute
  • Telegram
  • Contact us

Resources

  • Datasets
  • Writing
  • Download on Hugging Face
  • Third-party datasets
  • Privacy Policy
  • Terms of Service

2026 Dataset.ET.

Managed by Snapwre.

Supported by DigiCom.

Built with love for Ethiopian languages.