Dataset.ET
  • Mission
  • Languages
  • How It Works
  • Community
  • Dataset
  • Writing
Contribute
Writing

Notes from the work

What we are building, what we find when we look closely at the data, and what we get wrong along the way.

Release27 August 20267 min

We released 22.7 hours of Amharic speech, and Addis AI benchmarked it in days

Our first open dataset, what we found when we screened it, and why a quarter of the recordings humans approved were not good enough to ship.

Dataset.ET — The future of AI speaks Ethiopian

Open infrastructure for Ethiopian language AI. Dataset.ET is an open dataset initiative building speech, text, and translation datasets for 80+ Ethiopian languages.

Project

  • Mission
  • Languages
  • How It Works
  • Research (Coming Soon)

Community

  • Contribute
  • Telegram

Resources

  • Privacy Policy

2026 Dataset.ET.

Managed by Snapwre.

Built with love for Ethiopian languages.