Dataset.ET
  • Mission
  • Languages
  • How It Works
  • Community
  • Dataset
  • Directory
  • Writing
Contribute
Open Data

Datasets

Everything we publish is free under an open licence. Each release keeps its own permanent page, so a citation to a version never breaks.

v0.1.030 August 2026 · CC BY 4.0

Afaan Oromoo Speech

9.8 hours of read Afaan Oromoo speech from 74 contributors, screened and published free under CC BY 4.0.

9.8hours
3,594recordings
74speakers
v0.2.030 August 2026 · CC BY 4.0

Amharic Speech

51.5 hours of read Amharic speech from 493 contributors, screened and published free under CC BY 4.0.

51.5hours
16,866recordings
493speakers
Back to Home
Dataset.ET — The future of AI speaks Ethiopian

Open infrastructure for Ethiopian language AI. Dataset.ET is an open dataset initiative building speech, text, and translation datasets for 80+ Ethiopian languages.

Member ofNVIDIA Inception

Project

  • Mission
  • Languages
  • How It Works

Community

  • Contribute
  • Telegram

Resources

  • Datasets
  • Writing
  • Download on Hugging Face
  • Third-party datasets
  • Privacy Policy
  • Terms of Service

2026 Dataset.ET.

Managed by Snapwre.

Built with love for Ethiopian languages.