open-speech-corpora
A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
GraphCanon updated 3w · GitHub synced 3w
Decision brief
Open-Speech-Corpora indexes diverse speech datasets critical for ASR and TTS projects in various languages.
Good fit when
- You are working on projects requiring specific language corpora like Amharic, Hausa, or Wolof.
- Your project needs a mix of licenses and sizes to cater to different research or compliance needs.
Avoid when
- You exclusively need data for voice cloning with a high number of speakers.
- If your focus is on specialized voice emotion recognition datasets not covered in the repository.
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Dormant (783d since push)
- As of 3w
- Provenance
- Not a fork · Organization account
- As of 3w
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
git clone https://github.com/coqui-ai/open-speech-corporaSimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
This repository curates lists of speech corpora that can be used in various applications such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and more. The collection includes datasets with diverse languages and sizes to support multiple research areas in the speech technology field.
Capability facts
No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).
Categories
Tags
README
📜 GNU General Public License
| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |
|---|---|---|---|---|---|
| VoxForge | English | ~120 hours | ~2966 speakers | http://www.repository.voxforge1.org/downloads/en/Trunk/Audio/Main/16kHz_16bit/ https://voice.mozilla.org/en/datasets | GNU-GPL 3.0 |
| VoxForge | Russian | http://www.repository.voxforge1.org/downloads/ru/Trunk/Audio/Main/16kHz_16bit/ http://www.repository.voxforge1.org/downloads/Russian/Trunk/Audio/Main/16kHz_16bit/ | GNU-GPL 3.0 | ||
| VoxForge | German | http://www.repository.voxforge1.org/downloads/de/Trunk/Audio/Main/16kHz_16bit/ | GNU-GPL 3.0 |
📜 Apache License
| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |
|---|---|---|---|---|---|
| AISHELL-1 | Mandarin | 170 hours | 400 speakers | http://www.openslr.org/33/ | Apache 2.0 |
| Tunisian_MSA | Modern Standard Arabic (Tunisia) | 11.2 hours | 118 speakers | http://www.openslr.org/46/ | Apache 2.0 |
| African Accented French | French | 22 hours | 232 speakers | http://www.openslr.org/57/ | Apache 2.0 |
| THCHS-30 | Mandarin Chinese | 33.57 hours (13,389 utterances) | 40 speakers (31 female; 9 male) | http://www.openslr.org/18/ | Apache 2.0 |
| Living Audio Dataset - Dutch | Dutch | 57:49 min | 1 speaker | https://github.com/Idlak/Living-Audio-Dataset | Apache 2.0 |
| Living Audio Dataset - English | English | 50:50 min | 1 speaker | https://github.com/Idlak/Living-Audio-Dataset | Apache 2.0 |
| Living Audio Dataset - Irish | Irish | 61:56 min | 1 speaker | https://github.com/Idlak/Living-Audio-Dataset | Apache 2.0 |
| Living Audio Dataset - Russian | Russian | 34:58 min | 1 speaker | https://github.com/Idlak/Living-Audio-Dataset | Apache 2.0 |
📜 MIT License
| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |
|---|---|---|---|---|---|
| ALFFA | Amharic;Hausa (paid); Swahili; Wolof | http://www.openslr.org/25/ https://github.com/besacier/ALFFA_PUBLIC | MIT |
📜 BSD 3-Clause License
| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |
|---|---|---|---|---|---|
| M-AILABS German Corpus | German | 237 hours and 22 minutes | http://www.caito.de/data/Training/stt_tts/de_DE.tgz | M-AILABS LICENSE (a data-specific BSD 3-Clause License) | |
| M-AILABS Queen's English Corpus | Queen's English | 45 hours and 35 minutes | http://www.caito.de/data/Training/stt_tts/en_UK.tgz | M-AILABS LICENSE (a data-specific BSD 3-Clause License) | |
| M-AILABS US English Corpus | American English | 102 hours and 7 minutes | http://www.caito.de/data/Training/stt_tts/en_US.tgz | M-AILABS LICENSE (a data-specific BSD 3-Clause License) | |
| M-AILABS Spanish Corpus | Spanish Spanish | 108 hours and 34 minutes | http://www.caito.de/data/Training/stt_tts/es_ES.tgz | M-AILABS LICENSE (a data-specific BSD 3-Clause License) | |
| M-AILABS Italian Corpus | Italian | 127 hours and 40 minutes | http://www.caito.de/data/Training/stt_tts/it_IT.tgz | M-AILABS LICENSE (a data-specific BSD 3-Clause License) | |
| M-AILABS Ukrainian Corpus | Ukrainian | 87 hours and 8 minutes | http://www.caito.de/data/Training/stt_tts/uk_UK.tgz | [M-AILABS LICENSE](https: |
For agents
This page has a .md twin and JSON over the API.