Home/Speech & Audio/open-speech-corpora
open-speech-corpora logo

open-speech-corpora

coqui-ai/open-speech-corpora

A list of accessible speech corpora for ASR, TTS, and other Speech Technologies

GraphCanon updated 3w · GitHub synced 3w

1.4k stars151 forksLast push 2y MIT

Decision brief

Open-Speech-Corpora indexes diverse speech datasets critical for ASR and TTS projects in various languages.

Good fit when

  • You are working on projects requiring specific language corpora like Amharic, Hausa, or Wolof.
  • Your project needs a mix of licenses and sizes to cater to different research or compliance needs.

Avoid when

  • You exclusively need data for voice cloning with a high number of speakers.
  • If your focus is on specialized voice emotion recognition datasets not covered in the repository.

Observed Jul 17, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Dormant (783d since push)
As of 3w
Provenance
Not a fork · Organization account
As of 3w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/coqui-ai/open-speech-corpora

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

This repository curates lists of speech corpora that can be used in various applications such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and more. The collection includes datasets with diverse languages and sizes to support multiple research areas in the speech technology field.

Capability facts

No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).

Categories

Tags

README

📜 GNU General Public License

CORPUSLANGUAGES# HOURS# SPEAKERSDOWNLOADLICENSE
VoxForgeEnglish~120 hours~2966 speakershttp://www.repository.voxforge1.org/downloads/en/Trunk/Audio/Main/16kHz_16bit/ https://voice.mozilla.org/en/datasetsGNU-GPL 3.0
VoxForgeRussianhttp://www.repository.voxforge1.org/downloads/ru/Trunk/Audio/Main/16kHz_16bit/ http://www.repository.voxforge1.org/downloads/Russian/Trunk/Audio/Main/16kHz_16bit/GNU-GPL 3.0
VoxForgeGermanhttp://www.repository.voxforge1.org/downloads/de/Trunk/Audio/Main/16kHz_16bit/GNU-GPL 3.0

📜 Apache License

CORPUSLANGUAGES# HOURS# SPEAKERSDOWNLOADLICENSE
AISHELL-1Mandarin170 hours400 speakershttp://www.openslr.org/33/Apache 2.0
Tunisian_MSAModern Standard Arabic (Tunisia)11.2 hours118 speakershttp://www.openslr.org/46/Apache 2.0
African Accented FrenchFrench22 hours232 speakershttp://www.openslr.org/57/Apache 2.0
THCHS-30Mandarin Chinese33.57 hours (13,389 utterances)40 speakers (31 female; 9 male)http://www.openslr.org/18/Apache 2.0
Living Audio Dataset - DutchDutch57:49 min1 speakerhttps://github.com/Idlak/Living-Audio-DatasetApache 2.0
Living Audio Dataset - EnglishEnglish50:50 min1 speakerhttps://github.com/Idlak/Living-Audio-DatasetApache 2.0
Living Audio Dataset - IrishIrish61:56 min1 speakerhttps://github.com/Idlak/Living-Audio-DatasetApache 2.0
Living Audio Dataset - RussianRussian34:58 min1 speakerhttps://github.com/Idlak/Living-Audio-DatasetApache 2.0

📜 MIT License

CORPUSLANGUAGES# HOURS# SPEAKERSDOWNLOADLICENSE
ALFFAAmharic;Hausa (paid); Swahili; Wolofhttp://www.openslr.org/25/ https://github.com/besacier/ALFFA_PUBLICMIT

📜 BSD 3-Clause License

CORPUSLANGUAGES# HOURS# SPEAKERSDOWNLOADLICENSE
M-AILABS German CorpusGerman237 hours and 22 minuteshttp://www.caito.de/data/Training/stt_tts/de_DE.tgzM-AILABS LICENSE (a data-specific BSD 3-Clause License)
M-AILABS Queen's English CorpusQueen's English45 hours and 35 minuteshttp://www.caito.de/data/Training/stt_tts/en_UK.tgzM-AILABS LICENSE (a data-specific BSD 3-Clause License)
M-AILABS US English CorpusAmerican English102 hours and 7 minuteshttp://www.caito.de/data/Training/stt_tts/en_US.tgzM-AILABS LICENSE (a data-specific BSD 3-Clause License)
M-AILABS Spanish CorpusSpanish Spanish108 hours and 34 minuteshttp://www.caito.de/data/Training/stt_tts/es_ES.tgzM-AILABS LICENSE (a data-specific BSD 3-Clause License)
M-AILABS Italian CorpusItalian127 hours and 40 minuteshttp://www.caito.de/data/Training/stt_tts/it_IT.tgzM-AILABS LICENSE (a data-specific BSD 3-Clause License)
M-AILABS Ukrainian CorpusUkrainian87 hours and 8 minuteshttp://www.caito.de/data/Training/stt_tts/uk_UK.tgz[M-AILABS LICENSE](https:

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.