Home/Speech/Alternatives

Alternatives hub · graph-backed

Speech alternatives

In short

Top alternatives to Speech are CosyVoice and espnet, ranked by typed graph edges - model-training.

Not a popularity vote. Each alternative is a typed graph neighbor of Speech in Developer Tools, Model Training, Speech & Audio - ranked by edge type and constraint overlap, with live GitHub stats shown for context.

Speech trust report - maintenance, provenance, and scan signals for Speech.

GraphCanon updated 2w · GitHub pushed 2w

Speech alternatives (markdown)

Constraints24 of 24 match
CosyVoice logo
CosyVoicerelated

Multi-lingual large voice generation model with full-stack abilities for inference, training and deployment.

Pythonmodel-trainingspeech-audio
23k
stars
espnet logo
espnetrelated

End-to-End Speech Processing Toolkit

Pythonmodel-trainingspeech-audio
9.9k
stars
TTS logo
TTSrelated

🐸💬 - a deep learning toolkit for Text-to-Speech

Pythonmodel-trainingspeech-audio
46k
stars
AudioGPT logo
AudioGPTrelated

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Pythonspeech-audio
10k
stars
ChatTTS logo
ChatTTSrelated

A generative speech model for daily dialogue

Pythonspeech-audio
40k
stars
dc_tts logo
dc_ttsrelated

A TensorFlow Implementation of DC-TTS

Pythonspeech-audio
1.2k
stars
dia logo
diarelated

A TTS model for generating ultra-realistic dialogue

Pythonspeech-audio
19k
stars
EmotiVoice logo
EmotiVoicerelated

A Multi-Voice and Prompt-Controlled TTS Engine

Pythonspeech-audio
8.5k
stars
FluidAudio logo
FluidAudiorelated

CoreML audio models for text-to-speech, speech-to-text, voice activity detection and speaker diarization in Swift.

Swiftspeech-audio
2.6k
stars
Fun-ASR logo
Fun-ASRrelated

Fun-ASR-Nano LLM-ASR model supports 31 languages for real-time speech recognition tasks

FreemiumCspeech-audio
1.4k
stars
hifi-gan logo
hifi-ganrelated

Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Pythonspeech-audio
2.4k
stars
Kokoro-FastAPI logo
Kokoro-FastAPIrelated

Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model

Pythonspeech-audio
5.3k
stars
metavoice-src logo
metavoice-srcrelated

Foundational model for human-like, expressive TTS

Pythonspeech-audio
4.2k
stars
MockingBird logo
MockingBirdrelated

Clone a voice in 5 seconds to generate arbitrary speech in real-time

Pythonspeech-audio
37k
stars
sherpa-onnx logo
sherpa-onnxrelated

Speech-to-text and related audio processing tools using ONNX with cross-platform support

C++speech-audio
14k
stars
silero-models logo
silero-modelsrelated

Silero Models provide simple access to pre-trained text-to-speech models

Jupyter Notebookspeech-audio
6.0k
stars
Speech-AI-Forge logo
Speech-AI-Forgerelated

A project for TTS generation with API and WebUI implementations

FreemiumPythonspeech-audio
1.4k
stars
speech-to-speech logo
speech-to-speechrelated

Build local voice agents with open-source models

FreemiumPythonspeech-audio
8.2k
stars
speechbrain logo
speechbrainrelated

A PyTorch-based Speech Toolkit

Pythonspeech-audio
12k
stars
StreamSpeech logo
StreamSpeechrelated

All-in-one speech recognition and synthesis model for offline and simultaneous processing

Pythonspeech-audio
1.3k
stars
STT logo
STTrelated

A fast open-source deep-learning toolkit for speech-to-text

C++speech-audio
2.6k
stars
supertonic logo
supertonicrelated

Lightning-Fast On-Device Multilingual TTS via ONNX

Swiftspeech-audio
14k
stars
tensorflow-speech-recognition logo
tensorflow-speech-recognitionrelated

Speech recognition using TensorFlow deep learning framework

FreemiumPythonspeech-audio
2.2k
stars
TensorFlowTTS logo
TensorFlowTTSrelated

Real-Time State-of-the-art Speech Synthesis for Tensorflow 2

Pythonspeech-audio
4.0k
stars

When NOT to use Speech

Constraint-first guidance from category fit and live maintenance signals - not marketing copy.

  • For environments where GPU access is limited or unavailable since the toolkit highly recommends a GPU setup for both training and recommended for inference.
  • If your Python/PyTorch/CUDA versions fall below the specified requirements (Python 3.12+, PyTorch 2.7+), as lower versions will not be compatible with NeMo Speech.
  • In scenarios where you're working with models that do not require or benefit significantly from GPU acceleration, given its architecture optimized for GPU use.

Related alternatives hubs

High-intent OSS-vs-OSS alternatives pages elsewhere in the graph (including vector-DB picks for Pinecone-style queries).

Head-to-head comparisons

Common questions

What are the best alternatives to Speech?
Graph-backed alternatives to Speech include CosyVoice, espnet, TTS, AudioGPT, ChatTTS. GraphCanon ranks them by typed relationship edges and constraint overlap from decision_facts - not marketing votes or raw star sort.
How does GraphCanon rank Speech alternatives?
Direct alternative and successor edges from the knowledge graph come first, ordered by edge type and shared constraint facets (persona, runtime, hosting). Category neighbours fill the list only after curated edges. Stars are shown for context, not as the primary sort.
When should I avoid Speech?
For environments where GPU access is limited or unavailable since the toolkit highly recommends a GPU setup for both training and recommended for inference. If your Python/PyTorch/CUDA versions fall below the specified requirements (Python 3.12+, PyTorch 2.7+), as lower versions will not be compatible with NeMo Speech. In scenarios where you're working with models that do not require or benefit significantly from GPU acceleration, given its architecture optimized for GPU use.
Is Speech open source?
Yes. Speech is an open-source project on GitHub under the Apache-2.0 license, with 17,940 stars.
What is Speech used for?
NVIDIA-NeMo/Speech is a comprehensive toolkit for Automatic Speech Recognition (ASR), speaker diarization, speaker recognition, speech synthesis (TTS), and more. It is built on PyTorch and supports CUDA for efficient GPU utilization.
What category is Speech in?
Speech is categorized under Developer Tools, Model Training, Speech & Audio in the GraphCanon knowledge graph.
How do Speech alternatives compare head-to-head?
Each alternative has a neutral compare page against Speech, for example CosyVoice vs Speech, espnet vs Speech, TTS vs Speech. Stats come from live GitHub metadata.
Is there a machine-readable alternatives list?
Yes. The markdown twin at Speech alternatives lists direct alternatives and same-category tools with internal links to each tool markdown page.
Where are other high-intent alternatives hubs?
Related P0 OSS-vs-OSS hubs: LangChain alternatives, LlamaIndex alternatives, Qdrant alternatives, FinRobot alternatives, free-llm-api-resources alternatives, caveman alternatives, rtk alternatives, unsloth alternatives, ollama alternatives. Vector-database intent (including Pinecone-style queries) is covered at Qdrant alternatives.
Where can I see maintenance and security signals for Speech?
GraphCanon publishes a sourced trust report for Speech at Speech trust report - maintenance posture, fork provenance, and dependency/MCP scan status with methodology tags. Not a safety grade.

Was this helpful?

Anonymous feedback helps us improve pages and translations.