---
title: "Speech & Audio"
type: "category"
slug: "speech-audio"
canonical_url: "https://www.graphcanon.com/categories/speech-audio"
tool_count: 205
---

# Speech & Audio

*GraphCanon updated Aug 22, 2026*

Speech-to-text, text-to-speech, and audio models and pipelines (Whisper, TTS engines, ASR).

205 tools in this category (showing the top 60 by stars).

## Featured comparisons

- [rig vs LocalAI](/compare/0xplaygrounds-rig-vs-mudler-localai.md)
- [ChatTTS vs LLaMA-Omni](/compare/2noise-chattts-vs-ictnlp-llama-omni.md)
- [ChatTTS vs whisper](/compare/2noise-chattts-vs-openai-whisper.md)
- [voice-pro vs GPT-SoVITS](/compare/abus-aikorea-voice-pro-vs-rvc-boss-gpt-sovits.md)
- [llmfit vs LocalAI](/compare/alexsjones-llmfit-vs-mudler-localai.md)
- [yalm vs whisper.cpp](/compare/andrewkchan-yalm-vs-ggml-org-whisper-cpp.md)
- [bionic-gpt vs LocalAI](/compare/bionic-gpt-bionic-gpt-vs-mudler-localai.md)
- [NextChat vs py-gpt](/compare/chatgptnextweb-nextchat-vs-szczyglis-dev-py-gpt.md)

## Tools

- [transformers](/tools/huggingface-transformers.md) - Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models (★ 164,121) [Very active]
- [LocalAI](/tools/mudler-localai.md) - Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required. (★ 48,500) [Very active]
- [whisper.cpp](/tools/ggml-org-whisper-cpp.md) - Port of OpenAI's Whisper model in C/C++ for speech-to-text inference (★ 52,501) [Very active]
- [screenpipe](/tools/screenpipe-screenpipe.md) - AI that records and analyzes everything you do, say, hear locally (★ 20,534) [Very active]
- [FunASR](/tools/modelscope-funasr.md) - Industrial-grade speech recognition toolkit (★ 19,554) [Very active]
- [mastra](/tools/mastra-ai-mastra.md) - Modern TypeScript framework for AI-powered applications and agents (★ 27,229) [Very active]
- [CosyVoice](/tools/funaudiollm-cosyvoice.md) - Multi-lingual large voice generation model with full-stack abilities for inference, training and deployment. (★ 22,373) [Steady]
- [nuclear](/tools/nukeop-nuclear.md) - Streaming music player that finds free music for you (★ 18,297) [Very active]
- [cactus](/tools/cactus-compute-cactus.md) - Low-latency AI engine for mobile devices & wearables (★ 5,535) [Very active]
- [ChatTTS](/tools/2noise-chattts.md) - A generative speech model for daily dialogue (★ 39,768) [Slowing]
- [bootcamp](/tools/milvus-io-bootcamp.md) - Dealing with all unstructured data including reverse image search, audio search, molecular search, video analysis, and question-answer systems. (★ 2,443) [Active]
- [speechbrain](/tools/speechbrain-speechbrain.md) - A PyTorch-based Speech Toolkit (★ 11,725) [Steady]
- [whisper](/tools/openai-whisper.md) - Robust Speech Recognition via Large-Scale Weak Supervision (★ 106,740) [Active]
- [dograh](/tools/dograh-hq-dograh.md) - Self-hosted open source voice AI platform (★ 5,064) [Very active]
- [ODS](/tools/osmantic-ods.md) - Transform personal computers into AI servers. (★ 3,799) [Very active]
- [py-gpt](/tools/szczyglis-dev-py-gpt.md) - Desktop AI Assistant powered by multiple LLMs and various functionalities (★ 1,880) [Very active]
- [meetily](/tools/zackriya-solutions-meetily.md) - Privacy first, AI meeting assistant with local processing (★ 29,268) [Steady]
- [ailia-models](/tools/ailia-ai-ailia-models.md) - Repository of pre-trained AI models for ailia SDK (★ 2,365) [Very active]
- [mlx-audio](/tools/blaizzy-mlx-audio.md) - A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library on Apple's MLX framework. (★ 7,639) [Very active]
- [Kokoro-FastAPI](/tools/remsky-kokoro-fastapi.md) - Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model (★ 5,265) [Active]
- [Speech](/tools/nvidia-nemo-speech.md) - A scalable generative AI framework for Speech AI (★ 17,940) [Very active]
- [argmax-oss-swift](/tools/argmaxinc-argmax-oss-swift.md) - On-device Speech AI for Apple Silicon (★ 6,294) [Very active]
- [openwhispr](/tools/openwhispr-openwhispr.md) - Voice-to-text dictation app with local and cloud models (★ 5,018) [Very active]
- [WhisperLive](/tools/collabora-whisperlive.md) - A nearly-live implementation of OpenAI's Whisper for real-time voice recognition (★ 4,190) [Very active]
- [SmartSub](/tools/buxuku-smartsub.md) - A desktop subtitle tool supporting ASR and translation (★ 4,624) [Very active]
- [SenseVoice](/tools/funaudiollm-sensevoice.md) - Multilingual speech understanding toolkit with ASR, emotion recognition, and audio event detection. (★ 8,961) [Very active]
- [MOSS-TTS](/tools/openmoss-moss-tts.md) - An open-source speech and sound generation model family designed for high-fidelity scenarios including multi-speaker dialogue。 (★ 3,922) [Very active]
- [DiffSinger](/tools/moonintheriver-diffsinger.md) - Singing Voice Synthesis via Shallow Diffusion Mechanism (★ 4,834) [Very active]
- [espnet](/tools/espnet-espnet.md) - End-to-End Speech Processing Toolkit (★ 9,903) [Very active]
- [VoxCPM](/tools/openbmb-voxcpm.md) - Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning (★ 34,452) [Active]
- [ChatGPT-On-CS](/tools/cs-lazy-tools-chatgpt-on-cs.md) - 智能对话客服工具，支持多平台接入和多种AI模型 (★ 4,264) [Steady]
- [GPT-SoVITS](/tools/rvc-boss-gpt-sovits.md) - Voice Cloning and Text-to-Speech with Minimal Voice Data (★ 60,178) [Very active]
- [Handy](/tools/cjpais-handy.md) - Free, open source speech-to-text application for offline use (★ 27,869) [Very active]
- [openless](/tools/open-less-openless.md) - AI-polished text from voice input for macOS and Windows (★ 2,881) [Very active]
- [parlor](/tools/fikrikarim-parlor.md) - On-device real-time multimodal AI for voice and vision (★ 1,914) [Very active]
- [FunClip](/tools/modelscope-funclip.md) - A video transcription and subtitle generation tool with LLM-assisted functionality. (★ 6,085) [Very active]
- [speech-to-speech](/tools/huggingface-speech-to-speech.md) - Build local voice agents with open-source models (★ 8,219) [Very active]
- [pyvideotrans](/tools/jianchang512-pyvideotrans.md) - Translate video language and embed dubbing & subtitles (★ 18,478) [Very active]
- [vosk-api](/tools/alphacep-vosk-api.md) - Offline speech recognition API (★ 15,013) [Active]
- [faster-whisper](/tools/systran-faster-whisper.md) - Faster Whisper transcription with CTranslate2 (★ 24,689) [Slowing]
- [sherpa-onnx](/tools/k2-fsa-sherpa-onnx.md) - Speech-to-text and related audio processing tools using ONNX with cross-platform support (★ 13,842) [Very active]
- [FluidAudio](/tools/fluidinference-fluidaudio.md) - CoreML audio models for text-to-speech, speech-to-text, voice activity detection and speaker diarization in Swift. (★ 2,554) [Very active]
- [supertonic](/tools/supertone-inc-supertonic.md) - Lightning-Fast On-Device Multilingual TTS via ONNX (★ 13,543) [Very active]
- [Applio](/tools/iahispano-applio.md) - A simple high-quality voice conversion tool focused on ease of use and performance (★ 3,530) [Very active]
- [LLPlayer](/tools/umlx5h-llplayer.md) - Media player for language learning featuring dual subtitles, AI-generated subtitles and real-time translation (★ 4,027) [Active]
- [Open-LLM-VTuber](/tools/open-llm-vtuber-open-llm-vtuber.md) - Voice interaction with LLMs and Live2D visuals (★ 13,224) [Slowing]
- [pluely](/tools/iamsrikanthnani-pluely.md) - Privacy-first AI assistant for meetings and interviews (★ 2,354) [Active]
- [TTS-WebUI](/tools/rsxdalv-tts-webui.md) - A comprehensive WebUI for multiple TTS systems (★ 3,216) [Very active]
- [MockingBird](/tools/babysor-mockingbird.md) - Clone a voice in 5 seconds to generate arbitrary speech in real-time (★ 36,913) [Slowing]
- [OmniVoice-Studio](/tools/debpalash-omnivoice-studio.md) - The open-source ElevenLabs alternative for local voice cloning and related tasks (★ 9,209) [Very active]
- [Foundry-Local](/tools/microsoft-foundry-local.md) - SDK and CLI for local AI inference with GPU acceleration, supporting speech-to-text models like Whisper. (★ 2,479) [Very active]
- [AudioNotes](/tools/harry0703-audionotes.md) - Quickly extracts audio and video content into structured markdown notes (★ 2,259) [Very active]
- [espeak-ng](/tools/espeak-ng-espeak-ng.md) - Open source speech synthesizer supporting more than hundred languages and accents (★ 6,688) [Very active]
- [index-tts](/tools/index-tts-index-tts.md) - A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech (★ 22,239) [Active]
- [whisperX](/tools/m-bain-whisperx.md) - WhisperX for automatic speech recognition with word-level timestamps and diarization (★ 23,329) [Active]
- [silero-models](/tools/snakers4-silero-models.md) - Silero Models provide simple access to pre-trained text-to-speech models (★ 6,030) [Steady]
- [abogen](/tools/denizsafak-abogen.md) - Generate audiobooks from EPUBs, PDFs and text with synchronized captions. (★ 5,429) [Very active]
- [Home-AssistantConfig](/tools/ccostan-home-assistantconfig.md) - Home Assistant configuration and documentation for smart home setup. (★ 5,256) [Very active]
- [awesome-generative-ai](/tools/filipecalegario-awesome-generative-ai.md) - A comprehensive list of generative AI resources (★ 3,524) [Slowing]
- [voice-pro](/tools/abus-aikorea-voice-pro.md) - Gradio WebUI for TTS and voice cloning with audio processing capabilities (★ 11,348) [Active]

## Common questions

### What are the best speech & audio tools?

GraphCanon ranks Speech & Audio tools by GitHub adoption and freshness. transformers is the current leader (164,121 stars). See the full list on this page - sorted by stars, with [maintenance labels](/glossary/trust-and-signals/maintenance-label) and graph relationships.

### How does GraphCanon rank Speech & Audio tools?

We sort by GitHub stars and push recency on category pages, not paid placement. Alternatives and compare pages use [typed graph edges](/glossary/knowledge-graph/typed-edge) (alternative, successor, integrates_with) plus shared categories - constraint-first, not marketing votes.

### How many tools are in Speech & Audio?

205 published tools are tagged with Speech & Audio in the GraphCanon knowledge graph.

### What are popular Speech & Audio comparisons?

Head-to-head compare pages in this category include rig vs LocalAI, ChatTTS vs LLaMA-Omni, ChatTTS vs whisper. Each comparison uses live GitHub stats and optional [trust signals](/glossary/trust-and-signals/trust-signal) - see the comparisons block on this page.

### Is there a machine-readable Speech & Audio list?

Yes. Append `.md` to this URL or fetch [`/md/categories/speech-audio`](/md/categories/speech-audio) for a markdown twin. The JSON API exposes the same corpus at [`/api/graphcanon/categories/speech-audio`](/api/graphcanon/categories/speech-audio).

---

**Machine-readable endpoints**

- JSON: [`/api/graphcanon/categories/speech-audio`](/api/graphcanon/categories/speech-audio)
- LLM index: [/llms.txt](/llms.txt)
- Full corpus: [/llms-full.txt](/llms-full.txt)

_GraphCanon - The knowledge graph for AI development. https://www.graphcanon.com/_
