{"data":{"slug":"huggingface-speech-to-speech","name":"speech-to-speech","tagline":"Build local voice agents with open-source models","github_url":"https://github.com/huggingface/speech-to-speech","owner":"huggingface","repo":"speech-to-speech","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":8219,"forks":1025,"topics":["ai","assistant","language-model","machine-learning","python","speech","speech-synthesis","speech-to-text","speech-translation"],"archived":false,"github_pushed_at":"2026-07-30T11:35:50+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/huggingface-speech-to-speech","markdown_url":"https://www.graphcanon.com/tools/huggingface-speech-to-speech.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-speech-to-speech","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-speech-to-speech","description":"Build local voice agents with open-source models","homepage_url":null,"license":"Apache-2.0","open_issues":121,"watchers":94,"ai_summary":"Package for creating localized speech-to-speech systems utilising open-source components for real-time and pre-recorded audio processing.","readme_excerpt":"## Installation\n\nRequires Python 3.10+.\n\n```bash\npip install speech-to-speech\n```\n\nThe default install covers the standard realtime path:\n\n- Parakeet TDT for STT\n- OpenAI-compatible API for the language model\n- Qwen3-TTS for speech output, using the GGML backend by default on non-macOS platforms and `mlx-audio` on Apple Silicon\n- local audio and realtime server modes\n\nmacOS and non-macOS dependencies are resolved automatically via platform markers in `pyproject.toml`.\n\n---\n\n### Docker\n\nInstall the [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html), then:\n\n```bash\ndocker compose up\n```\n\nThe compose file starts a llama.cpp server with Gemma 4, starts the TCP socket server, and exposes ports `8080`, `12345`, and `12346`.","github_created_at":"2024-08-07T15:32:09+00:00","created_at":"2026-07-11T12:15:07.2673+00:00","updated_at":"2026-07-30T12:00:04.812988+00:00","categories":[{"slug":"speech-audio","name":"Speech & Audio","url":"https://www.graphcanon.com/categories/speech-audio","markdown_url":"https://www.graphcanon.com/categories/speech-audio.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/speech-audio"}],"tags":[{"slug":"ai","name":"ai"},{"slug":"assistant","name":"assistant"},{"slug":"language-model","name":"language-model"},{"slug":"machine-learning","name":"machine-learning"},{"slug":"python","name":"python"},{"slug":"speech","name":"speech"},{"slug":"speech-synthesis","name":"speech-synthesis"},{"slug":"speech-to-text","name":"speech-to-text"}],"trust":{"provenance":{"is_fork":false,"github_id":839428333,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-07-30T12:00:03.763Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":1,"days_since_push":0,"last_release_at":"2026-06-11T19:35:07Z"},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T12:15:10.543Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-07-30T12:00:04.274Z"},"deploy":{"source":"dockerfile:Dockerfile","self_host":true,"observed_at":"2026-07-30T12:00:04.274Z","managed_saas":false},"has_cli":{"value":true,"source":"pyproject.toml:[project.scripts]","observed_at":"2026-07-30T12:00:04.274Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-07-30T12:00:04.274Z"},"has_docker":{"value":true,"source":"dockerfile:Dockerfile","observed_at":"2026-07-30T12:00:04.274Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-07-30T12:00:04.274Z"}},"decision_facts":{"hosting":null,"pricing":{"model":"freemium","summary":"Free and open-source software under the Apache-2.0 license, with possible premium services based on usage or special features not covered in this repository."},"requirements":{"notes":["Docker setup may require additional resources and the installation of the NVIDIA Container Toolkit for non-standard setups."],"min_ram_gb":4,"requires_docker":true},"constraints":{"min_ram_gb":4,"pricing_model":"freemium","requires_docker":true},"when_to_use":["When you need to leverage open-source components for real-time speech processing in your projects, as speech-to-speech provides an integrated solution with Parakeet TDT for STT.","For developers who intend to develop voice agents that require flexible deployment options, including both local audio and real-time server modes.","If your project environment is based on Python 3.10+ and you can benefit from the streamlined installation process provided by speech-to-speech."],"when_not_to_use":["When the need arises for a voice agent solution that exclusively utilizes proprietary models or services, as speech-to-speech depends fully on open-source components.","For projects aiming to run exclusively under macOS without cross-platform capabilities, despite automatic dependency resolution between different platforms."],"source":"enrich:decision_facts","observed_at":"2026-07-16T19:36:09.991Z"},"constraint_facets":{"min_ram_gb":4,"pricing_model":"freemium","requires_docker":true},"decision_summary":[{"label":"Pricing","value":"freemium - Free and open-source software under the Apache-2.0 license, with possible premium services based on usage or special features not covered in this repository."},{"label":"Requirements","value":"Min 4 GB RAM; Requires Docker; Docker setup may require additional resources and the installation of the NVIDIA Container Toolkit for non-standard setups."},{"label":"Adopt for","value":"speech-to-speech is an open-source Python package geared towards building localized voice agents via real-time and pre-recorded audio processing."}]}}