{"data":{"slug":"yl4579-styletts2","name":"StyleTTS2","tagline":"StyleTTS 2 advances human-like text-to-speech using style diffusion and adversarial training.","github_url":"https://github.com/yl4579/StyleTTS2","owner":"yl4579","repo":"StyleTTS2","owner_avatar_url":"https://avatars.githubusercontent.com/u/71044569?v=4","primary_language":"Python","stars":6322,"forks":694,"topics":["adversarial-training","deep-learning","diffusion-models","gan","latent-diffusion","latent-diffusion-models","pytorch","speaker-adaptation","speech-synthesis","text-to-speech","tts","wavlm"],"archived":false,"github_pushed_at":"2024-08-10T00:48:18+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/yl4579-styletts2","markdown_url":"https://www.graphcanon.com/tools/yl4579-styletts2.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/yl4579-styletts2","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=yl4579-styletts2","description":"StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models","homepage_url":null,"license":"MIT","open_issues":118,"watchers":83,"ai_summary":"This project focuses on enhancing the quality of synthesized speech in text-to-speech applications through advanced models such as style diffusion and generative adversarial networks, specifically addressing speaker adaptation for a more natural speaking voice.","readme_excerpt":"## License\n\nCode: MIT License\n\nPre-Trained Models: Before using these pre-trained models, you agree to inform the listeners that the speech samples are synthesized by the pre-trained models, unless you have the permission to use the voice you synthesize. That is, you agree to only use voices whose speakers grant the permission to have their voice cloned, either directly or by license before making synthesized voices public, or you have to publicly announce that these voices are synthesized if you do not have the permission to use these voices.","github_created_at":"2023-06-14T00:48:11+00:00","created_at":"2026-07-11T12:06:08.469968+00:00","updated_at":"2026-07-29T12:00:05.663709+00:00","categories":[{"slug":"speech-audio","name":"Speech & Audio","url":"https://www.graphcanon.com/categories/speech-audio","markdown_url":"https://www.graphcanon.com/categories/speech-audio.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/speech-audio"}],"tags":[{"slug":"adversarial-training","name":"adversarial training"},{"slug":"deep-learning","name":"deep-learning"},{"slug":"diffusion-models","name":"diffusion-models"},{"slug":"gan","name":"gan"},{"slug":"latent-diffusion","name":"latent-diffusion"},{"slug":"pytorch","name":"pytorch"},{"slug":"speech-synthesis","name":"speech-synthesis"},{"slug":"tts","name":"tts"}],"trust":{"provenance":{"is_fork":false,"github_id":653386433,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-07-29T12:00:04.828Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":718,"last_release_at":null},"security_summary":{"status":"ok","scanner":"osv@v1","low_count":0,"high_count":0,"last_scan_at":"2026-07-11T12:06:10.030Z","medium_count":0,"scan_profile":"deps","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-07-29T12:00:05.298Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-07-29T12:00:05.298Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-07-29T12:00:05.298Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When you need highly natural speech synthesis with accurate speaker adaptation through advanced generative models","If your project allows for informing listeners that the voice is synthesized, adhering to ethical usage guidelines"],"when_not_to_use":["Avoid if requirements do not align with using models specifically trained on large speech language models","Do not use in contexts where explicit disclosure of synthesis is not feasible or appropriate as per ethical considerations"],"source":"enrich:decision_facts","observed_at":"2026-07-16T23:02:18.202Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"StyleTTS2 leverages style diffusion and GANs for superior speaker adaptation in text-to-speech applications."}]}}