{"data":{"slug":"vllm-project-semantic-router","name":"semantic-router","tagline":"System level intelligent runtime for Mixture-of-Models across edge, data center and cloud","github_url":"https://github.com/vllm-project/semantic-router","owner":"vllm-project","repo":"semantic-router","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Go","stars":5241,"forks":825,"topics":["ai-gateway","bert-classification","fine-tuning","golang","huggingface-candle","huggingface-transformers","kubernetes","llm","llmrouter","mixture-of-models","pii-detection","prompt-engineering","prompt-guard","rust","semantic-router","vllm"],"archived":false,"github_pushed_at":"2026-08-23T17:59:46+00:00","maintenance_label":"Very active","stars_delta_30d":202,"url":"https://www.graphcanon.com/tools/vllm-project-semantic-router","markdown_url":"https://www.graphcanon.com/tools/vllm-project-semantic-router.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-semantic-router","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-semantic-router","description":"A programmable Mixture-of-Models router for heterogeneous LLM inference","homepage_url":"https://vllm-sr.ai","license":"Apache-2.0","open_issues":357,"watchers":58,"ai_summary":"Provides a smart runtime environment for managing mixed models of AI systems across different deployment environments like edge devices, data centers, and cloud services.","readme_excerpt":"### Install\n\n```bash\ncurl -fsSL https://vllm-sr.ai/install.sh | bash\n```\n\nFor platform notes, detailed setup options, and troubleshooting, see the **[Installation Guide](https://vllm-sr.ai/docs/installation/)**.\n\n<details>\n<summary>Online playground credentials</summary>\n\n- Username: `love@vllm-sr.ai`\n- Password: `vllm-sr`\n\n</details>","github_created_at":"2025-08-26T21:49:50+00:00","created_at":"2026-07-11T11:37:27.800716+00:00","updated_at":"2026-08-23T18:01:11.369647+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"ai-gateway","name":"ai-gateway"},{"slug":"bert-classification","name":"bert-classification"},{"slug":"fine-tuning","name":"fine-tuning"},{"slug":"golang","name":"golang"},{"slug":"huggingface-candle","name":"huggingface-candle"},{"slug":"huggingface-transformers","name":"huggingface-transformers"},{"slug":"kubernetes","name":"kubernetes"},{"slug":"llmrouter","name":"llmrouter"}],"trust":{"provenance":{"is_fork":false,"github_id":1045247072,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-23T18:01:10.623Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":1,"days_since_push":0,"last_release_at":"2026-06-05T12:08:52Z","stars_delta_30d":202,"open_issues_delta_30d":90},"security_summary":{"status":"no_manifest","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:37:29.139Z","medium_count":0,"scan_profile":"mcp_manifest","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-23T18:01:11.089Z"},"languages":{"value":["go"],"source":"github.language","observed_at":"2026-08-23T18:01:11.089Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-23T18:01:11.089Z"}},"decision_facts":{"hosting":null,"pricing":{"model":"freemium","summary":"Open-source under Apache-2.0 license, enabling free usage and modification with no cost for the core service."},"requirements":{"notes":["Requires Go runtime and Docker setup, suitable environments include Kubernetes clusters."],"min_ram_gb":2,"requires_docker":true},"constraints":{"min_ram_gb":2,"pricing_model":"freemium","requires_docker":true},"when_to_use":["- Semantic-Router is ideal when you need an intelligent system to manage a mixture of AI models in different deployment contexts like edge, data center, or cloud.","- Use it if your project involves fine-tuning models and deploying them efficiently using Kubernetes, leveraging its support for HuggingFace transformers and Candle libraries."],"when_not_to_use":["- Avoid Semantic-Router if your organization strictly uses languages other than Go since the system is language-specific, which might complicate integration.","- For projects that do not require cross-environment deployment capabilities (such as those exclusively running on cloud infrastructure), Semantic-Router may introduce unnecessary complexity."],"source":"enrich:decision_facts","observed_at":"2026-07-12T16:29:24.149Z"},"constraint_facets":{"min_ram_gb":2,"pricing_model":"freemium","requires_docker":true},"decision_summary":[{"label":"Pricing","value":"freemium - Open-source under Apache-2.0 license, enabling free usage and modification with no cost for the core service."},{"label":"Requirements","value":"Min 2 GB RAM; Requires Docker; Requires Go runtime and Docker setup, suitable environments include Kubernetes clusters."},{"label":"Adopt for","value":"Semantic-Router is a Go-based intelligent runtime optimized for managing AI models across diverse deployment environments including edge devices, data centers, and cloud services."}]}}