{"data":{"slug":"peva3-smarterrouter","name":"SmarterRouter","tagline":"An intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI.","github_url":"https://github.com/peva3/SmarterRouter","owner":"peva3","repo":"SmarterRouter","owner_avatar_url":"https://avatars.githubusercontent.com/u/1492185?v=4","primary_language":"Python","stars":152,"forks":19,"topics":["ai-cache","ai-gateway","docker","fastapi","gpu-monitoring","llm","llm-proxy","llm-router","local-llm","model-serving","ollama","ollama-api","openai-proxy","self-hosted","self-hosted-ai","semantic-cache"],"archived":false,"github_pushed_at":"2026-05-10T02:47:50+00:00","maintenance_label":"Slowing","stars_delta_30d":6,"url":"https://www.graphcanon.com/tools/peva3-smarterrouter","markdown_url":"https://www.graphcanon.com/tools/peva3-smarterrouter.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/peva3-smarterrouter","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=peva3-smarterrouter","description":"SmarterRouter: An intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI. Features semantic caching, model profiling, and automatic failover for local AI labs.","homepage_url":null,"license":"MIT","open_issues":2,"watchers":5,"ai_summary":"SmarterRouter offers semantic caching, model profiling, automatic failover, and serves as an LLm proxy for local AI labs. Supports Docker, fastapi, GPU monitoring.","readme_excerpt":"## Quick Start (5 minutes)\n\nGet up and running with Docker in three commands:\n\n```bash\n\n---\n\n# 2. Start with Docker Compose\ndocker-compose up -d\n\n---\n\n# Generate optimal .env file based on hardware detection\npython -m smarterrouter generate-env\n```\n\nThe setup wizard automatically:\n- 🔍 Detects your Ollama installation (local, Docker, or remote)\n- ⚙️ Identifies GPU hardware (NVIDIA, AMD, Intel, Apple Silicon)\n- 📊 Analyzes available models and suggests optimal settings\n- 📝 Generates a tailored `.env` configuration file\n\n---\n\n### One-Line Docker Deployment (New in v2.1.5)\n\nFor the simplest deployment experience, use the included script:\n\n```bash\n\n---\n\n# Customize deployment\n./docker-run.sh --port 11436 --data-dir ./smarterrouter-data --env-file .env\n```\n\nThe script automatically:\n- 🐳 Detects GPU vendor and configures appropriate Docker device mounts\n- 📁 Creates persistent data directory\n- 🔧 Generates optimal configuration for your hardware\n- 🚀 Starts the container with proper restart policy\n\nFor production deployments, continue using `docker-compose.yml` with GPU-specific configurations.\n\n---\n\n## License\n\nMIT License - see [LICENSE](LICENSE) for details.","github_created_at":"2026-02-16T01:52:17+00:00","created_at":"2026-07-15T11:01:52.044566+00:00","updated_at":"2026-09-20T05:08:46.253316+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"ai-cache","name":"ai-cache"},{"slug":"ai-gateway","name":"ai-gateway"},{"slug":"docker-compose","name":"docker-compose"},{"slug":"fastapi","name":"fastapi"},{"slug":"gpu-monitoring","name":"gpu-monitoring"},{"slug":"llm-router","name":"llm-router"},{"slug":"model-serving","name":"model-serving"},{"slug":"ollama-api","name":"ollama-api"}],"trust":{"provenance":{"is_fork":false,"github_id":1158849733,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-09-20T05:08:43.532Z","maintenance":{"label":"Slowing","score":36,"methodology":"github_public_v1","releases_90d":0,"days_since_push":133,"last_release_at":"2026-04-18T16:59:50Z","stars_delta_30d":6,"open_issues_delta_30d":1},"security_summary":{"status":"findings","scanner":"osv@v1","low_count":3,"high_count":0,"last_scan_at":"2026-07-15T11:01:53.622Z","medium_count":0,"scan_profile":"deps","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-09-20T05:08:44.602Z"},"deploy":{"source":"dockerfile:Dockerfile","self_host":true,"observed_at":"2026-09-20T05:08:44.602Z","managed_saas":false},"has_cli":{"value":true,"source":"pyproject.toml:[project.scripts]","observed_at":"2026-09-20T05:08:44.602Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-09-20T05:08:44.602Z"},"has_docker":{"value":true,"source":"dockerfile:Dockerfile","observed_at":"2026-09-20T05:08:44.602Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-09-20T05:08:44.602Z"}},"decision_facts":{"hosting":null,"pricing":{"model":"freemium"},"requirements":{"min_ram_gb":4,"requires_docker":true},"constraints":{"min_ram_gb":4,"pricing_model":"freemium","requires_docker":true},"when_to_use":["If your project requires intelligent load balancing across various AI models including Ollama, llama.cpp, and OpenAI APIs due to VRAM limitations.","Optimal when deploying a local AI lab that benefits from the automation of environment generation tailored for GPU hardware."],"when_not_to_use":["Avoid if your setup strictly avoids Docker-based deployments or prefers a simpler proxy configuration without semantic caching capabilities.","Not recommended in scenarios where custom, non-supported models must be routed dynamically and lack VRAM-aware routing support."],"source":"enrich:decision_facts","observed_at":"2026-07-17T11:39:22.448Z"},"constraint_facets":{"min_ram_gb":4,"pricing_model":"freemium","requires_docker":true},"decision_summary":[{"label":"Pricing","value":"freemium"},{"label":"Requirements","value":"Min 4 GB RAM; Requires Docker"},{"label":"Adopt for","value":"SmarterRouter stands out with its support for semantic caching, automatic failover features, and the ability to integrate seamlessly with Docker via an automated setup wizard."}]}}