{"data":{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm","description":"A high-throughput and memory-efficient inference and serving engine for LLMs","homepage_url":"https://vllm.ai","license":"Apache-2.0","open_issues":6208,"watchers":585,"ai_summary":"vLLM is designed to provide efficient and scalable inference capabilities for large language models, supporting a variety of hardware backends including CUDA and TPU.","readme_excerpt":"## Getting Started\n\nInstall vLLM with [`uv`](https://docs.astral.sh/uv/) (recommended) or `pip`:\n\n```bash\nuv pip install vllm\n```\n\nOr [build from source](https://docs.vllm.ai/en/latest/getting_started/installation/gpu/index.html#build-wheel-from-source) for development.\n\nVisit our [documentation](https://docs.vllm.ai/en/latest/) to learn more.\n\n- [Installation](https://docs.vllm.ai/en/latest/getting_started/installation.html)\n- [Quickstart](https://docs.vllm.ai/en/latest/getting_started/quickstart.html)\n- [List of Supported Models](https://docs.vllm.ai/en/latest/models/supported_models.html)","github_created_at":"2023-02-09T11:23:20+00:00","created_at":"2026-07-07T15:09:03.849693+00:00","updated_at":"2026-08-01T12:00:15.56582+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"amd","name":"amd"},{"slug":"cuda","name":"cuda"},{"slug":"deepseek","name":"deepseek"},{"slug":"gpt","name":"gpt"},{"slug":"inference","name":"inference"},{"slug":"llama","name":"llama"},{"slug":"llm-serving","name":"llm-serving"},{"slug":"model-serving","name":"model-serving"}],"trust":{"provenance":{"is_fork":false,"github_id":599547518,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-01T12:00:14.448Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":10,"days_since_push":0,"last_release_at":"2026-07-27T01:06:58Z"},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T10:29:31.222Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-01T12:00:14.888Z"},"has_cli":{"value":true,"source":"pyproject.toml:[project.scripts]","observed_at":"2026-08-01T12:00:14.888Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-01T12:00:14.888Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-01T12:00:14.888Z"}},"decision_facts":{"hosting":null,"pricing":{"model":"freemium","summary":"vLLM operates under the Apache-2.0 license, so it's entirely free to use without direct monetary costs, but users might incur costs related to hardware and cloud services required for deployment."},"requirements":{"notes":["Installation can be done via `uv pip install vllm` or by building from source, allowing flexibility in how the tool is set up."],"min_ram_gb":null,"requires_docker":false},"constraints":{"min_ram_gb":null,"pricing_model":"freemium","requires_docker":false},"when_to_use":["When you need to deploy large language models with requirements for both high throughput and low resource consumption.","If your project involves using various hardware types like CUDA or TPU, vLLM provides versatile support that can adapt to different environments efficiently."],"when_not_to_use":["Avoid using vLLM if your application strictly limits itself to a single type of hardware without needing cross-platform compatibility, as it may introduce unnecessary complexity.","If memory efficiency is not a concern and you are optimizing for simplicity over resource management, alternatives with less configuration might be preferable."],"source":"enrich:decision_facts","observed_at":"2026-07-11T10:41:29.998Z"},"constraint_facets":{"min_ram_gb":null,"pricing_model":"freemium","requires_docker":false},"decision_summary":[{"label":"Pricing","value":"freemium - vLLM operates under the Apache-2.0 license, so it's entirely free to use without direct monetary costs, but users might incur costs related to hardware and cloud services required for deployment."},{"label":"Requirements","value":"Installation can be done via `uv pip install vllm` or by building from source, allowing flexibility in how the tool is set up."},{"label":"Adopt for","value":"vLLM is a specialized inference engine for large language models that prioritizes high throughput and memory efficiency, suitable for deployment across different hardware backends."}]}}