{"data":{"slug":"tiiny-ai-powerinfer","name":"PowerInfer","tagline":"High-speed Large Language Model Serving for Local Deployment","github_url":"https://github.com/Tiiny-AI/PowerInfer","owner":"Tiiny-AI","repo":"PowerInfer","owner_avatar_url":"https://avatars.githubusercontent.com/u/256922953?v=4","primary_language":"C++","stars":9718,"forks":591,"topics":["large-language-models","llama","llm","llm-inference","local-inference"],"archived":false,"github_pushed_at":"2026-05-11T06:48:06+00:00","maintenance_label":"Slowing","stars_delta_30d":76,"url":"https://www.graphcanon.com/tools/tiiny-ai-powerinfer","markdown_url":"https://www.graphcanon.com/tools/tiiny-ai-powerinfer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tiiny-ai-powerinfer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tiiny-ai-powerinfer","description":"High-speed Large Language Model Serving for Local Deployment","homepage_url":null,"license":"MIT","open_issues":129,"watchers":102,"ai_summary":"Tiiny-AI/PowerInfer is a C++ library aimed at providing high-speed inference services for large language models locally.","readme_excerpt":"## Getting Started\n\n- [Installation](#setup-and-installation)\n- [Model Weights](#model-weights)\n- [Inference](#inference)\n\n---\n\n# make sure that you have done `pip install -r requirements.txt`\npython convert.py --outfile /PATH/TO/POWERINFER/GGUF/REPO/MODELNAME.powerinfer.gguf /PATH/TO/ORIGINAL/MODEL /PATH/TO/PREDICTOR","github_created_at":"2023-12-15T02:24:10+00:00","created_at":"2026-07-07T17:34:08.22811+00:00","updated_at":"2026-08-17T06:02:02.305072+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"large-language-models","name":"large language models"},{"slug":"llama","name":"llama"},{"slug":"llm","name":"llm"},{"slug":"llm-inference","name":"llm-inference"},{"slug":"local-inference","name":"local-inference"}],"trust":{"provenance":{"is_fork":false,"github_id":731842419,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-17T06:02:01.577Z","maintenance":{"label":"Slowing","score":36,"methodology":"github_public_v1","releases_90d":0,"days_since_push":97,"last_release_at":null,"stars_delta_30d":76,"open_issues_delta_30d":0},"security_summary":{"status":"findings","scanner":"osv@v1","low_count":46,"high_count":0,"last_scan_at":"2026-07-11T11:02:31.958Z","medium_count":0,"scan_profile":"deps","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-17T06:02:02.033Z"},"languages":{"value":["c++"],"source":"github.language","observed_at":"2026-08-17T06:02:02.033Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-17T06:02:02.033Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["- If your deployment requires local handling of large language model inference with high-speed performance, PowerInfer excels in offering this capability using the C++ environment.","For projects valuing MIT license compatibility to support open-source or commercial use without restrictions on derivative works, PowerInfer's permissive license is a significant advantage."],"when_not_to_use":["- Consider alternatives if you prefer frameworks with more extensive Python support, as the setup and conversion scripts in PowerInfer primarily use Python to prepare models despite it being a C++-dr븐"],"source":"enrich:decision_facts","observed_at":"2026-07-11T16:02:26.542Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"PowerInfer is a C++ library designed for high-speed inference of large language models locally."}]}}