{"data":{"slug":"mlc-ai-mlc-llm","name":"mlc-llm","tagline":"Universal LLM Deployment Engine with ML Compilation","github_url":"https://github.com/mlc-ai/mlc-llm","owner":"mlc-ai","repo":"mlc-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/106173866?v=4","primary_language":"Python","stars":23063,"forks":2111,"topics":["language-model","llm","machine-learning-compilation","tvm"],"archived":false,"github_pushed_at":"2026-07-31T03:03:18+00:00","maintenance_label":"Active","stars_delta_30d":103,"url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm","markdown_url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mlc-ai-mlc-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mlc-ai-mlc-llm","description":"Universal LLM Deployment Engine with ML Compilation","homepage_url":"https://llm.mlc.ai/","license":"Apache-2.0","open_issues":334,"watchers":196,"ai_summary":"A tool for deploying large language models using efficient machine learning compilation techniques.","readme_excerpt":"<div align=\"center\">\n\n# MLC LLM\n\n\n\n\n\n\n**Universal LLM Deployment Engine with ML Compilation**\n\n[Get Started](https://llm.mlc.ai/docs/get_started/quick_start) | [Documentation](https://llm.mlc.ai/docs) | [Blog](https://blog.mlc.ai/)\n\n</div>\n\n## About\n\nMLC LLM is a machine learning compiler and high-performance deployment engine for large language models.  The mission of this project is to enable everyone to develop, optimize, and deploy AI models natively on everyone's platforms. \n\n<div align=\"center\">\n<table style=\"width:100%\">\n  <thead>\n    <tr>\n      <th style=\"width:15%\"> </th>\n      <th style=\"width:20%\">AMD GPU</th>\n      <th style=\"width:20%\">NVIDIA GPU</th>\n      <th style=\"width:20%\">Apple GPU</th>\n      <th style=\"width:24%\">Intel GPU</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <td>Linux / Win</td>\n      <td>✅ Vulkan, ROCm</td>\n      <td>✅ Vulkan, CUDA</td>\n      <td>N/A</td>\n      <td>✅ Vulkan</td>\n    </tr>\n    <tr>\n      <td>macOS</td>\n      <td>✅ Metal (dGPU)</td>\n      <td>N/A</td>\n      <td>✅ Metal</td>\n      <td>✅ Metal (iGPU)</td>\n    </tr>\n    <tr>\n      <td>Web Browser</td>\n      <td colspan=4>✅ WebGPU and WASM </td>\n    </tr>\n    <tr>\n      <td>iOS / iPadOS</td>\n      <td colspan=4>✅ Metal on Apple A-series GPU</td>\n    </tr>\n    <tr>\n      <td>Android</td>\n      <td colspan=2>✅ OpenCL on Adreno GPU</td>\n      <td colspan=2>✅ OpenCL on Mali GPU</td>\n    </tr>\n  </tbody>\n</table>\n</div>\n\nMLC LLM compiles and runs code on MLCEngine -- a unified high-performance LLM inference engine across the above platforms. MLCEngine provides OpenAI-compatible API available through REST server, python, javascript, iOS, Android, all backed by the same engine and compiler that we keep improving with the community.\n\n## Get Started\n\nPlease visit our [documentation](https://llm.mlc.ai/docs/) to get started with MLC LLM.\n- [Installation](https://llm.mlc.ai/docs/install/mlc_llm)\n- [Quick start](https://llm.mlc.ai/docs/get_started/quick_start)\n- [Introduction](https://llm.mlc.ai/docs/get_started/introduction)\n\n## Citation\n\nPlease consider citing our project if you find it useful:\n\n```bibtex\n@software{mlc-llm,\n    author = {{MLC team}},\n    title = {{MLC-LLM}},\n    url = {https://github.com/mlc-ai/mlc-llm},\n    year = {2023-2025}\n}\n```\n\nThe underlying techniques of MLC LLM include:\n\n<details>\n  <summary>References (Click to expand)</summary>\n\n  ```bibtex\n  @inproceedings{tensorir,\n      author = {Feng, Siyuan and Hou, Bohan and Jin, Hongyi and Lin, Wuwei and Shao, Junru and Lai, Ruihang and Ye, Zihao and Zheng, Lianmin and Yu, Cody Hao and Yu, Yong and Chen, Tianqi},\n      title = {TensorIR: An Abstraction for Automatic Tensorized Program Optimization},\n      year = {2023},\n      isbn = {9781450399166},\n      publisher = {Association for Computing Machinery},\n      address = {New York, NY, USA},\n      url = {https://doi.org/10.1145/3575693.3576933},\n      doi = {10.1145/3575693.3576933},\n      booktitle = {Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2},\n      pages = {804–817},\n      numpages = {14},\n      keywords = {Tensor Computation, Machine Learning Compiler, Deep Neural Network},\n      location = {Vancouver, BC, Canada},\n      series = {ASPLOS 2023}\n  }\n\n  @inproceedings{metaschedule,\n      author = {Shao, Junru and Zhou, Xiyou and Feng, Siyuan and Hou, Bohan and Lai, Ruihang and Jin, Hongyi and Lin, Wuwei and Masuda, Masahiro and Yu, Cody Hao and Chen, Tianqi},\n      booktitle = {Advances in Neural Information Processing Systems},\n      editor = {S. Koyejo and S. Mohamed and A. Agarwal and D. Belgrave and K. Cho and A. Oh},\n      pages = {35783--35796},\n      publisher = {Curran Associates, Inc.},\n      title = {Tensor Program Optimization with Probabilistic Programs},\n      url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/e894eafae43e68b4c8dfdacf742bcbf3-Paper-Conference.pdf},\n      volume = {35},\n      year = {2","github_created_at":"2023-04-29T01:59:25+00:00","created_at":"2026-07-07T17:32:58.503122+00:00","updated_at":"2026-08-17T00:01:42.386879+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"language-model","name":"language-model"},{"slug":"llm","name":"llm"},{"slug":"machine-learning-compilation","name":"machine-learning-compilation"},{"slug":"tvm","name":"tvm"}],"trust":{"provenance":{"is_fork":false,"github_id":634081686,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-17T00:01:41.595Z","maintenance":{"label":"Active","score":82,"methodology":"github_public_v1","releases_90d":0,"days_since_push":16,"last_release_at":"2023-04-29T03:31:41Z","stars_delta_30d":103,"open_issues_delta_30d":11},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T10:59:58.363Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-17T00:01:42.035Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-17T00:01:42.035Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-17T00:01:42.035Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":{"notes":["- Requires familiarity with Python and machine learning concepts.","- Efficient with large language models but may have higher initial setup complexity due to specialized features."]},"constraints":null,"when_to_use":["- When you need an efficient tool specifically designed with advanced compilation techniques that optimize performance for large language models (LLMs).","- If your team is working within the Python ecosystem and needs a solution that tightly integrates machine learning compilation without needing to switch languages."],"when_not_to_use":["- Avoid mlc-llm if you are looking for a broader suite of tools; this tool focuses intensely on deployment efficiency via ML compilation techniques.","- If you prefer tools with extensive third-party integrations or community-developed extensions, as mlc-llm's focus is narrow to deep optimization."],"source":"enrich:decision_facts","observed_at":"2026-07-11T13:56:23.627Z"},"constraint_facets":null,"decision_summary":[{"label":"Requirements","value":"- Requires familiarity with Python and machine learning concepts.; - Efficient with large language models but may have higher initial setup complexity due to specialized features."},{"label":"Adopt for","value":"Mature deployment engine for efficient large-scale model serving, leveraging advanced compilation techniques."},{"label":"License detail","value":"Open-source under the Apache-2.0 license, allowing for free use in both open source and commercial contexts while requiring acknowledgment of its use."}]}}