{"data":{"node":{"slug":"mlc-ai-mlc-llm","name":"mlc-llm","tagline":"Universal LLM Deployment Engine with ML Compilation","github_url":"https://github.com/mlc-ai/mlc-llm","owner":"mlc-ai","repo":"mlc-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/106173866?v=4","primary_language":"Python","stars":23063,"forks":2111,"topics":["language-model","llm","machine-learning-compilation","tvm"],"archived":false,"github_pushed_at":"2026-07-31T03:03:18+00:00","maintenance_label":"Active","stars_delta_30d":103,"url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm","markdown_url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mlc-ai-mlc-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mlc-ai-mlc-llm"},"categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"language-model","name":"language-model"},{"slug":"llm","name":"llm"},{"slug":"machine-learning-compilation","name":"machine-learning-compilation"},{"slug":"tvm","name":"tvm"}],"edges":[{"type":"integrates_with","direction":"out","explanation":"mlc-llm integrates with llmfit because llmfit evaluates the compatibility of large language models with different hardware, providing sizing and optimization insights, which mlc-llm can then use to optimize and deploy these models efficiently across various platforms.","successor_context":null,"tool":{"slug":"alexsjones-llmfit","name":"llmfit","tagline":"Hundreds of models & providers. One command to find what runs on your hardware.","github_url":"https://github.com/AlexsJones/llmfit","owner":"AlexsJones","repo":"llmfit","owner_avatar_url":"https://avatars.githubusercontent.com/u/1235925?v=4","primary_language":"Rust","stars":31867,"forks":1978,"topics":["gguf","llm","localai","mlx","skill","unsloth"],"archived":false,"github_pushed_at":"2026-08-14T07:36:41+00:00","maintenance_label":"Very active","stars_delta_30d":2339,"url":"https://www.graphcanon.com/tools/alexsjones-llmfit","markdown_url":"https://www.graphcanon.com/tools/alexsjones-llmfit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alexsjones-llmfit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alexsjones-llmfit"}},{"type":"integrates_with","direction":"out","explanation":"MLC LLM can compile and optimize models from transformers for deployment across various hardware platforms.","successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"alternative","direction":"out","explanation":"Both MLC-LLM and vllm serve the purpose of efficiently deploying large language models with a focus on performance optimization across various hardware platforms.","successor_context":null,"tool":{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm"}},{"type":"integrates_with","direction":"out","explanation":"mlc-llm integrates with FastGPT through its capability to optimize and deploy large language models, which can be utilized by FastGPT for enhancing its data processing, RAG retrieval, and visual AI workflows. MLC LLM's role as a deployment engine complements FastGPT's aim of offering an all-inclusive platform for developing complex question-answering systems.","successor_context":null,"tool":{"slug":"labring-fastgpt","name":"FastGPT","tagline":"A knowledge-based platform built on LLMs for developing and deploying complex question-answering systems","github_url":"https://github.com/labring/FastGPT","owner":"labring","repo":"FastGPT","owner_avatar_url":"https://avatars.githubusercontent.com/u/102226726?v=4","primary_language":"TypeScript","stars":29366,"forks":7264,"topics":["agent","claude","deepseek","llm","mcp","nextjs","openai","qwen","rag","workflow"],"archived":false,"github_pushed_at":"2026-08-16T15:38:02+00:00","maintenance_label":"Very active","stars_delta_30d":360,"url":"https://www.graphcanon.com/tools/labring-fastgpt","markdown_url":"https://www.graphcanon.com/tools/labring-fastgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/labring-fastgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=labring-fastgpt"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"microsoft-autogen","name":"autogen","tagline":"A programming framework for agentic AI","github_url":"https://github.com/microsoft/autogen","owner":"microsoft","repo":"autogen","owner_avatar_url":"https://avatars.githubusercontent.com/u/6154722?v=4","primary_language":"Python","stars":60139,"forks":9059,"topics":["agentic","agentic-agi","agents","ai","autogen","autogen-ecosystem","chatgpt","framework","llm-agent","llm-framework"],"archived":false,"github_pushed_at":"2026-04-15T11:59:09+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/microsoft-autogen","markdown_url":"https://www.graphcanon.com/tools/microsoft-autogen.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/microsoft-autogen","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=microsoft-autogen"}},{"type":"alternative","direction":"out","explanation":"Both MLC-LLM and llama.cpp are focused on LLM inference but with different hardware support and optimizations.","successor_context":null,"tool":{"slug":"ggml-org-llama-cpp","name":"llama.cpp","tagline":"LLM inference in C/C++","github_url":"https://github.com/ggml-org/llama.cpp","owner":"ggml-org","repo":"llama.cpp","owner_avatar_url":"https://avatars.githubusercontent.com/u/134263123?v=4","primary_language":"C++","stars":122941,"forks":21406,"topics":["ggml"],"archived":false,"github_pushed_at":"2026-08-07T05:28:54+00:00","maintenance_label":"Very active","stars_delta_30d":3353,"url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp","markdown_url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ggml-org-llama-cpp","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ggml-org-llama-cpp"}},{"type":"alternative","direction":"in","explanation":"SGLang and mlc-LLM both aim at deploying large language models efficiently across different hardware setups. They differ in their underlying technologies and deployment strategies, making them alternatives for model serving.","successor_context":null,"tool":{"slug":"sgl-project-sglang","name":"sglang","tagline":"High-performance serving framework for large language and multimodal models","github_url":"https://github.com/sgl-project/sglang","owner":"sgl-project","repo":"sglang","owner_avatar_url":"https://avatars.githubusercontent.com/u/147780389?v=4","primary_language":"Python","stars":31454,"forks":7720,"topics":["attention","blackwell","cuda","deepseek","diffusion","glm","gpt-oss","inference","llama","llm","minimax","moe","qwen","qwen-image","reinforcement-learning","transformer","vlm","wan"],"archived":false,"github_pushed_at":"2026-08-07T06:00:20+00:00","maintenance_label":"Very active","stars_delta_30d":1409,"url":"https://www.graphcanon.com/tools/sgl-project-sglang","markdown_url":"https://www.graphcanon.com/tools/sgl-project-sglang.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sgl-project-sglang","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sgl-project-sglang"}},{"type":"integrates_with","direction":"in","explanation":"SGLang could integrate with ML Compilation to facilitate deployment and optimization of various LLMs across different hardware platforms.","successor_context":null,"tool":{"slug":"sgl-project-sglang","name":"sglang","tagline":"High-performance serving framework for large language and multimodal models","github_url":"https://github.com/sgl-project/sglang","owner":"sgl-project","repo":"sglang","owner_avatar_url":"https://avatars.githubusercontent.com/u/147780389?v=4","primary_language":"Python","stars":31454,"forks":7720,"topics":["attention","blackwell","cuda","deepseek","diffusion","glm","gpt-oss","inference","llama","llm","minimax","moe","qwen","qwen-image","reinforcement-learning","transformer","vlm","wan"],"archived":false,"github_pushed_at":"2026-08-07T06:00:20+00:00","maintenance_label":"Very active","stars_delta_30d":1409,"url":"https://www.graphcanon.com/tools/sgl-project-sglang","markdown_url":"https://www.graphcanon.com/tools/sgl-project-sglang.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sgl-project-sglang","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sgl-project-sglang"}},{"type":"integrates_with","direction":"in","explanation":"PowerInfer could potentially integrate with MLC LLM as it provides an engine for deploying language models across various devices and contexts, which complements PowerInfer's deployment capabilities.","successor_context":null,"tool":{"slug":"tiiny-ai-powerinfer","name":"PowerInfer","tagline":"High-speed Large Language Model Serving for Local Deployment","github_url":"https://github.com/Tiiny-AI/PowerInfer","owner":"Tiiny-AI","repo":"PowerInfer","owner_avatar_url":"https://avatars.githubusercontent.com/u/256922953?v=4","primary_language":"C++","stars":9718,"forks":591,"topics":["large-language-models","llama","llm","llm-inference","local-inference"],"archived":false,"github_pushed_at":"2026-05-11T06:48:06+00:00","maintenance_label":"Slowing","stars_delta_30d":76,"url":"https://www.graphcanon.com/tools/tiiny-ai-powerinfer","markdown_url":"https://www.graphcanon.com/tools/tiiny-ai-powerinfer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tiiny-ai-powerinfer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tiiny-ai-powerinfer"}},{"type":"alternative","direction":"in","explanation":"MLC-LLM also provides a solution for deploying large language models, focusing on the ML compilation for universal deployment. LoRAX focuses more on dynamic serving of fine-tuned models using the LoRA technique.","successor_context":null,"tool":{"slug":"predibase-lorax","name":"lorax","tagline":"Multi-LoRA inference server for scalable fine-tuned LLMs","github_url":"https://github.com/predibase/lorax","owner":"predibase","repo":"lorax","owner_avatar_url":"https://avatars.githubusercontent.com/u/75280641?v=4","primary_language":"Python","stars":3826,"forks":326,"topics":["fine-tuning","gpt","llama","llm","llm-inference","llm-serving","llmops","lora","model-serving","pytorch","transformers"],"archived":false,"github_pushed_at":"2026-05-28T18:12:20+00:00","maintenance_label":"Steady","stars_delta_30d":10,"url":"https://www.graphcanon.com/tools/predibase-lorax","markdown_url":"https://www.graphcanon.com/tools/predibase-lorax.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/predibase-lorax","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=predibase-lorax"}},{"type":"related","direction":"in","explanation":"RTP-LLM and mlc-llm both deal with large language model deployment but approach the problem differently, RTP-LLM focusing on inference acceleration while mlc-llm focuses on universal deployment via ML compilation.","successor_context":null,"tool":{"slug":"alibaba-rtp-llm","name":"rtp-llm","tagline":"Alibaba's high-performance LLM inference engine for diverse applications.","github_url":"https://github.com/alibaba/rtp-llm","owner":"alibaba","repo":"rtp-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/1961952?v=4","primary_language":"Cuda","stars":1312,"forks":260,"topics":["gpt","inference","llama","llm","llm-serving","llmops","model-serving"],"archived":false,"github_pushed_at":"2026-08-20T17:20:55+00:00","maintenance_label":"Very active","stars_delta_30d":30,"url":"https://www.graphcanon.com/tools/alibaba-rtp-llm","markdown_url":"https://www.graphcanon.com/tools/alibaba-rtp-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alibaba-rtp-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alibaba-rtp-llm"}}],"neighbours":[{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm","shared_categories":["inference-serving"]},{"slug":"mlflow-mlflow","name":"mlflow","tagline":"AI engineering platform for debugging, evaluating, monitoring, and optimizing AI applications","github_url":"https://github.com/mlflow/mlflow","owner":"mlflow","repo":"mlflow","owner_avatar_url":"https://avatars.githubusercontent.com/u/39938107?v=4","primary_language":"Python","stars":27591,"forks":6189,"topics":["agentops","agents","ai","ai-governance","apache-spark","evaluation","langchain","llm-evaluation","llmops","machine-learning","ml","mlflow","mlops","model-management","observability","open-source","openai","prompt-engineering"],"archived":false,"github_pushed_at":"2026-08-20T00:54:28+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/mlflow-mlflow","markdown_url":"https://www.graphcanon.com/tools/mlflow-mlflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mlflow-mlflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mlflow-mlflow","shared_categories":["inference-serving"]},{"slug":"lyogavin-airllm","name":"airllm","tagline":"AirLLM 70B inference with single 4GB GPU","github_url":"https://github.com/lyogavin/airllm","owner":"lyogavin","repo":"airllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/1113905?v=4","primary_language":"Jupyter Notebook","stars":24183,"forks":2722,"topics":["chinese-llm","chinese-nlp","finetune","generative-ai","instruct-gpt","instruction-set","llama","llm","lora","open-models","open-source","open-source-models","qlora"],"archived":false,"github_pushed_at":"2026-07-23T08:29:43+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/lyogavin-airllm","markdown_url":"https://www.graphcanon.com/tools/lyogavin-airllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lyogavin-airllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lyogavin-airllm","shared_categories":["inference-serving"]},{"slug":"stas00-ml-engineering","name":"ml-engineering","tagline":"Machine Learning Engineering Open Book","github_url":"https://github.com/stas00/ml-engineering","owner":"stas00","repo":"ml-engineering","owner_avatar_url":"https://avatars.githubusercontent.com/u/10676103?v=4","primary_language":"Python","stars":18632,"forks":1200,"topics":["ai","debugging","gpus","inference","large-language-models","llm","machine-learning","machine-learning-engineering","mlops","network","pytorch","scalability","slurm","storage","training","transformers"],"archived":false,"github_pushed_at":"2026-08-14T19:59:44+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/stas00-ml-engineering","markdown_url":"https://www.graphcanon.com/tools/stas00-ml-engineering.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/stas00-ml-engineering","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=stas00-ml-engineering","shared_categories":["inference-serving"]},{"slug":"nvidia-tensorrt-llm","name":"TensorRT-LLM","tagline":"Python API for defining and optimizing Large Language Models (LLMs) on NVIDIA GPUs","github_url":"https://github.com/NVIDIA/TensorRT-LLM","owner":"NVIDIA","repo":"TensorRT-LLM","owner_avatar_url":"https://avatars.githubusercontent.com/u/1728152?v=4","primary_language":"Python","stars":14317,"forks":2641,"topics":["blackwell","cuda","llm-serving","moe","pytorch"],"archived":false,"github_pushed_at":"2026-08-07T05:40:26+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/nvidia-tensorrt-llm","markdown_url":"https://www.graphcanon.com/tools/nvidia-tensorrt-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nvidia-tensorrt-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nvidia-tensorrt-llm","shared_categories":["llm-frameworks","inference-serving"]},{"slug":"apache-tvm","name":"tvm","tagline":"Open Machine Learning Compiler Framework","github_url":"https://github.com/apache/tvm","owner":"apache","repo":"tvm","owner_avatar_url":"https://avatars.githubusercontent.com/u/47359?v=4","primary_language":"Python","stars":13642,"forks":3939,"topics":["compiler","deep-learning","gpu","javascript","machine-learning","metal","opencl","performance","rocm","spirv","tensor","tvm","vulkan"],"archived":false,"github_pushed_at":"2026-08-03T01:32:46+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/apache-tvm","markdown_url":"https://www.graphcanon.com/tools/apache-tvm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/apache-tvm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=apache-tvm","shared_categories":["llm-frameworks","inference-serving"]},{"slug":"lightning-ai-litgpt","name":"litgpt","tagline":"High-performance LLMs with recipes for pretraining, finetuning and deployment","github_url":"https://github.com/Lightning-AI/litgpt","owner":"Lightning-AI","repo":"litgpt","owner_avatar_url":"https://avatars.githubusercontent.com/u/58386951?v=4","primary_language":"Python","stars":13605,"forks":1483,"topics":["ai","artificial-intelligence","deep-learning","large-language-models","llm","llm-inference","llms"],"archived":false,"github_pushed_at":"2026-07-20T10:24:12+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/lightning-ai-litgpt","markdown_url":"https://www.graphcanon.com/tools/lightning-ai-litgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lightning-ai-litgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lightning-ai-litgpt","shared_categories":["llm-frameworks","inference-serving"]},{"slug":"bentoml-openllm","name":"OpenLLM","tagline":"Run any open-source LLMs as OpenAI compatible API endpoint in the cloud.","github_url":"https://github.com/bentoml/OpenLLM","owner":"bentoml","repo":"OpenLLM","owner_avatar_url":"https://avatars.githubusercontent.com/u/49176046?v=4","primary_language":"Python","stars":12454,"forks":828,"topics":["bentoml","fine-tuning","llama","llama2","llama3-1","llama3-2","llama3-2-vision","llm","llm-inference","llm-ops","llm-serving","llmops","mistral","mlops","model-inference","open-source-llm","openllm","vicuna"],"archived":false,"github_pushed_at":"2026-08-03T16:59:03+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/bentoml-openllm","markdown_url":"https://www.graphcanon.com/tools/bentoml-openllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/bentoml-openllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=bentoml-openllm","shared_categories":["inference-serving"]},{"slug":"kalyanks-nlp-llm-engineer-toolkit","name":"llm-engineer-toolkit","tagline":"A curated list of over 120 LLM libraries categorized.","github_url":"https://github.com/KalyanKS-NLP/llm-engineer-toolkit","owner":"KalyanKS-NLP","repo":"llm-engineer-toolkit","owner_avatar_url":"https://avatars.githubusercontent.com/u/202506543?v=4","primary_language":null,"stars":10767,"forks":1682,"topics":["ai-engineer","generative-ai","large-language-models","llm-engineer","llms"],"archived":false,"github_pushed_at":"2026-08-16T13:05:43+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/kalyanks-nlp-llm-engineer-toolkit","markdown_url":"https://www.graphcanon.com/tools/kalyanks-nlp-llm-engineer-toolkit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/kalyanks-nlp-llm-engineer-toolkit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=kalyanks-nlp-llm-engineer-toolkit","shared_categories":["inference-serving"]},{"slug":"internlm-lmdeploy","name":"lmdeploy","tagline":"Toolkit for compressing, deploying, and serving LLMs","github_url":"https://github.com/InternLM/lmdeploy","owner":"InternLM","repo":"lmdeploy","owner_avatar_url":"https://avatars.githubusercontent.com/u/135356492?v=4","primary_language":"Python","stars":7995,"forks":723,"topics":["codellama","cuda-kernels","deepspeed","fastertransformer","internlm","llama","llama2","llama3","llm","llm-inference","turbomind"],"archived":false,"github_pushed_at":"2026-08-06T09:17:57+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/internlm-lmdeploy","markdown_url":"https://www.graphcanon.com/tools/internlm-lmdeploy.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/internlm-lmdeploy","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=internlm-lmdeploy","shared_categories":["inference-serving"]},{"slug":"ericlbuehler-mistral-rs","name":"mistral.rs","tagline":"Fast flexible LLM inference","github_url":"https://github.com/EricLBuehler/mistral.rs","owner":"EricLBuehler","repo":"mistral.rs","owner_avatar_url":"https://avatars.githubusercontent.com/u/65165915?v=4","primary_language":"Rust","stars":7575,"forks":671,"topics":["llm","rust","uqff"],"archived":false,"github_pushed_at":"2026-07-29T20:21:17+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/ericlbuehler-mistral-rs","markdown_url":"https://www.graphcanon.com/tools/ericlbuehler-mistral-rs.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ericlbuehler-mistral-rs","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ericlbuehler-mistral-rs","shared_categories":["inference-serving"]},{"slug":"eth-sri-lmql","name":"lmql","tagline":"A language for constraint-guided and efficient LLM programming.","github_url":"https://github.com/eth-sri/lmql","owner":"eth-sri","repo":"lmql","owner_avatar_url":"https://avatars.githubusercontent.com/u/5363413?v=4","primary_language":"Python","stars":4203,"forks":221,"topics":["chatgpt","huggingface","language-model","programming-language"],"archived":false,"github_pushed_at":"2025-05-22T07:32:31+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/eth-sri-lmql","markdown_url":"https://www.graphcanon.com/tools/eth-sri-lmql.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/eth-sri-lmql","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=eth-sri-lmql","shared_categories":["llm-frameworks"]}]}}