{"data":{"node":{"slug":"predibase-lorax","name":"lorax","tagline":"Multi-LoRA inference server for scalable fine-tuned LLMs","github_url":"https://github.com/predibase/lorax","owner":"predibase","repo":"lorax","owner_avatar_url":"https://avatars.githubusercontent.com/u/75280641?v=4","primary_language":"Python","stars":3826,"forks":326,"topics":["fine-tuning","gpt","llama","llm","llm-inference","llm-serving","llmops","lora","model-serving","pytorch","transformers"],"archived":false,"github_pushed_at":"2026-05-28T18:12:20+00:00","maintenance_label":"Steady","stars_delta_30d":10,"url":"https://www.graphcanon.com/tools/predibase-lorax","markdown_url":"https://www.graphcanon.com/tools/predibase-lorax.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/predibase-lorax","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=predibase-lorax"},"categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"fine-tuning","name":"fine-tuning"},{"slug":"gpt","name":"gpt"},{"slug":"llama","name":"llama"},{"slug":"llm-inference","name":"llm-inference"},{"slug":"llm-serving","name":"llm-serving"},{"slug":"pytorch","name":"pytorch"},{"slug":"transformers","name":"transformers"}],"edges":[{"type":"depends_on","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"integrates_with","direction":"out","explanation":"LoRAX serves fine-tuned models using LoRA (Low-Rank Adaptation), which is facilitated through the HuggingFace transformers library. LoRAX dynamically loads adapters that are often available in the HuggingFace model hub.","successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"alternative","direction":"out","explanation":"Both SGLang and LoRAX are serving frameworks designed for large language models, differing in their approach to handling dynamic model loadings and integration with various LLM adapters.","successor_context":null,"tool":{"slug":"sgl-project-sglang","name":"sglang","tagline":"High-performance serving framework for large language and multimodal models","github_url":"https://github.com/sgl-project/sglang","owner":"sgl-project","repo":"sglang","owner_avatar_url":"https://avatars.githubusercontent.com/u/147780389?v=4","primary_language":"Python","stars":31454,"forks":7720,"topics":["attention","blackwell","cuda","deepseek","diffusion","glm","gpt-oss","inference","llama","llm","minimax","moe","qwen","qwen-image","reinforcement-learning","transformer","vlm","wan"],"archived":false,"github_pushed_at":"2026-08-07T06:00:20+00:00","maintenance_label":"Very active","stars_delta_30d":1409,"url":"https://www.graphcanon.com/tools/sgl-project-sglang","markdown_url":"https://www.graphcanon.com/tools/sgl-project-sglang.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sgl-project-sglang","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sgl-project-sglang"}},{"type":"alternative","direction":"out","explanation":"MLC-LLM also provides a solution for deploying large language models, focusing on the ML compilation for universal deployment. LoRAX focuses more on dynamic serving of fine-tuned models using the LoRA technique.","successor_context":null,"tool":{"slug":"mlc-ai-mlc-llm","name":"mlc-llm","tagline":"Universal LLM Deployment Engine with ML Compilation","github_url":"https://github.com/mlc-ai/mlc-llm","owner":"mlc-ai","repo":"mlc-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/106173866?v=4","primary_language":"Python","stars":23063,"forks":2111,"topics":["language-model","llm","machine-learning-compilation","tvm"],"archived":false,"github_pushed_at":"2026-07-31T03:03:18+00:00","maintenance_label":"Active","stars_delta_30d":103,"url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm","markdown_url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mlc-ai-mlc-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mlc-ai-mlc-llm"}},{"type":"alternative","direction":"out","explanation":"Both vLLM and LoRAX aim to provide efficient LLM serving solutions. While vLLM focuses on ease of use and cost-effectiveness, LoRAX is optimized for dynamic adapter loading that scales up to thousands of fine-tuned models.","successor_context":null,"tool":{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm"}},{"type":"integrates_with","direction":"out","explanation":"Lorax has an 'integrates with' relationship to LlamaFactory because Lorax serves as a multi-LoRA inference server that supports scaling fine-tuned models, which can include the myriad of large language models (LLMs) and vision-language models (VLMs) produced by LlamaFactory. This integration allows for efficient deployment and scaling of models fine-tuned using LlamaFactory's repository.","successor_context":null,"tool":{"slug":"hiyouga-llamafactory","name":"LlamaFactory","tagline":"Unified Efficient Fine-Tuning of 100+ LLMs & VLMs","github_url":"https://github.com/hiyouga/LlamaFactory","owner":"hiyouga","repo":"LlamaFactory","owner_avatar_url":"https://avatars.githubusercontent.com/u/16256802?v=4","primary_language":"Python","stars":74132,"forks":9071,"topics":["agent","ai","deepseek","fine-tuning","gemma","gpt","instruction-tuning","large-language-models","llama","llama3","llm","lora","moe","nlp","peft","qlora","quantization","qwen","rlhf","transformers"],"archived":false,"github_pushed_at":"2026-08-13T12:45:56+00:00","maintenance_label":"Very active","stars_delta_30d":803,"url":"https://www.graphcanon.com/tools/hiyouga-llamafactory","markdown_url":"https://www.graphcanon.com/tools/hiyouga-llamafactory.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/hiyouga-llamafactory","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=hiyouga-llamafactory"}}],"neighbours":[{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm","shared_categories":["inference-serving"]},{"slug":"tloen-alpaca-lora","name":"alpaca-lora","tagline":"Instruct-tune LLaMA on consumer hardware","github_url":"https://github.com/tloen/alpaca-lora","owner":"tloen","repo":"alpaca-lora","owner_avatar_url":"https://avatars.githubusercontent.com/u/4811103?v=4","primary_language":"Jupyter Notebook","stars":18912,"forks":2180,"topics":[],"archived":false,"github_pushed_at":"2024-07-29T13:37:49+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/tloen-alpaca-lora","markdown_url":"https://www.graphcanon.com/tools/tloen-alpaca-lora.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tloen-alpaca-lora","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tloen-alpaca-lora","shared_categories":["inference-serving"]},{"slug":"jundot-omlx","name":"omlx","tagline":"LLM inference server with continuous batching and SSD caching for Apple Silicon","github_url":"https://github.com/jundot/omlx","owner":"jundot","repo":"omlx","owner_avatar_url":"https://avatars.githubusercontent.com/u/64250138?v=4","primary_language":"Python","stars":18679,"forks":1617,"topics":["apple-silicon","inference-server","llm","macos","mlx","openai-api"],"archived":false,"github_pushed_at":"2026-08-14T09:12:48+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/jundot-omlx","markdown_url":"https://www.graphcanon.com/tools/jundot-omlx.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/jundot-omlx","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=jundot-omlx","shared_categories":["inference-serving"]},{"slug":"lightning-ai-litgpt","name":"litgpt","tagline":"High-performance LLMs with recipes for pretraining, finetuning and deployment","github_url":"https://github.com/Lightning-AI/litgpt","owner":"Lightning-AI","repo":"litgpt","owner_avatar_url":"https://avatars.githubusercontent.com/u/58386951?v=4","primary_language":"Python","stars":13605,"forks":1483,"topics":["ai","artificial-intelligence","deep-learning","large-language-models","llm","llm-inference","llms"],"archived":false,"github_pushed_at":"2026-07-20T10:24:12+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/lightning-ai-litgpt","markdown_url":"https://www.graphcanon.com/tools/lightning-ai-litgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lightning-ai-litgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lightning-ai-litgpt","shared_categories":["inference-serving"]},{"slug":"artidoro-qlora","name":"qlora","tagline":"QLoRA finetuning of quantized LLMs","github_url":"https://github.com/artidoro/qlora","owner":"artidoro","repo":"qlora","owner_avatar_url":"https://avatars.githubusercontent.com/u/11949572?v=4","primary_language":"Jupyter Notebook","stars":10979,"forks":876,"topics":[],"archived":false,"github_pushed_at":"2024-06-10T19:20:16+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/artidoro-qlora","markdown_url":"https://www.graphcanon.com/tools/artidoro-qlora.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/artidoro-qlora","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=artidoro-qlora","shared_categories":[]},{"slug":"fminference-flexllmgen","name":"FlexLLMGen","tagline":"Running large language models on a single GPU for throughput-oriented scenarios.","github_url":"https://github.com/FMInference/FlexLLMGen","owner":"FMInference","repo":"FlexLLMGen","owner_avatar_url":"https://avatars.githubusercontent.com/u/125944572?v=4","primary_language":"Python","stars":9361,"forks":590,"topics":["deep-learning","gpt-3","high-throughput","large-language-models","machine-learning","offloading","opt"],"archived":true,"github_pushed_at":"2024-10-28T03:05:41+00:00","maintenance_label":"Archived","url":"https://www.graphcanon.com/tools/fminference-flexllmgen","markdown_url":"https://www.graphcanon.com/tools/fminference-flexllmgen.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/fminference-flexllmgen","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=fminference-flexllmgen","shared_categories":["inference-serving"]},{"slug":"optimalscale-lmflow","name":"LMFlow","tagline":"An Extensible Toolkit for Finetuning and Inference of Large Foundation Models","github_url":"https://github.com/OptimalScale/LMFlow","owner":"OptimalScale","repo":"LMFlow","owner_avatar_url":"https://avatars.githubusercontent.com/u/128913633?v=4","primary_language":"Python","stars":8486,"forks":825,"topics":["chatgpt","deep-learning","instruction-following","language-model","pretrained-models","pytorch","transformer"],"archived":false,"github_pushed_at":"2026-05-22T02:57:26+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/optimalscale-lmflow","markdown_url":"https://www.graphcanon.com/tools/optimalscale-lmflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/optimalscale-lmflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=optimalscale-lmflow","shared_categories":["inference-serving"]},{"slug":"ai-dynamo-dynamo","name":"dynamo","tagline":"A Datacenter Scale Distributed Inference Serving Framework","github_url":"https://github.com/ai-dynamo/dynamo","owner":"ai-dynamo","repo":"dynamo","owner_avatar_url":"https://avatars.githubusercontent.com/u/201626793?v=4","primary_language":"Rust","stars":7575,"forks":1368,"topics":["diffusion","disaggregated-serving","kubernetes","llm-inference","omni","routing-engine","rust","sglang","tensorrt-llm","vllm"],"archived":false,"github_pushed_at":"2026-07-25T05:43:22+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/ai-dynamo-dynamo","markdown_url":"https://www.graphcanon.com/tools/ai-dynamo-dynamo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ai-dynamo-dynamo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ai-dynamo-dynamo","shared_categories":["inference-serving"]},{"slug":"cloneofsimo-lora","name":"lora","tagline":"Jupyter Notebook repository for fine-tuning diffusion models using Low-Rank Adaptation.","github_url":"https://github.com/cloneofsimo/lora","owner":"cloneofsimo","repo":"lora","owner_avatar_url":"https://avatars.githubusercontent.com/u/35953539?v=4","primary_language":"Jupyter Notebook","stars":7545,"forks":496,"topics":["diffusion","dreambooth","fine-tuning","lora","stable-diffusion"],"archived":false,"github_pushed_at":"2024-03-22T03:48:10+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/cloneofsimo-lora","markdown_url":"https://www.graphcanon.com/tools/cloneofsimo-lora.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/cloneofsimo-lora","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=cloneofsimo-lora","shared_categories":[]},{"slug":"flashinfer-ai-flashinfer","name":"flashinfer","tagline":"FlashInfer is a kernel library for serving large language models","github_url":"https://github.com/flashinfer-ai/flashinfer","owner":"flashinfer-ai","repo":"flashinfer","owner_avatar_url":"https://avatars.githubusercontent.com/u/145061914?v=4","primary_language":"Python","stars":6024,"forks":1196,"topics":["attention","cuda","distributed-inference","gpu","jit","large-large-models","llm-inference","moe","nvidia","pytorch"],"archived":false,"github_pushed_at":"2026-07-25T05:01:59+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/flashinfer-ai-flashinfer","markdown_url":"https://www.graphcanon.com/tools/flashinfer-ai-flashinfer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/flashinfer-ai-flashinfer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=flashinfer-ai-flashinfer","shared_categories":["inference-serving"]},{"slug":"turboderp-exllama","name":"exllama","tagline":"Memory-efficient rewrite of HF transformers for Llama with quantized weights","github_url":"https://github.com/turboderp/exllama","owner":"turboderp","repo":"exllama","owner_avatar_url":"https://avatars.githubusercontent.com/u/11859846?v=4","primary_language":"Python","stars":2937,"forks":220,"topics":[],"archived":false,"github_pushed_at":"2023-09-30T19:06:04+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/turboderp-exllama","markdown_url":"https://www.graphcanon.com/tools/turboderp-exllama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/turboderp-exllama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=turboderp-exllama","shared_categories":["inference-serving"]},{"slug":"michaelfeil-infinity","name":"infinity","tagline":"High-throughput, low-latency serving engine for text-embeddings and various models","github_url":"https://github.com/michaelfeil/infinity","owner":"michaelfeil","repo":"infinity","owner_avatar_url":"https://avatars.githubusercontent.com/u/63565275?v=4","primary_language":"Python","stars":2907,"forks":196,"topics":["bert-embeddings","llm","text-embeddings"],"archived":false,"github_pushed_at":"2026-03-24T03:59:47+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/michaelfeil-infinity","markdown_url":"https://www.graphcanon.com/tools/michaelfeil-infinity.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/michaelfeil-infinity","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=michaelfeil-infinity","shared_categories":["inference-serving"]}]}}