{"data":{"node":{"slug":"fla-org-flash-linear-attention","name":"flash-linear-attention","tagline":"🚀 Efficient implementations for emerging model architectures","github_url":"https://github.com/fla-org/flash-linear-attention","owner":"fla-org","repo":"flash-linear-attention","owner_avatar_url":"https://avatars.githubusercontent.com/u/40835596?v=4","primary_language":"Python","stars":5568,"forks":661,"topics":["large-language-models","machine-learning-systems","natural-language-processing","sequence-modeling"],"archived":false,"github_pushed_at":"2026-08-17T10:13:08+00:00","maintenance_label":"Very active","stars_delta_30d":208,"url":"https://www.graphcanon.com/tools/fla-org-flash-linear-attention","markdown_url":"https://www.graphcanon.com/tools/fla-org-flash-linear-attention.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/fla-org-flash-linear-attention","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=fla-org-flash-linear-attention"},"categories":[{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"large-language-models","name":"large language models"},{"slug":"machine-learning-systems","name":"machine-learning-systems"},{"slug":"natural-language-processing","name":"natural-language-processing"},{"slug":"sequence-modeling","name":"sequence-modeling"}],"edges":[{"type":"related","direction":"out","explanation":"Both projects involve working with LLMs for specific applications, Flash Linear Attention for efficient implementations and TradingAgents for financial trading frameworks.","successor_context":null,"tool":{"slug":"tauricresearch-tradingagents","name":"TradingAgents","tagline":"Multi-Agents LLM Financial Trading Framework","github_url":"https://github.com/TauricResearch/TradingAgents","owner":"TauricResearch","repo":"TradingAgents","owner_avatar_url":"https://avatars.githubusercontent.com/u/192884433?v=4","primary_language":"Python","stars":98335,"forks":18953,"topics":["agent","finance","llm","multiagent","trading"],"archived":false,"github_pushed_at":"2026-07-18T15:55:05+00:00","maintenance_label":"Active","stars_delta_30d":5040,"url":"https://www.graphcanon.com/tools/tauricresearch-tradingagents","markdown_url":"https://www.graphcanon.com/tools/tauricresearch-tradingagents.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tauricresearch-tradingagents","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tauricresearch-tradingagents"}},{"type":"related","direction":"out","explanation":"Both projects are about optimizing the use of AI models, but LLMFit offers sizing recommendations while Flash Linear Attention focuses on efficient implementation and hardware efficiency.","successor_context":null,"tool":{"slug":"alexsjones-llmfit","name":"llmfit","tagline":"Hundreds of models & providers. One command to find what runs on your hardware.","github_url":"https://github.com/AlexsJones/llmfit","owner":"AlexsJones","repo":"llmfit","owner_avatar_url":"https://avatars.githubusercontent.com/u/1235925?v=4","primary_language":"Rust","stars":31867,"forks":1978,"topics":["gguf","llm","localai","mlx","skill","unsloth"],"archived":false,"github_pushed_at":"2026-08-14T07:36:41+00:00","maintenance_label":"Very active","stars_delta_30d":2339,"url":"https://www.graphcanon.com/tools/alexsjones-llmfit","markdown_url":"https://www.graphcanon.com/tools/alexsjones-llmfit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alexsjones-llmfit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alexsjones-llmfit"}},{"type":"related","direction":"out","explanation":"Both projects aim to make AI models operational without the need for high-end hardware, though Flash Linear Attention specifically focuses on efficient linear attention mechanisms.","successor_context":null,"tool":{"slug":"mudler-localai","name":"LocalAI","tagline":"Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.","github_url":"https://github.com/mudler/LocalAI","owner":"mudler","repo":"LocalAI","owner_avatar_url":"https://avatars.githubusercontent.com/u/2420543?v=4","primary_language":"Go","stars":48500,"forks":4362,"topics":["agents","ai","api","audio-generation","decentralized","distributed","image-generation","libp2p","llama","llm","mamba","mcp","musicgen","object-detection","rerank","stable-diffusion","text-generation","tts"],"archived":false,"github_pushed_at":"2026-08-16T05:07:25+00:00","maintenance_label":"Very active","stars_delta_30d":924,"url":"https://www.graphcanon.com/tools/mudler-localai","markdown_url":"https://www.graphcanon.com/tools/mudler-localai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mudler-localai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mudler-localai"}},{"type":"integrates_with","direction":"out","explanation":"Flash Linear Attention can be employed in Haystack to optimize performance of sequence models used within the orchestration framework for LLM applications.","successor_context":null,"tool":{"slug":"deepset-ai-haystack","name":"haystack","tagline":"Open-source AI orchestration framework for building context-engineered LLM applications.","github_url":"https://github.com/deepset-ai/haystack","owner":"deepset-ai","repo":"haystack","owner_avatar_url":"https://avatars.githubusercontent.com/u/51827949?v=4","primary_language":"Python","stars":26073,"forks":2972,"topics":["agent","agents","ai","gemini","generative-ai","gpt-4","information-retrieval","large-language-models","llm","machine-learning","nlp","orchestration","python","pytorch","question-answering","rag","retrieval-augmented-generation","semantic-search","summarization","transformers"],"archived":false,"github_pushed_at":"2026-08-01T03:06:32+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/deepset-ai-haystack","markdown_url":"https://www.graphcanon.com/tools/deepset-ai-haystack.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/deepset-ai-haystack","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=deepset-ai-haystack"}},{"type":"integrates_with","direction":"out","explanation":"Flash Linear Attention provides efficient implementations of linear attention mechanisms and other components for modern sequence models, which can be integrated into transformers to improve their performance or reduce computational requirements. Transformers, as a comprehensive library, supports various model architectures and can incorporate advanced attention modules like those from FlashLinear","successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"related","direction":"out","explanation":"flash-linear-attention and AirLLM are related in that both aim for efficient large language model inference, with flash-linear-attention focusing more on the underlying attention mechanisms while AirLLM seeks to optimize for lightweight GPU usage.","successor_context":null,"tool":{"slug":"lyogavin-airllm","name":"airllm","tagline":"AirLLM 70B inference with single 4GB GPU","github_url":"https://github.com/lyogavin/airllm","owner":"lyogavin","repo":"airllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/1113905?v=4","primary_language":"Jupyter Notebook","stars":24183,"forks":2722,"topics":["chinese-llm","chinese-nlp","finetune","generative-ai","instruct-gpt","instruction-set","llama","llm","lora","open-models","open-source","open-source-models","qlora"],"archived":false,"github_pushed_at":"2026-07-23T08:29:43+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/lyogavin-airllm","markdown_url":"https://www.graphcanon.com/tools/lyogavin-airllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lyogavin-airllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lyogavin-airllm"}},{"type":"related","direction":"out","explanation":"Flash Linear Attention focuses on efficient implementations for modern sequence models, whereas llms-from-scratch is about building LLMs from scratch using PyTorch.","successor_context":null,"tool":{"slug":"rasbt-llms-from-scratch","name":"LLMs-from-scratch","tagline":"Implement a ChatGPT-like LLM in PyTorch from scratch, step by step","github_url":"https://github.com/rasbt/LLMs-from-scratch","owner":"rasbt","repo":"LLMs-from-scratch","owner_avatar_url":"https://avatars.githubusercontent.com/u/5618407?v=4","primary_language":"Jupyter Notebook","stars":102733,"forks":15748,"topics":["ai","artificial-intelligence","attention-mechanism","deep-learning","finetuning","from-scratch","generative-ai","gpt","instruction-tuning","language-model","large-language-models","llm","machine-learning","natural-language-processing","pretraining","python","pytorch","tokenizer","transformers"],"archived":false,"github_pushed_at":"2026-08-10T01:11:40+00:00","maintenance_label":"Very active","stars_delta_30d":3541,"url":"https://www.graphcanon.com/tools/rasbt-llms-from-scratch","markdown_url":"https://www.graphcanon.com/tools/rasbt-llms-from-scratch.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/rasbt-llms-from-scratch","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=rasbt-llms-from-scratch"}},{"type":"related","direction":"out","explanation":"Both flash-linear-attention and Lightning AI LIT-GPT are related in the space of large language model (LLM) optimization and efficient architectures, but they do not directly integrate or succeed each other.","successor_context":null,"tool":{"slug":"lightning-ai-litgpt","name":"litgpt","tagline":"High-performance LLMs with recipes for pretraining, finetuning and deployment","github_url":"https://github.com/Lightning-AI/litgpt","owner":"Lightning-AI","repo":"litgpt","owner_avatar_url":"https://avatars.githubusercontent.com/u/58386951?v=4","primary_language":"Python","stars":13605,"forks":1483,"topics":["ai","artificial-intelligence","deep-learning","large-language-models","llm","llm-inference","llms"],"archived":false,"github_pushed_at":"2026-07-20T10:24:12+00:00","maintenance_label":"Active","stars_delta_30d":137,"url":"https://www.graphcanon.com/tools/lightning-ai-litgpt","markdown_url":"https://www.graphcanon.com/tools/lightning-ai-litgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lightning-ai-litgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lightning-ai-litgpt"}},{"type":"depends_on","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"microsoft-autogen","name":"autogen","tagline":"A programming framework for agentic AI","github_url":"https://github.com/microsoft/autogen","owner":"microsoft","repo":"autogen","owner_avatar_url":"https://avatars.githubusercontent.com/u/6154722?v=4","primary_language":"Python","stars":60139,"forks":9059,"topics":["agentic","agentic-agi","agents","ai","autogen","autogen-ecosystem","chatgpt","framework","llm-agent","llm-framework"],"archived":false,"github_pushed_at":"2026-04-15T11:59:09+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/microsoft-autogen","markdown_url":"https://www.graphcanon.com/tools/microsoft-autogen.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/microsoft-autogen","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=microsoft-autogen"}},{"type":"related","direction":"out","explanation":"Both projects deal with large language models but from different angles; Flash Linear Attention is about efficient implementations, while LlamaFactory focuses on unified fine-tuning.","successor_context":null,"tool":{"slug":"hiyouga-llamafactory","name":"LlamaFactory","tagline":"Unified Efficient Fine-Tuning of 100+ LLMs & VLMs","github_url":"https://github.com/hiyouga/LlamaFactory","owner":"hiyouga","repo":"LlamaFactory","owner_avatar_url":"https://avatars.githubusercontent.com/u/16256802?v=4","primary_language":"Python","stars":74132,"forks":9071,"topics":["agent","ai","deepseek","fine-tuning","gemma","gpt","instruction-tuning","large-language-models","llama","llama3","llm","lora","moe","nlp","peft","qlora","quantization","qwen","rlhf","transformers"],"archived":false,"github_pushed_at":"2026-08-13T12:45:56+00:00","maintenance_label":"Very active","stars_delta_30d":803,"url":"https://www.graphcanon.com/tools/hiyouga-llamafactory","markdown_url":"https://www.graphcanon.com/tools/hiyouga-llamafactory.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/hiyouga-llamafactory","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=hiyouga-llamafactory"}},{"type":"depends_on","direction":"out","explanation":"flash-linear-attention likely depends on Megatron-LM for GPU optimization during the training of large-scale transformer models.","successor_context":null,"tool":{"slug":"nvidia-megatron-lm","name":"Megatron-LM","tagline":"Ongoing research training transformer models at scale","github_url":"https://github.com/NVIDIA/Megatron-LM","owner":"NVIDIA","repo":"Megatron-LM","owner_avatar_url":"https://avatars.githubusercontent.com/u/1728152?v=4","primary_language":"Python","stars":17341,"forks":4333,"topics":["large-language-models","model-para","transformers"],"archived":false,"github_pushed_at":"2026-08-06T23:12:52+00:00","maintenance_label":"Very active","stars_delta_30d":353,"url":"https://www.graphcanon.com/tools/nvidia-megatron-lm","markdown_url":"https://www.graphcanon.com/tools/nvidia-megatron-lm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nvidia-megatron-lm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nvidia-megatron-lm"}}],"neighbours":[{"slug":"alexsjones-llmfit","name":"llmfit","tagline":"Hundreds of models & providers. One command to find what runs on your hardware.","github_url":"https://github.com/AlexsJones/llmfit","owner":"AlexsJones","repo":"llmfit","owner_avatar_url":"https://avatars.githubusercontent.com/u/1235925?v=4","primary_language":"Rust","stars":31867,"forks":1978,"topics":["gguf","llm","localai","mlx","skill","unsloth"],"archived":false,"github_pushed_at":"2026-08-14T07:36:41+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/alexsjones-llmfit","markdown_url":"https://www.graphcanon.com/tools/alexsjones-llmfit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alexsjones-llmfit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alexsjones-llmfit","shared_categories":["model-training"]},{"slug":"lightning-ai-litgpt","name":"litgpt","tagline":"High-performance LLMs with recipes for pretraining, finetuning and deployment","github_url":"https://github.com/Lightning-AI/litgpt","owner":"Lightning-AI","repo":"litgpt","owner_avatar_url":"https://avatars.githubusercontent.com/u/58386951?v=4","primary_language":"Python","stars":13605,"forks":1483,"topics":["ai","artificial-intelligence","deep-learning","large-language-models","llm","llm-inference","llms"],"archived":false,"github_pushed_at":"2026-07-20T10:24:12+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/lightning-ai-litgpt","markdown_url":"https://www.graphcanon.com/tools/lightning-ai-litgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lightning-ai-litgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lightning-ai-litgpt","shared_categories":["model-training"]},{"slug":"eugeneyan-open-llms","name":"open-llms","tagline":"A list of open LLMs available for commercial use.","github_url":"https://github.com/eugeneyan/open-llms","owner":"eugeneyan","repo":"open-llms","owner_avatar_url":"https://avatars.githubusercontent.com/u/6831355?v=4","primary_language":null,"stars":12849,"forks":985,"topics":["commercial","large-language-models","llm","llms"],"archived":false,"github_pushed_at":"2025-02-13T06:37:12+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/eugeneyan-open-llms","markdown_url":"https://www.graphcanon.com/tools/eugeneyan-open-llms.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/eugeneyan-open-llms","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=eugeneyan-open-llms","shared_categories":[]},{"slug":"fareedkhan-dev-train-llm-from-scratch","name":"train-llm-from-scratch","tagline":"A straightforward method for training your LLM from raw text to aligned model generation","github_url":"https://github.com/FareedKhan-dev/train-llm-from-scratch","owner":"FareedKhan-dev","repo":"train-llm-from-scratch","owner_avatar_url":"https://avatars.githubusercontent.com/u/63067900?v=4","primary_language":"Python","stars":9141,"forks":1264,"topics":["gemini","large-language-models","llm","openai","training","transformers"],"archived":false,"github_pushed_at":"2026-08-17T05:07:26+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/fareedkhan-dev-train-llm-from-scratch","markdown_url":"https://www.graphcanon.com/tools/fareedkhan-dev-train-llm-from-scratch.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/fareedkhan-dev-train-llm-from-scratch","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=fareedkhan-dev-train-llm-from-scratch","shared_categories":["model-training"]},{"slug":"bitsandbytes-foundation-bitsandbytes","name":"bitsandbytes","tagline":"Large language model quantization toolkit for PyTorch.","github_url":"https://github.com/bitsandbytes-foundation/bitsandbytes","owner":"bitsandbytes-foundation","repo":"bitsandbytes","owner_avatar_url":"https://avatars.githubusercontent.com/u/175231607?v=4","primary_language":"Python","stars":8385,"forks":900,"topics":["llm","machine-learning","pytorch","qlora","quantization"],"archived":false,"github_pushed_at":"2026-07-29T18:27:51+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/bitsandbytes-foundation-bitsandbytes","markdown_url":"https://www.graphcanon.com/tools/bitsandbytes-foundation-bitsandbytes.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/bitsandbytes-foundation-bitsandbytes","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=bitsandbytes-foundation-bitsandbytes","shared_categories":[]},{"slug":"eleutherai-gpt-neox","name":"gpt-neox","tagline":"Implementation of model parallel autoregressive transformers on GPUs based on Megatron and DeepSpeed libraries","github_url":"https://github.com/EleutherAI/gpt-neox","owner":"EleutherAI","repo":"gpt-neox","owner_avatar_url":"https://avatars.githubusercontent.com/u/68924597?v=4","primary_language":"Python","stars":7452,"forks":1119,"topics":["deepspeed-library","gpt-3","language-model","transformers"],"archived":false,"github_pushed_at":"2026-06-11T19:25:44+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/eleutherai-gpt-neox","markdown_url":"https://www.graphcanon.com/tools/eleutherai-gpt-neox.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/eleutherai-gpt-neox","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=eleutherai-gpt-neox","shared_categories":["model-training"]},{"slug":"linkedin-liger-kernel","name":"Liger-Kernel","tagline":"Efficient Triton Kernels for LLM Training","github_url":"https://github.com/linkedin/Liger-Kernel","owner":"linkedin","repo":"Liger-Kernel","owner_avatar_url":"https://avatars.githubusercontent.com/u/357098?v=4","primary_language":"Python","stars":6555,"forks":573,"topics":["finetuning","gemma2","hacktoberfest","llama","llama3","llm-training","llms","mistral","phi3","triton","triton-kernels"],"archived":false,"github_pushed_at":"2026-08-07T08:48:09+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/linkedin-liger-kernel","markdown_url":"https://www.graphcanon.com/tools/linkedin-liger-kernel.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/linkedin-liger-kernel","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=linkedin-liger-kernel","shared_categories":["model-training"]},{"slug":"nvidia-fastertransformer","name":"FasterTransformer","tagline":"Transformer related optimization including BERT and GPT","github_url":"https://github.com/NVIDIA/FasterTransformer","owner":"NVIDIA","repo":"FasterTransformer","owner_avatar_url":"https://avatars.githubusercontent.com/u/1728152?v=4","primary_language":"C++","stars":6446,"forks":935,"topics":["bert","gpt","pytorch","transformer"],"archived":false,"github_pushed_at":"2024-03-27T11:25:30+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/nvidia-fastertransformer","markdown_url":"https://www.graphcanon.com/tools/nvidia-fastertransformer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nvidia-fastertransformer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nvidia-fastertransformer","shared_categories":[]},{"slug":"flashinfer-ai-flashinfer","name":"flashinfer","tagline":"FlashInfer is a kernel library for serving large language models","github_url":"https://github.com/flashinfer-ai/flashinfer","owner":"flashinfer-ai","repo":"flashinfer","owner_avatar_url":"https://avatars.githubusercontent.com/u/145061914?v=4","primary_language":"Python","stars":6024,"forks":1196,"topics":["attention","cuda","distributed-inference","gpu","jit","large-large-models","llm-inference","moe","nvidia","pytorch"],"archived":false,"github_pushed_at":"2026-07-25T05:01:59+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/flashinfer-ai-flashinfer","markdown_url":"https://www.graphcanon.com/tools/flashinfer-ai-flashinfer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/flashinfer-ai-flashinfer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=flashinfer-ai-flashinfer","shared_categories":[]},{"slug":"xlite-dev-awesome-llm-inference","name":"Awesome-LLM-Inference","tagline":"A curated list of LLM/VLM inference papers with codes","github_url":"https://github.com/xlite-dev/Awesome-LLM-Inference","owner":"xlite-dev","repo":"Awesome-LLM-Inference","owner_avatar_url":"https://avatars.githubusercontent.com/u/204302598?v=4","primary_language":"Python","stars":5415,"forks":428,"topics":["awesome-llm","deepseek","deepseek-r1","deepseek-v3","flash-attention","flash-attention-3","flash-mla","llm-inference","minimax-01","mla","paged-attention","qwen3","tensorrt-llm","vllm"],"archived":false,"github_pushed_at":"2026-06-23T03:48:43+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/xlite-dev-awesome-llm-inference","markdown_url":"https://www.graphcanon.com/tools/xlite-dev-awesome-llm-inference.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/xlite-dev-awesome-llm-inference","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=xlite-dev-awesome-llm-inference","shared_categories":[]},{"slug":"minimax-ai-minimax-m1","name":"MiniMax-M1","tagline":"Open-weight large-scale hybrid-attention reasoning model","github_url":"https://github.com/MiniMax-AI/MiniMax-M1","owner":"MiniMax-AI","repo":"MiniMax-M1","owner_avatar_url":"https://avatars.githubusercontent.com/u/194880281?v=4","primary_language":"Python","stars":3172,"forks":283,"topics":["large-language-models","llm","minimax-m1","reasoning-models"],"archived":false,"github_pushed_at":"2025-07-07T11:57:22+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/minimax-ai-minimax-m1","markdown_url":"https://www.graphcanon.com/tools/minimax-ai-minimax-m1.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/minimax-ai-minimax-m1","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=minimax-ai-minimax-m1","shared_categories":[]},{"slug":"b4rtaz-distributed-llama","name":"distributed-llama","tagline":"Distributed LLM inference using home devices cluster","github_url":"https://github.com/b4rtaz/distributed-llama","owner":"b4rtaz","repo":"distributed-llama","owner_avatar_url":"https://avatars.githubusercontent.com/u/12797776?v=4","primary_language":"C++","stars":3012,"forks":242,"topics":["distributed-computing","distributed-llm","llama2","llama3","llm","llm-inference","llms","neural-network","open-llm"],"archived":false,"github_pushed_at":"2026-07-05T16:47:20+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/b4rtaz-distributed-llama","markdown_url":"https://www.graphcanon.com/tools/b4rtaz-distributed-llama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/b4rtaz-distributed-llama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=b4rtaz-distributed-llama","shared_categories":[]}]}}