{"data":{"node":{"slug":"bigscience-workshop-petals","name":"petals","tagline":"Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading","github_url":"https://github.com/bigscience-workshop/petals","owner":"bigscience-workshop","repo":"petals","owner_avatar_url":"https://avatars.githubusercontent.com/u/82455566?v=4","primary_language":"Python","stars":10496,"forks":642,"topics":["bloom","chatbot","deep-learning","distributed-systems","falcon","gpt","guanaco","language-models","large-language-models","llama","machine-learning","mixtral","neural-networks","nlp","pipeline-parallelism","pretrained-models","pytorch","tensor-parallelism","transformer","volunteer-computing"],"archived":false,"github_pushed_at":"2024-09-07T11:54:28+00:00","maintenance_label":"Dormant","stars_delta_30d":212,"url":"https://www.graphcanon.com/tools/bigscience-workshop-petals","markdown_url":"https://www.graphcanon.com/tools/bigscience-workshop-petals.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/bigscience-workshop-petals","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=bigscience-workshop-petals"},"categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"bloom","name":"bloom"},{"slug":"chatbot","name":"chatbot"},{"slug":"deep-learning","name":"deep-learning"},{"slug":"distributed-systems","name":"distributed-systems"},{"slug":"falcon","name":"falcon"},{"slug":"gpt","name":"gpt"},{"slug":"guanaco","name":"guanaco"},{"slug":"language-models","name":"language-models"}],"edges":[{"type":"alternative","direction":"out","explanation":"Both Petals and Ollama provide ways to run LLMs locally with optimizations, but they do so using different approaches and infrastructure setups.","successor_context":null,"tool":{"slug":"ollama-ollama","name":"ollama","tagline":"Get up and running with various large language models using Ollama.","github_url":"https://github.com/ollama/ollama","owner":"ollama","repo":"ollama","owner_avatar_url":"https://avatars.githubusercontent.com/u/151674099?v=4","primary_language":"Go","stars":177524,"forks":17229,"topics":["deepseek","gemma","gemma3","glm","go","golang","gpt-oss","llama","llama3","llm","llms","minimax","mistral","ollama","qwen"],"archived":false,"github_pushed_at":"2026-07-31T23:59:29+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/ollama-ollama","markdown_url":"https://www.graphcanon.com/tools/ollama-ollama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ollama-ollama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ollama-ollama"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"unslothai-unsloth","name":"unsloth","tagline":"A web UI for training and running open models locally.","github_url":"https://github.com/unslothai/unsloth","owner":"unslothai","repo":"unsloth","owner_avatar_url":"https://avatars.githubusercontent.com/u/150920049?v=4","primary_language":"Python","stars":69621,"forks":6285,"topics":["agent","deepseek","fine-tuning","gemma","gemma3","gpt-oss","llama","llama3","llm","llms","mistral","openai","qwen","reinforcement-learning","self-hosted","text-to-speech","tts","ui","unsloth"],"archived":false,"github_pushed_at":"2026-08-06T06:01:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/unslothai-unsloth","markdown_url":"https://www.graphcanon.com/tools/unslothai-unsloth.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unslothai-unsloth","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unslothai-unsloth"}},{"type":"depends_on","direction":"out","explanation":"Petals uses transformers from huggingface for pre-trained models.","successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"integrates_with","direction":"out","explanation":"Petals uses Hugging Face's `transformers` library to interact with various large language models and integrates its distributed computing features on top of it.","successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"integrates_with","direction":"out","explanation":"`SGLang` serves as a language model serving framework, which can potentially integrate with Petals for serving large models using distributed methods.","successor_context":null,"tool":{"slug":"sgl-project-sglang","name":"sglang","tagline":"High-performance serving framework for large language and multimodal models","github_url":"https://github.com/sgl-project/sglang","owner":"sgl-project","repo":"sglang","owner_avatar_url":"https://avatars.githubusercontent.com/u/147780389?v=4","primary_language":"Python","stars":31454,"forks":7720,"topics":["attention","blackwell","cuda","deepseek","diffusion","glm","gpt-oss","inference","llama","llm","minimax","moe","qwen","qwen-image","reinforcement-learning","transformer","vlm","wan"],"archived":false,"github_pushed_at":"2026-08-07T06:00:20+00:00","maintenance_label":"Very active","stars_delta_30d":1409,"url":"https://www.graphcanon.com/tools/sgl-project-sglang","markdown_url":"https://www.graphcanon.com/tools/sgl-project-sglang.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sgl-project-sglang","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sgl-project-sglang"}},{"type":"alternative","direction":"out","explanation":"VL LM serves as an alternative to Petals for LLM serving with a focus on ease and speed of deployment.","successor_context":null,"tool":{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm"}},{"type":"related","direction":"out","explanation":"Both work on LLM inference but ggml focuses on low-level C/C++ implementation whereas petals provides a higher level distributed execution framework.","successor_context":null,"tool":{"slug":"ggml-org-llama-cpp","name":"llama.cpp","tagline":"LLM inference in C/C++","github_url":"https://github.com/ggml-org/llama.cpp","owner":"ggml-org","repo":"llama.cpp","owner_avatar_url":"https://avatars.githubusercontent.com/u/134263123?v=4","primary_language":"C++","stars":122941,"forks":21406,"topics":["ggml"],"archived":false,"github_pushed_at":"2026-08-07T05:28:54+00:00","maintenance_label":"Very active","stars_delta_30d":3353,"url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp","markdown_url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ggml-org-llama-cpp","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ggml-org-llama-cpp"}}],"neighbours":[{"slug":"nomic-ai-gpt4all","name":"gpt4all","tagline":"Run Local LLMs on Any Device","github_url":"https://github.com/nomic-ai/gpt4all","owner":"nomic-ai","repo":"gpt4all","owner_avatar_url":"https://avatars.githubusercontent.com/u/102670180?v=4","primary_language":"C++","stars":77396,"forks":8304,"topics":["ai-chat","llm-inference"],"archived":false,"github_pushed_at":"2025-05-27T20:05:19+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/nomic-ai-gpt4all","markdown_url":"https://www.graphcanon.com/tools/nomic-ai-gpt4all.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nomic-ai-gpt4all","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nomic-ai-gpt4all","shared_categories":["llm-frameworks","inference-serving"]},{"slug":"alexsjones-llmfit","name":"llmfit","tagline":"Hundreds of models & providers. One command to find what runs on your hardware.","github_url":"https://github.com/AlexsJones/llmfit","owner":"AlexsJones","repo":"llmfit","owner_avatar_url":"https://avatars.githubusercontent.com/u/1235925?v=4","primary_language":"Rust","stars":31867,"forks":1978,"topics":["gguf","llm","localai","mlx","skill","unsloth"],"archived":false,"github_pushed_at":"2026-08-14T07:36:41+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/alexsjones-llmfit","markdown_url":"https://www.graphcanon.com/tools/alexsjones-llmfit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alexsjones-llmfit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alexsjones-llmfit","shared_categories":["llm-frameworks"]},{"slug":"lightning-ai-pytorch-lightning","name":"pytorch-lightning","tagline":"Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.","github_url":"https://github.com/Lightning-AI/pytorch-lightning","owner":"Lightning-AI","repo":"pytorch-lightning","owner_avatar_url":"https://avatars.githubusercontent.com/u/58386951?v=4","primary_language":"Python","stars":31267,"forks":3768,"topics":["ai","artificial-intelligence","data-science","deep-learning","machine-learning","python","pytorch"],"archived":false,"github_pushed_at":"2026-08-03T01:13:09+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/lightning-ai-pytorch-lightning","markdown_url":"https://www.graphcanon.com/tools/lightning-ai-pytorch-lightning.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lightning-ai-pytorch-lightning","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lightning-ai-pytorch-lightning","shared_categories":["inference-serving"]},{"slug":"lightning-ai-litgpt","name":"litgpt","tagline":"High-performance LLMs with recipes for pretraining, finetuning and deployment","github_url":"https://github.com/Lightning-AI/litgpt","owner":"Lightning-AI","repo":"litgpt","owner_avatar_url":"https://avatars.githubusercontent.com/u/58386951?v=4","primary_language":"Python","stars":13605,"forks":1483,"topics":["ai","artificial-intelligence","deep-learning","large-language-models","llm","llm-inference","llms"],"archived":false,"github_pushed_at":"2026-07-20T10:24:12+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/lightning-ai-litgpt","markdown_url":"https://www.graphcanon.com/tools/lightning-ai-litgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lightning-ai-litgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lightning-ai-litgpt","shared_categories":["llm-frameworks","inference-serving"]},{"slug":"fminference-flexllmgen","name":"FlexLLMGen","tagline":"Running large language models on a single GPU for throughput-oriented scenarios.","github_url":"https://github.com/FMInference/FlexLLMGen","owner":"FMInference","repo":"FlexLLMGen","owner_avatar_url":"https://avatars.githubusercontent.com/u/125944572?v=4","primary_language":"Python","stars":9361,"forks":590,"topics":["deep-learning","gpt-3","high-throughput","large-language-models","machine-learning","offloading","opt"],"archived":true,"github_pushed_at":"2024-10-28T03:05:41+00:00","maintenance_label":"Archived","url":"https://www.graphcanon.com/tools/fminference-flexllmgen","markdown_url":"https://www.graphcanon.com/tools/fminference-flexllmgen.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/fminference-flexllmgen","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=fminference-flexllmgen","shared_categories":["inference-serving"]},{"slug":"fareedkhan-dev-train-llm-from-scratch","name":"train-llm-from-scratch","tagline":"A straightforward method for training your LLM from raw text to aligned model generation","github_url":"https://github.com/FareedKhan-dev/train-llm-from-scratch","owner":"FareedKhan-dev","repo":"train-llm-from-scratch","owner_avatar_url":"https://avatars.githubusercontent.com/u/63067900?v=4","primary_language":"Python","stars":9141,"forks":1264,"topics":["gemini","large-language-models","llm","openai","training","transformers"],"archived":false,"github_pushed_at":"2026-08-17T05:07:26+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/fareedkhan-dev-train-llm-from-scratch","markdown_url":"https://www.graphcanon.com/tools/fareedkhan-dev-train-llm-from-scratch.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/fareedkhan-dev-train-llm-from-scratch","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=fareedkhan-dev-train-llm-from-scratch","shared_categories":["inference-serving"]},{"slug":"bitsandbytes-foundation-bitsandbytes","name":"bitsandbytes","tagline":"Large language model quantization toolkit for PyTorch.","github_url":"https://github.com/bitsandbytes-foundation/bitsandbytes","owner":"bitsandbytes-foundation","repo":"bitsandbytes","owner_avatar_url":"https://avatars.githubusercontent.com/u/175231607?v=4","primary_language":"Python","stars":8385,"forks":900,"topics":["llm","machine-learning","pytorch","qlora","quantization"],"archived":false,"github_pushed_at":"2026-07-29T18:27:51+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/bitsandbytes-foundation-bitsandbytes","markdown_url":"https://www.graphcanon.com/tools/bitsandbytes-foundation-bitsandbytes.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/bitsandbytes-foundation-bitsandbytes","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=bitsandbytes-foundation-bitsandbytes","shared_categories":["llm-frameworks","inference-serving"]},{"slug":"b4rtaz-distributed-llama","name":"distributed-llama","tagline":"Distributed LLM inference using home devices cluster","github_url":"https://github.com/b4rtaz/distributed-llama","owner":"b4rtaz","repo":"distributed-llama","owner_avatar_url":"https://avatars.githubusercontent.com/u/12797776?v=4","primary_language":"C++","stars":3012,"forks":242,"topics":["distributed-computing","distributed-llm","llama2","llama3","llm","llm-inference","llms","neural-network","open-llm"],"archived":false,"github_pushed_at":"2026-07-05T16:47:20+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/b4rtaz-distributed-llama","markdown_url":"https://www.graphcanon.com/tools/b4rtaz-distributed-llama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/b4rtaz-distributed-llama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=b4rtaz-distributed-llama","shared_categories":["inference-serving"]},{"slug":"huggingface-nanotron","name":"nanotron","tagline":"Minimalistic large language model 3D-parallelism training","github_url":"https://github.com/huggingface/nanotron","owner":"huggingface","repo":"nanotron","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":2775,"forks":329,"topics":[],"archived":false,"github_pushed_at":"2026-05-26T10:32:37+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/huggingface-nanotron","markdown_url":"https://www.graphcanon.com/tools/huggingface-nanotron.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-nanotron","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-nanotron","shared_categories":[]},{"slug":"microsoft-pai","name":"pai","tagline":"Resource scheduling and cluster management for AI","github_url":"https://github.com/microsoft/pai","owner":"microsoft","repo":"pai","owner_avatar_url":"https://avatars.githubusercontent.com/u/6154722?v=4","primary_language":"JavaScript","stars":2686,"forks":549,"topics":["ai","artificial-intelligence","chainer","cloud","cluster-management","cluster-manager","gpu","gpu-cluster","gpu-computing","gpu-scheduler","jupyter","kubernetes","machine-learning","model-training","on-premise","pytorch","resource-management","scheduling","tensorflow"],"archived":true,"github_pushed_at":"2024-06-06T07:56:07+00:00","maintenance_label":"Archived","url":"https://www.graphcanon.com/tools/microsoft-pai","markdown_url":"https://www.graphcanon.com/tools/microsoft-pai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/microsoft-pai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=microsoft-pai","shared_categories":["inference-serving"]},{"slug":"rafska-awesome-local-llm","name":"awesome-local-llm","tagline":"Resources for running LLMs locally","github_url":"https://github.com/rafska/awesome-local-llm","owner":"rafska","repo":"awesome-local-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/17859377?v=4","primary_language":null,"stars":2518,"forks":316,"topics":["ai","awesome","awesome-list","llm","local","local-ai","local-llm","resources","self-hosted","selfhosted"],"archived":false,"github_pushed_at":"2026-08-04T23:08:15+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/rafska-awesome-local-llm","markdown_url":"https://www.graphcanon.com/tools/rafska-awesome-local-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/rafska-awesome-local-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=rafska-awesome-local-llm","shared_categories":["inference-serving"]},{"slug":"dstackai-dstack","name":"dstack","tagline":"Vendor-agnostic orchestration for AI workloads","github_url":"https://github.com/dstackai/dstack","owner":"dstackai","repo":"dstack","owner_avatar_url":"https://avatars.githubusercontent.com/u/54146142?v=4","primary_language":"Python","stars":2192,"forks":240,"topics":["agent-skills","agentic-orchestration","amd","cloud","containers","docker","fine-tuning","gpu","inference","k8s","kubernetes","llms","machine-learning","nvidia","orchestration","python","slurm","training"],"archived":false,"github_pushed_at":"2026-07-24T11:15:18+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/dstackai-dstack","markdown_url":"https://www.graphcanon.com/tools/dstackai-dstack.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/dstackai-dstack","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=dstackai-dstack","shared_categories":["inference-serving"]}]}}