{"data":{"slug":"lemonade-sdk-lemonade","name":"lemonade","tagline":"Serves optimized LLMs locally via GPUs and NPUs","github_url":"https://github.com/lemonade-sdk/lemonade","owner":"lemonade-sdk","repo":"lemonade","owner_avatar_url":"https://avatars.githubusercontent.com/u/209808453?v=4","primary_language":"C++","stars":5448,"forks":477,"topics":["ai","amd","genai","gpu","llama","llm","llm-inference","local-server","mcp","mcp-server","mistral","npu","onnxruntime","openai-api","qwen","radeon","rocm","ryzen","vulkan"],"archived":false,"github_pushed_at":"2026-08-24T17:58:28+00:00","maintenance_label":"Very active","stars_delta_30d":340,"url":"https://www.graphcanon.com/tools/lemonade-sdk-lemonade","markdown_url":"https://www.graphcanon.com/tools/lemonade-sdk-lemonade.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lemonade-sdk-lemonade","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lemonade-sdk-lemonade","description":"Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk","homepage_url":"https://lemonade-server.ai/","license":"Apache-2.0","open_issues":521,"watchers":33,"ai_summary":"Lemonade is a toolkit that focuses on deploying and running local AI applications by serving fine-tuned LLM models directly on GPU and NPU hardware.","readme_excerpt":"## Getting Started\n\n1. **Install**: [Windows](https://github.com/lemonade-sdk/lemonade/releases/latest/download/lemonade.msi) · [Linux](#supported-platforms) · [macOS](https://github.com/lemonade-sdk/lemonade/releases) · [Docker](https://lemonade-server.ai/docs/guide/install/docker) · [Source](./docs/dev/getting-started.md)\n2. **Get Models**: Browse and download with the [Model Manager](#model-library)\n3. **Generate**: Try models with the built-in interfaces for chat, image gen, speech gen, and more\n4. **Mobile**: Take your lemonade to go: [iOS](https://apps.apple.com/us/app/lemonade-mobile/id6757372210) · [Android](https://play.google.com/store/apps/details?id=com.lemonade.mobile.chat.ai&pli=1) · [Source](https://github.com/lemonade-sdk/lemonade-mobile)\n5. **Connect**: Use Lemonade with your [favorite apps](https://lemonade-server.ai/marketplace):\n\n\n<p align=\"center\">\n  <a href=\"https://lemonade-server.ai/docs/server/apps/claude-code/\" title=\"Claude Code\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/claude-code/logo.png\" alt=\"Claude Code\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://quickthoughts.ca/posts/firefox-chatback-lemonade-sdk/\" title=\"Firefox Chatbot\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/fx-chatbot/logo.png\" alt=\"Firefox Chatbot\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://lemonade-server.ai/docs/server/apps/anythingLLM/\" title=\"AnythingLLM\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/anythingllm/logo.png\" alt=\"AnythingLLM\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://marketplace.dify.ai/plugins/langgenius/lemonade\" title=\"Dify\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/dify/logo.png\" alt=\"Dify\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://github.com/amd/gaia?tab=readme-ov-file#getting-started-guide\" title=\"GAIA\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/gaia/logo.png\" alt=\"GAIA\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://admcpr.com/local-github-copilot-with-lemonade-server-on-windows\" title=\"GitHub Copilot\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/github-copilot/logo.png\" alt=\"GitHub Copilot\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://github.com/lemonade-sdk/infinity-arcade\" title=\"Infinity Arcade\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/infinity-arcade/logo.png\" alt=\"Infinity Arcade\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://n8n.io/integrations/lemonade-model/\" title=\"n8n\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/n8n/logo.png\" alt=\"n8n\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://lemonade-server.ai/docs/server/apps/open-webui/\" title=\"Open WebUI\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/open-webui/logo.png\" alt=\"Open WebUI\" width=\"60\" /></a>&nbsp;&nbsp;<a href=\"https://lemonade-server.ai/docs/server/apps/open-hands/\" title=\"OpenHands\"><img src=\"https://raw.githubusercontent.com/lemonade-sdk/marketplace/main/apps/openhands/logo.png\" alt=\"OpenHands\" width=\"60\" /></a>\n</p>\n\n<p align=\"center\"><em>Want your app featured here? <a href=\"https://github.com/lemonade-sdk/marketplace\">Just submit a marketplace PR!</a></em></p>\n\n---\n\n## License and Attribution\n\nThis project is:\n- Built with C++ (server) and React (app) with ❤️ for the open source community,\n- Standing on the shoulders of great tools from:\n  - [ggml/llama.cpp](https://github.com/ggml-org/llama.cpp)\n  - [ggml/whisper.cpp](https://github.com/ggerganov/whisper.cpp)\n  - [ggml/stable-diffusion.cpp](https://github.com/leejet/stable-diffusion.cpp)\n  - [kokoros](https://github.com/lucasjinreal/Kokoros)\n  - [OnnxRuntime GenAI](https://github.com/microsoft/onnxruntime-genai)\n  - [Hugging Face Hub](https://github.com/huggingface/huggingface_hub)\n  - [ModelScope](https://github.com/modelscope/modelscope)\n  - [OpenAI API](htt","github_created_at":"2025-05-15T19:17:39+00:00","created_at":"2026-07-11T11:42:53.823674+00:00","updated_at":"2026-08-24T18:01:22.93804+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"ai","name":"ai"},{"slug":"amd","name":"amd"},{"slug":"genai","name":"genai"},{"slug":"gpu","name":"gpu"},{"slug":"llama","name":"llama"},{"slug":"llm-inference","name":"llm-inference"},{"slug":"local-server","name":"local-server"},{"slug":"npu","name":"npu"}],"trust":{"provenance":{"is_fork":false,"github_id":984341155,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-24T18:01:22.105Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":11,"days_since_push":0,"last_release_at":"2026-08-19T19:39:54Z","stars_delta_30d":340,"open_issues_delta_30d":74},"security_summary":{"status":"no_manifest","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:42:54.943Z","medium_count":0,"scan_profile":"mcp_manifest","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-24T18:01:22.584Z"},"deploy":{"source":"dockerfile:Dockerfile","self_host":true,"observed_at":"2026-08-24T18:01:22.584Z","managed_saas":false},"languages":{"value":["c++"],"source":"github.language","observed_at":"2026-08-24T18:01:22.584Z"},"has_docker":{"value":true,"source":"dockerfile:Dockerfile","observed_at":"2026-08-24T18:01:22.584Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-24T18:01:22.584Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["- For users looking to leverage their own GPU or NPU hardware to serve fine-tuned language models.","- In environments that require high performance but have strict data privacy constraints, as it allows local deployment."],"when_not_to_use":["- If you do not have access to a compatible GPU or NPU device for running the LLMs locally.","- For projects requiring cloud-based services and APIs over local deployment, Lemonade may introduce additional complexity in setup and maintenance compared to fully managed solutions."],"source":"enrich:decision_facts","observed_at":"2026-07-14T18:29:01.549Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Lemonade specializes in serving optimized LLMs locally with support for both GPUs and NPUs."}]}}