{"data":{"node":{"slug":"vibrantlabsai-ragas","name":"ragas","tagline":"Supercharge Your LLM Application Evaluations 🚀","github_url":"https://github.com/vibrantlabsai/ragas","owner":"vibrantlabsai","repo":"ragas","owner_avatar_url":"https://avatars.githubusercontent.com/u/122604797?v=4","primary_language":"Python","stars":15388,"forks":1637,"topics":["evaluation","llm","llmops"],"archived":false,"github_pushed_at":"2026-02-24T07:47:19+00:00","maintenance_label":"Slowing","stars_delta_30d":470,"url":"https://www.graphcanon.com/tools/vibrantlabsai-ragas","markdown_url":"https://www.graphcanon.com/tools/vibrantlabsai-ragas.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vibrantlabsai-ragas","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vibrantlabsai-ragas"},"categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"evaluation","name":"evaluation"},{"slug":"llm","name":"llm"},{"slug":"llmops","name":"llmops"}],"edges":[{"type":"related","direction":"out","explanation":"`Ragas` focuses on evaluations of LLM applications, while `Giskard-OSS` covers evals and test generation for agentic systems; both serve different aspects in the evaluation process.","successor_context":null,"tool":{"slug":"giskard-ai-giskard-oss","name":"giskard-oss","tagline":"Open-Source Evaluation & Testing library for LLM Agents","github_url":"https://github.com/Giskard-AI/giskard-oss","owner":"Giskard-AI","repo":"giskard-oss","owner_avatar_url":"https://avatars.githubusercontent.com/u/71782571?v=4","primary_language":"Python","stars":5727,"forks":511,"topics":["agent-evaluation","ai-red-team","ai-security","ai-testing","fairness-ai","llm","llm-eval","llm-evaluation","llm-security","llmops","ml-testing","ml-validation","mlops","rag-evaluation","red-team-tools","responsible-ai","trustworthy-ai"],"archived":false,"github_pushed_at":"2026-08-01T23:22:37+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/giskard-ai-giskard-oss","markdown_url":"https://www.graphcanon.com/tools/giskard-ai-giskard-oss.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/giskard-ai-giskard-oss","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=giskard-ai-giskard-oss"}},{"type":"integrates_with","direction":"out","explanation":"Given that both tools focus on AI observability, monitoring and evaluation with likely overlapping use cases (especially within agent frameworks), it is reasonable to assume they could integrate well for comprehensive LLM application evaluations.","successor_context":null,"tool":{"slug":"raga-ai-hub-ragaai-catalyst","name":"RagaAI-Catalyst","tagline":"Python SDK for AI agent observability and evaluation","github_url":"https://github.com/raga-ai-hub/RagaAI-Catalyst","owner":"raga-ai-hub","repo":"RagaAI-Catalyst","owner_avatar_url":"https://avatars.githubusercontent.com/u/161833182?v=4","primary_language":"Python","stars":16148,"forks":3565,"topics":["agentic-ai","agentic-ai-development","agentneo","agents","ai-agent-monitoring","ai-application-debugging","ai-evaluation-tools","ai-performance-optimization","ai-tool-interaction-monitoring","llm-testing","llm-tracing","llmops"],"archived":false,"github_pushed_at":"2026-02-11T14:43:33+00:00","maintenance_label":"Slowing","stars_delta_30d":5,"url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst","markdown_url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/raga-ai-hub-ragaai-catalyst","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=raga-ai-hub-ragaai-catalyst"}},{"type":"related","direction":"out","explanation":"Both `Ragas` and `RagaAI-Catalyst` deal with AI observability, monitoring, and evaluation frameworks for LLM applications.","successor_context":null,"tool":{"slug":"raga-ai-hub-ragaai-catalyst","name":"RagaAI-Catalyst","tagline":"Python SDK for AI agent observability and evaluation","github_url":"https://github.com/raga-ai-hub/RagaAI-Catalyst","owner":"raga-ai-hub","repo":"RagaAI-Catalyst","owner_avatar_url":"https://avatars.githubusercontent.com/u/161833182?v=4","primary_language":"Python","stars":16148,"forks":3565,"topics":["agentic-ai","agentic-ai-development","agentneo","agents","ai-agent-monitoring","ai-application-debugging","ai-evaluation-tools","ai-performance-optimization","ai-tool-interaction-monitoring","llm-testing","llm-tracing","llmops"],"archived":false,"github_pushed_at":"2026-02-11T14:43:33+00:00","maintenance_label":"Slowing","stars_delta_30d":5,"url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst","markdown_url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/raga-ai-hub-ragaai-catalyst","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=raga-ai-hub-ragaai-catalyst"}},{"type":"alternative","direction":"out","explanation":"Both PromptFoo and RAGAS aim at evaluating LLM applications, suggesting they could be alternatives to each other as they solve similar problems but likely in different ways.","successor_context":null,"tool":{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo"}},{"type":"integrates_with","direction":"out","explanation":"`Ragas` can be used to evaluate LLM applications which makes it a potential integration partner for `promptfoo`, a tool focused on evaluating and red-teaming LLM apps.","successor_context":null,"tool":{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"truera-trulens","name":"trulens","tagline":"Evaluation and Tracking for LLM Experiments and AI Agents","github_url":"https://github.com/truera/trulens","owner":"truera","repo":"trulens","owner_avatar_url":"https://avatars.githubusercontent.com/u/51224128?v=4","primary_language":"Python","stars":3516,"forks":327,"topics":["agent-evaluation","agentops","ai-agents","ai-monitoring","ai-observability","evals","explainable-ml","llm-eval","llm-evaluation","llmops","llms","machine-learning","neural-networks"],"archived":false,"github_pushed_at":"2026-08-20T10:21:00+00:00","maintenance_label":"Very active","stars_delta_30d":68,"url":"https://www.graphcanon.com/tools/truera-trulens","markdown_url":"https://www.graphcanon.com/tools/truera-trulens.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/truera-trulens","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=truera-trulens"}},{"type":"related","direction":"out","explanation":"Both langfuse and ragas focus on observability, evaluation, and management of LLM applications, but their functionalities do not directly integrate or overlap to a level that would constitute an 'integrates_with' relationship.","successor_context":null,"tool":{"slug":"langfuse-langfuse","name":"langfuse","tagline":"Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets","github_url":"https://github.com/langfuse/langfuse","owner":"langfuse","repo":"langfuse","owner_avatar_url":"https://avatars.githubusercontent.com/u/134601687?v=4","primary_language":"TypeScript","stars":32271,"forks":3466,"topics":["analytics","autogen","evaluation","langchain","large-language-models","llama-index","llm","llm-evaluation","llm-observability","llmops","monitoring","observability","open-source","openai","playground","prompt-engineering","prompt-management","self-hosted","ycombinator"],"archived":false,"github_pushed_at":"2026-07-31T22:58:07+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/langfuse-langfuse","markdown_url":"https://www.graphcanon.com/tools/langfuse-langfuse.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/langfuse-langfuse","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=langfuse-langfuse"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"langwatch-langwatch","name":"langwatch","tagline":"The platform for LLM evaluations and AI agent testing","github_url":"https://github.com/langwatch/langwatch","owner":"langwatch","repo":"langwatch","owner_avatar_url":"https://avatars.githubusercontent.com/u/146763322?v=4","primary_language":"TypeScript","stars":3479,"forks":340,"topics":["ai","analytics","datasets","dspy","evaluation","gpt","llm","llm-ops","llmops","low-code","observability","openai","prompt-engineering"],"archived":false,"github_pushed_at":"2026-08-07T21:03:52+00:00","maintenance_label":"Very active","stars_delta_30d":152,"url":"https://www.graphcanon.com/tools/langwatch-langwatch","markdown_url":"https://www.graphcanon.com/tools/langwatch-langwatch.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/langwatch-langwatch","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=langwatch-langwatch"}},{"type":"related","direction":"out","explanation":"`Ragas` and `Opik` share a focus on AI observability, although their specific use cases might differ.","successor_context":null,"tool":{"slug":"comet-ml-opik","name":"opik","tagline":"Debug, evaluate, and monitor your LLM applications with comprehensive tracing and production-ready dashboards","github_url":"https://github.com/comet-ml/opik","owner":"comet-ml","repo":"opik","owner_avatar_url":"https://avatars.githubusercontent.com/u/31487821?v=4","primary_language":"Python","stars":21177,"forks":1681,"topics":["evaluation","hacktoberfest","hacktoberfest2025","langchain","llama-index","llm","llm-evaluation","llm-observability","llmops","open-source","openai","playground","prompt-engineering"],"archived":false,"github_pushed_at":"2026-08-07T11:46:34+00:00","maintenance_label":"Very active","stars_delta_30d":767,"url":"https://www.graphcanon.com/tools/comet-ml-opik","markdown_url":"https://www.graphcanon.com/tools/comet-ml-opik.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/comet-ml-opik","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=comet-ml-opik"}},{"type":"alternative","direction":"in","explanation":"Ragas by VibrantLabsAI focuses on evaluation and observability for LLM applications, which aligns with the features offered by Langtrace.","successor_context":null,"tool":{"slug":"scale3-labs-langtrace","name":"langtrace","tagline":"Open Telemetry based observability tool for LLM applications","github_url":"https://github.com/Scale3-Labs/langtrace","owner":"Scale3-Labs","repo":"langtrace","owner_avatar_url":"https://avatars.githubusercontent.com/u/110545750?v=4","primary_language":"TypeScript","stars":1228,"forks":126,"topics":["ai","datasets","evaluations","gpt","langchain","llm","llm-framework","llmops","observability","open-source","open-telemetry","openai","prompt-engineering","tracing"],"archived":false,"github_pushed_at":"2025-11-17T15:08:48+00:00","maintenance_label":"Slowing","stars_delta_30d":12,"url":"https://www.graphcanon.com/tools/scale3-labs-langtrace","markdown_url":"https://www.graphcanon.com/tools/scale3-labs-langtrace.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/scale3-labs-langtrace","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=scale3-labs-langtrace"}},{"type":"alternative","direction":"in","explanation":"Both `continuous-eval` and `ragas` aim to provide comprehensive evaluation capabilities for LLM applications, making them alternatives.","successor_context":null,"tool":{"slug":"relari-ai-continuous-eval","name":"continuous-eval","tagline":"Data-Driven Evaluation for LLM-Powered Applications","github_url":"https://github.com/relari-ai/continuous-eval","owner":"relari-ai","repo":"continuous-eval","owner_avatar_url":"https://avatars.githubusercontent.com/u/135984758?v=4","primary_language":"Python","stars":515,"forks":38,"topics":["evaluation-framework","evaluation-metrics","information-retrieval","llm-evaluation","llmops","rag","retrieval-augmented-generation"],"archived":false,"github_pushed_at":"2026-08-10T22:12:03+00:00","maintenance_label":"Active","stars_delta_30d":-1,"url":"https://www.graphcanon.com/tools/relari-ai-continuous-eval","markdown_url":"https://www.graphcanon.com/tools/relari-ai-continuous-eval.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/relari-ai-continuous-eval","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=relari-ai-continuous-eval"}},{"type":"successor","direction":"in","explanation":"Ragas provides domain-specific evaluations and optimization tools specifically for Large Language Models (LLMs), while UpTrain offers a broader platform for both evaluating and improving generative AI systems including LLMs. The successor relationship from Ragas to UpTrain suggests an expansion in functionality from specialized LLM evaluation to a more comprehensive solution for various types of G","successor_context":null,"tool":{"slug":"uptrain-ai-uptrain","name":"uptrain","tagline":"Unified platform for evaluating and improving Generative AI applications","github_url":"https://github.com/uptrain-ai/uptrain","owner":"uptrain-ai","repo":"uptrain","owner_avatar_url":"https://avatars.githubusercontent.com/u/114582870?v=4","primary_language":"Python","stars":2359,"forks":204,"topics":["autoevaluation","evaluation","experimentation","hallucination-detection","jailbreak-detection","llm-eval","llm-prompting","llm-test","llmops","machine-learning","monitoring","openai-evals","prompt-engineering","root-cause-analysis"],"archived":false,"github_pushed_at":"2024-08-18T13:30:44+00:00","maintenance_label":"Dormant","stars_delta_30d":4,"url":"https://www.graphcanon.com/tools/uptrain-ai-uptrain","markdown_url":"https://www.graphcanon.com/tools/uptrain-ai-uptrain.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/uptrain-ai-uptrain","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=uptrain-ai-uptrain"}}],"neighbours":[{"slug":"infiniflow-ragflow","name":"ragflow","tagline":"Retrieval-Augmented Generation engine with agent capabilities","github_url":"https://github.com/infiniflow/ragflow","owner":"infiniflow","repo":"ragflow","owner_avatar_url":"https://avatars.githubusercontent.com/u/69962740?v=4","primary_language":"Go","stars":86541,"forks":10167,"topics":["agent-harness","agentic-ai","agentic-retrieval","agentic-search","ai","ai-agents","context-engine","context-engineering","context-management","harness-engineering","knowledge-compilation","llm-apps","rag","retrieval-augmented-generation"],"archived":false,"github_pushed_at":"2026-07-31T14:59:12+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/infiniflow-ragflow","markdown_url":"https://www.graphcanon.com/tools/infiniflow-ragflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/infiniflow-ragflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=infiniflow-ragflow","shared_categories":[]},{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo","shared_categories":["evaluation-observability"]},{"slug":"comet-ml-opik","name":"opik","tagline":"Debug, evaluate, and monitor your LLM applications with comprehensive tracing and production-ready dashboards","github_url":"https://github.com/comet-ml/opik","owner":"comet-ml","repo":"opik","owner_avatar_url":"https://avatars.githubusercontent.com/u/31487821?v=4","primary_language":"Python","stars":21177,"forks":1681,"topics":["evaluation","hacktoberfest","hacktoberfest2025","langchain","llama-index","llm","llm-evaluation","llm-observability","llmops","open-source","openai","playground","prompt-engineering"],"archived":false,"github_pushed_at":"2026-08-07T11:46:34+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/comet-ml-opik","markdown_url":"https://www.graphcanon.com/tools/comet-ml-opik.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/comet-ml-opik","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=comet-ml-opik","shared_categories":["evaluation-observability"]},{"slug":"openai-evals","name":"evals","tagline":"Framework for evaluating LLMs and LLM systems with an open-source registry of benchmarks.","github_url":"https://github.com/openai/evals","owner":"openai","repo":"evals","owner_avatar_url":"https://avatars.githubusercontent.com/u/14957082?v=4","primary_language":"Python","stars":19127,"forks":3050,"topics":[],"archived":false,"github_pushed_at":"2026-04-14T15:29:57+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/openai-evals","markdown_url":"https://www.graphcanon.com/tools/openai-evals.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/openai-evals","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=openai-evals","shared_categories":["evaluation-observability"]},{"slug":"confident-ai-deepeval","name":"deepeval","tagline":"LLM Evaluation Framework.","github_url":"https://github.com/confident-ai/deepeval","owner":"confident-ai","repo":"deepeval","owner_avatar_url":"https://avatars.githubusercontent.com/u/130858411?v=4","primary_language":"Python","stars":17226,"forks":1736,"topics":["evaluation-framework","evaluation-metrics","llm-evaluation","llm-evaluation-framework","llm-evaluation-metrics","python"],"archived":false,"github_pushed_at":"2026-07-27T11:33:31+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/confident-ai-deepeval","markdown_url":"https://www.graphcanon.com/tools/confident-ai-deepeval.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/confident-ai-deepeval","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=confident-ai-deepeval","shared_categories":["evaluation-observability"]},{"slug":"raga-ai-hub-ragaai-catalyst","name":"RagaAI-Catalyst","tagline":"Python SDK for AI agent observability and evaluation","github_url":"https://github.com/raga-ai-hub/RagaAI-Catalyst","owner":"raga-ai-hub","repo":"RagaAI-Catalyst","owner_avatar_url":"https://avatars.githubusercontent.com/u/161833182?v=4","primary_language":"Python","stars":16148,"forks":3565,"topics":["agentic-ai","agentic-ai-development","agentneo","agents","ai-agent-monitoring","ai-application-debugging","ai-evaluation-tools","ai-performance-optimization","ai-tool-interaction-monitoring","llm-testing","llm-tracing","llmops"],"archived":false,"github_pushed_at":"2026-02-11T14:43:33+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst","markdown_url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/raga-ai-hub-ragaai-catalyst","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=raga-ai-hub-ragaai-catalyst","shared_categories":["evaluation-observability"]},{"slug":"run-llama-rags","name":"rags","tagline":"Build ChatGPT over your data with natural language","github_url":"https://github.com/run-llama/rags","owner":"run-llama","repo":"rags","owner_avatar_url":"https://avatars.githubusercontent.com/u/130722866?v=4","primary_language":"Python","stars":6549,"forks":656,"topics":["agent","chatbot","chatgpt","gpts","llamaindex","llm","openai","rag","streamlit"],"archived":false,"github_pushed_at":"2024-04-05T05:36:59+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/run-llama-rags","markdown_url":"https://www.graphcanon.com/tools/run-llama-rags.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/run-llama-rags","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=run-llama-rags","shared_categories":[]},{"slug":"giskard-ai-giskard-oss","name":"giskard-oss","tagline":"Open-Source Evaluation & Testing library for LLM Agents","github_url":"https://github.com/Giskard-AI/giskard-oss","owner":"Giskard-AI","repo":"giskard-oss","owner_avatar_url":"https://avatars.githubusercontent.com/u/71782571?v=4","primary_language":"Python","stars":5727,"forks":511,"topics":["agent-evaluation","ai-red-team","ai-security","ai-testing","fairness-ai","llm","llm-eval","llm-evaluation","llm-security","llmops","ml-testing","ml-validation","mlops","rag-evaluation","red-team-tools","responsible-ai","trustworthy-ai"],"archived":false,"github_pushed_at":"2026-08-01T23:22:37+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/giskard-ai-giskard-oss","markdown_url":"https://www.graphcanon.com/tools/giskard-ai-giskard-oss.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/giskard-ai-giskard-oss","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=giskard-ai-giskard-oss","shared_categories":["evaluation-observability"]},{"slug":"marker-inc-korea-autorag","name":"AutoRAG","tagline":"Open-source framework for RAG evaluation and optimization via AutoML","github_url":"https://github.com/Marker-Inc-Korea/AutoRAG","owner":"Marker-Inc-Korea","repo":"AutoRAG","owner_avatar_url":"https://avatars.githubusercontent.com/u/74290595?v=4","primary_language":"TypeScript","stars":4968,"forks":419,"topics":["analysis","automl","benchmarking","document-parser","embeddings","evaluation","llm","llm-evaluation","llm-ops","open-source","ops","optimization","pipeline","python","qa","rag","rag-evaluation","retrieval-augmented-generation"],"archived":false,"github_pushed_at":"2026-08-05T14:46:13+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/marker-inc-korea-autorag","markdown_url":"https://www.graphcanon.com/tools/marker-inc-korea-autorag.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/marker-inc-korea-autorag","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=marker-inc-korea-autorag","shared_categories":["evaluation-observability"]},{"slug":"huggingface-lighteval","name":"lighteval","tagline":"All-in-one toolkit for evaluating LLMs across multiple backends","github_url":"https://github.com/huggingface/lighteval","owner":"huggingface","repo":"lighteval","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":2508,"forks":523,"topics":["evaluation","evaluation-framework","evaluation-metrics","huggingface"],"archived":false,"github_pushed_at":"2026-06-29T13:03:33+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/huggingface-lighteval","markdown_url":"https://www.graphcanon.com/tools/huggingface-lighteval.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-lighteval","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-lighteval","shared_categories":["evaluation-observability"]},{"slug":"ray-project-llm-applications","name":"llm-applications","tagline":"Comprehensive guide to building RAG-based LLM applications for production","github_url":"https://github.com/ray-project/llm-applications","owner":"ray-project","repo":"llm-applications","owner_avatar_url":"https://avatars.githubusercontent.com/u/22125274?v=4","primary_language":"Jupyter Notebook","stars":1857,"forks":255,"topics":["anyscale","fine-tuning","llama2","llms","machine-learning","openai","ray","serving"],"archived":false,"github_pushed_at":"2024-08-02T00:27:10+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/ray-project-llm-applications","markdown_url":"https://www.graphcanon.com/tools/ray-project-llm-applications.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ray-project-llm-applications","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ray-project-llm-applications","shared_categories":[]},{"slug":"deepsense-ai-ragbits","name":"ragbits","tagline":"Building blocks for rapid development of GenAI applications","github_url":"https://github.com/deepsense-ai/ragbits","owner":"deepsense-ai","repo":"ragbits","owner_avatar_url":"https://avatars.githubusercontent.com/u/14294751?v=4","primary_language":"Python","stars":1668,"forks":143,"topics":["agents","document-search","evaluation","guardrails","llms","optimization","prompts","rag","vector-stores"],"archived":false,"github_pushed_at":"2026-05-18T07:51:25+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/deepsense-ai-ragbits","markdown_url":"https://www.graphcanon.com/tools/deepsense-ai-ragbits.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/deepsense-ai-ragbits","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=deepsense-ai-ragbits","shared_categories":["evaluation-observability"]}]}}