{"data":{"node":{"slug":"giskard-ai-giskard-oss","name":"giskard-oss","tagline":"Open-Source Evaluation & Testing library for LLM Agents","github_url":"https://github.com/Giskard-AI/giskard-oss","owner":"Giskard-AI","repo":"giskard-oss","owner_avatar_url":"https://avatars.githubusercontent.com/u/71782571?v=4","primary_language":"Python","stars":5727,"forks":511,"topics":["agent-evaluation","ai-red-team","ai-security","ai-testing","fairness-ai","llm","llm-eval","llm-evaluation","llm-security","llmops","ml-testing","ml-validation","mlops","rag-evaluation","red-team-tools","responsible-ai","trustworthy-ai"],"archived":false,"github_pushed_at":"2026-08-01T23:22:37+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/giskard-ai-giskard-oss","markdown_url":"https://www.graphcanon.com/tools/giskard-ai-giskard-oss.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/giskard-ai-giskard-oss","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=giskard-ai-giskard-oss"},"categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"agent-evaluation","name":"agent-evaluation"},{"slug":"ai-red-team","name":"ai-red-team"},{"slug":"ai-security","name":"ai-security"},{"slug":"ai-testing","name":"ai-testing"},{"slug":"fairness-ai","name":"fairness-ai"},{"slug":"llm","name":"llm"},{"slug":"llm-eval","name":"llm-eval"},{"slug":"llm-evaluation","name":"llm-evaluation"}],"edges":[{"type":"alternative","direction":"out","explanation":"Both Giskard and PromptFoo offer evaluation and red-teaming capabilities for LLM apps, targeting similar functionality but with different implementations.","successor_context":null,"tool":{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo"}},{"type":"related","direction":"out","explanation":"Both Langfuse and Giskard are involved in observability, evaluation, and management of LLMs, but they serve somewhat different purposes within the AI development lifecycle.","successor_context":null,"tool":{"slug":"langfuse-langfuse","name":"langfuse","tagline":"Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets","github_url":"https://github.com/langfuse/langfuse","owner":"langfuse","repo":"langfuse","owner_avatar_url":"https://avatars.githubusercontent.com/u/134601687?v=4","primary_language":"TypeScript","stars":32271,"forks":3466,"topics":["analytics","autogen","evaluation","langchain","large-language-models","llama-index","llm","llm-evaluation","llm-observability","llmops","monitoring","observability","open-source","openai","playground","prompt-engineering","prompt-management","self-hosted","ycombinator"],"archived":false,"github_pushed_at":"2026-07-31T22:58:07+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/langfuse-langfuse","markdown_url":"https://www.graphcanon.com/tools/langfuse-langfuse.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/langfuse-langfuse","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=langfuse-langfuse"}},{"type":"related","direction":"out","explanation":"Evidently also provides an open-source framework for observability of ML and LLM models, similar to Giskard. However, its focus is on general machine learning systems rather than specialized LLM evaluations.","successor_context":null,"tool":{"slug":"evidentlyai-evidently","name":"evidently","tagline":"An open-source ML and LLM observability framework.","github_url":"https://github.com/evidentlyai/evidently","owner":"evidentlyai","repo":"evidently","owner_avatar_url":"https://avatars.githubusercontent.com/u/75031056?v=4","primary_language":"Jupyter Notebook","stars":7790,"forks":895,"topics":["data-drift","data-quality","data-science","data-validation","generative-ai","hacktoberfest","html-report","jupyter-notebook","llm","llmops","machine-learning","mlops","model-monitoring","pandas-dataframe"],"archived":false,"github_pushed_at":"2026-08-05T16:29:57+00:00","maintenance_label":"Very active","stars_delta_30d":117,"url":"https://www.graphcanon.com/tools/evidentlyai-evidently","markdown_url":"https://www.graphcanon.com/tools/evidentlyai-evidently.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/evidentlyai-evidently","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=evidentlyai-evidently"}},{"type":"related","direction":"out","explanation":"Phoenix from Arize AI offers observability and evaluation for AI models, which is complementary to the testing and evaluation aspects provided by Giskard.","successor_context":null,"tool":{"slug":"arize-ai-phoenix","name":"phoenix","tagline":"AI Observability & Evaluation","github_url":"https://github.com/Arize-ai/phoenix","owner":"Arize-ai","repo":"phoenix","owner_avatar_url":"https://avatars.githubusercontent.com/u/59858760?v=4","primary_language":"Python","stars":10847,"forks":1028,"topics":["agents","ai-monitoring","ai-observability","aiengineering","anthropic","datasets","evals","langchain","llamaindex","llm-eval","llm-evaluation","llmops","llms","openai","prompt-engineering","smolagents"],"archived":false,"github_pushed_at":"2026-08-01T11:48:49+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/arize-ai-phoenix","markdown_url":"https://www.graphcanon.com/tools/arize-ai-phoenix.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/arize-ai-phoenix","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=arize-ai-phoenix"}},{"type":"integrates_with","direction":"out","explanation":"giskard-oss integrates with agenta because giskard-oss provides an open-source library for evaluating AI agents through dynamic and multi-turn testing, which complements agenta's LLMOps platform that includes evaluation capabilities and observability tools for managing prompts and monitoring performance of language models.","successor_context":null,"tool":{"slug":"agenta-ai-agenta","name":"agenta","tagline":"The open-source LLMOps platform for prompt management, evaluation, and observability.","github_url":"https://github.com/Agenta-AI/agenta","owner":"Agenta-AI","repo":"agenta","owner_avatar_url":"https://avatars.githubusercontent.com/u/127993667?v=4","primary_language":"TypeScript","stars":4445,"forks":609,"topics":["agent-builder","agent-observability","agent-orchestration","agent-workspace","agentic-ai","ai-agent","ai-agents","ai-automation","ai-skills-manager","ai-workflow-builder","harness","mcp","open-source","self-hosted","workflow-automation"],"archived":false,"github_pushed_at":"2026-08-07T10:41:36+00:00","maintenance_label":"Very active","stars_delta_30d":170,"url":"https://www.graphcanon.com/tools/agenta-ai-agenta","markdown_url":"https://www.graphcanon.com/tools/agenta-ai-agenta.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/agenta-ai-agenta","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=agenta-ai-agenta"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"traceloop-openllmetry","name":"openllmetry","tagline":"Open-source observability for GenAI and LLM applications based on OpenTelemetry.","github_url":"https://github.com/traceloop/openllmetry","owner":"traceloop","repo":"openllmetry","owner_avatar_url":"https://avatars.githubusercontent.com/u/125419530?v=4","primary_language":"Python","stars":7377,"forks":1047,"topics":["artifical-intelligence","datascience","generative-ai","good-first-issue","good-first-issues","help-wanted","llm","llmops","metrics","ml","model-monitoring","monitoring","observability","open-source","open-telemetry","opentelemetry","opentelemetry-python","python"],"archived":false,"github_pushed_at":"2026-08-10T08:49:01+00:00","maintenance_label":"Very active","stars_delta_30d":75,"url":"https://www.graphcanon.com/tools/traceloop-openllmetry","markdown_url":"https://www.graphcanon.com/tools/traceloop-openllmetry.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/traceloop-openllmetry","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=traceloop-openllmetry"}},{"type":"integrates_with","direction":"in","explanation":"Opik could integrate with Giskard OSS to improve testing and evaluation of AI agents due to overlapping focus areas on LLM/agent observability.","successor_context":null,"tool":{"slug":"comet-ml-opik","name":"opik","tagline":"Debug, evaluate, and monitor your LLM applications with comprehensive tracing and production-ready dashboards","github_url":"https://github.com/comet-ml/opik","owner":"comet-ml","repo":"opik","owner_avatar_url":"https://avatars.githubusercontent.com/u/31487821?v=4","primary_language":"Python","stars":21177,"forks":1681,"topics":["evaluation","hacktoberfest","hacktoberfest2025","langchain","llama-index","llm","llm-evaluation","llm-observability","llmops","open-source","openai","playground","prompt-engineering"],"archived":false,"github_pushed_at":"2026-08-07T11:46:34+00:00","maintenance_label":"Very active","stars_delta_30d":767,"url":"https://www.graphcanon.com/tools/comet-ml-opik","markdown_url":"https://www.graphcanon.com/tools/comet-ml-opik.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/comet-ml-opik","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=comet-ml-opik"}},{"type":"related","direction":"in","explanation":"`Ragas` focuses on evaluations of LLM applications, while `Giskard-OSS` covers evals and test generation for agentic systems; both serve different aspects in the evaluation process.","successor_context":null,"tool":{"slug":"vibrantlabsai-ragas","name":"ragas","tagline":"Supercharge Your LLM Application Evaluations 🚀","github_url":"https://github.com/vibrantlabsai/ragas","owner":"vibrantlabsai","repo":"ragas","owner_avatar_url":"https://avatars.githubusercontent.com/u/122604797?v=4","primary_language":"Python","stars":15388,"forks":1637,"topics":["evaluation","llm","llmops"],"archived":false,"github_pushed_at":"2026-02-24T07:47:19+00:00","maintenance_label":"Slowing","stars_delta_30d":470,"url":"https://www.graphcanon.com/tools/vibrantlabsai-ragas","markdown_url":"https://www.graphcanon.com/tools/vibrantlabsai-ragas.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vibrantlabsai-ragas","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vibrantlabsai-ragas"}},{"type":"alternative","direction":"in","explanation":"Both Promptfoo and Giskard OSS evaluate LLMs through red-teaming techniques and test generation for agentic systems.","successor_context":null,"tool":{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo"}},{"type":"successor","direction":"in","explanation":"While Giskard OSS could be seen as another tool for red-teaming, Promptfoo's continuous development and integration with major AI providers suggest it serves as a successor in terms of comprehensive testing frameworks.","successor_context":{"status":"recommended","reason":"Continual improvement and major provider affiliations signify recommendation over older tools."},"tool":{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo"}},{"type":"alternative","direction":"in","explanation":"Evidently and Giskard-OSS both serve to evaluate and test AI systems, but they differ in their primary focus; Evidently is an observability framework that monitors AI systems across various data types using a wide array of metrics, while Giskard-OSS specializes in evaluating AI agents through dynamic and multi-turn testing scenarios.","successor_context":null,"tool":{"slug":"evidentlyai-evidently","name":"evidently","tagline":"An open-source ML and LLM observability framework.","github_url":"https://github.com/evidentlyai/evidently","owner":"evidentlyai","repo":"evidently","owner_avatar_url":"https://avatars.githubusercontent.com/u/75031056?v=4","primary_language":"Jupyter Notebook","stars":7790,"forks":895,"topics":["data-drift","data-quality","data-science","data-validation","generative-ai","hacktoberfest","html-report","jupyter-notebook","llm","llmops","machine-learning","mlops","model-monitoring","pandas-dataframe"],"archived":false,"github_pushed_at":"2026-08-05T16:29:57+00:00","maintenance_label":"Very active","stars_delta_30d":117,"url":"https://www.graphcanon.com/tools/evidentlyai-evidently","markdown_url":"https://www.graphcanon.com/tools/evidentlyai-evidently.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/evidentlyai-evidently","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=evidentlyai-evidently"}},{"type":"integrates_with","direction":"in","explanation":"Given Giskard OSS focuses on evaluations and red-teaming which are also covered by RagaAI Catalyst's features, they can be integrated to complement each other for comprehensive testing.","successor_context":null,"tool":{"slug":"raga-ai-hub-ragaai-catalyst","name":"RagaAI-Catalyst","tagline":"Python SDK for AI agent observability and evaluation","github_url":"https://github.com/raga-ai-hub/RagaAI-Catalyst","owner":"raga-ai-hub","repo":"RagaAI-Catalyst","owner_avatar_url":"https://avatars.githubusercontent.com/u/161833182?v=4","primary_language":"Python","stars":16148,"forks":3565,"topics":["agentic-ai","agentic-ai-development","agentneo","agents","ai-agent-monitoring","ai-application-debugging","ai-evaluation-tools","ai-performance-optimization","ai-tool-interaction-monitoring","llm-testing","llm-tracing","llmops"],"archived":false,"github_pushed_at":"2026-02-11T14:43:33+00:00","maintenance_label":"Slowing","stars_delta_30d":5,"url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst","markdown_url":"https://www.graphcanon.com/tools/raga-ai-hub-ragaai-catalyst.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/raga-ai-hub-ragaai-catalyst","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=raga-ai-hub-ragaai-catalyst"}},{"type":"integrates_with","direction":"in","explanation":"Giskard-OSS focuses on evaluations, red-teaming, and test generation for agentic systems which can directly integrate with PySpur for rapid iteration and testing.","successor_context":null,"tool":{"slug":"pyspur-dev-pyspur","name":"pyspur","tagline":"A visual playground for agentic workflows","github_url":"https://github.com/PySpur-Dev/pyspur","owner":"PySpur-Dev","repo":"pyspur","owner_avatar_url":"https://avatars.githubusercontent.com/u/182547524?v=4","primary_language":"TypeScript","stars":5771,"forks":429,"topics":["agent","agents","ai","builder","deepseek","framework","gemini","graph","human-in-the-loop","llm","llms","loops","multimodal","ollama","python","rag","reasoning","tool","trace","workflow"],"archived":false,"github_pushed_at":"2026-06-29T17:53:12+00:00","maintenance_label":"Steady","stars_delta_30d":15,"url":"https://www.graphcanon.com/tools/pyspur-dev-pyspur","markdown_url":"https://www.graphcanon.com/tools/pyspur-dev-pyspur.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pyspur-dev-pyspur","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pyspur-dev-pyspur"}},{"type":"integrates_with","direction":"in","explanation":"Both repositories provide tools for evaluating and testing generative AI systems, with lmms-eval focusing on multimodal evaluations while Giskard OSS focuses broadly on evals and red teaming.","successor_context":null,"tool":{"slug":"evolvinglmms-lab-lmms-eval","name":"lmms-eval","tagline":"One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks","github_url":"https://github.com/EvolvingLMMs-Lab/lmms-eval","owner":"EvolvingLMMs-Lab","repo":"lmms-eval","owner_avatar_url":"https://avatars.githubusercontent.com/u/154951679?v=4","primary_language":"Python","stars":4368,"forks":639,"topics":["agi","audio-evaluation","benchmark","evaluation","large-language-models","llm-evaluation","multimodal","multimodal-evaluation","video-understanding","vision-language-model","vlm"],"archived":false,"github_pushed_at":"2026-08-06T02:22:23+00:00","maintenance_label":"Active","stars_delta_30d":52,"url":"https://www.graphcanon.com/tools/evolvinglmms-lab-lmms-eval","markdown_url":"https://www.graphcanon.com/tools/evolvinglmms-lab-lmms-eval.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/evolvinglmms-lab-lmms-eval","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=evolvinglmms-lab-lmms-eval"}},{"type":"alternative","direction":"in","explanation":"`Vigil` focuses on protecting against adversarial attacks, while `Giskard OSS` evaluates and red-teams agentic systems. Both tools aim at improving the security/robustness of AI models but in slightly different ways.","successor_context":null,"tool":{"slug":"deadbits-vigil-llm","name":"vigil-llm","tagline":"Detect prompt injections and other risky inputs in LLMs","github_url":"https://github.com/deadbits/vigil-llm","owner":"deadbits","repo":"vigil-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/1332757?v=4","primary_language":"Python","stars":496,"forks":56,"topics":["adversarial-attacks","adversarial-machine-learning","large-language-models","llm-security","llmops","prompt-injection","security-tools","yara-scanner"],"archived":false,"github_pushed_at":"2024-01-31T18:43:41+00:00","maintenance_label":"Dormant","stars_delta_30d":5,"url":"https://www.graphcanon.com/tools/deadbits-vigil-llm","markdown_url":"https://www.graphcanon.com/tools/deadbits-vigil-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/deadbits-vigil-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=deadbits-vigil-llm"}}],"neighbours":[{"slug":"microsoft-ai-agents-for-beginners","name":"ai-agents-for-beginners","tagline":"12 Lessons to Get Started Building AI Agents","github_url":"https://github.com/microsoft/ai-agents-for-beginners","owner":"microsoft","repo":"ai-agents-for-beginners","owner_avatar_url":"https://avatars.githubusercontent.com/u/6154722?v=4","primary_language":"Jupyter Notebook","stars":72665,"forks":24038,"topics":["agentic-ai","agentic-framework","agentic-rag","ai-agents","ai-agents-framework","autogen","foundry","foundry-local","generative-ai","microsoft-foundry","semantic-kernel"],"archived":false,"github_pushed_at":"2026-08-18T11:47:04+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/microsoft-ai-agents-for-beginners","markdown_url":"https://www.graphcanon.com/tools/microsoft-ai-agents-for-beginners.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/microsoft-ai-agents-for-beginners","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=microsoft-ai-agents-for-beginners","shared_categories":[]},{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo","shared_categories":["evaluation-observability"]},{"slug":"nirdiamant-agents-towards-production","name":"agents-towards-production","tagline":"End-to-end, code-first tutorials for building production-grade GenAI agents","github_url":"https://github.com/NirDiamant/agents-towards-production","owner":"NirDiamant","repo":"agents-towards-production","owner_avatar_url":"https://avatars.githubusercontent.com/u/28316913?v=4","primary_language":"Jupyter Notebook","stars":21298,"forks":2824,"topics":["agent","agent-framework","agentic-ai","agents","ai-agents","deployment","genai","generative-ai","langgraph","llm","llms","mcp","mlops","multi-agent-systems","observability","production","python","rag","tutorials"],"archived":false,"github_pushed_at":"2026-08-15T00:52:10+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/nirdiamant-agents-towards-production","markdown_url":"https://www.graphcanon.com/tools/nirdiamant-agents-towards-production.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nirdiamant-agents-towards-production","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nirdiamant-agents-towards-production","shared_categories":[]},{"slug":"microsoft-agent-framework","name":"agent-framework","tagline":"Framework for building and deploying AI agents and multi-agent workflows","github_url":"https://github.com/microsoft/agent-framework","owner":"microsoft","repo":"agent-framework","owner_avatar_url":"https://avatars.githubusercontent.com/u/6154722?v=4","primary_language":"Python","stars":12718,"forks":2143,"topics":["agent-framework","agentic-ai","agents","ai","dotnet","multi-agent","orchestration","python","sdk","workflows"],"archived":false,"github_pushed_at":"2026-08-10T23:32:05+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/microsoft-agent-framework","markdown_url":"https://www.graphcanon.com/tools/microsoft-agent-framework.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/microsoft-agent-framework","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=microsoft-agent-framework","shared_categories":[]},{"slug":"hegelai-prompttools","name":"prompttools","tagline":"Open-source tools for prompt testing and experimentation","github_url":"https://github.com/hegelai/prompttools","owner":"hegelai","repo":"prompttools","owner_avatar_url":"https://avatars.githubusercontent.com/u/136523567?v=4","primary_language":"Python","stars":3046,"forks":255,"topics":["deep-learning","developer-tools","embeddings","large-language-models","llms","machine-learning","prompt-engineering","python","vector-search"],"archived":false,"github_pushed_at":"2026-02-11T03:24:04+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/hegelai-prompttools","markdown_url":"https://www.graphcanon.com/tools/hegelai-prompttools.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/hegelai-prompttools","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=hegelai-prompttools","shared_categories":[]},{"slug":"future-agi-future-agi","name":"future-agi","tagline":"End-to-end platform for evaluating, observing, and improving LLM and AI agent applications","github_url":"https://github.com/future-agi/future-agi","owner":"future-agi","repo":"future-agi","owner_avatar_url":"https://avatars.githubusercontent.com/u/147392366?v=4","primary_language":"Python","stars":1559,"forks":449,"topics":["ai-agents","ai-evals","ai-gateway","ai-optimization","ai-simulations","evaluation-framework","guardrails","hallucination-detection","llm","llm-evaluation","llm-observability","llmops","model-evaluation","observability","opentelemetry","rag","rag-evaluation","simulation","telemetry","tracing"],"archived":false,"github_pushed_at":"2026-08-01T14:44:07+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/future-agi-future-agi","markdown_url":"https://www.graphcanon.com/tools/future-agi-future-agi.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/future-agi-future-agi","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=future-agi-future-agi","shared_categories":["evaluation-observability"]},{"slug":"e2b-dev-awesome-ai-sdks","name":"awesome-ai-sdks","tagline":"A database of SDKs for AI agents creation and management","github_url":"https://github.com/e2b-dev/awesome-ai-sdks","owner":"e2b-dev","repo":"awesome-ai-sdks","owner_avatar_url":"https://avatars.githubusercontent.com/u/129434473?v=4","primary_language":null,"stars":1213,"forks":361,"topics":["agent","agentops","agents","ai","ai-agents","awesome","awesome-list","chatgpt","e2b","framework","langchain","llama-index","llm","llmops","openai","sdk","tools","vercel"],"archived":false,"github_pushed_at":"2026-07-09T17:29:58+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/e2b-dev-awesome-ai-sdks","markdown_url":"https://www.graphcanon.com/tools/e2b-dev-awesome-ai-sdks.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/e2b-dev-awesome-ai-sdks","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=e2b-dev-awesome-ai-sdks","shared_categories":[]},{"slug":"pguso-agents-from-scratch","name":"agents-from-scratch","tagline":"Build AI agents locally without relying on frameworks or cloud APIs.","github_url":"https://github.com/pguso/agents-from-scratch","owner":"pguso","repo":"agents-from-scratch","owner_avatar_url":"https://avatars.githubusercontent.com/u/4007140?v=4","primary_language":"Python","stars":954,"forks":240,"topics":["agent-architecture","ai-agents","ai-education","ai-from-scratch","artificial-intelligence","llama","llm","local-llm","machine-learning","no-framework","open-source-education","prompt-engineering","python"],"archived":false,"github_pushed_at":"2026-07-25T15:42:36+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/pguso-agents-from-scratch","markdown_url":"https://www.graphcanon.com/tools/pguso-agents-from-scratch.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pguso-agents-from-scratch","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pguso-agents-from-scratch","shared_categories":[]},{"slug":"darkrishabh-agent-skills-eval","name":"agent-skills-eval","tagline":"A test runner for agentskills.io-style AI agent skills","github_url":"https://github.com/darkrishabh/agent-skills-eval","owner":"darkrishabh","repo":"agent-skills-eval","owner_avatar_url":"https://avatars.githubusercontent.com/u/812474?v=4","primary_language":"TypeScript","stars":637,"forks":36,"topics":["agent-evals","agent-skills","agentskills","ai-agents","cli","jsonl","llm-evals","llm-evaluation","openai-compatible","typescript","yaml"],"archived":false,"github_pushed_at":"2026-07-15T15:24:06+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/darkrishabh-agent-skills-eval","markdown_url":"https://www.graphcanon.com/tools/darkrishabh-agent-skills-eval.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/darkrishabh-agent-skills-eval","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=darkrishabh-agent-skills-eval","shared_categories":["evaluation-observability"]},{"slug":"framerslab-agentos","name":"agentos","tagline":"TypeScript AI agent framework providing cognitive memory and runtime tool forging with support for multi-agent orchestration","github_url":"https://github.com/framerslab/agentos","owner":"framerslab","repo":"agentos","owner_avatar_url":"https://avatars.githubusercontent.com/u/184314983?v=4","primary_language":"TypeScript","stars":601,"forks":89,"topics":["agent-framework","agent-memory","agentic-ai","ai-agent-framework","ai-agents","autonomous-agents","cognitive-memory","emergent-behavior","guardrails","hexaco","llm","llm-orchestration","long-term-memory","multi-agent","rag","runtime-tool-generation","tool-use","vector-search","voice-ai"],"archived":false,"github_pushed_at":"2026-07-22T07:15:26+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/framerslab-agentos","markdown_url":"https://www.graphcanon.com/tools/framerslab-agentos.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/framerslab-agentos","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=framerslab-agentos","shared_categories":[]},{"slug":"declare-lab-instruct-eval","name":"instruct-eval","tagline":"Quantitative evaluation for instruction-tuned language models","github_url":"https://github.com/declare-lab/instruct-eval","owner":"declare-lab","repo":"instruct-eval","owner_avatar_url":"https://avatars.githubusercontent.com/u/59164695?v=4","primary_language":"Python","stars":552,"forks":45,"topics":["instruct-tuning","llm"],"archived":false,"github_pushed_at":"2024-03-10T05:00:00+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/declare-lab-instruct-eval","markdown_url":"https://www.graphcanon.com/tools/declare-lab-instruct-eval.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/declare-lab-instruct-eval","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=declare-lab-instruct-eval","shared_categories":["evaluation-observability"]},{"slug":"rhesis-ai-rhesis","name":"rhesis","tagline":"Testing platform for AI teams to generate tests and evaluate system performance","github_url":"https://github.com/rhesis-ai/rhesis","owner":"rhesis-ai","repo":"rhesis","owner_avatar_url":"https://avatars.githubusercontent.com/u/168341335?v=4","primary_language":"Python","stars":381,"forks":31,"topics":["annotations","feedback-loop","hypothesis-testing","llmops","regression-testing","systematic-evaluation"],"archived":false,"github_pushed_at":"2026-07-28T15:13:49+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/rhesis-ai-rhesis","markdown_url":"https://www.graphcanon.com/tools/rhesis-ai-rhesis.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/rhesis-ai-rhesis","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=rhesis-ai-rhesis","shared_categories":["evaluation-observability"]}]}}