{"data":{"node":{"slug":"apache-hamilton","name":"hamilton","tagline":"Modular dataflow definition for Python environments","github_url":"https://github.com/apache/hamilton","owner":"apache","repo":"hamilton","owner_avatar_url":"https://avatars.githubusercontent.com/u/47359?v=4","primary_language":"Jupyter Notebook","stars":2557,"forks":203,"topics":["dag","data-analysis","data-engineering","data-science","dataframe","etl","etl-framework","etl-pipeline","feature-engineering","hacktoberfest","lineage","llmops","machine-learning","mlops","orchestration","pandas","python","rag","software-engineering"],"archived":false,"github_pushed_at":"2026-08-01T18:45:37+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/apache-hamilton","markdown_url":"https://www.graphcanon.com/tools/apache-hamilton.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/apache-hamilton","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=apache-hamilton"},"categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"},{"slug":"developer-tools","name":"Developer Tools","url":"https://www.graphcanon.com/categories/developer-tools","markdown_url":"https://www.graphcanon.com/categories/developer-tools.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/developer-tools"}],"tags":[{"slug":"dag","name":"dag"},{"slug":"data-analysis","name":"data-analysis"},{"slug":"data-engineering","name":"data-engineering"},{"slug":"data-science","name":"data-science"},{"slug":"feature-engineering","name":"feature-engineering"},{"slug":"llmops","name":"llmops"},{"slug":"machine-learning","name":"machine-learning"},{"slug":"mlops","name":"mlops"}],"edges":[{"type":"related","direction":"out","explanation":"MindsDB deals with automating ML processes and can work alongside Hamilton which focuses on dataflow management for machine learning use cases.","successor_context":null,"tool":{"slug":"mindsdb-minds","name":"mindshub","tagline":"Get real work done with AI. Easily swap models while keeping your builds intact.","github_url":"https://github.com/mindsdb/mindshub","owner":"mindsdb","repo":"mindshub","owner_avatar_url":"https://avatars.githubusercontent.com/u/31035808?v=4","primary_language":"Makefile","stars":39496,"forks":6224,"topics":["agents","ai","anton","artificial-inteligence","claude","claude-cowork","codex","cowork","deepseek","glm","hermes","hermes-agent","kimi","mcp","mindsdb"],"archived":false,"github_pushed_at":"2026-07-10T16:53:47+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/mindsdb-minds","markdown_url":"https://www.graphcanon.com/tools/mindsdb-minds.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mindsdb-minds","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mindsdb-minds"}},{"type":"integrates_with","direction":"out","explanation":"Hamilton creates testable, modular self-documenting dataflows which could be useful for integrating with RAG engines like RagFlow.","successor_context":null,"tool":{"slug":"infiniflow-ragflow","name":"ragflow","tagline":"Retrieval-Augmented Generation engine with agent capabilities","github_url":"https://github.com/infiniflow/ragflow","owner":"infiniflow","repo":"ragflow","owner_avatar_url":"https://avatars.githubusercontent.com/u/69962740?v=4","primary_language":"Go","stars":86541,"forks":10167,"topics":["agent-harness","agentic-ai","agentic-retrieval","agentic-search","ai","ai-agents","context-engine","context-engineering","context-management","harness-engineering","knowledge-compilation","llm-apps","rag","retrieval-augmented-generation"],"archived":false,"github_pushed_at":"2026-07-31T14:59:12+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/infiniflow-ragflow","markdown_url":"https://www.graphcanon.com/tools/infiniflow-ragflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/infiniflow-ragflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=infiniflow-ragflow"}},{"type":"integrates_with","direction":"out","explanation":"Hamilton can be used in conjunction with LLM-app RAG templates and AI pipelines for defining and scaling data flows involved in the applications.","successor_context":null,"tool":{"slug":"pathwaycom-llm-app","name":"llm-app","tagline":"Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data.","github_url":"https://github.com/pathwaycom/llm-app","owner":"pathwaycom","repo":"llm-app","owner_avatar_url":"https://avatars.githubusercontent.com/u/25750857?v=4","primary_language":"Jupyter Notebook","stars":59037,"forks":1466,"topics":["chatbot","hugging-face","llm","llm-local","llm-prompting","llm-security","llmops","machine-learning","open-ai","pathway","rag","real-time","retrieval-augmented-generation","vector-database","vector-index"],"archived":false,"github_pushed_at":"2026-07-05T17:59:07+00:00","maintenance_label":"Steady","stars_delta_30d":11,"url":"https://www.graphcanon.com/tools/pathwaycom-llm-app","markdown_url":"https://www.graphcanon.com/tools/pathwaycom-llm-app.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pathwaycom-llm-app","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pathwaycom-llm-app"}},{"type":"related","direction":"out","explanation":"Both Apache Hamilton and pandas-ai are tools that leverage Python for data transformation tasks, although pandas-ai integrates LLMs for conversational data analysis.","successor_context":null,"tool":{"slug":"sinaptik-ai-pandas-ai","name":"pandas-ai","tagline":"Chat with your database or your datalake using LLMs and RAG.","github_url":"https://github.com/sinaptik-ai/pandas-ai","owner":"sinaptik-ai","repo":"pandas-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/154438448?v=4","primary_language":"Python","stars":23746,"forks":2342,"topics":["ai","csv","data","data-analysis","data-science","data-visualization","database","datalake","gpt-4","llm","pandas","sql","text-to-sql"],"archived":false,"github_pushed_at":"2025-10-28T10:02:13+00:00","maintenance_label":"Slowing","stars_delta_30d":90,"url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai","markdown_url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sinaptik-ai-pandas-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sinaptik-ai-pandas-ai"}},{"type":"integrates_with","direction":"out","explanation":"Hamilton aids in defining modular dataflows, and Ray is a unified framework for scaling Python applications. They integrate to scale dataflow processing.","successor_context":null,"tool":{"slug":"ray-project-ray","name":"ray","tagline":"Ray is an AI compute engine with a core distributed runtime and AI Libraries for accelerating ML workloads.","github_url":"https://github.com/ray-project/ray","owner":"ray-project","repo":"ray","owner_avatar_url":"https://avatars.githubusercontent.com/u/22125274?v=4","primary_language":"Python","stars":43526,"forks":7929,"topics":["data-science","deep-learning","deployment","distributed","hyperparameter-optimization","hyperparameter-search","large-language-models","llm","llm-inference","llm-serving","machine-learning","optimization","parallel","python","pytorch","ray","reinforcement-learning","rllib","serving","tensorflow"],"archived":false,"github_pushed_at":"2026-08-16T00:26:16+00:00","maintenance_label":"Very active","stars_delta_30d":270,"url":"https://www.graphcanon.com/tools/ray-project-ray","markdown_url":"https://www.graphcanon.com/tools/ray-project-ray.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ray-project-ray","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ray-project-ray"}}],"neighbours":[{"slug":"sinaptik-ai-pandas-ai","name":"pandas-ai","tagline":"Chat with your database or your datalake using LLMs and RAG.","github_url":"https://github.com/sinaptik-ai/pandas-ai","owner":"sinaptik-ai","repo":"pandas-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/154438448?v=4","primary_language":"Python","stars":23746,"forks":2342,"topics":["ai","csv","data","data-analysis","data-science","data-visualization","database","datalake","gpt-4","llm","pandas","sql","text-to-sql"],"archived":false,"github_pushed_at":"2025-10-28T10:02:13+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai","markdown_url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sinaptik-ai-pandas-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sinaptik-ai-pandas-ai","shared_categories":["data-retrieval"]},{"slug":"argoproj-argo-workflows","name":"argo-workflows","tagline":"Workflow Engine for Kubernetes","github_url":"https://github.com/argoproj/argo-workflows","owner":"argoproj","repo":"argo-workflows","owner_avatar_url":"https://avatars.githubusercontent.com/u/30269780?v=4","primary_language":"Go","stars":16867,"forks":3599,"topics":["airflow","argo","argo-workflows","batch-processing","cloud-native","cncf","dag","data-engineering","gitops","hacktoberfest","k8s","knative","kubernetes","machine-learning","mlops","pipelines","workflow","workflow-engine"],"archived":false,"github_pushed_at":"2026-07-31T17:55:58+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/argoproj-argo-workflows","markdown_url":"https://www.graphcanon.com/tools/argoproj-argo-workflows.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/argoproj-argo-workflows","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=argoproj-argo-workflows","shared_categories":["developer-tools"]},{"slug":"dagster-io-dagster","name":"dagster","tagline":"An orchestration platform for data assets","github_url":"https://github.com/dagster-io/dagster","owner":"dagster-io","repo":"dagster","owner_avatar_url":"https://avatars.githubusercontent.com/u/40032576?v=4","primary_language":"Python","stars":15949,"forks":2232,"topics":["analytics","dagster","data-engineering","data-integration","data-orchestrator","data-pipelines","data-science","etl","metadata","mlops","orchestration","python","scheduler","workflow","workflow-automation"],"archived":false,"github_pushed_at":"2026-08-09T07:33:07+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/dagster-io-dagster","markdown_url":"https://www.graphcanon.com/tools/dagster-io-dagster.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/dagster-io-dagster","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=dagster-io-dagster","shared_categories":["data-retrieval"]},{"slug":"fivetran-great-expectations","name":"great_expectations","tagline":"Always know what to expect from your data","github_url":"https://github.com/fivetran/great_expectations","owner":"fivetran","repo":"great_expectations","owner_avatar_url":"https://avatars.githubusercontent.com/u/2722259?v=4","primary_language":"Python","stars":11690,"forks":1790,"topics":["cleandata","data-engineering","data-profilers","data-profiling","data-quality","data-science","data-unit-tests","datacleaner","datacleaning","dataquality","dataunittest","eda","exploratory-analysis","exploratory-data-analysis","exploratorydataanalysis","mlops","pipeline","pipeline-debt","pipeline-testing","pipeline-tests"],"archived":false,"github_pushed_at":"2026-08-02T02:39:28+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/fivetran-great-expectations","markdown_url":"https://www.graphcanon.com/tools/fivetran-great-expectations.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/fivetran-great-expectations","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=fivetran-great-expectations","shared_categories":["data-retrieval"]},{"slug":"mage-ai-mage-ai","name":"mage-ai","tagline":"Build, run and manage data pipelines for integrating and transforming data","github_url":"https://github.com/mage-ai/mage-ai","owner":"mage-ai","repo":"mage-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/69371472?v=4","primary_language":"Python","stars":8790,"forks":982,"topics":["artificial-intelligence","data","data-engineering","data-integration","data-pipelines","data-science","dbt","elt","etl","machine-learning","orchestration","pipeline","pipelines","python","reverse-etl","spark","sql","transformation"],"archived":false,"github_pushed_at":"2026-08-10T23:12:25+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/mage-ai-mage-ai","markdown_url":"https://www.graphcanon.com/tools/mage-ai-mage-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mage-ai-mage-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mage-ai-mage-ai","shared_categories":["data-retrieval"]},{"slug":"datajuicer-data-juicer","name":"data-juicer","tagline":"Data processing for and with foundation models","github_url":"https://github.com/datajuicer/data-juicer","owner":"datajuicer","repo":"data-juicer","owner_avatar_url":"https://avatars.githubusercontent.com/u/223222708?v=4","primary_language":"Python","stars":6897,"forks":404,"topics":["data","data-analysis","data-pipeline","data-processing","data-science","data-visualization","foundation-models","instruction-tuning","large-language-models","llm","llms","multi-modal","pre-training","synthetic-data"],"archived":false,"github_pushed_at":"2026-08-13T09:19:31+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/datajuicer-data-juicer","markdown_url":"https://www.graphcanon.com/tools/datajuicer-data-juicer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datajuicer-data-juicer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datajuicer-data-juicer","shared_categories":["data-retrieval"]},{"slug":"pditommaso-awesome-pipeline","name":"awesome-pipeline","tagline":"Curated list of pipeline toolkits","github_url":"https://github.com/pditommaso/awesome-pipeline","owner":"pditommaso","repo":"awesome-pipeline","owner_avatar_url":"https://avatars.githubusercontent.com/u/816968?v=4","primary_language":null,"stars":6616,"forks":650,"topics":["awesome-list","workflow"],"archived":false,"github_pushed_at":"2026-08-04T07:45:32+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/pditommaso-awesome-pipeline","markdown_url":"https://www.graphcanon.com/tools/pditommaso-awesome-pipeline.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pditommaso-awesome-pipeline","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pditommaso-awesome-pipeline","shared_categories":["developer-tools"]},{"slug":"pachyderm-pachyderm","name":"pachyderm","tagline":"Data-Centric Pipelines and Data Versioning","github_url":"https://github.com/pachyderm/pachyderm","owner":"pachyderm","repo":"pachyderm","owner_avatar_url":"https://avatars.githubusercontent.com/u/10432478?v=4","primary_language":"Go","stars":6300,"forks":577,"topics":["analytics","big-data","containers","data-analysis","data-science","distributed-systems","docker","go","kubernetes","pachyderm"],"archived":false,"github_pushed_at":"2025-02-03T22:27:18+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/pachyderm-pachyderm","markdown_url":"https://www.graphcanon.com/tools/pachyderm-pachyderm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pachyderm-pachyderm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pachyderm-pachyderm","shared_categories":["developer-tools"]},{"slug":"lemony-ai-cascadeflow","name":"cascadeflow","tagline":"Optimized runtime for AI agents with cost and quality considerations.","github_url":"https://github.com/lemony-ai/cascadeflow","owner":"lemony-ai","repo":"cascadeflow","owner_avatar_url":"https://avatars.githubusercontent.com/u/169823043?v=4","primary_language":"Python","stars":4015,"forks":922,"topics":["agent","ai","anthropic","api","budgets","claude","cost-optimization","cost-transparency","google-adk","gpt","huggingface","llm","model-cascading","n8n","ollama","openai","python","together-ai","typescript","vllm"],"archived":false,"github_pushed_at":"2026-08-06T19:29:00+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/lemony-ai-cascadeflow","markdown_url":"https://www.graphcanon.com/tools/lemony-ai-cascadeflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lemony-ai-cascadeflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lemony-ai-cascadeflow","shared_categories":[]},{"slug":"ucbepic-docetl","name":"docetl","tagline":"A system for agentic LLM-powered data processing and ETL","github_url":"https://github.com/ucbepic/docetl","owner":"ucbepic","repo":"docetl","owner_avatar_url":"https://avatars.githubusercontent.com/u/88680502?v=4","primary_language":"Python","stars":3961,"forks":421,"topics":["agents","data","data-pipelines","document-analysis","document-processing","elt","etl","llm","python","semantic-data","unstructured-data","unstructured-data-analysis","workflow"],"archived":false,"github_pushed_at":"2026-08-09T23:31:04+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/ucbepic-docetl","markdown_url":"https://www.graphcanon.com/tools/ucbepic-docetl.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ucbepic-docetl","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ucbepic-docetl","shared_categories":["data-retrieval"]},{"slug":"ploomber-ploomber","name":"ploomber","tagline":"The fastest way to build data pipelines. Develop iteratively, deploy anywhere.","github_url":"https://github.com/ploomber/ploomber","owner":"ploomber","repo":"ploomber","owner_avatar_url":"https://avatars.githubusercontent.com/u/60114551?v=4","primary_language":"Python","stars":3622,"forks":243,"topics":["data-engineering","data-science","jupyter","jupyter-notebooks","machine-learning","mlops","notebooks","papermill","pipelines","pycharm","vscode","workflow"],"archived":true,"github_pushed_at":"2025-05-29T22:02:03+00:00","maintenance_label":"Archived","url":"https://www.graphcanon.com/tools/ploomber-ploomber","markdown_url":"https://www.graphcanon.com/tools/ploomber-ploomber.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ploomber-ploomber","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ploomber-ploomber","shared_categories":["developer-tools"]},{"slug":"huggingface-datatrove","name":"datatrove","tagline":"Platform-agnostic customizable pipeline processing blocks for data processing and transformation.","github_url":"https://github.com/huggingface/datatrove","owner":"huggingface","repo":"datatrove","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":3250,"forks":288,"topics":[],"archived":false,"github_pushed_at":"2026-08-06T15:27:26+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/huggingface-datatrove","markdown_url":"https://www.graphcanon.com/tools/huggingface-datatrove.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-datatrove","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-datatrove","shared_categories":["data-retrieval"]}]}}