{"data":{"node":{"slug":"sinaptik-ai-pandas-ai","name":"pandas-ai","tagline":"Chat with your database or your datalake using LLMs and RAG.","github_url":"https://github.com/sinaptik-ai/pandas-ai","owner":"sinaptik-ai","repo":"pandas-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/154438448?v=4","primary_language":"Python","stars":23746,"forks":2342,"topics":["ai","csv","data","data-analysis","data-science","data-visualization","database","datalake","gpt-4","llm","pandas","sql","text-to-sql"],"archived":false,"github_pushed_at":"2025-10-28T10:02:13+00:00","maintenance_label":"Slowing","stars_delta_30d":90,"url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai","markdown_url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sinaptik-ai-pandas-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sinaptik-ai-pandas-ai"},"categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"ai","name":"ai"},{"slug":"csv","name":"csv"},{"slug":"data-analysis","name":"data-analysis"},{"slug":"database","name":"database"},{"slug":"datalake","name":"datalake"},{"slug":"llm","name":"llm"},{"slug":"pandas","name":"pandas"},{"slug":"sql","name":"sql"}],"edges":[{"type":"integrates_with","direction":"out","explanation":"PandasAI leverages RAG and LLMs to enable conversational data analysis. LangChain is an agent engineering platform that might integrate with PandasAI for more advanced scenarios.","successor_context":null,"tool":{"slug":"langchain-ai-langchain","name":"langchain","tagline":"The agent engineering platform.","github_url":"https://github.com/langchain-ai/langchain","owner":"langchain-ai","repo":"langchain","owner_avatar_url":"https://avatars.githubusercontent.com/u/126733545?v=4","primary_language":"Python","stars":143615,"forks":23930,"topics":["agents","ai","ai-agents","anthropic","chatgpt","deepagents","enterprise","framework","gemini","generative-ai","langchain","langgraph","llm","multiagent","open-source","openai","pydantic","python","rag","typescript"],"archived":false,"github_pushed_at":"2026-08-07T08:27:07+00:00","maintenance_label":"Very active","stars_delta_30d":2337,"url":"https://www.graphcanon.com/tools/langchain-ai-langchain","markdown_url":"https://www.graphcanon.com/tools/langchain-ai-langchain.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/langchain-ai-langchain","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=langchain-ai-langchain"}},{"type":"related","direction":"out","explanation":"Both PandasAI and Parlant focus on building reliable AI agents that can interact with users via natural language, though they serve different contexts.","successor_context":null,"tool":{"slug":"emcie-co-parlant","name":"parlant","tagline":"Build reliable customer-facing AI agents with Parlant: an interaction control harness optimized for controlled, consistent, and predictable LLM interactions.","github_url":"https://github.com/emcie-co/parlant","owner":"emcie-co","repo":"parlant","owner_avatar_url":"https://avatars.githubusercontent.com/u/160175171?v=4","primary_language":"Python","stars":18253,"forks":1551,"topics":["ai-agents","ai-alignment","customer-service","customer-success","gemini","genai","hacktoberfest","llama3","llm","openai","python"],"archived":false,"github_pushed_at":"2026-07-12T19:40:55+00:00","maintenance_label":"Steady","stars_delta_30d":74,"url":"https://www.graphcanon.com/tools/emcie-co-parlant","markdown_url":"https://www.graphcanon.com/tools/emcie-co-parlant.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/emcie-co-parlant","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=emcie-co-parlant"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"integrates_with","direction":"out","explanation":"Pandas-ai integrates with litellm because litellm serves as a gateway to various Language Model APIs, which pandas-ai can utilize for its natural language processing tasks. This integration allows pandas-ai access to a wide range of LLMs through a standardized interface provided by litellm, enhancing its capabilities in conversational data analysis.","successor_context":null,"tool":{"slug":"berriai-litellm","name":"litellm","tagline":"Python SDK and Proxy Server for calling multiple LLM APIs","github_url":"https://github.com/BerriAI/litellm","owner":"BerriAI","repo":"litellm","owner_avatar_url":"https://avatars.githubusercontent.com/u/121462774?v=4","primary_language":"Python","stars":55221,"forks":10231,"topics":["ai-gateway","anthropic","azure-openai","bedrock","gateway","langchain","litellm","llm","llm-gateway","llmops","mcp-gateway","openai","openai-proxy","rust","rust-ai","vertex-ai"],"archived":false,"github_pushed_at":"2026-08-01T05:53:28+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/berriai-litellm","markdown_url":"https://www.graphcanon.com/tools/berriai-litellm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/berriai-litellm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=berriai-litellm"}},{"type":"integrates_with","direction":"out","explanation":"RAGFlow is a Retrieval-Augmented Generation engine that could be integrated with PandasAI to enhance its context management and retrieval capabilities.","successor_context":null,"tool":{"slug":"infiniflow-ragflow","name":"ragflow","tagline":"Retrieval-Augmented Generation engine with agent capabilities","github_url":"https://github.com/infiniflow/ragflow","owner":"infiniflow","repo":"ragflow","owner_avatar_url":"https://avatars.githubusercontent.com/u/69962740?v=4","primary_language":"Go","stars":86541,"forks":10167,"topics":["agent-harness","agentic-ai","agentic-retrieval","agentic-search","ai","ai-agents","context-engine","context-engineering","context-management","harness-engineering","knowledge-compilation","llm-apps","rag","retrieval-augmented-generation"],"archived":false,"github_pushed_at":"2026-07-31T14:59:12+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/infiniflow-ragflow","markdown_url":"https://www.graphcanon.com/tools/infiniflow-ragflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/infiniflow-ragflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=infiniflow-ragflow"}},{"type":"integrates_with","direction":"out","explanation":"PandasAI can potentially integrate with LLM applications ready for RAG and AI pipelines from PathwayCOM to enhance data processing workflows.","successor_context":null,"tool":{"slug":"pathwaycom-llm-app","name":"llm-app","tagline":"Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data.","github_url":"https://github.com/pathwaycom/llm-app","owner":"pathwaycom","repo":"llm-app","owner_avatar_url":"https://avatars.githubusercontent.com/u/25750857?v=4","primary_language":"Jupyter Notebook","stars":59037,"forks":1466,"topics":["chatbot","hugging-face","llm","llm-local","llm-prompting","llm-security","llmops","machine-learning","open-ai","pathway","rag","real-time","retrieval-augmented-generation","vector-database","vector-index"],"archived":false,"github_pushed_at":"2026-07-05T17:59:07+00:00","maintenance_label":"Steady","stars_delta_30d":11,"url":"https://www.graphcanon.com/tools/pathwaycom-llm-app","markdown_url":"https://www.graphcanon.com/tools/pathwaycom-llm-app.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pathwaycom-llm-app","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pathwaycom-llm-app"}},{"type":"related","direction":"out","explanation":"PandasAI and OpenLit both deal with enhancing the observability and monitoring of AI engineering systems but approach this goal from different angles.","successor_context":null,"tool":{"slug":"openlit-openlit","name":"openlit","tagline":"A comprehensive open-source platform for AI Engineering with LLM Observability, Monitoring, and Management","github_url":"https://github.com/openlit/openlit","owner":"openlit","repo":"openlit","owner_avatar_url":"https://avatars.githubusercontent.com/u/149867240?v=4","primary_language":"TypeScript","stars":2664,"forks":342,"topics":["ai-observability","amd-gpu","clickhouse","distributed-tracing","genai","gpu-monitoring","grafana","langchain","llmops","llms","metrics","monitoring-tool","nvidia-smi","observability","open-source","openai","opentelemetry","otlp","python","tracing"],"archived":false,"github_pushed_at":"2026-07-31T18:39:37+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/openlit-openlit","markdown_url":"https://www.graphcanon.com/tools/openlit-openlit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/openlit-openlit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=openlit-openlit"}},{"type":"related","direction":"in","explanation":"Both tools leverage LLMs to process structured data (Pandas-ai operates on databases and datalakes, whereas OpenDataLoader PDF handles PDF documents) for analysis.","successor_context":null,"tool":{"slug":"opendataloader-project-opendataloader-pdf","name":"opendataloader-pdf","tagline":"PDF Parser for AI-ready data","github_url":"https://github.com/opendataloader-project/opendataloader-pdf","owner":"opendataloader-project","repo":"opendataloader-pdf","owner_avatar_url":"https://avatars.githubusercontent.com/u/211280852?v=4","primary_language":"Java","stars":28528,"forks":2724,"topics":["a11y","accessibility","ai","bounding-box","document-parsing","eaa","html","json","markdown","ocr","ocr-recognition","pdf","pdf-accessibility","pdf-converter","pdf-extraction","pdf-parser","pdf-ua","rag","tables","tagged-pdf"],"archived":false,"github_pushed_at":"2026-08-18T04:01:45+00:00","maintenance_label":"Very active","stars_delta_30d":1078,"url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf","markdown_url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opendataloader-project-opendataloader-pdf","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opendataloader-project-opendataloader-pdf"}},{"type":"related","direction":"in","explanation":"Both WrenAI and pandas-ai use natural language processing to aid in data analysis, but they target different parts of the process: WrenAI focuses on generating BI insights across multiple databases, while pandas-ai is more about querying datasets using chat-based natural language.","successor_context":null,"tool":{"slug":"canner-wrenai","name":"WrenAI","tagline":"GenBI for AI agents, turns natural-language questions into trusted dashboards and SQL","github_url":"https://github.com/Canner/WrenAI","owner":"Canner","repo":"WrenAI","owner_avatar_url":"https://avatars.githubusercontent.com/u/7250217?v=4","primary_language":"Python","stars":17295,"forks":1955,"topics":["ai-agents","bigquery","business-intelligence","charts","clickhouse","context-engineering","dashboard","databricks","duckdb","genbi","generative-ai","llm","mcp","postgresql","rag","semantic-layer","snowflake","sql","text-to-sql","text2sql"],"archived":false,"github_pushed_at":"2026-08-18T05:19:19+00:00","maintenance_label":"Very active","stars_delta_30d":1328,"url":"https://www.graphcanon.com/tools/canner-wrenai","markdown_url":"https://www.graphcanon.com/tools/canner-wrenai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/canner-wrenai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=canner-wrenai"}},{"type":"integrates_with","direction":"in","explanation":"Both Data-Juicer and pandas-ai involve data processing for AI tasks with specific emphasis on making it accessible via LLMs.","successor_context":null,"tool":{"slug":"datajuicer-data-juicer","name":"data-juicer","tagline":"Data processing for and with foundation models","github_url":"https://github.com/datajuicer/data-juicer","owner":"datajuicer","repo":"data-juicer","owner_avatar_url":"https://avatars.githubusercontent.com/u/223222708?v=4","primary_language":"Python","stars":6897,"forks":404,"topics":["data","data-analysis","data-pipeline","data-processing","data-science","data-visualization","foundation-models","instruction-tuning","large-language-models","llm","llms","multi-modal","pre-training","synthetic-data"],"archived":false,"github_pushed_at":"2026-08-13T09:19:31+00:00","maintenance_label":"Very active","stars_delta_30d":166,"url":"https://www.graphcanon.com/tools/datajuicer-data-juicer","markdown_url":"https://www.graphcanon.com/tools/datajuicer-data-juicer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datajuicer-data-juicer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datajuicer-data-juicer"}},{"type":"integrates_with","direction":"in","explanation":"Promptify can be used with pandas-ai for chat-based data analysis, which requires well-crafted prompts and outputs in a structured format to communicate effectively with the dataset.","successor_context":null,"tool":{"slug":"promptslab-promptify","name":"Promptify","tagline":"Task-based NLP engine with Pydantic structured outputs","github_url":"https://github.com/promptslab/Promptify","owner":"promptslab","repo":"Promptify","owner_avatar_url":"https://avatars.githubusercontent.com/u/120981762?v=4","primary_language":"Python","stars":4630,"forks":363,"topics":["chatgpt","chatgpt-api","chatgpt-python","gpt-3","gpt-3-prompts","gpt-4","gpt-4-api","gpt3-library","large-language-models","machine-learning","nlp","openai","prompt-engineering","prompt-toolkit","prompt-tuning","prompt-versioning","prompting","prompts","promptversioning","transformers"],"archived":false,"github_pushed_at":"2026-03-27T01:09:45+00:00","maintenance_label":"Slowing","stars_delta_30d":11,"url":"https://www.graphcanon.com/tools/promptslab-promptify","markdown_url":"https://www.graphcanon.com/tools/promptslab-promptify.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptslab-promptify","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptslab-promptify"}},{"type":"related","direction":"in","explanation":"Both Apache Hamilton and pandas-ai are tools that leverage Python for data transformation tasks, although pandas-ai integrates LLMs for conversational data analysis.","successor_context":null,"tool":{"slug":"apache-hamilton","name":"hamilton","tagline":"Modular dataflow definition for Python environments","github_url":"https://github.com/apache/hamilton","owner":"apache","repo":"hamilton","owner_avatar_url":"https://avatars.githubusercontent.com/u/47359?v=4","primary_language":"Jupyter Notebook","stars":2557,"forks":203,"topics":["dag","data-analysis","data-engineering","data-science","dataframe","etl","etl-framework","etl-pipeline","feature-engineering","hacktoberfest","lineage","llmops","machine-learning","mlops","orchestration","pandas","python","rag","software-engineering"],"archived":false,"github_pushed_at":"2026-08-01T18:45:37+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/apache-hamilton","markdown_url":"https://www.graphcanon.com/tools/apache-hamilton.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/apache-hamilton","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=apache-hamilton"}},{"type":"related","direction":"in","explanation":"Neum AI and pandas-ai both involve conversational interfaces driven by LLMs for interacting with data or datalakes but serve different use cases.","successor_context":null,"tool":{"slug":"neumtry-neumai","name":"NeumAI","tagline":"Framework to manage creation and synchronization of vector embeddings at large scale","github_url":"https://github.com/NeumTry/NeumAI","owner":"NeumTry","repo":"NeumAI","owner_avatar_url":"https://avatars.githubusercontent.com/u/129831068?v=4","primary_language":"Python","stars":867,"forks":50,"topics":["ai","chatgpt","data","data-engineering","database","embeddings","etl","llm","llmops","mlops","ops","pipeline","python","rag","retrieval","vector-database","vectors"],"archived":false,"github_pushed_at":"2024-01-15T23:00:58+00:00","maintenance_label":"Dormant","stars_delta_30d":3,"url":"https://www.graphcanon.com/tools/neumtry-neumai","markdown_url":"https://www.graphcanon.com/tools/neumtry-neumai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/neumtry-neumai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=neumtry-neumai"}}],"neighbours":[{"slug":"pathwaycom-llm-app","name":"llm-app","tagline":"Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data.","github_url":"https://github.com/pathwaycom/llm-app","owner":"pathwaycom","repo":"llm-app","owner_avatar_url":"https://avatars.githubusercontent.com/u/25750857?v=4","primary_language":"Jupyter Notebook","stars":59037,"forks":1466,"topics":["chatbot","hugging-face","llm","llm-local","llm-prompting","llm-security","llmops","machine-learning","open-ai","pathway","rag","real-time","retrieval-augmented-generation","vector-database","vector-index"],"archived":false,"github_pushed_at":"2026-07-05T17:59:07+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/pathwaycom-llm-app","markdown_url":"https://www.graphcanon.com/tools/pathwaycom-llm-app.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pathwaycom-llm-app","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pathwaycom-llm-app","shared_categories":["data-retrieval","llm-frameworks"]},{"slug":"huggingface-datasets","name":"datasets","tagline":"Largest hub of ready-to-use datasets for AI models","github_url":"https://github.com/huggingface/datasets","owner":"huggingface","repo":"datasets","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":21791,"forks":3322,"topics":["ai","artificial-intelligence","computer-vision","dataset-hub","datasets","deep-learning","huggingface","llm","machine-learning","natural-language-processing","nlp","numpy","pandas","pytorch","speech","tensorflow"],"archived":false,"github_pushed_at":"2026-07-30T11:23:49+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/huggingface-datasets","markdown_url":"https://www.graphcanon.com/tools/huggingface-datasets.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-datasets","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-datasets","shared_categories":["data-retrieval"]},{"slug":"canner-wrenai","name":"WrenAI","tagline":"GenBI for AI agents, turns natural-language questions into trusted dashboards and SQL","github_url":"https://github.com/Canner/WrenAI","owner":"Canner","repo":"WrenAI","owner_avatar_url":"https://avatars.githubusercontent.com/u/7250217?v=4","primary_language":"Python","stars":17295,"forks":1955,"topics":["ai-agents","bigquery","business-intelligence","charts","clickhouse","context-engineering","dashboard","databricks","duckdb","genbi","generative-ai","llm","mcp","postgresql","rag","semantic-layer","snowflake","sql","text-to-sql","text2sql"],"archived":false,"github_pushed_at":"2026-08-18T05:19:19+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/canner-wrenai","markdown_url":"https://www.graphcanon.com/tools/canner-wrenai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/canner-wrenai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=canner-wrenai","shared_categories":["data-retrieval"]},{"slug":"fivetran-great-expectations","name":"great_expectations","tagline":"Always know what to expect from your data","github_url":"https://github.com/fivetran/great_expectations","owner":"fivetran","repo":"great_expectations","owner_avatar_url":"https://avatars.githubusercontent.com/u/2722259?v=4","primary_language":"Python","stars":11690,"forks":1790,"topics":["cleandata","data-engineering","data-profilers","data-profiling","data-quality","data-science","data-unit-tests","datacleaner","datacleaning","dataquality","dataunittest","eda","exploratory-analysis","exploratory-data-analysis","exploratorydataanalysis","mlops","pipeline","pipeline-debt","pipeline-testing","pipeline-tests"],"archived":false,"github_pushed_at":"2026-08-02T02:39:28+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/fivetran-great-expectations","markdown_url":"https://www.graphcanon.com/tools/fivetran-great-expectations.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/fivetran-great-expectations","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=fivetran-great-expectations","shared_categories":["data-retrieval"]},{"slug":"databendlabs-databend","name":"databend","tagline":"All-in-One Data Warehouse: Analytics, Search, AI, and Python Sandboxing Reimagined From Scratch.","github_url":"https://github.com/databendlabs/databend","owner":"databendlabs","repo":"databend","owner_avatar_url":"https://avatars.githubusercontent.com/u/80994548?v=4","primary_language":"Rust","stars":9420,"forks":891,"topics":["ai","bigdata","cloud-native","database","elasticsearch","geospatial","lakehouse","olap","rust","serverless","snowflake","sql","vector-database","vector-search"],"archived":false,"github_pushed_at":"2026-08-21T05:00:06+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/databendlabs-databend","markdown_url":"https://www.graphcanon.com/tools/databendlabs-databend.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/databendlabs-databend","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=databendlabs-databend","shared_categories":["data-retrieval"]},{"slug":"mage-ai-mage-ai","name":"mage-ai","tagline":"Build, run and manage data pipelines for integrating and transforming data","github_url":"https://github.com/mage-ai/mage-ai","owner":"mage-ai","repo":"mage-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/69371472?v=4","primary_language":"Python","stars":8790,"forks":982,"topics":["artificial-intelligence","data","data-engineering","data-integration","data-pipelines","data-science","dbt","elt","etl","machine-learning","orchestration","pipeline","pipelines","python","reverse-etl","spark","sql","transformation"],"archived":false,"github_pushed_at":"2026-08-10T23:12:25+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/mage-ai-mage-ai","markdown_url":"https://www.graphcanon.com/tools/mage-ai-mage-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mage-ai-mage-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mage-ai-mage-ai","shared_categories":["data-retrieval"]},{"slug":"alteryx-featuretools","name":"featuretools","tagline":"An open source python library for automated feature engineering","github_url":"https://github.com/alteryx/featuretools","owner":"alteryx","repo":"featuretools","owner_avatar_url":"https://avatars.githubusercontent.com/u/12972388?v=4","primary_language":"Python","stars":7665,"forks":915,"topics":["automated-feature-engineering","automated-machine-learning","automl","data-science","feature-engineering","machine-learning","python","scikit-learn"],"archived":false,"github_pushed_at":"2026-07-27T13:04:38+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/alteryx-featuretools","markdown_url":"https://www.graphcanon.com/tools/alteryx-featuretools.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alteryx-featuretools","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alteryx-featuretools","shared_categories":[]},{"slug":"datajuicer-data-juicer","name":"data-juicer","tagline":"Data processing for and with foundation models","github_url":"https://github.com/datajuicer/data-juicer","owner":"datajuicer","repo":"data-juicer","owner_avatar_url":"https://avatars.githubusercontent.com/u/223222708?v=4","primary_language":"Python","stars":6897,"forks":404,"topics":["data","data-analysis","data-pipeline","data-processing","data-science","data-visualization","foundation-models","instruction-tuning","large-language-models","llm","llms","multi-modal","pre-training","synthetic-data"],"archived":false,"github_pushed_at":"2026-08-13T09:19:31+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/datajuicer-data-juicer","markdown_url":"https://www.graphcanon.com/tools/datajuicer-data-juicer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datajuicer-data-juicer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datajuicer-data-juicer","shared_categories":["data-retrieval"]},{"slug":"run-llama-rags","name":"rags","tagline":"Build ChatGPT over your data with natural language","github_url":"https://github.com/run-llama/rags","owner":"run-llama","repo":"rags","owner_avatar_url":"https://avatars.githubusercontent.com/u/130722866?v=4","primary_language":"Python","stars":6549,"forks":656,"topics":["agent","chatbot","chatgpt","gpts","llamaindex","llm","openai","rag","streamlit"],"archived":false,"github_pushed_at":"2024-04-05T05:36:59+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/run-llama-rags","markdown_url":"https://www.graphcanon.com/tools/run-llama-rags.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/run-llama-rags","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=run-llama-rags","shared_categories":["data-retrieval"]},{"slug":"ucbepic-docetl","name":"docetl","tagline":"A system for agentic LLM-powered data processing and ETL","github_url":"https://github.com/ucbepic/docetl","owner":"ucbepic","repo":"docetl","owner_avatar_url":"https://avatars.githubusercontent.com/u/88680502?v=4","primary_language":"Python","stars":3961,"forks":421,"topics":["agents","data","data-pipelines","document-analysis","document-processing","elt","etl","llm","python","semantic-data","unstructured-data","unstructured-data-analysis","workflow"],"archived":false,"github_pushed_at":"2026-08-09T23:31:04+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/ucbepic-docetl","markdown_url":"https://www.graphcanon.com/tools/ucbepic-docetl.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ucbepic-docetl","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ucbepic-docetl","shared_categories":["data-retrieval"]},{"slug":"ploomber-ploomber","name":"ploomber","tagline":"The fastest way to build data pipelines. Develop iteratively, deploy anywhere.","github_url":"https://github.com/ploomber/ploomber","owner":"ploomber","repo":"ploomber","owner_avatar_url":"https://avatars.githubusercontent.com/u/60114551?v=4","primary_language":"Python","stars":3622,"forks":243,"topics":["data-engineering","data-science","jupyter","jupyter-notebooks","machine-learning","mlops","notebooks","papermill","pipelines","pycharm","vscode","workflow"],"archived":true,"github_pushed_at":"2025-05-29T22:02:03+00:00","maintenance_label":"Archived","url":"https://www.graphcanon.com/tools/ploomber-ploomber","markdown_url":"https://www.graphcanon.com/tools/ploomber-ploomber.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ploomber-ploomber","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ploomber-ploomber","shared_categories":[]},{"slug":"spiceai-spiceai","name":"spiceai","tagline":"A real-time analytics node for data-grounded AI applications","github_url":"https://github.com/spiceai/spiceai","owner":"spiceai","repo":"spiceai","owner_avatar_url":"https://avatars.githubusercontent.com/u/73862742?v=4","primary_language":"Rust","stars":3047,"forks":212,"topics":["artificial-intelligence","data","data-federation","developers","full-text-search","infrastructure","llm-inference","machine-learning","sql"],"archived":false,"github_pushed_at":"2026-07-25T05:47:50+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/spiceai-spiceai","markdown_url":"https://www.graphcanon.com/tools/spiceai-spiceai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/spiceai-spiceai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=spiceai-spiceai","shared_categories":["data-retrieval","llm-frameworks"]}]}}