{"data":{"node":{"slug":"adbar-trafilatura","name":"trafilatura","tagline":"Python & Command-line tool for web crawling, scraping and text extraction","github_url":"https://github.com/adbar/trafilatura","owner":"adbar","repo":"trafilatura","owner_avatar_url":"https://avatars.githubusercontent.com/u/2125866?v=4","primary_language":"Python","stars":6657,"forks":415,"topics":["article-extractor","corpus-builder","corpus-tools","crawler","html-to-markdown","html2text","llm","news-aggregator","news-crawler","nlp","rag","readability","rss-feed","scraping","tei","text-cleaning","text-extraction","text-mining","text-preprocessing","web-scraping"],"archived":false,"github_pushed_at":"2026-08-15T16:13:21+00:00","maintenance_label":"Very active","stars_delta_30d":343,"url":"https://www.graphcanon.com/tools/adbar-trafilatura","markdown_url":"https://www.graphcanon.com/tools/adbar-trafilatura.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/adbar-trafilatura","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=adbar-trafilatura"},"categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"}],"tags":[{"slug":"article-extractor","name":"article-extractor"},{"slug":"corpus-builder","name":"corpus-builder"},{"slug":"text-mining","name":"text-mining"},{"slug":"web-scraping","name":"web-scraping"}],"edges":[{"type":"alternative","direction":"out","explanation":"Both tools are designed for web data extraction, but Trafilatura focuses on text and metadata extraction using a more traditional approach, while Scrapegraph-AI leverages AI techniques.","successor_context":null,"tool":{"slug":"scrapegraphai-scrapegraph-ai","name":"Scrapegraph-ai","tagline":"Python scraper based on AI","github_url":"https://github.com/ScrapeGraphAI/Scrapegraph-ai","owner":"ScrapeGraphAI","repo":"Scrapegraph-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/171017415?v=4","primary_language":"Python","stars":29618,"forks":2925,"topics":["ai-crawler","ai-scraping","ai-search","crawler","data-extraction","firecrawl-alternative","large-language-model","llm","markdown","rag","scraping","scraping-python","web-crawler","web-crawlers","web-data","web-data-extraction","web-scraper","web-scraping","web-search","webscraping"],"archived":false,"github_pushed_at":"2026-07-20T14:22:20+00:00","maintenance_label":"Active","stars_delta_30d":1203,"url":"https://www.graphcanon.com/tools/scrapegraphai-scrapegraph-ai","markdown_url":"https://www.graphcanon.com/tools/scrapegraphai-scrapegraph-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/scrapegraphai-scrapegraph-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=scrapegraphai-scrapegraph-ai"}},{"type":"related","direction":"out","explanation":"Headroom compresses context for AI agents which could be related to how trafilatura extracts and cleans text content by removing noise from HTML documents.","successor_context":null,"tool":{"slug":"headroomlabs-ai-headroom","name":"headroom","tagline":"Compress tool outputs and data to reduce tokens before reaching the LLM.","github_url":"https://github.com/headroomlabs-ai/headroom","owner":"headroomlabs-ai","repo":"headroom","owner_avatar_url":"https://avatars.githubusercontent.com/u/294291659?v=4","primary_language":"Python","stars":66470,"forks":5103,"topics":["agent","ai","anthropic","claude-code","compression","context-engineering","context-window","cursor","fastapi","langchain","llm","mcp","openai","prompt-engineering","proxy","python","rag","token-optimization","tokens","typescript"],"archived":false,"github_pushed_at":"2026-08-16T00:57:42+00:00","maintenance_label":"Very active","stars_delta_30d":6941,"url":"https://www.graphcanon.com/tools/headroomlabs-ai-headroom","markdown_url":"https://www.graphcanon.com/tools/headroomlabs-ai-headroom.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/headroomlabs-ai-headroom","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=headroomlabs-ai-headroom"}},{"type":"alternative","direction":"out","explanation":"Trafilatura and chatWeb both deal with web content extraction. However, chatWeb extends beyond simple extraction to include summarization and question-answering capabilities.","successor_context":null,"tool":{"slug":"skywalkerdarren-chatweb","name":"chatWeb","tagline":"ChatWeb can crawl web pages and various document types for content extraction and summarization.","github_url":"https://github.com/SkywalkerDarren/chatWeb","owner":"SkywalkerDarren","repo":"chatWeb","owner_avatar_url":"https://avatars.githubusercontent.com/u/20706299?v=4","primary_language":"Python","stars":916,"forks":137,"topics":["ai","chatgpt","crawler","docx","embedding","faiss","gpt","gpt-35-turbo","news-extractor","newspaper","openai","pdf","pgvector","postgresql","vector-database"],"archived":false,"github_pushed_at":"2026-05-25T16:56:25+00:00","maintenance_label":"Steady","stars_delta_30d":0,"url":"https://www.graphcanon.com/tools/skywalkerdarren-chatweb","markdown_url":"https://www.graphcanon.com/tools/skywalkerdarren-chatweb.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/skywalkerdarren-chatweb","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=skywalkerdarren-chatweb"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"graphify-labs-graphify","name":"graphify","tagline":"Turn any code or documentation into a queryable knowledge graph","github_url":"https://github.com/Graphify-Labs/graphify","owner":"Graphify-Labs","repo":"graphify","owner_avatar_url":"https://avatars.githubusercontent.com/u/297659074?v=4","primary_language":"Python","stars":107507,"forks":10441,"topics":["ai-agents","antigravity","ast","claude-code","code-analysis","code-search","codex","cursor","developer-tools","gemini","graphrag","knowledge-graph","leiden","llm","mcp","openclaw","rag","skills","tree-sitter"],"archived":false,"github_pushed_at":"2026-08-17T18:42:58+00:00","maintenance_label":"Very active","stars_delta_30d":16918,"url":"https://www.graphcanon.com/tools/graphify-labs-graphify","markdown_url":"https://www.graphcanon.com/tools/graphify-labs-graphify.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/graphify-labs-graphify","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=graphify-labs-graphify"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"langchain-ai-langgraph","name":"langgraph","tagline":"Low-level orchestration framework for building stateful agents.","github_url":"https://github.com/langchain-ai/langgraph","owner":"langchain-ai","repo":"langgraph","owner_avatar_url":"https://avatars.githubusercontent.com/u/126733545?v=4","primary_language":"Python","stars":38352,"forks":6458,"topics":["agents","ai","ai-agents","chatgpt","deepagents","enterprise","framework","gemini","generative-ai","langchain","langgraph","llm","multiagent","open-source","openai","pydantic","python","rag"],"archived":false,"github_pushed_at":"2026-07-28T14:49:41+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/langchain-ai-langgraph","markdown_url":"https://www.graphcanon.com/tools/langchain-ai-langgraph.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/langchain-ai-langgraph","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=langchain-ai-langgraph"}},{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"chatchat-space-langchain-chatchat","name":"Langchain-Chatchat","tagline":"Local knowledge-based RAG and Agent app using Langchain and various LLMs","github_url":"https://github.com/chatchat-space/Langchain-Chatchat","owner":"chatchat-space","repo":"Langchain-Chatchat","owner_avatar_url":"https://avatars.githubusercontent.com/u/139558948?v=4","primary_language":"Python","stars":38522,"forks":6266,"topics":["chatbot","chatchat","chatglm","chatgpt","embedding","faiss","fastchat","gpt","knowledge-base","langchain","langchain-chatglm","llama","llm","milvus","ollama","qwen","rag","retrieval-augmented-generation","streamlit","xinference"],"archived":false,"github_pushed_at":"2025-11-10T09:27:42+00:00","maintenance_label":"Slowing","stars_delta_30d":254,"url":"https://www.graphcanon.com/tools/chatchat-space-langchain-chatchat","markdown_url":"https://www.graphcanon.com/tools/chatchat-space-langchain-chatchat.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/chatchat-space-langchain-chatchat","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=chatchat-space-langchain-chatchat"}},{"type":"integrates_with","direction":"out","explanation":"Browser-Use automates tasks online with AI interaction; it could integrate with trafilatura for more automated web scraping and extraction processes.","successor_context":null,"tool":{"slug":"browser-use-browser-use","name":"browser-use","tagline":"Make websites accessible for AI agents. Automate tasks online with ease.","github_url":"https://github.com/browser-use/browser-use","owner":"browser-use","repo":"browser-use","owner_avatar_url":"https://avatars.githubusercontent.com/u/192012301?v=4","primary_language":"Python","stars":109348,"forks":12022,"topics":["ai-agents","ai-tools","browser-automation","browser-use","llm","playwright","python"],"archived":false,"github_pushed_at":"2026-08-15T17:07:06+00:00","maintenance_label":"Very active","stars_delta_30d":4268,"url":"https://www.graphcanon.com/tools/browser-use-browser-use","markdown_url":"https://www.graphcanon.com/tools/browser-use-browser-use.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/browser-use-browser-use","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=browser-use-browser-use"}},{"type":"related","direction":"out","explanation":"Daytona and trafilatura both deal with running AI tasks but in different contexts; Daytona is for infrastructure to run AI-generated code, while trafilatura focuses on extracting data from the web.","successor_context":null,"tool":{"slug":"daytonaio-daytona","name":"daytona","tagline":"Secure and Elastic Infrastructure for Running AI-Generated Code","github_url":"https://github.com/daytonaio/daytona","owner":"daytonaio","repo":"daytona","owner_avatar_url":"https://avatars.githubusercontent.com/u/130513197?v=4","primary_language":null,"stars":71966,"forks":5653,"topics":["agentic-workflow","ai","ai-agents","ai-runtime","ai-sandboxes","code-execution","code-interpreter","developer-tools"],"archived":false,"github_pushed_at":"2026-07-24T07:12:07+00:00","maintenance_label":"Active","stars_delta_30d":-278,"url":"https://www.graphcanon.com/tools/daytonaio-daytona","markdown_url":"https://www.graphcanon.com/tools/daytonaio-daytona.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/daytonaio-daytona","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=daytonaio-daytona"}},{"type":"alternative","direction":"in","explanation":"Both ScrapeGraphAI and trafilatura are tools designed for extracting text and metadata from the web, indicating an alternative approach to similar problems.","successor_context":null,"tool":{"slug":"scrapegraphai-scrapegraph-ai","name":"Scrapegraph-ai","tagline":"Python scraper based on AI","github_url":"https://github.com/ScrapeGraphAI/Scrapegraph-ai","owner":"ScrapeGraphAI","repo":"Scrapegraph-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/171017415?v=4","primary_language":"Python","stars":29618,"forks":2925,"topics":["ai-crawler","ai-scraping","ai-search","crawler","data-extraction","firecrawl-alternative","large-language-model","llm","markdown","rag","scraping","scraping-python","web-crawler","web-crawlers","web-data","web-data-extraction","web-scraper","web-scraping","web-search","webscraping"],"archived":false,"github_pushed_at":"2026-07-20T14:22:20+00:00","maintenance_label":"Active","stars_delta_30d":1203,"url":"https://www.graphcanon.com/tools/scrapegraphai-scrapegraph-ai","markdown_url":"https://www.graphcanon.com/tools/scrapegraphai-scrapegraph-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/scrapegraphai-scrapegraph-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=scrapegraphai-scrapegraph-ai"}},{"type":"alternative","direction":"in","explanation":"Both 'Scrapling' and 'trafilatura' provide tools for web scraping and text extraction, with different implementations and focus areas.","successor_context":null,"tool":{"slug":"d4vinci-scrapling","name":"Scrapling","tagline":"An adaptive Web Scraping framework","github_url":"https://github.com/D4Vinci/Scrapling","owner":"D4Vinci","repo":"Scrapling","owner_avatar_url":"https://avatars.githubusercontent.com/u/20604835?v=4","primary_language":"Python","stars":71247,"forks":7067,"topics":["ai","ai-scraping","automation","crawler","crawling","crawling-python","data","data-extraction","mcp","mcp-server","playwright","python","scraping","selectors","stealth","web-scraper","web-scraping","web-scraping-python","webscraping","xpath"],"archived":false,"github_pushed_at":"2026-07-25T16:07:12+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/d4vinci-scrapling","markdown_url":"https://www.graphcanon.com/tools/d4vinci-scrapling.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/d4vinci-scrapling","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=d4vinci-scrapling"}}],"neighbours":[{"slug":"graphify-labs-graphify","name":"graphify","tagline":"Turn any code or documentation into a queryable knowledge graph","github_url":"https://github.com/Graphify-Labs/graphify","owner":"Graphify-Labs","repo":"graphify","owner_avatar_url":"https://avatars.githubusercontent.com/u/297659074?v=4","primary_language":"Python","stars":107507,"forks":10441,"topics":["ai-agents","antigravity","ast","claude-code","code-analysis","code-search","codex","cursor","developer-tools","gemini","graphrag","knowledge-graph","leiden","llm","mcp","openclaw","rag","skills","tree-sitter"],"archived":false,"github_pushed_at":"2026-08-17T18:42:58+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/graphify-labs-graphify","markdown_url":"https://www.graphcanon.com/tools/graphify-labs-graphify.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/graphify-labs-graphify","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=graphify-labs-graphify","shared_categories":["data-retrieval"]},{"slug":"d4vinci-scrapling","name":"Scrapling","tagline":"An adaptive Web Scraping framework","github_url":"https://github.com/D4Vinci/Scrapling","owner":"D4Vinci","repo":"Scrapling","owner_avatar_url":"https://avatars.githubusercontent.com/u/20604835?v=4","primary_language":"Python","stars":71247,"forks":7067,"topics":["ai","ai-scraping","automation","crawler","crawling","crawling-python","data","data-extraction","mcp","mcp-server","playwright","python","scraping","selectors","stealth","web-scraper","web-scraping","web-scraping-python","webscraping","xpath"],"archived":false,"github_pushed_at":"2026-07-25T16:07:12+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/d4vinci-scrapling","markdown_url":"https://www.graphcanon.com/tools/d4vinci-scrapling.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/d4vinci-scrapling","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=d4vinci-scrapling","shared_categories":["data-retrieval"]},{"slug":"scrapegraphai-scrapegraph-ai","name":"Scrapegraph-ai","tagline":"Python scraper based on AI","github_url":"https://github.com/ScrapeGraphAI/Scrapegraph-ai","owner":"ScrapeGraphAI","repo":"Scrapegraph-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/171017415?v=4","primary_language":"Python","stars":29618,"forks":2925,"topics":["ai-crawler","ai-scraping","ai-search","crawler","data-extraction","firecrawl-alternative","large-language-model","llm","markdown","rag","scraping","scraping-python","web-crawler","web-crawlers","web-data","web-data-extraction","web-scraper","web-scraping","web-search","webscraping"],"archived":false,"github_pushed_at":"2026-07-20T14:22:20+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/scrapegraphai-scrapegraph-ai","markdown_url":"https://www.graphcanon.com/tools/scrapegraphai-scrapegraph-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/scrapegraphai-scrapegraph-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=scrapegraphai-scrapegraph-ai","shared_categories":["data-retrieval"]},{"slug":"opendataloader-project-opendataloader-pdf","name":"opendataloader-pdf","tagline":"PDF Parser for AI-ready data","github_url":"https://github.com/opendataloader-project/opendataloader-pdf","owner":"opendataloader-project","repo":"opendataloader-pdf","owner_avatar_url":"https://avatars.githubusercontent.com/u/211280852?v=4","primary_language":"Java","stars":28528,"forks":2724,"topics":["a11y","accessibility","ai","bounding-box","document-parsing","eaa","html","json","markdown","ocr","ocr-recognition","pdf","pdf-accessibility","pdf-converter","pdf-extraction","pdf-parser","pdf-ua","rag","tables","tagged-pdf"],"archived":false,"github_pushed_at":"2026-08-18T04:01:45+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf","markdown_url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opendataloader-project-opendataloader-pdf","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opendataloader-project-opendataloader-pdf","shared_categories":["data-retrieval"]},{"slug":"unstructured-io-unstructured","name":"unstructured","tagline":"Convert documents to structured data effortlessly","github_url":"https://github.com/Unstructured-IO/unstructured","owner":"Unstructured-IO","repo":"unstructured","owner_avatar_url":"https://avatars.githubusercontent.com/u/108372208?v=4","primary_language":"HTML","stars":15238,"forks":1284,"topics":["data-pipelines","deep-learning","document-image-analysis","document-image-processing","document-parser","document-parsing","docx","donut","information-retrieval","langchain","llm","machine-learning","ml","natural-language-processing","nlp","ocr","pdf","pdf-to-json","pdf-to-text","preprocessing"],"archived":false,"github_pushed_at":"2026-07-31T20:54:17+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/unstructured-io-unstructured","markdown_url":"https://www.graphcanon.com/tools/unstructured-io-unstructured.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unstructured-io-unstructured","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unstructured-io-unstructured","shared_categories":["data-retrieval"]},{"slug":"yusufkaraaslan-skill-seekers","name":"Skill_Seekers","tagline":"Automation tool for converting documentation and code into Claude AI skills","github_url":"https://github.com/yusufkaraaslan/Skill_Seekers","owner":"yusufkaraaslan","repo":"Skill_Seekers","owner_avatar_url":"https://avatars.githubusercontent.com/u/11597362?v=4","primary_language":"Python","stars":14565,"forks":1477,"topics":["ai-tools","ast-parser","automation","claude-ai","claude-skills","code-analysis","conflict-detection","documentation","documentation-generator","github","github-scraper","mcp","mcp-server","multi-source","ocr","pdf","python","web-scraping"],"archived":false,"github_pushed_at":"2026-07-20T13:09:44+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/yusufkaraaslan-skill-seekers","markdown_url":"https://www.graphcanon.com/tools/yusufkaraaslan-skill-seekers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/yusufkaraaslan-skill-seekers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=yusufkaraaslan-skill-seekers","shared_categories":[]},{"slug":"ntegrals-openbrowser","name":"openbrowser","tagline":"AI-powered autonomous web browsing framework","github_url":"https://github.com/ntegrals/openbrowser","owner":"ntegrals","repo":"openbrowser","owner_avatar_url":"https://avatars.githubusercontent.com/u/26648900?v=4","primary_language":"TypeScript","stars":9510,"forks":864,"topics":["ai-agents","automation","claude","playwright","puppeteer","sandbox"],"archived":false,"github_pushed_at":"2026-04-02T12:55:42+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/ntegrals-openbrowser","markdown_url":"https://www.graphcanon.com/tools/ntegrals-openbrowser.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ntegrals-openbrowser","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ntegrals-openbrowser","shared_categories":[]},{"slug":"alirezamika-autoscraper","name":"autoscraper","tagline":"A Smart Automatic Fast and Lightweight Web Scraper for Python","github_url":"https://github.com/alirezamika/autoscraper","owner":"alirezamika","repo":"autoscraper","owner_avatar_url":"https://avatars.githubusercontent.com/u/17881612?v=4","primary_language":"Python","stars":7839,"forks":808,"topics":["ai","artificial-intelligence","automation","crawler","machine-learning","python","scrape","scraper","scraping","web-scraping","webautomation","webscraping"],"archived":false,"github_pushed_at":"2026-07-29T17:17:55+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/alirezamika-autoscraper","markdown_url":"https://www.graphcanon.com/tools/alirezamika-autoscraper.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alirezamika-autoscraper","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alirezamika-autoscraper","shared_categories":[]},{"slug":"firecrawl-firecrawl-mcp-server","name":"firecrawl-mcp-server","tagline":"Adding web scraping and search capabilities to LLM clients like Cursor and Claude","github_url":"https://github.com/firecrawl/firecrawl-mcp-server","owner":"firecrawl","repo":"firecrawl-mcp-server","owner_avatar_url":"https://avatars.githubusercontent.com/u/135057108?v=4","primary_language":"JavaScript","stars":7047,"forks":821,"topics":["batch-processing","claude","content-extraction","data-collection","firecrawl","firecrawl-ai","javascript-rendering","llm-tools","mcp","mcp-server","model-context-protocol","search-api","web-crawler","web-scraping"],"archived":false,"github_pushed_at":"2026-07-26T06:47:42+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/firecrawl-firecrawl-mcp-server","markdown_url":"https://www.graphcanon.com/tools/firecrawl-firecrawl-mcp-server.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/firecrawl-firecrawl-mcp-server","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=firecrawl-firecrawl-mcp-server","shared_categories":["data-retrieval"]},{"slug":"zipstack-unstract","name":"unstract","tagline":"LLM-Driven Extraction of Unstructured Data for API Deployments and ETL Pipeline Workflows","github_url":"https://github.com/Zipstack/unstract","owner":"Zipstack","repo":"unstract","owner_avatar_url":"https://avatars.githubusercontent.com/u/89070934?v=4","primary_language":"Python","stars":6932,"forks":663,"topics":["ai-agents","data-engineering","document-ai","generative-ai","idp","json-extraction","llm","mcp-server","ocr","pdf-extraction","prompt-engineering","structured-output"],"archived":false,"github_pushed_at":"2026-07-27T22:23:41+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/zipstack-unstract","markdown_url":"https://www.graphcanon.com/tools/zipstack-unstract.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/zipstack-unstract","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=zipstack-unstract","shared_categories":["data-retrieval"]},{"slug":"exa-labs-exa-mcp-server","name":"exa-mcp-server","tagline":"Platform for web search and crawling using Model Context Protocol.","github_url":"https://github.com/exa-labs/exa-mcp-server","owner":"exa-labs","repo":"exa-mcp-server","owner_avatar_url":"https://avatars.githubusercontent.com/u/77906174?v=4","primary_language":"TypeScript","stars":4777,"forks":362,"topics":["code-search","codesearch","crawling","mcp","mcp-server","model-context-protocol","web-search","websearch"],"archived":false,"github_pushed_at":"2026-07-24T17:07:01+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/exa-labs-exa-mcp-server","markdown_url":"https://www.graphcanon.com/tools/exa-labs-exa-mcp-server.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/exa-labs-exa-mcp-server","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=exa-labs-exa-mcp-server","shared_categories":["data-retrieval"]},{"slug":"ucbepic-docetl","name":"docetl","tagline":"A system for agentic LLM-powered data processing and ETL","github_url":"https://github.com/ucbepic/docetl","owner":"ucbepic","repo":"docetl","owner_avatar_url":"https://avatars.githubusercontent.com/u/88680502?v=4","primary_language":"Python","stars":3961,"forks":421,"topics":["agents","data","data-pipelines","document-analysis","document-processing","elt","etl","llm","python","semantic-data","unstructured-data","unstructured-data-analysis","workflow"],"archived":false,"github_pushed_at":"2026-08-09T23:31:04+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/ucbepic-docetl","markdown_url":"https://www.graphcanon.com/tools/ucbepic-docetl.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ucbepic-docetl","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ucbepic-docetl","shared_categories":["data-retrieval"]}]}}