{"data":{"node":{"slug":"opendataloader-project-opendataloader-pdf","name":"opendataloader-pdf","tagline":"PDF Parser for AI-ready data","github_url":"https://github.com/opendataloader-project/opendataloader-pdf","owner":"opendataloader-project","repo":"opendataloader-pdf","owner_avatar_url":"https://avatars.githubusercontent.com/u/211280852?v=4","primary_language":"Java","stars":28528,"forks":2724,"topics":["a11y","accessibility","ai","bounding-box","document-parsing","eaa","html","json","markdown","ocr","ocr-recognition","pdf","pdf-accessibility","pdf-converter","pdf-extraction","pdf-parser","pdf-ua","rag","tables","tagged-pdf"],"archived":false,"github_pushed_at":"2026-08-18T04:01:45+00:00","maintenance_label":"Very active","stars_delta_30d":1078,"url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf","markdown_url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opendataloader-project-opendataloader-pdf","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opendataloader-project-opendataloader-pdf"},"categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"a11y","name":"a11y"},{"slug":"accessibility","name":"accessibility"},{"slug":"ai","name":"ai"},{"slug":"bounding-box","name":"bounding-box"},{"slug":"document-parsing","name":"document-parsing"},{"slug":"ocr","name":"ocr"},{"slug":"pdf-accessibility","name":"pdf-accessibility"},{"slug":"pdf-ua","name":"pdf-ua"}],"edges":[{"type":"alternative","direction":"out","explanation":"Both OpenDataLoader PDF and unstructured parse various document types into structured data, focusing on accessibility and ease of use for AI applications.","successor_context":null,"tool":{"slug":"unstructured-io-unstructured","name":"unstructured","tagline":"Convert documents to structured data effortlessly","github_url":"https://github.com/Unstructured-IO/unstructured","owner":"Unstructured-IO","repo":"unstructured","owner_avatar_url":"https://avatars.githubusercontent.com/u/108372208?v=4","primary_language":"HTML","stars":15238,"forks":1284,"topics":["data-pipelines","deep-learning","document-image-analysis","document-image-processing","document-parser","document-parsing","docx","donut","information-retrieval","langchain","llm","machine-learning","ml","natural-language-processing","nlp","ocr","pdf","pdf-to-json","pdf-to-text","preprocessing"],"archived":false,"github_pushed_at":"2026-07-31T20:54:17+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/unstructured-io-unstructured","markdown_url":"https://www.graphcanon.com/tools/unstructured-io-unstructured.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unstructured-io-unstructured","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unstructured-io-unstructured"}},{"type":"integrates_with","direction":"out","explanation":"OpenDataLoader PDF extracts data that can be used in graph-based RAG systems like graphrag, making them complementary tools.","successor_context":null,"tool":{"slug":"microsoft-graphrag","name":"graphrag","tagline":"A modular graph-based Retrieval-Augmented Generation (RAG) system","github_url":"https://github.com/microsoft/graphrag","owner":"microsoft","repo":"graphrag","owner_avatar_url":"https://avatars.githubusercontent.com/u/6154722?v=4","primary_language":"Python","stars":35519,"forks":3734,"topics":["gpt","gpt-4","gpt4","graphrag","llm","llms","rag"],"archived":false,"github_pushed_at":"2026-08-14T18:16:11+00:00","maintenance_label":"Very active","stars_delta_30d":1049,"url":"https://www.graphcanon.com/tools/microsoft-graphrag","markdown_url":"https://www.graphcanon.com/tools/microsoft-graphrag.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/microsoft-graphrag","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=microsoft-graphrag"}},{"type":"integrates_with","direction":"out","explanation":"PaddleOCR's OCR capabilities can significantly enhance the functionality of OpenDataLoader PDF in handling scanned and handwritten documents, which are important use cases for both tools.","successor_context":null,"tool":{"slug":"paddlepaddle-paddleocr","name":"PaddleOCR","tagline":"A powerful, lightweight OCR toolkit to convert images and PDFs into structured data","github_url":"https://github.com/PaddlePaddle/PaddleOCR","owner":"PaddlePaddle","repo":"PaddleOCR","owner_avatar_url":"https://avatars.githubusercontent.com/u/23534030?v=4","primary_language":"Python","stars":87808,"forks":11187,"topics":["ai4science","chineseocr","document-parsing","document-translation","kie","ocr","paddleocr-vl","pdf-extractor-rag","pdf-parser","pdf2markdown","pp-ocr","pp-structure","rag"],"archived":false,"github_pushed_at":"2026-07-22T11:59:34+00:00","maintenance_label":"Active","stars_delta_30d":2062,"url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr","markdown_url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/paddlepaddle-paddleocr","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=paddlepaddle-paddleocr"}},{"type":"alternative","direction":"out","explanation":"Both OpenDataLoader PDF and langextract focus on extracting structured information from unstructured text, but they approach the task through different mechanisms (PDF parsing vs. LLM-based extraction).","successor_context":null,"tool":{"slug":"google-langextract","name":"langextract","tagline":"A Python library for extracting structured information from unstructured text using LLMs.","github_url":"https://github.com/google/langextract","owner":"google","repo":"langextract","owner_avatar_url":"https://avatars.githubusercontent.com/u/1342004?v=4","primary_language":"Python","stars":38400,"forks":2693,"topics":["gemini","gemini-ai","gemini-api","gemini-flash","gemini-pro","information-extration","large-language-models","llm","nlp","python","structured-data"],"archived":false,"github_pushed_at":"2026-08-11T15:31:39+00:00","maintenance_label":"Very active","stars_delta_30d":1241,"url":"https://www.graphcanon.com/tools/google-langextract","markdown_url":"https://www.graphcanon.com/tools/google-langextract.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/google-langextract","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=google-langextract"}},{"type":"related","direction":"out","explanation":"Both tools leverage LLMs to process structured data (Pandas-ai operates on databases and datalakes, whereas OpenDataLoader PDF handles PDF documents) for analysis.","successor_context":null,"tool":{"slug":"sinaptik-ai-pandas-ai","name":"pandas-ai","tagline":"Chat with your database or your datalake using LLMs and RAG.","github_url":"https://github.com/sinaptik-ai/pandas-ai","owner":"sinaptik-ai","repo":"pandas-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/154438448?v=4","primary_language":"Python","stars":23746,"forks":2342,"topics":["ai","csv","data","data-analysis","data-science","data-visualization","database","datalake","gpt-4","llm","pandas","sql","text-to-sql"],"archived":false,"github_pushed_at":"2025-10-28T10:02:13+00:00","maintenance_label":"Slowing","stars_delta_30d":90,"url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai","markdown_url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sinaptik-ai-pandas-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sinaptik-ai-pandas-ai"}},{"type":"alternative","direction":"in","explanation":"Both LangExtract and opendataloader-pdf focus on extracting structured data from unstructured sources, with opendataloader specifically targeting PDFs.","successor_context":null,"tool":{"slug":"google-langextract","name":"langextract","tagline":"A Python library for extracting structured information from unstructured text using LLMs.","github_url":"https://github.com/google/langextract","owner":"google","repo":"langextract","owner_avatar_url":"https://avatars.githubusercontent.com/u/1342004?v=4","primary_language":"Python","stars":38400,"forks":2693,"topics":["gemini","gemini-ai","gemini-api","gemini-flash","gemini-pro","information-extration","large-language-models","llm","nlp","python","structured-data"],"archived":false,"github_pushed_at":"2026-08-11T15:31:39+00:00","maintenance_label":"Very active","stars_delta_30d":1241,"url":"https://www.graphcanon.com/tools/google-langextract","markdown_url":"https://www.graphcanon.com/tools/google-langextract.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/google-langextract","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=google-langextract"}},{"type":"alternative","direction":"in","explanation":"Both opendataloader-pdf and unstructured deal with parsing PDFs, but they likely have different approaches or capabilities.","successor_context":null,"tool":{"slug":"unstructured-io-unstructured","name":"unstructured","tagline":"Convert documents to structured data effortlessly","github_url":"https://github.com/Unstructured-IO/unstructured","owner":"Unstructured-IO","repo":"unstructured","owner_avatar_url":"https://avatars.githubusercontent.com/u/108372208?v=4","primary_language":"HTML","stars":15238,"forks":1284,"topics":["data-pipelines","deep-learning","document-image-analysis","document-image-processing","document-parser","document-parsing","docx","donut","information-retrieval","langchain","llm","machine-learning","ml","natural-language-processing","nlp","ocr","pdf","pdf-to-json","pdf-to-text","preprocessing"],"archived":false,"github_pushed_at":"2026-07-31T20:54:17+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/unstructured-io-unstructured","markdown_url":"https://www.graphcanon.com/tools/unstructured-io-unstructured.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unstructured-io-unstructured","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unstructured-io-unstructured"}},{"type":"alternative","direction":"in","explanation":"PaddleOCR and opendataloader-pdf both aim at parsing documents into AI-ready data, serving a similar purpose in different ways.","successor_context":null,"tool":{"slug":"paddlepaddle-paddleocr","name":"PaddleOCR","tagline":"A powerful, lightweight OCR toolkit to convert images and PDFs into structured data","github_url":"https://github.com/PaddlePaddle/PaddleOCR","owner":"PaddlePaddle","repo":"PaddleOCR","owner_avatar_url":"https://avatars.githubusercontent.com/u/23534030?v=4","primary_language":"Python","stars":87808,"forks":11187,"topics":["ai4science","chineseocr","document-parsing","document-translation","kie","ocr","paddleocr-vl","pdf-extractor-rag","pdf-parser","pdf2markdown","pp-ocr","pp-structure","rag"],"archived":false,"github_pushed_at":"2026-07-22T11:59:34+00:00","maintenance_label":"Active","stars_delta_30d":2062,"url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr","markdown_url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/paddlepaddle-paddleocr","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=paddlepaddle-paddleocr"}},{"type":"integrates_with","direction":"in","explanation":"xberg can extract PDF content which could be fed into opendataloader-pdf for further AI processing, suggesting they integrate well together.","successor_context":null,"tool":{"slug":"xberg-io-xberg","name":"xberg","tagline":"A polyglot document intelligence framework with a Rust core","github_url":"https://github.com/xberg-io/xberg","owner":"xberg-io","repo":"xberg","owner_avatar_url":"https://avatars.githubusercontent.com/u/241328462?v=4","primary_language":"Rust","stars":9128,"forks":562,"topics":["bun","csharp","document-intelligence","elixir","ffi","golang","java","metadata-extraction","node","pdf-extraction","pdfium","php","python","rag","ruby","rust","table-extraction","tesseract","text-extraction","wasm"],"archived":false,"github_pushed_at":"2026-08-16T16:00:11+00:00","maintenance_label":"Very active","stars_delta_30d":460,"url":"https://www.graphcanon.com/tools/xberg-io-xberg","markdown_url":"https://www.graphcanon.com/tools/xberg-io-xberg.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/xberg-io-xberg","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=xberg-io-xberg"}},{"type":"integrates_with","direction":"in","explanation":"llama_cloud_services and opendataloader-pdf both process PDF documents, making it logical for them to integrate, especially when focusing on AI-ready data processing.","successor_context":null,"tool":{"slug":"run-llama-llama-cloud-services","name":"llama_cloud_services","tagline":"Knowledge Agents and Management in the Cloud","github_url":"https://github.com/run-llama/llama_cloud_services","owner":"run-llama","repo":"llama_cloud_services","owner_avatar_url":"https://avatars.githubusercontent.com/u/130722866?v=4","primary_language":"TypeScript","stars":4258,"forks":467,"topics":["document","document-parser","document-parsing","docx-to-markdown","parsing","pdf","pdf-document-processor","pdf-to-excel","pdf-to-json","pdf-to-markdown","pdf-to-text","ppt-to-json","ppt-to-markdown","pptx","structured-data","tables"],"archived":false,"github_pushed_at":"2026-05-18T19:09:10+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/run-llama-llama-cloud-services","markdown_url":"https://www.graphcanon.com/tools/run-llama-llama-cloud-services.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/run-llama-llama-cloud-services","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=run-llama-llama-cloud-services"}},{"type":"integrates_with","direction":"in","explanation":"ChatWeb processes PDFs and other documents which could be pre-processed by opendataloader-pdf, making them likely to integrate well.","successor_context":null,"tool":{"slug":"skywalkerdarren-chatweb","name":"chatWeb","tagline":"ChatWeb can crawl web pages and various document types for content extraction and summarization.","github_url":"https://github.com/SkywalkerDarren/chatWeb","owner":"SkywalkerDarren","repo":"chatWeb","owner_avatar_url":"https://avatars.githubusercontent.com/u/20706299?v=4","primary_language":"Python","stars":916,"forks":137,"topics":["ai","chatgpt","crawler","docx","embedding","faiss","gpt","gpt-35-turbo","news-extractor","newspaper","openai","pdf","pgvector","postgresql","vector-database"],"archived":false,"github_pushed_at":"2026-05-25T16:56:25+00:00","maintenance_label":"Steady","stars_delta_30d":0,"url":"https://www.graphcanon.com/tools/skywalkerdarren-chatweb","markdown_url":"https://www.graphcanon.com/tools/skywalkerdarren-chatweb.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/skywalkerdarren-chatweb","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=skywalkerdarren-chatweb"}}],"neighbours":[{"slug":"graphify-labs-graphify","name":"graphify","tagline":"Turn any code or documentation into a queryable knowledge graph","github_url":"https://github.com/Graphify-Labs/graphify","owner":"Graphify-Labs","repo":"graphify","owner_avatar_url":"https://avatars.githubusercontent.com/u/297659074?v=4","primary_language":"Python","stars":107507,"forks":10441,"topics":["ai-agents","antigravity","ast","claude-code","code-analysis","code-search","codex","cursor","developer-tools","gemini","graphrag","knowledge-graph","leiden","llm","mcp","openclaw","rag","skills","tree-sitter"],"archived":false,"github_pushed_at":"2026-08-17T18:42:58+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/graphify-labs-graphify","markdown_url":"https://www.graphcanon.com/tools/graphify-labs-graphify.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/graphify-labs-graphify","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=graphify-labs-graphify","shared_categories":["data-retrieval"]},{"slug":"iofficeai-officecli","name":"OfficeCLI","tagline":"First and best Office suite for AI agents to automate Word, Excel, and PowerPoint files.","github_url":"https://github.com/iOfficeAI/OfficeCLI","owner":"iOfficeAI","repo":"OfficeCLI","owner_avatar_url":"https://avatars.githubusercontent.com/u/145246968?v=4","primary_language":"C#","stars":28777,"forks":1953,"topics":["agent","ai","claude-code","cli","codex","docx","excel","office","openclaw","pptx","presentation","skills","word","xlsx"],"archived":false,"github_pushed_at":"2026-08-13T10:28:53+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/iofficeai-officecli","markdown_url":"https://www.graphcanon.com/tools/iofficeai-officecli.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/iofficeai-officecli","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=iofficeai-officecli","shared_categories":[]},{"slug":"promtengineer-localgpt","name":"localGPT","tagline":"Chat with your documents locally using GPT models","github_url":"https://github.com/PromtEngineer/localGPT","owner":"PromtEngineer","repo":"localGPT","owner_avatar_url":"https://avatars.githubusercontent.com/u/134474669?v=4","primary_language":"Python","stars":22209,"forks":2467,"topics":[],"archived":false,"github_pushed_at":"2026-07-18T07:15:01+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/promtengineer-localgpt","markdown_url":"https://www.graphcanon.com/tools/promtengineer-localgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promtengineer-localgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promtengineer-localgpt","shared_categories":["data-retrieval"]},{"slug":"unstructured-io-unstructured","name":"unstructured","tagline":"Convert documents to structured data effortlessly","github_url":"https://github.com/Unstructured-IO/unstructured","owner":"Unstructured-IO","repo":"unstructured","owner_avatar_url":"https://avatars.githubusercontent.com/u/108372208?v=4","primary_language":"HTML","stars":15238,"forks":1284,"topics":["data-pipelines","deep-learning","document-image-analysis","document-image-processing","document-parser","document-parsing","docx","donut","information-retrieval","langchain","llm","machine-learning","ml","natural-language-processing","nlp","ocr","pdf","pdf-to-json","pdf-to-text","preprocessing"],"archived":false,"github_pushed_at":"2026-07-31T20:54:17+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/unstructured-io-unstructured","markdown_url":"https://www.graphcanon.com/tools/unstructured-io-unstructured.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unstructured-io-unstructured","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unstructured-io-unstructured","shared_categories":["model-training","data-retrieval"]},{"slug":"yusufkaraaslan-skill-seekers","name":"Skill_Seekers","tagline":"Automation tool for converting documentation and code into Claude AI skills","github_url":"https://github.com/yusufkaraaslan/Skill_Seekers","owner":"yusufkaraaslan","repo":"Skill_Seekers","owner_avatar_url":"https://avatars.githubusercontent.com/u/11597362?v=4","primary_language":"Python","stars":14565,"forks":1477,"topics":["ai-tools","ast-parser","automation","claude-ai","claude-skills","code-analysis","conflict-detection","documentation","documentation-generator","github","github-scraper","mcp","mcp-server","multi-source","ocr","pdf","python","web-scraping"],"archived":false,"github_pushed_at":"2026-07-20T13:09:44+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/yusufkaraaslan-skill-seekers","markdown_url":"https://www.graphcanon.com/tools/yusufkaraaslan-skill-seekers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/yusufkaraaslan-skill-seekers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=yusufkaraaslan-skill-seekers","shared_categories":["model-training"]},{"slug":"xberg-io-xberg","name":"xberg","tagline":"A polyglot document intelligence framework with a Rust core","github_url":"https://github.com/xberg-io/xberg","owner":"xberg-io","repo":"xberg","owner_avatar_url":"https://avatars.githubusercontent.com/u/241328462?v=4","primary_language":"Rust","stars":9128,"forks":562,"topics":["bun","csharp","document-intelligence","elixir","ffi","golang","java","metadata-extraction","node","pdf-extraction","pdfium","php","python","rag","ruby","rust","table-extraction","tesseract","text-extraction","wasm"],"archived":false,"github_pushed_at":"2026-08-16T16:00:11+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/xberg-io-xberg","markdown_url":"https://www.graphcanon.com/tools/xberg-io-xberg.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/xberg-io-xberg","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=xberg-io-xberg","shared_categories":["data-retrieval"]},{"slug":"zipstack-unstract","name":"unstract","tagline":"LLM-Driven Extraction of Unstructured Data for API Deployments and ETL Pipeline Workflows","github_url":"https://github.com/Zipstack/unstract","owner":"Zipstack","repo":"unstract","owner_avatar_url":"https://avatars.githubusercontent.com/u/89070934?v=4","primary_language":"Python","stars":6932,"forks":663,"topics":["ai-agents","data-engineering","document-ai","generative-ai","idp","json-extraction","llm","mcp-server","ocr","pdf-extraction","prompt-engineering","structured-output"],"archived":false,"github_pushed_at":"2026-07-27T22:23:41+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/zipstack-unstract","markdown_url":"https://www.graphcanon.com/tools/zipstack-unstract.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/zipstack-unstract","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=zipstack-unstract","shared_categories":["data-retrieval"]},{"slug":"datajuicer-data-juicer","name":"data-juicer","tagline":"Data processing for and with foundation models","github_url":"https://github.com/datajuicer/data-juicer","owner":"datajuicer","repo":"data-juicer","owner_avatar_url":"https://avatars.githubusercontent.com/u/223222708?v=4","primary_language":"Python","stars":6897,"forks":404,"topics":["data","data-analysis","data-pipeline","data-processing","data-science","data-visualization","foundation-models","instruction-tuning","large-language-models","llm","llms","multi-modal","pre-training","synthetic-data"],"archived":false,"github_pushed_at":"2026-08-13T09:19:31+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/datajuicer-data-juicer","markdown_url":"https://www.graphcanon.com/tools/datajuicer-data-juicer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datajuicer-data-juicer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datajuicer-data-juicer","shared_categories":["model-training","data-retrieval"]},{"slug":"clusterzx-paperless-ai","name":"paperless-ai","tagline":"Automated document analyzer for Paperless-ngx using OpenAI API and compatible services to tag documents","github_url":"https://github.com/clusterzx/paperless-ai","owner":"clusterzx","repo":"paperless-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/32274973?v=4","primary_language":"JavaScript","stars":5882,"forks":322,"topics":["ai","automation","gemma","llama","mistral","ollama","paperless","paperless-ngx","phi"],"archived":false,"github_pushed_at":"2026-08-02T04:51:37+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/clusterzx-paperless-ai","markdown_url":"https://www.graphcanon.com/tools/clusterzx-paperless-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/clusterzx-paperless-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=clusterzx-paperless-ai","shared_categories":["model-training"]},{"slug":"guangzhengli-chatfiles","name":"ChatFiles","tagline":"Document Chatbot powered by GPT and Embedding","github_url":"https://github.com/guangzhengli/ChatFiles","owner":"guangzhengli","repo":"ChatFiles","owner_avatar_url":"https://avatars.githubusercontent.com/u/39078991?v=4","primary_language":"TypeScript","stars":3341,"forks":464,"topics":["chatbot","chatfile","chatgpt","chatgpt-api","chatpdf"],"archived":false,"github_pushed_at":"2024-12-17T10:26:50+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/guangzhengli-chatfiles","markdown_url":"https://www.graphcanon.com/tools/guangzhengli-chatfiles.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/guangzhengli-chatfiles","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=guangzhengli-chatfiles","shared_categories":[]},{"slug":"icereed-paperless-gpt","name":"paperless-gpt","tagline":"Document Digitalization powered by AI","github_url":"https://github.com/icereed/paperless-gpt","owner":"icereed","repo":"paperless-gpt","owner_avatar_url":"https://avatars.githubusercontent.com/u/444269?v=4","primary_language":"Go","stars":2614,"forks":195,"topics":["ai","chatgpt","llm","mistral","ocr","ollama","paperless","paperless-ngx"],"archived":false,"github_pushed_at":"2026-08-13T00:53:41+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/icereed-paperless-gpt","markdown_url":"https://www.graphcanon.com/tools/icereed-paperless-gpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/icereed-paperless-gpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=icereed-paperless-gpt","shared_categories":[]},{"slug":"agentset-ai-agentset","name":"agentset","tagline":"The open-source RAG platform with built-in citations and support for deep research","github_url":"https://github.com/agentset-ai/agentset","owner":"agentset-ai","repo":"agentset","owner_avatar_url":"https://avatars.githubusercontent.com/u/200139246?v=4","primary_language":"TypeScript","stars":2035,"forks":183,"topics":["agentic-rag","ai","ai-agents","ai-sdk","chatbots","embeddings","genai","llms","memory","memory-management","rag","vercel-ai-sdk"],"archived":false,"github_pushed_at":"2026-07-16T13:11:34+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/agentset-ai-agentset","markdown_url":"https://www.graphcanon.com/tools/agentset-ai-agentset.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/agentset-ai-agentset","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=agentset-ai-agentset","shared_categories":["data-retrieval"]}]}}