{"data":{"node":{"slug":"paddlepaddle-paddleocr","name":"PaddleOCR","tagline":"A powerful, lightweight OCR toolkit to convert images and PDFs into structured data","github_url":"https://github.com/PaddlePaddle/PaddleOCR","owner":"PaddlePaddle","repo":"PaddleOCR","owner_avatar_url":"https://avatars.githubusercontent.com/u/23534030?v=4","primary_language":"Python","stars":87808,"forks":11187,"topics":["ai4science","chineseocr","document-parsing","document-translation","kie","ocr","paddleocr-vl","pdf-extractor-rag","pdf-parser","pdf2markdown","pp-ocr","pp-structure","rag"],"archived":false,"github_pushed_at":"2026-07-22T11:59:34+00:00","maintenance_label":"Active","stars_delta_30d":2062,"url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr","markdown_url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/paddlepaddle-paddleocr","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=paddlepaddle-paddleocr"},"categories":[{"slug":"computer-vision","name":"Computer Vision","url":"https://www.graphcanon.com/categories/computer-vision","markdown_url":"https://www.graphcanon.com/categories/computer-vision.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/computer-vision"}],"tags":[{"slug":"ai4science","name":"ai4science"},{"slug":"chineseocr","name":"chineseocr"},{"slug":"document-parsing","name":"document-parsing"},{"slug":"document-translation","name":"document-translation"},{"slug":"kie","name":"kie"},{"slug":"ocr","name":"ocr"},{"slug":"pdf-extractor-rag","name":"pdf-extractor-rag"},{"slug":"pdf-parser","name":"pdf-parser"}],"edges":[{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"run-llama-llama-index","name":"llama_index","tagline":"Leading document agent and OCR platform","github_url":"https://github.com/run-llama/llama_index","owner":"run-llama","repo":"llama_index","owner_avatar_url":"https://avatars.githubusercontent.com/u/130722866?v=4","primary_language":"Python","stars":51442,"forks":7885,"topics":["agents","application","data","fine-tuning","framework","llamaindex","llm","multi-agents","rag","vector-database"],"archived":false,"github_pushed_at":"2026-08-06T21:24:16+00:00","maintenance_label":"Very active","stars_delta_30d":719,"url":"https://www.graphcanon.com/tools/run-llama-llama-index","markdown_url":"https://www.graphcanon.com/tools/run-llama-llama-index.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/run-llama-llama-index","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=run-llama-llama-index"}},{"type":"related","direction":"out","explanation":"Both PaddleOCR and pdfmux deal with PDF processing and extraction, but while PaddleOCR focuses on OCR and document parsing to generate structured data from images/PDFs, pdfmux emphasizes self-healing mechanisms for cost-aware PDF extraction.","successor_context":null,"tool":{"slug":"nameetp-pdfmux","name":"pdfmux","tagline":"PDF extraction with self-healing and cost-aware mechanisms","github_url":"https://github.com/NameetP/pdfmux","owner":"NameetP","repo":"pdfmux","owner_avatar_url":"https://avatars.githubusercontent.com/u/93118951?v=4","primary_language":"Python","stars":79,"forks":12,"topics":["ai-agent","docling","document-ai","document-parsing","langchain","llamaindex","llm","mcp","mcp-server","ocr","opendataloader","pdf","pdf-extraction","pdf-parser","pdf-to-json","pdf-to-markdown","python","rag","self-healing","structured-extraction"],"archived":false,"github_pushed_at":"2026-08-13T08:03:22+00:00","maintenance_label":"Very active","stars_delta_30d":3,"url":"https://www.graphcanon.com/tools/nameetp-pdfmux","markdown_url":"https://www.graphcanon.com/tools/nameetp-pdfmux.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nameetp-pdfmux","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nameetp-pdfmux"}},{"type":"integrates_with","direction":"out","explanation":"PaddleOCR extracts text and structured data from images and PDFs, which can be processed by graphify to create a queryable knowledge graph, integrating OCR-extracted information into a larger corpus of interconnected data.","successor_context":null,"tool":{"slug":"graphify-labs-graphify","name":"graphify","tagline":"Turn any code or documentation into a queryable knowledge graph","github_url":"https://github.com/Graphify-Labs/graphify","owner":"Graphify-Labs","repo":"graphify","owner_avatar_url":"https://avatars.githubusercontent.com/u/297659074?v=4","primary_language":"Python","stars":107507,"forks":10441,"topics":["ai-agents","antigravity","ast","claude-code","code-analysis","code-search","codex","cursor","developer-tools","gemini","graphrag","knowledge-graph","leiden","llm","mcp","openclaw","rag","skills","tree-sitter"],"archived":false,"github_pushed_at":"2026-08-17T18:42:58+00:00","maintenance_label":"Very active","stars_delta_30d":16918,"url":"https://www.graphcanon.com/tools/graphify-labs-graphify","markdown_url":"https://www.graphcanon.com/tools/graphify-labs-graphify.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/graphify-labs-graphify","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=graphify-labs-graphify"}},{"type":"alternative","direction":"out","explanation":"PaddleOCR and opendataloader-pdf both aim at parsing documents into AI-ready data, serving a similar purpose in different ways.","successor_context":null,"tool":{"slug":"opendataloader-project-opendataloader-pdf","name":"opendataloader-pdf","tagline":"PDF Parser for AI-ready data","github_url":"https://github.com/opendataloader-project/opendataloader-pdf","owner":"opendataloader-project","repo":"opendataloader-pdf","owner_avatar_url":"https://avatars.githubusercontent.com/u/211280852?v=4","primary_language":"Java","stars":28528,"forks":2724,"topics":["a11y","accessibility","ai","bounding-box","document-parsing","eaa","html","json","markdown","ocr","ocr-recognition","pdf","pdf-accessibility","pdf-converter","pdf-extraction","pdf-parser","pdf-ua","rag","tables","tagged-pdf"],"archived":false,"github_pushed_at":"2026-08-18T04:01:45+00:00","maintenance_label":"Very active","stars_delta_30d":1078,"url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf","markdown_url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opendataloader-project-opendataloader-pdf","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opendataloader-project-opendataloader-pdf"}},{"type":"depends_on","direction":"out","explanation":"PaddleOCR could potentially depend on paddler for serving its OCR models at scale and load balancing, although this is speculative without more specific integration details.","successor_context":null,"tool":{"slug":"intentee-paddler","name":"paddler","tagline":"Open-source LLM/VLM load balancer and serving platform for self-hosting at scale","github_url":"https://github.com/intentee/paddler","owner":"intentee","repo":"paddler","owner_avatar_url":"https://avatars.githubusercontent.com/u/215040511?v=4","primary_language":"Rust","stars":1663,"forks":97,"topics":["ai","llamacpp","llm","llmops","load-balancer"],"archived":false,"github_pushed_at":"2026-07-19T19:36:21+00:00","maintenance_label":"Steady","stars_delta_30d":21,"url":"https://www.graphcanon.com/tools/intentee-paddler","markdown_url":"https://www.graphcanon.com/tools/intentee-paddler.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/intentee-paddler","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=intentee-paddler"}},{"type":"related","direction":"out","explanation":"PaddleOCR and paddler are both related to the PaddlePaddle ecosystem but serve different purposes. While PaddleOCR focuses on OCR and document parsing, paddler is a platform for serving LLMs. They do not directly integrate or depend on each other.","successor_context":null,"tool":{"slug":"intentee-paddler","name":"paddler","tagline":"Open-source LLM/VLM load balancer and serving platform for self-hosting at scale","github_url":"https://github.com/intentee/paddler","owner":"intentee","repo":"paddler","owner_avatar_url":"https://avatars.githubusercontent.com/u/215040511?v=4","primary_language":"Rust","stars":1663,"forks":97,"topics":["ai","llamacpp","llm","llmops","load-balancer"],"archived":false,"github_pushed_at":"2026-07-19T19:36:21+00:00","maintenance_label":"Steady","stars_delta_30d":21,"url":"https://www.graphcanon.com/tools/intentee-paddler","markdown_url":"https://www.graphcanon.com/tools/intentee-paddler.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/intentee-paddler","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=intentee-paddler"}},{"type":"alternative","direction":"out","explanation":"PaddleOCR and TurboOCR both serve OCR purposes, though TurboOCR specifically emphasizes speed with GPU acceleration.","successor_context":null,"tool":{"slug":"aiptimizer-turboocr","name":"TurboOCR","tagline":"Fast GPU-based OCR server for document parsing","github_url":"https://github.com/aiptimizer/TurboOCR","owner":"aiptimizer","repo":"TurboOCR","owner_avatar_url":"https://avatars.githubusercontent.com/u/274521723?v=4","primary_language":"C++","stars":986,"forks":96,"topics":["document-ai","document-parsing","easyocr","fastapi","fp16","gpu-ocr","grpc","inference-server","nvidia","ocr","paddleocr","pdf-extraction","qwen-vl","rag","tensorrt","text-detection","text-recognition"],"archived":false,"github_pushed_at":"2026-08-14T11:29:16+00:00","maintenance_label":"Very active","stars_delta_30d":604,"url":"https://www.graphcanon.com/tools/aiptimizer-turboocr","markdown_url":"https://www.graphcanon.com/tools/aiptimizer-turboocr.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/aiptimizer-turboocr","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=aiptimizer-turboocr"}},{"type":"related","direction":"in","explanation":null,"successor_context":null,"tool":{"slug":"koala73-worldmonitor","name":"worldmonitor","tagline":"Real-time global intelligence dashboard with AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking","github_url":"https://github.com/koala73/worldmonitor","owner":"koala73","repo":"worldmonitor","owner_avatar_url":"https://avatars.githubusercontent.com/u/996596?v=4","primary_language":"TypeScript","stars":82968,"forks":12378,"topics":["agent","ai","dashboard","geopolitics","mcp","mcp-server","monitoring","news","opensource","osint","palantir","situation"],"archived":false,"github_pushed_at":"2026-08-18T22:53:21+00:00","maintenance_label":"Very active","stars_delta_30d":20953,"url":"https://www.graphcanon.com/tools/koala73-worldmonitor","markdown_url":"https://www.graphcanon.com/tools/koala73-worldmonitor.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/koala73-worldmonitor","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=koala73-worldmonitor"}},{"type":"related","direction":"in","explanation":null,"successor_context":null,"tool":{"slug":"unstructured-io-unstructured","name":"unstructured","tagline":"Convert documents to structured data effortlessly","github_url":"https://github.com/Unstructured-IO/unstructured","owner":"Unstructured-IO","repo":"unstructured","owner_avatar_url":"https://avatars.githubusercontent.com/u/108372208?v=4","primary_language":"HTML","stars":15238,"forks":1284,"topics":["data-pipelines","deep-learning","document-image-analysis","document-image-processing","document-parser","document-parsing","docx","donut","information-retrieval","langchain","llm","machine-learning","ml","natural-language-processing","nlp","ocr","pdf","pdf-to-json","pdf-to-text","preprocessing"],"archived":false,"github_pushed_at":"2026-07-31T20:54:17+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/unstructured-io-unstructured","markdown_url":"https://www.graphcanon.com/tools/unstructured-io-unstructured.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unstructured-io-unstructured","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unstructured-io-unstructured"}},{"type":"integrates_with","direction":"in","explanation":"PaddleOCR's OCR capabilities can significantly enhance the functionality of OpenDataLoader PDF in handling scanned and handwritten documents, which are important use cases for both tools.","successor_context":null,"tool":{"slug":"opendataloader-project-opendataloader-pdf","name":"opendataloader-pdf","tagline":"PDF Parser for AI-ready data","github_url":"https://github.com/opendataloader-project/opendataloader-pdf","owner":"opendataloader-project","repo":"opendataloader-pdf","owner_avatar_url":"https://avatars.githubusercontent.com/u/211280852?v=4","primary_language":"Java","stars":28528,"forks":2724,"topics":["a11y","accessibility","ai","bounding-box","document-parsing","eaa","html","json","markdown","ocr","ocr-recognition","pdf","pdf-accessibility","pdf-converter","pdf-extraction","pdf-parser","pdf-ua","rag","tables","tagged-pdf"],"archived":false,"github_pushed_at":"2026-08-18T04:01:45+00:00","maintenance_label":"Very active","stars_delta_30d":1078,"url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf","markdown_url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opendataloader-project-opendataloader-pdf","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opendataloader-project-opendataloader-pdf"}},{"type":"integrates_with","direction":"in","explanation":"Metaflow can be used to manage and deploy models like those from PaddleOCR, making it possible to streamline the development lifecycle for OCR systems.","successor_context":null,"tool":{"slug":"netflix-metaflow","name":"metaflow","tagline":"Build, Manage and Deploy AI/ML Systems","github_url":"https://github.com/Netflix/metaflow","owner":"Netflix","repo":"metaflow","owner_avatar_url":"https://avatars.githubusercontent.com/u/913567?v=4","primary_language":"Python","stars":10228,"forks":1330,"topics":["agents","ai","aws","azure","cost-optimization","datascience","distributed-training","gcp","generative-ai","high-performance-computing","kubernetes","llm","llmops","machine-learning","ml","ml-infrastructure","ml-platform","mlops","model-management","python"],"archived":false,"github_pushed_at":"2026-08-18T09:41:43+00:00","maintenance_label":"Very active","stars_delta_30d":38,"url":"https://www.graphcanon.com/tools/netflix-metaflow","markdown_url":"https://www.graphcanon.com/tools/netflix-metaflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/netflix-metaflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=netflix-metaflow"}},{"type":"integrates_with","direction":"in","explanation":"MineContext could use PaddleOCR's capabilities to process and understand text in images or documents for better context engineering, providing a way for the AI partner to gain insights from visual content.","successor_context":null,"tool":{"slug":"volcengine-minecontext","name":"MineContext","tagline":"Proactive context-aware AI partner","github_url":"https://github.com/volcengine/MineContext","owner":"volcengine","repo":"MineContext","owner_avatar_url":"https://avatars.githubusercontent.com/u/67365215?v=4","primary_language":"Python","stars":5476,"forks":406,"topics":["agent","context-engineering","electron","embedding-models","javascript","memory","proactive-ai","python","python3","rag","react","typescript","vector-database","vision-language-model"],"archived":false,"github_pushed_at":"2026-05-07T13:23:05+00:00","maintenance_label":"Slowing","stars_delta_30d":40,"url":"https://www.graphcanon.com/tools/volcengine-minecontext","markdown_url":"https://www.graphcanon.com/tools/volcengine-minecontext.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/volcengine-minecontext","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=volcengine-minecontext"}},{"type":"integrates_with","direction":"in","explanation":"cube-studio supports PaddlePaddle, and since PaddleOCR is built on top of PaddlePaddle, it integrates well with the functionalities provided by cube-studio.","successor_context":null,"tool":{"slug":"tencentmusic-cube-studio","name":"cube-studio","tagline":"一站式机器学习/深度学习/AI开发平台","github_url":"https://github.com/tencentmusic/cube-studio","owner":"tencentmusic","repo":"cube-studio","owner_avatar_url":"https://avatars.githubusercontent.com/u/53810446?v=4","primary_language":null,"stars":5077,"forks":880,"topics":["ai","aihub","argo","automl","deepseek","gpt","inference","kubeflow","kubernetes","llmops","mlops","notebook","pipeline","pytorch","spark","vgpu","workflow"],"archived":false,"github_pushed_at":"2026-07-11T06:55:31+00:00","maintenance_label":"Steady","stars_delta_30d":2,"url":"https://www.graphcanon.com/tools/tencentmusic-cube-studio","markdown_url":"https://www.graphcanon.com/tools/tencentmusic-cube-studio.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tencentmusic-cube-studio","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tencentmusic-cube-studio"}},{"type":"integrates_with","direction":"in","explanation":"PaddleOCR could be used within NExT-GPT to preprocess text-based images as it supports multimodal inputs including visual data.","successor_context":null,"tool":{"slug":"next-gpt-next-gpt","name":"NExT-GPT","tagline":"Code and models for ICML 2024 paper on multimodal large language model","github_url":"https://github.com/NExT-GPT/NExT-GPT","owner":"NExT-GPT","repo":"NExT-GPT","owner_avatar_url":"https://avatars.githubusercontent.com/u/143576855?v=4","primary_language":"Python","stars":3637,"forks":359,"topics":["chatgpt","foundation-models","gpt-4","instruction-tuning","large-language-models","llm","mllm","multi-modal-chatgpt","multimodal","visual-language-learning"],"archived":false,"github_pushed_at":"2025-05-13T09:57:47+00:00","maintenance_label":"Dormant","stars_delta_30d":-1,"url":"https://www.graphcanon.com/tools/next-gpt-next-gpt","markdown_url":"https://www.graphcanon.com/tools/next-gpt-next-gpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/next-gpt-next-gpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=next-gpt-next-gpt"}},{"type":"related","direction":"in","explanation":"The repository contains general AI resources and might include OCR-related materials under data engineering or practices, thus related to paddlepaddle/PaddleOCR.","successor_context":null,"tool":{"slug":"liguodongiot-llm-resource","name":"llm-resource","tagline":"LLM全栈优质资源汇总","github_url":"https://github.com/liguodongiot/llm-resource","owner":"liguodongiot","repo":"llm-resource","owner_avatar_url":"https://avatars.githubusercontent.com/u/13220186?v=4","primary_language":"Shell","stars":725,"forks":85,"topics":["llm","llmops"],"archived":false,"github_pushed_at":"2025-07-15T16:52:24+00:00","maintenance_label":"Dormant","stars_delta_30d":2,"url":"https://www.graphcanon.com/tools/liguodongiot-llm-resource","markdown_url":"https://www.graphcanon.com/tools/liguodongiot-llm-resource.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/liguodongiot-llm-resource","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=liguodongiot-llm-resource"}},{"type":"related","direction":"in","explanation":"PaddleOCR is an OCR toolkit which can potentially use language models for text understanding or improvement. It's adjacent since Gateway connects with LLMs but doesn't specifically support OCR tools.","successor_context":null,"tool":{"slug":"adaline-gateway","name":"gateway","tagline":"Unified SDK for calling over 200 LLMs","github_url":"https://github.com/adaline/gateway","owner":"adaline","repo":"gateway","owner_avatar_url":"https://avatars.githubusercontent.com/u/382430?v=4","primary_language":"TypeScript","stars":605,"forks":26,"topics":["ai","ai-agents","anthropic","language-model","llm","llmops","openai","prompt-engineering","togetherai","typescript"],"archived":false,"github_pushed_at":"2026-07-29T19:11:29+00:00","maintenance_label":"Active","stars_delta_30d":4,"url":"https://www.graphcanon.com/tools/adaline-gateway","markdown_url":"https://www.graphcanon.com/tools/adaline-gateway.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/adaline-gateway","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=adaline-gateway"}},{"type":"alternative","direction":"in","explanation":"Llama_index is an open-source framework designed to build applications that engage with documents using large language models (LLMs) and OCR technology among other integrations. PaddleOCR, on the other hand, specializes in Optical Character Recognition, converting images and PDFs into structured data formats efficiently and accurately. The alternative relationship between llama_index and PaddleOCR","successor_context":null,"tool":{"slug":"run-llama-llama-index","name":"llama_index","tagline":"Leading document agent and OCR platform","github_url":"https://github.com/run-llama/llama_index","owner":"run-llama","repo":"llama_index","owner_avatar_url":"https://avatars.githubusercontent.com/u/130722866?v=4","primary_language":"Python","stars":51442,"forks":7885,"topics":["agents","application","data","fine-tuning","framework","llamaindex","llm","multi-agents","rag","vector-database"],"archived":false,"github_pushed_at":"2026-08-06T21:24:16+00:00","maintenance_label":"Very active","stars_delta_30d":719,"url":"https://www.graphcanon.com/tools/run-llama-llama-index","markdown_url":"https://www.graphcanon.com/tools/run-llama-llama-index.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/run-llama-llama-index","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=run-llama-llama-index"}},{"type":"related","direction":"in","explanation":"Both OASIS and PaddleOCR deal with AI in a broad sense, but their focuses are different; OASIS is about agent-based social interaction simulations while PaddleOCR involves OCR and document AI.","successor_context":null,"tool":{"slug":"camel-ai-oasis","name":"oasis","tagline":"OASIS: Open Agent Social Interaction Simulations with One Million Agents","github_url":"https://github.com/camel-ai/oasis","owner":"camel-ai","repo":"oasis","owner_avatar_url":"https://avatars.githubusercontent.com/u/134388954?v=4","primary_language":"Python","stars":5031,"forks":618,"topics":["agent-based-framework","agent-based-simulation","ai-societies","deep-learning","large-language-models","large-scale","llm-agents","multi-agent-systems","natural-language-processing"],"archived":false,"github_pushed_at":"2026-08-14T17:35:02+00:00","maintenance_label":"Very active","stars_delta_30d":93,"url":"https://www.graphcanon.com/tools/camel-ai-oasis","markdown_url":"https://www.graphcanon.com/tools/camel-ai-oasis.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/camel-ai-oasis","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=camel-ai-oasis"}},{"type":"related","direction":"in","explanation":"Both PaddleOCR and OpenContracts process documents, but PaddleOCR focuses on OCR toolkit while OpenContracts offers a broader document intelligence platform.","successor_context":null,"tool":{"slug":"open-source-legal-opencontracts","name":"OpenContracts","tagline":"The open document intelligence platform for builders and hackers - DMS for the agentic world","github_url":"https://github.com/Open-Source-Legal/OpenContracts","owner":"Open-Source-Legal","repo":"OpenContracts","owner_avatar_url":"https://avatars.githubusercontent.com/u/219575010?v=4","primary_language":"Python","stars":1443,"forks":178,"topics":["agent","agentic-ai","ai","ai-agents","etl","etl-pipeline","llm","prompt-engineering","unstructured-data","vector-database"],"archived":false,"github_pushed_at":"2026-08-21T05:19:19+00:00","maintenance_label":"Very active","stars_delta_30d":40,"url":"https://www.graphcanon.com/tools/open-source-legal-opencontracts","markdown_url":"https://www.graphcanon.com/tools/open-source-legal-opencontracts.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/open-source-legal-opencontracts","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=open-source-legal-opencontracts"}}],"neighbours":[{"slug":"pdfmathtranslate-pdfmathtranslate","name":"PDFMathTranslate","tagline":"PDF scientific paper translation with preserved formats","github_url":"https://github.com/PDFMathTranslate/PDFMathTranslate","owner":"PDFMathTranslate","repo":"PDFMathTranslate","owner_avatar_url":"https://avatars.githubusercontent.com/u/205132128?v=4","primary_language":"Python","stars":35794,"forks":3186,"topics":["chinese","document","edit","english","japanese","korean","latex","math","mcp","modify","obsidian","openai","pdf","pdf2zh","python","russian","translate","translation","zotero"],"archived":false,"github_pushed_at":"2026-05-25T19:00:26+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/pdfmathtranslate-pdfmathtranslate","markdown_url":"https://www.graphcanon.com/tools/pdfmathtranslate-pdfmathtranslate.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pdfmathtranslate-pdfmathtranslate","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pdfmathtranslate-pdfmathtranslate","shared_categories":[]},{"slug":"opendataloader-project-opendataloader-pdf","name":"opendataloader-pdf","tagline":"PDF Parser for AI-ready data","github_url":"https://github.com/opendataloader-project/opendataloader-pdf","owner":"opendataloader-project","repo":"opendataloader-pdf","owner_avatar_url":"https://avatars.githubusercontent.com/u/211280852?v=4","primary_language":"Java","stars":28528,"forks":2724,"topics":["a11y","accessibility","ai","bounding-box","document-parsing","eaa","html","json","markdown","ocr","ocr-recognition","pdf","pdf-accessibility","pdf-converter","pdf-extraction","pdf-parser","pdf-ua","rag","tables","tagged-pdf"],"archived":false,"github_pushed_at":"2026-08-18T04:01:45+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf","markdown_url":"https://www.graphcanon.com/tools/opendataloader-project-opendataloader-pdf.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opendataloader-project-opendataloader-pdf","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opendataloader-project-opendataloader-pdf","shared_categories":[]},{"slug":"sinaptik-ai-pandas-ai","name":"pandas-ai","tagline":"Chat with your database or your datalake using LLMs and RAG.","github_url":"https://github.com/sinaptik-ai/pandas-ai","owner":"sinaptik-ai","repo":"pandas-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/154438448?v=4","primary_language":"Python","stars":23746,"forks":2342,"topics":["ai","csv","data","data-analysis","data-science","data-visualization","database","datalake","gpt-4","llm","pandas","sql","text-to-sql"],"archived":false,"github_pushed_at":"2025-10-28T10:02:13+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai","markdown_url":"https://www.graphcanon.com/tools/sinaptik-ai-pandas-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sinaptik-ai-pandas-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sinaptik-ai-pandas-ai","shared_categories":[]},{"slug":"unstructured-io-unstructured","name":"unstructured","tagline":"Convert documents to structured data effortlessly","github_url":"https://github.com/Unstructured-IO/unstructured","owner":"Unstructured-IO","repo":"unstructured","owner_avatar_url":"https://avatars.githubusercontent.com/u/108372208?v=4","primary_language":"HTML","stars":15238,"forks":1284,"topics":["data-pipelines","deep-learning","document-image-analysis","document-image-processing","document-parser","document-parsing","docx","donut","information-retrieval","langchain","llm","machine-learning","ml","natural-language-processing","nlp","ocr","pdf","pdf-to-json","pdf-to-text","preprocessing"],"archived":false,"github_pushed_at":"2026-07-31T20:54:17+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/unstructured-io-unstructured","markdown_url":"https://www.graphcanon.com/tools/unstructured-io-unstructured.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unstructured-io-unstructured","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unstructured-io-unstructured","shared_categories":[]},{"slug":"data-privacy-stack-presidio","name":"presidio","tagline":"A framework for detecting and anonymizing sensitive data","github_url":"https://github.com/data-privacy-stack/presidio","owner":"data-privacy-stack","repo":"presidio","owner_avatar_url":"https://avatars.githubusercontent.com/u/275623515?v=4","primary_language":"Python","stars":10395,"forks":1237,"topics":["anonymization","data-anonymization","data-masking","data-obfuscation","data-privacy","data-redaction","de-identification","guardrails","image-redactor","named-entity-recognition","nlp","personally-identifiable-information","phi","pii","pii-detection","privacy","python","sensitive-data","spacy","transformers"],"archived":false,"github_pushed_at":"2026-08-08T21:25:09+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/data-privacy-stack-presidio","markdown_url":"https://www.graphcanon.com/tools/data-privacy-stack-presidio.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/data-privacy-stack-presidio","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=data-privacy-stack-presidio","shared_categories":[]},{"slug":"xberg-io-xberg","name":"xberg","tagline":"A polyglot document intelligence framework with a Rust core","github_url":"https://github.com/xberg-io/xberg","owner":"xberg-io","repo":"xberg","owner_avatar_url":"https://avatars.githubusercontent.com/u/241328462?v=4","primary_language":"Rust","stars":9128,"forks":562,"topics":["bun","csharp","document-intelligence","elixir","ffi","golang","java","metadata-extraction","node","pdf-extraction","pdfium","php","python","rag","ruby","rust","table-extraction","tesseract","text-extraction","wasm"],"archived":false,"github_pushed_at":"2026-08-16T16:00:11+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/xberg-io-xberg","markdown_url":"https://www.graphcanon.com/tools/xberg-io-xberg.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/xberg-io-xberg","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=xberg-io-xberg","shared_categories":[]},{"slug":"datajuicer-data-juicer","name":"data-juicer","tagline":"Data processing for and with foundation models","github_url":"https://github.com/datajuicer/data-juicer","owner":"datajuicer","repo":"data-juicer","owner_avatar_url":"https://avatars.githubusercontent.com/u/223222708?v=4","primary_language":"Python","stars":6897,"forks":404,"topics":["data","data-analysis","data-pipeline","data-processing","data-science","data-visualization","foundation-models","instruction-tuning","large-language-models","llm","llms","multi-modal","pre-training","synthetic-data"],"archived":false,"github_pushed_at":"2026-08-13T09:19:31+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/datajuicer-data-juicer","markdown_url":"https://www.graphcanon.com/tools/datajuicer-data-juicer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datajuicer-data-juicer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datajuicer-data-juicer","shared_categories":[]},{"slug":"nyldn-claude-octopus","name":"claude-octopus","tagline":"Surface AI blindspots before you ship","github_url":"https://github.com/nyldn/claude-octopus","owner":"nyldn","repo":"claude-octopus","owner_avatar_url":"https://avatars.githubusercontent.com/u/4805949?v=4","primary_language":"Shell","stars":3962,"forks":374,"topics":["ai-agents","ai-orchestration","claude-code","claude-code-plugin","codex","copilot","developer-tools","double-diamond","gemini","multi-ai","multi-llm","ollama"],"archived":false,"github_pushed_at":"2026-08-13T23:28:23+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/nyldn-claude-octopus","markdown_url":"https://www.graphcanon.com/tools/nyldn-claude-octopus.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nyldn-claude-octopus","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nyldn-claude-octopus","shared_categories":[]},{"slug":"huggingface-datatrove","name":"datatrove","tagline":"Platform-agnostic customizable pipeline processing blocks for data processing and transformation.","github_url":"https://github.com/huggingface/datatrove","owner":"huggingface","repo":"datatrove","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":3250,"forks":288,"topics":[],"archived":false,"github_pushed_at":"2026-08-06T15:27:26+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/huggingface-datatrove","markdown_url":"https://www.graphcanon.com/tools/huggingface-datatrove.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-datatrove","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-datatrove","shared_categories":[]},{"slug":"intentee-paddler","name":"paddler","tagline":"Open-source LLM/VLM load balancer and serving platform for self-hosting at scale","github_url":"https://github.com/intentee/paddler","owner":"intentee","repo":"paddler","owner_avatar_url":"https://avatars.githubusercontent.com/u/215040511?v=4","primary_language":"Rust","stars":1663,"forks":97,"topics":["ai","llamacpp","llm","llmops","load-balancer"],"archived":false,"github_pushed_at":"2026-07-19T19:36:21+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/intentee-paddler","markdown_url":"https://www.graphcanon.com/tools/intentee-paddler.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/intentee-paddler","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=intentee-paddler","shared_categories":[]},{"slug":"pixeltable-pixeltable","name":"pixeltable","tagline":"Unified multimodal backend for AI data apps","github_url":"https://github.com/pixeltable/pixeltable","owner":"pixeltable","repo":"pixeltable","owner_avatar_url":"https://avatars.githubusercontent.com/u/160283145?v=4","primary_language":"Python","stars":1613,"forks":219,"topics":["ai","computer-vision","data-science","database","feature-engineering","feature-store","genai","llm","machine-learning","ml","multimodal","vector-database"],"archived":false,"github_pushed_at":"2026-08-21T06:36:51+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/pixeltable-pixeltable","markdown_url":"https://www.graphcanon.com/tools/pixeltable-pixeltable.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/pixeltable-pixeltable","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=pixeltable-pixeltable","shared_categories":["computer-vision"]},{"slug":"unum-cloud-uform","name":"UForm","tagline":"Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and video","github_url":"https://github.com/unum-cloud/UForm","owner":"unum-cloud","repo":"UForm","owner_avatar_url":"https://avatars.githubusercontent.com/u/56397513?v=4","primary_language":"Python","stars":1243,"forks":78,"topics":["bert","clip","clustering","contrastive-learning","cross-attention","huggingface-transformers","image-search","language-vision","llava","multi-lingual","multimodal","neural-network","openai","openclip","pretrained-models","pytorch","representation-learning","semantic-search","transformer","vector-search"],"archived":false,"github_pushed_at":"2025-10-30T23:39:54+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/unum-cloud-uform","markdown_url":"https://www.graphcanon.com/tools/unum-cloud-uform.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unum-cloud-uform","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unum-cloud-uform","shared_categories":["computer-vision"]}]}}