{"data":{"node":{"slug":"paulpierre-markdown-crawler","name":"markdown-crawler","tagline":"A multithreaded web crawler for creating markdown files from website pages","github_url":"https://github.com/paulpierre/markdown-crawler","owner":"paulpierre","repo":"markdown-crawler","owner_avatar_url":"https://avatars.githubusercontent.com/u/142327?v=4","primary_language":"Python","stars":467,"forks":54,"topics":["html-to-markdown","html-to-markdown-converter","html2md","llm","llmops","markdown","markdown-crawler","markdown-parser","markdown-scraper","md-crawler","rag","web-scraper"],"archived":false,"github_pushed_at":"2026-06-26T14:23:38+00:00","maintenance_label":"Steady","stars_delta_30d":2,"url":"https://www.graphcanon.com/tools/paulpierre-markdown-crawler","markdown_url":"https://www.graphcanon.com/tools/paulpierre-markdown-crawler.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/paulpierre-markdown-crawler","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=paulpierre-markdown-crawler"},"categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"}],"tags":[{"slug":"html-to-markdown","name":"html-to-markdown"},{"slug":"html2md","name":"html2md"},{"slug":"llmops","name":"llmops"},{"slug":"markdown-crawler","name":"markdown-crawler"},{"slug":"web-scraper","name":"web-scraper"}],"edges":[{"type":"alternative","direction":"out","explanation":"markdown-crawler and firecrawl both enable web scraping, but markdown-crawler focuses on converting HTML into markdown format for LLM RAG use cases, whereas firecrawl is more general-purpose.","successor_context":null,"tool":{"slug":"firecrawl-firecrawl","name":"firecrawl","tagline":"The API to search, scrape, and interact with the web at scale. 🔥","github_url":"https://github.com/firecrawl/firecrawl","owner":"firecrawl","repo":"firecrawl","owner_avatar_url":"https://avatars.githubusercontent.com/u/135057108?v=4","primary_language":"TypeScript","stars":167794,"forks":9398,"topics":["ai","ai-agents","ai-crawler","ai-scraping","ai-search","crawler","data-extraction","html-to-markdown","llm","markdown","scraper","scraping","web-crawler","web-data","web-data-extraction","web-scraper","web-scraping","web-search","webscraping"],"archived":false,"github_pushed_at":"2026-08-15T19:52:46+00:00","maintenance_label":"Very active","stars_delta_30d":15884,"url":"https://www.graphcanon.com/tools/firecrawl-firecrawl","markdown_url":"https://www.graphcanon.com/tools/firecrawl-firecrawl.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/firecrawl-firecrawl","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=firecrawl-firecrawl"}},{"type":"integrates_with","direction":"out","explanation":"markdown-crawler creates Markdown which can be processed by Transformers, as part of a larger pipeline to facilitate LLM operations like RAG use cases.","successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"integrates_with","direction":"out","explanation":"markdown-crawler can be used to preprocess content into markdown format, which could then be ingested and processed by an RAG engine like ragflow for information retrieval and augmentation.","successor_context":null,"tool":{"slug":"infiniflow-ragflow","name":"ragflow","tagline":"Retrieval-Augmented Generation engine with agent capabilities","github_url":"https://github.com/infiniflow/ragflow","owner":"infiniflow","repo":"ragflow","owner_avatar_url":"https://avatars.githubusercontent.com/u/69962740?v=4","primary_language":"Go","stars":86541,"forks":10167,"topics":["agent-harness","agentic-ai","agentic-retrieval","agentic-search","ai","ai-agents","context-engine","context-engineering","context-management","harness-engineering","knowledge-compilation","llm-apps","rag","retrieval-augmented-generation"],"archived":false,"github_pushed_at":"2026-07-31T14:59:12+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/infiniflow-ragflow","markdown_url":"https://www.graphcanon.com/tools/infiniflow-ragflow.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/infiniflow-ragflow","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=infiniflow-ragflow"}}],"neighbours":[{"slug":"jatinkrmalik-llmfeeder","name":"LLMFeeder","tagline":"Browser extension for converting web pages to clean Markdown","github_url":"https://github.com/jatinkrmalik/LLMFeeder","owner":"jatinkrmalik","repo":"LLMFeeder","owner_avatar_url":"https://avatars.githubusercontent.com/u/7387945?v=4","primary_language":"JavaScript","stars":451,"forks":54,"topics":["ai","automation","browser","chrome-extension","developer-tools","firefox-addon","llm","markdown"],"archived":false,"github_pushed_at":"2026-08-02T18:58:53+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/jatinkrmalik-llmfeeder","markdown_url":"https://www.graphcanon.com/tools/jatinkrmalik-llmfeeder.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/jatinkrmalik-llmfeeder","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=jatinkrmalik-llmfeeder","shared_categories":[]}]}}