{"data":{"node":{"slug":"conardli-easy-dataset","name":"easy-dataset","tagline":"A powerful tool for creating datasets for LLM fine-tuning, RAG, and evaluation","github_url":"https://github.com/ConardLi/easy-dataset","owner":"ConardLi","repo":"easy-dataset","owner_avatar_url":"https://avatars.githubusercontent.com/u/30708545?v=4","primary_language":"JavaScript","stars":14792,"forks":1523,"topics":["dataset","fine-tuning","javascript","llm","rag"],"archived":false,"github_pushed_at":"2026-05-01T15:03:32+00:00","maintenance_label":"Slowing","stars_delta_30d":125,"url":"https://www.graphcanon.com/tools/conardli-easy-dataset","markdown_url":"https://www.graphcanon.com/tools/conardli-easy-dataset.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/conardli-easy-dataset","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=conardli-easy-dataset"},"categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"dataset","name":"dataset"},{"slug":"fine-tuning","name":"fine-tuning"},{"slug":"javascript","name":"javascript"},{"slug":"llm","name":"llm"},{"slug":"rag","name":"rag"}],"edges":[{"type":"related","direction":"out","explanation":"While promptfoo focuses on evaluating LLM apps through prompts, easy-dataset assists in creating datasets which could potentially enhance the evaluation process by providing comprehensive data for testing.","successor_context":null,"tool":{"slug":"promptfoo-promptfoo","name":"promptfoo","tagline":"Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.","github_url":"https://github.com/promptfoo/promptfoo","owner":"promptfoo","repo":"promptfoo","owner_avatar_url":"https://avatars.githubusercontent.com/u/137907881?v=4","primary_language":"TypeScript","stars":23838,"forks":2147,"topics":["ci","ci-cd","cicd","evaluation","evaluation-framework","llm","llm-eval","llm-evaluation","llm-evaluation-framework","llmops","pentesting","prompt-engineering","prompt-testing","prompts","rag","red-teaming","testing","vulnerability-scanners"],"archived":false,"github_pushed_at":"2026-08-01T23:47:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/promptfoo-promptfoo","markdown_url":"https://www.graphcanon.com/tools/promptfoo-promptfoo.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/promptfoo-promptfoo","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=promptfoo-promptfoo"}},{"type":"integrates_with","direction":"out","explanation":"easy-dataset creates datasets for fine-tuning and evaluation, which complements transformers' capacity for training large language models.","successor_context":null,"tool":{"slug":"huggingface-transformers","name":"transformers","tagline":"Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models","github_url":"https://github.com/huggingface/transformers","owner":"huggingface","repo":"transformers","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":164121,"forks":34249,"topics":["audio","deep-learning","deepseek","gemma","glm","hacktoberfest","llm","machine-learning","model-hub","natural-language-processing","nlp","pretrained-models","python","pytorch","pytorch-transformers","qwen","speech-recognition","transformer","vlm"],"archived":false,"github_pushed_at":"2026-08-15T22:28:12+00:00","maintenance_label":"Very active","stars_delta_30d":1457,"url":"https://www.graphcanon.com/tools/huggingface-transformers","markdown_url":"https://www.graphcanon.com/tools/huggingface-transformers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-transformers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-transformers"}},{"type":"integrates_with","direction":"out","explanation":"RAG_Techniques focuses on retrieval-augmented generation systems, and easy-dataset is useful for creating datasets that can enrich the RAG system's knowledge base.","successor_context":null,"tool":{"slug":"nirdiamant-rag-techniques","name":"RAG_Techniques","tagline":"Showcases advanced techniques for Retrieval-Augmented Generation (RAG) systems with detailed notebook tutorials.","github_url":"https://github.com/NirDiamant/RAG_Techniques","owner":"NirDiamant","repo":"RAG_Techniques","owner_avatar_url":"https://avatars.githubusercontent.com/u/28316913?v=4","primary_language":"Jupyter Notebook","stars":29076,"forks":3540,"topics":["agentic-rag","ai","embeddings","generative-ai","gpt","langchain","llama-index","llm","llms","machine-learning","nlp","openai","python","rag","retrieval-augmented-generation","semantic-search","tutorials","vector-database"],"archived":false,"github_pushed_at":"2026-08-15T00:52:05+00:00","maintenance_label":"Very active","stars_delta_30d":455,"url":"https://www.graphcanon.com/tools/nirdiamant-rag-techniques","markdown_url":"https://www.graphcanon.com/tools/nirdiamant-rag-techniques.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nirdiamant-rag-techniques","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nirdiamant-rag-techniques"}},{"type":"related","direction":"out","explanation":"The llm-cookbook provides tutorials for developing with LLMs, and easy-dataset could serve as a tool to prepare datasets needed in those examples.","successor_context":null,"tool":{"slug":"datawhalechina-llm-cookbook","name":"llm-cookbook","tagline":"面向开发者的 LLM 入门教程，吴恩达大模型系列课程中文版","github_url":"https://github.com/datawhalechina/llm-cookbook","owner":"datawhalechina","repo":"llm-cookbook","owner_avatar_url":"https://avatars.githubusercontent.com/u/46047812?v=4","primary_language":"Jupyter Notebook","stars":24544,"forks":2953,"topics":["cookbook","llm"],"archived":false,"github_pushed_at":"2025-06-12T14:48:07+00:00","maintenance_label":"Dormant","stars_delta_30d":123,"url":"https://www.graphcanon.com/tools/datawhalechina-llm-cookbook","markdown_url":"https://www.graphcanon.com/tools/datawhalechina-llm-cookbook.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datawhalechina-llm-cookbook","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datawhalechina-llm-cookbook"}},{"type":"integrates_with","direction":"out","explanation":"Khoj acts as an AI second brain and integrates knowledge retrieval; easy-dataset can help prepare the datasets Khoj uses to enhance its capabilities.","successor_context":null,"tool":{"slug":"khoj-ai-khoj","name":"khoj","tagline":"Your AI second brain. Self-hostable.","github_url":"https://github.com/khoj-ai/khoj","owner":"khoj-ai","repo":"khoj","owner_avatar_url":"https://avatars.githubusercontent.com/u/134046886?v=4","primary_language":"Python","stars":36512,"forks":2386,"topics":["agent","ai","assistant","chat","chatgpt","emacs","image-generation","llama3","llamacpp","llm","obsidian","obsidian-md","offline-llm","productivity","rag","research","self-hosted","semantic-search","stt","whatsapp-ai"],"archived":false,"github_pushed_at":"2026-08-02T01:55:40+00:00","maintenance_label":"Active","stars_delta_30d":702,"url":"https://www.graphcanon.com/tools/khoj-ai-khoj","markdown_url":"https://www.graphcanon.com/tools/khoj-ai-khoj.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/khoj-ai-khoj","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=khoj-ai-khoj"}},{"type":"related","direction":"out","explanation":"Ray is a unified framework that provides capabilities for scaling out various components needed in an LLM ecosystem, including tools like Easy-Dataset which would benefit from scalable computing resources.","successor_context":null,"tool":{"slug":"ray-project-ray","name":"ray","tagline":"Ray is an AI compute engine with a core distributed runtime and AI Libraries for accelerating ML workloads.","github_url":"https://github.com/ray-project/ray","owner":"ray-project","repo":"ray","owner_avatar_url":"https://avatars.githubusercontent.com/u/22125274?v=4","primary_language":"Python","stars":43526,"forks":7929,"topics":["data-science","deep-learning","deployment","distributed","hyperparameter-optimization","hyperparameter-search","large-language-models","llm","llm-inference","llm-serving","machine-learning","optimization","parallel","python","pytorch","ray","reinforcement-learning","rllib","serving","tensorflow"],"archived":false,"github_pushed_at":"2026-08-16T00:26:16+00:00","maintenance_label":"Very active","stars_delta_30d":270,"url":"https://www.graphcanon.com/tools/ray-project-ray","markdown_url":"https://www.graphcanon.com/tools/ray-project-ray.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ray-project-ray","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ray-project-ray"}},{"type":"related","direction":"in","explanation":null,"successor_context":null,"tool":{"slug":"evidentlyai-evidently","name":"evidently","tagline":"An open-source ML and LLM observability framework.","github_url":"https://github.com/evidentlyai/evidently","owner":"evidentlyai","repo":"evidently","owner_avatar_url":"https://avatars.githubusercontent.com/u/75031056?v=4","primary_language":"Jupyter Notebook","stars":7790,"forks":895,"topics":["data-drift","data-quality","data-science","data-validation","generative-ai","hacktoberfest","html-report","jupyter-notebook","llm","llmops","machine-learning","mlops","model-monitoring","pandas-dataframe"],"archived":false,"github_pushed_at":"2026-08-05T16:29:57+00:00","maintenance_label":"Very active","stars_delta_30d":117,"url":"https://www.graphcanon.com/tools/evidentlyai-evidently","markdown_url":"https://www.graphcanon.com/tools/evidentlyai-evidently.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/evidentlyai-evidently","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=evidentlyai-evidently"}}],"neighbours":[{"slug":"huggingface-datasets","name":"datasets","tagline":"Largest hub of ready-to-use datasets for AI models","github_url":"https://github.com/huggingface/datasets","owner":"huggingface","repo":"datasets","owner_avatar_url":"https://avatars.githubusercontent.com/u/25720743?v=4","primary_language":"Python","stars":21791,"forks":3322,"topics":["ai","artificial-intelligence","computer-vision","dataset-hub","datasets","deep-learning","huggingface","llm","machine-learning","natural-language-processing","nlp","numpy","pandas","pytorch","speech","tensorflow"],"archived":false,"github_pushed_at":"2026-07-30T11:23:49+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/huggingface-datasets","markdown_url":"https://www.graphcanon.com/tools/huggingface-datasets.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/huggingface-datasets","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=huggingface-datasets","shared_categories":["data-retrieval"]},{"slug":"datajuicer-data-juicer","name":"data-juicer","tagline":"Data processing for and with foundation models","github_url":"https://github.com/datajuicer/data-juicer","owner":"datajuicer","repo":"data-juicer","owner_avatar_url":"https://avatars.githubusercontent.com/u/223222708?v=4","primary_language":"Python","stars":6897,"forks":404,"topics":["data","data-analysis","data-pipeline","data-processing","data-science","data-visualization","foundation-models","instruction-tuning","large-language-models","llm","llms","multi-modal","pre-training","synthetic-data"],"archived":false,"github_pushed_at":"2026-08-13T09:19:31+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/datajuicer-data-juicer","markdown_url":"https://www.graphcanon.com/tools/datajuicer-data-juicer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datajuicer-data-juicer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datajuicer-data-juicer","shared_categories":["model-training","data-retrieval"]},{"slug":"zjh-819-llmdatahub","name":"LLMDataHub","tagline":"Curated Collection of Datasets for LLM Training","github_url":"https://github.com/Zjh-819/LLMDataHub","owner":"Zjh-819","repo":"LLMDataHub","owner_avatar_url":"https://avatars.githubusercontent.com/u/75557474?v=4","primary_language":null,"stars":3413,"forks":234,"topics":["chatbot","chatgpt","dataset","llm"],"archived":false,"github_pushed_at":"2023-11-28T09:41:28+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/zjh-819-llmdatahub","markdown_url":"https://www.graphcanon.com/tools/zjh-819-llmdatahub.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/zjh-819-llmdatahub","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=zjh-819-llmdatahub","shared_categories":["model-training"]},{"slug":"nvidia-nemo-curator","name":"Curator","tagline":"Scalable data pre-processing and curation toolkit for LLMs","github_url":"https://github.com/NVIDIA-NeMo/Curator","owner":"NVIDIA-NeMo","repo":"Curator","owner_avatar_url":"https://avatars.githubusercontent.com/u/213689629?v=4","primary_language":"Python","stars":1681,"forks":306,"topics":["data","data-curation","data-prep","data-preparation","data-processing","data-processing-pipelines","data-quality","datacuration","datarecipes","deduplication","fast-data-processing","fine-tuning","large-language-models","large-scale-data-processing","llm","llm-data-quality","llmapps","python","semantic-deduplication"],"archived":false,"github_pushed_at":"2026-07-23T23:30:09+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/nvidia-nemo-curator","markdown_url":"https://www.graphcanon.com/tools/nvidia-nemo-curator.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nvidia-nemo-curator","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nvidia-nemo-curator","shared_categories":["model-training","data-retrieval"]},{"slug":"datadreamer-dev-datadreamer","name":"DataDreamer","tagline":"Prompt. Generate Synthetic Data. Train & Align Models.","github_url":"https://github.com/datadreamer-dev/DataDreamer","owner":"datadreamer-dev","repo":"DataDreamer","owner_avatar_url":"https://avatars.githubusercontent.com/u/154913957?v=4","primary_language":"Python","stars":1117,"forks":58,"topics":["alignment","deep-learning","fine-tuning","gpt","instruction-tuning","llm","llmops","llms","machine-learning","natural-language-processing","nlp","nlp-library","openai","python","pytorch","synthetic-data","synthetic-dataset-generation","transformers"],"archived":false,"github_pushed_at":"2025-02-02T21:23:50+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/datadreamer-dev-datadreamer","markdown_url":"https://www.graphcanon.com/tools/datadreamer-dev-datadreamer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/datadreamer-dev-datadreamer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=datadreamer-dev-datadreamer","shared_categories":["model-training","data-retrieval"]},{"slug":"radi-cho-datasetgpt","name":"datasetGPT","tagline":"A command-line tool for generating textual and conversational datasets with LLMs.","github_url":"https://github.com/radi-cho/datasetGPT","owner":"radi-cho","repo":"datasetGPT","owner_avatar_url":"https://avatars.githubusercontent.com/u/12954909?v=4","primary_language":"Python","stars":300,"forks":20,"topics":["cli","dataset-generation","large-language-models","python3"],"archived":false,"github_pushed_at":"2023-08-25T16:39:10+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/radi-cho-datasetgpt","markdown_url":"https://www.graphcanon.com/tools/radi-cho-datasetgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/radi-cho-datasetgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=radi-cho-datasetgpt","shared_categories":["model-training","data-retrieval"]},{"slug":"zhulinsen-fastdatasets","name":"FastDatasets","tagline":"A powerful tool for creating high-quality training datasets for Large Language Models (LLMs)","github_url":"https://github.com/ZhuLinsen/FastDatasets","owner":"ZhuLinsen","repo":"FastDatasets","owner_avatar_url":"https://avatars.githubusercontent.com/u/42829555?v=4","primary_language":"Python","stars":222,"forks":43,"topics":["asyncio","dataset-generation","datasets","llm","python"],"archived":false,"github_pushed_at":"2025-08-31T06:40:43+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/zhulinsen-fastdatasets","markdown_url":"https://www.graphcanon.com/tools/zhulinsen-fastdatasets.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/zhulinsen-fastdatasets","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=zhulinsen-fastdatasets","shared_categories":["model-training","data-retrieval"]},{"slug":"ahammadmejbah-awesome-datasets-hub","name":"Awesome-Datasets-Hub","tagline":"Curated collection of datasets for Large Language Models (LLMs)","github_url":"https://github.com/ahammadmejbah/Awesome-Datasets-Hub","owner":"ahammadmejbah","repo":"Awesome-Datasets-Hub","owner_avatar_url":"https://avatars.githubusercontent.com/u/56669333?v=4","primary_language":null,"stars":146,"forks":40,"topics":["benchmark","benchmarking","deep-learning","deep-neural-networks","deeplearning","genetic-algorithm","llm","llm-evaluation","llm-inference","machine-learning","machine-learning-algorithms","machinelearning","neural-network"],"archived":false,"github_pushed_at":"2026-06-20T07:06:51+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/ahammadmejbah-awesome-datasets-hub","markdown_url":"https://www.graphcanon.com/tools/ahammadmejbah-awesome-datasets-hub.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ahammadmejbah-awesome-datasets-hub","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ahammadmejbah-awesome-datasets-hub","shared_categories":["data-retrieval"]}]}}