Home/Model Training/LLMDataHub
LLMDataHub logo

LLMDataHub

Zjh-819/LLMDataHub

Curated Collection of Datasets for LLM Training

GraphCanon updated 2w · GitHub synced 2w

3.4k stars234 forksLast push 2y MIT

Decision brief

LLMDataHub offers a curated repository of datasets specifically designed for training large language models, including general alignment, domain-specific, pretraining, and multimodal datasets. It aids in the improvement,

Good fit when

  • - When you are looking to improve chatbot dialogue quality with specific datasets for instruction fine-tuning.
  • - If you need access to specialized chatbot training datasets that cover a variety of conversation domains and styles.

Avoid when

  • - Avoid using LLMDataHub if your project requires datasets not specifically curated for chatbot or language model training, as the focus here is on dialogue and instruction-specific data.
  • - Don't rely solely on this repository if you need real-time dataset curation; it may not always have the most recent or niche datasets compared to more dynamic sources.
Pricing:
freemium - Free access under MIT License, suitable for non-commercial use. Consult licensing terms if planning commercial usage.
Requirements:
The repository is accessible in various languages, though the specific dataset languages are detailed individually.

Observed Jul 11, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Dormant (982d since push)
As of 2w
Provenance
Not a fork · Personal account
As of 2w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/Zjh-819/LLMDataHub

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

LLMDataHub is a repository curating high-quality training datasets for large language models (LLMs), covering general alignment, domain-specific, pretraining, and multimodal datasets. It aids researchers and practitioners in easily finding relevant datasets to improve chatbot dialogue quality and language understanding.

Capability facts

No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).

Categories

Tags

README

LLMDataHub: Awesome Datasets for LLM Training


🔥 Alignment Datasets • 💡 Domain-specific Datasets • :atom: Pretraining Datasets 🖼️ Multimodal Datasets

GitHub last commit GitHub Repo stars

Introduction 📄

Large language models (LLMs), such as OpenAI's GPT series, Google's Bard, and Baidu's Wenxin Yiyan, are driving profound technological changes. Recently, with the emergence of open-source large model frameworks like LlaMa and ChatGLM, training an LLM is no longer the exclusive domain of resource-rich companies. Training LLMs by small organizations or individuals has become an important interest in the open-source community, with some notable works including Alpaca, Vicuna, and Luotuo. In addition to large model frameworks, large-scale and high-quality training corpora are also essential for training large language models. Currently, relevant open-source corpora in the community are still scattered. Therefore, the goal of this repository is to continuously collect high-quality training corpora for LLMs in the open-source community.

Training a chatbot LLM that can follow human instruction effectively requires access to high-quality datasets that cover a range of conversation domains and styles. In this repository, we provide a curated collection of datasets specifically designed for chatbot training, including links, size, language, usage, and a brief description of each dataset. Our goal is to make it easier for researchers and practitioners to identify and select the most relevant and useful datasets for their chatbot LLM training needs. Whether you're working on improving chatbot dialogue quality, response generation, or language understanding, this repository has something for you.

Contact 📬

If you want to contribute, you can contact:

Junhao Zhao 📧
Advised by Prof. Wanyun Cui

General Open Access Datasets for Alignment 🟢:

Type Tags 🏷️:

  • SFT: Supervised Finetune
    • Dialog: Each entry contains continuous conversations
    • Pairs: Each entry is an input-output pair
    • Context: Each entry has a context text and related QA pairs
  • PT: pretrain
  • CoT: Chain-of-Thought Finetune
  • RLHF: train reward model in Reinforcement Learning with Human Feedback

Datasets released in November 2023

Dataset nameUsed byTypeLanguageSizeDescription ️
helpSteer/RLHFEnglish37k instancesAn RLHF dataset that is annotated by human with helpfulness, correctness, coherence, complexity and verbosity measures
no_robots/SFTEnglish10k instanceHigh-quality human-created STF data, single turn.

Datasets released in September 2023

| Dataset name

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.