Awesome-Datasets-Hub
Curated collection of datasets for Large Language Models (LLMs)
GraphCanon updated 3w · GitHub synced 3w
Decision brief
Awesome-Datasets-Hub offers a curated selection of datasets focusing particularly on medical AI, NLP, and multimodal applications, essential for training large language models.
Good fit when
- You need comprehensive datasets for clinical evaluation or specialized biomedical QA tasks.
- Specific domain expertise is required in medical AI or health-related LLM instruction tuning.
Avoid when
- Your focus is on domains outside of healthcare and medicine, where this tool might not provide adequate data diversity.
- You seek real-time dataset updates, as the specific update cadence for Awesome-Datasets-Hub isn't publicly specified.
Observed Jul 12, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Steady (38d since push)
- As of 3w
- Provenance
- Not a fork · Personal account
- As of 3w
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
git clone https://github.com/ahammadmejbah/Awesome-Datasets-HubSimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
A repository featuring a variety of datasets essential for training and evaluating large language models in domains such as medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.
Capability facts
No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).
Categories
Tags
README
⚡ Medical Datasets for LLM
| Serial | Dataset | Domain | Field / Task | Scale | Strength | Language | License |
|---|---|---|---|---|---|---|---|
| 01 | MedQA (USMLE) Jin et al. · 2021 | Medical QA · Licensing Exam | 12,723 Q | ||||
| 02 | MedMCQA Pal et al. · 2022 | Medical MCQ · Indian Licensing | 194K Q | ||||
| 03 | PubMedQA Jin et al. · 2019 | Biomedical QA · Yes/No/Maybe | 273K QA | ||||
| 04 | BioASQ Tsatsaronis et al. | Biomedical QA · Semantic Indexing | 5,600+ Q | ||||
| 05 | MASH-QA | Healthcare QA · Multi-span | 35K QA | ||||
| 06 | MedQuAD | Consumer Medical QA | 47K QA | ||||
| 07 | LiveQA Medical | Consumer Health QA | 634 Q | ||||
| 08 | MedS-Bench | Clinical Evaluation | 13 tasks | ||||
| 09 | COVID-QA | COVID-19 Reading QA | 2,019 QA | ||||
| 10 | MIMIC-III | ICU EHR + Clinical Notes | 46K patients | ||||
| 11 | MIMIC-IV | Hospital EHR Records | 299K patients | ||||
| 12 | n2c2 / i2b2 Challenges | Clinical NLP Challenges | Multi-year | ||||
| 13 | MedNLI | Clinical Textual Entailment | 14K pairs | ||||
| 14 | emrQA | Clinical QA over EHR | 1M+ QA | ||||
| 15 | MTSamples | Clinical Transcription Notes | 5K+ reports | ||||
| 16 | EHRNoteQA PhysioNet · 2023 | Clinical Note Question Answering | 2.4K QA | ||||
| 17 | Asclepius STAR-MPCC · 2023 | Synthetic Clinical Notes | 158K Notes | ||||
| 18 | [GatorTron Corpus](https://arx |
For agents
This page has a .md twin and JSON over the API.