Home/Data & Retrieval/Awesome-Datasets-Hub
Awesome-Datasets-Hub logo

Awesome-Datasets-Hub

ahammadmejbah/Awesome-Datasets-Hub

Curated collection of datasets for Large Language Models (LLMs)

GraphCanon updated 3w · GitHub synced 3w

146 stars40 forksLast push 2mo

Decision brief

Awesome-Datasets-Hub offers a curated selection of datasets focusing particularly on medical AI, NLP, and multimodal applications, essential for training large language models.

Good fit when

  • You need comprehensive datasets for clinical evaluation or specialized biomedical QA tasks.
  • Specific domain expertise is required in medical AI or health-related LLM instruction tuning.

Avoid when

  • Your focus is on domains outside of healthcare and medicine, where this tool might not provide adequate data diversity.
  • You seek real-time dataset updates, as the specific update cadence for Awesome-Datasets-Hub isn't publicly specified.

Observed Jul 12, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Steady (38d since push)
As of 3w
Provenance
Not a fork · Personal account
As of 3w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/ahammadmejbah/Awesome-Datasets-Hub

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

A repository featuring a variety of datasets essential for training and evaluating large language models in domains such as medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.

Capability facts

No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).

Categories

Tags

README

Description

Badge image Badge image Badge image Badge image Badge image Badge image
Badge image Badge image

⚡ Medical Datasets for LLM

SerialDatasetDomainField / TaskScaleStrengthLanguageLicense
01MedQA (USMLE)
Jin et al. · 2021
Medical QA · Licensing Exam12,723 Q
02MedMCQA
Pal et al. · 2022
Medical MCQ · Indian Licensing194K Q
03PubMedQA
Jin et al. · 2019
Biomedical QA · Yes/No/Maybe273K QA
04BioASQ
Tsatsaronis et al.
Biomedical QA · Semantic Indexing5,600+ Q
05MASH-QAHealthcare QA · Multi-span35K QA
06MedQuADConsumer Medical QA47K QA
07LiveQA MedicalConsumer Health QA634 Q
08MedS-BenchClinical Evaluation13 tasks
09COVID-QACOVID-19 Reading QA2,019 QA
10MIMIC-IIIICU EHR + Clinical Notes46K patients
11MIMIC-IVHospital EHR Records299K patients
12n2c2 / i2b2 ChallengesClinical NLP ChallengesMulti-year
13MedNLIClinical Textual Entailment14K pairs
14emrQAClinical QA over EHR1M+ QA
15MTSamplesClinical Transcription Notes5K+ reports
16EHRNoteQA
PhysioNet · 2023
Clinical Note Question Answering2.4K QA
17Asclepius
STAR-MPCC · 2023
Synthetic Clinical Notes158K Notes
18[GatorTron Corpus](https://arx

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.