Awesome-Multimodal-Large-Language-Models
Latest Advances on Multimodal Large Language Models
GraphCanon updated 3d · GitHub synced 3d
Decision brief
Awesome-Multimodal-Large-Language-Models is a curated collection of surveys and benchmarks focused on multimodal large language models (MLLMs), encompassing evaluation frameworks, interactive Omni MLLMs, and benchmarking
Good fit when
- - You need comprehensive resources for evaluating multimodal LLMs and want access to the latest research findings in this area.
- - You are interested in state-of-the-art benchmarks for testing and comparing different aspects of MLLM performance on tasks such as video understanding or inter-modal interaction.
Avoid when
- - If your primary focus is on single-modality language models, without a need to integrate visual or audio elements.
- - If you prefer tools that provide hands-on implementation guidance rather than surveys and benchmarks for theoretical exploration.
Observed Jul 11, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Very active (2d since push)
- As of 3d
- Provenance
- Not a fork · Personal account
- As of 3d
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
git clone https://github.com/BradyFU/Awesome-Multimodal-Large-Language-ModelsHow it fits your stack(9)
Typed graph edges - alternatives, integrations, successors, and dependencies. Ranked by relationship type, not raw GitHub stars.
Related
Relationship graph
Optional deeper exploration of typed edges and category neighbours.
Similar tools
Same-category neighbours not already linked as typed edges.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
Compilation of surveys and benchmarks related to multimodal large language models (MLLMs) including evaluation frameworks, interactive Omni MLLMs, and comprehensive benchmark datasets.
Capability facts
No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).
Categories
Graph entities
Tags
README
Awesome-Multimodal-Large-Language-Models
✨ Highlights of NJU-MiG
🔥🔥 Surveys of MLLMs | 💬 WeChat (MLLM微信交流群)
-
🌟 MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
arXiv 2025, Paper, Project -
🌟 A Survey of Unified Multimodal Understanding and Generation: Advances and Challenges
arXiv 2025, Paper, Project -
A Survey on Multimodal Large Language Models
NSR 2024, Paper, Project
🔥🔥 VITA Series Omni MLLMs | 💬 WeChat (VITA微信交流群)
-
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
NeurIPS 2025 Highlight, Paper, Project -
VITA: Towards Open-Source Interactive Omni Multimodal LLM
arXiv 2024, Paper, Project -
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
NeurIPS 2025, Paper, Project
🔥🔥 MME Series MLLM Benchmarks
- 🔥 Video-MME-v2: Towards the Next Stage in Video Understanding Evaluation
-
🌟 MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
arXiv 2025, Paper, Project -
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
NeurIPS 2025 DB Highlight, Paper, Dataset, Eval Tool, ✒️ Citation -
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
CVPR 2025, Paper, Project, Dataset
Table of Contents
- Awesome Papers
- Multimodal Instruction Tuning (& Latest Works)
- Multimodal Hallucination
- Multimodal In-Context Learning
- Multimodal Chain-of-Thought
- LLM-Aided Visual Reasoning
- Foundation Models
- Evaluation
- Multimodal RLHF
- Others
- Awesome Datasets
- Datasets of Pre-Training for Alignment
- Datasets of Multimodal Instruction Tuning
- Datasets of In-Context Learning
- [Datasets of Multimod
For agents
This page has a .md twin and JSON over the API.