Home/Evaluation & Observability/Awesome-Multimodal-Large-Language-Models
Awesome-Multimodal-Large-Language-Models logo

Awesome-Multimodal-Large-Language-Models

BradyFU/Awesome-Multimodal-Large-Language-Models

Latest Advances on Multimodal Large Language Models

GraphCanon updated 3d · GitHub synced 3d

18k stars1.1k forksLast push 5d

Decision brief

Awesome-Multimodal-Large-Language-Models is a curated collection of surveys and benchmarks focused on multimodal large language models (MLLMs), encompassing evaluation frameworks, interactive Omni MLLMs, and benchmarking

Good fit when

  • - You need comprehensive resources for evaluating multimodal LLMs and want access to the latest research findings in this area.
  • - You are interested in state-of-the-art benchmarks for testing and comparing different aspects of MLLM performance on tasks such as video understanding or inter-modal interaction.

Avoid when

  • - If your primary focus is on single-modality language models, without a need to integrate visual or audio elements.
  • - If you prefer tools that provide hands-on implementation guidance rather than surveys and benchmarks for theoretical exploration.

Observed Jul 11, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Very active (2d since push)
As of 3d
Provenance
Not a fork · Personal account
As of 3d
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models

How it fits your stack(9)

Typed graph edges - alternatives, integrations, successors, and dependencies. Ranked by relationship type, not raw GitHub stars.

Related

Relationship graph

Optional deeper exploration of typed edges and category neighbours.

Similar tools

Same-category neighbours not already linked as typed edges.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

Compilation of surveys and benchmarks related to multimodal large language models (MLLMs) including evaluation frameworks, interactive Omni MLLMs, and comprehensive benchmark datasets.

Capability facts

No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).

Categories

Graph entities

Tags

README

Awesome-Multimodal-Large-Language-Models

✨ Highlights of NJU-MiG

🔥🔥 Surveys of MLLMs | 💬 WeChat (MLLM微信交流群)

  • 🌟 MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
    arXiv 2025, Paper, Project

  • 🌟 A Survey of Unified Multimodal Understanding and Generation: Advances and Challenges
    arXiv 2025, Paper, Project

  • A Survey on Multimodal Large Language Models
    NSR 2024, Paper, Project


🔥🔥 VITA Series Omni MLLMs | 💬 WeChat (VITA微信交流群)

  • VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
    NeurIPS 2025 Highlight, Paper, Project

  • VITA: Towards Open-Source Interactive Omni Multimodal LLM
    arXiv 2024, Paper, Project

  • VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
    NeurIPS 2025, Paper, Project


🔥🔥 MME Series MLLM Benchmarks

  • 🔥 Video-MME-v2: Towards the Next Stage in Video Understanding Evaluation

  • 🌟 MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
    arXiv 2025, Paper, Project

  • MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
    NeurIPS 2025 DB Highlight, Paper, Dataset, Eval Tool, ✒️ Citation

  • Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
    CVPR 2025, Paper, Project, Dataset


Table of Contents

  • Awesome Papers
    • Multimodal Instruction Tuning (& Latest Works)
    • Multimodal Hallucination
    • Multimodal In-Context Learning
    • Multimodal Chain-of-Thought
    • LLM-Aided Visual Reasoning
    • Foundation Models
    • Evaluation
    • Multimodal RLHF
    • Others
  • Awesome Datasets
    • Datasets of Pre-Training for Alignment
    • Datasets of Multimodal Instruction Tuning
    • Datasets of In-Context Learning
    • [Datasets of Multimod

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.