{"data":{"slug":"bradyfu-awesome-multimodal-large-language-models","name":"Awesome-Multimodal-Large-Language-Models","tagline":"Latest Advances on Multimodal Large Language Models","github_url":"https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models","owner":"BradyFU","repo":"Awesome-Multimodal-Large-Language-Models","owner_avatar_url":"https://avatars.githubusercontent.com/u/54254631?v=4","primary_language":null,"stars":17978,"forks":1133,"topics":["chain-of-thought","in-context-learning","instruction-following","instruction-tuning","large-language-models","large-vision-language-model","large-vision-language-models","multi-modality","multimodal-chain-of-thought","multimodal-in-context-learning","multimodal-instruction-tuning","multimodal-large-language-models","visual-instruction-tuning"],"archived":false,"github_pushed_at":"2026-08-14T17:17:50+00:00","maintenance_label":"Very active","stars_delta_30d":29,"url":"https://www.graphcanon.com/tools/bradyfu-awesome-multimodal-large-language-models","markdown_url":"https://www.graphcanon.com/tools/bradyfu-awesome-multimodal-large-language-models.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/bradyfu-awesome-multimodal-large-language-models","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=bradyfu-awesome-multimodal-large-language-models","description":":sparkles::sparkles:Latest Advances on Multimodal Large Language Models","homepage_url":null,"license":null,"open_issues":111,"watchers":287,"ai_summary":"Compilation of surveys and benchmarks related to multimodal large language models (MLLMs) including evaluation frameworks, interactive Omni MLLMs, and comprehensive benchmark datasets.","readme_excerpt":"# Awesome-Multimodal-Large-Language-Models\n\n<p align=\"center\">\n    <img src=\"./images/mig_logo.png\" width=\"90%\" height=\"90%\">\n</p>\n\n## ✨ Highlights of NJU-MiG\n\n> 🔥🔥 **Surveys of MLLMs**  |  **[💬 WeChat (MLLM微信交流群)](./images/wechat-group.png)**\n\n- 🌟 **MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs**  \narXiv 2025, [Paper](https://arxiv.org/pdf/2411.15296.pdf), [Project](https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Benchmarks) \n\n- 🌟 **A Survey of Unified Multimodal Understanding and Generation: Advances and Challenges**  \narXiv 2025, [Paper](https://www.techrxiv.org/doi/pdf/10.36227/techrxiv.176289261.16802577), [Project](https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Unified) \n\n- **A Survey on Multimodal Large Language Models**  \nNSR 2024, [Paper](https://arxiv.org/pdf/2306.13549.pdf), [Project](https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models)\n\n\n---\n\n\n> 🔥🔥 **VITA Series Omni MLLMs** | **[💬 WeChat (VITA微信交流群)](https://github.com/VITA-MLLM/VITA/blob/main/asset/wechat-group.jpg)**\n\n- **VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction**  \nNeurIPS 2025 Highlight, [Paper](https://arxiv.org/pdf/2501.01957.pdf), [Project](https://github.com/VITA-MLLM/VITA)\n\n- **VITA: Towards Open-Source Interactive Omni Multimodal LLM**  \narXiv 2024, [Paper](https://arxiv.org/pdf/2408.05211.pdf), [Project](https://vita-home.github.io/)\n\n- **VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model**  \nNeurIPS 2025, [Paper](https://arxiv.org/pdf/2505.03739.pdf), [Project](https://github.com/VITA-MLLM/VITA-Audio)\n\n\n---\n\n\n> 🔥🔥 **MME Series MLLM Benchmarks**\n\n- 🔥 **Video-MME-v2: Towards the Next Stage in Video Understanding Evaluation**\n\n<p align=\"center\">\n    <img src=\"./images/video-mme-v2-logo.png\" width=\"100%\" height=\"100%\">\n</p>\n\n<font size=7><div align='center' > [[🍎 Project Page](https://video-mme-v2.netlify.app/)] [[📖 Paper](https://arxiv.org/pdf/2604.05015)] [[🤗 Dataset](https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2)] [[🏆 Leaderboard](https://video-mme-v2.netlify.app/#leaderboard)]  </div></font>\n\n- 🌟 **MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs**  \narXiv 2025, [Paper](https://arxiv.org/pdf/2411.15296.pdf), [Project](https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Benchmarks)\n\n- **MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models**  \nNeurIPS 2025 DB Highlight, [Paper](https://arxiv.org/pdf/2306.13394.pdf), [Dataset](https://huggingface.co/datasets/lmms-lab/MME), [Eval Tool](https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/blob/Evaluation/tools/eval_tool.zip), [✒️ Citation](./images/bib_mme.txt)\n\n- **Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis**  \nCVPR 2025, [Paper](https://arxiv.org/pdf/2405.21075.pdf), [Project](https://video-mme.github.io/), [Dataset](https://github.com/BradyFU/Video-MME?tab=readme-ov-file#-dataset)\n\n\n---\n\n<font size=5><center><b> Table of Contents </b> </center></font>\n- [Awesome Papers](#awesome-papers)\n  - [Multimodal Instruction Tuning (& Latest Works)](#multimodal-instruction-tuning--latest-works)\n  - [Multimodal Hallucination](#multimodal-hallucination)\n  - [Multimodal In-Context Learning](#multimodal-in-context-learning)\n  - [Multimodal Chain-of-Thought](#multimodal-chain-of-thought)\n  - [LLM-Aided Visual Reasoning](#llm-aided-visual-reasoning)\n  - [Foundation Models](#foundation-models)\n  - [Evaluation](#evaluation)\n  - [Multimodal RLHF](#multimodal-rlhf)\n  - [Others](#others)\n- [Awesome Datasets](#awesome-datasets)\n  - [Datasets of Pre-Training for Alignment](#datasets-of-pre-training-for-alignment)\n  - [Datasets of Multimodal Instruction Tuning](#datasets-of-multimodal-instruction-tuning)\n  - [Datasets of In-Context Learning](#datasets-of-in-context-learning)\n  - [Datasets of Multimod","github_created_at":"2023-05-19T03:02:29+00:00","created_at":"2026-07-07T17:33:26.899967+00:00","updated_at":"2026-08-17T00:02:12.20846+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"chain-of-thought","name":"chain-of-thought"},{"slug":"in-context-learning","name":"in-context-learning"},{"slug":"instruction-following","name":"instruction-following"},{"slug":"instruction-tuning","name":"instruction-tuning"},{"slug":"large-language-models","name":"large language models"},{"slug":"multi-modality","name":"multi-modality"},{"slug":"multimodal-large-language-models","name":"multimodal-large-language-models"},{"slug":"visual-instruction-tuning","name":"visual-instruction-tuning"}],"trust":{"provenance":{"is_fork":false,"github_id":642642539,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-17T00:02:11.483Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":0,"days_since_push":2,"last_release_at":null,"stars_delta_30d":29,"open_issues_delta_30d":4},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:01:03.154Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-17T00:02:11.929Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["- You need comprehensive resources for evaluating multimodal LLMs and want access to the latest research findings in this area.","- You are interested in state-of-the-art benchmarks for testing and comparing different aspects of MLLM performance on tasks such as video understanding or inter-modal interaction."],"when_not_to_use":["- If your primary focus is on single-modality language models, without a need to integrate visual or audio elements.","- If you prefer tools that provide hands-on implementation guidance rather than surveys and benchmarks for theoretical exploration."],"source":"enrich:decision_facts","observed_at":"2026-07-11T14:35:40.821Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Awesome-Multimodal-Large-Language-Models is a curated collection of surveys and benchmarks focused on multimodal large language models (MLLMs), encompassing evaluation frameworks, interactive Omni MLLMs, and benchmarking"}]}}