{"data":{"slug":"open-compass-vlmevalkit","name":"VLMEvalKit","tagline":"An open-source evaluation toolkit for large vision-language models","github_url":"https://github.com/open-compass/VLMEvalKit","owner":"open-compass","repo":"VLMEvalKit","owner_avatar_url":"https://avatars.githubusercontent.com/u/143521324?v=4","primary_language":"Python","stars":4345,"forks":745,"topics":["chatgpt","claude","clip","computer-vision","evaluation","gemini","gpt","gpt-4v","gpt4","large-language-models","llava","llm","multi-modal","openai","openai-api","pytorch","qwen","vit","vqa"],"archived":false,"github_pushed_at":"2026-08-17T17:16:26+00:00","maintenance_label":"Very active","stars_delta_30d":60,"url":"https://www.graphcanon.com/tools/open-compass-vlmevalkit","markdown_url":"https://www.graphcanon.com/tools/open-compass-vlmevalkit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/open-compass-vlmevalkit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=open-compass-vlmevalkit","description":"Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks","homepage_url":"https://huggingface.co/spaces/opencompass/open_vlm_leaderboard","license":"Apache-2.0","open_issues":285,"watchers":13,"ai_summary":"VLMEvalKit is an open-source Python-based evaluation toolkit for large vision-language models (LVLMs). It allows for streamlined evaluation on various benchmarks without the heavy workload of data preparation.","readme_excerpt":"<b>A Toolkit for Evaluating Large Vision-Language Models. </b>\n\n[![][github-contributors-shield]][github-contributors-link] • [![][github-forks-shield]][github-forks-link] • [![][github-stars-shield]][github-stars-link] • [![][github-issues-shield]][github-issues-link] • [![][github-license-shield]][github-license-link]\n\nEnglish | [简体中文](/docs/zh-CN/README_zh-CN.md) | [日本語](/docs/ja/README_ja.md)\n\n<a href=\"https://rank.opencompass.org.cn/leaderboard-multimodal\">🏆 OC Learderboard </a> •\n<a href=\"#%EF%B8%8F-quickstart\">🏗️Quickstart </a> •\n<a href=\"#-datasets-models-and-evaluation-results\">📊Datasets & Models </a> •\n<a href=\"#%EF%B8%8F-development-guide\">🛠️Development </a>\n\n<a href=\"https://huggingface.co/spaces/opencompass/open_vlm_leaderboard\">🤗 HF Leaderboard</a> •\n<a href=\"https://huggingface.co/datasets/VLMEval/OpenVLMRecords\">🤗 Evaluation Records</a> •\n<a href=\"https://huggingface.co/spaces/opencompass/openvlm_video_leaderboard\">🤗 HF Video Leaderboard</a> •\n\n<a href=\"https://discord.gg/evDT4GZmxN\">🔊 Discord</a> •\n<a href=\"https://www.arxiv.org/abs/2407.11691\">📝 Report</a> •\n<a href=\"#-the-goal-of-vlmevalkit\">🎯Goal </a> •\n<a href=\"#%EF%B8%8F-citation\">🖊️Citation </a>\n</div>\n\n**VLMEvalKit** (the python package name is **vlmeval**) is an **open-source evaluation toolkit** of **large vision-language models (LVLMs)**. It enables **one-command evaluation** of LVLMs on various benchmarks, without the heavy workload of data preparation under multiple repositories. In VLMEvalKit, we adopt **generation-based evaluation** for all LVLMs, and provide the evaluation results obtained with both **exact matching** and **LLM-based answer extraction**.\n\n## Recent Codebase Changes\n- **[2025-09-12]** **Major Update: Improved Handling for Models with Thinking Mode**\n\n    A new feature in [PR 1229](https://github.com/open-compass/VLMEvalKit/pull/1175) that improves support for models with thinking mode. VLMEvalKit now allows for the use of a custom `split_thinking` function. **We strongly recommend this for models with thinking mode to ensure the accuracy of evaluation**.  To use this new functionality, please enable the Environment Variable: `SPLIT_THINK=True`. By default, the function will parse content within `<think>...</think>` tags and store it in the `thinking` key of the output. For more advanced customization, you can also create a `split_think` function for model. Please see the InternVL implementation for an example.\n- **[2025-09-12]** **Major Update: Improved Handling for Long Response(More than 16k/32k)**\n\n    A new feature in [PR 1229](https://github.com/open-compass/VLMEvalKit/pull/1175) that improves support for models with long response outputs. VLMEvalKit can now save prediction files in TSV format. **Since individual cells in an `.xlsx` file are limited to 32,767 characters, we strongly recommend using this feature for models that generate long responses (e.g., exceeding 16k or 32k tokens) to prevent data truncation.** To use this new functionality, please enable the Environment Variable: `PRED_FORMAT=tsv`.\n- **[2025-08-04]** In [PR 1175](https://github.com/open-compass/VLMEvalKit/pull/1175), we refine the `can_infer_option` and `can_infer_text`, which increasingly route the evaluation to LLM choice extractors and empirically leads to slight performance improvement for MCQ benchmarks.\n\n## 🆕 News\n\n- **[2026-04-08]** Supported [**Video-MME-v2**](https://github.com/MME-Benchmarks/Video-MME-v2). Video-MME-v2 is an authoritative benchmark towards the next stage in video understanding evaluation. 🔥🔥🔥\n- **[2025-07-07]** Supported [**SeePhys**](https://seephys.github.io/), which is a ​full spectrum multimodal benchmark for evaluating physics reasoning across different knowledge levels. thanks to [**Quinn777**](https://github.com/Quinn777) 🔥🔥🔥\n- **[2025-07-02]** Supported [**OvisU1**](https://huggingface.co/AIDC-AI/Ovis-U1-3B), thanks to [**liyang-7**](https://github.com/liyang-7) 🔥🔥🔥\n- **[2025-06-16]** Supported [**","github_created_at":"2023-12-01T08:18:11+00:00","created_at":"2026-07-07T17:35:22.627696+00:00","updated_at":"2026-08-17T18:01:09.769791+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"computer-vision","name":"computer-vision"},{"slug":"evaluation","name":"evaluation"},{"slug":"large-language-models","name":"large language models"},{"slug":"llm","name":"llm"},{"slug":"multi-modal","name":"multi-modal"}],"trust":{"provenance":{"is_fork":false,"github_id":725956246,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-17T18:01:07.432Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":0,"days_since_push":0,"last_release_at":"2025-06-21T17:45:35Z","stars_delta_30d":60,"open_issues_delta_30d":21},"security_summary":{"status":"findings","scanner":"osv@v1","low_count":16,"high_count":0,"last_scan_at":"2026-07-11T11:05:19.139Z","medium_count":0,"scan_profile":"deps","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-17T18:01:08.652Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-08-17T18:01:08.652Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-17T18:01:08.652Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When you need to evaluate models supporting thinking mode, as it provides a custom split_thinking function improving accuracy.","For handling long response outputs exceeding 16k or 32k tokens without data truncation by enabling TSV format saving for prediction files."],"when_not_to_use":["If your project requires evaluation tools that generate Excel files with individual cells larger than the default support of 32,767 characters and cannot switch to TSV format.","When you do not need generation-based evaluation methods with exact matching and LLM-based answer extraction."],"source":"enrich:decision_facts","observed_at":"2026-07-14T19:32:32.013Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"VLMEvalKit is an open-source Python evaluation toolkit for large vision-language models that offers one-command evaluation with support for various benchmarks and models."}]}}