{"data":{"slug":"opengvlab-ask-anything","name":"Ask-Anything","tagline":"ChatGPT with enhanced video understanding capabilities","github_url":"https://github.com/OpenGVLab/Ask-Anything","owner":"OpenGVLab","repo":"Ask-Anything","owner_avatar_url":"https://avatars.githubusercontent.com/u/94522163?v=4","primary_language":"Python","stars":3345,"forks":268,"topics":["big-model","captioning-videos","chat","chatgpt","foundation-models","gradio","langchain","large-language-models","large-model","stablelm","video","video-question-answering","video-understanding"],"archived":false,"github_pushed_at":"2026-07-17T10:31:09+00:00","maintenance_label":"Steady","stars_delta_30d":1,"url":"https://www.graphcanon.com/tools/opengvlab-ask-anything","markdown_url":"https://www.graphcanon.com/tools/opengvlab-ask-anything.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opengvlab-ask-anything","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opengvlab-ask-anything","description":"[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.","homepage_url":"https://vchat.opengvlab.com/","license":"MIT","open_issues":75,"watchers":32,"ai_summary":"A repository focused on end-to-end chatbots that leverage large language models (LLMs) to process and interact with videos. Supports various LLMs including ChatGPT, miniGPT4, StableLM, and MOSS.","readme_excerpt":"# 🦜 VideoChat Family: Ask-Anything \n\n\n | \n<a src=\"https://img.shields.io/discord/1099920215724277770?label=Discord&logo=discord\" href=\"https://discord.gg/A2Ex6Pph6A\">\n    <img src=\"https://img.shields.io/discord/1099920215724277770?label=Discord&logo=discord\">\n</a> | \n<a src=\"https://img.shields.io/badge/cs.CV-2305.06355-b31b1b?logo=arxiv&logoColor=red\" href=\"https://arxiv.org/abs/2305.06355\"> <img src=\"https://img.shields.io/badge/cs.CV-2305.06355-b31b1b?logo=arxiv&logoColor=red\">\n</a>| <a src=\"https://img.shields.io/badge/cs.CV-2311.17005-b31b1b?logo=arxiv&logoColor=red\" href=\"https://arxiv.org/abs/2311.17005\"> <img src=\"https://img.shields.io/badge/cs.CV-2311.17005-b31b1b?logo=arxiv&logoColor=red\">\n</a>| \n<a src=\"https://img.shields.io/twitter/follow/opengvlab?style=social\" href=\"https://twitter.com/opengvlab\">\n    <img src=\"https://img.shields.io/twitter/follow/opengvlab?style=social\"> </a>\n</a>\n<br>\n<a href=\"https://huggingface.co/spaces/OpenGVLab/VideoChatGPT\"><img src=\"https://huggingface.co/datasets/huggingface/badges/raw/main/open-in-hf-spaces-sm-dark.svg\" alt=\"Open in Spaces\"> [VideoChat-7B-8Bit] End2End ChatBOT for video and image. </a> <a href=\"https://huggingface.co/spaces/OpenGVLab/InternVideo2-Chat-8B-HD\"><img src=\"https://huggingface.co/datasets/huggingface/badges/raw/main/open-in-hf-spaces-sm-dark.svg\" alt=\"Open in Spaces\"> [InternVideo2-Chat-8B-HD]</a>\n\n\n[中文 README 及 中文交流群](README_cn.md) | [Paper](https://arxiv.org/abs/2305.06355)\n\n\n\n⭐️: We are also working on a updated version, stay tuned! \n    \n\n\n\n# :fire: Updates\n- **2026/07/17**: 🚀🚀 We release [VideoChat3](https://github.com/MCG-NJU/VideoChat3), a fully open, efficient 4B Video MLLM for general, long-form, and streaming video understanding. VideoChat3 improves 18/19 offline and 10/11 streaming metrics over Qwen3-VL-4B. We release the model weights, code, training recipes, and complete datasets. Check out our [paper](https://arxiv.org/abs/2607.14935) and [homepage](https://mcg-nju.github.io/VideoChat3/)!\n- **2025/01/18**: We release [videochat-flash](https://github.com/OpenGVLab/VideoChat-Flash) and [videochat-tpo](https://github.com/OpenGVLab/TPO) to extend MLLMs' capabilities on both long and accurate video understanding. [videochat-flash](https://github.com/OpenGVLab/VideoChat-Flash) sets new records in mutiple video benchmarks (for both short and long videos), improving code usability by leveaging [LLaVA](https://github.com/LLaVA-VL/LLaVA-NeXT) and others. [videochat-tpo](https://github.com/OpenGVLab/TPO) exploits classical vision task annotations (e.g. tracking) to optimize MLLMs in a DPO manner, enhancing MLLMs' performance and enabling capabilities in tracking, segmentation, and more.\n- **2024/06/25**: We release the [branch of videochat2 using `vllm`](https://github.com/OpenGVLab/Ask-Anything/tree/vllm), speed up the inference of videochat2.\n- **2024/06/19**: 🎉🎉 Our VideoChat2 achieves the best performances among the open-sourced VideoLLMs on [MLVU](https://github.com/JUNJIE99/MLVU), a multi-task long video understanding benchmark.\n- **2024/06/13**: Fix some bug and give testing scripts/\n    - :warning: We replace some repeated  (~30) QAs in MVBench, which may only affect the results by 0.5%.\n    - :loudspeaker: We give the scripts for testing [EgoSchema](https://github.com/egoschema/EgoSchema/tree/main) and [Video-MME](https://github.com/BradyFU/Video-MME/tree/main), please check the [demo_mistral.ipynb](./video_chat2/demo/demo_mistral.ipynb) and [demo_mistral_hd.ipynb](./video_chat2/demo/demo_mistral_hd.ipynb).\n- **2024/06/07**: :fire::fire::fire: We release **VideoChat2_HD**, which is fine-tuned with high-resolution data and is capable of handling more diverse tasks. It showcases better performance on different benchmarks, especially for detailed captioning. Furthermore, it achieves **54.8% on [Video-MME](https://github.com/BradyFU/Video-MME/tree/main)**, the best score among 7B MLLMs. Have a try! 🏃🏻‍♀️🏃🏻\n- **2024/06/06**: We release **","github_created_at":"2023-04-19T09:49:10+00:00","created_at":"2026-07-07T17:35:49.883309+00:00","updated_at":"2026-08-18T00:02:00.908251+00:00","categories":[{"slug":"computer-vision","name":"Computer Vision","url":"https://www.graphcanon.com/categories/computer-vision","markdown_url":"https://www.graphcanon.com/categories/computer-vision.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/computer-vision"},{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"chatbot","name":"chatbot"},{"slug":"langchain","name":"langchain"},{"slug":"large-language-models","name":"large language models"},{"slug":"video-understanding","name":"video-understanding"}],"trust":{"provenance":{"is_fork":false,"github_id":629922458,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-18T00:02:00.206Z","maintenance":{"label":"Steady","score":60,"methodology":"github_public_v1","releases_90d":0,"days_since_push":31,"last_release_at":null,"stars_delta_30d":1,"open_issues_delta_30d":-1},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:06:21.240Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-18T00:02:00.622Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-08-18T00:02:00.622Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-18T00:02:00.622Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When you need advanced video and image processing with large language models for captioning and QA tasks","If your projects require handling both long videos and detailed task annotations with optimized performance"],"when_not_to_use":["Avoid if only text-based interactions are needed, as Ask-Anything focuses on video understanding","Not suitable for real-time applications requiring ultra-fast inference without compromising on accuracy"],"source":"enrich:decision_facts","observed_at":"2026-07-12T13:44:46.821Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Ask-Anything is an end-to-end video chatbot framework leveraging LLMs like ChatGPT, miniGPT4, StableLM for enhanced video understanding."}]}}