{"data":{"slug":"jia-lab-research-mgm","name":"MGM","tagline":"Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models","github_url":"https://github.com/JIA-Lab-research/MGM","owner":"JIA-Lab-research","repo":"MGM","owner_avatar_url":"https://avatars.githubusercontent.com/u/64006090?v=4","primary_language":"Python","stars":3331,"forks":276,"topics":["generation","large-language-models","vision-language-model"],"archived":false,"github_pushed_at":"2024-05-04T14:36:51+00:00","maintenance_label":"Dormant","stars_delta_30d":1,"url":"https://www.graphcanon.com/tools/jia-lab-research-mgm","markdown_url":"https://www.graphcanon.com/tools/jia-lab-research-mgm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/jia-lab-research-mgm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=jia-lab-research-mgm","description":"Official repo for \"Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models\"","homepage_url":null,"license":"Apache-2.0","open_issues":61,"watchers":26,"ai_summary":"A repository focused on developing a multi-modal vision-language model for generation tasks. It requires specific version of Transformers and additional packages for full functionality.","readme_excerpt":"## Install\nPlease follow the instructions below to install the required packages.\n\nNOTE: If you want to use the 2B version, please ensure to install the latest version Transformers (>=4.38.0).\n\n1. Clone this repository\n```bash\ngit clone https://github.com/dvlab-research/MGM.git\n```\n\n2. Install Package\n```bash\nconda create -n mgm python=3.10 -y\nconda activate mgm\ncd MGM\npip install --upgrade pip  # enable PEP 660 support\npip install -e .\n```\n\n3. Install additional packages for training cases\n```bash\npip install ninja\npip install flash-attn --no-build-isolation\n```\n\n---\n\n## License\n\n\n\n\nThe data and checkpoint is intended and licensed for research use only. They are also restricted to uses that follow the license agreement of LLaVA, LLaMA, Vicuna and GPT-4. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes.","github_created_at":"2024-03-26T14:48:45+00:00","created_at":"2026-07-07T17:35:51.534667+00:00","updated_at":"2026-08-18T00:02:02.878198+00:00","categories":[{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"additional-packages-training-cases","name":"additional-packages-training-cases"},{"slug":"generation","name":"generation"},{"slug":"large-language-models","name":"large language models"},{"slug":"multi-modality","name":"multi-modality"},{"slug":"research-only-use","name":"research-only-use"},{"slug":"transformers-versioning","name":"transformers-versioning"},{"slug":"vision-language-model","name":"vision-language-model"}],"trust":{"provenance":{"is_fork":false,"github_id":777812139,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-18T00:02:02.084Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":835,"last_release_at":null,"stars_delta_30d":1,"open_issues_delta_30d":0},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:06:28.384Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-18T00:02:02.592Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-18T00:02:02.592Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-18T00:02:02.592Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When working on projects requiring integration of text and visual data for generation tasks.","If the latest version (>=4.38.0) of Transformers can be installed without conflicts."],"when_not_to_use":["Avoid if your project requires commercial licensing, as MGM is strictly research-use only under CC BY NC 4.0.","Not suitable if you are unable to update or ensure the availability of required Python packages like flash-attn and ninja for training purposes."],"source":"enrich:decision_facts","observed_at":"2026-07-12T13:56:09.001Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"MGM offers a focused approach on multi-modal vision-language generation tasks with specific dependency requirements."}]}}