{"data":{"slug":"coobiw-mpp-llava","name":"MPP-LLaVA","tagline":"Multimodal Pipeline Parallel based on Qwen-LM for training large language models with support for video and image inputs.","github_url":"https://github.com/Coobiw/MPP-LLaVA","owner":"Coobiw","repo":"MPP-LLaVA","owner_avatar_url":"https://avatars.githubusercontent.com/u/48615375?v=4","primary_language":"Jupyter Notebook","stars":685,"forks":34,"topics":["deepspeed","fine-tuning","mllm","model-parallel","multimodal-large-language-models","pipeline-parallelism","pretraining","qwen","video-language-model","video-large-language-models"],"archived":false,"github_pushed_at":"2025-03-10T14:16:21+00:00","maintenance_label":"Dormant","stars_delta_30d":0,"url":"https://www.graphcanon.com/tools/coobiw-mpp-llava","markdown_url":"https://www.graphcanon.com/tools/coobiw-mpp-llava.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/coobiw-mpp-llava","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=coobiw-mpp-llava","description":"Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.","homepage_url":null,"license":null,"open_issues":9,"watchers":5,"ai_summary":"A project supporting the fine-tuning of multimodal large language models (MLLMs) like Qwen14B using pipeline parallelism to handle video/image/multi-image data. It allows training on consumer-grade GPUs such as RTX3090/4090 with 24GB VRAM.","readme_excerpt":"## Installation\n\n```bash\nconda create -n minigpt4qwen python=3.8 && conda activate minigpt4qwen\npip install -e .\n```\n\n---\n\n## License\n\n- 本仓库的许多代码是基于[Lavis](https://github.com/salesforce/LAVIS) 的，其采用 [BSD 3-Clause License](https://github.com/Vision-CAIR/MiniGPT-4/blob/main/LICENSE_Lavis.md).\n- 本仓库采用Qwen-7B-Chat，支持商用和科研、开发用途，其License为[LICENSE](https://github.com/QwenLM/Qwen/blob/main/LICENSE)","github_created_at":"2023-10-24T17:27:32+00:00","created_at":"2026-07-11T11:40:40.75995+00:00","updated_at":"2026-08-24T06:01:23.243372+00:00","categories":[{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"deepspeed","name":"deepspeed"},{"slug":"fine-tuning","name":"fine-tuning"},{"slug":"model-parallel","name":"model-parallel"},{"slug":"multimodal-large-language-models","name":"multimodal-large-language-models"},{"slug":"pipeline-parallelism","name":"pipeline-parallelism"},{"slug":"pretraining","name":"pretraining"},{"slug":"qwen","name":"qwen"},{"slug":"video-language-model","name":"video-language-model"}],"trust":{"provenance":{"is_fork":false,"github_id":709422137,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-24T06:01:22.502Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":531,"last_release_at":"2024-06-13T14:23:36Z","stars_delta_30d":0,"open_issues_delta_30d":0},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:40:41.875Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-24T06:01:22.979Z"},"languages":{"value":["jupyter notebook","python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-24T06:01:22.979Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["You are working with a limited GPU budget but need to fine-tune large MLLMs like Qwen14B using pipeline parallelism.","Handling video or image data is essential for your project but high-end hardware is not available."],"when_not_to_use":["High-performance and high-capacity GPUs are readily accessible, allowing other tools to leverage more comprehensive parallelisms beyond consumer-grade GPUs limitations.","The project does not require the handling of video or image data as inputs for MLLM fine-tuning."],"source":"enrich:decision_facts","observed_at":"2026-07-12T17:49:01.250Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"MPP-LLaVA enables efficient fine-tuning of Qwen-based multimodal language models on consumer-grade GPUs for video, image, or multiple images inputs."}]}}