{"data":{"slug":"nvidia-megatron-lm","name":"Megatron-LM","tagline":"Ongoing research training transformer models at scale","github_url":"https://github.com/NVIDIA/Megatron-LM","owner":"NVIDIA","repo":"Megatron-LM","owner_avatar_url":"https://avatars.githubusercontent.com/u/1728152?v=4","primary_language":"Python","stars":17341,"forks":4333,"topics":["large-language-models","model-para","transformers"],"archived":false,"github_pushed_at":"2026-08-06T23:12:52+00:00","maintenance_label":"Very active","stars_delta_30d":353,"url":"https://www.graphcanon.com/tools/nvidia-megatron-lm","markdown_url":"https://www.graphcanon.com/tools/nvidia-megatron-lm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nvidia-megatron-lm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nvidia-megatron-lm","description":"Ongoing research training transformer models at scale","homepage_url":"https://docs.nvidia.com/megatron-core/developer-guide/latest/get-started/quickstart.html","license":"Other","open_issues":1112,"watchers":166,"ai_summary":"Megatron-LM is a repository from NVIDIA focused on the development and training of large-scale language models using transformer architectures. It provides tools for efficient parallelism strategies across multiple GPUs.","readme_excerpt":"## Getting Started\n\n**Install from PyPI:**\n\n```bash\nuv pip install megatron-core\n```\n\n**Or clone and install from source:**\n\n```bash\ngit clone https://github.com/NVIDIA/Megatron-LM.git\ncd Megatron-LM\nuv pip install -e .\n```\n\n> **Note:** Building from source can use a lot of memory. If the build runs out of memory, limit parallel compilation jobs by setting `MAX_JOBS` (for example, `MAX_JOBS=4 uv pip install -e .`).\n\nFor NVIDIA GPU Cloud (NGC) container setup and all installation options, review the **[Installation Guide](https://docs.nvidia.com/megatron-core/developer-guide/latest/get-started/install.html)**.\n\n- **[Your First Training Run](https://docs.nvidia.com/megatron-core/developer-guide/latest/get-started/quickstart.html)** - End-to-end training examples with data preparation\n- **[Parallelism Strategies](https://docs.nvidia.com/megatron-core/developer-guide/latest/user-guide/parallelism-guide.html)** - Scale training across GPUs with TP, PP, DP, EP, and CP\n- **[Contribution Guide](https://docs.nvidia.com/megatron-core/developer-guide/latest/developer/contribute.html)** - How to contribute to Megatron Core","github_created_at":"2019-03-21T16:15:52+00:00","created_at":"2026-07-07T17:33:33.267662+00:00","updated_at":"2026-08-07T00:02:05.765003+00:00","categories":[{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"large-language-models","name":"large language models"},{"slug":"model-para","name":"model-para"},{"slug":"transformers","name":"transformers"}],"trust":{"provenance":{"is_fork":false,"github_id":176982014,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-07T00:02:04.957Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":3,"days_since_push":0,"last_release_at":"2026-07-21T10:57:10Z","stars_delta_30d":353,"open_issues_delta_30d":122},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T10:35:41.370Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-07T00:02:05.473Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-07T00:02:05.473Z"},"license_spdx":{"value":"Other","source":"github.license","observed_at":"2026-08-07T00:02:05.473Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":{"notes":["Requires NVIDIA GPUs for optimized performance. Non-GPU usage is not supported or recommended.","Installation from source can be resource-intensive and may require limiting parallel compilation jobs to avoid running out of memory."],"min_ram_gb":32,"requires_docker":false},"constraints":{"min_ram_gb":32,"requires_docker":false},"when_to_use":["The tool is particularly beneficial when your project is GPU-centric and benefits from advanced parallelism techniques such as Tensor, Pipeline, Data, Expert, and Cluster Parallelisms (TP, PP, DP, EP,","and CP) provided by NVIDIA's ecosystem.","Megatron-LM is best utilized in scenarios where leveraging the specific optimizations for CUDA GPUs can lead to significant performance gains."],"when_not_to_use":["Avoid Megatron-LM if your computational setup does not include NVIDIA GPUs as it leverages GPU-specific features and parallelisms that may not be available or efficient on non-NVIDIA hardware.","If you need portability across various hardware without depending on proprietary optimizations, other tools might better serve your needs."],"source":"enrich:decision_facts","observed_at":"2026-07-11T14:44:07.028Z"},"constraint_facets":{"min_ram_gb":32,"requires_docker":false},"decision_summary":[{"label":"Requirements","value":"Min 32 GB RAM; Requires NVIDIA GPUs for optimized performance. Non-GPU usage is not supported or recommended.; Installation from source can be resource-intensive and may require limiting parallel compilation jobs to avoid running out of memory."},{"label":"Adopt for","value":"Megatron-LM from NVIDIA is a research-focused tool for developing and training large-scale language models with transformer architectures, emphasizing efficient parallelism across multiple GPUs."}]}}