{"data":{"slug":"amberljc-llmsys-paperlist","name":"LLMSys-PaperList","tagline":"Curated list of academic papers related to Large Language Model systems","github_url":"https://github.com/AmberLJC/LLMSys-PaperList","owner":"AmberLJC","repo":"LLMSys-PaperList","owner_avatar_url":"https://avatars.githubusercontent.com/u/42296458?v=4","primary_language":"Python","stars":2220,"forks":120,"topics":[],"archived":false,"github_pushed_at":"2026-07-25T02:03:12+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/amberljc-llmsys-paperlist","markdown_url":"https://www.graphcanon.com/tools/amberljc-llmsys-paperlist.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/amberljc-llmsys-paperlist","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=amberljc-llmsys-paperlist","description":"Large Language Model (LLM) Systems Paper List","homepage_url":null,"license":null,"open_issues":1,"watchers":49,"ai_summary":"A repository containing a curated collection of academic papers, articles, tutorials, and projects related to Large Language Model (LLM) systems in categories such as training, serving, multi-modal systems, LLM frameworks, ML conferences, survey papers, benchmarks, and more.","readme_excerpt":"# Awesome LLM Systems Papers\n\nA curated list of Large Language Model systems related academic papers, articles, tutorials, slides and projects. Star this repository, and then you can keep abreast of the latest developments of this booming research field.\n\n## Trends at a Glance (2024 → 2026)\n\nThree eras, one lens: **what unit of work the system optimizes** — a request (2024), a session or\nreasoning trace (2025), a whole agent trajectory (2026).\n\n\n\nServing is still the largest area in absolute terms, but its share of the list fell from **49% to\n33%** — the growth went to kernel/model co-design (3 → 29 papers), agentic systems (4 → 22),\nAI-for-systems (5 → 16) and edge (2 → 14).\n\n\n\n\n\nThe fastest-rising techniques by share of the year's papers: **agentic / multi-agent** (1.0% → 13.1%),\n**compiler / kernel / megakernel** (0.0% → 10.2%), **speculative decoding** (1.0% → 5.7%),\n**sparse attention** (1.0% → 4.5%) and **energy / power** (1.0% → 3.3%).\n\n**→ Full analysis, per-technique numbers and reproduction scripts: [`trends/`](trends/)**\n\n## Table of Contents\n\n- [Trends at a Glance](#trends-at-a-glance-2024--2026)\n- [LLM Systems](#llm-systems)\n  - [Training](#training)\n    - [Pre-training](#pre-training)\n    - [Post Training](#systems-for-post-training--rlhf)\n    - [Fault Tolerance / Straggler Mitigation](#fault-tolerance--straggler-mitigation)\n  - [Serving](#serving)\n    - [LLM serving](#llm-serving)\n    - [Agent Systems](#agent-systems)\n    - [Serving at the edge](#serving-at-the-edge)\n    - [System Efficiency Optimization - Model Co-design](#system-efficiency-optimization---model-co-design)\n  - [Multi-Modal Training Systems](#multi-modal-training-systems)\n  - [Multi-Modal Serving Systems](#multi-modal-serving-systems)\n- [LLM for Systems](#llm-for-systems)\n- [Industrial LLM Technical Report](#industrial-llm-technical-report)\n- [ML Conferences](#ml-conferences)\n  - [NeurIPS 2025](#neurips-2025)\n- [LLM Frameworks](#llm-frameworks)\n  - [Training](#training-1)\n  - [Post-Training](#post-training)\n  - [Serving](#serving-1)\n- [ML Systems](#ml-systems)\n- [Survey Paper](#survey-paper)\n- [LLM Benchmark / Leaderboard / Traces](#llm-benchmark--leaderboard--traces)\n- [Related ML Readings](#related-ml-readings)\n- [MLSys Courses](#mlsys-courses)\n- [Other Reading](#other-reading)\n\n\n## LLM Systems\n### Training\n#### Pre-training\n\n<details>\n<summary><b>Before 2024</b></summary>\n\n- [Megatron-LM](https://arxiv.org/pdf/1909.08053.pdf): Training Multi-Billion Parameter Language Models Using Model Parallelism\n- [Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM](https://arxiv.org/pdf/2104.04473.pdf)\n- [Reducing Activation Recomputation in Large Transformer Models](https://arxiv.org/pdf/2205.05198.pdf)\n- [Optimized Network Architectures for Large Language Model Training with Billions of Parameters](https://arxiv.org/pdf/2307.12169.pdf) | MIT\n- [Carbon Emissions and Large Neural Network Training](https://arxiv.org/pdf/2104.10350.pdf?fbclid=IwAR2o0_3HCtTnMxKbXka0OPrHzl8sCzQSSOYp0AOav76-zVWl_pYek2jX8Pk) | Google, UCB\n\n</details>\n\n<details>\n<summary><b>2024</b></summary>\n\n- [Perseus](https://arxiv.org/abs/2312.06902v1): Removing Energy Bloat from Large Model Training | SOSP' 24\n- [MegaScale](https://arxiv.org/abs/2402.15627): Scaling Large Language Model Training to More Than 10,000 GPUs | ByteDance\n- [DISTMM](https://www.usenix.org/conference/nsdi24/presentation/huang): Accelerating distributed multimodal model training | NSDI' 24\n- [Pipeline Parallelism with Controllable Memory](https://arxiv.org/abs/2405.15362) | Sea AI Lab\n- [Boosting Large-scale Parallel Training Efficiency with C4](https://arxiv.org/abs/2406.04594): A Communication-Driven Approach\n- [Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model Training](https://openreview.net/pdf?id=uLpyWQPyF9) | ICML' 24\n- [Alibaba HPN:](https://ennanzhai.github.io/pub/sigcomm24-hpn.pdf) A Data Center Network for Large Language ModelTraining\n- [The Llama 3 Herd o","github_created_at":"2023-06-06T02:34:49+00:00","created_at":"2026-07-11T10:32:44.487284+00:00","updated_at":"2026-08-06T18:01:02.626983+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"academic-sources","name":"academic-sources"},{"slug":"framework-overview","name":"framework-overview"},{"slug":"inference-techniques","name":"inference-techniques"},{"slug":"research-papers","name":"research papers"},{"slug":"training-methodologies","name":"training-methodologies"}],"trust":{"provenance":{"is_fork":false,"github_id":649955694,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-06T18:01:01.750Z","maintenance":{"label":"Active","score":82,"methodology":"github_public_v1","releases_90d":0,"days_since_push":12,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T10:32:45.556Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-06T18:01:02.254Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-08-06T18:01:02.254Z"}},"decision_facts":{"hosting":{"model":"unknown","summary":"(repository does not specify hosting environment)"},"pricing":null,"requirements":null,"constraints":{"hosting_model":"unknown"},"when_to_use":["- When you need a curated list focusing on technical advancements in pre-training, post-training, serving, and multi-modal LLM systems.","- If your interest lies specifically in recent developments and cutting-edge research by leading industry players like Google, ByteDance, and Sea AI Lab, from conferences like SOSP' 24 and NSDI' 24.","- For systematic access to resources that span from network optimization for training, such as Alibaba's HPN, to energy reduction techniques in LLM training."],"when_not_to_use":["- If you are looking for a general repository of machine learning papers rather than specific developments related to Large Language Models.","- When your primary need is documentation or code examples rather than academic papers and project insights.","- For applications where real-time updates and active community support are imperative, as LLMSys-PaperList primarily serves as a static list without user interaction features like commenting or liveＱ"],"source":"enrich:decision_facts","observed_at":"2026-07-11T11:17:07.356Z"},"constraint_facets":{"hosting_model":"unknown"},"decision_summary":[{"label":"Hosting","value":"unknown - (repository does not specify hosting environment)"},{"label":"Adopt for","value":"LLMSys-PaperList offers a comprehensive list of papers and resources tailored specifically to Large Language Model (LLM) systems."},{"label":"License detail","value":"(unknown)"}]}}