{"data":{"slug":"opencoder-llm-opencoder-llm","name":"OpenCoder-llm","tagline":"The Open Cookbook for Top-Tier Code Large Language Models","github_url":"https://github.com/OpenCoder-llm/OpenCoder-llm","owner":"OpenCoder-llm","repo":"OpenCoder-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/186387526?v=4","primary_language":"Python","stars":2103,"forks":125,"topics":[],"archived":false,"github_pushed_at":"2024-12-08T16:46:00+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/opencoder-llm-opencoder-llm","markdown_url":"https://www.graphcanon.com/tools/opencoder-llm-opencoder-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/opencoder-llm-opencoder-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=opencoder-llm-opencoder-llm","description":"The Open Cookbook for Top-Tier Code Large Language Model","homepage_url":"https://opencoder-llm.github.io/","license":"MIT","open_issues":11,"watchers":32,"ai_summary":"OpenCoder is an open-source project focusing on large language models designed to generate high-quality code. It includes various resources such as datasets, evaluation frameworks, and data pipelines.","readme_excerpt":"<div align=\"center\">\n  <img src=\"https://github.com/OpenCoder-llm/opencoder-llm.github.io/blob/main/static/images/opencoder_icon.jpg?raw=true\" width=\"30%\" alt=\"OpenCoder-Icon\" />\n</div>\n\n\n<p align=\"center\">\n    <h1 align=\"center\">\n\n        OpenCoder\n    </h1>\n     <p align=\"center\">⚡ The Open Cookbook for Top-Tier Code Large Language Models ⚡</p>\n</p>\n\n<p align=\"center\">\n        🏠<a href=\"https://opencoder-llm.github.io/\">Home Page</a>&nbsp&nbsp | &nbsp&nbsp🤗<a href=\"https://huggingface.co/collections/infly/opencoder-672cec44bbb86c39910fb55e\">Model</a>&nbsp&nbsp | &nbsp&nbsp📊<a href=\"https://huggingface.co/collections/OpenCoder-LLM/opencoder-datasets-672e6db6a0fed24bd69ef1c2\">Dataset</a>&nbsp&nbsp | &nbsp&nbsp📄<a href=\"https://arxiv.org/abs/2411.04905\">Paper</a>&nbsp ｜ 🚀<a href=\"https://huggingface.co/spaces/OpenCoder-LLM/OpenCoder-8B-Instruct\">Demo</a>&nbsp&nbsp\n</p>\n\n\n\n## News\n- 🔥🔥🔥 ```2024/12/08``` We have released our pretraining data cleaning pipeline: [opc_data_filtering](https://github.com/OpenCoder-llm/opc_data_filtering). Try to use this pipeline to create your own high-quality code pretraining corpus!\n- 🔥 ```2024/11/19``` We have released intermedidate checkpoints during our pretraining stage: 🤗 [OpenCoder-1.5B-Base-Checkpoints](https://huggingface.co/OpenCoder-LLM/OpenCoder-1.5B-Base-Checkpoints) and 🤗 [OpenCoder-8B-Base-Checkpoints](https://huggingface.co/OpenCoder-LLM/OpenCoder-8B-Base-Checkpoints).\n- 🔥 ```2024/11/15``` We have released meta data of **RefineCode** 📊 [RefineCode-code-corpus-meta](https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-raw-code-meta). You can collect your own **RefineCode** referring to this dataset!\n- 🔥 ```2024/11/12``` We have released our efficient CodeLLM evaluation framework: [OpenCodeEval](https://github.com/OpenCoder-llm/OpenCoder-llm/tree/main/OpenCodeEval).\n- 🔥 ```2024/11/12``` We have released high-quality annealing data 📊 [opc-annealing-corpus](https://huggingface.co/datasets/OpenCoder-LLM/opc-annealing-corpus), which includes algorithmic-corpus along with corresponding synthetic data.\n- 🔥 ```2024/11/11``` We have released 55B of recalled pages from [Fineweb](https://huggingface.co/datasets/HuggingFaceFW/fineweb), including 📊 [fineweb-code-corpus](https://huggingface.co/datasets/OpenCoder-LLM/fineweb-code-corpus) and 📊 [fineweb-math-corpus](https://huggingface.co/datasets/OpenCoder-LLM/fineweb-math-corpus).\n- 🔥 ```2024/11/09``` We have released 4.5M Post-training data: 📊 [Dataset](https://huggingface.co/collections/OpenCoder-LLM/opencoder-datasets-672e6db6a0fed24bd69ef1c2).\n- 🔥 ```2024/11/08``` We have released our models! Please download them from 🤗 [Model](https://huggingface.co/collections/infly/opencoder-672cec44bbb86c39910fb55e).\n- 🔥 ```2024/11/07``` We have released our paper on Arxiv: 📄 [OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models](https://arxiv.org/abs/2411.04905).\n\n\n## Releases\n- [x] Data cleaning pipeline\n- [x] **RefineCode**: Code-related web data\n- [x] **RefineCode**: Metadata of raw code data \n- [x] Intermedidate Checkpoints\n- [x] CodeLLM evaluation framework: OpenCodeEval\n- [x] High-quality annealing data\n- [x] Post-training data\n- [x] Final model weights\n- [x] Paper\n\nWe are working hard to release all those resources! 💪 \n\n\n## Introduction\n\n**OpenCoder** is an open and reproducible code LLM family which includes 1.5B and 8B base and chat models, supporting both English and Chinese languages. Starting from scratch, OpenCoder is pretrained on 2.5 trillion tokens composed of 90% raw code and 10% code-related web data, and supervised finetuned on over 4.5M high-quality SFT examples, finally reaching the performance of top-tier code LLMs. We provide not only model weights and inference code, but also the reproducible training data, the complete data processing pipeline, rigorous experimental ablation results, and detailed training protocols. Empowering researchers to build and innovate, OpenCoder is your open fo","github_created_at":"2024-10-26T11:07:49+00:00","created_at":"2026-07-11T23:43:45.127652+00:00","updated_at":"2026-08-05T12:01:50.760648+00:00","categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"},{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"code-generation","name":"code generation"},{"slug":"data-filtering","name":"data filtering"},{"slug":"dataset","name":"dataset"},{"slug":"evaluation-framework","name":"evaluation-framework"},{"slug":"large-language-models","name":"large language models"}],"trust":{"provenance":{"is_fork":false,"github_id":878879204,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-05T12:01:49.769Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":604,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T23:43:47.162Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-05T12:01:50.265Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-08-05T12:01:50.265Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-05T12:01:50.265Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When you need access to both English and Chinese language support in your code generation tasks.","If you require a reproducible framework that provides intermediate checkpoints such as the OpenCoder-1.5B-Base-Checkpoints and OpenCoder-8B-Base-Checkpoints for model development.","For organizations looking to leverage high-quality annealing data along with their synthetic counterparts, available through this tool.","When your project involves pretraining large language models and you need a robust data cleaning pipeline like opc_data_filtering."],"when_not_to_use":["If your primary focus is on natural language processing tasks that do not involve code generation or require languages other than English or Chinese.","For scenarios where the availability of intermediate checkpoints during pretraining stages does not add value to your development process.","If you are working with datasets that already provide synthetic annealing data, and additional resources for this type of data are unnecessary.","When a tool without an open-source data cleaning pipeline is sufficient for your code generation tasks."],"source":"enrich:decision_facts","observed_at":"2026-07-16T18:58:34.868Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"OpenCoder-llm offers comprehensive resources for generating high-quality code through its large language models, including datasets, evaluation frameworks, and data pipelines."}]}}