---
title: "cceval vs HLCE"
type: "comparison"
canonical_url: "https://www.graphcanon.com/compare/amazon-science-cceval-vs-humanity-s-last-code-exam-hlce"
tools: ["amazon-science-cceval", "humanity-s-last-code-exam-hlce"]
---

# cceval vs HLCE

*GraphCanon updated Aug 8, 2026*

## Verdict

Pick cceval if cceval is designed for assessing cross-file code completion capabilities in multilingual environments across different setups and retrieval methods; pick HLCE if hLCE offers evaluation scripts to assess code generation using LLMs, specifically for research purposes.

[cceval](https://crosscodeeval.github.io/) reports 182 GitHub stars, 28 forks, and 5 open issues, last pushed Aug 15, 2025. [HLCE](https://humanity-s-last-code-exam.github.io/website/) has 96 stars, 8 forks, and 1 open issues, last pushed Aug 21, 2025. Figures are from public GitHub metadata via [cceval's repository](https://github.com/amazon-science/cceval) and [HLCE's repository](https://github.com/Humanity-s-Last-Code-Exam/HLCE).

| | [cceval](/tools/amazon-science-cceval.md) | [HLCE](/tools/humanity-s-last-code-exam-hlce.md) |
| --- | --- | --- |
| Tagline | CrossCodeEval Benchmark for Cross-File Code Completion | Source Evaluation scripts for Humanity's Last Code Exam |
| Stars | 182 | 96 |
| Forks | 28 | 8 |
| Open issues | 5 | 1 |
| Language | Python | Python |
| Adopt for | cceval is designed for assessing cross-file code completion capabilities in multilingual environments across different setups and retrieval methods. | HLCE offers evaluation scripts to assess code generation using LLMs, specifically for research purposes. |
| Persona | - | - |
| Runtime | - | - |
| License | Apache-2.0 | - |
| Categories | Evaluation & Observability | Evaluation & Observability, LLM Frameworks |

## Trust and health

_Sourced signals - not a safety guarantee. No winner column._

| | [cceval](/tools/amazon-science-cceval.md) | [HLCE](/tools/humanity-s-last-code-exam-hlce.md) |
| --- | --- | --- |
| Days since push | 354d | 352d |
| Open issues (now) | 5 | 1 |
| Full report | [trust report](/tools/amazon-science-cceval/trust.md) | [trust report](/tools/humanity-s-last-code-exam-hlce/trust.md) |

## Decision facts: cceval

- **Adopt for:** cceval is designed for assessing cross-file code completion capabilities in multilingual environments across different setups and retrieval methods.

## Decision facts: HLCE

- **Adopt for:** HLCE offers evaluation scripts to assess code generation using LLMs, specifically for research purposes.

## Choose when

### Choose cceval if…

- Tags unique to cceval: code-completion, cross-file, evaluation-tool, multilingual.
- You need to evaluate the performance of cross-file code completion systems that can handle multiple programming languages simultaneously
- More GitHub stars (182 vs 96) - visibility, not fit.

### Choose HLCE if…

- Tags unique to HLCE: codegen, codellm, llm-evaluation.
- Also covers LLM Frameworks.
- When you are researching the capabilities of language models in generating code and need benchmarking tools that focus on this aspect exclusively.

## When NOT to use cceval

- Your evaluation needs are focused solely on single-file or intralingual code completion benchmarks
- You require real-time data access or live updates; cceval provides pre-packaged datasets that need to be manually obtained and uncompressed

## When NOT to use HLCE

- If you require tools that cater to general-purpose evaluation beyond the scope of LLM code generation in a research context.
- When proprietary or non-research licenses are necessary, since HLCE does not detail its licensing beyond being for research purposes only.

## Common questions

### What is the difference between cceval and HLCE?

cceval: CrossCodeEval Benchmark for Cross-File Code Completion. HLCE: Source Evaluation scripts for Humanity's Last Code Exam. See the comparison table for live GitHub stats and shared categories.

### When should I choose cceval over HLCE?

Choose cceval over HLCE when Tags unique to cceval: code-completion, cross-file, evaluation-tool, multilingual; You need to evaluate the performance of cross-file code completion systems that can handle multiple programming languages simultaneously; More GitHub stars (182 vs 96) - visibility, not fit.

### When should I choose HLCE over cceval?

Choose HLCE over cceval when Tags unique to HLCE: codegen, codellm, llm-evaluation; Also covers LLM Frameworks; When you are researching the capabilities of language models in generating code and need benchmarking tools that focus on this aspect exclusively.

### When should I avoid cceval?

Your evaluation needs are focused solely on single-file or intralingual code completion benchmarks You require real-time data access or live updates; cceval provides pre-packaged datasets that need to be manually obtained and uncompressed

### When should I avoid HLCE?

If you require tools that cater to general-purpose evaluation beyond the scope of LLM code generation in a research context. When proprietary or non-research licenses are necessary, since HLCE does not detail its licensing beyond being for research purposes only.

### Is cceval or HLCE more popular on GitHub?

cceval has more GitHub stars (182 vs 96). Stars measure visibility, not whether either tool fits your constraints.

### Are cceval and HLCE open source?

Yes - both are open-source projects on GitHub.

### Where can I find alternatives to cceval or HLCE?

GraphCanon lists graph-backed alternatives at [cceval alternatives](/tools/amazon-science-cceval/alternatives) and [HLCE alternatives](/tools/humanity-s-last-code-exam-hlce/alternatives) ([cceval markdown twin](/tools/amazon-science-cceval/alternatives.md), [HLCE markdown twin](/tools/humanity-s-last-code-exam-hlce/alternatives.md)), ranked by typed relationship edges rather than popularity votes.

### Is there a machine-readable version of this comparison?

Yes. The markdown twin at [this comparison](/compare/amazon-science-cceval-vs-humanity-s-last-code-exam-hlce.md) mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.

### Which is better maintained, cceval or HLCE?

cceval: Slowing. HLCE: Slowing. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.

### Where are the full trust reports for cceval and HLCE?

GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: [cceval trust report](/tools/amazon-science-cceval/trust); [HLCE trust report](/tools/humanity-s-last-code-exam-hlce/trust).

---

**Machine-readable endpoints**

- JSON: [`/api/graphcanon/graph?tool=amazon-science-cceval`](/api/graphcanon/graph?tool=amazon-science-cceval)
- LLM index: [/llms.txt](/llms.txt)
- Full corpus: [/llms-full.txt](/llms-full.txt)

_GraphCanon - The knowledge graph for AI development. https://www.graphcanon.com/_
