---
title: "WeaveBench vs LLM-Agent-Paper-List"
type: "comparison"
canonical_url: "https://www.graphcanon.com/compare/weavebench-weavebench-vs-woooodyy-llm-agent-paper-list"
tools: ["weavebench-weavebench", "woooodyy-llm-agent-paper-list"]
---

# WeaveBench vs LLM-Agent-Paper-List

*GraphCanon updated Aug 17, 2026*

## Verdict

Pick WeaveBench if weaveBench is designed for evaluating computer-use agents that integrate both GUI and CLI interactions in real-world scenarios across various work domains; pick LLM-Agent-Paper-List if lists essential papers on LLM-based agents with integrated tools like AgentGym for RL training.

[WeaveBench](https://weavebench.github.io) reports 157 GitHub stars, 1 forks, and 4 open issues, last pushed Jul 22, 2026. [LLM-Agent-Paper-List](https://arxiv.org/abs/2309.07864) has 8.2k stars, 495 forks, and 31 open issues, last pushed Sep 12, 2025. Figures are from public GitHub metadata via [WeaveBench's repository](https://github.com/weavebench/WeaveBench) and [LLM-Agent-Paper-List's repository](https://github.com/WooooDyy/LLM-Agent-Paper-List).

| | [WeaveBench](/tools/weavebench-weavebench.md) | [LLM-Agent-Paper-List](/tools/woooodyy-llm-agent-paper-list.md) |
| --- | --- | --- |
| Tagline | A Long-Horizon Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces | Must-read papers for LLM-based agents. |
| Stars | 157 | 8,172 |
| Forks | 1 | 495 |
| Open issues | 4 | 31 |
| Language | Python | - |
| Adopt for | WeaveBench is designed for evaluating computer-use agents that integrate both GUI and CLI interactions in real-world scenarios across various work domains. | Lists essential papers on LLM-based agents with integrated tools like AgentGym for RL training. |
| Persona | - | - |
| Runtime | - | - |
| License | MIT | - |
| Categories | AI Agents, Evaluation & Observability | AI Agents, Evaluation & Observability |

## Trust and health

_Sourced signals - not a safety guarantee. No winner column._

| | [WeaveBench](/tools/weavebench-weavebench.md) | [LLM-Agent-Paper-List](/tools/woooodyy-llm-agent-paper-list.md) |
| --- | --- | --- |
| Maintenance | Very active (96%) | Slowing (36%) |
| Days since push | 6d | 339d |
| Open issues (now) | 4 | 31 |
| Stars delta | Unknown | +4 (30d) |
| Open issues delta | Unknown | +2 (30d) |
| Owner type | Organization | User |
| Full report | [trust report](/tools/weavebench-weavebench/trust.md) | [trust report](/tools/woooodyy-llm-agent-paper-list/trust.md) |

## Decision facts: WeaveBench

- **Adopt for:** WeaveBench is designed for evaluating computer-use agents that integrate both GUI and CLI interactions in real-world scenarios across various work domains.

## Decision facts: LLM-Agent-Paper-List

- **Adopt for:** Lists essential papers on LLM-based agents with integrated tools like AgentGym for RL training.

## Choose when

### Choose WeaveBench if…

- Tags unique to WeaveBench: agent-as-judge, benchmark, computer-use-agent, gui-agent.
- Use WeaveBench if you need to assess agents capable of handling tasks that require intermingling graphical user interface operations with command-line or code-based actions.
- More recently updated (last pushed Jul 22, 2026).

### Choose LLM-Agent-Paper-List if…

- Tags unique to LLM-Agent-Paper-List: agent, large language models, llm, nlp.
- Looking to survey key advancements in LLM-based agent research, specifically through papers endorsed by authors.
- More GitHub stars (8.2k vs 157) - visibility, not fit.

## When NOT to use WeaveBench

- Avoid WeaveBench if your testing needs do not involve scenarios that require the integration of both GUI and CLI operations.
- Do not use it when you are specifically interested only in benchmarking agents designed for single-channel tasks, either strictly CLI-based or purely graphical interface-driven.

## When NOT to use LLM-Agent-Paper-List

- Seeking real-time interactive debugging tools; focuses more on paper reviews and general frameworks than coding sandbox features.
- Require support documentation in languages other than English or project-specific code details, as licensing and detailed documentation are currently unverified.

## Common questions

### What is the difference between WeaveBench and LLM-Agent-Paper-List?

WeaveBench: A Long-Horizon Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces. LLM-Agent-Paper-List: Must-read papers for LLM-based agents.. See the comparison table for live GitHub stats and shared categories.

### When should I choose WeaveBench over LLM-Agent-Paper-List?

Choose WeaveBench over LLM-Agent-Paper-List when Tags unique to WeaveBench: agent-as-judge, benchmark, computer-use-agent, gui-agent; Use WeaveBench if you need to assess agents capable of handling tasks that require intermingling graphical user interface operations with command-line or code-based actions; More recently updated (last pushed Jul 22, 2026).

### When should I choose LLM-Agent-Paper-List over WeaveBench?

Choose LLM-Agent-Paper-List over WeaveBench when Tags unique to LLM-Agent-Paper-List: agent, large language models, llm, nlp; Looking to survey key advancements in LLM-based agent research, specifically through papers endorsed by authors; More GitHub stars (8.2k vs 157) - visibility, not fit.

### When should I avoid WeaveBench?

Avoid WeaveBench if your testing needs do not involve scenarios that require the integration of both GUI and CLI operations. Do not use it when you are specifically interested only in benchmarking agents designed for single-channel tasks, either strictly CLI-based or purely graphical interface-driven.

### When should I avoid LLM-Agent-Paper-List?

Seeking real-time interactive debugging tools; focuses more on paper reviews and general frameworks than coding sandbox features. Require support documentation in languages other than English or project-specific code details, as licensing and detailed documentation are currently unverified.

### Is WeaveBench or LLM-Agent-Paper-List more popular on GitHub?

LLM-Agent-Paper-List has more GitHub stars (8,172 vs 157). Stars measure visibility, not whether either tool fits your constraints.

### Are WeaveBench and LLM-Agent-Paper-List open source?

Yes - both are open-source projects on GitHub.

### Where can I find alternatives to WeaveBench or LLM-Agent-Paper-List?

GraphCanon lists graph-backed alternatives at [WeaveBench alternatives](/tools/weavebench-weavebench/alternatives) and [LLM-Agent-Paper-List alternatives](/tools/woooodyy-llm-agent-paper-list/alternatives) ([WeaveBench markdown twin](/tools/weavebench-weavebench/alternatives.md), [LLM-Agent-Paper-List markdown twin](/tools/woooodyy-llm-agent-paper-list/alternatives.md)), ranked by typed relationship edges rather than popularity votes.

### Is there a machine-readable version of this comparison?

Yes. The markdown twin at [this comparison](/compare/weavebench-weavebench-vs-woooodyy-llm-agent-paper-list.md) mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.

### Which is better maintained, WeaveBench or LLM-Agent-Paper-List?

WeaveBench: Very active. LLM-Agent-Paper-List: Slowing. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.

### Where are the full trust reports for WeaveBench and LLM-Agent-Paper-List?

GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: [WeaveBench trust report](/tools/weavebench-weavebench/trust); [LLM-Agent-Paper-List trust report](/tools/woooodyy-llm-agent-paper-list/trust).

---

**Machine-readable endpoints**

- JSON: [`/api/graphcanon/graph?tool=weavebench-weavebench`](/api/graphcanon/graph?tool=weavebench-weavebench)
- LLM index: [/llms.txt](/llms.txt)
- Full corpus: [/llms-full.txt](/llms-full.txt)

_GraphCanon - The knowledge graph for AI development. https://www.graphcanon.com/_
