Comparison
WeaveBench vs LLM-Agent-Paper-List
Verdict
Pick WeaveBench if weaveBench is designed for evaluating computer-use agents that integrate both GUI and CLI interactions in real-world scenarios across various work domains; pick LLM-Agent-Paper-List if lists essential papers on LLM-based agents with integrated tools like AgentGym for RL training.
Markdown twin · WeaveBench alternatives · LLM-Agent-Paper-List alternatives
GraphCanon updated 1w
Trust & integrity
| Signal | WeaveBench | LLM-Agent-Paper-List |
|---|---|---|
| Maintenance | Very active (6d since push) As of 3w · github_public_v1 | Slowing (339d since push) As of 1w · github_public_v1 |
| Provenance | Not a fork · Organization account As of 3w · github_public_v1 | Not a fork · Personal account As of 1w · github_public_v1 |
| OSV dependency advisories | Published findings As of 1mo · osv@v1 | No lockfile (source not queried) As of 1mo · osv@v1 |
| deps.dev advisories | Not queried deps.dev@v1 | Not queried deps.dev@v1 |
| OpenSSF Scorecard | Not queried openssf-scorecard@v1 | Not queried openssf-scorecard@v1 |
Tagline
- WeaveBench
- A Long-Horizon Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
- LLM-Agent-Paper-List
- Must-read papers for LLM-based agents.
Stars
- WeaveBench
- 157
- LLM-Agent-Paper-List
- 8.2k
Forks
- WeaveBench
- 1
- LLM-Agent-Paper-List
- 495
Open issues
- WeaveBench
- 4
- LLM-Agent-Paper-List
- 31
Language
- WeaveBench
- Python
- LLM-Agent-Paper-List
- -
Adopt for
- WeaveBench
- WeaveBench is designed for evaluating computer-use agents that integrate both GUI and CLI interactions in real-world scenarios across various work domains.
- LLM-Agent-Paper-List
- Lists essential papers on LLM-based agents with integrated tools like AgentGym for RL training.
Persona
- WeaveBench
- -
- LLM-Agent-Paper-List
- -
Runtime
- WeaveBench
- -
- LLM-Agent-Paper-List
- -
License
- WeaveBench
- MIT
- LLM-Agent-Paper-List
- -
Last pushed
- WeaveBench
- Jul 22, 2026
- LLM-Agent-Paper-List
- Sep 12, 2025
Categories
- WeaveBench
- AI Agents, Evaluation & Observability
- LLM-Agent-Paper-List
- AI Agents, Evaluation & Observability
Trust and health
Maintenance
- WeaveBench
- Very active (96%)
- LLM-Agent-Paper-List
- Slowing (36%)
Days since push
- WeaveBench
- 6d
- LLM-Agent-Paper-List
- 339d
Open issues (now)
- WeaveBench
- 4
- LLM-Agent-Paper-List
- 31
Stars delta
- WeaveBench
- Unknown
- LLM-Agent-Paper-List
- +4 (30d)
Open issues delta
- WeaveBench
- Unknown
- LLM-Agent-Paper-List
- +2 (30d)
Owner type
- WeaveBench
- Organization
- LLM-Agent-Paper-List
- User
OSV dependency advisories
- WeaveBench
- Published findings
- LLM-Agent-Paper-List
- No lockfile (source not queried)
Full report
- WeaveBench
- Trust report
- LLM-Agent-Paper-List
- Trust report
Choose WeaveBench if…
- Tags unique to WeaveBench: agent-as-judge, benchmark, computer-use-agent, gui-agent.
- Use WeaveBench if you need to assess agents capable of handling tasks that require intermingling graphical user interface operations with command-line or code-based actions.
- More recently updated (last pushed Jul 22, 2026).
When NOT to use WeaveBench
- Avoid WeaveBench if your testing needs do not involve scenarios that require the integration of both GUI and CLI operations.
- Do not use it when you are specifically interested only in benchmarking agents designed for single-channel tasks, either strictly CLI-based or purely graphical interface-driven.
Choose LLM-Agent-Paper-List if…
- Tags unique to LLM-Agent-Paper-List: agent, large language models, llm, nlp.
- Looking to survey key advancements in LLM-based agent research, specifically through papers endorsed by authors.
- More GitHub stars (8.2k vs 157) - visibility, not fit.
When NOT to use LLM-Agent-Paper-List
- Seeking real-time interactive debugging tools; focuses more on paper reviews and general frameworks than coding sandbox features.
- Require support documentation in languages other than English or project-specific code details, as licensing and detailed documentation are currently unverified.
Explore
Sources
Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.
- GitHub stars (weavebench/WeaveBench) · observed Jul 29, 2026
- GitHub forks (weavebench/WeaveBench) · observed Jul 29, 2026
- Last push (weavebench/WeaveBench) · observed Jul 22, 2026
- License file (MIT) · observed Jul 29, 2026
- Decision facts (enrichment) · observed Jul 15, 2026
- Trust scan (lockfile / OSV) · observed Jul 11, 2026
- GitHub stars (WooooDyy/LLM-Agent-Paper-List) · observed Aug 17, 2026
- GitHub forks (WooooDyy/LLM-Agent-Paper-List) · observed Aug 17, 2026
- Last push (WooooDyy/LLM-Agent-Paper-List) · observed Sep 12, 2025
- License file (unknown) · observed Aug 17, 2026
- Decision facts (enrichment) · observed Jul 12, 2026
- Trust scan (lockfile / OSV) · observed Jul 11, 2026
GitHub stars on cards: WeaveBench 157 · LLM-Agent-Paper-List 8.2k (synced Jul 29, 2026).
Common questions
- What is the difference between WeaveBench and LLM-Agent-Paper-List?
- WeaveBench: A Long-Horizon Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces. LLM-Agent-Paper-List: Must-read papers for LLM-based agents.. See the comparison table for live GitHub stats and shared categories.
- When should I choose WeaveBench over LLM-Agent-Paper-List?
- Choose WeaveBench over LLM-Agent-Paper-List when Tags unique to WeaveBench: agent-as-judge, benchmark, computer-use-agent, gui-agent; Use WeaveBench if you need to assess agents capable of handling tasks that require intermingling graphical user interface operations with command-line or code-based actions; More recently updated (last pushed Jul 22, 2026).
- When should I choose LLM-Agent-Paper-List over WeaveBench?
- Choose LLM-Agent-Paper-List over WeaveBench when Tags unique to LLM-Agent-Paper-List: agent, large language models, llm, nlp; Looking to survey key advancements in LLM-based agent research, specifically through papers endorsed by authors; More GitHub stars (8.2k vs 157) - visibility, not fit.
- When should I avoid WeaveBench?
- Avoid WeaveBench if your testing needs do not involve scenarios that require the integration of both GUI and CLI operations. Do not use it when you are specifically interested only in benchmarking agents designed for single-channel tasks, either strictly CLI-based or purely graphical interface-driven.
- When should I avoid LLM-Agent-Paper-List?
- Seeking real-time interactive debugging tools; focuses more on paper reviews and general frameworks than coding sandbox features. Require support documentation in languages other than English or project-specific code details, as licensing and detailed documentation are currently unverified.
- Is WeaveBench or LLM-Agent-Paper-List more popular on GitHub?
- LLM-Agent-Paper-List has more GitHub stars (8,172 vs 157). Stars measure visibility, not whether either tool fits your constraints.
- Are WeaveBench and LLM-Agent-Paper-List open source?
- Yes - both are open-source projects on GitHub.
- Where can I find alternatives to WeaveBench or LLM-Agent-Paper-List?
- GraphCanon lists graph-backed alternatives at WeaveBench alternatives and LLM-Agent-Paper-List alternatives (WeaveBench markdown twin, LLM-Agent-Paper-List markdown twin), ranked by typed relationship edges rather than popularity votes.
- Is there a machine-readable version of this comparison?
- Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
- Which is better maintained, WeaveBench or LLM-Agent-Paper-List?
- WeaveBench: Very active. LLM-Agent-Paper-List: Slowing. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
- Where are the full trust reports for WeaveBench and LLM-Agent-Paper-List?
- GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: WeaveBench trust report; LLM-Agent-Paper-List trust report.