Home/Compare/agentic-vbench vs ClawBench

Comparison

agentic-vbench vs ClawBench

Verdict

Pick agentic-vbench if agenticVBench evaluates AI agents' real-world post-production capabilities with specific task prompts for activities like audio restoration; pick ClawBench if clawBench offers an open-source benchmark framework for evaluating browser-based AI agents on real-world online tasks.

Markdown twin · agentic-vbench alternatives · ClawBench alternatives

GraphCanon updated Sep 20, 2026

20views this month

agentic-vbench logo

agentic-vbench

PhiloLabs/agentic-vbench

96pushed Sep 2, 2026
vs
ClawBench logo

ClawBench

TIGER-AI-Lab/ClawBench

793pushed Sep 15, 2026

Trust & integrity

Signalagentic-vbenchClawBench
Maintenance
Very active (6d since push)
As of Sep 9, 2026 · github_public_v1
Very active (4d since push)
As of Sep 20, 2026 · github_public_v1
Provenance
Not a fork · Organization account
As of Sep 9, 2026 · github_public_v1
Not a fork · Organization account
As of Sep 20, 2026 · github_public_v1
OSV dependency advisories
No lockfile (source not queried)
As of Jul 15, 2026 · osv@v1
No lockfile (source not queried)
As of Jul 11, 2026 · osv@v1
deps.dev advisories
Not queried
deps.dev@v1
Not queried
deps.dev@v1
OpenSSF Scorecard
Not queried
openssf-scorecard@v1
Not queried
openssf-scorecard@v1

Tagline

agentic-vbench
A benchmark for evaluating AI agents in performing real-world post-production tasks like audio and video editing.
ClawBench
Open-source benchmark for browser AI agents on daily tasks

Stars

agentic-vbench
96
ClawBench
793

Forks

agentic-vbench
27
ClawBench
58

Open issues

agentic-vbench
37
ClawBench
47

Language

agentic-vbench
Python
ClawBench
Python

Adopt for

agentic-vbench
AgenticVBench evaluates AI agents' real-world post-production capabilities with specific task prompts for activities like audio restoration.
ClawBench
ClawBench offers an open-source benchmark framework for evaluating browser-based AI agents on real-world online tasks.

Persona

agentic-vbench
-
ClawBench
-

Runtime

agentic-vbench
-
ClawBench
-

License

agentic-vbench
Apache-2.0
ClawBench
ClawBench operates under the Apache-2.0 license, offering a permissive free software license that encourages software reuse and interoperability.

Last pushed

agentic-vbench
Sep 2, 2026
ClawBench
Sep 15, 2026

Categories

agentic-vbench
AI Agents, Evaluation & Observability
ClawBench
AI Agents, Evaluation & Observability

Trust and health

Days since push

agentic-vbench
6d
ClawBench
4d

Open issues (now)

agentic-vbench
37
ClawBench
47

Stars delta

agentic-vbench
+14 (30d)
ClawBench
+261 (30d)

Open issues delta

agentic-vbench
-20 (30d)
ClawBench
0 (30d)

Full report

agentic-vbench
Trust report
ClawBench
Trust report

Shared compatibility

  • Python · agentic-vbench: Python runtime · ClawBench: Python runtime

Choose agentic-vbench if…

  • Requirements: Requires Docker; Install via scripts provided in the repository.; Python virtual environment setup for reproducibility..
  • Tags unique to agentic-vbench: ai-agents, benchmark, harbor, video-editing.
  • When you need to benchmark the performance of AI agents in handling specialized tasks such as audio and video editing that require precise restorative actions.

When NOT to use agentic-vbench

  • When the focus is on generic performance evaluations rather than on real-world, task-specific benchmarks that assess handling complex post-production scenarios.
  • If your budget or timeline cannot accommodate a per-task wall clock time of ~10 minutes and cost ranging from $0.10 to $2 based on agent token usage.

Choose ClawBench if…

  • Requirements: Requires setup for browser automation tasks within the Chrome environment..
  • Tags unique to ClawBench: agent-evaluation, agentic-ai, browser-agent, real-world-benchmark.
  • You are developing a browser AI agent and wish to measure its performance against everyday online activities, as ClawBench specifically simulates these scenarios.

When NOT to use ClawBench

  • If your focus is solely on backend server-based AI evaluations without a browser interface involvement, ClawBench will not be the appropriate choice.
  • For those developing standalone applications or mobile agents, ClawBench’s browser-centric tasks will not reflect their operational capabilities accurately.

Explore

Sources

Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.

GitHub stars on cards: agentic-vbench 96 · ClawBench 793 (synced Sep 20, 2026).

Common questions

What is the difference between agentic-vbench and ClawBench?
agentic-vbench: A benchmark for evaluating AI agents in performing real-world post-production tasks like audio and video editing.. ClawBench: Open-source benchmark for browser AI agents on daily tasks. See the comparison table for live GitHub stats and shared categories.
When should I choose agentic-vbench over ClawBench?
Choose agentic-vbench over ClawBench when Requirements: Requires Docker; Install via scripts provided in the repository.; Python virtual environment setup for reproducibility.; Tags unique to agentic-vbench: ai-agents, benchmark, harbor, video-editing; When you need to benchmark the performance of AI agents in handling specialized tasks such as audio and video editing that require precise restorative actions.
When should I choose ClawBench over agentic-vbench?
Choose ClawBench over agentic-vbench when Requirements: Requires setup for browser automation tasks within the Chrome environment.; Tags unique to ClawBench: agent-evaluation, agentic-ai, browser-agent, real-world-benchmark; You are developing a browser AI agent and wish to measure its performance against everyday online activities, as ClawBench specifically simulates these scenarios.
When should I avoid agentic-vbench?
When the focus is on generic performance evaluations rather than on real-world, task-specific benchmarks that assess handling complex post-production scenarios. If your budget or timeline cannot accommodate a per-task wall clock time of ~10 minutes and cost ranging from $0.10 to $2 based on agent token usage.
When should I avoid ClawBench?
If your focus is solely on backend server-based AI evaluations without a browser interface involvement, ClawBench will not be the appropriate choice. For those developing standalone applications or mobile agents, ClawBench’s browser-centric tasks will not reflect their operational capabilities accurately.
Is agentic-vbench or ClawBench more popular on GitHub?
ClawBench has more GitHub stars (793 vs 96). Stars measure visibility, not whether either tool fits your constraints.
Are agentic-vbench and ClawBench open source?
Yes - both are open-source projects on GitHub (agentic-vbench: Apache-2.0, ClawBench: Apache-2.0).
Where can I find alternatives to agentic-vbench or ClawBench?
GraphCanon lists graph-backed alternatives at agentic-vbench alternatives and ClawBench alternatives (agentic-vbench markdown twin, ClawBench markdown twin), ranked by typed relationship edges rather than popularity votes.
Is there a machine-readable version of this comparison?
Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
Which is better maintained, agentic-vbench or ClawBench?
agentic-vbench: Very active. ClawBench: Very active. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
Where are the full trust reports for agentic-vbench and ClawBench?
GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: agentic-vbench trust report; ClawBench trust report.

Was this helpful?

Anonymous feedback helps us improve pages and translations.