agentic-vbench
A benchmark for evaluating AI agents in performing real-world post-production tasks like audio and video editing.
GraphCanon updated Sep 9, 2026 · GitHub synced Sep 9, 2026
28views this month
Decision brief
AgenticVBench evaluates AI agents' real-world post-production capabilities with specific task prompts for activities like audio restoration.
Good fit when
- When you need to benchmark the performance of AI agents in handling specialized tasks such as audio and video editing that require precise restorative actions.
- If your development team requires metrics on processing time, cost, and accuracy when evaluating different AI agent algorithms designed for post-production applications.
Avoid when
- When the focus is on generic performance evaluations rather than on real-world, task-specific benchmarks that assess handling complex post-production scenarios.
- If your budget or timeline cannot accommodate a per-task wall clock time of ~10 minutes and cost ranging from $0.10 to $2 based on agent token usage.
- Requirements:
- Requires Docker; Install via scripts provided in the repository.; Python virtual environment setup for reproducibility.
Observed Jul 16, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Very active (6d since push)
- As of Sep 9, 2026
- Provenance
- Not a fork · Organization account
- As of Sep 9, 2026
- Security (OSV)
- No lockfile
- As of Jul 15, 2026
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
pip install agentic-vbench PyPISimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
AgenticVBench evaluates AI agent capabilities to handle real-world post-production tasks such as audio restoration by measuring performance metrics, processing time, and cost. It provides specific task prompts that guide the agents through various restorative processes.
Capability facts
- Languages
- python
Source: github.language · Sep 9, 2026
Categories
Compatibility
Sourced claims from the README excerpt - not unsourced marketing copy.
Source: README excerpt (regex_v1, Sep 9, 2026)
python3 -m venv .venv && .venv/bin/pip install --upgrade pipSource link
Tags
README
1. Install reward.json → ≈ 1.0, 30 s on a cached image, zero agent cost bash ./avb results show rewards from the latest job cat jobs/<job name / /steps/solve/verifier/reward.json
For agents
This page has a .md twin and JSON over the API.