Project note
Why we are building VetoBench
Generative-video benchmarks measure output quality. Far fewer measure whether an allowed request produces anything at all.
The missing result
A video model can rank well on visual quality and still waste a user's time. Safe prompts may be refused, identical requests may alternate between success and failure, and user-owned clips may become impossible to extend after a filter change.
Those outcomes are usually treated as support issues. We think they belong in evaluation. A model that does not reliably begin the task has a capability problem, even when its successful outputs look excellent.
What the benchmark should measure
The first VetoBench pilot will focus on benign completion, repeated-run stability, user-content extension and refusal calibration. It will also separate explicit policy refusals from generic errors, provider failures and outputs that complete technically but miss the requested task.
The objective is not to reward a model for saying yes to everything. A useful safety system must block genuinely disallowed requests while allowing ordinary creative work to proceed.
Current status
VetoBench is in protocol development. No model scores or provider partnerships have been announced. Capvise Labs is preparing a small pilot, an evaluation-access covenant and a process for publishing provider-supplied access without allowing provider control over the result.
Document status
This page records the lab's position at publication. It does not claim an unreleased model, completed benchmark run or peer-reviewed result. Material changes will be dated.