Capvise Labs / System 01

VetoBench

Will it generate or will it veto?

A benchmark in development for false refusals, inconsistent failures and user-content extension in generative-video systems.

VetoBench / pilot viewRepeated-run matrix
Illustrative
S-014Benign scene
CompletedCompletedRefused
S-027Owned clip extension
System errorCompletedCompleted
S-041Everyday action
CompletedTask missCompleted
C-006Control request
RefusedRefusedRefused
CompletedRefusedErrorTask miss
Status
Protocol development
Release
Pilot 0.1
Scores
None published

The missing measurement

Most video benchmarks begin after a video exists.

That leaves out a common failure: the request never reaches a useful output. The model refuses a benign scene, a filter blocks a user-owned clip, or the same prompt works once and fails on the next attempt.

VetoBench will measure that front door. Output quality still matters, but first the system has to perform the allowed task reliably.

Pilot tracks

Five ways a generation can become unreliable

01

Benign completion

Does a clearly allowed request reach generation on the first attempt?

02

Repeated-run stability

Do identical settings produce a stable success or refusal pattern across runs?

03

User-content extension

Can a supplied clip be continued or edited when the request is allowed?

04

Refusal calibration

Does the system block the control set while allowing the safe set?

05

Operational reliability

How often do provider errors, timeouts and retries prevent a usable result?

Outcome taxonomy

“It failed” is not one result.

The pilot will preserve the difference between a policy decision, an infrastructure fault and a video that technically renders but misses the task.

Completed

A video is returned and the requested task is materially fulfilled.

Explicit refusal

The system states that a policy or safety rule prevents generation.

Generic refusal

The system declines without a clear policy reason.

Provider error

The request fails because of an API, capacity, timeout or unknown system error.

Task miss

A video is returned, but it does not complete the requested action or edit.

Pilot shape

Small enough to run. Clear enough to challenge.

Prompt setSafe, control and extension cases
Repeated runsThree per case in the first pilot
Access recordEndpoint, version, settings and funding source
PublicationAggregate score plus inspectable failure cases

Founding access

Providers can supply the compute. They cannot supply the result.

Capvise Labs is seeking production-equivalent API access, capped credits and shared inference infrastructure for the pilot. Every material source of support will be disclosed on the relevant model page.

Support does not provide control over prompts, scoring, publication or rank. A provider may review factual endpoint details before publication, not rewrite the evaluation.

Participation routes

Leaderboard

No rankings yet.

The public table will open after the pilot protocol is frozen and the first model runs pass review. Until then, an empty leaderboard is more honest than invented numbers.

View leaderboard status