Safety

Safety work starts before a release note.

Capvise Labs does not operate a frontier model today. This page sets the standard we intend to apply as our systems, access and capability grow.

Position

Useful systems need a well-calibrated boundary, not a blanket yes or no.

Under-refusal can enable harm. Over-refusal can make ordinary work unreliable and drive users toward less controlled alternatives. Both outcomes deserve measurement.

Our first public safety project, VetoBench, applies this view to generative video. It will count benign failures beside appropriate refusals rather than treating any blocked request as a safety success.

See VetoBench

Commitments

What a release should make inspectable

01

Define the boundary

State what a system is intended to do, which failure classes matter and who owns the release decision.

02

Record the conditions

Keep model versions, settings, access source, retries and material human intervention with the result.

03

Test both directions

Measure harmful behaviour that should be blocked and useful behaviour that safety systems should allow.

04

Change the claim

When evidence weakens, narrow the release or revise the public statement instead of defending old copy.

Impact

Evidence

Release test

The burden of evidence should rise with the cost of being wrong.

Capability

What does the system complete under controlled and adversarial conditions?

Reliability

Does the behaviour repeat across time, accounts, settings and near-identical tasks?

Control

Can an operator understand, constrain, interrupt and recover the system?

Impact

Who pays when the system fails, and how quickly can the damage be reversed?