Capvise Labs / AI research and engineering

Advancingmachine intelligence.

We are building a research program around reasoning, multimodal systems, autonomous agents and the methods needed to test them honestly.

Lab stageFormation
Public systemVetoBench
Applied ventureMarka
Research / Systems / EvaluationCurrent work below

Research agenda

The lab is early. The problems are already here.

Long tasks lose state. Agents take actions they cannot explain. Multimodal models confuse recognition with reasoning. Safety filters block useful work and miss other risks. These are engineering problems with research inside them.

See the questions we are pursuing

Featured system / VetoBench

Did the model do the work, refuse it, or simply fail?

Video leaderboards usually start after a video exists. VetoBench starts one step earlier. It records whether a safe request completes, whether that result repeats, and what the user pays in retries when it does not.

VetoBench / pilot viewRepeated-run matrix
Illustrative
S-014Benign scene
CompletedCompletedRefused
S-027Owned clip extension
System errorCompletedCompleted
S-041Everyday action
CompletedTask missCompleted
C-006Control request
RefusedRefusedRefused
CompletedRefusedErrorTask miss

Safety and evaluation

A refusal is not automatically safe. A completion is not automatically useful.

We will score both sides of the boundary: harmful requests that should be blocked and benign requests that should work. The funding and access behind each evaluation will be disclosed with the result.

Report the exact endpoint and model versionSeparate policy refusals from system errorsPublish limits beside the scoreDo not sell influence over public rankings
Read the lab's safety position

Work with the lab

Bring a model, a hard evaluation problem or a serious research proposal.

We are not running a broad hiring campaign. We are open to focused collaboration, evaluation access and infrastructure support.