Research agenda

We start with questions that force systems to show their work.

Capvise Labs is at formation stage. These pages describe the problems we intend to pursue and the standards we will use to separate a promising idea from a real result.

Working thesis

Intelligence becomes useful when it can hold state, revise a plan and admit uncertainty.

Scale has produced systems that can answer difficult questions and generate rich media. The same systems still lose track of constraints, invent confidence and fail unpredictably when a task crosses tools, modalities or time.

We want to study those failures in complete systems. That means building models, environments and evaluations close enough to real use that weak assumptions become visible.

01

Research direction

Reasoning systems

We want to know whether a system can hold a plan together, notice when an assumption fails and explain why it changed course. The work will focus on complete tasks where mistakes compound, rather than isolated answers that hide the path taken.

Questions on the table

What should a model remember, discard and revisit?

When should a system stop and ask for evidence?

Can verification improve reliability without hiding extra cost?

02

Research direction

Multimodal intelligence

Perception is useful only when it changes reasoning. We are interested in systems that can combine evidence across modalities, preserve source information and remain useful when inputs are incomplete or contradictory.

Questions on the table

How should evidence from different modalities be reconciled?

Which tasks expose grounded reasoning rather than recognition?

How should uncertainty move across a multimodal pipeline?

03

Research direction

Agents and autonomy

Agent performance depends on state, permissions, recovery and control. We are interested in systems that verify consequential actions, expose their current plan and leave a clear path for interruption or rollback.

Questions on the table

What should be checked before an external action is taken?

How can operators understand a long-running task quickly?

Which failures should trigger pause, recovery or escalation?

04

Research direction

Safety and evaluation

A model can be impressive and still be unreliable. We study evaluation methods that separate capability, consistency, calibration and operational failure. VetoBench is the first public system in this direction.

Questions on the table

Which tests predict real product failure?

How should over-refusal and under-refusal be scored together?

What access and funding rules keep an evaluation credible?

Research standard

A claim should leave a trail.

Question

Name the failure

Describe the behaviour precisely enough that another person could disagree.

Instrument

Build what is missing

Create the environment, model or benchmark required to observe the problem.

Evidence

Keep the comparison fair

Record settings, cost, versions, retries and negative results beside the score.

Release

Publish the boundary

State what the result supports, what it does not and what changed after review.