Capvise Labs / AI research and engineering
Advancingmachine intelligence.
We are building a research program around reasoning, multimodal systems, autonomous agents and the methods needed to test them honestly.
Current work
Real projects, plain status labels. No placeholder launches.
Research agenda
The lab is early. The problems are already here.
Long tasks lose state. Agents take actions they cannot explain. Multimodal models confuse recognition with reasoning. Safety filters block useful work and miss other risks. These are engineering problems with research inside them.
See the questions we are pursuingResearch directions
Questions worth building around
These are directions, not claims of completed programs. Each one names a problem we intend to test in working systems.
Systems and ventures
Two projects. Two different jobs.
VetoBench is public research infrastructure. Marka is an applied venture for human data and evaluation work.
Featured system / VetoBench
Did the model do the work, refuse it, or simply fail?
Video leaderboards usually start after a video exists. VetoBench starts one step earlier. It records whether a safe request completes, whether that result repeats, and what the user pays in retries when it does not.
Safety and evaluation
A refusal is not automatically safe. A completion is not automatically useful.
We will score both sides of the boundary: harmful requests that should be blocked and benign requests that should work. The funding and access behind each evaluation will be disclosed with the result.
Notes
Work we can stand behind
The first releases are project and governance notes. Papers, datasets and benchmark results will appear only after they exist.
Work with the lab
Bring a model, a hard evaluation problem or a serious research proposal.
We are not running a broad hiring campaign. We are open to focused collaboration, evaluation access and infrastructure support.