Research direction
Reasoning systems
We want to know whether a system can hold a plan together, notice when an assumption fails and explain why it changed course. The work will focus on complete tasks where mistakes compound, rather than isolated answers that hide the path taken.
Questions on the table
What should a model remember, discard and revisit?
When should a system stop and ask for evidence?
Can verification improve reliability without hiding extra cost?