Date: September 23, 2026
Speaker: Ankit Goyal, Head of AI, Proception
Title: Reasoning Recipe for Robot Foundation Models
Host: Kuan Fang

A color photo of a man smiling for a photo.

Abstract: What does it take for a robot foundation model to reason beyond its training demonstrations? Robot foundation models are primarily trained to imitate, and can struggle when instructions are underspecified, tasks require unstated steps, or situations depart from their training data. In this talk, I will present a recipe for instilling reasoning into pretrained robot foundation models through their existing language-conditioning interface, without changing their architecture or training from scratch. The key ingredient is action-grounded causal reasoning: linking each action to why it is needed and what it is intended to achieve. I will describe how we enrich robot demonstrations with this supervision and incorporate it through lightweight fine-tuning. Across both world action models and vision-language-action models, we find that stronger grounding of causal language leads to better task execution. These findings suggest a practical path toward more capable robot foundation models: teaching existing models to reason about their actions.

Bio: Ankit Goyal is the Head of AI at Proception and an incoming Assistant Professor in the Robotics Department at the University of Michigan. His research focuses on robot learning, dexterous manipulation, robot foundation models, and 3D perception. Previously, he was a Senior Research Scientist at NVIDIA. He received his Ph.D. in Computer Science from Princeton University, his M.S. in Computer Science and Engineering from the University of Michigan, and his B.Tech. in Electrical Engineering from IIT Kanpur. His honors include recognition as an RSS Pioneer and a Qualcomm Innovation Fellowship.