Actions, Events & Long-Horizon Reasoning from Videos
Beyond recognizing what is in a single frame, video understanding requires reasoning about how content unfolds over time — segmenting a stream into meaningful units, anticipating what comes next, locating precise moments of change, answering questions that require following a narrative, and tracking specific objects reliably as they move and occlude. The lab has a sustained program across all of these temporal reasoning problems, and a defining trend across the group's recent work is treating generative modeling — diffusion in particular — and large language models not as end tasks in themselves, but as flexible backbones for representing and reasoning about temporal structure.