Active Perception

본문
At the lab, this paradigm is most directly instantiated by embodied and robot vision research, where perception is coupled to action rather than passively describing a fully given scene. Affordance-grounding work builds actionable, task-driven representations from partial observations — learning open-vocabulary affordances at scale, and actively reconstructing the unseen 3D structure needed to ground an affordance through generative completion. Compact world models and 3D grounding support the planning and exploration side of the same loop, letting an agent reason about where to move or look next, while extracting relevant visual information rather than processing every view of a scene.
Across these works, the shared commitment is building perception systems that actively reconstruct, ground, and select information in service of action, rather than passively consuming a fully observed input.
Related papers
Chunghyun Park, Seunghyeon Lee, Minsu Cho. Affostruction: 3D Affordance Grounding with Generative Reconstruction. CVPR, 2026
Dongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho, Suha Kwak. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model. CVPR, 2026
Junha Lee*, Eunha Park*, Chunghyun Park, Dahyun Kang, Minsu Cho. Affogato: Learning Open-Vocabulary Affordance Grounding with Automatic Data Generation at Scale. ECCV, 2026
Seongmin Jung, Seongho Choi, Gunwoo Jeon, Minsu Cho, Jongwoo Lim. PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding. ECCV, 2026
Jinsoo Park, Donggyu Choi, Ahyun Seo, Minsu Cho, Jeany Son. Diversity-Aware View Partitioning for Scalable VGGT. ECCV, 2026