Research

본문 바로가기

Research

Active Perception

Active perception is grounded in a paradigm shift that computer vision underwent as it moved beyond passive, all-at-once scene analysis. Rather than treating a camera or a model as a fixed device that must take in an entire scene and try to make sense of everything it sees, active perception treats the observer as something that can choose — where to look, how to move through and interact with a scene, or which view to prioritize — in order to answer a specific question as efficiently as possible.
Active Perception

본문

At the lab, this paradigm is most directly instantiated by embodied and robot vision research, where perception is coupled to action rather than passively describing a fully given scene. Affordance-grounding work builds actionable, task-driven representations from partial observations — learning open-vocabulary affordances at scale, and actively reconstructing the unseen 3D structure needed to ground an affordance through generative completion. Compact world models and 3D grounding support the planning and exploration side of the same loop, letting an agent reason about where to move or look next, while extracting relevant visual information rather than processing every view of a scene.

Across these works, the shared commitment is building perception systems that actively reconstruct, ground, and select information in service of action, rather than passively consuming a fully observed input.

Related papers

  • Chunghyun Park, Seunghyeon Lee, Minsu Cho. Affostruction: 3D Affordance Grounding with Generative Reconstruction. CVPR, 2026

  • Dongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho, Suha Kwak. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model. CVPR, 2026

  • Junha Lee*, Eunha Park*, Chunghyun Park, Dahyun Kang, Minsu Cho. Affogato: Learning Open-Vocabulary Affordance Grounding with Automatic Data Generation at Scale. ECCV, 2026

  • Seongmin Jung, Seongho Choi, Gunwoo Jeon, Minsu Cho, Jongwoo Lim. PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding. ECCV, 2026

  • Jinsoo Park, Donggyu Choi, Ahyun Seo, Minsu Cho, Jeany Son. Diversity-Aware View Partitioning for Scalable VGGT. ECCV, 2026

Computer Vision Laboratory E2 302, Dept. of CSE, POSTECH 77 Cheongam Rd, Nam-gu, Pohang, Gyeongbuk, 37673 Korea