Hyunse Yoon

dblp:303/5893 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0004-4423-8544ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 SDAS: Semantic Data Acquisition System for Minimizing Redundancy and Maximizing Diversity
abstract
In this paper, we propose SDAS, a new motion data assessment and storage system designed to acquire new motion data with reduced redundancy and maximizing diversity. SDAS collects data in the field, retrieves the most similar data from the database in real-time, and provides visualization tools that allow for the comparison of differences between the capture data and the stored data. Through this system, researchers can efficiently build and manage a database. The demonstration video is available at https://youtu.be/vqW0uMDnZTw.
Yeseung Park, Hyunse Yoon, Jungwoo Huh, Jungsu Kim, Jeongwook Choi, Sanghoon Lee 0001
AAAI2
2025 Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA
abstract
Amodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and complex purposes. While methods like LISA integrate multi-modal large language models (LLMs) with segmentation for reasoning tasks, they are limited to predicting only visible object regions and face challenges in handling complex occlusion scenarios. To address these limitations, we propose a novel task named amodal reasoning segmentation, aiming to predict the complete amodal shape of occluded objects while providing answers with elaborations based on user text input. We develop a generalizable dataset generation pipeline and introduce a new dataset focusing on daily life scenarios, encompassing diverse real-world occlusions. Furthermore, we present AURA (Amodal Understanding and Reasoning Assistant), a novel model with advanced global and spatial-level designs specifically tailored to handle complex occlusions. Extensive experiments validate AURA's effectiveness on the proposed dataset.
Hyunse Yoon, Sanghoon Lee 0001, Weisi Lin
ICCV2
2023 Fusing Explicit and Implicit Flow for Optical Flow Estimation
abstract
Estimating optical flow for large movement remains a challenging issue due to inconsistency in features between frames. To resolve this challenge, we propose a novel sequence-based deep learning network that jointly trains explicit flow and implicit flow to accurately estimate optical flow. To do this, we implemented three submodules: explicit flow embedder, implicit flow embedder, and flow fusion network. Explicit flow embedder learns the pair-wise correlation between visible pixels based on the spatial attention made per image. Implicit flow embedder learns implicit flow based on the temporal context of motion from all frames in the sequence. To effectively learn the implicit flow, we give a longer sequence of frames as input. Flow fusion network fuses features from explicit and implicit embedder to output the final optical flow. Through extensive experiments, our model demonstrates its robustness against the large motion while providing accurate flow estimation for pixels without pairs in the next frame.
Hyunse Yoon, Seongmin Lee 0002, Sanghoon Lee 0001
ICIP1
2023 Video-Based Stabilized 3D Face Alignment Using Temporal Multi-Discrimination
abstract
Existing 3D face alignment primarily aim to achieve accurate face alignment result for a static facial image. While these methods have strong alignment performance under large poses, occlusion, and extreme lighting conditions, they often result in trembling artifacts in video-based sequential 3D face alignment. Reducing temporal misalignment remains a challenging task because a single misaligned frame can propagate errors to other frames along the temporal axis. To address this issue, we propose a novel temporal discriminating scheme that learns the distribution gap between the face alignment results and ground truth face animation. By leveraging the discrimination results as a guide, the proposed method can effectively align the 3D faces to the input video by reducing temporal trembling artifacts. To effectively learn the distribution gap, we introduce a multi-discriminating scheme that separately discriminates facial animation based on identity and expression changes. It enables the proposed method to produce a stabilized alignment result, especially in dynamic and fast movement. Through extensive experiments in both qualitative and quantitative evaluations, it is confirmed that our method outperforms state-of-the-art 3D face alignment methods by animating stabilized results in the video.
Seongmin Lee 0002, Hyunse Yoon, Jiwoo Kang 0001, Jungsu Kim, Jiwan Son, Jungwoo Huh, Sanghoon Lee 0001
MMSP2
2021 WarpingFusion: Accurate Multi-View TSDF Fusion with Local Perspective Warp
abstract
In this paper, we propose the novel 3D reconstruction framework, where the surface of a target object is reconstructed accurately and robustly from multi-view depth maps. A depth map of a moving object tends to have the spatially-varying perspective warps due to motion blur and rolling shutter artifacts. Incorporating those misaligned points from the views into the world coordinate leads to significant artifacts in the reconstructed shape. We address the mismatches by the patch-based depth-to-surface alignment using implicit surface-based distance measurement. The patch-based minimization finds spatial warps on the depth map fast and accurately with the global transformation preserved. The proposed framework efficiently optimizes the local alignments against depth occlusions and local variants thanks to the point to surface distance based on an implicit representation. The proposed method shows significant improvements over the other reconstruction methods, demonstrating efficiency and benefits of our method in the multi-view reconstruction.
Jiwoo Kang 0001, Seongmin Lee 0002, Mingyu Jang, Hyunse Yoon, Sanghoon Lee 0001
ICIP4
2021 Deep Chessboard Corner Detection Using Multi-task Learning
abstract
Camera calibration is an indispensable step in the fields of robotics and computer vision, which includes augmented reality, 3D reconstruction, and camera motion estimation. Before camera calibration, detecting matching correspondence is necessary to understand the structure of the world from multiple images. For an accurate result, a calibration object, such as a chessboard, is used. Existing handcrafted feature methods precisely detect chessboard corners but are weak against blurs, noises, and severe lens distortion. Conversely, neural network-based methods can detect corners regardless of noises in the image. Both methods do not utilize the information of camera priors, which are lens distortion and intrinsic parameters, affecting the location of chessboard corners. Learning of lens distortion and intrinsic parameters enables the proposed network to understand the alignment of corners more precisely. Therefore, in this paper, we propose a novel multi-task learning framework to detect chessboard corners and simultaneously estimate lens distortion and intrinsic parameters. In order to train these three tasks, synthetic images of the chessboard are generated with ground-truth labels corresponding to each task. Hence, by learning the camera priors, the proposed network can more precisely locate the corners than other state-of-the-art corner detection methods while robust to noises, blurs, and distortion.
Hyunse Yoon, Seongmin Lee 0002, Jiwoo Kang 0001, Sanghoon Lee 0001
MMSP1