Heansung Lee 0001

dblp:258/3748-1 · also Hean Sung Lee 0001 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0003-4012-3219ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Spatio-temporal Feature-level Augmentation Vision Transformer for video-based person re-identification
Minjung Kim 0002, MyeongAh Cho, Heansung Lee 0001, Sangyoun Lee
Pattern Recognit.3
2024 Multi-Scale Structural Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCNs) have attracted considerable interest in skeleton-based action recognition. Existing GCN-based models have proposed methods to learn dynamic graph topologies generated from the feature information of vertices to capture inherent relationships. However, these models have two main limitations. Firstly, they struggle to effectively utilize high-dimensional or structural information, which limits their capacity for feature representation and consequently hinders performance improvement. Secondly, among these models, the multi-scale methods that aggregate information at different scales often over-capture unnecessary relationships between vertices. This leads to an over-smoothing problem where smoothed features are extracted, making it difficult to distinguish the features of each vertex. To address these limitations, we propose the multi-scale structural graph convolutional network (MSS-GCN) for skeleton-based action recognition. Within the MSS-GCN framework, the common intersection graph convolution (CI-GC) leverages the overlapped neighbor information, indicating the overlap between neighboring vertices for a given pair of root vertices. The graph topology of CI-GC is designed to compute the structural correlation between neighboring vertices corresponding to each hop, thereby enriching the context of inter-vertex relationships. Then, our proposed multi-scale spatio-temporal modeling aggregates local-global features to provide a comprehensive representation. In addition, we propose a Graph Weight Annealing (GWA) method, which is a graph scheduling method to mitigate the over-smoothing caused by multi-scale aggregation. By varying the importance between a vertex and its neighbors, we demonstrate that the over-smoothing problem can be effectively mitigated. Moreover, our proposed GWA method can easily be adapted to different GCN models to enhance performance. Combining the MSS-GCN model and the GWA method, we propose a powerful feature extractor that effectively classifies actions for skeleton-based action recognition in various datasets. We evaluate our approach on three benchmark datasets: NTU RGB+D, NTU RGB+D 120, and NW-UCLA. The proposed MSS-GCN achieves state-of-the-art performance on all three datasets, further validating the effectiveness of our approach.
Sungjun Jang, Heansung Lee 0001, Woo Jin Kim, Sungmin Woo, Sangyoun Lee
IEEE Trans. Circuits Syst. Video Technol.2
2022 Tackling Background Distraction in Video Object Segmentation
Suhwan Cho, Heansung Lee 0001, Minhyeok Lee, Sungjun Jang, Minjung Kim 0002, Sangyoun Lee
ECCV (22)2
2022 Occluded Person Re-Identification Via Relational Adaptive Feature Correction Learning
abstract
Occluded person re-identification (Re-ID) in images captured by multiple cameras is challenging because the target person is occluded by pedestrians or objects, especially in crowded scenes. In addition to the processes performed during holistic person Re-ID, occluded person Re-ID involves the removal of obstacles and the detection of partially visible body parts. Most existing methods utilize the off-the-shelf pose or parsing networks as pseudo labels, which are prone to error. To address these issues, we propose a novel Occlusion Correction Network (OCNet) that corrects features through relational-weight learning and obtains diverse and representative features without using external networks. In addition, we present a simple concept of a center feature in order to provide an intuitive solution to pedestrian occlusion scenarios. Furthermore, we suggest the idea of Separation Loss (SL) for focusing on different parts between global features and part features. We conduct extensive experiments on five challenging benchmark datasets for occluded and holistic Re-ID tasks to demonstrate that our method achieves superior performance to state-of-the-art methods especially on occluded scene.
Minjung Kim 0002, MyeongAh Cho, Heansung Lee 0001, Suhwan Cho, Sangyoun Lee
ICASSP3
2022 Detection-Identification Balancing Margin Loss for One-Stage Multi-Object Tracking
abstract
In recent years, one-stage multi-object tracking (MOT) methods, which jointly learn detection and identification in a single network, have attracted extensive attention, due to their efficiency. However, the negative transfer effects caused by the two conflicting objectives of detection and identification have rarely been explored. In this paper, we propose a Detection-Identification Balancing Margin (DIM) loss for minimizing the adverse effects caused by these two different objectives. The proposed DIM loss consists of Detection Margin (DM) loss and Identification Margin (IM) loss. DM loss forces features that are farther from the center of the foreground features than the defined margin due to identification learning to be converged to ensure accurate detection. IM loss enables the various feature representations that are essential for identification by intentionally spreading features that become overly clustered due to detection learning. The proposed DIM loss demonstrates competitive and balanced performance for MOT by providing a positive transfer for features that had a strong negative impact on detection and identification, respectively. (HOTA 61.5, MOTA 75.3, IDF1 75.6 on MOT16, and real-time rates of 25.9 fps were achieved)
Heansung Lee 0001, Suhwan Cho, Sungjun Jang, Sungmin Woo, Sangyoun Lee
ICIP1
2022 Pixel-Level Bijective Matching for Video Object Segmentation
abstract
Semi-supervised video object segmentation (VOS) aims to track the designated objects present in the initial frame of a video at the pixel level. To fully exploit the appearance information of an object, pixel-level feature matching is widely used in VOS. Conventional feature matching runs in a surjective manner, i.e., only the best matches from the query frame to the reference frame are considered. Each location in the query frame refers to the optimal location in the reference frame regardless of how often each reference frame location is referenced. This works well in most cases and is robust against rapid appearance variations, but may cause critical errors when the query frame contains background distractors that look similar to the target object. To mitigate this concern, we introduce a bijective matching mechanism to find the best matches from the query frame to the reference frame and vice versa. Before finding the best matches for the query frame pixels, the optimal matches for the reference frame pixels are first considered to prevent each reference frame pixel from being overly referenced. As this mechanism operates in a strict manner, i.e., pixels are connected if and only if they are the sure matches for each other, it can effectively eliminate background distractors. In addition, we propose a mask embedding module to improve the existing mask propagation method. By embedding multiple historic masks with coordinate information, it can effectively capture the position information of a target object. Code and models are available at https://github.com/suhwan-cho/BMVOS.
Suhwan Cho, Heansung Lee 0001, Minjung Kim 0002, Sungjun Jang, Sangyoun Lee
WACV2
2022 SSAT: Self-Supervised Associating Network for Multiobject Tracking
abstract
Multi-object tracking (MOT), which is crucial for computer vision and video processing, has immense potential for improvement. Traditional tracking-by-detection approaches include feature-based object re-identification methods that use trained features, but these methods suffer from a lack of suitable training data. In training datasets used for MOT, every object in a video sequence must have its own location and ID. However, assigning IDs to each object in every sequence is considerably labor-intensive, and hence current MOT datasets are unsuitable for training re-identification networks. To resolve this issue, this paper proposes a novel self-supervised learning method using several short videos that contain no human-added labels, based on the idea that each video is a set of temporally corresponding image frames. We then describe how to improve tracking performance using a re-identification network trained in a self-supervised manner. In addition, ablation studies were conducted in order to define the optimal parameters, such as number of clips, data augmentation, and appropriate matching algorithms. The proposed approach achieved competitive performance compared with current best-practice methods including supervised methods, achieving MOT accuracy = 62.0% and ID F1-score = 62.7% on the MOT17 benchmark.
Tae-Young Chung, MyeongAh Cho, Heansung Lee 0001, Sangyoun Lee
IEEE Trans. Circuits Syst. Video Technol.3
2020 Crvos: Clue Refining Network For Video Object Segmentation
abstract
The encoder-decoder based methods for semi-supervised video object segmentation (Semi-VOS) have received extensive attention due to their superior performances. However, most of them have complex intermediate networks which generate strong specifiers to be robust against challenging scenarios, and this is quite inefficient when dealing with relatively simple scenarios. To solve this problem, we propose a real-time network, Clue Refining Network for Video Object Segmentation (CRVOS), that does not have any intermediate network to efficiently deal with these scenarios. In this work, we propose a simple specifier, referred to as the Clue, which consists of the previous frame’s coarse mask and coordinates information. We also propose a novel refine module which shows the better performance compared with the general ones by using a deconvolution layer instead of a bilinear upsampling layer. Our proposed method shows the fastest speed among the existing methods with a competitive accuracy. On DAVIS 2016 validation set, our method achieves 63.5 fps and $\mathcal{J} \& \mathcal{F}$ score of 81.6%.
Suhwan Cho, MyeongAh Cho, Tae-Young Chung, Heansung Lee 0001, Sangyoun Lee
ICIP4
2019 Depth map upsampling with a confidence-based joint guided filter
Yoonmo Yang, Heansung Lee 0001, Byung Tae Oh
Signal Process. Image Commun.2