Sungmin Woo

dblp:73/7651 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-1378-6722ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 DualFocus: Depth from Focus with Spatio-Focal Dual Variational Constraints
abstract
Depth-from-Focus (DFF) enables precise depth estimation by analyzing focus cues across a stack of images captured at varying focal lengths. While recent learning-based approaches have advanced this field, they often struggle in complex scenes with fine textures or abrupt depth changes, where focus cues may become ambiguous or misleading. We present DualFocus, a novel DFF framework that leverages the focal stack’s unique gradient patterns induced by focus variation, jointly modeling focus changes over spatial and focal dimensions. Our approach introduces a variational formulation with dual constraints tailored to DFF: spatial constraints exploit gradient pattern changes across focus levels to distinguish true depth edges from texture artifacts, while focal constraints enforce unimodal, monotonic focus probabilities aligned with physical focus behavior. These inductive biases improve robustness and accuracy in challenging regions. Comprehensive experiments on four public datasets demonstrate that DualFocus consistently outperforms state-of-the-art methods in both depth accuracy and perceptual quality.
Sungmin Woo, Sangyoun Lee
NeurIPS1
2024 ProDepth: Boosting Self-supervised Multi-frame Monocular Depth with Probabilistic Fusion
Sungmin Woo, Wonjoon Lee, Woo Jin Kim, Dogyoon Lee, Sangyoun Lee
ECCV (3)1
2024 FIMP: Future Interaction Modeling for Multi-Agent Motion Prediction
abstract
Multi-agent motion prediction is a crucial concern in autonomous driving, yet it remains a challenge owing to the ambiguous intentions of dynamic agents and their intricate interactions. Existing studies have attempted to capture interactions between road entities by using the definite data in history timesteps, as future information is not available and involves high uncertainty. However, without sufficient guidance for capturing future states of interacting agents, they frequently produce unrealistic trajectory overlaps. In this work, we propose Future Interaction modeling for Motion Prediction (FIMP), which captures potential future interactions in an end-to-end manner. FIMP adopts a future decoder that implicitly extracts the potential future information in an intermediate feature-level, and identifies the interacting entity pairs through future affinity learning and top-k filtering strategy. Experiments show that our future interaction modeling improves the performance remarkably, leading to superior performance on the Argoverse motion forecasting benchmark.
Sungmin Woo, Minjung Kim 0002, Donghyeong Kim, Sungjun Jang, Sangyoun Lee
ICRA1
2024 Multi-Scale Structural Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCNs) have attracted considerable interest in skeleton-based action recognition. Existing GCN-based models have proposed methods to learn dynamic graph topologies generated from the feature information of vertices to capture inherent relationships. However, these models have two main limitations. Firstly, they struggle to effectively utilize high-dimensional or structural information, which limits their capacity for feature representation and consequently hinders performance improvement. Secondly, among these models, the multi-scale methods that aggregate information at different scales often over-capture unnecessary relationships between vertices. This leads to an over-smoothing problem where smoothed features are extracted, making it difficult to distinguish the features of each vertex. To address these limitations, we propose the multi-scale structural graph convolutional network (MSS-GCN) for skeleton-based action recognition. Within the MSS-GCN framework, the common intersection graph convolution (CI-GC) leverages the overlapped neighbor information, indicating the overlap between neighboring vertices for a given pair of root vertices. The graph topology of CI-GC is designed to compute the structural correlation between neighboring vertices corresponding to each hop, thereby enriching the context of inter-vertex relationships. Then, our proposed multi-scale spatio-temporal modeling aggregates local-global features to provide a comprehensive representation. In addition, we propose a Graph Weight Annealing (GWA) method, which is a graph scheduling method to mitigate the over-smoothing caused by multi-scale aggregation. By varying the importance between a vertex and its neighbors, we demonstrate that the over-smoothing problem can be effectively mitigated. Moreover, our proposed GWA method can easily be adapted to different GCN models to enhance performance. Combining the MSS-GCN model and the GWA method, we propose a powerful feature extractor that effectively classifies actions for skeleton-based action recognition in various datasets. We evaluate our approach on three benchmark datasets: NTU RGB+D, NTU RGB+D 120, and NW-UCLA. The proposed MSS-GCN achieves state-of-the-art performance on all three datasets, further validating the effectiveness of our approach.
Sungjun Jang, Heansung Lee 0001, Woo Jin Kim, Sungmin Woo, Sangyoun Lee
IEEE Trans. Circuits Syst. Video Technol.5
2023 Leveraging Spatio-Temporal Dependency for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has attracted considerable attention due to its compact representation of the human body’s skeletal sructure. Many recent methods have achieved remarkable performance using graph convolutional networks (GCNs) and convolutional neural networks (CNNs), which extract spatial and temporal features, respectively. Although spatial and temporal dependencies in the human skeleton have been explored separately, spatio-temporal dependency is rarely considered. In this paper, we propose the Spatio-Temporal Curve Network (STC-Net) to effectively leverage the spatio-temporal dependency of the human skeleton. Our proposed network consists of two novel elements: 1) The Spatio-Temporal Curve (STC) module; and 2) Dilated Kernels for Graph Convolution (DK-GC). The STC module dynamically adjusts the receptive field by identifying meaningful node connections between every adjacent frame and generating spatio-temporal curves based on the identified node connections, providing an adaptive spatio-temporal coverage. In addition, we propose DK-GC to consider long-range dependencies, which results in a large receptive field without any additional parameters by applying an extended kernel to the given adjacency matrices of the graph. Our STC-Net combines these two modules and achieves state-of-the-art performance on four skeleton-based action recognition benchmarks. Code is available at https://github.com/Jho-Yonsei/STC-Net.
Minhyeok Lee, Suhwan Cho, Sungmin Woo, Sungjun Jang, Sangyoun Lee
ICCV4
2023 MSV-RGNN: Multiscale Voxel Graph Neural Network for 3D Object Detection
abstract
This paper proposes a two-stage 3D object detection framework, multiscale voxel graph neural network (MSV-RGNN) which aims to fully exploit multiple scale graph features by establishing global and local relationships between voxel features at different 3D convolutional neural network (CNN) layers. In contrast to conventional graph-based methods, our proposed multiscale-voxel-graph region-of-interest (RoI) pooling module constructs graphs across diverse voxel resolutions to obtain geometric structure information on voxel features. Initially, our multiscale-voxel-graph RoI pooling module sample voxel center points with voxel-wise feature vectors and 3D region proposals from backbone network. Subsequently, graphs are constructed at different scales and graph features are aggregated for second-stage refinement. The experimental results demonstrate the potential of using multiscale graphs across different voxel resolutions for 3D object detection, achieving decent experimental results with state-of-the-art methods.
Wonjoon Lee, Sungmin Woo, Donghyeong Kim, Sangyoun Lee
ICIP2
2023 MKConv: Multidimensional feature representation for point cloud analysis
abstract
Despite the remarkable success of deep learning , an optimal convolution operation on point clouds remains elusive owing to their irregular data structure . Existing methods mainly focus on designing an effective continuous kernel function that can handle an arbitrary point in continuous space. Various approaches exhibiting high performance have been proposed, but we observe that the standard pointwise feature is represented by 1D channels and can become more informative when its representation involves additional spatial feature dimensions. In this paper, we present Multidimensional Kernel Convolution (MKConv), a novel convolution operator that learns to transform the point feature representation from a vector to a multidimensional matrix. Unlike standard point convolution, MKConv proceeds via two steps. (i) It first activates the spatial dimensions of local feature representation by exploiting multidimensional kernel weights. These spatially expanded features can represent their embedded information through spatial correlation as well as channel correlation in feature space , carrying more detailed local structure information. (ii) Then, discrete convolutions are applied to the multidimensional features which can be regarded as a grid-structured matrix. In this way, we can utilize the discrete convolutions for point cloud data without voxelization that suffers from information loss. Furthermore, we propose a spatial attention module, Multidimensional Local Attention (MLA), to provide comprehensive structure awareness within the local point set by reweighting the spatial feature dimensions. We demonstrate that MKConv has excellent applicability to point cloud processing tasks including object classification, object part segmentation, and scene semantic segmentation with superior results.
Sungmin Woo, Dogyoon Lee, Woo Jin Kim, Sangyoun Lee
Pattern Recognit.1
2022 Detection-Identification Balancing Margin Loss for One-Stage Multi-Object Tracking
abstract
In recent years, one-stage multi-object tracking (MOT) methods, which jointly learn detection and identification in a single network, have attracted extensive attention, due to their efficiency. However, the negative transfer effects caused by the two conflicting objectives of detection and identification have rarely been explored. In this paper, we propose a Detection-Identification Balancing Margin (DIM) loss for minimizing the adverse effects caused by these two different objectives. The proposed DIM loss consists of Detection Margin (DM) loss and Identification Margin (IM) loss. DM loss forces features that are farther from the center of the foreground features than the defined margin due to identification learning to be converged to ensure accurate detection. IM loss enables the various feature representations that are essential for identification by intentionally spreading features that become overly clustered due to detection learning. The proposed DIM loss demonstrates competitive and balanced performance for MOT by providing a positive transfer for features that had a strong negative impact on detection and identification, respectively. (HOTA 61.5, MOTA 75.3, IDF1 75.6 on MOT16, and real-time rates of 25.9 fps were achieved)
Heansung Lee 0001, Suhwan Cho, Sungjun Jang, Sungmin Woo, Sangyoun Lee
ICIP5
2022 LiDAR Depth Completion Using Color-Embedded Information via Knowledge Distillation
abstract
Depth completion is the task of reconstructing dense depth images from sparse LiDAR data. LiDAR depth completion, for which LiDAR data is the only input, is an ill-posed and challenging problem owing to the underlying properties of LiDAR data: extremely few points, presence of discontinuities, and absence of texture information. Accordingly, most approaches are heavily dependent on guided color images, which leads to unsatisfactory results when the color images are degraded. To alleviate the dependency on color images but leverage this information during training, we present a deep convolutional neural network (CNN) consisting of depth and edge CNNs via transferring of knowledge. In order to compensate for the limitations of LiDAR data, we design the edge CNN to learn a gradient depth image from a powerful teacher network through theKnowledge-Distillationmethod. Since the teacher network is trained with color images, color-embedded information can be obtained in the test phase even if color images are not used as an input. We further propose aSelf-Distillationmethod for transferring the color-embedded features from the edge CNN to the depth CNN. Enforcing the depth features to contain edge information hardly observed in LiDAR data enables the depth CNN to generate more edge-attentive and structure-preserving results. Our novel methods show remarkable results in outdoor and indoor environments for KITTI and NYU-Depth-V2 datasets. Experiments performed with low-channel LiDAR data in KITTI and few depth points in the NYU-Depth-V2 dataset show that our method is robust to data sparsity and applicable in various scenarios.
Junhyeop Lee, Woo Jin Kim, Sungmin Woo, Kyungjae Lee 0003, Sangyoun Lee
IEEE Trans. Intell. Transp. Syst.4
2022 AIBM: Accurate and Instant Background Modeling for Moving Object Detection
abstract
Detecting moving objects has been widely studied since it plays vital many applications, such as video surveillance and intelligent transportation systems. It is necessary to accurately differentiate the foreground and the background in this technology to analyze object motions in the scene. Conventional detection methods use many reference frames to model the background to detect moving objects; however, the detection is inaccurate when immediate changes occur in the scene because the instant update of the background model is impossible. To be robust illumination changes and dynamic backgrounds, we propose an accurate and instant background modeling (AIBM) method that inpaints the background with superpixels and enhances it in detail with pixel-levels. Unlike the previous approaches, the proposed AIBM method utilizes spatio-temporal information of only two consecutive frames to eliminate the lengthy initialization and update period of the model. In this paper, we illustrate the importance of accurate and instant background modeling in detecting moving objects. The performance of our method is evaluated with three benchmark datasets (CDnet2014, LASIESTA, and SBI). The experimental results show that our AIBM method is robust to sudden changes in the scene and outperforms the other conventional methods with F-measure of 0.8911 and 0.9682 in detection accuracy in CDNet2014 and LASIESTA datasets, respectively. The accuracy of the generated background is also measured on the SBI dataset, which demonstrates the importance of high-quality background modeling.
Woo Jin Kim, Junhyeop Lee, Sungmin Woo, Sangyoun Lee
IEEE Trans. Intell. Transp. Syst.4
2021 Regularization Strategy for Point Cloud via Rigidly Mixed Sample
abstract
Data augmentation is an effective regularization strategy to alleviate the overfitting, which is an inherent drawback of the deep neural networks. However, data augmentation is rarely considered for point cloud processing despite many studies proposing various augmentation methods for image data. Actually, regularization is essential for point clouds since lack of generality is more likely to occur in point cloud due to small datasets. This paper proposes a Rigid Subset Mix (RSMix)1, a novel data augmentation method for point clouds that generates a virtual mixed sample by replacing part of the sample with shape-preserved subsets from another sample. RSMix preserves structural information of the point cloud sample by extracting subsets from each sample without deformation using a neighboring function. The neighboring function was carefully designed considering unique properties of point cloud, unordered structure and non-grid. Experiments verified that RSMix successfully regularized the deep neural networks with remarkable improvement for shape classification. We also analyzed various combinations of data augmentations including RSMix with single and multi-view evaluations, based on abundant ablation studies.
Dogyoon Lee, Jaeha Lee, Junhyeop Lee, Hyeongmin Lee, Minhyeok Lee, Sungmin Woo, Sangyoun Lee
CVPR6
2020 False Positive Removal for 3D Vehicle Detection With Penetrated Point Classifier
abstract
Recently, researchers have been leveraging LiDAR point cloud for higher accuracy in 3D vehicle detection. Most state-of-the-art methods are deep learning based, but are easily affected by the number of points generated on the object. This vulnerability leads to numerous false positive boxes at high recall positions, where objects are occasionally predicted with few points. To address the issue, we introduce Penetrated Point Classifier (PPC) based on the underlying property of LiDAR that points cannot be generated behind vehicles. It determines whether a point exists behind the vehicle of the predicted box, and if does, the box is distinguished as false positive. Our straightforward yet unprecedented approach is evaluated on KITTI dataset and achieved performance improvement of PointRCNN, one of the state-of-the-art methods. The experiment results show that precision at the highest recall position is dramatically increased by 15.46 percentage points and 14.63 percentage points on the moderate and hard difficulty of car class, respectively.
Sungmin Woo, Woo Jin Kim, Junhyeop Lee, Dogyoon Lee, Sangyoun Lee
ICIP1