Wonjun Kim 0001

dblp:07/4324-1 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0001-5121-5931ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Identity-Aware Fusion for Interactive Human Mesh Reconstruction
abstract
Even though recent studies have demonstrated significant improvements in 3D human mesh reconstruction, most methods still face challenges under occlusions. To mitigate this issue, multi-view human mesh reconstruction methods have been proposed, which leverage additional visual cues by aggregating the complementary information. However, in scenes with close human interactions, the identity of each person is hard to be preserved across different views, thus leading to incorrect merging of visual features between individuals. To address this limitation, we propose an identity-aware fusion scheme for interactive human mesh reconstruction. The key idea is to maintain the identity of each person across different views based on the shape feature from the anchor view, which is determined by the highest visibility of each person. This helps the model to distinguish features encoded from each person during the aggregation process, even under close interactions with severe occlusions, leading to the reliable reconstruction of 3D human meshes. Furthermore, we design a weighting scheme to adaptively fuse joints belonging to the same identity across different views while allowing the model to focus more on visible joints. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of interactive human mesh reconstruction.
SungMin Jang, Wonjun Kim 0001
IEEE Signal Process. Lett.2
2025 Clustering-Based Adaptive Query Generation for Semantic Segmentation
abstract
Semantic segmentation is one of the crucial tasks in the field of computer vision, aiming to label each pixel according to its class. Most recently, several semantic segmentation methods, which adopt the transformer decoder with learnable queries, have achieved the impressive improvement. However, since learnable queries are primarily determined by the distribution of training samples, discriminative characteristics of the input image often have been disregarded. In this letter, we propose a novel clustering-based query generation method for semantic segmentation. The key idea of the proposed method is to adaptively generate queries based on the clustering scheme, which leverages semantic affinities in the latent space. By aggregating latent features that represent the same class in a given input, the semantic information of each class can be efficiently encoded into the query. Furthermore, we propose to apply the auxiliary loss function to predict the segmentation result in a coarse scale during the process of query generation. This enables each query to grasp spatial information of the target object in a given image. Experimental results on various benchmarks show that the proposed method effectively improves the performance of semantic segmentation. The code and model are publicly available at:https://github.com/DCVL-Seg/CQG-release.
Yeong Woo Kim, Wonjun Kim 0001
IEEE Signal Process. Lett.2
2025 SVD-Guided Diffusion for Training-Free Low-Light Image Enhancement
abstract
Low-light image enhancement aims to improve the visibility and the contrast of images captured under poor lighting conditions while preserving contextual details. In this context, most previous methods have relied on the paired training data, which often leads to overfitting to specific data distributions. Although recent approaches have adopted generative priors of the diffusion model to avoid such learning bias, the stochastic nature of the diffusion model restricts the precise control over luminance-related features. To address these challenges, we propose a novel and training-free method that integrates the Singular Value Decomposition (SVD) with a pretrained diffusion model. Based on our observation that SVD tends to separate an image into luminance and structural components, we propose to leverage the decomposition capability of SVD and the generative prior of the diffusion model simultaneously. Specifically, our approach effectively guides the restoration process of lighting conditions by adaptively combining singular values of the intermediate result, which is obtained from each denoising step, with those of low-light input. For this combination, we define a semantic-aware scaling scheme based on a vision-language model. Experimental results on benchmark datasets demonstrate that the proposed method efficiently improves the performance of low-light image enhancement compared to other training-free methods.
Jingi Kim, Wonjun Kim 0001
IEEE Signal Process. Lett.2
2025 Shape-Selective Splatting: Regularizing the Shape of Gaussian for Sparse-View Rendering
abstract
In recent years, 3D Gaussian splatting (3DGS) has shown high-fidelity rendering results in real-time. However, 3DGS often encounters the overfitting problem under sparse-view conditions due to insufficient cross-view constraints. In this letter, to mitigate this limitation, we focus on the effect of Gaussian shapes on the scene reconstruction from sparse input views. The key idea is to allow each Gaussian to adaptively select its shape in accordance with the scene structure. Specifically, we propose to put a learnable parameter into Gaussian attributes, which indicates the probability of each shape. This indicator is optimized with other attributes while making each Gaussian change its shape to 1D, 2D, and 3D for representing edges, planar surfaces, and volumetric regions, respectively. Based on a geometrically accurate representation, the proposed method consequently alleviates the model from overfitting to a limited set of training views. Furthermore, we apply a depth regularization scheme within a set of selected pixels to precisely constrain positions of Gaussians. Experimental results on benchmark datasets show that the proposed method effectively improves the performance of novel view synthesis under sparse input views. The code and model are publicly available at: https://github.com/DCVL3D/SSS-release.
Gun Ryu, Wonjun Kim 0001
IEEE Signal Process. Lett.2
2025 Eigenpose: Occlusion-Robust 3D Human Mesh Reconstruction
abstract
A new approach for occlusion-robust 3D human mesh reconstruction from a single image is introduced in this paper. Since occlusion has emerged as a major problem to be resolved in this field, there have been meaningful efforts to deal with various types of occlusions (e.g., person-to-person occlusion, person-to-object occlusion, self-occlusion, etc.). Although many recent studies have shown the remarkable progress, previous regression-based methods still have respective limitations to handle occlusion problems due to the lack of the appearance information. To address this problem, we propose a novel method for human mesh reconstruction based on the pose-relevant subspace analysis. Specifically, we first generate a set of eigenvectors, so-called eigenposes, by conducting the singular value decomposition (SVD) of the pose matrix, which contains diverse poses sampled from the training set. These eigenposes are then linearly combined to construct a target body pose according to fusing coefficients, which are learned through the proposed network. Such combination of principal body postures (i.e., eigenposes) in a global manner gives a great help to cope with partial ambiguities by occlusions. Furthermore, we also propose to exploit a joint injection module that efficiently incorporates the spatial information of visible joints into the encoded feature during the estimation process of fusing coefficients. Experimental results on benchmark datasets demonstrate the ability of the proposed method to robustly reconstruct the human mesh under various occlusions occurring in real-world scenarios. The code and model are publicly available at: https://github.com/DCVL-3D/Eigenpose_release.
Mi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, Wonjun Kim 0001
IEEE Trans. Image Process.4
2024 Instance-Aware Contrastive Learning for Occluded Human Mesh Reconstruction
abstract
A simple yet effective method for occlusion-robust 3D human mesh reconstruction from a single image is presented in this paper. Although many recent studies have shown the remarkable improvement in human mesh reconstruction, it is still difficult to generate accurate meshes when person-to-person occlusion occurs due to the ambigu-ity of who a body part belongs to. To address this problem, we propose an instance-aware contrastive learning scheme. Specifically, joint features belonging to the target human are trained to be proximate with the center feature (i.e., feature extracted from the body center position). On the other hand, center features of different human instances are forced to be far apart so that joint features of each person can be clearly distinguished from others. By interpreting the joint possession based on such contrastive learning scheme, the proposed method easily understands the spatial occupancy of body parts for each person in a given image, thus can reconstruct reliable human meshes even with severely overlapped cases between multiple persons. Ex-perimental results on benchmark datasets demonstrate the robustness of the proposed method compared to previous approaches under person-to-person occlusions. The code and model are publicly available at: https://github.com/DCVL-3D/InstanceHMR_release.
Mi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, Wonjun Kim 0001
CVPR4
2024 GCN-assisted attention-guided UNet for automated retinal OCT segmentation
Dongsuk Oh, Jonghyeon Moon, Kyoungtae Park, Wonjun Kim 0001, Seungho Yoo, Hyungwoo Lee, Jiho Yoo
Expert Syst. Appl.4
2024 Learning scale-aware relationships via Laplacian decomposition-based transformer for 3D human pose estimation
Hyukmin Kwon, Seong Yong Lim, Wonjun Kim 0001
Multim. Syst.4
2024 One-class learning for face anti-spoofing via pseudo-negative sampling
Mi-Gyeong Gwon, Wonjun Kim 0001
Multim. Tools Appl.2
2023 Sampling is Matter: Point-Guided 3D Human Mesh Reconstruction
abstract
This paper presents a simple yet powerful method for 3D human mesh reconstruction from a single RGB image. Most recently, the non-local interactions of the whole mesh vertices have been effectively estimated in the transformer while the relationship between body parts also has begun to be handled via the graph model. Even though those approaches have shown the remarkable progress in 3D human mesh reconstruction, it is still difficult to directly infer the relationship between features, which are encoded from the 2D input image, and 3D coordinates of each vertex. To resolve this problem, we propose to design a simple feature sampling scheme. The key idea is to sample features in the embedded space by following the guide of points, which are estimated as projection results of 3D mesh vertices (i.e., ground truth). This helps the model to concentrate more on vertex-relevant features in the 2D space, thus leading to the reconstruction of the natural human pose. Furthermore, we apply progressive attention masking to precisely estimate local interactions between vertices even under severe occlusions. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of 3D human mesh reconstruction. The code and model are publicly available at: https://github.com/DCVL-3D/PointHMR_release.
Mi-Gyeong Gwon, Hyukmin Kwon, Gi-Mun Um, Wonjun Kim 0001
CVPR6
2023 Multi-scale and Multi-view Feature Blending for Free-viewpoint Rendering of Dynamic Humans
abstract
With the great success in neural radiance fields (NeRF), human-specific NeRF has been actively introduced in recent years. However, such human rendering techniques often have difficulties to recover subtle details of textural surfaces such as wrinkles on clothes, which are mostly generated by complicated human motions. To solve this problem, we propose a new method for human-specific NeRF based on multi-scale and multi-view feature blending. Specifically, multi-scale image features, which are sampled from different views, are combined into the transformer architecture. This blended feature successfully guides the NeRF network to render photo-realistic results by enhancing local details of the human performer. Experimental results on the ZJU-MoCap dataset show that the proposed method outperforms previous methods both in qualitative and quantitative evaluations.
Hyukmin Kwon, Seong Yong Lim, Wonjun Kim 0001
VCIP4
2023 Part-attentive kinematic chain-based regressor for 3D human modeling
Gi-Mun Um, Jeongil Seo, Wonjun Kim 0001
J. Vis. Commun. Image Represent.4
2023 Dynamic Residual Filtering With Laplacian Pyramid for Instance Segmentation
abstract
Various studies have been conducted on instance segmentation and made great strides over the past few years. Most recently, instance-specific mask generation via dynamic kernel predictions has shown the significant performance improvement even without bounding boxes as well as anchors. However, this scheme still does not fully consider dynamic properties since the size of the receptive field is not enough to cover the spatially-meaningful range due to memory limitations. Furthermore, the single-fused feature often fails to grasp complicated boundaries for objects of different sizes. In this article, we propose the dynamic residual filtering method with the Laplacian pyramid, which separately restores the global layout and local boundaries of instance masks. Specifically, we firstly apply the Laplacian pyramid-based decomposition scheme to features encoded from the backbone and subsequently restore sub-band mask residuals from coarse to fine pyramid levels. To do this, we design spatially-aware convolution filters to progressively capture the residual form of mask features at each level of the Laplacian pyramid while holding deformable receptive fields with dynamic offset information. This is fairly desirable since global and local properties of mask features can be accurately restored with keeping the spatial flexibility through the invertible process of the Laplacian reconstruction. Experimental results on the COCO dataset demonstrate that our proposed method achieves the state-of-the-art performance, i.e., 42.7% AP. The code and model are publicly available at:https://github.com/tjqansthd/LapMask.
Minsoo Song, Gi-Mun Um, Heekyung Lee, Jeongil Seo, Wonjun Kim 0001
IEEE Trans. Multim.5
2022 Decomposition and replacement: Spatial knowledge distillation for monocular depth estimation
Minsoo Song, Wonjun Kim 0001
J. Vis. Commun. Image Represent.2
2022 Direction-aware feedback network for robust lane detection
Jinhee Kim, Wonjun Kim 0001
Multim. Tools Appl.2
2021 Low-light image enhancement by diffusion pyramid with residuals
Wonjun Kim 0001
J. Vis. Commun. Image Represent.1
2021 Monocular Depth Estimation Using Laplacian Pyramid-Based Depth Residuals
abstract
With a great success of the generative model via deep neural networks, monocular depth estimation has been actively studied by exploiting various encoder-decoder architectures. However, the decoding process in most previous methods, which repeats simple up-sampling operations, probably fails to fully utilize underlying properties of well-encoded features for monocular depth estimation. To resolve this problem, we propose a simple but effective scheme by incorporating the Laplacian pyramid into the decoder architecture. Specifically, encoded features are fed into different streams for decoding depth residuals, which are defined by decomposition of the Laplacian pyramid, and corresponding outputs are progressively combined to reconstruct the final depth map from coarse to fine scales. This is fairly desirable to precisely estimate the depth boundary as well as the global layout. We also propose to apply weight standardization to pre-activation convolution blocks of the decoder architecture, which gives a great help to improve the flow of gradients and thus makes optimization easier. Experimental results on benchmark datasets constructed under various indoor and outdoor environments demonstrate that the proposed method is effective for monocular depth estimation compared to state-of-the-art models. The code and model are publicly available at: |https://github.com/tjqansthd/LapDepth-release|.
Minsoo Song, Seokjae Lim, Wonjun Kim 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 DSLR: Deep Stacked Laplacian Restorer for Low-Light Image Enhancement
abstract
Various images captured in complicated lighting conditions often suffer from deterioration of the image quality. Such poor quality not only dissatisfies the user expectation but also may lead to a significant performance drop in many applications. In this paper, anovel method for low-light image enhancement is proposed by leveraging useful propertiesof the Laplacian pyramid both in image and feature spaces. Specifically, the proposed method, so-called a deep stacked Laplacian restorer (DSLR), is capable of separately recovering the global illumination and local details from the original input, and progressively combining them in the image space. Moreover, the Laplacian pyramid defined in the feature space makes such recovering processes more efficient based on abundant connectionsof higher-order residuals in a multiscale structure. This decomposition-based scheme is fairly desirable for learning the highly nonlinear relation between degraded images and their enhanced results. Experimental results on various datasets demonstrate that the proposed DSLR outperforms state-of-the-art methods. The code and model are publicly available at: https://github.com/SeokjaeLIM/DSLR-release .
Seokjae Lim, Wonjun Kim 0001
IEEE Trans. Multim.2
2020 Top-down thermal tracking based on rotatable elliptical motion model for intelligent livestock breeding
Minji Kim 0002, Wonjun Kim 0001
Multim. Syst.2
2020 Attentive Feedback Feature Pyramid Network for Shadow Detection
abstract
Shadow detection is one of the most challenging issues in computer vision. Inspired by the great success of the convolutional neural network (CNN) for the problem of image restoration, learned features have been widely adopted for shadow detection. However, most existing methods still suffer from ambiguities driven by black-colored objects, which are not actually shaded, as well as the background clutter. In this letter, we propose the attentive feedback feature pyramid network (AFFPN) for shadow detection in a single image. The key idea of the proposed method is to extract shadow-relevant features based on multiple feedback modules, which are defined in the feature pyramid network. Specifically, attentive features extracted from each level of the encoder are progressively refined via connections between feedback modules from high-level to low-level layers for learning properties of shadow more accurately. Experimental results on benchmark datasets show that the proposed method is effective for shadow detection under complicated real-world environments. The code and model are publicly available at: https://github.com/JinheeKIM94/AFFPN_release.
Jinhee Kim, Wonjun Kim 0001
IEEE Signal Process. Lett.2
2020 Deep Spectral-Spatial Network for Single Image Deblurring
abstract
Inspired by the great success of the deep neural networks in various fields of computer vision, studies for image deblurring have begun to become more active in recent days. However, most previous approaches often fail to accurately remove the blur artifacts, e.g., ghosting effects at the object boundaries and degradation of local details, in restored results. In this paper, we propose a deep spectral-spatial network (DSSN) for resolving the problem of single image deblurring. Specifically, the proposed method is able to efficiently recover scene characteristics in a global manner by minimizing differences of the frequency magnitude between the blurred input and corresponding sharp image via the spectral restorer, and the spatial restorer fine-tunes local details of the intermediate result, which is estimated by the spectral one, based on the intensity similarity. This cascaded scheme of deblurring processes is fairly desirable for clearly restoring edge-like structures as well as the textural information in a coarse-to-fine manner. Experimental results on benchmark datasets demonstrate that the proposed DSSN outperforms state-of-the-art methods. The code and model are publicly available at: https://github.com/SeokjaeLIM/DSSN_release.
Seokjae Lim, Wonjun Kim 0001
IEEE Signal Process. Lett.3
2019 Multiple object tracking in soccer videos using topographic surface analysis
Wonjun Kim 0001
J. Vis. Commun. Image Represent.1
2019 Moving object detection using edges of residuals under varying illuminations
Wonjun Kim 0001
Multim. Syst.1
2019 Skeleton-Based Gait Recognition via Robust Frame-Level Matching
abstract
Gait is a useful biometric feature for human identification in video surveillance applications since it can be obtained without subject cooperation. In recent years, model-based gait recognition using a 3D skeleton has been widely studied through view-invariant modeling and kinematic gait analysis. However, existing methods integrate all frame-level feature vectors using the same criterion, even though skeleton information is highly sensitive to changes in covariate conditions such as clothing, carrying, and occlusion. The scheme inevitably reduces the frame-level discriminative power and eventually degrades performance. Instead, we propose a robust frame-level matching method for gait recognition that minimizes the influence of noisy patterns as well as secures the frame-level discriminative power. To this end, we measure the skeleton quality in terms of body symmetry for each frame. Based on the quality, we construct a quality-adjusted cost matrix between input frames and registered frames to prevent matching with noisy patterns. Our two-stage linear matching is then applied to the cost matrix to compute a frame-level discriminative score including similarity and margin. In the end, the identity of a probe is determined by a weighted majority voting scheme via frame-level scores. It enhances the robustness against inaccurate skeleton estimation results by assigning different weights for each frame based on the score. Our approach outperforms the state-of-the-art methods on three public datasets (UPCVgait, UPCVgaitK2, and SDUgait) and a new gait dataset which we create with consideration of unpredictable behaviors while walking. In addition, we demonstrate that our method is robust to skeleton estimation error, partial occlusion, and data loss. The CILgait dataset and MATLAB code are available at https://sites.google.com/site/seokeonchoi/gait-recognition.
Seokeon Choi, Jonghee Kim, Wonjun Kim 0001, Changick Kim
IEEE Trans. Inf. Forensics Secur.3
2018 Multiple player tracking in soccer videos: an adaptive multiscale sampling approach
Wonjun Kim 0001, Sung-Won Moon, Ji Won Lee, Do-Won Nam, Chanho Jung
Multim. Syst.1
2018 Background subtraction with variable illumination in outdoor scenes
Wonjun Kim 0001
Multim. Tools Appl.1
2018 Vein Enhancement Using a Dark Diffusion Prior
abstract
Vein recognition is emerging as a good alternative to the high-level biometric authentication due to its robustness against forgery. However, the poor quality of the vein image still poses a great difficulty to the extension of its usability. In this letter, we propose a novel method for vein enhancement based on a dark diffusion prior. The key idea of the proposed method is to reveal the characteristics of veins, which absorbs the near infrared well compared to background, by exploring multiple diffusion spaces. In contrast to previous approaches, the proposed method does not yield the blurring effect while efficiently suppressing unwanted noise. Experimental results on the benchmark database show that the proposed method is effective for vein recognition compared to other enhancement schemes introduced in the literature.
Wonjun Kim 0001
IEEE Signal Process. Lett.1
2017 Directional coherence-based spatiotemporal descriptor for object detection in static and dynamic scenes
Wonjun Kim 0001, Jae-Joon Han
Mach. Vis. Appl.1
2017 Fingerprint Liveness Detection Using Local Coherence Patterns
abstract
In this letter, we propose a novel image descriptor for fingerprint liveness detection using the local coherence of a given image. Based on the observation that materials employed for making fake fingerprints (e.g., silicone, wood glue, etc.) tend to yield the nonuniformity in the captured image due to the replica fabrication process, we focus on the difference of the dispersion in the image gradient field between live and fake fingerprints. More specifically, we propose to define the local patterns of the coherence along the dominant direction, the so-called local coherence patterns, as our features, which are fed into the linear support vector machine (SVM) classifier to determine whether a given fingerprint is fake or not. Experimental results on various datasets show that the proposed image descriptor is effective for fingerprint liveness detection compared to other approaches employed in the literature.
Wonjun Kim 0001
IEEE Signal Process. Lett.1
2016 Background Subtraction Using Illumination-Invariant Structural Complexity
abstract
In this letter, we propose a novel method for background subtraction in outdoor scenes. Inspired by the observation that the orthogonal decomposition onto a set of pixel intensities efficiently reveals illumination effects, we exploit a simple, yet powerful feature for describing the underlying structure of the local region in a given video, the so-called illumination-invariant structural complexity (IISC). In contrast to previous approaches still suffering from high-level false positives driven by varying illuminations in outdoor environments, our IISC feature has an ability to greatly discriminate structural changes by moving objects from those by illumination effects. We also provide the theoretical analysis to confirm that the proposed IISC feature is useful for modeling the background under diverse lighting conditions. Moreover, our framework does not require any preprocessing task. Experimental results on various datasets demonstrate that the proposed method is effective for video surveillance in a wide range of outdoor environments.
Wonjun Kim 0001, Youngsung Kim
IEEE Signal Process. Lett.1