Hanul Kim 0001

dblp:314/4344-1 · also Han-Ul Kim 0001 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-7450-6600ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Dual-stream feature aggregation and dual guided upsampling for efficient multi-exposure correction
Jong-Hyeon Baek, Hyo-Jun Lee, Hanul Kim 0001, Yeong Jun Koh
Knowl. Based Syst.3
2025 GRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object Tracking
abstract
Recently, 3D multi-object tracking (MOT) has widely adopted the standard tracking-by-detection paradigm, which solves the association problem between detections and tracks. Many tracking-by-detection approaches establish constrained relationships between detections and tracks using a distance threshold to reduce confusion during association. However, this approach does not effectively and comprehensively utilize the information regarding objects due to the constraints of the distance threshold. In this paper, we propose GRAE-3DMOT, Geometry Relation-Aware Encoder 3D Multi-Object Tracking, which contains a geometric relation-aware encoder to produce informative features for association. The geometric relation-aware encoder consists of three components: a spatial relation-aware encoder, a spatiotemporal relation-aware encoder, and a distance-aware feature fusion layer. The spatial relation-aware encoder effectively aggregates detection features by comprehensively exploiting as many detections as possible. The spatiotemporal relation-aware encoder provides spatiotemporal relation-aware features by combing spatial and temporal relation features, where the spatiotemporal relation-aware features are transformed into association scores for MOT. The distance-aware feature fusion layer is integrated into both encoders to enhance the relation features of physically proximate objects. Experimental results demonstrate that the proposed GRAE-3DMOT outperforms the state-of-the-art on the nuScenes. Our approach achieves 73.7% and 70.2% AMOTA on the nuScenes validation and test sets using CenterPoint detections. Code is available at https://github.com/altkddhfcjs/GRAE-3DMOT.
Hyunseop Kim, Hyo-Jun Lee, Yonguk Lee, Hanul Kim 0001, Yeong Jun Koh
CVPR5
2025 SOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View Projection
abstract
Existing view transformations in vision-centric 3D Semantic Scene Completion (SSC) inevitably experience erroneous feature duplication in the reconstructed voxel space due to occlusions, leading to a dilution of informative contexts. Furthermore, semantic classes exhibit high variability in their appearance in real-world driving scenarios. To address these issues, we introduce a novel 3D SSC method, called SOAP, including two key components: an occluded region-aware view projection and a scene-adaptive decoder. The occluded region-aware view projection effectively converts 2D image features into voxel space, refining the duplicated features of occluded regions using information gathered from previous observations. The sceneadaptive decoder guides query embeddings to learn diverse driving environments based on a comprehensive semantic repository. Extensive experiments validate that the proposed SOAP significantly outperforms existing methods for the vision-centric 3D SSC on automated driving datasets, SemanticKITTI and SSCBench. Code is available at https://github.com/gywns6287/SOAP.
Hyo-Jun Lee, Yeong Jun Koh, Hanul Kim 0001, Hyunseop Kim, Yonguk Lee
CVPR3
2025 Dual-Path Temporal Decoder for End-to-End Multi-Object Tracking
abstract
We present a novel end-to-end transformer-based framework for Multiple Object Tracking (MOT) that advances temporal modeling and identity preservation. Despite recent progress in transformer-based MOT, existing methods still struggle to maintain consistent object identities across frames, especially under occlusions, appearance changes, or detection failures. We propose a dual-path temporal decoder that explicitly separates appearance adaptation and identity preservation. The appearance-adaptive decoder dynamically updates query features using current frame information, while the identity-preserving decoder freezes query features and reuses historical sampling offsets to maintain long-term temporal consistency. To further enhance stability, we introduce a confidence-guided update suppression strategy that retains previously reliable features when predictions are unreliable. Extensive experiments on MOT benchmarks demonstrate that our approach achieves state-of-the-art performance across major tracking metrics, with significant gains in association accuracy and identity consistency. Our results demonstrate the importance of decoupling dynamic appearance modeling from static identity cues, and provide a scalable foundation for robust tracking in complex scenarios.
Hyunseop Kim, Juheon Jeong, Hanul Kim 0001, Yeong Jun Koh
NeurIPS3
2023 BAAM: Monocular 3D pose and shape reconstruction with bi-contextual attention module and attention-guided modeling
abstract
3D traffic scene comprises various 3D information about car objects, including their pose and shape. However, most recent studies pay relatively less attention to reconstructing detailed shapes. Furthermore, most of them treat each 3D object as an independent one, resulting in losses of relative context inter-objects and scene context reflecting road circumstances. A novel monocular 3D pose and shape reconstruction algorithm, based on bi-contextual attention and attention-guided modeling (BAAM), is proposed in this work. First, given 2D primitives, we reconstruct 3D object shape based on attention-guided modeling that considers the relevance between detected objects and vehicle shape priors. Next, we estimate 3D object pose through bi-contextual attention, which leverages relation-context inter objects and scene-context between an object and road environment. Finally, we propose a 3D nonmaximum suppression algorithm to eliminate spurious objects based on their Bird-Eye-View distance. Extensive experiments demonstrate that the proposed BAAM yields state-of-the-art performance on ApolloCar3D. Also, they show that the proposed BAAM can be plugged into any mature monocular 3D object detector on KITTI and significantly boost their performance. Code is available at https://github.com/gywns6287/BAAM.
Hyo-Jun Lee, Hanul Kim 0001, Su-Min Choi, Seong-Gyun Jeong, Yeong Jun Koh
CVPR2
2023 Luminance-aware Color Transform for Multiple Exposure Correction
abstract
Images captured with irregular exposures inevitably present unsatisfactory visual effects, such as distorted hue and color tone. However, most recent studies mainly focus on underexposure correction, which limits their applicability to real-world scenarios where exposure levels vary. Furthermore, some works to tackle multiple exposure rely on the encoder-decoder architecture, resulting in losses of details in input images during down-sampling and up-sampling processes. With this regard, a novel correction algorithm for multiple exposure, called luminance-aware color transform (LACT), is proposed in this study. First, we reason the relative exposure condition between images to obtain luminance features based on a luminance comparison module. Next, we encode the set of transformation functions from the luminance features, which enable complex color transformations for both overexposure and underexposure images. Finally, we project the transformed representation onto RGB color space to produce exposure correction results. Extensive experiments demonstrate that the proposed LACT yields new state-of-the-arts on two multiple exposure datasets. Code is available at https://github.com/whdgusdl48/LACT.
Jong-Hyeon Baek, Su-Min Choi, Hyo-Jun Lee, Hanul Kim 0001, Yeong Jun Koh
ICCV5
2023 Multiple transformation function estimation for image enhancement
abstract
Most deep learning-based image enhancement algorithms have been developed based on the image-to-image translation approach, in which enhancement processes are difficult to interpret. In this paper, we propose a novel interpretable image enhancement algorithm that estimates multiple transformation functions to describe complex color mapping. First, we develop a histogram-based multiple transformation function estimation network (HMTF-Net) to estimate multiple transformation functions by exploiting both the spatial and statistical information of the input images. Second, we estimate pixel-wise weight maps, which indicate the contribution of each transformation function at each pixel, based on the local structures of the input image and the transformed images obtained by each transformation function. Finally, we obtain the enhanced image as the weighted sum of the transformed images using the estimated weight maps. Extensive experiments confirm the effectiveness of the proposed approach and demonstrate that the proposed algorithm outperforms state-of-the-art image enhancement algorithms for different image enhancement tasks.
Vien Gia An, Minhee Cha, Thuy Thi Pham, Hanul Kim 0001, Chul Lee
J. Vis. Commun. Image Represent.5
2021 Representative Color Transform for Image Enhancement
abstract
Recently, the encoder-decoder and intensity transformation approaches lead to impressive progress in image enhancement. However, the encoder-decoder often loses details in input images during down-sampling and up-sampling processes. Also, the intensity transformation has a limited capacity to cover color transformation between low-quality and high-quality images. In this paper, we propose a novel approach, called representative color transform (RCT), to tackle these issues in existing methods. RCT determines different representative colors specialized in input images and estimates transformed colors for the representative colors. It then determines enhanced colors using these transformed colors based on the similarity between input and representative colors. Extensive experiments demonstrate that the proposed algorithm outperforms recent state-of-the-art algorithms on various image enhancement problems.
Hanul Kim 0001, Su-Min Choi, Chang-Su Kim 0001, Yeong Jun Koh
ICCV1
2021 Efficient Action Recognition via Dynamic Knowledge Propagation
abstract
Efficient action recognition has become crucial to extend the success of action recognition to many real-world applications. Contrary to most existing methods, which mainly focus on selecting salient frames to reduce the computation cost, we focus more on making the most of the selected frames. To this end, we employ two networks of different capabilities that operate in tandem to efficiently recognize actions. Given a video, the lighter network processes more frames while the heavier one only processes a few. In order to enable the effective interaction between the two, we propose dynamic knowledge propagation based on a cross-attention mechanism. This is the main component of our framework that is essentially a student-teacher architecture, but as the teacher model continues to interact with the student model during inference, we call it a dynamic student-teacher framework. Through extensive experiments, we demonstrate the effectiveness of each component of our framework. Our method outperforms competing state-of-the-art methods on two video datasets: ActivityNet-v1.3 and Mini-Kinetics.
Hanul Kim 0001, Mihir Jain, Juntae Lee, Sungrack Yun, Fatih Porikli
ICCV1
2020 PieNet: Personalized Image Enhancement Network
Hanul Kim 0001, Yeong Jun Koh, Chang-Su Kim 0001
ECCV (30)1
2020 Global and Local Enhancement Networks for Paired and Unpaired Image Enhancement
Hanul Kim 0001, Yeong Jun Koh, Chang-Su Kim 0001
ECCV (25)1
2019 Meta Learning for Unsupervised Clustering
Hanul Kim 0001, Yeong Jun Koh, Chang-Su Kim 0001
BMVC1
2018 Monocular Depth Estimation Using Whole Strip Masking and Reliability-Based Refinement
Minhyeok Heo, Jaehan Lee, Kyung-Rae Kim, Hanul Kim 0001, Chang-Su Kim 0001
ECCV (4)4
2018 Photographic composition classification and dominant geometric element detection for outdoor scenes
Juntae Lee, Hanul Kim 0001, Chul Lee, Chang-Su Kim 0001
J. Vis. Commun. Image Represent.2
2017 Semantic Line Detection and Its Applications
abstract
Semantic lines characterize the layout of an image. Despite their importance in image analysis and scene understanding, there is no reliable research for semantic line detection. In this paper, we propose a semantic line detector using a convolutional neural network with multi-task learning, by regarding the line detection as a combination of classification and regression tasks. We use convolution and max-pooling layers to obtain multi-scale feature maps for an input image. Then, we develop the line pooling layer to extract a feature vector for each candidate line from the feature maps. Next, we feed the feature vector into the parallel classification and regression layers. The classification layer decides whether the line candidate is semant ic or not. In case of a semantic line, the regression layer determines the offset for refining the line location. Experimental results show that the proposed detector extracts semantic lines accurately and reliably. Moreover, we demonstrate that the proposed detector can be used successfully in three applications: horizon estimation, composition enhancement, and image simplification.
Juntae Lee, Hanul Kim 0001, Chul Lee, Chang-Su Kim 0001
ICCV2
2017 Locator-Checker-Scaler Object Tracking Using Spatially Ordered and Weighted Patch Descriptor
abstract
In this paper, we propose a simple yet effective object descriptor and a novel tracking algorithm to track a target object accurately. For the object description, we divide the bounding box of a target object into multiple patches and describe them with color and gradient histograms. Then, we determine the foreground weight of each patch to alleviate the impacts of background information in the bounding box. To this end, we perform random walk with restart (RWR) simulation. We then concatenate the weighted patch descriptors to yield the spatially ordered and weighted patch (SOWP) descriptor. For the object tracking, we incorporate the proposed SOWP descriptor into a novel tracking algorithm, which has three components: locator, checker, and scaler (LCS). The locator and the scaler estimate the center location and the size of a target, respectively. The checker determines whether it is safe to adjust the target scale in a current frame. These three components cooperate with one another to achieve robust tracking. Experimental results demonstrate that the proposed LCS tracker achieves excellent performance on recent benchmarks.
Hanul Kim 0001, Chang-Su Kim 0001
IEEE Trans. Image Process.1
2016 CDT: Cooperative Detection and Tracking for Tracing Multiple Objects in Video Sequences
Hanul Kim 0001, Chang-Su Kim 0001
ECCV (6)1
2015 SOWP: Spatially Ordered and Weighted Patch Descriptor for Visual Tracking
abstract
A simple yet effective object descriptor for visual tracking is proposed in this paper. We first decompose the bounding box of a target object into multiple patches, which are described by color and gradient histograms. Then, we concatenate the features of the spatially ordered patches to represent the object appearance. Moreover, to alleviate the impacts of background information possibly included in the bounding box, we determine patch weights using random walk with restart (RWR) simulations. The patch weights represent the importance of each patch in the description of foreground information, and are used to construct an object descriptor, called spatially ordered and weighted patch (SOWP) descriptor. We incorporate the proposed SOWP descriptor into the structured output tracking framework. Experimental results demonstrate that the proposed algorithm yields significantly better performance than the state-of-the-art trackers on a recent benchmark dataset, and also excels in another recent benchmark dataset.
Hanul Kim 0001, Dae-Youn Lee, Jae-Young Sim, Chang-Su Kim 0001
ICCV1