Youngseok Kim 0001

dblp:32/40-1 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-9984-2416ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation
abstract
Recently, camera-radar fusion-based 3D object detection methods in bird's eye view (BEV) have gained attention due to the complementary characteristics and cost-effectiveness of these sensors. Previous approaches using forward projection struggle with sparse BEV feature generation, while those employing backward projection overlook depth ambiguity, leading to false positives. In this paper, to address the aforementioned limitations, we propose a novel camera-radar fusion-based 3D object detection and segmentation model named CRAB (Camera-Radar fusion for reducing depth Ambiguity in Backward projection-based view transformation), using a backward projection that leverages radar to mitigate depth ambiguity. During the view transformation, CRAB aggregates perspective view image context features into BEV queries. It improves depth distinction among queries along the same ray by combining the dense but unreliable depth distribution from images with the sparse yet precise depth information from radar occupancy. We further introduce spatial cross-attention with a feature map containing radar context information to enhance the comprehension of the 3D scene. When evaluated on the nuScenes open dataset, our proposed approach achieves a state-of-the-art performance among backward projection-based camera-radar fusion methods with 62.4% NDS and 54.0% mAP in 3D object detection.
In-Jae Lee, Sihwan Hwang, Youngseok Kim 0001, Wonjune Kim, Sanmin Kim, Dongsuk Kum
ICRA3
2024 LabelDistill: Label-Guided Cross-Modal Knowledge Distillation for Camera-Based 3D Object Detection
Sanmin Kim, Youngseok Kim 0001, Sihwan Hwang, Hyeonjun Jeong, Dongsuk Kum
ECCV (56)2
2023 CRAFT: Camera-Radar 3D Object Detection with Spatio-Contextual Fusion Transformer
abstract
Camera and radar sensors have significant advantages in cost, reliability, and maintenance compared to LiDAR. Existing fusion methods often fuse the outputs of single modalities at the result-level, called the late fusion strategy. This can benefit from using off-the-shelf single sensor detection algorithms, but late fusion cannot fully exploit the complementary properties of sensors, thus having limited performance despite the huge potential of camera-radar fusion. Here we propose a novel proposal-level early fusion approach that effectively exploits both spatial and contextual properties of camera and radar for 3D object detection. Our fusion framework first associates image proposal with radar points in the polar coordinate system to efficiently handle the discrepancy between the coordinate system and spatial properties. Using this as a first stage, following consecutive cross-attention based feature fusion layers adaptively exchange spatio-contextual information between camera and radar, leading to a robust and attentive fusion. Our camera-radar fusion approach achieves the state-of-the-art 41.1% mAP and 52.3% NDS on the nuScenes test set, which is 8.7 and 10.8 points higher than the camera-only baseline, as well as yielding competitive performance on the LiDAR method.
Youngseok Kim 0001, Sanmin Kim, Dongsuk Kum
AAAI1
2023 Predict to Detect: Prediction-guided 3D Object Detection using Sequential Images
abstract
Recent camera-based 3D object detection methods have introduced sequential frames to improve the detection performance hoping that multiple frames would mitigate the large depth estimation error. Despite improved detection performance, prior works rely on naive fusion methods (e.g., concatenation) or are limited to static scenes (e.g., temporal stereo), neglecting the importance of the motion cue of objects. These approaches do not fully exploit the potential of sequential images and show limited performance improvements. To address this limitation, we propose a novel 3D object detection model, P2D (Predict to Detect), that integrates a prediction scheme into a detection framework to explicitly extract and leverage motion features. P2D predicts object information in the current frame using solely past frames to learn temporal motion features. We then introduce a novel temporal feature aggregation method that attentively exploits Bird’s-Eye-View (BEV) features based on predicted object information, resulting in accurate 3D object detection. Experimental results demonstrate that P2D improves mAP and NDS by 3.0% and 3.7% compared to the sequential image-based baseline, proving that incorporating a prediction scheme can significantly improve detection accuracy.
Sanmin Kim, Youngseok Kim 0001, In-Jae Lee, Dongsuk Kum
ICCV2
2023 CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception
abstract
Autonomous driving requires an accurate and fast 3D perception system that includes 3D object detection, tracking, and segmentation. Although recent low-cost camera-based approaches have shown promising results, they are susceptible to poor illumination or bad weather conditions and have a large localization error. Hence, fusing camera with low-cost radar, which provides precise long-range measurement and operates reliably in all environments, is promising but has not yet been thoroughly investigated. In this paper, we propose Camera Radar Net (CRN), a novel camera-radar fusion framework that generates a semantically rich and spatially accurate bird’s-eye-view (BEV) feature map for various tasks. To overcome the lack of spatial information in an image, we transform perspective view image features to BEV with the help of sparse but accurate radar points. We further aggregate image and radar feature maps in BEV using multi-modal deformable attention designed to tackle the spatial misalignment between inputs. CRN with real-time setting operates at 20 FPS while achieving comparable performance to LiDAR detectors on nuScenes, and even outperforms at a far distance on 100m setting. Moreover, CRN with offline setting yields 62.4% NDS, 57.5% mAP on nuScenes test set and ranks first among all camera and camera-radar 3D object detectors.
Youngseok Kim 0001, Juyeb Shin, Sanmin Kim, In-Jae Lee, Dongsuk Kum
ICCV1
2023 Joint Semi-Supervised and Active Learning via 3D Consistency for 3D Object Detection
abstract
Autonomous driving powered by deep learning requires large-scale, high-quality training data from diverse driving environments to operate effectively worldwide. However, collecting and annotating such data is costly and time-consuming. To address this challenge, active learning methods have been explored to select the most informative data samples for training. Nevertheless, most existing methods focus on 2D tasks and do not fully exploit the value of unlabeled data. In this paper, we propose a semi-supervised active learning approach for 3D object detection tasks that leverages the potential of collected data and reduces annotation costs. Our method considers the 3D consistency of bounding box predictions in both semi-supervised and active learning processes, thereby improving the performance of point cloud-based 3D object detection models. Our framework specifically utilizes self-supervision to decrease bounding box uncertainties. Moreover, it selects objects that are either occluded or distant and still exhibit high uncertainty for annotation even after semi-supervised training has decreased their uncertainty. Experiments on the KITTI dataset demonstrate that our semi-supervised active learning approach selects objects with high measurement uncertainties and enhances the model's ability to detect occluded objects. Our approach improves the baseline by more than 60% (+17.12 mAP) when using only 1500 annotated frames.
Sihwan Hwang, Sanmin Kim, Youngseok Kim 0001, Dongsuk Kum
ICRA3
2023 Boosting Monocular 3D Object Detection With Object-Centric Auxiliary Depth Supervision
abstract
Recent advances in monocular 3D detection leverage a depth estimation network explicitly as an intermediate stage of the 3D detection network. Depth map approaches yield more accurate depth to objects than other methods thanks to the depth estimation network trained on a large-scale dataset. However, depth map approaches can be limited by the accuracy of the depth map, and sequentially using two separated networks for depth estimation and 3D detection significantly increases computation cost and inference time. In this work, we propose a method to boost the RGB image-based 3D detector by jointly training the detection network with a depth prediction loss analogous to the depth estimation task. In this way, our 3D detection network can be supervised by more depth supervision from raw LiDAR points, which does not require any human annotation cost, to estimate accurate depth without explicitly predicting the depth map. Our novel object-centric depth prediction loss focuses on depth around foreground objects, which is important for 3D object detection, to leverage pixel-wise depth supervision in an object-centric manner. Our depth regression model is further trained to predict the uncertainty of depth to represent the 3D confidence of objects. To effectively train the 3D detector with raw LiDAR points and to enable end-to-end training, we revisit the regression target of 3D objects and design a network architecture. Extensive experiments on KITTI and nuScenes benchmarks show that our method can significantly boost the monocular image-based 3D detector to outperform depth map approaches while maintaining the real-time inference speed.
Youngseok Kim 0001, Sanmin Kim, Sangmin Sim, Dongsuk Kum
IEEE Trans. Intell. Transp. Syst.1
2022 Sequential Image-based 3D Object Detection with Location Refinement
abstract
Recent advances in object detection tasks enable the detection network to predict 3D objects from a monocular image, but the performance of monocular 3D object detectors is inferior due to the depth information lost in the image. Most monocular 3D detectors do not utilize sequential information from multi-frame images, even though the object’s temporal motion is very informative for 3D object detection. In this paper, we propose a sequential image-based 3D object detection architecture that focuses on improving the localization performance of 3D detectors using temporal information for autonomous driving applications. To this end, the proposed network is trained with a pair of sequential images to predict 3D objects with their localization uncertainties on each image. Afterward, the object detected from sequential images is associated, and paired object features are fed to the sub-network to predict the depth displacement between frames. Finally, paired objects and their predicted depths and depth displacement are refined to minimize residuals between predictions and output the final 3D location of objects. The experimental results on challenging the nuScenes dataset demonstrate that our method improves the performance of the 3D detector by reducing the localization error.
Sangmin Sim, Youngseok Kim 0001, Dongsuk Kum
ICPR2
2020 Low-Level Sensor Fusion for 3D Vehicle Detection Using Radar Range-Azimuth Heatmap and Monocular Image
Jinhyeong Kim, Youngseok Kim 0001, Dongsuk Kum
ACCV (3)2
2020 GRIF Net: Gated Region of Interest Fusion Network for Robust 3D Object Detection from Radar Point Cloud and Monocular Image
abstract
Robust and accurate scene representation is essential for advanced driver assistance systems (ADAS) such as automated driving. The radar and camera are two widely used sensors for commercial vehicles due to their low-cost, high-reliability, and low-maintenance. Despite their strengths, radar and camera have very limited performance when used individually. In this paper, we propose a low-level sensor fusion 3D object detector that combines two Region of Interest (RoI) from radar and camera feature maps by a Gated RoI Fusion (GRIF) to perform robust vehicle detection. To take advantage of sensors and utilize a sparse radar point cloud, we design a GRIF that employs the explicit gating mechanism to adaptively select the appropriate data when one of the sensors is abnormal. Our experimental evaluations on nuScenes show that our fusion method GRIF not only has significant performance improvement over single radar and image method but achieves comparable performance to the LiDAR detection method. We also observe that the proposed GRIF achieve higher recall than mean or concatenation fusion operation when points are sparse.
Youngseok Kim 0001, Dongsuk Kum
IROS1
2019 Deep Learning based Vehicle Position and Orientation Estimation via Inverse Perspective Mapping Image
abstract
In this paper, we present a method for estimating a position, size, and orientation using a single monocular image. The proposed method makes use of an inverse perspective mapping to effectively estimate the distance from the image. The proposed method consists of two stages: 1) cancel the pitch and roll motion of the camera using inertial measurement unit and project the corrected front view image onto the bird's eye view using inverse perspective mapping. 2) detect the position, size, and orientation of the vehicle using a convolutional neural network. The camera motion cancellation process makes vanishing point to be located at the same point regardless of the ego vehicle attitude change. Through this process, the projected bird's eye view image can be parallel and linear to the x-y plane of the vehicle coordinate system. The convolutional neural network predicts not only the position and size but also the orientation of the vehicle for the 3D localization. The predicted oriented bounding box from the bird's eye view image is converted in the meter unit by the inverse projection matrix. The proposed method was evaluated on the KITTI raw dataset on the metric of the root mean square error, mean average percentage error, and average precision. Despite the conceptually simple architecture, the proposed method achieves promising performance compared to other image based approaches. The video demonstration is available online [1].
Youngseok Kim 0001, Dongsuk Kum
IV1