VLDB 2026 Research / reviewers in the wild / expert
Sanmin Kim
dblp:283/0107
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-4042-6570ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing Annotation Costs for Autonomous Driving: Scalable Perception via Multi-Level Active Domain Adaptation
Sihwan Hwang, Sanmin Kim, Hyeonjun Jeong, Dongsuk Kum |
IV | 2 |
| 2025 | CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View TransformationabstractRecently, camera-radar fusion-based 3D object detection methods in bird's eye view (BEV) have gained attention due to the complementary characteristics and cost-effectiveness of these sensors. Previous approaches using forward projection struggle with sparse BEV feature generation, while those employing backward projection overlook depth ambiguity, leading to false positives. In this paper, to address the aforementioned limitations, we propose a novel camera-radar fusion-based 3D object detection and segmentation model named CRAB (Camera-Radar fusion for reducing depth Ambiguity in Backward projection-based view transformation), using a backward projection that leverages radar to mitigate depth ambiguity. During the view transformation, CRAB aggregates perspective view image context features into BEV queries. It improves depth distinction among queries along the same ray by combining the dense but unreliable depth distribution from images with the sparse yet precise depth information from radar occupancy. We further introduce spatial cross-attention with a feature map containing radar context information to enhance the comprehension of the 3D scene. When evaluated on the nuScenes open dataset, our proposed approach achieves a state-of-the-art performance among backward projection-based camera-radar fusion methods with 62.4% NDS and 54.0% mAP in 3D object detection. In-Jae Lee, Sihwan Hwang, Youngseok Kim 0001, Wonjune Kim, Sanmin Kim, Dongsuk Kum |
ICRA | 5 |
| 2025 | REOcc: Camera-Radar Fusion with Radar Feature Enrichment for 3D Occupancy PredictionabstractVision-based 3D occupancy prediction has made significant advancements, but its reliance on cameras alone struggles in challenging environments. This limitation has driven the adoption of sensor fusion, among which camera-radar fusion stands out as a promising solution due to their complementary strengths. However, the sparsity and noise of the radar data limits its effectiveness, leading to suboptimal fusion performance. In this paper, we propose REOcc, a novel camera-radar fusion network designed to enrich radar feature representations for 3D occupancy prediction. Our approach introduces two main components, a Radar Densifier and a Radar Amplifier, which refine radar features by integrating spatial and contextual information, effectively enhancing spatial density and quality. Extensive experiments on the Occ3D-nuScenes benchmark demonstrate that REOcc achieves significant performance gains over the camera-only baseline model, particularly in dynamic object classes. These results underscore REOcc’s capability to mitigate the sparsity and noise of the radar data. Consequently, radar complements camera data more effectively, unlocking the full potential of camera-radar fusion for robust and reliable 3D occupancy prediction. Chaehee Song, Sanmin Kim, Hyeonjun Jeong, Juyeb Shin, Joonhee Lim, Dongsuk Kum |
IROS | 2 |
| 2024 | Continual Learning for Motion Prediction Model via Meta-Representation Learning and Optimal Memory Buffer Retention StrategyabstractEmbodied AI, such as autonomous vehicles, suffers from insufficient, long-tailed data because it must be obtained from the physical world. In fact, data must be continuously obtained in a series of small batches, and the model must also be continuously trained to achieve generalizability and scalability by improving the biased data distribution. This paper addresses the training cost and catastrophic forgetting problems when continuously updating models to adapt to incoming small batches from various environments for real-world motion prediction in autonomous driving. To this end, we propose a novel continual motion prediction (CMP) learning framework based on sparse meta-representation learning and an optimal memory buffer retention strategy. In meta-representation learning, a model explicitly learns a sparse representation of each driving environment, from road geometry to vehicle states, by training to reduce catastrophic forgetting based on an augmented modulation network with sparsity regularization. Also, in the adaptation phase, We develop an Optimal Memory Buffer Retention strategy that smartly preserves diverse samples by focusing on representation similarity. This approach handles the nuanced task distribution shifts characteristic of motion prediction datasets, ensuring our model stays responsive to evolving input variations without requiring extensive resources. The experiment results demonstrate that the proposed method shows superior adaptation performance to the conventional continual learning approach, which is developed using a synthetic dataset for the continual learning problem. Daejun Kang, Dongsuk Kum, Sanmin Kim |
CVPR | 3 |
| 2024 | Beyond the Data Imbalance: Employing the Heterogeneous Datasets for Vehicle Maneuver Prediction
Hyeong-Seok Jeon, Sanmin Kim, Abi Rahman Syamil, Dongsuk Kum |
ECCV (56) | 2 |
| 2024 | LabelDistill: Label-Guided Cross-Modal Knowledge Distillation for Camera-Based 3D Object Detection
Sanmin Kim, Youngseok Kim 0001, Sihwan Hwang, Hyeonjun Jeong, Dongsuk Kum |
ECCV (56) | 1 |
| 2023 | CRAFT: Camera-Radar 3D Object Detection with Spatio-Contextual Fusion TransformerabstractCamera and radar sensors have significant advantages in cost, reliability, and maintenance compared to LiDAR. Existing fusion methods often fuse the outputs of single modalities at the result-level, called the late fusion strategy. This can benefit from using off-the-shelf single sensor detection algorithms, but late fusion cannot fully exploit the complementary properties of sensors, thus having limited performance despite the huge potential of camera-radar fusion. Here we propose a novel proposal-level early fusion approach that effectively exploits both spatial and contextual properties of camera and radar for 3D object detection. Our fusion framework first associates image proposal with radar points in the polar coordinate system to efficiently handle the discrepancy between the coordinate system and spatial properties. Using this as a first stage, following consecutive cross-attention based feature fusion layers adaptively exchange spatio-contextual information between camera and radar, leading to a robust and attentive fusion. Our camera-radar fusion approach achieves the state-of-the-art 41.1% mAP and 52.3% NDS on the nuScenes test set, which is 8.7 and 10.8 points higher than the camera-only baseline, as well as yielding competitive performance on the LiDAR method. Youngseok Kim 0001, Sanmin Kim, Dongsuk Kum |
AAAI | 2 |
| 2023 | Predict to Detect: Prediction-guided 3D Object Detection using Sequential ImagesabstractRecent camera-based 3D object detection methods have introduced sequential frames to improve the detection performance hoping that multiple frames would mitigate the large depth estimation error. Despite improved detection performance, prior works rely on naive fusion methods (e.g., concatenation) or are limited to static scenes (e.g., temporal stereo), neglecting the importance of the motion cue of objects. These approaches do not fully exploit the potential of sequential images and show limited performance improvements. To address this limitation, we propose a novel 3D object detection model, P2D (Predict to Detect), that integrates a prediction scheme into a detection framework to explicitly extract and leverage motion features. P2D predicts object information in the current frame using solely past frames to learn temporal motion features. We then introduce a novel temporal feature aggregation method that attentively exploits Bird’s-Eye-View (BEV) features based on predicted object information, resulting in accurate 3D object detection. Experimental results demonstrate that P2D improves mAP and NDS by 3.0% and 3.7% compared to the sequential image-based baseline, proving that incorporating a prediction scheme can significantly improve detection accuracy. Sanmin Kim, Youngseok Kim 0001, In-Jae Lee, Dongsuk Kum |
ICCV | 1 |
| 2023 | CRN: Camera Radar Net for Accurate, Robust, Efficient 3D PerceptionabstractAutonomous driving requires an accurate and fast 3D perception system that includes 3D object detection, tracking, and segmentation. Although recent low-cost camera-based approaches have shown promising results, they are susceptible to poor illumination or bad weather conditions and have a large localization error. Hence, fusing camera with low-cost radar, which provides precise long-range measurement and operates reliably in all environments, is promising but has not yet been thoroughly investigated. In this paper, we propose Camera Radar Net (CRN), a novel camera-radar fusion framework that generates a semantically rich and spatially accurate bird’s-eye-view (BEV) feature map for various tasks. To overcome the lack of spatial information in an image, we transform perspective view image features to BEV with the help of sparse but accurate radar points. We further aggregate image and radar feature maps in BEV using multi-modal deformable attention designed to tackle the spatial misalignment between inputs. CRN with real-time setting operates at 20 FPS while achieving comparable performance to LiDAR detectors on nuScenes, and even outperforms at a far distance on 100m setting. Moreover, CRN with offline setting yields 62.4% NDS, 57.5% mAP on nuScenes test set and ranks first among all camera and camera-radar 3D object detectors. Youngseok Kim 0001, Juyeb Shin, Sanmin Kim, In-Jae Lee, Dongsuk Kum |
ICCV | 3 |
| 2023 | Joint Semi-Supervised and Active Learning via 3D Consistency for 3D Object DetectionabstractAutonomous driving powered by deep learning requires large-scale, high-quality training data from diverse driving environments to operate effectively worldwide. However, collecting and annotating such data is costly and time-consuming. To address this challenge, active learning methods have been explored to select the most informative data samples for training. Nevertheless, most existing methods focus on 2D tasks and do not fully exploit the value of unlabeled data. In this paper, we propose a semi-supervised active learning approach for 3D object detection tasks that leverages the potential of collected data and reduces annotation costs. Our method considers the 3D consistency of bounding box predictions in both semi-supervised and active learning processes, thereby improving the performance of point cloud-based 3D object detection models. Our framework specifically utilizes self-supervision to decrease bounding box uncertainties. Moreover, it selects objects that are either occluded or distant and still exhibit high uncertainty for annotation even after semi-supervised training has decreased their uncertainty. Experiments on the KITTI dataset demonstrate that our semi-supervised active learning approach selects objects with high measurement uncertainties and enhances the model's ability to detect occluded objects. Our approach improves the baseline by more than 60% (+17.12 mAP) when using only 1500 annotated frames. Sihwan Hwang, Sanmin Kim, Youngseok Kim 0001, Dongsuk Kum |
ICRA | 2 |
| 2023 | Are Reactions to Ego Vehicles Predictable Without Data?: A Semi-Supervised ApproachabstractTo make intelligent decisions in an autonomous vehicle, the system must predict the future reactions of surrounding vehicles for any given action plan of the ego vehicle. However, learning reactive trajectories is challenging due to scant action-reaction pair data. That is, building a dataset with multiple action-reaction pairs for an identical scene history is impossible in reality. Here, we propose a semi-supervised learning framework with auxiliary structures to handle this problem. The proposed training framework has two modules: Action Reconstructor and Identifier modules with corresponding loss functions referred to as the Reconstruction Loss and Association Loss. In addition to the conventional supervised approach pertaining to readily available data, the Action Reconstructor module is employed to learn the dependencies on the ego vehicle in an unsupervised manner. Furthermore, reaction trajectory data corresponding to the augmented future trajectories of the ego vehicle are not available, meaning that the model must be trained in an unsupervised manner as well. The main idea of the proposed unsupervised learning method is to find the identity feature vector from both history and future trajectories and associate these features for each vehicle. This idea is realized by introducing the Identifier network and the Association Loss, which are used only during the training process. Interestingly, experimental results show that plausible reaction can be predicted for the augmented future trajectory of the ego vehicle, which indicates that the network can generalize the interactive behavior of vehicles from a partially labelled dataset. Hyeong-Seok Jeon, Sanmin Kim, Kibeom Lee, Daejun Kang, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Boosting Monocular 3D Object Detection With Object-Centric Auxiliary Depth SupervisionabstractRecent advances in monocular 3D detection leverage a depth estimation network explicitly as an intermediate stage of the 3D detection network. Depth map approaches yield more accurate depth to objects than other methods thanks to the depth estimation network trained on a large-scale dataset. However, depth map approaches can be limited by the accuracy of the depth map, and sequentially using two separated networks for depth estimation and 3D detection significantly increases computation cost and inference time. In this work, we propose a method to boost the RGB image-based 3D detector by jointly training the detection network with a depth prediction loss analogous to the depth estimation task. In this way, our 3D detection network can be supervised by more depth supervision from raw LiDAR points, which does not require any human annotation cost, to estimate accurate depth without explicitly predicting the depth map. Our novel object-centric depth prediction loss focuses on depth around foreground objects, which is important for 3D object detection, to leverage pixel-wise depth supervision in an object-centric manner. Our depth regression model is further trained to predict the uncertainty of depth to represent the 3D confidence of objects. To effectively train the 3D detector with raw LiDAR points and to enable end-to-end training, we revisit the regression target of 3D objects and design a network architecture. Extensive experiments on KITTI and nuScenes benchmarks show that our method can significantly boost the monocular image-based 3D detector to outperform depth map approaches while maintaining the real-time inference speed. Youngseok Kim 0001, Sanmin Kim, Sangmin Sim, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 2 |