EDBT 2026 Demo / reviewers in the wild / expert
Dongsuk Kum
dblp:123/5638
· DBLP profile ↗
40ranked-venue papers
0as first author
22since 2021 · last 2026
0000-0002-2590-4845ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 since 2021Systems, architecture and hardware · 7 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing Annotation Costs for Autonomous Driving: Scalable Perception via Multi-Level Active Domain Adaptation
Sihwan Hwang, Sanmin Kim, Hyeonjun Jeong, Dongsuk Kum |
IV | 4 |
| 2025 | CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View TransformationabstractRecently, camera-radar fusion-based 3D object detection methods in bird's eye view (BEV) have gained attention due to the complementary characteristics and cost-effectiveness of these sensors. Previous approaches using forward projection struggle with sparse BEV feature generation, while those employing backward projection overlook depth ambiguity, leading to false positives. In this paper, to address the aforementioned limitations, we propose a novel camera-radar fusion-based 3D object detection and segmentation model named CRAB (Camera-Radar fusion for reducing depth Ambiguity in Backward projection-based view transformation), using a backward projection that leverages radar to mitigate depth ambiguity. During the view transformation, CRAB aggregates perspective view image context features into BEV queries. It improves depth distinction among queries along the same ray by combining the dense but unreliable depth distribution from images with the sparse yet precise depth information from radar occupancy. We further introduce spatial cross-attention with a feature map containing radar context information to enhance the comprehension of the 3D scene. When evaluated on the nuScenes open dataset, our proposed approach achieves a state-of-the-art performance among backward projection-based camera-radar fusion methods with 62.4% NDS and 54.0% mAP in 3D object detection. In-Jae Lee, Sihwan Hwang, Youngseok Kim 0001, Wonjune Kim, Sanmin Kim, Dongsuk Kum |
ICRA | 6 |
| 2025 | REOcc: Camera-Radar Fusion with Radar Feature Enrichment for 3D Occupancy PredictionabstractVision-based 3D occupancy prediction has made significant advancements, but its reliance on cameras alone struggles in challenging environments. This limitation has driven the adoption of sensor fusion, among which camera-radar fusion stands out as a promising solution due to their complementary strengths. However, the sparsity and noise of the radar data limits its effectiveness, leading to suboptimal fusion performance. In this paper, we propose REOcc, a novel camera-radar fusion network designed to enrich radar feature representations for 3D occupancy prediction. Our approach introduces two main components, a Radar Densifier and a Radar Amplifier, which refine radar features by integrating spatial and contextual information, effectively enhancing spatial density and quality. Extensive experiments on the Occ3D-nuScenes benchmark demonstrate that REOcc achieves significant performance gains over the camera-only baseline model, particularly in dynamic object classes. These results underscore REOcc’s capability to mitigate the sparsity and noise of the radar data. Consequently, radar complements camera data more effectively, unlocking the full potential of camera-radar fusion for robust and reliable 3D occupancy prediction. Chaehee Song, Sanmin Kim, Hyeonjun Jeong, Juyeb Shin, Joonhee Lim, Dongsuk Kum |
IROS | 6 |
| 2025 | InstaGraM: Instance-Level Graph Modeling for Vectorized HD Map LearningabstractFor scalable autonomous driving, a robust map-based localization system, independent of GPS, is fundamental. To achieve such map-based localization, online high-definition (HD) map construction plays a significant role in accurate estimation of the pose. Although recent advancements in online HD map construction have predominantly investigated on vectorized representation due to its effectiveness, they suffer from computational cost and fixed parametric model, which limit scalability. To alleviate these limitations, we propose a novel HD map learning framework that leverages graph modeling. This framework is designed to learn the construction of diverse geometric shapes, thereby enhancing the scalability of HD map construction. Our approach involves representing the map elements as an instance-level graph by decomposing them into vertices and edges to facilitate accurate and efficient end-to-end vectorized HD map learning. Furthermore, we introduce an association strategy using a Graph Neural Network to efficiently handle the complex geometry of various map elements, while maintaining scalability. Comprehensive experiments on public open dataset show that our proposed network outperforms state-of-the-art model by$1.6$mAP. We further showcase the superior scalability of our approach compared to state-of-the-art methods, achieving a$4.8$mAP improvement in long range configuration. Our code is available at https://github.com/juyebshin/InstaGraM. Juyeb Shin, Hyeonjun Jeong, François Rameau, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | RadarDistill: Boosting Radar-Based Object Detection Performance via Knowledge Distillation from LiDAR FeaturesabstractThe inherent noisy and sparse characteristics of radar data pose challenges in finding effective representations for 3D object detection. In this paper, we propose RadarD-istill, a novel knowledge distillation (KD) method, which can improve the representation of radar data by leveraging LiDAR data. RadarDistill successfully transfers de-sirable characteristics of LiDAR features into radar features using three key components: Cross-Modality Alignment (CMA), Activation-based Feature Distillation (AFD), and Proposal-based Feature Distillation (PFD). CMA en-hances the density of radar features by employing multiple layers of dilation operations, effectively addressing the challenge of inefficient knowledge transfer from LiDAR to radar. AFD selectively transfers knowledge based on regions of the LiDAR features, with a specific focus on areas where activation intensity exceeds a predefined thresh-old. P FD similarly guides the radar network to selectively mimic features from the LiDAR network within the object proposals. Our comparative analyses conducted on the nuScenes datasets demonstrate that RadarDistill achieves state-of-the-art (SOTA) performance for radar-only object detection task, recording 20.5% in mAP and 43.7% in NDS. Also, RadarDistill significantly improves the performance of the camera-radar fusion model. Geonho Bang, Kwangjin Choi, Jisong Kim, Dongsuk Kum |
CVPR | 4 |
| 2024 | Continual Learning for Motion Prediction Model via Meta-Representation Learning and Optimal Memory Buffer Retention StrategyabstractEmbodied AI, such as autonomous vehicles, suffers from insufficient, long-tailed data because it must be obtained from the physical world. In fact, data must be continuously obtained in a series of small batches, and the model must also be continuously trained to achieve generalizability and scalability by improving the biased data distribution. This paper addresses the training cost and catastrophic forgetting problems when continuously updating models to adapt to incoming small batches from various environments for real-world motion prediction in autonomous driving. To this end, we propose a novel continual motion prediction (CMP) learning framework based on sparse meta-representation learning and an optimal memory buffer retention strategy. In meta-representation learning, a model explicitly learns a sparse representation of each driving environment, from road geometry to vehicle states, by training to reduce catastrophic forgetting based on an augmented modulation network with sparsity regularization. Also, in the adaptation phase, We develop an Optimal Memory Buffer Retention strategy that smartly preserves diverse samples by focusing on representation similarity. This approach handles the nuanced task distribution shifts characteristic of motion prediction datasets, ensuring our model stays responsive to evolving input variations without requiring extensive resources. The experiment results demonstrate that the proposed method shows superior adaptation performance to the conventional continual learning approach, which is developed using a synthetic dataset for the continual learning problem. Daejun Kang, Dongsuk Kum, Sanmin Kim |
CVPR | 2 |
| 2024 | Beyond the Data Imbalance: Employing the Heterogeneous Datasets for Vehicle Maneuver Prediction
Hyeong-Seok Jeon, Sanmin Kim, Abi Rahman Syamil, Dongsuk Kum |
ECCV (56) | 5 |
| 2024 | LabelDistill: Label-Guided Cross-Modal Knowledge Distillation for Camera-Based 3D Object Detection
Sanmin Kim, Youngseok Kim 0001, Sihwan Hwang, Hyeonjun Jeong, Dongsuk Kum |
ECCV (56) | 5 |
| 2024 | RCM-Fusion: Radar-Camera Multi-Level Fusion for 3D Object DetectionabstractWhile LiDAR sensors have been successfully applied to 3D object detection, the affordability of radar and camera sensors has led to a growing interest in fusing radars and cameras for 3D object detection. However, previous radar-camera fusion models could not fully utilize the potential of radar information. In this paper, we propose Radar-Camera Multi-level fusion (RCM-Fusion), which attempts to fuse both modalities at feature and instance levels. For feature-level fusion, we propose a Radar Guided BEV Encoder which transforms camera features into precise BEV representations using the guidance of radar Bird’s-Eye-View (BEV) features and combines the radar and camera BEV features. For instance-level fusion, we propose a Radar Grid Point Refinement module that reduces localization error by accounting for the characteristics of the radar point clouds. The experiments on the public nuScenes dataset demonstrate that our proposed RCM-Fusion achieves state-of-the-art performances among single frame-based radar-camera fusion methods in the nuScenes 3D object detection benchmark. The code will be made publicly available. Jisong Kim, Minjae Seong, Geonho Bang, Dongsuk Kum |
ICRA | 4 |
| 2024 | Learning Terminal State of the Trajectory Planner: Application for Collision Scenarios of Autonomous VehiclesabstractCollision Avoidance/Mitigation System (CAMS) for autonomous vehicles is a crucial technology that ensures the safety and reliability of autonomous driving systems. Conventional collision avoidance approaches struggle in complex and various scenarios by avoiding collisions based on rules for specific collision scenarios. This has led to learning-based methods using neural networks for adaptive collision avoidance. However, the approaches directly outputting control inputs through neural networks have drawbacks in interpretability and stability. To address these limitations, we propose a trajectory planning method for CAMS that combines deep reinforcement learning (DRL) and quintic polynomial (QP) trajectory planning. The proposed method determines the terminal state and confidence of the trajectory using DRL and plans a QP trajectory based on them. By utilizing the terminal state and confidence of the trajectory rather than direct control inputs as the output of the neural network, it generates a more realistic and continuous path. Moreover, this approach considers collision avoidance and mitigation in an integrated manner through the reward function of RL. Our experimental results demonstrate that the proposed method not only improves interpretability and stability compared to existing learning-based methods but also upholds performance in complex and various collision scenarios. Joonhee Lim, Kibeom Lee, Jangho Shin, Dongsuk Kum |
ICRA | 4 |
| 2024 | An Analytical Approach to the Predictive Energy Management of Connected HEVs: What Information Do We Need to Guarantee Global Optimality?abstractThe predictive energy management (PEM) of hybrid electric vehicles (HEVs) is a challenging problem of trajectory optimization involving future information. Most previous studies have presented intuitive methods for selecting parameters that represent future information. However, such intuitive methods lack theoretical analysis and do not guarantee global optimality. This study adopts the novel perspective that the PEM problem can be analytically solved so as to ensure near-optimal efficiency. The key idea is to reformulate the trajectory optimization problem as a quadratic programming (QP) problem based on optimal control principles. The equivalence between the original and QP problems is theoretically derived. The proposed PEM strategy is then implemented by solving the QP problem at every sampling time to obtain the optimal control input. Whereas the original problem requires information on the entire future trajectory, which is difficult to predict accurately, the QP problem requires only lumped parameters, specifically, the energy demands and time durations of each segment of the future path. These lumped future driving parameters can potentially be predicted using information obtained through vehicle connectivity. Simulation results obtained under a real-world driving scenario show that the proposed PEM strategy provides a control result in real time (within 2 ms) that is very close to the globally optimal solution both qualitatively and quantitatively, with a loss of optimality of only 0.14%. Kyunghwan Choi, Geunyoung Park, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | CRAFT: Camera-Radar 3D Object Detection with Spatio-Contextual Fusion TransformerabstractCamera and radar sensors have significant advantages in cost, reliability, and maintenance compared to LiDAR. Existing fusion methods often fuse the outputs of single modalities at the result-level, called the late fusion strategy. This can benefit from using off-the-shelf single sensor detection algorithms, but late fusion cannot fully exploit the complementary properties of sensors, thus having limited performance despite the huge potential of camera-radar fusion. Here we propose a novel proposal-level early fusion approach that effectively exploits both spatial and contextual properties of camera and radar for 3D object detection. Our fusion framework first associates image proposal with radar points in the polar coordinate system to efficiently handle the discrepancy between the coordinate system and spatial properties. Using this as a first stage, following consecutive cross-attention based feature fusion layers adaptively exchange spatio-contextual information between camera and radar, leading to a robust and attentive fusion. Our camera-radar fusion approach achieves the state-of-the-art 41.1% mAP and 52.3% NDS on the nuScenes test set, which is 8.7 and 10.8 points higher than the camera-only baseline, as well as yielding competitive performance on the LiDAR method. Youngseok Kim 0001, Sanmin Kim, Dongsuk Kum |
AAAI | 4 |
| 2023 | Predict to Detect: Prediction-guided 3D Object Detection using Sequential ImagesabstractRecent camera-based 3D object detection methods have introduced sequential frames to improve the detection performance hoping that multiple frames would mitigate the large depth estimation error. Despite improved detection performance, prior works rely on naive fusion methods (e.g., concatenation) or are limited to static scenes (e.g., temporal stereo), neglecting the importance of the motion cue of objects. These approaches do not fully exploit the potential of sequential images and show limited performance improvements. To address this limitation, we propose a novel 3D object detection model, P2D (Predict to Detect), that integrates a prediction scheme into a detection framework to explicitly extract and leverage motion features. P2D predicts object information in the current frame using solely past frames to learn temporal motion features. We then introduce a novel temporal feature aggregation method that attentively exploits Bird’s-Eye-View (BEV) features based on predicted object information, resulting in accurate 3D object detection. Experimental results demonstrate that P2D improves mAP and NDS by 3.0% and 3.7% compared to the sequential image-based baseline, proving that incorporating a prediction scheme can significantly improve detection accuracy. Sanmin Kim, Youngseok Kim 0001, In-Jae Lee, Dongsuk Kum |
ICCV | 4 |
| 2023 | CRN: Camera Radar Net for Accurate, Robust, Efficient 3D PerceptionabstractAutonomous driving requires an accurate and fast 3D perception system that includes 3D object detection, tracking, and segmentation. Although recent low-cost camera-based approaches have shown promising results, they are susceptible to poor illumination or bad weather conditions and have a large localization error. Hence, fusing camera with low-cost radar, which provides precise long-range measurement and operates reliably in all environments, is promising but has not yet been thoroughly investigated. In this paper, we propose Camera Radar Net (CRN), a novel camera-radar fusion framework that generates a semantically rich and spatially accurate bird’s-eye-view (BEV) feature map for various tasks. To overcome the lack of spatial information in an image, we transform perspective view image features to BEV with the help of sparse but accurate radar points. We further aggregate image and radar feature maps in BEV using multi-modal deformable attention designed to tackle the spatial misalignment between inputs. CRN with real-time setting operates at 20 FPS while achieving comparable performance to LiDAR detectors on nuScenes, and even outperforms at a far distance on 100m setting. Moreover, CRN with offline setting yields 62.4% NDS, 57.5% mAP on nuScenes test set and ranks first among all camera and camera-radar 3D object detectors. Youngseok Kim 0001, Juyeb Shin, Sanmin Kim, In-Jae Lee, Dongsuk Kum |
ICCV | 6 |
| 2023 | Joint Semi-Supervised and Active Learning via 3D Consistency for 3D Object DetectionabstractAutonomous driving powered by deep learning requires large-scale, high-quality training data from diverse driving environments to operate effectively worldwide. However, collecting and annotating such data is costly and time-consuming. To address this challenge, active learning methods have been explored to select the most informative data samples for training. Nevertheless, most existing methods focus on 2D tasks and do not fully exploit the value of unlabeled data. In this paper, we propose a semi-supervised active learning approach for 3D object detection tasks that leverages the potential of collected data and reduces annotation costs. Our method considers the 3D consistency of bounding box predictions in both semi-supervised and active learning processes, thereby improving the performance of point cloud-based 3D object detection models. Our framework specifically utilizes self-supervision to decrease bounding box uncertainties. Moreover, it selects objects that are either occluded or distant and still exhibit high uncertainty for annotation even after semi-supervised training has decreased their uncertainty. Experiments on the KITTI dataset demonstrate that our semi-supervised active learning approach selects objects with high measurement uncertainties and enhances the model's ability to detect occluded objects. Our approach improves the baseline by more than 60% (+17.12 mAP) when using only 1500 annotated frames. Sihwan Hwang, Sanmin Kim, Youngseok Kim 0001, Dongsuk Kum |
ICRA | 4 |
| 2023 | Are Reactions to Ego Vehicles Predictable Without Data?: A Semi-Supervised ApproachabstractTo make intelligent decisions in an autonomous vehicle, the system must predict the future reactions of surrounding vehicles for any given action plan of the ego vehicle. However, learning reactive trajectories is challenging due to scant action-reaction pair data. That is, building a dataset with multiple action-reaction pairs for an identical scene history is impossible in reality. Here, we propose a semi-supervised learning framework with auxiliary structures to handle this problem. The proposed training framework has two modules: Action Reconstructor and Identifier modules with corresponding loss functions referred to as the Reconstruction Loss and Association Loss. In addition to the conventional supervised approach pertaining to readily available data, the Action Reconstructor module is employed to learn the dependencies on the ego vehicle in an unsupervised manner. Furthermore, reaction trajectory data corresponding to the augmented future trajectories of the ego vehicle are not available, meaning that the model must be trained in an unsupervised manner as well. The main idea of the proposed unsupervised learning method is to find the identity feature vector from both history and future trajectories and associate these features for each vehicle. This idea is realized by introducing the Identifier network and the Association Loss, which are used only during the training process. Interestingly, experimental results show that plausible reaction can be predicted for the augmented future trajectory of the ego vehicle, which indicates that the network can generalize the interactive behavior of vehicles from a partially labelled dataset. Hyeong-Seok Jeon, Sanmin Kim, Kibeom Lee, Daejun Kang, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Boosting Monocular 3D Object Detection With Object-Centric Auxiliary Depth SupervisionabstractRecent advances in monocular 3D detection leverage a depth estimation network explicitly as an intermediate stage of the 3D detection network. Depth map approaches yield more accurate depth to objects than other methods thanks to the depth estimation network trained on a large-scale dataset. However, depth map approaches can be limited by the accuracy of the depth map, and sequentially using two separated networks for depth estimation and 3D detection significantly increases computation cost and inference time. In this work, we propose a method to boost the RGB image-based 3D detector by jointly training the detection network with a depth prediction loss analogous to the depth estimation task. In this way, our 3D detection network can be supervised by more depth supervision from raw LiDAR points, which does not require any human annotation cost, to estimate accurate depth without explicitly predicting the depth map. Our novel object-centric depth prediction loss focuses on depth around foreground objects, which is important for 3D object detection, to leverage pixel-wise depth supervision in an object-centric manner. Our depth regression model is further trained to predict the uncertainty of depth to represent the 3D confidence of objects. To effectively train the 3D detector with raw LiDAR points and to enable end-to-end training, we revisit the regression target of 3D objects and design a network architecture. Extensive experiments on KITTI and nuScenes benchmarks show that our method can significantly boost the monocular image-based 3D detector to outperform depth map approaches while maintaining the real-time inference speed. Youngseok Kim 0001, Sanmin Kim, Sangmin Sim, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Joint 3D Object Detection and Tracking Using Spatio-Temporal Representation of Camera Image and LiDAR Point CloudsabstractIn this paper, we propose a new joint object detection and tracking (JoDT) framework for 3D object detection and tracking based on camera and LiDAR sensors. The proposed method, referred to as 3D DetecTrack, enables the detector and tracker to cooperate to generate a spatio-temporal representation of the camera and LiDAR data, with which 3D object detection and tracking are then performed. The detector constructs the spatio-temporal features via the weighted temporal aggregation of the spatial features obtained by the camera and LiDAR fusion. Then, the detector reconfigures the initial detection results using information from the tracklets maintained up to the previous time step. Based on the spatio-temporal features generated by the detector, the tracker associates the detected objects with previously tracked objects using a graph neural network (GNN). We devise a fully-connected GNN facilitated by a combination of rule-based edge pruning and attention-based edge gating, which exploits both spatial and temporal object contexts to improve tracking performance. The experiments conducted on both KITTI and nuScenes benchmarks demonstrate that the proposed 3D DetecTrack achieves significant improvements in both detection and tracking performances over baseline methods and achieves state-of-the-art performance among existing methods through collaboration between the detector and tracker. Junho Koh, Jaekyum Kim, Jin Hyeok Yoo, Yecheol Kim, Dongsuk Kum |
AAAI | 5 |
| 2022 | Sequential Image-based 3D Object Detection with Location RefinementabstractRecent advances in object detection tasks enable the detection network to predict 3D objects from a monocular image, but the performance of monocular 3D object detectors is inferior due to the depth information lost in the image. Most monocular 3D detectors do not utilize sequential information from multi-frame images, even though the object’s temporal motion is very informative for 3D object detection. In this paper, we propose a sequential image-based 3D object detection architecture that focuses on improving the localization performance of 3D detectors using temporal information for autonomous driving applications. To this end, the proposed network is trained with a pair of sequential images to predict 3D objects with their localization uncertainties on each image. Afterward, the object detected from sequential images is associated, and paired object features are fed to the sub-network to predict the depth displacement between frames. Finally, paired objects and their predicted depths and depth displacement are refined to minimize residuals between predictions and output the final 3D location of objects. The experimental results on challenging the nuScenes dataset demonstrate that our method improves the performance of the 3D detector by reducing the localization error. Sangmin Sim, Youngseok Kim 0001, Dongsuk Kum |
ICPR | 3 |
| 2022 | Autonomous Vehicle Cut-In Algorithm for Lane-Merging Scenarios via Policy-Based Reinforcement Learning Nested Within Finite-State MachineabstractLane-merging scenarios pose highly challenging problems for autonomous vehicles due to conflicts of interest between the human-driven and cutting-in autonomous vehicles. Such conflicts become severe when traffic increases, and cut-in algorithms suffer from a steep trade-off between safety and cut-in performance. In this study, a reinforcement learning (RL)-based cut-in policy network nested within a finite state machine (FSM)—which is a high-level decision maker, is proposed to achieve high cut-in performance without sacrificing safety. This FSM-RL hybrid approach is proposed to obtain 1) a strategic and adjustable algorithm, 2) optimal safety and cut-in performance, and 3) robust and consistent performance. In the high-level decision making algorithm, the FSM provides a framework for four cut-in phases (ready for safe gap selection, gap approach, negotiation, and lane-change execution) and handles the transitions between these phases by calculating the collision risks associated with target vehicles. For the lane-change phase, a policy-based deep-RL approach with a soft actor–critic network is employed to get optimal cut-in performance. The results of simulations show that the proposed FSM-RL cut-in algorithm consistently achieves a high cut-in success rate without sacrificing safety. In particular, as the traffic increases, the cut-in success rate and safety are significantly improved over existing optimized rule-based cut-in algorithms and end-to-end RL algorithm. Seulbin Hwang, Kibeom Lee, Hyeong-Seok Jeon, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Efficient Design Space Exploration of Multi-Mode, Two-Planetary-Gear, Power-Split Hybrid Electric Powertrains via Virtual LeversabstractThe recent industry trend of incorporating multiple planetary gears (PG) and clutches in power-split hybrid electric vehicles has enhanced their potential fuel economy and acceleration performance. However, this increase in performance potential comes at the cost of increased design and control complexity. In this paper, we propose a highly efficient design methodology that finds the optimal multi-mode, two-PG powertrain by extending our recently developed virtual lever, a modeling tool that eliminates the redundancy in the physical design space and hence minimizes the computational load associated with optimizing overmultipledifferent types of PGs. First, every operating mode of every powertrain architecture is modeled using the virtual lever, whose parameters comprise the virtual design space. Second, the fuel economy and acceleration time are evaluated for every design in this continuous, virtual design space. Then, the powertrain architectures are compared by their respective Pareto frontiers in the fuel economy — acceleration time plane. Furthermore, a design space conversion method is used to group the evaluated designs by their respective feasible physical realizations to compare the 16 possible physical realizations of the generic Volt 2nd, which is found to be the best powertrain architecture. The proposed method uncovered several different physical realizations of the generic Volt 2nd that outperforms General Motors’ existing Chevy Volt 2nd powertrain. Chonghyuk Song, Jaeho Hwang, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | LaPred: Lane-Aware Prediction of Multi-Modal Future Trajectories of Dynamic AgentsabstractIn this paper, we address the problem of predicting the future motion of a dynamic agent (called a target agent) given its current and past states as well as the information on its environment. It is paramount to develop a prediction model that can exploit the contextual information in both static and dynamic environments surrounding the target agent and generate diverse trajectory samples that are meaningful in a traffic context. We propose a novel prediction model, referred to as the lane-aware prediction (LaPred) network, which uses the instance-level lane entities extracted from a semantic map to predict the multi-modal future trajectories. For each lane candidate found in the neighborhood of the target agent, LaPred extracts the joint features relating the lane and the trajectories of the neighboring agents. Then, the features for all lane candidates are fused with the attention weights learned through a self-supervised learning task that identifies the lane candidate likely to be followed by the target agent. Using the instance-level lane information, LaPred can produce the trajectories compliant with the surroundings better than 2D raster image-based methods and generate the diverse future trajectories given multiple lane candidates. The experiments conducted on the public nuScenes dataset and Argo- verse dataset demonstrate that the proposed LaPred method significantly outperforms the existing prediction models, achieving state-of-the-art performance in the benchmarks. Byeoungdo Kim, Seong Hyeon Park, Seokhwan Lee, Elbek Khoshimjonov, Dongsuk Kum, Jeong Soo Kim |
CVPR | 5 |
| 2020 | Low-Level Sensor Fusion for 3D Vehicle Detection Using Radar Range-Azimuth Heatmap and Monocular Image
Jinhyeong Kim, Youngseok Kim 0001, Dongsuk Kum |
ACCV (3) | 3 |
| 2020 | ScarfNet: Multi-scale Features with Deeply Fused and Redistributed Semantics for Enhanced Object DetectionabstractConvolutional neural networks (CNNs) have led us to achieve significant progress in object detection research. To detect objects of various sizes, object detectors often exploit the hierarchy of the multiscale feature maps called feature pyramids, which are readily obtained by the CNN architecture. However, the performance of these object detectors is limited because the bottom-level feature maps, which experience fewer convolutional layers, lack the semantic information needed to capture the characteristics of the small objects. To address such problems, various methods have been proposed to increase the depth for the bottom-level features used for object detection. While most approaches are based on the generation of additional features through the top-down pathway with lateral connections, our approach directly fuses multi-scale feature maps using bidirectional long short-term memory (biLSTM) in an effort to leverage the gating functions and parameter-sharing in generating deeply fused semantics. The resulting semantic information is redistributed to the individual pyramidal feature at each scale through the channel-wise attention model. We integrate our semantic combining and attentive redistribution feature network (ScarfNet) with the baseline object detectors, i.e., Faster R-CNN, single-shot multibox detector (SSD), and RetinaNet. Experimental results show that our method offers a significant performance gain over the baseline detectors and outperforms the competing multiscale fusion methods in the PASCAL VOC and COCO detection benchmarks. Jin Hyeok Yoo, Dongsuk Kum |
ICPR | 2 |
| 2020 | SCALE-Net: Scalable Vehicle Trajectory Prediction Network under Random Number of Interacting Vehicles via Edge-enhanced Graph Convolutional Neural NetworkabstractPredicting the future trajectory of surrounding vehicles in a randomly varying traffic level is one of the most challenging problems in developing an autonomous vehicle. Since there is no pre-defined number of interacting vehicles participated in, the prediction network has to be scalable with respect to the number of vehicles in order to guarantee consistent performance in terms of both accuracy and computational load. In this paper, the first fully scalable trajectory prediction network, SCALE-Net, is proposed that can ensure both high prediction performance while keeping the computational load low regardless of the number of surrounding vehicles. The SCALE-Net employs the Edge-enhanced Graph Convolutional Neural Network (EGCN) for the inter-vehicular interaction embedding network. Since the proposed EGCN is inherently scalable with respect to the graph node (an agent in this study), the model can be operated independently from the total number of vehicles considered. We evaluated the scalability of the SCALE-Net on the publically available NGSIM datasets by comparing variations on computation time and prediction accuracy per single driving scene with respect to the varying vehicle number. The experimental test shows that both computation time and prediction performance of the SCALE-Net consistently outperform those of previous models regardless of the level of traffic complexities. Hyeong-Seok Jeon, Dongsuk Kum |
IROS | 3 |
| 2020 | GRIF Net: Gated Region of Interest Fusion Network for Robust 3D Object Detection from Radar Point Cloud and Monocular ImageabstractRobust and accurate scene representation is essential for advanced driver assistance systems (ADAS) such as automated driving. The radar and camera are two widely used sensors for commercial vehicles due to their low-cost, high-reliability, and low-maintenance. Despite their strengths, radar and camera have very limited performance when used individually. In this paper, we propose a low-level sensor fusion 3D object detector that combines two Region of Interest (RoI) from radar and camera feature maps by a Gated RoI Fusion (GRIF) to perform robust vehicle detection. To take advantage of sensors and utilize a sparse radar point cloud, we design a GRIF that employs the explicit gating mechanism to adaptively select the appropriate data when one of the sensors is abnormal. Our experimental evaluations on nuScenes show that our fusion method GRIF not only has significant performance improvement over single radar and image method but achieves comparable performance to the LiDAR detection method. We also observe that the proposed GRIF achieve higher recall than mean or concatenation fusion operation when points are sparse. Youngseok Kim 0001, Dongsuk Kum |
IROS | 3 |
| 2020 | Systematic Design of Input- and Output-Split Hybrid Electric Vehicles With a Speed Reduction/Multiplication Gear Using Simplified-Lever ModelabstractPower-split hybrid electric vehicles (HEV) employ a speed reduction gear (SRG) mainly to enhance acceleration performance. However, the full potentials of an additional gear on fuel economy and acceleration performance have not yet been thoroughly investigated due to a vast design space: 432 configurations with three design variables (i.e. two planetary gear ratios and a final drive gear ratio). In this paper, a systematic speed reduction gear design methodology is proposed to analyze the impact of a speed reduction (multiplication) gear on the performance of input- and output-split HEVs and select an optimal configuration. First, the physical and virtual design spaces of SRG (SMG) are defined and the relationship between two design spaces are identified. Second, performance metrics are evaluated within the virtual design space using the proposed simplified lever. The proposed approach completely eliminates the redundancy present in the physical design spaces, and thus, the impact of SRG (SMG) on the performance of input- and output-split configurations can be analyzed with the minimum computational burden. Lastly, the selected gear ratio is converted back to the physical planetary gear connections. The results confirm that SRG (SMG) can potentially either improve or deteriorate the acceleration performance and the fuel economy of split HEVs. Therefore, the full potential of speed reduction or multiplication gear should be thoroughly analyzed when designing split hybrid electric vehicles. Jingeon Kang, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Deep Learning based Vehicle Position and Orientation Estimation via Inverse Perspective Mapping ImageabstractIn this paper, we present a method for estimating a position, size, and orientation using a single monocular image. The proposed method makes use of an inverse perspective mapping to effectively estimate the distance from the image. The proposed method consists of two stages: 1) cancel the pitch and roll motion of the camera using inertial measurement unit and project the corrected front view image onto the bird's eye view using inverse perspective mapping. 2) detect the position, size, and orientation of the vehicle using a convolutional neural network. The camera motion cancellation process makes vanishing point to be located at the same point regardless of the ego vehicle attitude change. Through this process, the projected bird's eye view image can be parallel and linear to the x-y plane of the vehicle coordinate system. The convolutional neural network predicts not only the position and size but also the orientation of the vehicle for the 3D localization. The predicted oriented bounding box from the bird's eye view image is converted in the meter unit by the inverse projection matrix. The proposed method was evaluated on the KITTI raw dataset on the metric of the root mean square error, mean average percentage error, and average precision. Despite the conceptually simple architecture, the proposed method achieves promising performance compared to other image based approaches. The video demonstration is available online [1]. Youngseok Kim 0001, Dongsuk Kum |
IV | 2 |
| 2019 | Robust multi-lane detection and tracking using adaptive threshold and lane classification
Yeongho Son, Elijah S. Lee, Dongsuk Kum |
Mach. Vis. Appl. | 3 |
| 2019 | Charging Automation for Electric Vehicles: Is a Smaller Battery Good for the Wireless Charging Electric Vehicles?abstractDynamic wireless charging (DWC) is an emerging technology that enables the batteries of electric vehicles (EVs) to charge automatically while the vehicles are in motion. The DWC-EV system addresses the challenges inherent in battery technology, such as the short driving range, long recharging time, and high price. Compared with conventional plug-in EVs, the DWC-EV can charge a battery more frequently because it can be done while the EV is in motion from the charging infrastructure installed on the road. In this paper, we analyze how this frequent-charging characteristic of DWC-EV can affect the battery lifetime in the DWC-EV. We first introduce a mathematical model to evaluate the economic cost of the DWC-EV for a given battery size. A battery degradation model is incorporated to account for the quantitative relationship between the installation of the charging infrastructure and battery life extension. We then use the model to analyze how the economic cost varies with the size of the battery. Our preliminary findings provide insight into the relationship between DWC from the charging infrastructure and the battery's lifetime. Seungmin Jeong, Young Jae Jang, Dongsuk Kum, Min-Seok Lee |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2019 | Guest Editorial Introduction to the Special Issue on Intelligent Transportation Systems Empowered by AI TechnologiesabstractThere has been an increasing level of demand for faster, safer and greener transportation systems with higher levels of capacity and convenience, though the implementation of transportation systems overall is often restricted by geographical limitations, presenting a challenge to scientists and engineers in the field. However, we have been witnessing the evolution of the transportation systems over the last few decades, and at present we are facing a new era of intelligent transportation systems (ITS) empowered by artificial intelligence (AI) technologies. There have been classification, deep learning, and reinforcement learning techniques, to name a few, which collectively have enabled almost all technical elements of the ITS. For example, autonomous vehicle technologies are now mature enough to introduce self-driving cars, taxis, buses, and trucks on the roads and streets; traffic signals are controlled by AI-based systems for far more enhanced traffic efficiency; and machine learning based on big data is improving the operational performance of transportation systems to the next level of safety, efficiency, and sustainability. Seung-Hyun Kong, Hai Le Vu 0001, Juan-Carlos Cano, Dongsuk Kum, Brendan Tran Morris |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2019 | Synthesis of Robust Lane Keeping Systems: Impact of Controller and Design Parameters on System PerformanceabstractThe lane keeping system (LKS), a promising driver assistance system, is essential for autonomous vehicles. In realworld road conditions, it can be quite challenging because LKS must stay within the lane without causing passenger discomfort while both disturbances (e.g., road curvature, wind gusts, and hydroplaning) and model uncertainties in parameters [e.g., vehicle mass, center of gravity (CG), and tire cornering stiffness] are present. In this paper, the performance limits and tradeoffs between three performance criteria (lane tracking, stability robustness, and passenger comfort) are first investigated by exploring the entire design space of three prominent controllers, i.e., proportional-integral-derivative, linear-quadratic- Gaussian, and H-infinity (H∞). Then, a sensitivity study on the vehicle parameters is conducted in order to investigate the impact of the parameters on the three performance metrics. Based on the aforementioned studies, this paper concludes that a robust controller can provide the maximum performance limit with respect to the lane tracking and stability robustness, when properly designed. However, it is observed that the robust controller is still sensitive to a few design and model parameters, such as look-ahead distance and CG. Therefore, the sensitivity study suggests that for vehicles with excessive mass and CG changes, such as SUVs and trucks, the adaptation of controller and look-ahead distance may be necessary to maximize both tracking performance and passenger comfort over a wide range of vehicle speeds. Kibeom Lee, Shengbo Eben Li, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Predictive Cruise Control Using Radial Basis Function Network-Based Vehicle Motion Prediction and Chance Constrained Model Predictive ControlabstractPredicting future motions of surrounding vehicles and driver's intentions are essential to avoid future potential risks. The predicting future motions, however, is very challenging because the future cannot be deterministically known a priori and there are infinitely many possible future trajectories. Prediction becomes far more challenging when trying to foresee distant future. This paper proposes a probabilistic motion prediction algorithm that can accurately compute the likelihood of multiple target lanes and trajectories of surrounding vehicles by using the artificial neural network; more specifically radial base function network (RBFN). The RBFN prediction algorithm estimates the likelihood of each lane being the driver's target lane in categorical distributions and the corresponding future trajectories in parallel. In order to demonstrate the effectiveness of the proposed prediction algorithm, it is applied for the predictive cruise control problem. Chance-constrained model predictive control (CCMPC) is utilized because the chance constraints in CCMPC can handle collision uncertainties associated with future uncertainties from the proposed prediction algorithm. The RBFN-based CCMPC simulation is conducted for several risky cut-in scenarios and compared with the state-of-the-art Interactive Multiple Model (IMM)-based prediction algorithm. The simulation results show that the RBFN-based CCMPC achieves higher collision avoidance success rate than that of the IMM-based CCMPC while using smaller actuator inputs and providing higher passenger comforts. Furthermore, the RBFN-based CCMPC showed high robustness to false braking during near lane-change (lane-keeping) scenarios. Seungje Yoon, Hyeong-Seok Jeon, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Traffic Scene Prediction via Deep Learning: Introduction of Multi-Channel Occupancy Grid Map as a Scene RepresentationabstractWhen predicting future motions of surrounding vehicles for autonomous vehicles, the inter-vehicular interaction must be considered in order to predict future risks and to make safe and intelligent decisions. This becomes critical when it comes to conflicting driving situation such as lane merge, tollgate area, and unsignalized intersections. Previously developed future prediction algorithms show limited performance when handling interactions and conflicts between vehicles because they focused on predicting individual vehicle motion and/or interaction between a single pair of vehicles rather than the entire traffic scene. In this paper, a scene representation method, namely multi-channel Occupancy Grid Map (OGM), is proposed to describe the entire traffic scene, which is then utilized for the deep learning architecture that predicts the future traffic scene or OGM. Multi-channel OGM represents entire traffic scene as a manner of image-like structure from bird’s eye view composed with dynamic layer and static layer depicting the occupancy of the dynamic and static objects. By using this 2D traffic scene representation, future prediction can be modeled as a video processing problem, where future time-serial image sequence need to be predicted. In order to predict future traffic scenes based on past traffic scenes, a deep learning architecture is proposed using Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) Networks. With the proposed deep learning architecture, future prediction accuracy in highly conflicting traffic situation is guaranteed up to 90 percent with 3 seconds of prediction horizon. A video ofthe traffic scene prediction results is available online [1]. Hyeong-Seok Jeon, Dongsuk Kum, Woo-Yeol Jeong |
Intelligent Vehicles Symposium | 2 |
| 2018 | Collision Risk Assessment Algorithm via Lane-Based Probabilistic Motion Prediction of Surrounding VehiclesabstractIn order to ensure reliable autonomous driving, the system must be able to detect future dangers in sufficient time to avoid or mitigate collisions. In this paper, we propose a collision risk assessment algorithm that can quantitatively assess collision risks for a set of local path candidates via the lane-based probabilistic motion prediction of surrounding vehicles. First, we compute target lane probabilities, which represent how likely a driver is to drive or move toward each lane, based on lateral position and lateral velocity in curvilinear coordinates. And then, collision risks are computed by incorporating both model probability distribution of lanes and a time-to-collision between a pair of predicted trajectories. Finally, collision risks are plotted on a trajectory plane that represents each set of the tangential acceleration and the final lateral offset of local path candidates. This collision risk map provides intuitive risk measures, and can also be utilized to determine a control strategy for a collision avoidance maneuver. Validation of the model is conducted by comparing the model probabilities with the maneuver probabilities derived from the next generation simulation database. Furthermore, the effectiveness of the proposed algorithm is verified in two driving scenarios, preceding vehicle braking and cut-in, on a curved highway with multiple vehicles. Jae-Hwan Kim, Dongsuk Kum |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Feature-based lateral position estimation of surrounding vehicles using stereo visionabstractDriver behavior Prediction has become an important topic in the recent development of Advanced Driver Assistance Systems (ADAS). To predict future behavior and potential risks associated with surrounding vehicles, their lateral position information is required. However, existing computer vision algorithms tend to either focus on longitudinal measurements or provide lateral position information with limited performance for limited scenes (i.e no viewpoint change and occlusion). In this paper, feature-based lateral position estimation algorithm is proposed using stereo vision and provides lateral position regardless of viewpoint change and occlusion by extracting a pixel-wise feature. In the preprocessing step, v-disparity from stereo depth map is calculated and used for ground detection. Then, vehicle candidates are created based on image thresholding and filtering, removing the ground portion from the camera image. These generated candidates are verified as vehicles by using deep convolutional neural network. In order to track and estimate the lateral position of the detected vehicles, speeded up robust feature (SURF) points are matched in consecutive image frames, and the feature point is projected onto the ground; defined as the grounded feature point. Finally, inverse perspective mapping (IPM) is applied on the original image to estimate the lateral position of the grounded feature point. The proposed algorithm successfully detects a feature point of neighboring vehicle and estimates its lateral position by tracking the grounded feature point. For testing the algorithm, the datasets in a highway and an urban setting are used and provide zero mean error and 0.25m standard deviation error in lateral position estimation. Elijah S. Lee, Dongsuk Kum |
Intelligent Vehicles Symposium | 2 |
| 2016 | The multilayer perceptron approach to lateral motion prediction of surrounding vehicles for autonomous vehiclesabstractFor safe and reliable autonomous driving systems, prediction of surrounding vehicles' future behavior and potential risks are critical. The state-of-the-art prediction algorithms tend to show limited performance on long-term predictions due to their deterministic nature. In this paper, a probabilistic lateral motion prediction algorithm is proposed based on multilayer perceptron (MLP) approach. The MLP model consists of two parts; target lane and trajectory models. In order to develop an intuitive and accurate prediction algorithm, a lane-based trajectory prediction model is introduced based on the fact that vehicles drive within a lane except for during lane changes. More specifically, a set of three representative trajectories with different levels of lane-change positions are generated for each target lane, and real-world traffic data is categorized by each trajectory for MLP training. These target lane and trajectory models enable the stochastic MLP modeling and training. The proposed MLP model outputs probabilities of how likely a vehicle will follow each trajectory and each lane for a given input of vehicle position history including current position. For training the MLP model, Next Generation Simulation traffic data are used. Simulation results show that the proposed algorithm detects lane-changes one to one and a half second earlier than existing methods and three seconds before lane crossing with about ninety percentages accuracy. Seungje Yoon, Dongsuk Kum |
Intelligent Vehicles Symposium | 2 |
| 2016 | Efficient and accurate computation of model predictive control using pseudospectral discretization
Shengbo Eben Li, Shaobing Xu, Dongsuk Kum |
Neurocomputing | 3 |
| 2015 | Threat prediction algorithm based on local path candidates and surrounding vehicle trajectory predictions for automated driving vehiclesabstractAmong others, a reliable threat prediction algorithm is one of the key enabling technologies for the commercialization of the automated driving systems and other driver assistance systems. Previous algorithms that use Time-to-Collision (TTC) as a measure of threat tend to assume constant state and constant input; e.g. constant yaw rate and constant acceleration. Although the predictability of these algorithms is acceptable within a one second time horizon, it becomes invalid for predictions over one second because yaw rate and acceleration are highly unlikely to be constant. Therefore, in this paper, we propose a threat prediction algorithm that can accurately predict TTC over a longer time horizon based on future trajectory predictions of a surrounding vehicle. First, a comprehensive set of local path candidates is generated along the curvilinear coordinates using a quintic (5thorder) polynomial with respect to the arc-length corresponding to the different lateral offsets. Trajectory prediction of a surrounding vehicle is accomplished by introducing target lane detection, which is estimated according to the amount of difference between the current motion and the centerline of the driving lane. Based on these future vehicle trajectories, TTC is computed by comparing the entrance and exit time of two vehicles into and out of the conflict area where the occupied spaces of two vehicles overlap. Finally, in order to provide threat assessment results, the inverse TTC values obtained above are plotted on a 2-dimensional trajectory plane where each set of the tangential acceleration and the initial yaw acceleration values represents each local path candidate. Thus, these threat assessment results can be directly utilized to determine a driving strategy of autonomous vehicles. Jae-Hwan Kim, Dongsuk Kum |
Intelligent Vehicles Symposium | 2 |
| 2015 | Synthesis of multiple model switching controllers using H∞ theory for systems with large uncertainties
Feng Gao 0007, Shengbo Eben Li, Dongsuk Kum, Hui Zhang 0019 |
Neurocomputing | 3 |