EDBT 2026 Demo / reviewers in the wild / expert
In-Jae Lee
dblp:349/4614
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0004-3654-8614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 54% Robot navigation and mapping · 23% Autonomous driving · 13% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d object detection |
2.2 | 3 | 2025 | CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025 CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023 Predict to Detect: Prediction-guided 3D Object Detection using Sequential Images · ICCV 2023 |
Robotics › Robot navigation and mapping › sensor fusion
radar-camera fusion |
1.5 | 2 | 2025 | CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025 CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023 |
Robotics › Autonomous driving › perception
3d perception |
0.7 | 1 | 2023 | CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023 |
Computer vision › 3D vision › 3d object detection
bird's-eye-view detection |
0.7 | 1 | 2023 | CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023 |
Computer vision › 3D vision
depth estimation |
0.3 | 1 | 2025 | CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025 |
Computer vision › 3D vision
view transformation |
0.3 | 1 | 2025 | CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025 |
Computer vision › 3D vision
multimodal perception |
0.2 | 1 | 2023 | CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023 |
Robotics › Autonomous driving
perception |
0.2 | 1 | 2023 | Predict to Detect: Prediction-guided 3D Object Detection using Sequential Images · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
spatial cross-attention · 0.9radar occupancy · 0.9backward projection · 0.9temporal feature aggregation · 0.7perspective-to-BEV transformation · 0.7multimodal fusion · 0.7deformable attention · 0.7bird's-eye-view representation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View TransformationabstractRecently, camera-radar fusion-based 3D object detection methods in bird's eye view (BEV) have gained attention due to the complementary characteristics and cost-effectiveness of these sensors. Previous approaches using forward projection struggle with sparse BEV feature generation, while those employing backward projection overlook depth ambiguity, leading to false positives. In this paper, to address the aforementioned limitations, we propose a novel camera-radar fusion-based 3D object detection and segmentation model named CRAB (Camera-Radar fusion for reducing depth Ambiguity in Backward projection-based view transformation), using a backward projection that leverages radar to mitigate depth ambiguity. During the view transformation, CRAB aggregates perspective view image context features into BEV queries. It improves depth distinction among queries along the same ray by combining the dense but unreliable depth distribution from images with the sparse yet precise depth information from radar occupancy. We further introduce spatial cross-attention with a feature map containing radar context information to enhance the comprehension of the 3D scene. When evaluated on the nuScenes open dataset, our proposed approach achieves a state-of-the-art performance among backward projection-based camera-radar fusion methods with 62.4% NDS and 54.0% mAP in 3D object detection. In-Jae Lee, Sihwan Hwang, Youngseok Kim 0001, Wonjune Kim, Sanmin Kim, Dongsuk Kum |
ICRA | 1 |
| 2025 | OpenBox: Annotate Any Bounding Boxes in 3DabstractUnsupervised and open-vocabulary 3D object detection has recently gained attention, particularly in autonomous driving, where reducing annotation costs and recognizing unseen objects are critical for both safety and scalability. However, most existing approaches uniformly annotate 3D bounding boxes, ignore objects’ physical states, and require multiple self-training iterations for annotation refinement, resulting in suboptimal quality and substantial computational overhead. To address these challenges, we propose OpenBox, a two-stage automatic annotation pipeline that leverages a 2D vision foundation model. In the first stage, OpenBox associates instance-level cues from 2D images processed by a vision foundation model with the corresponding 3D point clouds via context-aware refinement. In the
second stage, it categorizes instances by rigidity and motion state, then generates adaptive bounding boxes with class-specific size statistics. As a result, OpenBox produces high-quality 3D bounding box annotations without requiring self-training.
Experiments on the Waymo Open Dataset (WOD), the Lyft Level 5 Perception dataset, and the nuScenes dataset demonstrate improved accuracy and efficiency over baselines. In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park |
NeurIPS | 1 |
| 2023 | Predict to Detect: Prediction-guided 3D Object Detection using Sequential ImagesabstractRecent camera-based 3D object detection methods have introduced sequential frames to improve the detection performance hoping that multiple frames would mitigate the large depth estimation error. Despite improved detection performance, prior works rely on naive fusion methods (e.g., concatenation) or are limited to static scenes (e.g., temporal stereo), neglecting the importance of the motion cue of objects. These approaches do not fully exploit the potential of sequential images and show limited performance improvements. To address this limitation, we propose a novel 3D object detection model, P2D (Predict to Detect), that integrates a prediction scheme into a detection framework to explicitly extract and leverage motion features. P2D predicts object information in the current frame using solely past frames to learn temporal motion features. We then introduce a novel temporal feature aggregation method that attentively exploits Bird’s-Eye-View (BEV) features based on predicted object information, resulting in accurate 3D object detection. Experimental results demonstrate that P2D improves mAP and NDS by 3.0% and 3.7% compared to the sequential image-based baseline, proving that incorporating a prediction scheme can significantly improve detection accuracy. Sanmin Kim, Youngseok Kim 0001, In-Jae Lee, Dongsuk Kum |
ICCV | 3 |
| 2023 | CRN: Camera Radar Net for Accurate, Robust, Efficient 3D PerceptionabstractAutonomous driving requires an accurate and fast 3D perception system that includes 3D object detection, tracking, and segmentation. Although recent low-cost camera-based approaches have shown promising results, they are susceptible to poor illumination or bad weather conditions and have a large localization error. Hence, fusing camera with low-cost radar, which provides precise long-range measurement and operates reliably in all environments, is promising but has not yet been thoroughly investigated. In this paper, we propose Camera Radar Net (CRN), a novel camera-radar fusion framework that generates a semantically rich and spatially accurate bird’s-eye-view (BEV) feature map for various tasks. To overcome the lack of spatial information in an image, we transform perspective view image features to BEV with the help of sparse but accurate radar points. We further aggregate image and radar feature maps in BEV using multi-modal deformable attention designed to tackle the spatial misalignment between inputs. CRN with real-time setting operates at 20 FPS while achieving comparable performance to LiDAR detectors on nuScenes, and even outperforms at a far distance on 100m setting. Moreover, CRN with offline setting yields 62.4% NDS, 57.5% mAP on nuScenes test set and ranks first among all camera and camera-radar 3D object detectors. Youngseok Kim 0001, Juyeb Shin, Sanmin Kim, In-Jae Lee, Dongsuk Kum |
ICCV | 4 |