Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

In-Jae Lee

dblp:349/4614 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0004-3654-8614ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 54% Robot navigation and mapping · 23% Autonomous driving · 13%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
2.232025
CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025
CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023
Predict to Detect: Prediction-guided 3D Object Detection using Sequential Images · ICCV 2023
Robotics › Robot navigation and mapping › sensor fusion
radar-camera fusion
1.522025
CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025
CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023
Robotics › Autonomous driving › perception
3d perception
0.712023
CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023
Computer vision › 3D vision › 3d object detection
bird's-eye-view detection
0.712023
CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023
Computer vision › 3D vision
depth estimation
0.312025
CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025
Computer vision › 3D vision
view transformation
0.312025
CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation · ICRA 2025
Computer vision › 3D vision
multimodal perception
0.212023
CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception · ICCV 2023
Robotics › Autonomous driving
perception
0.212023
Predict to Detect: Prediction-guided 3D Object Detection using Sequential Images · ICCV 2023

Methods — techniques the papers use, named apart from their topics

spatial cross-attention · 0.9radar occupancy · 0.9backward projection · 0.9temporal feature aggregation · 0.7perspective-to-BEV transformation · 0.7multimodal fusion · 0.7deformable attention · 0.7bird's-eye-view representation · 0.7
YearPublicationVenuePosition
2025 CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection Based View Transformation
abstract
Recently, camera-radar fusion-based 3D object detection methods in bird's eye view (BEV) have gained attention due to the complementary characteristics and cost-effectiveness of these sensors. Previous approaches using forward projection struggle with sparse BEV feature generation, while those employing backward projection overlook depth ambiguity, leading to false positives. In this paper, to address the aforementioned limitations, we propose a novel camera-radar fusion-based 3D object detection and segmentation model named CRAB (Camera-Radar fusion for reducing depth Ambiguity in Backward projection-based view transformation), using a backward projection that leverages radar to mitigate depth ambiguity. During the view transformation, CRAB aggregates perspective view image context features into BEV queries. It improves depth distinction among queries along the same ray by combining the dense but unreliable depth distribution from images with the sparse yet precise depth information from radar occupancy. We further introduce spatial cross-attention with a feature map containing radar context information to enhance the comprehension of the 3D scene. When evaluated on the nuScenes open dataset, our proposed approach achieves a state-of-the-art performance among backward projection-based camera-radar fusion methods with 62.4% NDS and 54.0% mAP in 3D object detection.
In-Jae Lee, Sihwan Hwang, Youngseok Kim 0001, Wonjune Kim, Sanmin Kim, Dongsuk Kum
ICRA1
2025 OpenBox: Annotate Any Bounding Boxes in 3D
abstract
Unsupervised and open-vocabulary 3D object detection has recently gained attention, particularly in autonomous driving, where reducing annotation costs and recognizing unseen objects are critical for both safety and scalability. However, most existing approaches uniformly annotate 3D bounding boxes, ignore objects’ physical states, and require multiple self-training iterations for annotation refinement, resulting in suboptimal quality and substantial computational overhead. To address these challenges, we propose OpenBox, a two-stage automatic annotation pipeline that leverages a 2D vision foundation model. In the first stage, OpenBox associates instance-level cues from 2D images processed by a vision foundation model with the corresponding 3D point clouds via context-aware refinement. In the second stage, it categorizes instances by rigidity and motion state, then generates adaptive bounding boxes with class-specific size statistics. As a result, OpenBox produces high-quality 3D bounding box annotations without requiring self-training. Experiments on the Waymo Open Dataset (WOD), the Lyft Level 5 Perception dataset, and the nuScenes dataset demonstrate improved accuracy and efficiency over baselines.
In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park
NeurIPS1
2023 Predict to Detect: Prediction-guided 3D Object Detection using Sequential Images
abstract
Recent camera-based 3D object detection methods have introduced sequential frames to improve the detection performance hoping that multiple frames would mitigate the large depth estimation error. Despite improved detection performance, prior works rely on naive fusion methods (e.g., concatenation) or are limited to static scenes (e.g., temporal stereo), neglecting the importance of the motion cue of objects. These approaches do not fully exploit the potential of sequential images and show limited performance improvements. To address this limitation, we propose a novel 3D object detection model, P2D (Predict to Detect), that integrates a prediction scheme into a detection framework to explicitly extract and leverage motion features. P2D predicts object information in the current frame using solely past frames to learn temporal motion features. We then introduce a novel temporal feature aggregation method that attentively exploits Bird’s-Eye-View (BEV) features based on predicted object information, resulting in accurate 3D object detection. Experimental results demonstrate that P2D improves mAP and NDS by 3.0% and 3.7% compared to the sequential image-based baseline, proving that incorporating a prediction scheme can significantly improve detection accuracy.
Sanmin Kim, Youngseok Kim 0001, In-Jae Lee, Dongsuk Kum
ICCV3
2023 CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception
abstract
Autonomous driving requires an accurate and fast 3D perception system that includes 3D object detection, tracking, and segmentation. Although recent low-cost camera-based approaches have shown promising results, they are susceptible to poor illumination or bad weather conditions and have a large localization error. Hence, fusing camera with low-cost radar, which provides precise long-range measurement and operates reliably in all environments, is promising but has not yet been thoroughly investigated. In this paper, we propose Camera Radar Net (CRN), a novel camera-radar fusion framework that generates a semantically rich and spatially accurate bird’s-eye-view (BEV) feature map for various tasks. To overcome the lack of spatial information in an image, we transform perspective view image features to BEV with the help of sparse but accurate radar points. We further aggregate image and radar feature maps in BEV using multi-modal deformable attention designed to tackle the spatial misalignment between inputs. CRN with real-time setting operates at 20 FPS while achieving comparable performance to LiDAR detectors on nuScenes, and even outperforms at a far distance on 100m setting. Moreover, CRN with offline setting yields 62.4% NDS, 57.5% mAP on nuScenes test set and ranks first among all camera and camera-radar 3D object detectors.
Youngseok Kim 0001, Juyeb Shin, Sanmin Kim, In-Jae Lee, Dongsuk Kum
ICCV4