EDBT 2026 Demo / reviewers in the wild / expert
Junho Koh
dblp:223/5511
· DBLP profile ↗
12ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-2318-9128ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal FusionabstractRadar-camera fusion methods have emerged as a cost-effective approach for 3D object detection but still lag behind LiDAR-based methods in performance. Recent works have focused on employing temporal fusion and Knowledge Distillation (KD) strategies to overcome these limitations. However, existing approaches have not sufficiently accounted for uncertainties arising from object motion or sensor-specific errors inherent in radar and camera modalities. In this work, we propose RCTDistill, a novel cross-modal KD method based on temporal fusion, comprising three key modules: Range-Azimuth Knowledge Distillation (RAKD), Temporal Knowledge Distillation (TKD), and Region-Decoupled Knowledge Distillation (RDKD). RAKD is designed to consider the inherent errors in the range and azimuth directions, enabling effective knowledge transfer from LiDAR features to refine inaccurate BEV representations. TKD mitigates temporal misalignment caused by dynamic objects by aligning historical radar-camera BEV features with current LiDAR representations. RDKD enhances feature discrimination by distilling relational knowledge from the teacher model, allowing the student to differentiate foreground and background features. RCTDistill achieves state-of-the-art radar-camera fusion performance on both the nuScenes and View-of-Delft (VoD) datasets, with the fastest inference speed of 26.2 FPS. Geonho Bang, Minjae Seong, Jisong Kim, Geunju Baek, Daye Oh, Junhyung Kim, Junho Koh |
ICCV | 7 |
| 2025 | UCFFormer: Recognizing human actions from multimodal sensors using unified contrastive fusion transformer
Kyoung Ok Yang, Junho Koh |
Neurocomputing | 2 |
| 2025 | OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D PerceptionabstractMulti-view camera-based 3D perception can be conducted using bird’s eye view (BEV) features obtained through perspective view-to-BEV transformations. Several studies have shown that the performance of these 3D perception methods can be further enhanced by combining sequential BEV features obtained from multiple camera frames. However, even after compensating for the ego-motion of an autonomous agent, the performance gain from temporal aggregation is limited when combining a large number of image frames. This limitation arises due to dynamic changes in BEV features over time caused by object motion. In this paper, we introduce a novel temporal 3D perception method called OnlineBEV, which combines BEV features over time using a recurrent structure. This structure increases the effective number of combined features with minimal memory usage. However, it is critical to spatially align the features over time to maintain strong performance. OnlineBEV employs the Motion-guided BEV Fusion Network (MBFNet) to achieve temporal feature alignment. MBFNet extracts motion features from consecutive BEV frames and dynamically aligns historical BEV features with current ones using these motion features. To enforce temporal feature alignment explicitly, we use Temporal Consistency Learning Loss, which captures discrepancies between historical and target BEV features. Experiments conducted on the nuScenes benchmark demonstrate that OnlineBEV achieves significant performance gains over the current best method, SOLOFusion. OnlineBEV achieves 63.9% NDS on the nuScenes test set, recording state-of-the-art performance in the camera-only 3D object detection task. Junho Koh, Youngwoo Lee |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object DetectionabstractDeveloping high-performance, real-time architectures for LiDAR-based 3D object detectors is essential for the successful commercialization of autonomous vehicles. Pillar-based methods stand out as a practical choice for onboard deployment due to their computational efficiency. However, despite their efficiency, these methods can sometimes underperform compared to alternative point encoding techniques such as Voxel-encoding or PointNet++. We argue that current pillar-based methods have not sufficiently captured the fine-grained distributions of LiDAR points within each pillar structure. Consequently, there exists considerable room for improvement in pillar feature encoding. In this paper, we introduce a novel pillar encoding architecture referred to as Fine-Grained Pillar Feature Encoding (FG-PFE). FG-PFE utilizes Spatio-Temporal Virtual (STV) grids to capture the distribution of point clouds within each pillar across vertical, temporal, and horizontal dimensions. Through STV grids, points within each pillar are individually encoded using Vertical PFE (V-PFE), Temporal PFE (T-PFE), and Horizontal PFE (H-PFE). These encoded features are then aggregated through an Attentive Pillar Aggregation method. Our experiments conducted on the nuScenes dataset demonstrate that FG-PFE achieves significant performance improvements over baseline models such as PointPillar, CenterPoint-Pillar, and PillarNet, with only a minor increase in computational overhead. Konyul Park, Yecheol Kim, Junho Koh, Byungwoo Park |
ICRA | 3 |
| 2023 | MGTANet: Encoding Sequential LiDAR Points Using Long Short-Term Motion-Guided Temporal Attention for 3D Object DetectionabstractMost scanning LiDAR sensors generate a sequence of point clouds in real-time. While conventional 3D object detectors use a set of unordered LiDAR points acquired over a fixed time interval, recent studies have revealed that substantial performance improvement can be achieved by exploiting the spatio-temporal context present in a sequence of LiDAR point sets. In this paper, we propose a novel 3D object detection architecture, which can encode LiDAR point cloud sequences acquired by multiple successive scans. The encoding process of the point cloud sequence is performed on two different time scales. We first design a short-term motion-aware voxel encoding that captures the short-term temporal changes of point clouds driven by the motion of objects in each voxel. We also propose long-term motion-guided bird’s eye view (BEV) feature enhancement that adaptively aligns and aggregates the BEV feature maps obtained by the short-term voxel encoding by utilizing the dynamic motion context inferred from the sequence of the feature maps. The experiments conducted on the public nuScenes benchmark demonstrate that the proposed 3D object detector offers significant improvements in performance compared to the baseline methods and that it sets a state-of-the-art performance for certain 3D object detection categories. Code is available at https://github.com/HYjhkoh/MGTANet.git. Junho Koh, Junhyung Lee, Youngwoo Lee, Jaekyum Kim |
AAAI | 1 |
| 2023 | D-Align: Dual Query Co-attention Network for 3D Object Detection Based on Multi-frame Point Cloud SequenceabstractLiDAR sensors are widely used for 3D object detection in various mobile robotics applications. LiDAR sensors continuously generate point cloud data in real-time. Conventional 3D object detectors detect objects using a set of points acquired over a fixed duration. However, recent studies have shown that the performance of object detection can be further enhanced by utilizing spatio-temporal information obtained from point cloud sequences. In this paper, we propose a new 3D object detector, named D-Align, which can effectively produce strong bird's-eye-view (BEV) features by aligning and aggregating the features obtained from a sequence of point sets. The proposed method includes a novel dual-query co-attention network that uses two types of queries, including target query set (T-QS) and support query set (S-QS), to update the features of target and support frames, respectively. D-Align aligns S-QS to T-QS based on the temporal context features extracted from the adjacent feature maps and then aggregates S-QS with T-QS using a gated fusion mechanism. The dual queries are updated through multiple attention layers to progressively enhance the target frame features used to produce the detection results. Our experiments on the nuScenes dataset show that the proposed D-Align method greatly improved the performance of a single frame-based baseline method and significantly outperformed the latest 3D object detectors. Code is available at https://github.com/junhyung-SPALab/D-Align. Junhyung Lee, Junho Koh, Youngwoo Lee |
ICRA | 2 |
| 2022 | Joint 3D Object Detection and Tracking Using Spatio-Temporal Representation of Camera Image and LiDAR Point CloudsabstractIn this paper, we propose a new joint object detection and tracking (JoDT) framework for 3D object detection and tracking based on camera and LiDAR sensors. The proposed method, referred to as 3D DetecTrack, enables the detector and tracker to cooperate to generate a spatio-temporal representation of the camera and LiDAR data, with which 3D object detection and tracking are then performed. The detector constructs the spatio-temporal features via the weighted temporal aggregation of the spatial features obtained by the camera and LiDAR fusion. Then, the detector reconfigures the initial detection results using information from the tracklets maintained up to the previous time step. Based on the spatio-temporal features generated by the detector, the tracker associates the detected objects with previously tracked objects using a graph neural network (GNN). We devise a fully-connected GNN facilitated by a combination of rule-based edge pruning and attention-based edge gating, which exploits both spatial and temporal object contexts to improve tracking performance. The experiments conducted on both KITTI and nuScenes benchmarks demonstrate that the proposed 3D DetecTrack achieves significant improvements in both detection and tracking performances over baseline methods and achieves state-of-the-art performance among existing methods through collaboration between the detector and tracker. Junho Koh, Jaekyum Kim, Jin Hyeok Yoo, Yecheol Kim, Dongsuk Kum |
AAAI | 1 |
| 2021 | Joint Representation of Temporal Image Sequences and Object Motion for Video Object DetectionabstractIn this paper, we propose a new video object detection (VoD) method, referred to as temporal feature aggregation and motion-aware VoD (TM-VoD), that produces a joint representation of temporal image sequences and object motion. The TM-VoD generates strong spatiotemporal features for VOD by temporally redundant information in an image sequence and the motion context. These are produced at the feature level in the region proposal stage and at the instance level in the refinement stage. In the region proposal stage, visual features are temporally fused with appropriate weights at the pixel level via gated attention model. Furthermore, pixel level motion features are obtained by capturing the changes between adjacent visual feature maps. In the refinement stage, the visual features are aligned and aggregated at the instance level. We propose a novel feature alignment method, which uses the initial region proposals as anchors to predict the box coordinates for all video frames. Moreover, the instance level motion features are obtained by applying the region of interest (RoI) pooling to the pixel level motion features and by encoding the sequential changes in the box coordinates. Finally, all these instance level features are concatenated to produce a joint representation of the objects. Experiments on the ImageNet VID dataset demonstrate that the proposed method significantly outperforms existing VoDs and achieves performance comparable with that of state-of-the-art VoDs. Junho Koh, Jaekyum Kim, Younji Shin, Byeongwon Lee, Seungji Yang |
ICRA | 1 |
| 2020 | Video Object Detection Using Object's Motion Context and Spatio-Temporal Feature AggregationabstractThe deep learning technique has recently led to significant improvement in object detection accuracy. In many applications, object detection is performed on video data consisting of a sequence of two-dimensional (2D) image frames. Numerous object detection schemes have been designed to detect objects independently in each video frame. Though temporal information within adjacent image frames can be exploited in subsequent object tracking stage, it has been shown that the object detection accuracy can be significantly improved by exploiting the temporal structure in the image sequence in the object detection stage. In this paper, we propose a novel video object detection method that exploits both the motion context inferred from the adjacent frames and the spatio-temporal features aggregated over the image sequence. First, correlation between the spatial feature maps over two adjacent frames are computed and the embedding vector, representing the motion context, is obtained by encoding the N correlation maps using long short term memory (LSTM). In addition to utilizing the motion context, the spatial feature maps for ( N+1) consecutive frames are aggregated to boost the quality of the feature map. The gated attention network is employed to selectively combine the temporal feature maps based on their relevance to the feature map in the present image frame. While most video object detectors have been developed for two-stage object detectors, our proposed idea applies to one-stage detectors with the advantage of low computational complexity in practical real-time applications. Our numerical evaluation conducted on the ImageNet object detection from video (VID) dataset demonstrates that our proposed network achieves significant performance gain over the baseline algorithms and outperforms the existing one-stage video object detectors. Jaekyum Kim, Junho Koh, Byeongwon Lee, Seungji Yang |
ICPR | 2 |
| 2019 | Enhanced Object Detection in Bird's Eye View Using 3D Global Context Inferred From Lidar Point DataabstractIn this paper, we present a new deep neural network architecture, which detects objects in bird's eye view (BEV) using Lidar sensor data in autonomous driving scenarios. The key idea of the proposed method is to improve the accuracy of the object detection by exploiting the 3D global context provided by the whole set of Lidar points. The overall structure of the proposed method consists of two parts: 1) the detection core network (DetNet) and 2) the context extraction network (ConNet). First, the DetNet generates the BEV representation by projecting the Lidar points into the BEV plane and applies the CNN to extract the feature maps locally activated on the objects. The ConNet directly processes the whole set of the Lidar points to produce the 1 × 1 × k feature vector capturing the 3D geometrical structure of the surrounding in the global scale. The context vector produced by the ConNet is concatenated to each pixel of the feature maps obtained by the DetNet. The combined feature maps are used to regress the oriented bounding box and identify the category of the object. The experiments evaluated on the public KITTI dataset show that the use of the context feature offers the significant performance gain over the baseline and the proposed object detector achieves the competitive performance as compared to the state of the art 3D object detectors. Yecheol Kim, Jaekyum Kim, Junho Koh |
IV | 3 |
| 2018 | Robust Deep Multi-modal Learning Based on Gated Information Fusion Network
Jaekyum Kim, Junho Koh, Yecheol Kim, Jaehyung Choi, Youngbae Hwang |
ACCV (4) | 2 |
| 2018 | Robust Camera Lidar Sensor Fusion Via Deep Gated Information Fusion NetworkabstractIn this paper, we introduce a new deep learning architecture for camera and Lidar sensor fusion. The proposed scheme performs 2D object detection using the RGB camera image and the depth, height, and intensity images generated by projecting the 3D Lidar point cloud into camera image plane. The proposed object detector consists of two convolutional neural networks (CNNs) that process the RGB and Lidar images separately as well as the fusion network that combines the feature maps produced at the intermediate layers of the CNNs. We aim to develop a robust object detector that maintains good object detection accuracy even when the quality of the sensor signals is degraded for object detection. Towards this end, we devise the gated fusion unit (GFU) that adjusts the contribution of the feature maps generated by two CNN structures via gating mechanism. Using the GFU, the proposed object detector can fuse the high level feature maps drawn from two modalities with appropriate weights to achieve robust performance. Experiments conducted on the challenging KITTI benchmark show that the proposed camera and Lidar fusion network outperforms the conventional sensor fusion methods even when either of the camera and Lidar sensor signals is corrupted by missing data, occlusion, noise, and illumination change. Jaekyum Kim, Jaehyung Choi, Yecheol Kim, Junho Koh, Chung Choo Chung |
Intelligent Vehicles Symposium | 4 |