Daehyun Ji

dblp:274/9684 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0002-9081-9786ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Mamba-VOS: Efficient Video Object Segmentation with Selective State Space Models
Cheolhun Jang, Wontae Kim 0002, Daehyun Ji, Nam Ik Cho
ICPR (5)3
2025 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
abstract
The resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the practical necessity for real-time deployment require smaller query resolutions, which inevitably leads to an information loss. Therefore, it is essential to encode and preserve rich visual details within limited query sizes while ensuring a comprehensive representation of 3D occupancy. To this end, we introduce ProtoOcc, a novel occupancy network that leverages prototypes of clustered image segments in view transformation to enhance low-resolution context. In particular, the mapping of 2D prototypes onto 3D voxel queries encodes high-level visual geometries and complements the loss of spatial information from reduced query resolutions. Additionally, we design a multi-perspective decoding strategy to efficiently disentangle the densely compressed visual cues into a high-dimensional 3D occupancy scene. Experimental results on both Occ3D and SemanticKITTI benchmarks demonstrate the effectiveness of the proposed method, showing clear improvements over the baselines. More importantly, ProtoOcc achieves competitive performance against the baselines even with 75% reduced voxel resolution. Project page: https://kuai-lab.github.io/cvpr2025protoocc.
Gyeongrok Oh, Sungjune Kim, Heeju Ko, Hyung-gun Chi, Jinkyu Kim 0001, Daehyun Ji, Sujin Jang, Sangpil Kim
CVPR7
2025 Diffusion Transformer Meets Multi-Level Wavelet Spectrum for Single Image Super-Resolution
abstract
Discrete Wavelet Transform (DWT) has been widely explored to enhance the performance of image superresolution (SR). Despite some DWT-based methods improving SR by capturing fine-grained frequency signals, most existing approaches neglect the interrelations among multiscale frequency sub-bands, resulting in inconsistencies and unnatural artifacts in the reconstructed images. To address this challenge, we propose a Diffusion Transformer model based on image Wavelet spectra for SR (DTWSR). DTWSR incorporates the superiority of diffusion models and transformers to capture the interrelations among multiscale frequency sub-bands, leading to a more consistence and realistic SR image. Specifically, we use a Multi-level Discrete Wavelet Transform to decompose images into wavelet spectra. A pyramid tokenization method is proposed which embeds the spectra into a sequence of tokens for transformer model, facilitating to capture features from both spatial and frequency domain. A dual-decoder is designed elaborately to handle the distinct variances in low-frequency and high-frequency sub-bands, without omitting their alignment in image generation. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of our method, with high performance on both perception quality and fidelity.
Paul Barom Jeon, Daehyun Ji
ICCV6
2025 Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning
abstract
Navigating in an unfamiliar environment during deployment poses a critical challenge for a vision-language navigation (VLN) agent. Yet, test-time adaptation (TTA) remains relatively underexplored in robotic navigation, leading us to the fundamental question: what are the key properties of TTA for online VLN? In our view, effective adaptation requires three qualities: 1) flexibility in handling different navigation outcomes, 2) interactivity with external environment, and 3) maintaining a harmony between plasticity and stability. To address this, we introduce FeedTTA, a novel TTA framework for online VLN utilizing feedback-based reinforcement learning. Specifically, FeedTTA learns by maximizing binary episodic feedback, a practical setup in which the agent receives a binary scalar after each episode that indicates the success or failure of the navigation. Additionally, we propose a gradient regularization technique that leverages the binary structure of FeedTTA to achieve a balance between plasticity and stability during adaptation. Our extensive experiments on challenging VLN benchmarks demonstrate the superior adaptability of FeedTTA, even outperforming the state-of-the-art offline training methods in REVERIE benchmark with a single stream of learning.
Sungjune Kim, Gyeongrok Oh, Heeju Ko, Daehyun Ji, Sujin Jang, Sangpil Kim
ICML4
2025 CDP: Constrained Diffusion Policies with Mirror Diffusion Model for Safety-Assured Imitation Learning
abstract
This paper presents a novel imitation learning framework, called constrained diffusion policy (CDP). The primary objective of CDP is to ensure that learned policies strictly adhere to safety constraints while imitating expert demonstrations. To achieve this, we define a polytopic constraint that represents the safe boundary for obstacle-free region. We introduce a novel mirror map and its inverse function to incorporate a generalized polytopic constraint manifold into the mirror diffusion model. By mapping sampled data onto a constrained manifold, the mirror diffusion model generates actions that satisfy safety constraints. This approach successfully addresses the safety issues commonly encountered in conventional imitation learning models. We apply the proposed framework to mobile navigation tasks in robotics, using the Isaac Gym simulator and the Unitree Go2 quadrupedal robot. Experimental results demonstrate that the proposed framework can successfully train policies that imitate expert behaviors while strictly maintaining safety constraints, thereby achieving safety-assured imitation learning.
Taeoh Ha, Hyunsoo Cha 0003, Daehyun Ji
IROS3
2024 CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection
abstract
Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised domain adaptation (UDA) method, called CMDA, which (i) leverages visual semantic cues from an image modality (i.e., camera images) as an effective semantic bridge to close the domain gap in the cross-modal Bird's Eye View (BEV) representations. Further, (ii) we also introduce a self-training-based learning strategy, wherein a model is adversarially trained to generate domain-invariant features, which disrupt the discrimination of whether a feature instance comes from a source or an unseen target domain. Overall, our CMDA framework guides the 3DOD model to generate highly informative and domain-adaptive features for novel data distributions. In our extensive experiments with large-scale benchmarks, such as nuScenes, Waymo, and KITTI, those mentioned above provide significant performance gains for UDA tasks, achieving state-of-the-art performance.
Gyusam Chang, Wonseok Roh, Sujin Jang, Daehyun Ji, Gyeongrok Oh, Jinsun Park, Jinkyu Kim 0001, Sangpil Kim
AAAI5
2024 Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection
abstract
Recent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks. However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target datasets (i.e., direct transfer) due to the inevitable geometric misalignment between the source and target domains. In practice, we also encounter constraints on resources for training models and collecting annotations for the successful deployment of 3D object detectors. In this paper, we propose Unified Domain Generalization and Adaptation (UDGA), a practical solution to mitigate those drawbacks. We first propose Multi-view Overlap Depth Constraint that leverages the strong association between multi-view, significantly alleviating geometric gaps due to perspective view changes. Then, we present a Label-Efficient Domain Adaptation approach to handle unfamiliar targets with significantly fewer amounts of labels (i.e., 1$\%$ and 5$\%)$, while preserving well-defined source knowledge for training efficiency. Overall, UDGA framework enables stable detection performance in both source and target domains, effectively bridging inevitable domain gaps, while demanding fewer annotations. We demonstrate the robustness of UDGA with large-scale benchmarks: nuScenes, Lyft, and Waymo, where our framework outperforms the current state-of-the-art methods.
Gyusam Chang, Donghyun Kim 0006, Jinkyu Kim 0001, Daehyun Ji, Sujin Jang, Sangpil Kim
NeurIPS6
2024 Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation
abstract
Predicting and constructing road geometric information (e.g., lane lines, road markers) is a crucial task for safe autonomous driving, while such static map elements can be repeatedly occluded by various dynamic objects on the road. Recent studies have shown significantly improved vectorized high-definition (HD) map construction performance, but there has been insufficient investigation of temporal information across adjacent input frames (i.e., clips), which may lead to inconsistent and suboptimal prediction results. To tackle this, we introduce a novel paradigm of clip-level vectorized HD map construction, MapUnveiler, which explicitly unveils the occluded map elements within a clip input by relating dense image representations with efficient clip tokens. Additionally, MapUnveiler associates inter-clip information through clip token propagation, effectively utilizing long- term temporal map information. MapUnveiler runs efficiently with the proposed clip-level pipeline by avoiding redundant computation with temporal stride while building a global map relationship. Our extensive experiments demonstrate that MapUnveiler achieves state-of-the-art performance on both the nuScenes and Argoverse2 benchmark datasets. We also showcase that MapUnveiler significantly outperforms state-of-the-art approaches in a challenging setting, achieving +10.7% mAP improvement in heavily occluded driving road scenes. The project page can be found at https://mapunveiler.github.io.
Nayeon Kim 0006, Hongje Seong, Daehyun Ji, Sujin Jang
NeurIPS3
2023 D-3DLD: Depth-Aware Voxel Space Mapping for Monocular 3D Lane Detection with Uncertainty
abstract
The estimation of 3D lanes from monocular RGB images is a fundamentally ill-posed problem. Previous studies have assumed that all lanes are on a flat ground plane. However, we argue that the algorithms based on this assumption have difficulty in detecting various lanes in actual driving environments. Contrary to previous approaches, we expand rich contextual features from an image domain to a 3D space by utilizing depth-aware voxel mapping. In addition, we determine 3D lanes based on voxelized features. We design a new lane representation combined with uncertainties and predict the confidence intervals of 3D lane points using Laplace loss. Experimental results show that the proposed method achieves state-of-the-art detection accuracy on three challenging datasets, including two real-world datasets, and significantly outperforms existing methods with reasonable computation load.
Nayeon Kim 0006, Moonsub Byeon, Daehyun Ji, Dokwan Oh
ICASSP3
2023 STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection
abstract
3D object detection (3DOD) from multi-view images is an economically appealing alternative to expensive LiDAR-based detectors, but also an extremely challenging task due to the absence of precise spatial cues. Recent studies have leveraged the teacher-student paradigm for cross-modal distillation, where a strong LiDAR-modality teacher transfers useful knowledge to a multi-view-based image-modality student. However, prior approaches have only focused on minimizing global distances between cross-modal features, which may lead to suboptimal knowledge distillation results. Based on these insights, we propose a novel structural and temporal cross-modal knowledge distillation (STXD) framework for multi-view 3DOD. First, STXD reduces redundancy of the feature components of the student by regularizing the cross-correlation of cross-modal features, while maximizing their similarities. Second, to effectively transfer temporal knowledge, STXD encodes temporal relations of features across a sequence of frames via similarity maps. Lastly, STXD also adopts a response distillation method to further enhance the quality of knowledge distillation at the output-level. Our extensive experiments demonstrate that STXD significantly improves the NDS and mAP of the based student detectors by 2.8%~4.5% on the nuScenes testing dataset.
Sujin Jang, Sung Ju Hwang, Daehyun Ji
NeurIPS5
2022 Semi-Supervised 360° Depth Estimation from Multiple Fisheye Cameras with Pixel-Level Selective Loss
abstract
In this paper, we study a practical omnidirectional depth estimation with neural networks that enables effective learning on real world data obtained using wide-baseline multiple fish-eye cameras. Most previous approaches only used synthetic data providing dense and accurate depth ground truth (GT). However, it is unrealistic to acquire such high quality GT data in real world due to limitations of the existing depth sensors. We first introduce two critical problems that can reduce the accuracy of depth estimation: depth GT sparsity and sensor calibration error. We then propose a novel semi-supervised learning method using pixel-level loss that selectively uses supervised loss and unsupervised re-projection loss according to existence of GT. Empirical results demonstrate that our method efficiently reduces the performance degradation in both simulation on synthetic data and real world data using sparse depth sensor.
Daeul Park, Daehyun Ji
ICASSP4
2022 Stability and dissipativity criteria for neural networks with time-varying delays via an augmented zero equality approach
Seung-Hoon Lee 0001, Myeong-Jin Park, Daehyun Ji, Oh-Min Kwon 0001
Neural Networks3
2021 How to handle noisy labels for robust learning from uncertainty
Daehyun Ji, Dokwan Oh, Yoonsuk Hyun, Oh-Min Kwon 0001, Myeong-Jin Park
Neural Networks1
2020 Segmenting 2K-Videos at 36.5 FPS with 24.3 GFLOPs: Accurate and Lightweight Realtime Semantic Segmentation Network
abstract
We propose a fast and lightweight end-to-end convolutional network architecture for real-time segmentation of high resolution videos, NfS-SegNet, that can segement 2K-videos at 36.5 FPS with 24.3 GFLOPS. This speed and computation-efficiency is due to following reasons: 1) The encoder network, NfS-Net, is optimized for speed with simple building blocks without memory-heavy operations such as depthwise convolutions, and outperforms state-of-the-art lightweight CNN architectures such as SqueezeNet [2], Mo- bileNet v1 [3] & v2 [4] and ShuffleNet v1 [5] & v2 [6] on image classification with significantly higher speed. 2) The NfS- SegNet has an asymmetric architecture with deeper encoder and shallow decoder, whose design is based on our empirical finding that the decoder is the main bottleneck in computation with relatively small contribution to the final performance. 3) Our novel uncertainty-aware knowledge distillation method guides the teacher model to focus its knowledge transfer on the most difficult image regions. We validate the performance of NfS-SegNet with the CITYSCAPE [1] benchmark, on which it achieves state-of-the-art performance among lightweight segementation models in terms of both accuracy and speed.
Dokwan Oh, Daehyun Ji, Cheolhun Jang, Yoonsuk Hyun, Hong S. Bae, Sung Ju Hwang
ICRA2