Jaehyup Lee

dblp:255/6314 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-2137-2650ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Mechanistic Dissection of Cross-Attention Subspaces in Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models utilize cross-attention to integrate textual information into the visual latent space, yet the transformation from text embeddings to latent features remains largely unexplored. We provide a mechanistic analysis of the output-value (OV) circuits within cross-attention layers through spectral analysis via singular value decomposition. Our analysis reveals that semantic concepts are encoded in low-dimensional subspaces spanned by singular vectors in OV circuits across cross-attention heads. To verify this, we intervene on concept-related components in the diffusion process, demonstrating that intervention on identified spectral components affects conceptual changes. We further validate these findings by examining visual outputs of isolated subspaces and their alignment with text embedding space. Through this mechanistic understanding, we demonstrate that only nullifying these spectral components can achieve targeted concept removal with performance comparable to existing methods while providing interpretability. Our work reveals how cross-attention layers encode semantic concepts in spectral subspaces of OV circuits, providing mechanistic insights and enabling precise concept manipulation without retraining.
Jun-Hyun Bae, Wonyong Jo, Jaehyup Lee, Heechul Jung
AAAI3
2026 From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
abstract
Cross-view image retrieval, particularly street-to-satellite matching, is a critical task for applications such as autonomous navigation, urban planning, and localization in GPS-denied environments. However, existing approaches often require supervised training on curated datasets and rely on panoramic or UAV-based images, which limits real-world deployment. In this paper, we present a simple yet effective cross-view image retrieval framework that leverages a pretrained vision encoder and a large language model (LLM), requiring no additional training. Given a monocular street-view image, our method extracts geographic cues through web-based image search and LLM-based location inference, generates a satellite query via geocoding API, and retrieves matching tiles using a pretrained vision encoder (e.g., DINOv2) with PCA-based whitening feature refinement. Despite not using ground-truth supervision or finetuning, our proposed method outperforms prior learning-based approaches on the benchmark dataset under zero-shot settings. Moreover, our pipeline enables automatic construction of semantically aligned street-to-satellite datasets, which is offering a scalable and cost-efficient alternative to manual annotation. All source codes will be made publicly available at https://jeonghomin.github.io/ street2orbit.github.io/.
Jeongho Min, Dongyoung Kim, Jaehyup Lee
WACV3
2025 U-Know-DiffPAN: An Uncertainty-aware Knowledge Distillation Diffusion Framework with Details Enhancement for PAN-Sharpening
abstract
Conventional methods for PAN-sharpening often struggle to restore fine details due to limitations in leveraging high-frequency information. Moreover, diffusion-based approaches lack sufficient conditioning to fully utilize Panchromatic (PAN) images and low-resolution multi-spectral (LRMS) inputs effectively. To address these challenges, we propose an uncertainty-aware knowledge distillation diffusion framework with details enhancement for PAN-sharpening, called U-Know-DiffPAN. The U-Know-DiffPAN incorporates uncertainty-aware knowledge distillation for effective transfer of feature details from our teacher model to a student one. The teacher model in our U-Know-DiffPAN captures frequency details through freqeuncy selective attention, facilitating accurate reverse process learning. By conditioning the encoder on compact vector representations of PAN and LRMS and the decoder on Wavelet transforms, we enable rich frequency utilization. So, the high-capacity teacher model distills frequency-rich features into a lightweight student model aided by an un certainty map. From this, the teacher model can guide the student model to focus on difficult image regions for PAN-sharpening via the usage of the uncertainty map. Extensive experiments on diverse datasets demonstrate the robustness and superior performance of our U-Know-DiffPAN over very recent state-of-the-art PAN-sharpening methods. The project page is available at https://kaist-viclab.github.io/U-Know-DiffPAN-site/.
Sungpyo Kim, Jeonghyeok Do, Jaehyup Lee, Munchurl Kim
CVPR3
2025 ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects
abstract
Weakly supervised Oriented Object Detection (WS-OOD) has gained attention as a cost-effective alternative to fully supervised methods, providing efficiency and high accuracy. Among weakly supervised approaches, horizontal bounding box (HBox) supervised OOD stands out for its ability to directly leverage existing HBox annotations while achieving the highest accuracy under weak supervision settings. This paper introduces adaptive bounding box scaling and symmetry-prior-based orientation prediction, called ABBSPO that is a framework for WS-OOD. Our ABBSPO addresses the limitations of previous HBox-supervised OOD methods, which compare ground truth (GT) HBoxes directly with predicted RBoxes’ minimum circumscribed rectangles, often leading to inaccuracies. To overcome this, we propose: (i) Adaptive Bounding Box Scaling (ABBS) that appropriately scales the GT HBoxes to optimize for the size of each predicted RBox, ensuring more accurate prediction for RBoxes’ scales; and (ii) a Symmetric Prior Angle (SPA) loss that uses the inherent symmetry of aerial objects for self-supervised learning, addressing the issue in previous methods where learning fails if they consistently make incorrect predictions for all three augmented views (original, rotated, and flipped). Extensive experimental results demonstrate that our ABBSPO achieves state-of-the-art results, outperforming existing methods.
Hyugjae Chang, Jaeho Moon, Jaehyup Lee, Munchurl Kim
CVPR4
2025 PAN-Crafter: Learning Modality-Consistent Alignment for Pan-Sharpening
abstract
PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment -- caused by sensor placement, acquisition timing, and resolution disparity -- induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11$\times$ faster inference time and 0.63$\times$ the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.
Jeonghyeok Do, Sungpyo Kim, Geunhyuk Youk, Jaehyup Lee, Munchurl Kim
ICCV4
2025 Uncertainty-Guided Face Matting for Occlusion-Aware Face Transformation
abstract
Face filters have become a key element of short-form video content, enabling a wide array of visual effects such as stylization and face swapping. However, their performance often degrades in the presence of occlusions, where objects like hands, hair, or accessories obscure the face. To address this limitation, we introduce the novel task of face matting, which estimates fine-grained alpha mattes to separate occluding elements from facial regions. We further present FaceMat, a trimap-free, uncertainty-aware framework that predicts high-quality alpha mattes under complex occlusions. Our approach leverages a two-stage training pipeline: a teacher model is trained to jointly estimate alpha mattes and per-pixel uncertainty using a negative log-likelihood (NLL) loss, and this uncertainty is then used to guide the student model through spatially adaptive knowledge distillation. This formulation enables the student to focus on ambiguous or occluded regions, improving generalization and preserving semantic consistency. Unlike previous approaches that rely on trimaps or segmentation masks, our framework requires no auxiliary inputs making it well-suited for real-time applications. In addition, we reformulate the matting objective by explicitly treating skin as foreground and occlusions as background, enabling clearer compositing strategies. To support this task, we newly constructed CelebAMat, a large-scale synthetic dataset specifically designed for occlusion-aware face matting. Extensive experiments show that FaceMat outperforms state-of-the-art methods across multiple benchmarks, enhancing the visual quality and robustness of face filters in real-world, unconstrained video scenarios. The source code and CelebAMat dataset are available at https://github.com/hyebin-c/FaceMat.git
Hyebin Cho, Jaehyup Lee
ACM Multimedia2
2025 3DEKD: 3D Explanation-Based Knowledge Distillation for Pillar-Based 3D Object Detection
abstract
LiDAR-based 3D object detection has been widely utilized in fields such as autonomous driving and robotics. However, the black-box nature of 3D models limits their interpretability, making it difficult to understand their predictions and evaluate significant feature contributions, which are essential for safety-critical applications. To address this, we propose a novel knowledge distillation method for 3D object detection that integrates explanation-based and aggregation techniques to achieve effective knowledge transfer and, as a result, enhance the model’s interpretability. Our method generates attribution maps that highlight the importance of 3D points in the teacher model and aggregates them into a single map. This map is aligned with the student model’s pillar features using corresponding coordinates, allowing pillar-wise feature mapping. Building on this feature mapping, to the best of our knowledge, this is the first study to propose a distillation process that effectively transfers the teacher model’s explanations of critical regions to the student model. Experimental results demonstrate that the proposed method increases 3D and BEV mAP by up to 2.09% and 0.84%, respectively, compared to the existing models.
Heejung Choi, Dabin Kang, Jaehyup Lee, Sanghyo Park 0001
IEEE Signal Process. Lett.4
2025 U-SET: Uncertainty-Aware SAR-to-EO Translation
abstract
Synthetic aperture radar (SAR) imagery has become increasingly vital across diverse applications, including military surveillance, environmental monitoring, and disaster response. However, interpreting SAR images poses challenges for non-experts owing to their distinct imaging characteristics, such as speckle noise and structural distortions. Additionally, inherent properties of SAR imaging and temporal disparities frequently lead to local misalignments between paired SAR and electro-optical (EO) images. To mitigate these issues, we introduce U-SET, an uncertainty-aware framework for SAR-to-EO translation that explicitly models pixel-wise uncertainty, facilitating locally adaptive learning. By leveraging the uncertainty estimation capabilities of deep-learning models, U-SET effectively prioritizes structurally complex and ambiguous regions during training. Comprehensive evaluations on our newly compiled KOMPSAT dataset demonstrate that U-SET achieves state-of-the-art performance, outperforming existing methods across six image quality metrics in both quantitative and qualitative assessments.
Minyoung Jeon, Hyun-Ho Kim, Juheon Park, Jaehyup Lee
IEEE Signal Process. Lett.5
2024 Segmentation-Guided Context Learning Using EO Object Labels for Stable SAR-to-EO Translation
abstract
Recently, the analysis and use of synthetic aperture radar (SAR) imagery have become crucial for surveillance, military operations, and environmental monitoring. A common challenge with SAR images is the presence of speckle noise, which can hinder their interpretability. To enhance the clarity of SAR images, this letter introduces a novel SAR-to-electro-optical (EO) image translation (SET) network, called SGCL-SET, which first incorporates EO object label information for stable translation. We use a pretrained segmentation network to provide the segmentation regions with their labels into learning the SET. Our SGCL-SET can be trained to effectively learn the translation for the regions of confusing contexts using the segmentation and label information. Through comprehensive experiments on our KOMPSAT dataset, our SGCL-SET significantly outperforms all the previous methods with large margins across nine image quality evaluation metrics.
Jaehyup Lee, Hyun-Ho Kim, Munchurl Kim
IEEE Geosci. Remote. Sens. Lett.1
2023 CFCA-SET: Coarse-to-Fine Context-Aware SAR-to-EO Translation With Auxiliary Learning of SAR-to-NIR Translation
abstract
Satellite Synthetic Aperture Radar (SAR) images are immensely valuable because they can be obtained regardless of weather and time conditions. However, SAR images have fatal noise and less contextual information, thus making it harder and less interpretable. So, translation of SAR to Electro-Optical (EO) images is highly required for easier interpretation. In this paper, we propose a novel coarse-to-fine context-aware SAR-to-EO image translation (CFCA-SET) framework and a misalignment-resistant loss for the misaligned pairs of SAR-EO images. With our auxiliary learning of SAR-to-Near-Infrared translation, CFCA-SET consists of a two-stage training: (i) the low-resolution SAR-to-EO translation is learned in the coarse stage via a local self-attention module that helps diminish the SAR noise, and (ii) the resulting output is used as guidance in the fine stage to generate the SAR colorization of high resolution. Our proposed auxiliary learning of SAR-to-NIR translation can successfully lead CFCA-SET to learn distinguishable characteristics of various SAR objects with less confusion in a context-aware manner. To handle the inevitable misalignment problem between SAR and EO images, we newly design a misalignment-resistant loss function. Extensive experimental results show that our CFCA-SET can generate more recognizable and understandable EO-like images compared to other methods in terms of nine image quality metrics. Our CFCA-SET surpasses the state-of-the-art methods for two (QXS and CASET) datasets with the improvements: PSNR (3.6%, 29%), ERGAS (7.4%, 30%), SSIM (15%, 15%), SAM (21%, 38%), Ds (16%, 13%), QNR (1.5%, 3.1%), CHD (18%, 12%), LPIPS (4.2%, 8%), and FID (9.0%, 33%).
Jaehyup Lee, Hyebin Cho, Hyun-Ho Kim, Munchurl Kim
IEEE Trans. Geosci. Remote. Sens.1
2021 SIPSA-Net: Shift-Invariant Pan Sharpening With Moving Object Alignment for Satellite Imagery
abstract
Pan-sharpening is a process of merging a high-resolution (HR) panchromatic (PAN) image and its corresponding low-resolution (LR) multi-spectral (MS) image to create an HR-MS and pan-sharpened image. However, due to the different sensors’ locations, characteristics and acquisition time, PAN and MS image pairs often tend to have various amounts of misalignment. Conventional deep-learning-based methods that were trained with such misaligned PAN-MS image pairs suffer from diverse artifacts such as double-edge and blur artifacts in the resultant PAN-sharpened images. In this paper, we propose a novel framework called shift-invariant pan-sharpening with moving object alignment (SIPSA-Net) which is the first method to take into account such large misalignment of moving object regions for PAN sharpening. The SISPA-Net has a feature alignment module (FAM) that can adjust one feature to be aligned to another feature, even between the two different PAN and MS domains. For better alignment in pan-sharpened images, a shift-invariant spectral loss is newly designed, which ignores the inherent misalignment in the original MS input, thereby having the same effect as optimizing the spectral loss with a well-aligned MS image. Extensive experimental results show that our SIPSA-Net can generate pan-sharpened images with remarkable improvements in terms of visual quality and alignment, compared to the state-of-the-art methods.
Jaehyup Lee, Soomin Seo, Munchurl Kim
CVPR1
2020 A CNN-Based Multi-scale Super-Resolution Architecture on FPGA for 4K/8K UHD Applications
Jae-Seok Choi, Jaehyup Lee, Munchurl Kim
MMM (2)3