Taeheon Kim

dblp:182/7467 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 An Information Theoretic Evaluation Metric for Strong Unlearning
abstract
Machine unlearning (MU) aims to remove the influence of specific data from trained models, addressing privacy concerns and ensuring compliance with regulations such as the "right to be forgotten." Evaluating strong unlearning, where the unlearned model is indistinguishable from one retrained without the forgetting data, remains a significant challenge in deep neural networks (DNNs). Common black-box metrics, such as variants of membership inference attacks and accuracy comparisons, primarily assess model outputs but often fail to capture residual information in intermediate layers. To bridge this gap, we introduce the Information Difference Index (IDI), a novel white-box metric inspired by information theory. IDI quantifies retained information in intermediate features by measuring mutual information between those features and the labels to be forgotten, offering a more comprehensive assessment of unlearning efficacy. Our experiments demonstrate that IDI effectively measures the degree of unlearning across various datasets and architectures, providing a reliable tool for evaluating strong unlearning in DNNs.
Dongjae Jeon, Wonje Jeung, Taeheon Kim, Albert No
AAAI3
2026 OSCAR: Optical-Aware Semantic Control for Aleatoric Refinement in Sar-to-Optical Translation
Hyunseo Lee, Sang Min Kim, Ho Kyung Shin, Taeheon Kim, Woo-Jeoung Nam
ICPR (10)4
2025 Representation Bending for Large Language Model Safety
abstract
Large Language Models (LLMs) have emerged as powerful tools, but their inherent safety risks - ranging from harmful content generation to broader societal harms - pose significant challenges. These risks can be amplified by the recent adversarial attacks, fine-tuning vulnerabilities, and the increasing deployment of LLMs in high-stakes environments. Existing safety-enhancing techniques, such as fine-tuning with human feedback or adversarial training, are still vulnerable as they address specific threats and often fail to generalize across unseen attacks, or require manual system-level defenses. This paper introduces REPBEND, a novel approach that fundamentally disrupts the representations underlying harmful behaviors in LLMs, offering a scalable solution to enhance (potentially inherent) safety. REPBEND brings the idea of activation steering - simple vector arithmetic for steering model's behavior during inference - to loss-based fine-tuning. Through extensive evaluation, REPBEND achieves state-of-the-art performance, outperforming prior methods such as Circuit Breaker, RMU, and NPO, with up to 95% reduction in attack success rates across diverse jailbreak benchmarks, all with negligible reduction in model usability and general capabilities.
Ashkan Yousefpour, Taeheon Kim, Ryan Sungmo Kwon, Seungbeen Lee, Wonje Jeung, Seungju Han 0002, Alvin Wan, Harrison Ngan, Youngjae Yu
ACL (1)2
2025 MSCoTDet: Language-Driven Multi-Modal Fusion for Improved Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection is attractive for around-the-clock applications due to the complementary information between RGB and thermal modalities. However, current models often fail to detect pedestrians in certain cases (e.g., thermal-obscured pedestrians), particularly due to the modality bias learned from statistically biased datasets. In this paper, we investigate how to mitigate modality bias in multispectral pedestrian detection using a Large Language Model (LLM). Accordingly, we design a Multispectral Chain-of-Thought (MSCoT) prompting strategy, which prompts the LLM to perform multispectral pedestrian detection. Moreover, we propose a novel Multispectral Chain-of-Thought Detection (MSCoTDet) framework that integrates MSCoT prompting into multispectral pedestrian detection. To this end, we design a Language-driven Multi-modal Fusion (LMF) strategy that enables fusing the outputs of MSCoT prompting with the detection results of vision-based multispectral pedestrian detection models. Extensive experiments validate that MSCoTDet effectively mitigates modality biases and improves multispectral pedestrian detection.
Taeheon Kim, Sangyun Chung, Damin Yeom, Youngjoon Yu, Hak Gu Kim, Yong Man Ro
IEEE Trans. Circuits Syst. Video Technol.1
2024 Causal Mode Multiplexer: A Novel Framework for Unbiased Multispectral Pedestrian Detection
abstract
RGBT multispectral pedestrian detection has emerged as a promising solution for safety-critical applications that require day/night operations. However, the modality bias problem remains unsolved as multispectral pedestrian detectors learn the statistical bias in datasets. Specifically, datasets in multispectral pedestrian detection mainly distribute between ROTO11R⋆T⋆ refers to the visibility (O/X) in each modality. Generally, ROTO refers to daytime images, and RXTO refers to nighttime images. ROTX refers to daytime images in obscured situations. (day) and RXTO (night) data; the majority of the pedestrian labels statistically co-occur with their thermal features. As a result, multispectral pedestrian detectors show poor generalization ability on examples beyond this statistical correlation, such as ROTX data. To address this problem, we propose a novel Causal Mode Multiplexer (CMM) framework that effectively learns the causalities between multispectral inputs and predictions. Moreover, we construct a new dataset (ROTX-MP) to evaluate modality bias in multispectral pedestrian detection. ROTX-MP mainly includes ROTX examples not presented in previous datasets. Extensive experiments demonstrate that our proposed CMM framework generalizes well on existing datasets (KAIST, CVC-14, FLIR) and the new ROTX-MP. Our code and dataset are available at: https://github.com/ssbin0914/Causal-Mode-Multiplexer.git.
Taeheon Kim, Sebin Shin, Youngjoon Yu, Hak Gu Kim, Yong Man Ro
CVPR1
2024 Deep Learning Framework for Semantic Change Detection in Urban Green Spaces Along With Overall Urban Areas
abstract
Urban green spaces, crucial for ecological balance, face global degradation from natural disasters and rapid urbanization. Manual deforestation monitoring is laborious, prompting a shift to remote sensing and bitemporal satellite imagery. Traditional change detection (CD) methods have limitations, but deep learning, especially in semantic CD, shows promise. This study addresses challenges in semantic CD techniques, advocating for comprehensive training on datasets covering both semantic change masks and binary change masks. We propose a novel semantic CD network for urban changes while additionally providing urban greenery increased and decreased regions, integrating deep bitemporal features with an encoder-decoder structure, Atrous spatial pyramid pooling, and a spatial attention module with parallel dilated convolutions. Quantitative assessment, especially with pre-trained VGG16 as a backbone and parallel convolutional layers, demonstrates the proposed method's superiority, showcasing substantial improvements in urban greenery CD alongside overall urban changes. The proposed method holds potential for monitoring climate change, rapid urbanization, and the impact of natural disasters on urban environments, particularly urban greenery.
Aisha Javed, Taeheon Kim, Changhui Lee, Youkyung Han
IGARSS2
2024 Vision-based motion prediction for construction workers safety in real-time multi-camera system
Yuntae Jeon, Dai Quoc Tran, Almo Senja Kulinan, Taeheon Kim, Minsoo Park, Seunghee Park
Adv. Eng. Informatics4
2023 Multispectral Invisible Coating: Laminated Visible-Thermal Physical Attack against Multispectral Object Detectors Using Transparent Low-E Films
abstract
Multispectral object detection plays a vital role in safety-critical vision systems that require an around-the-clock operation and encounter dynamic real-world situations(e.g., self-driving cars and autonomous surveillance systems). Despite its crucial competence in safety-related applications, its security against physical attacks is severely understudied. We investigate the vulnerability of multispectral detectors against physical attacks by proposing a new physical method: Multispectral Invisible Coating. Utilizing transparent Low-e films, we realize a laminated visible-thermal physical attack by attaching Low-e films over a visible attack printing. Moreover, we apply our physical method to manufacture a Multispectral Invisible Suit that hides persons from the multiple view angles of Multispectral detectors. To simulate our attack under various surveillance scenes, we constructed a large-scale multispectral pedestrian dataset which we will release in public. Extensive experiments show that our proposed method effectively attacks the state-of-the-art multispectral detector both in the digital space and the physical world.
Taeheon Kim, Youngjoon Yu, Yong Man Ro
AAAI1
2023 Image Registration Between Kompsat-3a Mid-Wave Infrared And Electric Optical Images Using Hybrid Pyramid Matching Method
abstract
Korean multi-purpose satellite 3A (KOMPSAT-3A) can acquire electric optical (EO) and mid-wave infrared (MIR) images. Since MIR and EO images provide different information, they can be used together to effectively observe various phenomena on the Earth's surface. However, geometric misalignments exist between the EO and MIR images as the difference in the positions of each sensor when acquiring the images. In this study, we propose a hybrid pyramid matching (HPM) method to conduct the image registration between heterogeneous EO and MIR images with different spatial and spectral characteristics. The HPM method extracts reliable tie points (TPs) by iteratively adjusting the location of local templates in pyramid image pairs. Then, the image registration is conducted using transformation matrix estimated based on the TPs. The HPM method achieved superior accuracy and performance at three different sites.
Taeheon Kim, Yerin Yun, Changhui Lee, Youkyung Han
IGARSS1
2023 Deep Learning-Based Cloud Detection in High-Resolution Satellite Imagery Using Various Open-Source Cloud Images
abstract
Cloud cover is a significant obstacle to use optical satellite imagery. Therefore, various studies have been proposed to accurately detect clouds and evaluate satellite image quality. In particular, with the advancement of deep learning technology, many cloud detection studies are being conducted. However, a large volume of high-quality data is required to develop an effective deep learning model training. Thus, in this study, we compare the performance of deep learning cloud detection models for according to the diversity of sensors and resolutions of training data. For conducting the study, five case dataset combinations were constructed and trained with HRNet (High-Resolution Network). The performance evaluation of the trained models was conducted using test images from the KOMPSAT and PlanetScope satellites. As a mean of achieving high cloud detection results, it was found that selecting and using high-quality data is more effective than simply increasing the number of training data.
Yerin Yun, Taeheon Kim, Changhui Lee, Youkyung Han
IGARSS2
2023 FMPR-Net: False Matching Point Removal Network for Very-High-Resolution Satellite Image Registration
abstract
Image registration is the most basic preprocessing method used to unify coordinates among multitemporal very-high-resolution (VHR) satellite images, thus allowing the acquisition of reliable data of the Earth’s surface. Although image registration requires multiple matching points (MP), false matching points (FMP) are included because of the similar spectral patterns and noise. However, removing FMPs from VHR satellite image pairs is challenging, especially when the images are directly affected by complex factors, such as shadow, relief displacement, and terrain shielding. Therefore, we propose a false matching point removal network (FMPR-net) based on deep learning to eliminate effectively the FMPs to improve registration accuracy. The training dataset is produced by a semi-automatic method. It involves the generation of image patch pairs based on a matching process of scale-invariant feature transform and the assignment of labels referring to the characteristics of true matching points (TMP) and FMPs. The FMPR-net is designed in a Siamese format consisting of two matching point deep feature extractors (MDFE). The architecture of the MDFE consists of one main network and three branch networks to achieve robust extraction of meaningful deep features describing the characteristics of MPs. The FMPR-net removes the FMPs using a true matching probability calculated based on the similarity between deep features. Experiments conducted on four pairs of VHR satellite images have demonstrated that the FMPR-net can effectively remove the FMPs. Consequently, accurate VHR satellite image registration is possible by reducing uncertainty caused by the FMPs.
Taeheon Kim, Yerin Yun, Changhui Lee, Francesca Bovolo, Youkyung Han
IEEE Trans. Geosci. Remote. Sens.1
2022 Map: Multispectral Adversarial Patch to Attack Person Detection
abstract
Recently, multispectral person detection has shown great performance in real world applications such as autonomous driving and security systems. However, the reliability of person detection against physical attacks has not been fully explored yet in multispectral person detectors. To evaluate the robustness of multispectral person detectors in the physical world, we propose a novel Multispectral Adversarial Patch (MAP) generation framework. MAP is optimized with a Cross-spectral Mapping(CSM) and Material Emissivity(ME) loss. This paper is the first to evaluate the reliability of a multispectral person detector against physical attack. Throughout experiment, our proposed adversarial patch successfully attacks the person detector and the Average Precision (AP) score is dropped by 90.79% in digital space and 73.34% in physical space.
Taeheon Kim, Hong Joo Lee 0001, Yong Man Ro
ICASSP1
2022 Image Registration of Very-High-Resolution Satellite Images Using Deep Learning Model for Outlier Elimination
abstract
Very-high-resolution (VHR) satellite image contains reliable various information over large areas, so that, it has been used as key data in the field of remote sensing. Image registration must be conducted to effectively use the multitemporal VHR satellite images. Conjugate points (CPs) extracted from the same region between images are required to perform image registration. However, outliers included in the CPs cause distortion when they were used for the image registration. Here we propose a deep learning-based technique to effectively remove the outliers. A Siamese network was built as a purpose of an outlier removal, and the network was trained using data based on the patch pirs centered on each CP. Experimental results demonstrate that the proposed method can remove outliers more effectively than a random sample consensus (RANSAC) technique thus and achieves improved registration accuracy.
Taeheon Kim, Yerin Yun, Changhui Lee, Junho Yeom, Youkyung Han
IGARSS1
2022 Building Impact Analysis for Very-High-Resolution Image Co-Registration
abstract
Since multi-temporal very-high-resolution (VHR) satellite images generally have geometric misalignment, image co-registration process is required to minimize it. To perform precise image co-registration, extraction of reliable conjugate points (CPs) is an important process. Moreover, CPs extracted from elevated objects can cause severe relief displacements according to acquisition angles of images. In this study, the effect of CPs extracted from buildings on co-registration performance was analyzed. To this end, CPs were extracted using a method that combines feature-based and area-based matching methods, and digital map was used to remove CPs extracted from the buildings. Root mean square errors (RMSE) was calculated using manually obtained checkpoints to evaluate the accuracy of co-registration according to the presence or absence of CPs extracted on buildings. When CPs extracted from buildings were removed, the RMSE of the checkpoints extracted from the dense-building area was improved by more than 4 pixels.
Jueon Park, Taeheon Kim, Aisha Javed, Changhui Lee, Youkyung Han
IGARSS2
2022 Defending Physical Adversarial Attack on Object Detection via Adversarial Patch-Feature Energy
abstract
Object detection plays an important role in security-critical systems such as autonomous vehicles but has shown to be vulnerable to adversarial patch attacks. Existing defense methods are restricted to localized noise patches by removing noisy regions in the input image. However, adversarial patches have developed into natural-looking patterns which evade existing defenses. To address this issue, we propose a defense method based on a novel concept "Adversarial Patch- Feature Energy" (APE) which exploits common deep feature characteristics of an adversarial patch. Our proposed defense consists of APE-masking and APE-refinement which can be employed to defend against any adversarial patch on literature. Extensive experiments demonstrate that APE-based defense achieves impressive robustness against adversarial patches both in the digital space and the physical world.
Taeheon Kim, Youngjoon Yu, Yong Man Ro
ACM Multimedia1
2016 Redirected head gaze to support AR meetings distributed over heterogeneous environments
abstract
We demonstrate a method for redirecting gaze of virtual avatars in distributed augmented reality (AR) meetings. As social cues are a necessity for effective communication, our method tries to preserve gaze awareness, one of the key elements of a face-to-face meeting. When using AR to bring multiple sites together in a distributed meeting, with different numbers of participants and physical arrangements across sites, gaze awareness is maintained regardless of the seating topology. By maintaining gaze, we hope to enhance the presence of remote attendees and improve communication among the users, making meetings in AR a practical option for teleconferencing.
Taeheon Kim, Ashwin Kachhara, Blair MacIntyre
VR1