EDBT 2026 Demo / reviewers in the wild / expert
Hong Joo Lee 0001
dblp:93/6986-1
· DBLP profile ↗
24ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0001-6626-5683ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Textual Compositional Reasoning for Robust Change CaptioningabstractChange captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and compositional semantics. To alleviate this, we present CORTEX (COmpositional Reasoning-aware TEXt-guided), a novel framework that integrates complementary textual cues to enhance change understanding. In addition to capturing cues from pixel-level differences, CORTEX utilizes scene-level textual knowledge provided by Vision Language Models (VLMs) to extract richer image text signals that reveal underlying compositional reasoning. CORTEX consists of three key modules: (i) an Image-level Change Detector that identifies low-level visual differences between paired images, (ii) a Reasoning-aware Text Extraction (RTE) module that use VLMs to generate compositional reasoning descriptions implicit in visual features, and (iii) an Image-Text Dual Alignment (ITDA) module that aligns visual and textual features for fine-grained relational reasoning. This enables CORTEX to reason over visual and textual features and capture changes that are otherwise ambiguous in visual features alone. Kyu Ri Park, Seong Tae Kim 0001, Hong Joo Lee 0001, Jung Uk Kim |
AAAI | 4 |
| 2026 | Where It Moves, It Matters: Referring Surgical Instrument Segmentation via MotionabstractEnabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexplored in surgical videos, with existing approaches struggling to generalize due to reliance on static visual cues and predefined instrument names. In this work, we introduce SurgRef, a novel motion-guided framework that grounds free-form language expressions in instrument motion, capturing how tools move and interact across time, rather than what they look like. This allows models to understand and segment instruments even under occlusion, ambiguity, or unfamiliar terminology. To train and evaluate SurgRef, we present Ref-IMotion, a diverse, multi-institutional video dataset with dense spatiotemporal masks and rich motion-centric expressions. SurgRef achieves state-of-the-art accuracy and generalization across surgical procedures, setting a new benchmark for robust, language-driven surgical video segmentation. Kun Yuan 0004, Long Bai 0008, Nassir Navab, Hongliang Ren 0001, Hong Joo Lee 0001, Tom Vercauteren, Nicolas Padoy |
AAAI | 8 |
| 2026 | Unsupervised domain adaptation for medical image segmentation using adaptogen-perturbationabstractDomains shift originated from differences in devices or patients in the medical field, poses a significant challenge when applying pre-trained models to clinical applications. To tackle this challenge, domain adaptation methods have been explored. However, most existing methods are designed for a single target domain adaptation or require sharing all target domain data for adaptation, which is infeasible in the medical field due to privacy issues. In this paper, we propose a novel unsupervised multi-target domain adaptation method without requiring data sharing. To this end, we introduce an additional signal, termed Adaptogen-Perturbation (AP) optimized to bridge the gap between the source and target domains. The optimized AP is injected into the latent feature and facilitates the adaptation of the pre-trained model to the target domain. Moreover, we propose a Spectral/Geometric Consistency learning framework to optimize the AP in an unsupervised manner. This promotes consistent predictions across two types of transformations: geometric and frequency-space spectral transformations, enhancing robustness to both variations. Extensive experiments with multiple medical segmentation datasets demonstrate the effectiveness of APs. Hong Joo Lee 0001, Yuan Bi, Sangmin Lee 0001, Gyeong-Moon Park, Jung Uk Kim, Seong Tae Kim 0001, Zhongliang Jiang, Nassir Navab |
Medical Image Anal. | 1 |
| 2026 | Adversarial Wear and Tear: Exploiting Natural Damage for Generating Physical-World Adversarial Examples
Samra Irshad, Seungkyu Lee 0001, Nassir Navab, Hong Joo Lee 0001, Seong Tae Kim 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | PRADA: Protecting and Detecting Dataset Abuse for Open-Source Medical Dataset
Jinhyeok Jang, Hong Joo Lee 0001, Nassir Navab, Seong Tae Kim 0001 |
MICCAI (14) | 2 |
| 2025 | Understanding adversarial robustness of deep neural networks via decision reliance
Soyoun Won, Hyeon Bae Kim, Yong Hyun Ahn, Hong Joo Lee 0001, Seong Tae Kim 0001 |
Image Vis. Comput. | 4 |
| 2024 | Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
Kyu Ri Park, Hong Joo Lee 0001, Jung Uk Kim |
ECCV (15) | 2 |
| 2024 | Defending Video Recognition Model Against Adversarial Perturbations via Defense PatternsabstractDeep Neural Networks (DNNs) have been widely successful in various domains, but they are vulnerable to adversarial attacks. Recent studies have also demonstrated that video recognition models are susceptible to adversarial perturbations, but the existing defense strategies in the image domain do not transfer well to the video domain due to the lack of considering temporal development and require a high computational cost for training video recognition models. This paper, first, investigates the temporal vulnerability of video recognition models by quantifying the effect of temporal perturbations on the model's performance. Based on these investigations, we propose Defense Patterns (DPs) that can effectively protect video recognition models by adding them to the input video frames. The DPs are generated on top of a pre-trained model, eliminating the need for retraining or fine-tuning, which significantly reduces the computational cost. Experimental results on two benchmark datasets and various action recognition models demonstrate the effectiveness of the proposed method in enhancing the robustness of video recognition models. Hong Joo Lee 0001, Yong Man Ro |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Advancing Adversarial Training by Injecting Booster SignalabstractRecent works have demonstrated that deep neural networks (DNNs) are highly vulnerable to adversarial attacks. To defend against adversarial attacks, many defense strategies have been proposed, among which adversarial training (AT) has been demonstrated to be the most effective strategy. However, it has been known that AT sometimes hurts natural accuracy. Then, many works focus on optimizing model parameters to handle the problem. Different from the previous approaches, in this article, we propose a new approach to improve the adversarial robustness using an external signal rather than model parameters. In the proposed method, a well-optimized universal external signal called a booster signal is injected into the outside of the image which does not overlap with the original content. Then, it boosts both adversarial robustness and natural accuracy. The booster signal is optimized in parallel to model parameters step by step collaboratively. Experimental results show that the booster signal can improve both the natural and robust accuracies over the recent state-of-the-art AT methods. Also, optimizing the booster signal is general and flexible enough to be adopted on any existing AT methods. Hong Joo Lee 0001, Youngjoon Yu, Yong Man Ro |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Robust Proxy: Improving Adversarial Robustness by Robust Proxy LearningabstractRecently, it has been widely known that deep neural networks are highly vulnerable and easily broken by adversarial attacks. To mitigate the adversarial vulnerability, many defense algorithms have been proposed. Recently, to improve adversarial robustness, many works try to enhance feature representation by imposing more direct supervision on the discriminative feature. However, existing approaches lack an understanding of learning adversarially robust feature representation. In this paper, we propose a novel training framework called Robust Proxy Learning. In the proposed method, the model explicitly learns robust feature representations with robust proxies. To this end, firstly, we demonstrate that we can generate class-representative robust features by adding class-wise robust perturbations. Then, we use the class representative features as robust proxies. With the class-wise robust features, the model explicitly learns adversarially robust features through the proposed robust proxy learning framework. Through extensive experiments, we verify that we can manually generate robust features, and our proposed learning framework could increase the robustness of the DNNs. Hong Joo Lee 0001, Yong Man Ro |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Map: Multispectral Adversarial Patch to Attack Person DetectionabstractRecently, multispectral person detection has shown great performance in real world applications such as autonomous driving and security systems. However, the reliability of person detection against physical attacks has not been fully explored yet in multispectral person detectors. To evaluate the robustness of multispectral person detectors in the physical world, we propose a novel Multispectral Adversarial Patch (MAP) generation framework. MAP is optimized with a Cross-spectral Mapping(CSM) and Material Emissivity(ME) loss. This paper is the first to evaluate the reliability of a multispectral person detector against physical attack. Throughout experiment, our proposed adversarial patch successfully attacks the person detector and the Average Precision (AP) score is dropped by 90.79% in digital space and 73.34% in physical space. Taeheon Kim, Hong Joo Lee 0001, Yong Man Ro |
ICASSP | 2 |
| 2022 | Defending Person Detection Against Adversarial Patch Attack by Using Universal Defensive FrameabstractPerson detection has attracted great attention in the computer vision area and is an imperative element in human-centric computer vision. Although the predictive performances of person detection networks have been improved dramatically, they are vulnerable to adversarial patch attacks. Changing the pixels in a restricted region can easily fool the person detection network in safety-critical applications such as autonomous driving and security systems. Despite the necessity of countering adversarial patch attacks, very few efforts have been dedicated to defending person detection against adversarial patch attack. In this paper, we propose a novel defense strategy that defends against an adversarial patch attack by optimizing a defensive frame for person detection. The defensive frame alleviates the effect of the adversarial patch while maintaining person detection performance with clean person. The proposed defensive frame in the person detection is generated with a competitive learning algorithm which makes an iterative competition between detection threatening module and detection shielding module in person detection. Comprehensive experimental results demonstrate that the proposed method effectively defends person detection against adversarial patch attacks. Youngjoon Yu, Hong Joo Lee 0001, Hakmin Lee, Yong Man Ro |
IEEE Trans. Image Process. | 2 |
| 2021 | Towards Robust Training of Multi-Sensor Data Fusion Network Against Adversarial Examples in Semantic SegmentationabstractThe success of multi-sensor data fusions in deep learning appears to be attributed to the use of complementary information among multiple sensor datasets. Compared to their predictive performance, relatively less attention has been devoted to the adversarial robustness of multi-sensor data fusion models. To achieve adversarial robust multi-sensor data fusion networks, we propose here a novel robust training scheme called Multi-Sensor Cumulative Learning (MSCL). The motivation behind the MSCL method is based on the way human beings learn new skills. The MSCL allows the multi-sensor fusion network to learn robust features from individual sensors, and then learn complex joint features from multiple sensors just as people learn to walk before they run. The step wise framework of MSCL enables the network to incorporate pre-trained knowledge of robustness with new joint information from multiple sensors. Extensive experimental evidence validated that the MSCL outperforms other multi-sensor fusion training in defending against adversarial examples. Youngjoon Yu, Hong Joo Lee 0001, Byeong Cheon Kim, Jung Uk Kim, Yong Man Ro |
ICASSP | 2 |
| 2021 | Adversarially Robust Multi-Sensor Fusion Model Training Via Random Feature Fusion For Semantic SegmentationabstractMulti-sensor data fusion model aims to improve the model performance by fusing multiple types of sensor data. Although multi-sensor data fusion models have been developed for remarkable performance, there is a lack of studies on the adversarial vulnerability of the multi-sensor data fusion models. In this paper, we propose a robust multi-sensor data fusion method that is not vulnerable to adversarial attacks. To this end, we devise a random feature fusion method to preserve multi-sensor fusion features. Through the random feature fusion, we could explicitly hide the information about which features are being used for the fusion. In experiments, we verify that our proposed random feature fusion method shows the adversarial robustness considerably under diverse adversarial settings. Hong Joo Lee 0001, Yong Man Ro |
ICIP | 1 |
| 2021 | CUA Loss: Class Uncertainty-Aware Gradient Modulation for Robust Object DetectionabstractRecently, a wide range of research on object detection has shown breakthrough performance. However, in a challenging environment, such as occlusion and small object cases, object detectors still produce inaccurate or erroneous predictions. To effectively cope with such conditions, most of the existing methods have suggested loss functions to guide the object detectors by modulating the magnitude of their loss. However, when modulating the loss function, they are highly dependent on the classification score of the object detector. It is a known fact that deep neural networks tend to be overconfident in their predictions. In this article, to alleviate the problem of the object detectors which heavily rely on the prediction in the training phase, we devise a novel loss function called class uncertainty-aware (CUA) loss. CUA loss considers the predictive ambiguity as well as the predictions on classification score when modulating loss function. In addition to the classification score, CUA loss further modulates the loss gradient in an increasing way when the object detectors output an uncertain prediction. Therefore, object detectors with CUA loss effectively cope with challenging environments where prediction results are uncertain. With comprehensive experiments on three public datasets (i.e. PASCAL VOC, MS COCO, and Berkeley DeepDrive), we verified that our CUA loss enhanced the accuracy of the object detectors and outperformed previous state-of-the-art loss functions. Jung Uk Kim, Seong Tae Kim 0001, Hong Joo Lee 0001, Sangmin Lee 0001, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Robust Ensemble Model Training via Random Layer Sampling Against Adversarial Attack
Hakmin Lee, Hong Joo Lee 0001, Seong Tae Kim 0001, Yong Man Ro |
BMVC | 2 |
| 2020 | Structure Boundary Preserving Segmentation for Medical Image With Ambiguous BoundaryabstractIn this paper, we propose a novel image segmentation method to tackle two critical problems of medical image, which are (i) ambiguity of structure boundary in the medical image domain and (ii) uncertainty of the segmented region without specialized domain knowledge. To solve those two problems in automatic medical segmentation, we propose a novel structure boundary preserving segmentation framework. To this end, the boundary key point selection algorithm is proposed. In the proposed algorithm, the key points on the structural boundary of the target object are estimated. Then, a boundary preserving block (BPB) with the boundary key point map is applied for predicting the structure boundary of the target object. Further, for embedding experts' knowledge in the fully automatic segmentation, we propose a novel shape boundary-aware evaluator (SBE) with the ground-truth structure information indicated by experts. The proposed SBE could give feedback to the segmentation network based on the structure boundary key point. The proposed method is general and flexible enough to be built on top of any deep learning-based segmentation network. We demonstrate that the proposed method could surpass the state-of-the-art segmentation network and improve the accuracy of three different segmentation network models on different types of medical image datasets. Hong Joo Lee 0001, Jung Uk Kim, Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro |
CVPR | 1 |
| 2020 | Fake Video Detection With Certainty-Based Attention NetworkabstractDeepFake synthesizes realistic fake videos that could be used maliciously such as manipulation and harassment. In order to prevent such malicious usages, detecting fake videos is immediately needed. In this paper, we propose a novel fake video detection method by adopting predictive uncertainty in detection. We devise the certainty-based attention network which guides to focus certainty-key frames in detecting fake videos. In addition, certainty-based attention is proposed for refining the features with consideration for frame-level certainty. Experiments are performed to validate the effectiveness of the proposed method by comparing the existing methods on Celeb-DF, the latest DeepFake dataset. Dae Hwi Choi, Hong Joo Lee 0001, Sangmin Lee 0001, Jung Uk Kim, Yong Man Ro |
ICIP | 2 |
| 2020 | Learning Style Correlation for Elaborate Few-Shot ClassificationabstractFew-shot classification is defined as a task where the network aims to classify unseen classes given only a few samples. Recent approaches, especially metric-based methods, have great progress in few-shot classification. However, the existing metric-based methods have a limitation in deploying discriminative features for elaborate comparison. They usually extract features from the embedding network without direct consideration of the relationship between support and query sets. To address the relationship, we propose a novel architecture, Style Correlated Module (SCM) to learn style correlation between support and query sets for few-shot classification. The proposed module leads support and query feature maps to focus on significant style correlated features and encourage the metric network to conduct an elaborate comparison. Furthermore, the proposed module can be generally applied to the existing metric-based approaches by adding the SCM behind the embedding network. We evaluate our proposed method with comprehensive experiments on two publicly available datasets and demonstrate its effectiveness with comparable results. Minsu Kim 0001, Jung Uk Kim, Hong Joo Lee 0001, Sangmin Lee 0001, Joanna Hong, Yong Man Ro |
ICIP | 4 |
| 2020 | Robust Video Facial Authentication With Unsupervised Mode DisentanglementabstractDeep learning-based video facial authentication has limitations when it comes to real-world applications, due to large mode variations such as illumination, pose, and eyeglasses variations in real-life situations. Many of existing mode-invariant facial authentication methods need labels of each mode. However, the label information could not be always available in practice. To alleviate this problem, we develop an unsupervised mode disentangling method for video facial authentication. By matching both disentangled identity features and dynamic features of two facial videos, our proposed method shows significant face verification and identification performances on three publicly available datasets, KAIST-MPMI, UVA-NEMO, and YTF. Minsu Kim 0001, Hong Joo Lee 0001, Sangmin Lee 0001, Yong Man Ro |
ICIP | 2 |
| 2020 | Unsupervised Disentangling of Viewpoint and Residues Variations by Substituting Representations for Robust Face RecognitionabstractIt is well-known that identity-unrelated variations (e.g., viewpoint or illumination) degrade the performances of face recognition methods. In order to handle this challenge, a robust method for disentangling the identity and view representations has drawn an attention in the machine learning area. However, existing methods learn discriminative features which require a manual supervision of such factors of variations. In this paper, we propose a novel disentangling framework through modeling three representations of identity, viewpoint, and residues (i.e., identity and pose unrelated) which do not require supervision of the variations. By jointly modeling the three representations, we enhance the disentanglement of each representation and achieve robust face recognition performance. Further, the learned viewpoint representation can be utilized for pose estimation or editing of a posed facial image. Extensive quantitative and qualitative evaluations verify the effectiveness of our proposed method which disentangles identity, viewpoint, and residues of facial images. Minsu Kim 0001, Joanna Hong, Hong Joo Lee 0001, Yong Man Ro |
ICPR | 4 |
| 2020 | Face Tells Detailed Expression: Generating Comprehensive Facial Expression Sentence Through Facial Action Units
Joanna Hong, Hong Joo Lee 0001, Yelin Kim, Yong Man Ro |
MMM (2) | 2 |
| 2020 | Lightweight and Effective Facial Landmark Detection using Adversarial Learning with Face Geometric Map Generative NetworkabstractFacial landmark detection plays an important role in face analysis tasks. Moreover, it is used as a prerequisite in many facial related applications, the simplicity, as well as effectiveness, is essential in the facial landmark detection. In this paper, we propose an effective facial landmark detection network and an associated learning framework with the geometric prior-generative adversarial network. The geometric prior-generative adversarial network consists of one generator and two discriminators. The generator consists of an encoder and two decoders. The encoder predicts facial landmark points. The decoders generate a facial inner and contour geometric map from predicted landmark points. Generating face geometric maps from predicted landmark points helps the predicted landmark points to represent the face geometric information, including shape and configuration. The discriminators determine that the given geometric maps are generated from actual landmark points or estimated landmark points. Our proposed network is end-to-end trainable, and only the encoder part is used simply as the facial landmark detector in the testing stage. To verify the effectiveness of the proposed method, we have conducted comprehensive experiments with benchmark data sets. The results have shown that the proposed method achieves comparable performances over recently proposed facial landmark detection methods with a simple and effective facial landmark detection network. Hong Joo Lee 0001, Seong Tae Kim 0001, Hakmin Lee, Yong Man Ro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Teacher and Student Joint Learning for Compact Facial Landmark Detection Network
Hong Joo Lee 0001, Wissam J. Baddar, Hak Gu Kim, Seong Tae Kim 0001, Yong Man Ro |
MMM (1) | 1 |