Jung Uk Kim

dblp:205/3068 · DBLP profile ↗
← Back
44ranked-venue papers
13as first author
32since 2021 · last 2026
0000-0003-4533-4875ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 12 first-author · 28 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Do We Need Perfect Data? Leveraging Noise for Domain Generalized Segmentation
abstract
Domain generalization in semantic segmentation faces challenges from domain shifts, particularly under adverse conditions. While diffusion-based data generation methods show promise, they introduce inherent misalignment between generated images and semantic masks. This paper presents FLEX-Seg (FLexible Edge eXploitation for Segmentation), a framework that transforms this limitation into an opportunity for robust learning. FLEX-Seg comprises three key components: (1) Granular Adaptive Prototypes that captures boundary characteristics across multiple scales, (2) Uncertainty Boundary Emphasis that dynamically adjusts learning emphasis based on prediction entropy, and (3) Hardness-Aware Sampling that progressively focuses on challenging examples. By leveraging inherent misalignment rather than enforcing strict alignment, FLEX-Seg learns robust representations while capturing rich stylistic variations. Experiments across five real-world datasets demonstrate consistent improvements over state-of-the-art methods, achieving 2.44% and 2.63% mIoU gains on ACDC and Dark Zurich. Our findings validate that adaptive strategies for handling imperfect synthetic data lead to superior domain generalization.
Taeyeong Kim, SeungJoon Lee, Jung Uk Kim, MyeongAh Cho
AAAI3
2026 See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
abstract
Video moment retrieval (MR) and highlight detection (HD) with natural language queries aim to localize relevant moments and key highlights in a video clips. However, existing methods overlook the importance of individual words, treating the entire text query and video clips as a black-box, which hinders contextual understanding. In this paper, we propose a novel approach that enables fine-grained clip filtering by identifying and prioritizing important words in the query. Our method integrates image-text scene understanding through Multimodal Large Language Models (MLLMs) and enhances the semantic understanding of video clips. We introduce a feature enhancement module (FEM) to capture important words from the query and a ranking-based filtering module (RFM) to iteratively refine video clips based on their relevance to these important words. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art methods, achieving superior performance in both MR and HD tasks.
YuEun Lee, Jung Uk Kim
AAAI2
2026 Task Prototype-Based Knowledge Retrieval for Multi-Task Learning from Partially Annotated Data
abstract
Multi-task learning (MTL) is critical in real-world applications such as autonomous driving and robotics, enabling simultaneous handling of diverse tasks. However, obtaining fully annotated data for all tasks is impractical due to labeling costs. Existing methods for partially labeled MTL typically rely on predictions from unlabeled tasks, making it difficult to establish reliable task associations and potentially leading to negative transfer and suboptimal performance. To address these issues, we propose a prototype-based knowledge retrieval framework that achieves robust MTL instead of relying on predictions from unlabeled tasks. Our framework consists of two key components: (1) a task prototype embedding task-specific characteristics and quantifying task associations, and (2) a knowledge retrieval transformer that adaptively refines feature representations based on these associations. To achieve this, we introduce an association knowledge generating (AKG) loss to ensure the task prototype consistently captures task-specific characteristics. Extensive experiments demonstrate the effectiveness of our framework, highlighting its potential for robust multi-task learning, even when only a subset of tasks is annotated.
Youngmin Oh 0003, Hyungil Kim, Jung Uk Kim
AAAI3
2026 Leveraging Textual Compositional Reasoning for Robust Change Captioning
abstract
Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and compositional semantics. To alleviate this, we present CORTEX (COmpositional Reasoning-aware TEXt-guided), a novel framework that integrates complementary textual cues to enhance change understanding. In addition to capturing cues from pixel-level differences, CORTEX utilizes scene-level textual knowledge provided by Vision Language Models (VLMs) to extract richer image text signals that reveal underlying compositional reasoning. CORTEX consists of three key modules: (i) an Image-level Change Detector that identifies low-level visual differences between paired images, (ii) a Reasoning-aware Text Extraction (RTE) module that use VLMs to generate compositional reasoning descriptions implicit in visual features, and (iii) an Image-Text Dual Alignment (ITDA) module that aligns visual and textual features for fine-grained relational reasoning. This enables CORTEX to reason over visual and textual features and capture changes that are otherwise ambiguous in visual features alone.
Kyu Ri Park, Seong Tae Kim 0001, Hong Joo Lee 0001, Jung Uk Kim
AAAI5
2026 Unsupervised domain adaptation for medical image segmentation using adaptogen-perturbation
abstract
Domains shift originated from differences in devices or patients in the medical field, poses a significant challenge when applying pre-trained models to clinical applications. To tackle this challenge, domain adaptation methods have been explored. However, most existing methods are designed for a single target domain adaptation or require sharing all target domain data for adaptation, which is infeasible in the medical field due to privacy issues. In this paper, we propose a novel unsupervised multi-target domain adaptation method without requiring data sharing. To this end, we introduce an additional signal, termed Adaptogen-Perturbation (AP) optimized to bridge the gap between the source and target domains. The optimized AP is injected into the latent feature and facilitates the adaptation of the pre-trained model to the target domain. Moreover, we propose a Spectral/Geometric Consistency learning framework to optimize the AP in an unsupervised manner. This promotes consistent predictions across two types of transformations: geometric and frequency-space spectral transformations, enhancing robustness to both variations. Extensive experiments with multiple medical segmentation datasets demonstrate the effectiveness of APs.
Hong Joo Lee 0001, Yuan Bi, Sangmin Lee 0001, Gyeong-Moon Park, Jung Uk Kim, Seong Tae Kim 0001, Zhongliang Jiang, Nassir Navab
Medical Image Anal.5
2026 Adverse Weather Removal via Dynamic Enhancement Diffusion With Weather-Adaptive Prompting
Youngmin Oh 0003, Sungyoung Lee 0001, MyeongAh Cho, Jung Uk Kim
IEEE Trans. Image Process.4
2026 SSMPD: Semi-Supervised Learning for Multispectral Pedestrian Detection
abstract
Pedestrian detection is a crucial task in computer vision. Utilizing multispectral knowledge, especially, is essential to effectively detect the pedestrians. Existing multispectral pedestrian detection methods, however, perform only in fully-supervised situations. Although studies on semi-supervised object detection have been conducted, they focus only on single modality environments. Therefore, we propose novel semi-supervised multispectral pedestrian detector (SSMPD) that effectively utilizes multispectral knowledge. Our SSMPD consists of three methods that effectively address the pseudo-labels in the multispectral domain and a novel data selection method. First, we introduce a Pedestrian Appearance-Aware (PAA) weight to consider the quality of the pseudo-label by adjusting the multispectral knowledge transfer from the teacher model to the student model. Second, we propose a Unified Modal-Aware Simultaneous (UMAS) learning to consider the single modality (visible or thermal) and multispectral modalities when learning with the pseudo-label. Finally, we introduce a Similarity-based Contrastive (SC) loss to guide the teacher model in enhancing the quality of pseudo-labels. In addition, we provide diverse data selection for more effective semi-supervised learning. Extensive experimental results on the KAIST and LLVIP datasets demonstrate the effectiveness of our method.
Seungho Shin, Gyeong-Moon Park, Jung Uk Kim
IEEE Trans. Multim.4
2025 Multispectral Pedestrian Detection with Sparsely Annotated Label
abstract
Although existing Sparsely Annotated Object Detection (SAOD) approches have made progress in handling sparsely annotated environments in multispectral domain, where only some pedestrians are annotated, they still have the following limitations: (i) they lack considerations for improving the quality of pseudo-labels for missing annotations, and (ii) they rely on fixed ground truth annotations, which leads to learning only a limited range of pedestrian visual appearances in the multispectral domain. To address these issues, we propose a novel framework called Sparsely Annotated Multispectral Pedestrian Detection (SAMPD). For limitation (i), we introduce Multispectral Pedestrian-aware Adaptive Weight (MPAW) and Positive Pseudo-label Enhancement (PPE) module. Utilizing multispectral knowledge, these modules ensure the generation of high-quality pseudo-labels and enable effective learning by increasing weights for high-quality pseudo-labels based on modality characteristics. To address limitation (ii), we propose an Adaptive Pedestrian Retrieval Augmentation (APRA) module, which adaptively incorporates pedestrian patches from ground-truth and dynamically integrates high-quality pseudo-labels with the ground-truth, facilitating a more diverse learning pool of pedestrians. Extensive experimental results demonstrate that our SAMPD significantly enhances performance in sparsely annotated environments within the multispectral domain.
Seungho Shin, Gyeong-Moon Park, Jung Uk Kim
AAAI4
2025 Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
abstract
The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed both simultaneously. However, they still struggle to fully capture the overall video context, making it challenging to determine which words are most relevant. In this paper, we present a novel Video Context-aware Keyword Attention module that overcomes this limitation by capturing keyword variation within the context of the entire video. To achieve this, we introduce a video context clustering module that provides concise representations of the overall video context, thereby enhancing the understanding of keyword dynamics. Furthermore, we propose a keyword weight detection module with keyword-aware contrastive learning that incorporates keyword information to enhance fine-grained alignment between visual and textual features. Extensive experiments on the QVHighlights, TVSum, and Charades-STA benchmarks demonstrate that our proposed method significantly improves performance in moment retrieval and highlight detection tasks compared to existing approaches.
Sung Jin Um, Sangmin Lee 0001, Jung Uk Kim
AAAI4
2025 Object-aware Sound Source Localization via Audio-Visual Scene Understanding
abstract
Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objects coexist. This limitation arises primarily from their reliance on simple audio-visual correspondence, which does not capture fine-grained semantic differences between sound-making and silent objects. To address these challenges, we propose a novel sound source localization framework leveraging Multimodal Large Language Models (MLLMs) to generate detailed contextual information that explicitly distinguishes between sound-making foreground objects and silent background objects. To effectively integrate this detailed information, we introduce two novel loss functions: Object-aware Contrastive Alignment (OCA) loss and Object Region Isolation (ORI) loss. Extensive experimental results on MUSIC and VGGSound datasets demonstrate the effectiveness of our approach, significantly outperforming existing methods in both single-source and multi-source localization scenarios. Code and generated detailed contextual information are available at: https://github.com/VisualAIKHU/OA-SSL.
Sung Jin Um, Sangmin Lee 0001, Jung Uk Kim
CVPR4
2025 Unified link prediction modeling for enhanced knowledge graph completion task
Tri D. T. Nguyen, Ubaid Ur Rehman 0002, Musarrat Hussain, Rao Faizan, Jamil Hussain, Sung-Ho Bae, Jung Uk Kim, Seong Tae Kim 0001, Sungyoung Lee 0001
Expert Syst. Appl.7
2025 Spatial Mask-Based Adaptive Robust Training for Video Object Segmentation With Noisy Labels
abstract
Recent advances in video object segmentation (VOS) highlight its potential across various applications. Semi-supervised VOS aims to segment target objects in video frames based on annotations from the initial frame. Collecting a large-scale video segmentation dataset is challenging, which could induce noisy labels. However, it has been overlooked and most of the research efforts have been devoted to training VOS models by assuming the training dataset is clean. In this study, we first explore the effect of VOS models under noisy labels in the training dataset. To investigate the effect of noisy labels, we simulate the noisy annotations on DAVIS 2017 and YouTubeVOS datasets. Experiments show that the traditional training strategy is vulnerable to noisy annotations. To address this issue, we propose a novel noise-robust training method, named SMART (Spatial Mask-based Adaptive Robust Training), which is designed to train models effectively in the presence of noisy annotations. The proposed method employs two key strategies. Firstly, the model focuses on the common spatial areas from clean knowledge-based predictions and annotations. Secondly, the model is trained with adaptive balancing losses based on their reliability. Comparative experiments have demonstrated the effectiveness of our approach by outperforming other noise handling methods over various noise degrees.
Enki Cho, Jung Uk Kim, Seong Tae Kim 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Enabling Visual Object Detection With Object Sounds via Visual Modality Recalling Memory
abstract
When humans hear the sound of an object, they recall associated visual information and integrate the sound with recalled visual modality to detect the object. In this article, we present a novel sound-based object detector that mimics this process. We design a visual modality recalling (VMR) memory to recall information of a visual modality based on an audio modal input (i.e., sound). To achieve this goal, we propose a VMR loss and an audio-visual association loss to guide the VMR memory to memorize visual modal information by establishing associations between audio and visual modalities. With the visual modal information recalled through the VMR memory along with the original audio input, we perform audio-visual integration. In this step, we introduce an integrated feature contrastive loss that allows the integrated feature to be embedded as if it were encoded using both audio and visual modal inputs. This guidance enables our sound-based object detector to effectively perform visual object detection even when only sound is provided. We believe that our work is a cornerstone study that offers a new perspective to conventional object detection studies that solely rely on the visual modality. Comprehensive experimental results demonstrate the effectiveness of the proposed method with the VMR memory.
Jung Uk Kim, Yong Man Ro
IEEE Trans. Neural Networks Learn. Syst.1
2024 Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
abstract
The goal of the multi-sound source localization task is to localize sound sources from the mixture individu-ally. While recent multi-sound source localization meth-ods have shown improved performance, they face chal-lenges due to their reliance on prior information about the number of objects to be separated. In this paper, to overcome this limitation, we present a novel multi-sound source localization method that can perform localization without prior knowledge of the number of sound sources. To achieve this goal, we propose an iterative object iden-tification (101) module, which can recognize sound-making objects in an iterative manner. After finding the regions of sound-making objects, we devise object similarity-aware clustering (OSC) loss to guide the 101 module to effectively combine regions of the same object but also dis-tinguish between different objects and backgrounds. It enables our method to perform accurate localization of sound-making objects without any prior knowledge. Exten-sive experimental results on the MUSIC and VGGSound benchmarks show the significant performance improve-ments of the proposed method over the existing methods for both single and multi-source. Our code is available at: https://github.comNisuaIAIKHUINoPrior_MultiSSL.
Sung Jin Um, Sangmin Lee 0001, Jung Uk Kim
CVPR4
2024 Towards Model-Agnostic Dataset Condensation by Heterogeneous Models
Jun-Yeong Moon, Jung Uk Kim, Gyeong-Moon Park
ECCV (29)2
2024 MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
Youngmin Oh 0003, Hyungil Kim, Seong Tae Kim 0001, Jung Uk Kim
ECCV (10)4
2024 Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
Kyu Ri Park, Hong Joo Lee 0001, Jung Uk Kim
ECCV (15)3
2024 Enhancing Audio-Visual Question Answering with Missing Modality via Trans-Modal Associative Learning
abstract
We present a novel method for Audio-Visual Question Answering (AVQA) in real-world scenarios where one modality (audio or visual) can be missing. Inspired by human cognitive processes, we introduce a Trans-Modal Associative (TMA) memory that recalls missing modal information (i.e., pseudo modal feature) by establishing associations between available modal features and textual cues. During training phase, we employ a Trans-Modal Recalling (TMR) loss to guide the TMA memory in generating the pseudo modal feature that closely matches the real modal feature. This allows our method to robustly answer the question, even when one modality is missing during inference. We believe that our approach, which effectively copes with missing modalities, can be broadly applied to a variety of multimodal applications.
Kyu Ri Park, Youngmin Oh 0003, Jung Uk Kim
ICASSP3
2024 Chain-of-Factors: A Zero-Shot Prompting Methodology Enabling Factor-Centric Reasoning in Large Language Models
abstract
Large language models (LLMs) have significantly improved numerous natural language processing tasks. However, their performance relies heavily on the provided instructions or prompts. Recently, several prompting methodologies have been developed to enhance the reasoning abilities of LLMs. Notably, the Chain-of-Thought (CoT) approach provides examples that help break down tasks into sub-steps, resulting in more accurate solutions. However, the process of generating detailed examples may not be user-friendly, as end users prefer providing task descriptions rather than a set of examples. In this study, we introduce Chain-of-Factors (CoF), an innovative zero-shot prompting methodology that incorporates task-specific instructions as a chain of factors into the prompt, aimed at enhancing the factor-centric reasoning abilities of LLMs. Experiments on three LLMs, including ChatGPT-3.5, Gemini, and GPT-4, show performance improvements ranging from 0.01% to 40.2% in accuracy on various symbolic reasoning and logical reasoning tasks compared with zero-shot and few-shot CoT. In summary, CoF enhances LLMs' reasoning abilities by including task-specific steps and instructions, while also decreasing the necessity for fine-tuning specific to each task.
Musarrat Hussain, Ubaid Ur Rehman 0002, Tri D. T. Nguyen, Sungyoung Lee 0001, Seong Tae Kim 0001, Sung-Ho Bae, Jung Uk Kim
ICMLA7
2023 Towards Robust Audio-Based Vehicle Detection Via Importance-Aware Audio-Visual Learning
abstract
Although audio modality has the potential to solve various visually challenging conditions of visual modality, there are few studies on audio-based detection. This is because the audio modality itself contains less accurate spatial information. To alleviate this issue, the existing audio-based methods adopt the visual modality in the training phase to transfer more precise spatial knowledge to the audio modality. However, they do not consider the case where the visual modality is less informative. In this paper, we present a new audio-based vehicle detector that can transfer multimodal knowledge of vehicles to the audio modality during training. To this end, we combine the audio-visual modal knowledge according to the importance of each modality to generate integrated audiovisual feature. Also, we introduce an audio-visual distillation (AVD) loss that guides representation of the audio modal feature to resemble that of the integrated audio-visual feature. As a result, our audio-based detector can perform robust vehicle detection as if it were utilizing both modalities, even if it only receives audio modality as input in the inference. Comprehensive experimental results demonstrate that our method exhibits consistent improvements over the existing methods.
Jung Uk Kim, Seong Tae Kim 0001
ICASSP1
2023 Similarity Relation Preserving Cross-Modal Learning for Multispectral Pedestrian Detection Against Adversarial Attacks
abstract
Although multispectral pedestrian detection studies have shown remarkable detection performances, they are still vulnerable to adversarial attacks. We see the similarity relations between object candidates were not maintained because of the adversarial attacks, resulting in performance degradation. In this paper, we introduce a new method that can preserve the similarity relation between candidates against adversarial attacks using multispectral knowledge. First, we propose Similarity Relation Generation (SRG) module to generate the optimal similarity relation between clean candidates by referring to the two modalities (color and thermal). Second, we propose Adversarial Similarity Relation Preserving (ASRP) module to guide the similarity relation between adversarial candidates to be similar to that of the clean candidates. By maintaining the relationship between candidates, our multispectral detector can distinguish between pedestrian/background classes even in adversarial attacks. Comprehensive experimental results show that our method conspicuously improves the adversarial robustness.
Jung Uk Kim, Yong Man Ro
ICASSP1
2023 Online Class Incremental Learning on Stochastic Blurry Task Boundary via Mask and Visual Prompt Tuning
abstract
Continual learning aims to learn a model from a continuous stream of data, but it mainly assumes a fixed number of data and tasks with clear task boundaries. However, in real-world scenarios, the number of input data and tasks is constantly changing in a statistical way, not a static way. Although recently introduced incremental learning scenarios having blurry task boundaries somewhat address the above issues, they still do not fully reflect the statistical properties of real-world situations because of the fixed ratio of disjoint and blurry samples. In this paper, we propose a new Stochastic incremental Blurry task boundary scenario, called Si-Blurry, which reflects the stochastic properties of the real-world. We find that there are two major challenges in the Si-Blurry scenario: (1) intra- and inter-task forget-tings and (2) class imbalance problem. To alleviate them, we introduce Mask and Visual Prompt tuning (MVP). In MVP, to address the intra- and inter-task forgetting issues, we propose a novel instance-wise logit masking and contrastive visual prompt tuning loss. Both of them help our model discern the classes to be learned in the current batch. It results in consolidating the previous knowledge. In addition, to alleviate the class imbalance problem, we introduce a new gradient similarity-based focal loss and adaptive feature scaling to ease overfitting to the major classes and underfitting to the minor classes. Extensive experiments show that our proposed MVP significantly outperforms the existing state-of-the-art methods in our challenging Si-Blurry scenario. The code is available at https://github.com/moonjunyyy/Si-Blurry
Jun-Yeong Moon, Keon-Hee Park, Jung Uk Kim, Gyeong-Moon Park
ICCV3
2023 Robust Multispectral Pedestrian Detection Via Spectral Position-Free Feature Mapping
abstract
Recently, although multispectral pedestrian detection has achieved remarkable performances, there is still a problem to be handled, position shift problem. Due to the problem, a pedestrian looks like existing in different positions between each modal image. Then, a single bounding box usually fails to capture an entire pedestrian properly in both modal images at the same time, which means it would not contain some parts of a pedestrian and includes noisy backgrounds instead. In this paper, we propose a novel approach, that is, a pedestrian feature mapping from mis-captured pedestrian features to well-captured pedestrian features which encode an entire pedestrian properly in both modal images. To this end, we utilize a memory architecture which stores well-captured pedestrian features, and then, the well-captured features can enhance the quality of pedestrian representation by providing the distinctive information of a pedestrian. We validate the effectiveness of our approach with comprehensive experiments on two multispectral pedestrian detection datasets, achieving state-of-the-art performances.
Sungjune Park, Jung Uk Kim, Jin Mo Song, Yong Man Ro
ICIP2
2023 Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source Localization
abstract
The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches only use audio as an auxiliary role to compare spatial regions of the visual modality. Humans, on the other hand, utilize both audio and visual modalities as spatial cues to locate sound sources. In this paper, we propose an audio-visual spatial integration network that integrates spatial cues from both modalities to mimic human behavior when detecting sound-making objects. Additionally, we introduce a recursive attention network to mimic human behavior of iterative focusing on objects, resulting in more accurate attention regions. To effectively encode spatial information from both modalities, we propose audio-visual pair matching loss and spatial region alignment loss. By utilizing the spatial cues of audio-visual modalities and recursively focusing objects, our method can perform more robust sound source localization. Comprehensive experimental results on the Flickr SoundNet and VGG-Sound Source datasets demonstrate the superiority of our proposed method over existing approaches. Our code is available at: https://github.com/VisualAIKHU/SIRA-SSL.
Sung Jin Um, Jung Uk Kim
ACM Multimedia3
2023 Stereoscopic Vision Recalling Memory for Monocular 3D Object Detection
abstract
Monocular 3D object detection has drawn increasing attention in various human-related applications, such as autonomous vehicles, due to its cost-effective property. On the other hand, a monocular image alone inherently contains insufficient information to infer the 3D information. In this paper, we propose a new monocular 3D object detector that can recall the stereoscopic visual information about an object, given a left-view monocular image. Here, we devise a location embedding module to handle each object by being aware of its location. Next, given the object appearance of the left-view monocular image, we devise Monocular-to-Stereoscopic (M2S) memory that can recall the object appearance of the right-view and depth information. For this purpose, we introduce a stereoscopic vision memorizing loss that guides the M2S memory to store the stereoscopic visual information. Furthermore, we propose a binocular vision association loss to guide the M2S memory that can associate the information of the left-right view about the object when estimating the depth. As a result, our monocular 3D object detector with the M2S memory can effectively exploit the recalled stereoscopic visual information in the inference phase. The comprehensive experimental results on two public datasets, KITTI 3D Object Detection Benchmark and Waymo Open Dataset, demonstrate the effectiveness of the proposed method. We claim that our method is a step-forward method that follows the behaviors of humans that can recall the stereoscopic visual information even when one eye is closed.
Jung Uk Kim, Hyungil Kim, Yong Man Ro
IEEE Trans. Image Process.1
2022 Towards Versatile Pedestrian Detector with Multisensory-Matching and Multispectral Recalling Memory
abstract
Recently, automated surveillance cameras can change a visible sensor and a thermal sensor for all-day operation. However, existing single-modal pedestrian detectors mainly focus on detecting pedestrians in only one specific modality (i.e., visible or thermal), so they cannot cope with other modal inputs. In addition, recent multispectral pedestrian detectors have shown remarkable performance by adopting multispectral modalities, but they also have limitations in practical applications (e.g., different Field-of-View (FoV) and frame rate). In this paper, we introduce a versatile pedestrian detector that shows robust detection performance in any single modality. We propose a multisensory-matching contrastive loss to reduce the difference between the visual representation of pedestrians in the visible and thermal modalities. Moreover, for the robust detection on a single modality, we design a Multispectral Recalling (MSR) Memory. The MSR Memory enhances the visual representation of the single modal features by recalling that of the multispectral modalities. To guide the MSR Memory to store the multispectral modal contexts, we introduce a multispectral recalling loss. It enables the pedestrian detector to encode more discriminative features with a single input modality. We believe our method is a step forward detector that can be applied to a variety of real-world applications. The comprehensive experimental results verify the effectiveness of the proposed method.
Jung Uk Kim, Sungjune Park, Yong Man Ro
AAAI1
2022 Robust Thermal Infrared Pedestrian Detection By Associating Visible Pedestrian Knowledge
abstract
Recently, pedestrian detection on thermal infrared images has shown the robust pedestrian detection performance. In this paper, we propose a novel thermal infrared pedestrian detection framework which can associate and utilize the complementary pedestrian knowledge from visible images. Motivated by that humans can associate useful information from other sensors to perform a more reliable decision, we devise a Visible-sensory Pedestrian Associating (VPA) Memory to conduct the robust pedestrian detection by utilizing complementary visible-sensory pedestrian knowledge explicitly. The VPA Memory is trained to store the pedestrian information of visible images and associate it with a given thermal infrared pedestrian knowledge via the memory associating learning. We verify the effectiveness of the proposed framework with extensive experiments, and it achieves state-of-the-art pedestrian detection performance on thermal infrared images.
Sungjune Park, Dae Hwi Choi, Jung Uk Kim, Yong Man Ro
ICASSP3
2022 Uncertainty-Guided Cross-Modal Learning for Robust Multispectral Pedestrian Detection
abstract
Multispectral pedestrian detection has received great attention in recent years as multispectral modalities (i.e. color and thermal) can provide complementary visual information. However, there are major inherent issues in multispectral pedestrian detection. First, the cameras of the two modalities have different field-of-views (FoVs), so that image pairs are often miscalibrated. Second, modality discrepancy is observed, because image pairs are captured at different wavelengths. In this paper, to alleviate these issues, we propose a new uncertainty-aware multispectral pedestrian detection framework. In our framework, we consider two types of uncertainties: 1) Region of Interest (RoI) uncertainty and 2) predictive uncertainty. For the miscalibration issue, we propose RoI uncertainty which represents the reliability of the RoI candidates. With the RoI uncertainty, when combining two modal features, we devise uncertainty-aware feature fusion (UFF) module to reduce the effect of RoI features with high RoI uncertainty. We also propose uncertainty-aware cross-modal guiding (UCG) module for the modality discrepancy. In the UCG module, we use the predictive uncertainty, which indicates how reliable the prediction of the RoI feature is. Based on the predictive uncertainty, the UCG module guides the feature distribution of high predictive uncertain (less reliable) modality to resemble that of low predictive uncertain (more reliable) modality. The UCG module can encode more discriminative features by guiding feature distributions of two modalities to be similar. With comprehensive experiments on the public multispectral datasets, we verified that our method reduces the effect of the miscalibration and alleviates the modality discrepancy, outperforming existing state-of-the-art methods.
Jung Uk Kim, Sungjune Park, Yong Man Ro
IEEE Trans. Circuits Syst. Video Technol.1
2021 Towards Robust Training of Multi-Sensor Data Fusion Network Against Adversarial Examples in Semantic Segmentation
abstract
The success of multi-sensor data fusions in deep learning appears to be attributed to the use of complementary information among multiple sensor datasets. Compared to their predictive performance, relatively less attention has been devoted to the adversarial robustness of multi-sensor data fusion models. To achieve adversarial robust multi-sensor data fusion networks, we propose here a novel robust training scheme called Multi-Sensor Cumulative Learning (MSCL). The motivation behind the MSCL method is based on the way human beings learn new skills. The MSCL allows the multi-sensor fusion network to learn robust features from individual sensors, and then learn complex joint features from multiple sensors just as people learn to walk before they run. The step wise framework of MSCL enables the network to incorporate pre-trained knowledge of robustness with new joint information from multiple sensors. Extensive experimental evidence validated that the MSCL outperforms other multi-sensor fusion training in defending against adversarial examples.
Youngjoon Yu, Hong Joo Lee 0001, Byeong Cheon Kim, Jung Uk Kim, Yong Man Ro
ICASSP4
2021 Robust Small-scale Pedestrian Detection with Cued Recall via Memory Learning
abstract
Although the visual appearances of small-scale objects are not well observed, humans can recognize them by associating the visual cues of small objects from their memorized appearance. It is called cued recall. In this paper, motivated by the memory process of humans, we introduce a novel pedestrian detection framework that imitates cued recall in detecting small-scale pedestrians. We propose a large-scale embedding learning with the large-scale pedestrian recalling memory (LPR Memory). The purpose of the proposed large-scale embedding learning is to memorize and recall the large-scale pedestrian appearance via the LPR Memory. To this end, we employ the large-scale pedestrian exemplar set, so that, the LPR Memory can recall the information of the large-scale pedestrians from the small-scale pedestrians. Comprehensive quantitative and qualitative experimental results validate the effectiveness of the proposed framework with the LPR Memory.
Jung Uk Kim, Sungjune Park, Yong Man Ro
ICCV1
2021 Robust Multispectral Pedestrian Detection via Uncertainty-Aware Cross-Modal Learning
Sungjune Park, Jung Uk Kim, Yeongyun Kim, Sang-Keun Moon, Yong Man Ro
MMM (1)2
2021 CUA Loss: Class Uncertainty-Aware Gradient Modulation for Robust Object Detection
abstract
Recently, a wide range of research on object detection has shown breakthrough performance. However, in a challenging environment, such as occlusion and small object cases, object detectors still produce inaccurate or erroneous predictions. To effectively cope with such conditions, most of the existing methods have suggested loss functions to guide the object detectors by modulating the magnitude of their loss. However, when modulating the loss function, they are highly dependent on the classification score of the object detector. It is a known fact that deep neural networks tend to be overconfident in their predictions. In this article, to alleviate the problem of the object detectors which heavily rely on the prediction in the training phase, we devise a novel loss function called class uncertainty-aware (CUA) loss. CUA loss considers the predictive ambiguity as well as the predictions on classification score when modulating loss function. In addition to the classification score, CUA loss further modulates the loss gradient in an increasing way when the object detectors output an uncertain prediction. Therefore, object detectors with CUA loss effectively cope with challenging environments where prediction results are uncertain. With comprehensive experiments on three public datasets (i.e. PASCAL VOC, MS COCO, and Berkeley DeepDrive), we verified that our CUA loss enhanced the accuracy of the object detectors and outperformed previous state-of-the-art loss functions.
Jung Uk Kim, Seong Tae Kim 0001, Hong Joo Lee 0001, Sangmin Lee 0001, Yong Man Ro
IEEE Trans. Circuits Syst. Video Technol.1
2020 Structure Boundary Preserving Segmentation for Medical Image With Ambiguous Boundary
abstract
In this paper, we propose a novel image segmentation method to tackle two critical problems of medical image, which are (i) ambiguity of structure boundary in the medical image domain and (ii) uncertainty of the segmented region without specialized domain knowledge. To solve those two problems in automatic medical segmentation, we propose a novel structure boundary preserving segmentation framework. To this end, the boundary key point selection algorithm is proposed. In the proposed algorithm, the key points on the structural boundary of the target object are estimated. Then, a boundary preserving block (BPB) with the boundary key point map is applied for predicting the structure boundary of the target object. Further, for embedding experts' knowledge in the fully automatic segmentation, we propose a novel shape boundary-aware evaluator (SBE) with the ground-truth structure information indicated by experts. The proposed SBE could give feedback to the segmentation network based on the structure boundary key point. The proposed method is general and flexible enough to be built on top of any deep learning-based segmentation network. We demonstrate that the proposed method could surpass the state-of-the-art segmentation network and improve the accuracy of three different segmentation network models on different types of medical image datasets.
Hong Joo Lee 0001, Jung Uk Kim, Sangmin Lee 0001, Hak Gu Kim, Yong Man Ro
CVPR2
2020 SACA Net: Cybersickness Assessment of Individual Viewers for VR Content via Graph-Based Symptom Relation Embedding
Sangmin Lee 0001, Jung Uk Kim, Hak Gu Kim, Seongyeop Kim, Yong Man Ro
ECCV (23)2
2020 Towards High-Performance Object Detection: Task-Specific Design Considering Classification and Localization Separation
abstract
Object detection performs two tasks (classification and localization) simultaneously. Two tasks share a similarity: they need robust features that effectively represent the visual appearance of the objects. However, two tasks also have different properties. First, classification mainly requires features from discriminative parts of an object to determine the object category, whereas localization mainly requires features from the entire object regions for localizing by drawing a bounding box. Second, classification has a translation invariant property, whereas localization has a translation variant property. In order to increase the efficiency of object detection, it is necessary to design a network in consideration of the commonalities and differences of two tasks. In this work, we simply modified layers of the existing object detection networks into three parts by considering such characteristics: lower-layer feature sharing part, layer separation part, and feature fusion part. As a result, the performance of the proposed method was noticeably improved by properly sharing, separating, and fusing layers of the existing object detection networks.
Jung Uk Kim, Seong Tae Kim 0001, Eun Sung Kim, Sang-Keun Moon, Yong Man Ro
ICASSP1
2020 Fake Video Detection With Certainty-Based Attention Network
abstract
DeepFake synthesizes realistic fake videos that could be used maliciously such as manipulation and harassment. In order to prevent such malicious usages, detecting fake videos is immediately needed. In this paper, we propose a novel fake video detection method by adopting predictive uncertainty in detection. We devise the certainty-based attention network which guides to focus certainty-key frames in detecting fake videos. In addition, certainty-based attention is proposed for refining the features with consideration for frame-level certainty. Experiments are performed to validate the effectiveness of the proposed method by comparing the existing methods on Celeb-DF, the latest DeepFake dataset.
Dae Hwi Choi, Hong Joo Lee 0001, Sangmin Lee 0001, Jung Uk Kim, Yong Man Ro
ICIP4
2020 Comprehensive Facial Expression Synthesis Using Human-Interpretable Language
abstract
Recent advances in facial expression synthesis have shown promising results using diverse expression representations including facial action units. Facial action units for an elaborate facial expression synthesis need to be intuitively represented for human comprehension, not a numeric categorization of facial action units. To address this issue, we utilize human-friendly approach: use of natural language where language helps human grasp conceptual contexts. In this paper, therefore, we propose a new facial expression synthesis model from language-based facial expression description. Our method can synthesize the facial image with detailed expressions. In addition, effectively embedding language features on facial features, our method can control individual word to handle each part of facial movement. Extensive qualitative and quantitative evaluations were conducted to verify the effectiveness of the natural language.
Joanna Hong, Jung Uk Kim, Sangmin Lee 0001, Yong Man Ro
ICIP2
2020 Learning Style Correlation for Elaborate Few-Shot Classification
abstract
Few-shot classification is defined as a task where the network aims to classify unseen classes given only a few samples. Recent approaches, especially metric-based methods, have great progress in few-shot classification. However, the existing metric-based methods have a limitation in deploying discriminative features for elaborate comparison. They usually extract features from the embedding network without direct consideration of the relationship between support and query sets. To address the relationship, we propose a novel architecture, Style Correlated Module (SCM) to learn style correlation between support and query sets for few-shot classification. The proposed module leads support and query feature maps to focus on significant style correlated features and encourage the metric network to conduct an elaborate comparison. Furthermore, the proposed module can be generally applied to the existing metric-based approaches by adding the SCM behind the embedding network. We evaluate our proposed method with comprehensive experiments on two publicly available datasets and demonstrate its effectiveness with comparable results.
Minsu Kim 0001, Jung Uk Kim, Hong Joo Lee 0001, Sangmin Lee 0001, Joanna Hong, Yong Man Ro
ICIP3
2020 Class Incremental Learning With Task-Selection
abstract
Despite the success of the deep neural networks (DNNs), in case of incremental learning, DNNs are known to suffer from catastrophic forgetting problems which are the phenomenon of entirely forgetting previously learned task information upon learning current task information. To alleviate this problem, we propose a novel knowledge distillation-based class incremental learning method with a task-selective autoencoder (TsAE). By learning the TsAE to reconstruct the feature map of each task, the proposed method effectively memorizes not only the classes of the current task but also the classes of previously learned tasks. Since the proposed TsAE has a simple but powerful architecture, it can be easily generalized to other knowledge distillation-based class incremental learning methods. Our experimental results on various datasets, including iCIFAR-100 and iILSVRC-small, demonstrated that the proposed method achieves higher classification accuracy and less forgetting compared to the stateof-the-art methods.
Eun Sung Kim, Jung Uk Kim, Sangmin Lee 0001, Sang-Keun Moon, Yong Man Ro
ICIP2
2020 Revisiting Role of Autoencoders in Adversarial Settings
abstract
To combat against adversarial attacks, autoencoder structure is widely used to perform denoising which is regarded as gradient masking. In this paper, we revisit the role of autoencoders in adversarial settings. Through the comprehensive experimental results and analysis, this paper presents the inherent property of adversarial robustness in the autoencoders. We also found that autoencoders may use robust features that cause inherent adversarial robustness. We believe that our discovery of the adversarial robustness of the autoencoders can provide clues to the future research and applications for adversarial defense.
Byeong Cheon Kim, Jung Uk Kim, Hakmin Lee, Yong Man Ro
ICIP2
2020 Towards Human-Like Interpretable Object Detection Via Spatial Relation Encoding
abstract
The performance of recent deep neural networks in various computer vision areas such as object detection has increased significantly. Along with such advances, attempts to visualize and interpret the networks have been made in order to understand how a network predicts a certain result. However, there is a lack of research on ways to improve the interpretability of networks’ features. In this paper, we propose a spatial relation reasoning (SRR) framework to encode interpretable networks’ features, especially an object detector, by mimicking the human visual cognition system. The SRR consists of the spatial feature encoder (SFE) and the graph-based spatial relation encoder (GSRE) to consider spatial relationships between different parts of an object. So that, object detectors can encode spatially-related object features enabling humanlike visual interpretation. We verified the proposed framework with general object detectors on public datasets-PAS-CAL VOC and MS COCO.
Jung Uk Kim, Sungjune Park, Yong Man Ro
ICIP1
2020 BBC Net: Bounding-Box Critic Network for Occlusion-Robust Object Detection
abstract
Object detection has received significant interest in the research field of computer vision and is widely used in human-centric applications. The occlusion problem is a frequent obstacle that degrades detection quality. In this paper, we propose a novel object detection framework targeting robust object detection in occlusion. The proposed deep learning-based network consists mainly of two parts: 1) object detection framework, which classifies the object categories and localizes the object location and 2) plug-in bounding-box (BB) estimator, which estimates the object and occlusion region from the feature map of the backbone network and the corresponding critic network for evaluating the predicted BB map. The BB estimator and the critic network are the plug-in modules added to the object detection framework and learned competitively with adversarial manner. As the plug-in BB estimator is learned to estimate the BB map containing the object and occlusion pattern information, the backbone network can embed this information to enable robust detection under occlusion in the test phase. The comprehensive experimental results on the PASCAL VOC, MS COCO, and KITTI dataset showed that the performance is improved with the plug-in BB-Critic network by predicting and criticizing object and occlusion in general generic object detection framework.
Jung Uk Kim, Jungsu Kwon, Hak Gu Kim, Yong Man Ro
IEEE Trans. Circuits Syst. Video Technol.1
2019 Attentive Layer Separation for Object Classification and Object Localization in Object Detection
abstract
Object detection became one of the major fields in computer vision. In object detection, object classification and object localization tasks are conducted. Previous deep learning-based object detection networks perform with feature maps generated by completely shared networks. However, object classification focuses on the most discriminative object part of the feature map. Whereas, object localization requires a feature map that is focused on the entire area of the object. In this paper, we propose a novel object detection network by considering the difference between the two tasks. The proposed deep learning-based network mainly consists of two parts; 1) Attention network part where task-specific attention maps are generated, 2) Layer separation part where layers for estimating two tasks are separated. Comprehensive experimental results based on PASCAL VOC dataset and MS COCO dataset showed that proposed object detection network outperformed the state-of-the-art methods.
Jung Uk Kim, Yong Man Ro
ICIP1
2018 Object Bounding Box-Critic Networks for Occlusion-Robust Object Detection in Road Scene
abstract
Object detection in a road scene has received a significant attention from research fields of developing autonomous vehicle and automatic road monitoring systems. However, object occlusion problems frequently occur in generic road scenes. Due to such occlusion problems, previous object detection methods have limitations of not being able to detect objects accurately. In this paper, we propose a novel object detection network which is robust in occlusions. For effective object detection even with occlusion, the proposed network mainly consists of two parts; 1) Object detection framework, 2) Multiple object bounding box (OBB)-Critic network for predicting a BB map which estimates both object region and occlusion region. Comprehensive experimental results on a KITTI Vision Benchmark Suite dataset showed that the proposed object detection network outperformed the state-of-the-art methods.
Jung Uk Kim, Jungsu Kwon, Hak Gu Kim, Haesung Lee, Yong Man Ro
ICIP1