Seong Tae Kim 0001

dblp:70/4799-1 · also Seong-Tae Kim 0001 · DBLP profile ↗
← Back
53ranked-venue papers
6as first author
34since 2021 · last 2026
0000-0002-2132-6021ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 11 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging Textual Compositional Reasoning for Robust Change Captioning
abstract
Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and compositional semantics. To alleviate this, we present CORTEX (COmpositional Reasoning-aware TEXt-guided), a novel framework that integrates complementary textual cues to enhance change understanding. In addition to capturing cues from pixel-level differences, CORTEX utilizes scene-level textual knowledge provided by Vision Language Models (VLMs) to extract richer image text signals that reveal underlying compositional reasoning. CORTEX consists of three key modules: (i) an Image-level Change Detector that identifies low-level visual differences between paired images, (ii) a Reasoning-aware Text Extraction (RTE) module that use VLMs to generate compositional reasoning descriptions implicit in visual features, and (iii) an Image-Text Dual Alignment (ITDA) module that aligns visual and textual features for fine-grained relational reasoning. This enables CORTEX to reason over visual and textual features and capture changes that are otherwise ambiguous in visual features alone.
Kyu Ri Park, Seong Tae Kim 0001, Hong Joo Lee 0001, Jung Uk Kim
AAAI3
2026 Unsupervised domain adaptation for medical image segmentation using adaptogen-perturbation
abstract
Domains shift originated from differences in devices or patients in the medical field, poses a significant challenge when applying pre-trained models to clinical applications. To tackle this challenge, domain adaptation methods have been explored. However, most existing methods are designed for a single target domain adaptation or require sharing all target domain data for adaptation, which is infeasible in the medical field due to privacy issues. In this paper, we propose a novel unsupervised multi-target domain adaptation method without requiring data sharing. To this end, we introduce an additional signal, termed Adaptogen-Perturbation (AP) optimized to bridge the gap between the source and target domains. The optimized AP is injected into the latent feature and facilitates the adaptation of the pre-trained model to the target domain. Moreover, we propose a Spectral/Geometric Consistency learning framework to optimize the AP in an unsupervised manner. This promotes consistent predictions across two types of transformations: geometric and frequency-space spectral transformations, enhancing robustness to both variations. Extensive experiments with multiple medical segmentation datasets demonstrate the effectiveness of APs.
Hong Joo Lee 0001, Yuan Bi, Sangmin Lee 0001, Gyeong-Moon Park, Jung Uk Kim, Seong Tae Kim 0001, Zhongliang Jiang, Nassir Navab
Medical Image Anal.6
2026 Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
abstract
Reliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding.
Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim 0001, Gonçalo Arantes, Kehan Song, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Oluwatosin Alabi, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang 0007, Long Bai 0008, Hongliang Ren 0001, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Andrés Arbeláez, Yiping Li 0002, Yasmina Alkhalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feußner, Dirk Wilhelm, Christoph Palm
Medical Image Anal.16
2026 Adversarial Wear and Tear: Exploiting Natural Damage for Generating Physical-World Adversarial Examples
Samra Irshad, Seungkyu Lee 0001, Nassir Navab, Hong Joo Lee 0001, Seong Tae Kim 0001
IEEE Trans. Dependable Secur. Comput.5
2025 LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic Pathologies
abstract
Generating Natural Language Explanations (NLEs) for model predictions on medical images, particularly those depicting thoracic pathologies, remains a critical and challenging task. Existing methodologies often struggle due to general models' insufficient domain-specific medical knowledge and privacy concerns associated with retrieval-based augmentation techniques. To address these issues, we propose a novel Vision-Language framework augmented with a Knowledge Graph (KG)-based datastore, which enhances the model's understanding by incorporating additional domain-specific medical knowledge essential for generating accurate and informative NLEs. Our framework employs a KG-based retrieval mechanism that not only improves the precision of the generated explanations but also preserves data privacy by avoiding direct data retrieval. The KG datastore is designed as a plug-and-play module, allowing for seamless integration with various model architectures. We introduce and evaluate three distinct frameworks within this paradigm: KG-LLaVA, which integrates the pre-trained LLaVA model with KG-RAG; Med-XPT, a custom framework combining MedCLIP, a transformer-based projector, and GPT-2; and Bio-LLaVA, which adapts LLaVA by incorporating the Bio-ViT-L vision model. These frameworks are validated on the MIMIC-NLE dataset, where they achieve state-of-the-art results, underscoring the effectiveness of KG augmentation in generating high-quality NLEs for thoracic pathologies.
Yong Hyun Ahn, Sungyoung Lee 0001, Seong Tae Kim 0001
AAAI5
2025 HiCM²: Hierarchical Compact Memory Modeling for Dense Video Captioning
abstract
With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical dense memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical dense memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets.
Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 0001, Seong Tae Kim 0001
AAAI5
2025 When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series
abstract
Recently, forecasting future abnormal events has emerged as an important scenario to tackle realworld necessities. However, the solution of predicting specific future time points when anomalies will occur, known as Anomaly Prediction (AP), remains under-explored. Existing methods dealing with time series data fail in AP, focusing only on immediate anomalies or failing to provide precise predictions for future anomalies. To address AP, we propose a novel framework called Anomaly to Prompt (A2P), comprised of Anomaly-Aware Forecasting (AAF) and Synthetic Anomaly Prompting (SAP). To enable the forecasting model to forecast abnormal time points, we adopt a strategy to learn the relationships of anomalies. For the robust detection of anomalies, our proposed SAP introduces a learnable Anomaly Prompt Pool (APP) that simulates diverse anomaly patterns using signal-adaptive prompt. Comprehensive experiments on multiple real-world datasets demonstrate the superiority of A2P over state-of-the-art methods, showcasing its ability to predict future anomalies.
Min-Yeong Park, Won-Jeong Lee, Seong Tae Kim 0001, Gyeong-Moon Park
ICML3
2025 PRADA: Protecting and Detecting Dataset Abuse for Open-Source Medical Dataset
Jinhyeok Jang, Hong Joo Lee 0001, Nassir Navab, Seong Tae Kim 0001
MICCAI (14)4
2025 SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition
Ka Young Kim, Hyeon Bae Kim, Seong Tae Kim 0001
MICCAI (10)3
2025 Towards Holistic Surgical Scene Graph
Enki Cho, Ka Young Kim, Jung Yong Kim, Seong Tae Kim 0001, Namkee Oh
MICCAI (9)5
2025 Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
abstract
Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods—based on saliency—produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Language-based approaches offer structure but often fail to explain motions due to their tacit nature—intuitively understood but difficult to verbalize. To address these challenges, we propose Disentangled Action aNd Context concept-based Explainable (DANCE) video action recognition, a framework that predicts actions through disentangled concept types: motion dynamics, objects, and scenes. We define motion dynamics concepts as human pose sequences. We employ a large language model to automatically extract object and scene concepts. Built on an ante-hoc concept bottleneck design, DANCE enforces prediction through these concepts. Experiments on four datasets—KTH, Penn Action, HAA500, and UCF101—demonstrate that DANCE significantly improves explanation clarity with competitive performance. Through a user study, we validate the superior interpretability of DANCE. Experimental results also show that DANCE is beneficial for model debugging, editing, and failure analysis.
Jongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim 0001, Jinwoo Choi 0001
NeurIPS4
2025 Unified link prediction modeling for enhanced knowledge graph completion task
Tri D. T. Nguyen, Ubaid Ur Rehman 0002, Musarrat Hussain, Rao Faizan, Jamil Hussain, Sung-Ho Bae, Jung Uk Kim, Seong Tae Kim 0001, Sungyoung Lee 0001
Expert Syst. Appl.8
2025 Understanding adversarial robustness of deep neural networks via decision reliance
Soyoun Won, Hyeon Bae Kim, Yong Hyun Ahn, Hong Joo Lee 0001, Seong Tae Kim 0001
Image Vis. Comput.5
2025 PRISM: Pseudo-Labeling and Region-Based Inpainting for Synthetic Change Detection Modeling
abstract
With the rapid advances in deep learning, there have been significant advances in the research of remote sensing, especially in the area of change detection. Change detection is crucial for monitoring urban development and environmental changes, but constructing high-quality, region-specific datasets remains a costly and labor-intensive challenge. Moreover, models trained on specific regions often underperform in new domains due to environmental differences. To address these limitations, we propose a novel framework, PRISM (Pseudo-labeling and Region-based Inpainting for Synthetic Change Detection Modeling), that generates synthetic post-images directly from single pre-images, eliminating the need for region-specific annotations or bitemporal image pairs. Our method segments building areas in the input image based on the foundation model, and the segmented regions are removed. The generative inpainting is applied to simulate realistic landscapes (e.g., grasslands, plains) that reflect hypothetical future changes. By leveraging morphological operations and descriptive text prompts, our approach ensures seamless integration of generated content with the surrounding context, producing realistic and region-adaptable datasets. Experimental results demonstrate that our method could reduce the dependency on annotated data, enhance adaptability across diverse regions, and enable efficient and scalable change detection modeling in an unsupervised setting.
Enki Cho, Soyoun Won, Sung-Sik Choo, Seong Tae Kim 0001
IEEE Geosci. Remote. Sens. Lett.4
2025 Effects of mixed sample data augmentation on interpretability of neural networks
Soyoun Won, Sung-Ho Bae, Seong Tae Kim 0001
Neural Networks3
2025 Spatial Mask-Based Adaptive Robust Training for Video Object Segmentation With Noisy Labels
abstract
Recent advances in video object segmentation (VOS) highlight its potential across various applications. Semi-supervised VOS aims to segment target objects in video frames based on annotations from the initial frame. Collecting a large-scale video segmentation dataset is challenging, which could induce noisy labels. However, it has been overlooked and most of the research efforts have been devoted to training VOS models by assuming the training dataset is clean. In this study, we first explore the effect of VOS models under noisy labels in the training dataset. To investigate the effect of noisy labels, we simulate the noisy annotations on DAVIS 2017 and YouTubeVOS datasets. Experiments show that the traditional training strategy is vulnerable to noisy annotations. To address this issue, we propose a novel noise-robust training method, named SMART (Spatial Mask-based Adaptive Robust Training), which is designed to train models effectively in the presence of noisy annotations. The proposed method employs two key strategies. Firstly, the model focuses on the common spatial areas from clean knowledge-based predictions and annotations. Secondly, the model is trained with adaptive balancing losses based on their reliability. Comparative experiments have demonstrated the effectiveness of our approach by outperforming other noise handling methods over various noise degrees.
Enki Cho, Jung Uk Kim, Seong Tae Kim 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 WWW: A Unified Framework for Explaining what, Where and why of Neural Networks by Interpretation of Neuron Concepts
abstract
Recent advancements in neural networks have show-cased their remarkable capabilities across various do-mains. Despite these successes, the “black box” problem still remains. To address this, we propose a novel frame-work, www, that offers the ‘what’, ‘where’, and ‘why’ of the neural network decisions in human-understandable terms. Specifically, WWW utilizes adaptive selection for concept discovery, employing adaptive cosine simi-larity and thresholding techniques to effectively explain ‘what’. To address the ‘where’ and ‘why’, we proposed a novel combination of neuron activation maps (NAMs) with Shapley values, generating localized concept maps and heatmaps for individual inputs. Furthermore, WWW in-troduces a method for predicting uncertainty, leveraging heatmap similarities to estimate the prediction's reliability. Experimental evaluations of WWW demonstrate superior performance in both quantitative and qualitative metrics, outperforming existing methods in interpretability. WWW provides a unified solution for explaining ‘what’, ‘where’, and ‘why’, introducing a methodfor localized explanations from global interpretations and offering a plug-and-play so-lution adaptable to various architectures. Code is available at: https://github.comlailab-kyungheelWWW
Yong Hyun Ahn, Hyeon Bae Kim, Seong Tae Kim 0001
CVPR3
2024 Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
abstract
There has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a multitasking problem of event localization and event captioning to consider inter-task relations. However, addressing both tasks using only visual input is challenging due to the lack of semantic content. In this study, we address this by proposing a novel framework inspired by the cognitive information processing of humans. Our model utilizes external memory to incorporate prior knowledge. The memory retrieval method is proposed with cross-modal video-to-text matching. To effectively incorporate retrieved text features, the versatile encoder and the decoder with visual and textual cross-attention modules are designed. Comparative experiments have been conducted to show the effectiveness of the proposed method on ActivityNet Captions and YouCook2 datasets. Experimental results show promising performance of our model without extensive pretraining from a large video dataset. Our code is available at https://github.com/ailab-kyunghee/CM2_DVC.
Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 0001, Seong Tae Kim 0001
CVPR5
2024 MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
Youngmin Oh 0003, Hyungil Kim, Seong Tae Kim 0001, Jung Uk Kim
ECCV (10)3
2024 Chain-of-Factors: A Zero-Shot Prompting Methodology Enabling Factor-Centric Reasoning in Large Language Models
abstract
Large language models (LLMs) have significantly improved numerous natural language processing tasks. However, their performance relies heavily on the provided instructions or prompts. Recently, several prompting methodologies have been developed to enhance the reasoning abilities of LLMs. Notably, the Chain-of-Thought (CoT) approach provides examples that help break down tasks into sub-steps, resulting in more accurate solutions. However, the process of generating detailed examples may not be user-friendly, as end users prefer providing task descriptions rather than a set of examples. In this study, we introduce Chain-of-Factors (CoF), an innovative zero-shot prompting methodology that incorporates task-specific instructions as a chain of factors into the prompt, aimed at enhancing the factor-centric reasoning abilities of LLMs. Experiments on three LLMs, including ChatGPT-3.5, Gemini, and GPT-4, show performance improvements ranging from 0.01% to 40.2% in accuracy on various symbolic reasoning and logical reasoning tasks compared with zero-shot and few-shot CoT. In summary, CoF enhances LLMs' reasoning abilities by including task-specific steps and instructions, while also decreasing the necessity for fine-tuning specific to each task.
Musarrat Hussain, Ubaid Ur Rehman 0002, Tri D. T. Nguyen, Sungyoung Lee 0001, Seong Tae Kim 0001, Sung-Ho Bae, Jung Uk Kim
ICMLA5
2024 Mask-Free Neuron Concept Annotation for Interpreting Neural Networks in Medical Domain
Hyeon Bae Kim, Yong Hyun Ahn, Seong Tae Kim 0001
MICCAI (10)3
2024 OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
Chu Myaet Thwal, Minh N. H. Nguyen, Ye Lin Tun 0001, Seong Tae Kim 0001, My T. Thai, Choong Seon Hong
Neural Networks4
2023 LINe: Out-of-Distribution Detection by Leveraging Important Neurons
abstract
It is important to quantify the uncertainty of input samples, especially in mission-critical domains such as autonomous driving and healthcare, where failure predictions on out-of-distribution (OOD) data are likely to cause big problems. OOD detection problem fundamentally begins in that the model cannot express what it is not aware of. Post-hoc OOD detection approaches are widely explored because they do not require an additional re-training process which might degrade the model's performance and increase the training cost. In this study, from the perspective of neurons in the deep layer of the model representing high-level features, we introduce a new aspect for analyzing the difference in model outputs between in-distribution data and OOD data. We propose a novel method, Leveraging Important Neurons (LINe), for post-hoc Out of distribution detection. Shapley value-based pruning reduces the effects of noisy outputs by selecting only high-contribution neurons for predicting specific classes of input data and masking the rest. Activation clipping fixes all values above a certain threshold into the same value, allowing LINe to treat all the class-specific features equally and just consider the difference between the number of activated feature differences between in-distribution and OOD data. Comprehensive experiments verify the effectiveness of the proposed method by outperforming state-of-the-art post-hoc OOD detection methods on CIFAR-10, CIFAR-100, and ImageNet datasets. Code is available on https://github.com/LINe-OOD
Yong Hyun Ahn, Gyeong-Moon Park, Seong Tae Kim 0001
CVPR3
2023 Towards Robust Audio-Based Vehicle Detection Via Importance-Aware Audio-Visual Learning
abstract
Although audio modality has the potential to solve various visually challenging conditions of visual modality, there are few studies on audio-based detection. This is because the audio modality itself contains less accurate spatial information. To alleviate this issue, the existing audio-based methods adopt the visual modality in the training phase to transfer more precise spatial knowledge to the audio modality. However, they do not consider the case where the visual modality is less informative. In this paper, we present a new audio-based vehicle detector that can transfer multimodal knowledge of vehicles to the audio modality during training. To this end, we combine the audio-visual modal knowledge according to the importance of each modality to generate integrated audiovisual feature. Also, we introduce an audio-visual distillation (AVD) loss that guides representation of the audio modal feature to resemble that of the integrated audio-visual feature. As a result, our audio-based detector can perform robust vehicle detection as if it were utilizing both modalities, even if it only receives audio modality as input in the inference. Comprehensive experimental results demonstrate that our method exhibits consistent improvements over the existing methods.
Jung Uk Kim, Seong Tae Kim 0001
ICASSP2
2023 Exploiting recollection effects for memory-based video object segmentation
Enki Cho, Minkuk Kim, Hyungil Kim, Jinyoung Moon, Seong Tae Kim 0001
Image Vis. Comput.5
2022 Exploiting Diversity of Unlabeled Data for Label-Efficient Semi-Supervised Active Learning
abstract
The availability of large labeled datasets is the key component for the success of deep learning. However, annotating labels on large datasets is generally time-consuming and expensive. Active learning is a research area that addresses the issues of expensive labeling by selecting the most important samples for labeling. Diversity-based sampling algorithms are known as integral components of representation-based approaches for active learning. In this paper, we introduce a new diversity-based initial dataset selection algorithm to select the most informative set of samples for initial labeling in the active learning setting. Self-supervised representation learning is used to consider the diversity of samples in the initial dataset selection algorithm. Also, we propose a novel active learning query strategy, which uses diversity-based sampling on consistency-based embeddings. By considering the consistency information with the diversity in the consistency-based embedding scheme, the proposed method could select more informative samples for labeling in the semi-supervised learning setting. Comparative experiments show that the proposed method achieves compelling results on CIFAR-10 and Caltech-101 datasets compared with previous active learning approaches by utilizing the diversity of unlabeled data.
Felix Buchert, Nassir Navab, Seong Tae Kim 0001
ICPR3
2022 Robust Perturbation for Visual Explanation: Cross-Checking Mask Optimization to Avoid Class Distortion
abstract
Along with the outstanding performance of the deep neural networks (DNNs), considerable research efforts have been devoted to finding ways to understand the decision of DNNs structures. In the computer vision domain, visualizing the attribution map is one of the most intuitive and understandable ways to achieve human-level interpretation. Among them, perturbation-based visualization can explain the "black box" property of the given network by optimizing perturbation masks that alter the network prediction of the target class the most. However, existing perturbation methods could make unexpected changes to network predictions after applying a perturbation mask to the input image, resulting in a loss of robustness and fidelity of the perturbation mechanisms. In this paper, we define class distortion as the unexpected changes of the network prediction during the perturbation process. To handle that, we propose a novel visual interpretation framework, Robust Perturbation, which shows robustness against the unexpected class distortion during the mask optimization. With a new cross-checking mask optimization strategy, our proposed framework perturbs the target prediction of the network while upholding the non-target predictions, providing more reliable and accurate visual explanations. We evaluate our framework on three different public datasets through extensive experiments. Furthermore, we propose a new metric for class distortion evaluation. In both quantitative and qualitative experiments, tackling the class distortion problem turns out to enhance the quality and fidelity of the visual explanation in comparison with the existing perturbation-based methods.
Seongyeop Kim, Seong Tae Kim 0001, Yong Man Ro
IEEE Trans. Image Process.3
2021 Neural Response Interpretation Through the Lens of Critical Pathways
abstract
Is critical input information encoded in specific sparse pathways within the neural network? In this work, we discuss the problem of identifying these critical pathways and subsequently leverage them for interpreting the network’s response to an input. The pruning objective — selecting the smallest group of neurons for which the response remains equivalent to the original network — has been previously proposed for identifying critical pathways. We demonstrate that sparse pathways derived from pruning do not necessarily encode critical input information. To ensure sparse pathways include critical fragments of the encoded input information, we propose pathway selection via neurons’ contribution to the response. We proceed to explain how critical pathways can reveal critical input features. We prove that pathways selected via neuron contribution are locally linear (in an ℓ2-ball), a property that we use for proposing a feature attribution method: "pathway gradient". We validate our interpretation method using mainstream evaluation experiments. The validation of pathway gradient interpretation method further confirms that selected pathways using neuron contributions correspond to critical input features. The code12is publicly available.
Ashkan Khakzar, Soroosh Baselizadeh, Saurabh Khanduja, Christian Rupprecht 0001, Seong Tae Kim 0001, Nassir Navab
CVPR5
2021 OperA: Attention-Regularized Transformers for Surgical Phase Recognition
Tobias Czempiel, Magdalini Paschali, Daniel Ostler, Seong Tae Kim 0001, Benjamin Busam, Nassir Navab
MICCAI (4)4
2021 Towards Semantic Interpretation of Thoracic Disease and COVID-19 Diagnosis Models
Ashkan Khakzar, Sabrina Musatian, Jonas Buchberger, Icxel Valeriano Quiroz, Nikolaus Pinger, Soroosh Baselizadeh, Seong Tae Kim 0001, Nassir Navab
MICCAI (3)7
2021 Explaining COVID-19 and Thoracic Pathology Model Predictions by Identifying Informative Input Features
Ashkan Khakzar, Wejdene Mansour, Yuezhi Cai, Seong Tae Kim 0001, Nassir Navab
MICCAI (3)7
2021 Longitudinal Quantitative Assessment of COVID-19 Infection Progression from Chest CTs
Seong Tae Kim 0001, Leili Goli, Magdalini Paschali, Ashkan Khakzar, Matthias Keicher, Tobias Czempiel, Egon Burian, Rickmer Braren, Nassir Navab, Thomas Wendler 0001
MICCAI (7)1
2021 Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive Information
abstract
One principal approach for illuminating a black-box neural network is feature attribution, i.e. identifying the importance of input features for the network’s prediction. The predictive information of features is recently proposed as a proxy for the measure of their importance. So far, the predictive information is only identified for latent features by placing an information bottleneck within the network. We propose a method to identify features with predictive information in the input domain. The method results in fine-grained identification of input features' information and is agnostic to network architecture. The core idea of our method is leveraging a bottleneck on the input that only lets input features associated with predictive latent features pass through. We compare our method with several feature attribution methods using mainstream feature attribution evaluation experiments. The code is publicly available.
Ashkan Khakzar, Azade Farshad, Seong Tae Kim 0001, Nassir Navab
NeurIPS5
2021 CUA Loss: Class Uncertainty-Aware Gradient Modulation for Robust Object Detection
abstract
Recently, a wide range of research on object detection has shown breakthrough performance. However, in a challenging environment, such as occlusion and small object cases, object detectors still produce inaccurate or erroneous predictions. To effectively cope with such conditions, most of the existing methods have suggested loss functions to guide the object detectors by modulating the magnitude of their loss. However, when modulating the loss function, they are highly dependent on the classification score of the object detector. It is a known fact that deep neural networks tend to be overconfident in their predictions. In this article, to alleviate the problem of the object detectors which heavily rely on the prediction in the training phase, we devise a novel loss function called class uncertainty-aware (CUA) loss. CUA loss considers the predictive ambiguity as well as the predictions on classification score when modulating loss function. In addition to the classification score, CUA loss further modulates the loss gradient in an increasing way when the object detectors output an uncertain prediction. Therefore, object detectors with CUA loss effectively cope with challenging environments where prediction results are uncertain. With comprehensive experiments on three public datasets (i.e. PASCAL VOC, MS COCO, and Berkeley DeepDrive), we verified that our CUA loss enhanced the accuracy of the object detectors and outperformed previous state-of-the-art loss functions.
Jung Uk Kim, Seong Tae Kim 0001, Hong Joo Lee 0001, Sangmin Lee 0001, Yong Man Ro
IEEE Trans. Circuits Syst. Video Technol.2
2020 Robust Ensemble Model Training via Random Layer Sampling Against Adversarial Attack
Hakmin Lee, Hong Joo Lee 0001, Seong Tae Kim 0001, Yong Man Ro
BMVC3
2020 Towards High-Performance Object Detection: Task-Specific Design Considering Classification and Localization Separation
abstract
Object detection performs two tasks (classification and localization) simultaneously. Two tasks share a similarity: they need robust features that effectively represent the visual appearance of the objects. However, two tasks also have different properties. First, classification mainly requires features from discriminative parts of an object to determine the object category, whereas localization mainly requires features from the entire object regions for localizing by drawing a bounding box. Second, classification has a translation invariant property, whereas localization has a translation variant property. In order to increase the efficiency of object detection, it is necessary to design a network in consideration of the commonalities and differences of two tasks. In this work, we simply modified layers of the existing object detection networks into three parts by considering such characteristics: lower-layer feature sharing part, layer separation part, and feature fusion part. As a result, the performance of the proposed method was noticeably improved by properly sharing, separating, and fusing layers of the existing object detection networks.
Jung Uk Kim, Seong Tae Kim 0001, Eun Sung Kim, Sang-Keun Moon, Yong Man Ro
ICASSP2
2020 TeCNO: Surgical Phase Recognition with Multi-stage Temporal Convolutional Networks
Tobias Czempiel, Magdalini Paschali, Matthias Keicher, Walter Simson, Hubertus Feußner, Seong Tae Kim 0001, Nassir Navab
MICCAI (3)6
2020 Multimodal facial biometrics recognition: Dual-stream convolutional neural networks with multi-feature fusion layers
Leslie Ching Ow Tiong, Seong Tae Kim 0001, Yong Man Ro
Image Vis. Comput.2
2020 Lightweight and Effective Facial Landmark Detection using Adversarial Learning with Face Geometric Map Generative Network
abstract
Facial landmark detection plays an important role in face analysis tasks. Moreover, it is used as a prerequisite in many facial related applications, the simplicity, as well as effectiveness, is essential in the facial landmark detection. In this paper, we propose an effective facial landmark detection network and an associated learning framework with the geometric prior-generative adversarial network. The geometric prior-generative adversarial network consists of one generator and two discriminators. The generator consists of an encoder and two decoders. The encoder predicts facial landmark points. The decoders generate a facial inner and contour geometric map from predicted landmark points. Generating face geometric maps from predicted landmark points helps the predicted landmark points to represent the face geometric information, including shape and configuration. The discriminators determine that the given geometric maps are generated from actual landmark points or estimated landmark points. Our proposed network is end-to-end trainable, and only the encoder part is used simply as the facial landmark detector in the testing stage. To verify the effectiveness of the proposed method, we have conducted comprehensive experiments with benchmark data sets. The results have shown that the proposed method achieves comparable performances over recently proposed facial landmark detection methods with a simple and effective facial landmark detection network.
Hong Joo Lee 0001, Seong Tae Kim 0001, Hakmin Lee, Yong Man Ro
IEEE Trans. Circuits Syst. Video Technol.2
2019 Probenet: Probing Deep Networks
abstract
Despite the rapid progress of deep learning research in recent years, interpreting deep network is still quite challenging. Interpreting deep networks is essential to both end-users and developers since it gives confidence in the usage of the deep network. This paper deals with a method for interpreting deep networks, especially visual interpretation. In order to get visual interpretation from a target deep network, we propose a ProbeNet that provides a decomposed visual interpretation of the target deep network. The ProbeNet decomposes the feature representations of the point of the target deep network into human interpretable units. Furthermore, the ProbeNet provides kernel-level analysis about the target deep network. In experiments, visual interpretation of two different target deep networks showed the usefulness of the ProbeNet to interpret target deep networks.
Jae-Hyeok Lee 0001, Seong Tae Kim 0001, Yong Man Ro
ICIP2
2019 Realistic Breast Mass Generation Through BIRADS Category
Hakmin Lee, Seong Tae Kim 0001, Jae-Hyeok Lee 0001, Yong Man Ro
MICCAI (6)2
2019 Implementation of multimodal biometric recognition via multi-feature deep learning networks and feature fusion
Leslie Ching Ow Tiong, Seong Tae Kim 0001, Yong Man Ro
Multim. Tools Appl.2
2019 Attended Relation Feature Representation of Facial Dynamics for Facial Authentication
abstract
In psychology, it is known that facial dynamics benefit the perception of identity. This paper proposes a novel deep network framework to capture identity information from facial dynamics and their relations. In the proposed method, facial dynamics occurred from a smile expression are analyzed and utilized for facial authentication. Detailed changes in the local regions of a face such as wrinkles and dimples are encoded in the facial dynamic feature representation. The latent relationships of the facial dynamic features are learned by the facial dynamic relational network. In the facial dynamic relational network, the relation features of the facial dynamic are encoded and the relational importance is encoded based on the relation features. As a result, the proposed method has more attention on the important relation features in facial authentication. Through comprehensive and comparative experiments, the effectiveness of the proposed method has been verified in facial authentication.
Seong Tae Kim 0001, Yong Man Ro
IEEE Trans. Inf. Forensics Secur.1
2018 Facial Dynamics Interpreter Network: What Are the Important Relations Between Local Dynamics for Facial Trait Estimation?
Seong Tae Kim 0001, Yong Man Ro
ECCV (12)1
2018 Teacher and Student Joint Learning for Compact Facial Landmark Detection Network
Hong Joo Lee 0001, Wissam J. Baddar, Hak Gu Kim, Seong Tae Kim 0001, Yong Man Ro
MMM (1)4
2018 Convolution with Logarithmic Filter Groups for Efficient Shallow CNN
Tae Kwan Lee, Wissam J. Baddar, Seong Tae Kim 0001, Yong Man Ro
MMM (1)3
2017 Adaptive attention fusion network for visual question answering
abstract
Automatic understanding of the content of a reference image and natural language questions is needed in Visual Question Answering (VQA). Generating a visual attention map that focuses on the regions related to the context of the question can improve performance of VQA. In this paper, we propose adaptive attention-based VQA network. The proposed method utilizes the complementary information from the attention maps depending on three levels of word embedding (word level, phrase level, and question level embedding), and adaptively fuses the information to represent the image-question pair appropriately. Comparative experiments have been conducted on the public COCO-QA database to validate the proposed method. Experimental results have shown that the proposed method outperforms previous methods in terms of accuracy.
Geonmo Gu, Seong Tae Kim 0001, Yong Man Ro
ICME2
2016 Latent feature representation with 3-D multi-view deep convolutional neural network for bilateral analysis in digital breast tomosynthesis
abstract
In clinical studies of breast cancer, masses appear as asymmetric densities between the left and the right breasts, which show different breast tissue structures. For classifying breast masses, most researchers have developed hand-crafted bilateral features by extracting the asymmetric information in 2-D mammograms. In digital breast tomosynthesis (DBT), which has 3D volume data, effective bilateral features are needed to detect masses. In this paper, we propose latent bilateral feature representation with 3-D multi-view deep convolutional neural network (DCNN) in the DBT reconstructed volume. The proposed DCNN is designed to discover hidden or latent bilateral feature representation of masses in self-taught learning. Experimental results show that the proposed latent bilateral feature representation outperforms conventional hand-crafted features by achieving a high area under the receiver operating characteristic curve.
Dae Hoe Kim, Seong Tae Kim 0001, Yong Man Ro
ICASSP2
2016 A deep facial landmarks detection with facial contour and facial components constraint
abstract
In this paper, we propose a new facial landmarks detection method based on deep learning with facial contour and facial components constraints. The proposed deep convolutional neural networks (DCNNs) for facial landmark detection consists of two deep networks: one DCNN is to detect landmarks constrained on the facial contour and the other is to detect landmarks constrained on facial components. A novel DCNN structure for the landmarks detection with facial component constraints is proposed, which branches the network at higher layers in order to capture the intricate local facial components features. Moreover, a novel learning strategy is proposed to learn the DCNN for detecting the landmarks on the facial contour by exploiting the relationship between facial contour landmarks and those on facial components. Experimental results have shown that the proposed method outperforms the state-of-the-art FLD methods.
Wissam J. Baddar, Jisoo Son, Dae Hoe Kim, Seong Tae Kim 0001, Yong Man Ro
ICIP4
2016 Spatio-temporal representation for face authentication by using multi-task learning with human attributes
abstract
For human identification, facial motion is useful in representing specific dynamic signature. In this paper, we present an effective spatio-temporal representation from facial motion as well as appearance by devising a 3D convolutional neural network (CNN). To maintain the intra-class invariance with limited number of training samples, a multi-task learning approach with human attributes, which are high-level semantic descriptions for identity, has been proposed. Identity-related human attributes can be leveraged to learn the 3D CNN. Comparative experiment has showed that the proposed method improves the performance of the face-based authentication system compared to conventional methods by effectively encoding facial appearance and motions with identity-related human attributes.
Seong Tae Kim 0001, Dae Hoe Kim, Yong Man Ro
ICIP1
2015 Feature extraction from bilateral dissimilarity in digital breast tomosynthesis reconstructed volume
abstract
In this paper, we propose bilateral features for classifying breast masses by extracting the asymmetric information of both the left and the right breasts in the digital breast tomosynthesis (DBT) reconstructed volume. Clinically, it is known that the left and the right breast of the same patient tend to present a high degree of symmetry of internal structures over broad areas. On the other hand, masses appear as asymmetric densities which show different breast tissue structures between the left and the right breasts. Based on that clinical fact, bilateral features are proposed to measure the dissimilarity of texture or intensity characteristics between volumes-of-interest (VOIs) in a given breast, and the corresponding VOIs in the bilateral breast. Experimental results show that the proposed bilateral features in conjunction with single-view mass features can achieve higher level of classification performance in terms of the area under the receiver operating characteristic (ROC) curve (AUC) compared to the performance of the single-view features only.
Dae Hoe Kim, Seong Tae Kim 0001, Wissam J. Baddar, Yong Man Ro
ICIP2
2015 Region matching based on local structure information in ipsilateral digital breast tomosynthesis views
abstract
Digital breast tomosynthesis (DBT) is an emerging 3D x-ray imaging modality in breast cancer screening. Clinical studies have been reported that sensitivity of breast cancer detection can be increased by using two ipsilateral DBT views. Matching corresponding regions in the ipsilateral DBT views is important to achieve high detection sensitivity. In this paper, we propose a novel and effective region matching method based on local structure information in the ipsilateral DBT views. In the proposed method, for a given query region, we find corresponding region on target slices by using local descriptor and associated similarity. Experimental results showed that the proposed region matching method achieved the improvement in accuracy of region matching compared with the existing breast compression model-based region matching method.
Seong Tae Kim 0001, Dae Hoe Kim, Dong Jin Ji, Yong Man Ro
ICIP1
2008 Safety-Ensuring Systematic Design for Service Robots
Seong Tae Kim 0001, Pyung Hun Chang
ICOST1