Gyeong-Moon Park

dblp:166/0276 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0003-4011-9981ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 5 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction
abstract
Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text. This omits crucial visual cues, such as facial expressions and gestures, as well as broader conversational contexts, which are necessary for accurate prediction. In this paper, we introduce Context-Aware Multimodal Alignment for Backchannel Prediction (CAMA-BC), a novel framework that leverages visual information through Multi-Layer Multimodal Alignment (MMA). Our alignment process comprises two stages. First, Context Alignment (MMA-CA) utilizes unlabeled dialogues with videos to capture conversational contexts. Next, Backchannel Alignment (MMA-BA) fine-tunes the representations specifically for backchannel prediction. Experimental results show that CAMA-BC significantly outperforms both existing methods and simple multimodal baselines, with particular effectiveness in recognizing complex backchannels such as empathy.
Min-Jae Kim, Jun-Yeong Moon, Mujeen Sung, Gyeong-Moon Park
ACL (1)4
2026 Unsupervised domain adaptation for medical image segmentation using adaptogen-perturbation
abstract
Domains shift originated from differences in devices or patients in the medical field, poses a significant challenge when applying pre-trained models to clinical applications. To tackle this challenge, domain adaptation methods have been explored. However, most existing methods are designed for a single target domain adaptation or require sharing all target domain data for adaptation, which is infeasible in the medical field due to privacy issues. In this paper, we propose a novel unsupervised multi-target domain adaptation method without requiring data sharing. To this end, we introduce an additional signal, termed Adaptogen-Perturbation (AP) optimized to bridge the gap between the source and target domains. The optimized AP is injected into the latent feature and facilitates the adaptation of the pre-trained model to the target domain. Moreover, we propose a Spectral/Geometric Consistency learning framework to optimize the AP in an unsupervised manner. This promotes consistent predictions across two types of transformations: geometric and frequency-space spectral transformations, enhancing robustness to both variations. Extensive experiments with multiple medical segmentation datasets demonstrate the effectiveness of APs.
Hong Joo Lee 0001, Yuan Bi, Sangmin Lee 0001, Gyeong-Moon Park, Jung Uk Kim, Seong Tae Kim 0001, Zhongliang Jiang, Nassir Navab
Medical Image Anal.4
2026 SSMPD: Semi-Supervised Learning for Multispectral Pedestrian Detection
abstract
Pedestrian detection is a crucial task in computer vision. Utilizing multispectral knowledge, especially, is essential to effectively detect the pedestrians. Existing multispectral pedestrian detection methods, however, perform only in fully-supervised situations. Although studies on semi-supervised object detection have been conducted, they focus only on single modality environments. Therefore, we propose novel semi-supervised multispectral pedestrian detector (SSMPD) that effectively utilizes multispectral knowledge. Our SSMPD consists of three methods that effectively address the pseudo-labels in the multispectral domain and a novel data selection method. First, we introduce a Pedestrian Appearance-Aware (PAA) weight to consider the quality of the pseudo-label by adjusting the multispectral knowledge transfer from the teacher model to the student model. Second, we propose a Unified Modal-Aware Simultaneous (UMAS) learning to consider the single modality (visible or thermal) and multispectral modalities when learning with the pseudo-label. Finally, we introduce a Similarity-based Contrastive (SC) loss to guide the teacher model in enhancing the quality of pseudo-labels. In addition, we provide diverse data selection for more effective semi-supervised learning. Extensive experimental results on the KAIST and LLVIP datasets demonstrate the effectiveness of our method.
Seungho Shin, Gyeong-Moon Park, Jung Uk Kim
IEEE Trans. Multim.3
2025 Multispectral Pedestrian Detection with Sparsely Annotated Label
abstract
Although existing Sparsely Annotated Object Detection (SAOD) approches have made progress in handling sparsely annotated environments in multispectral domain, where only some pedestrians are annotated, they still have the following limitations: (i) they lack considerations for improving the quality of pseudo-labels for missing annotations, and (ii) they rely on fixed ground truth annotations, which leads to learning only a limited range of pedestrian visual appearances in the multispectral domain. To address these issues, we propose a novel framework called Sparsely Annotated Multispectral Pedestrian Detection (SAMPD). For limitation (i), we introduce Multispectral Pedestrian-aware Adaptive Weight (MPAW) and Positive Pseudo-label Enhancement (PPE) module. Utilizing multispectral knowledge, these modules ensure the generation of high-quality pseudo-labels and enable effective learning by increasing weights for high-quality pseudo-labels based on modality characteristics. To address limitation (ii), we propose an Adaptive Pedestrian Retrieval Augmentation (APRA) module, which adaptively incorporates pedestrian patches from ground-truth and dynamically integrates high-quality pseudo-labels with the ground-truth, facilitating a more diverse learning pool of pedestrians. Extensive experimental results demonstrate that our SAMPD significantly enhances performance in sparsely annotated environments within the multispectral domain.
Seungho Shin, Gyeong-Moon Park, Jung Uk Kim
AAAI3
2025 Universal Domain Adaptation for Semantic Segmentation
abstract
Unsupervised domain adaptation for semantic segmentation (UDA-SS) aims to transfer knowledge from labeled source data to unlabeled target data. However, traditional UDA-SS methods assume that category settings between source and target domains are known, which is unrealistic in real-world scenarios. This leads to performance degradation if private private classes exist. To address this limitation, we propose Universal Domain Adaptation for Semantic Segmentation (UniDA-SS), achieving robust adaptation even without prior knowledge of category settings. We define the problem in the UniDA-SS scenario as low confidence scores of common classes in the target domain, which leads to confusion with private classes. To solve this problem, we propose UniMAP: UniDA-SS with Image Matching and Prototype-based Distinction, a novel framework composed of two key components. First, Domain-Specific Prototype-based Distinction (DSPD) divides each class into two domain-specific prototypes, enabling finer separation of domain-specific features and enhancing the identification of common classes across domains. Second, Target-based Image Matching (TIM) selects a source image containing the most common-class pixels based on the target pseudo-label and pairs it in a batch to promote effective learning of common classes. We also introduce a new UniDA-SS benchmark and demonstrate through various experiments that UniMAP significantly outperforms baselines. The code is available at https://github.com/KU-VGI/UniMAP.
Seun-An Choe, Keon-Hee Park, Jinwoo Choi 0001, Gyeong-Moon Park
CVPR4
2025 ESC: Erasing Space Concept for Knowledge Deletion
abstract
As concerns regarding privacy in deep learning continue to grow, individuals are increasingly apprehensive about the potential exploitation of their personal knowledge in trained models. Despite several research efforts to address this, they often fail to consider the real-world demand from users for complete knowledge erasure. Furthermore, our investigation reveals that existing methods have a risk of leaking personal knowledge through embedding features. To address these issues, we introduce a novel concept of Knowledge Deletion (KD), an advanced task that considers both concerns, and provides an appropriate metric, named Knowledge Retention score (KR), for assessing knowledge retention in feature space. To achieve this, we propose a novel training-free erasing approach named Erasing Space Concept (ESC), which restricts the important subspace for the forgetting knowledge by eliminating the relevant activations in the feature. In addition, we suggest ESC with Training (ESC-T), which uses a learnable mask to better balance the trade-off between forgetting and preserving knowledge in KD. Our extensive experiments on various datasets and models demonstrate that our proposed methods achieve the fastest and state-of-the-art performance. Notably, our methods are applicable to diverse forgetting scenarios, such as facial domain setting, demonstrating the generalizability of our methods. The code is available at https://github.com/KU-VGI/ESC.
Tae-Young Lee, Sundong Park, Minwoo Jeon, Hyoseok Hwang, Gyeong-Moon Park
CVPR5
2025 Test-Time Fine-Tuning of Image Compression Models for Multi-Task Adaptability
abstract
The field of computer vision was initially inspired by the human visual system and has progressively expanded to include a broader range of machine vision applications. Consequently, image compressors should be designed to effectively accommodate not only human visual perception but also machine vision tasks, including closed-set scenarios that enable pre-training and open-set scenarios that involve previously unseen tasks at test time. Many recent studies effectively address both human visual perception and closed-set machine vision tasks simultaneously but struggle to handle open-set machine vision tasks. To address this issue, this paper proposes a fully instance-specific test time fine-tuning (TTFT) for adapting learned image compression (LIC) to both closed-set and open-set machine vision tasks effectively. With our method, a large-scale LIC model, originally trained for human perception, is adapted to the target task through TTFT using Singular Value Decomposition based Low Rank Adaptation (SVD-LoRA). During TTFT, the decoder adopts a modified learning scheme that focuses exclusively on training the singular values, which helps prevent excessive bitstream overhead. This enables fully instance-specific optimization for the target task, even for open-set tasks. Experimental results demonstrate that the proposed method effectively adapts the backbone compressor to diverse machine vision tasks, outperforming competing methods. The code is available at project page.
Unki Park, Seongmoon Jeong, Youngchan Jang, Gyeong-Moon Park, Jong Hwan Ko
CVPR4
2025 ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
Jongseo Lee, Kyungho Bae, Kyle Min 0001, Gyeong-Moon Park, Jinwoo Choi 0001
ICCV4
2025 GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head Avatar
abstract
Despite recent progress in 3D head avatar generation, balancing identity preservation, i.e., reconstruction, with novel poses and expressions, i.e., animation, remains a challenge. Existing methods struggle to adapt Gaussians to varying geometrical deviations across facial regions, resulting in suboptimal quality. To address this, we propose GeoAvatar, a framework for adaptive geometrical Gaussian Splatting. GeoAvatar leverages Adaptive Pre-allocation Stage (APS), an unsupervised method that segments Gaussians into rigid and flexible sets for adaptive offset regularization. Then, based on mouth anatomy and dynamics, we introduce a novel mouth structure and the part-wise deformation strategy to enhance the animation fidelity of the mouth. Finally, we propose a regularization loss for precise rigging between Gaussians and 3DMM faces. Moreover, we release DynamicFace, a video dataset with highly expressive facial motions. Extensive experiments show the superiority of GeoAvatar compared to state-of-the-art methods in reconstruction and novel animation scenarios.
Seungjun Moon, Hah Min Lew, Seungeun Lee, Ji-Su Kang, Gyeong-Moon Park
ICCV5
2025 SFUOD: Source-Free Unknown Object Detection
abstract
Source-free object detection adapts a detector pre-trained on a source domain to an unlabeled target domain without requiring access to labeled source data. While this setting is practical as it eliminates the need for the source dataset during domain adaptation, it operates under the restrictive assumption that only pre-defined objects from the source domain exist in the target domain. This closed-set setting prevents the detector from detecting undefined objects. To ease this assumption, we propose Source-Free Unknown Object Detection (SFUOD), a novel scenario which enables the detector to not only recognize known objects but also detect undefined objects as unknown objects. To this end, we propose CollaPAUL (Collaborative tuning and Principal Axis-based Unknown Labeling), a novel framework for SFUOD. Collaborative tuning enhances knowledge adaptation by integrating target-dependent knowledge from the auxiliary encoder with source-dependent knowledge from the pre-trained detector through a cross-domain attention mechanism. Additionally, principal axes-based unknown labeling assigns pseudo-labels to unknown objects by estimating objectness via principal axes projection and confidence scores from model predictions. The proposed CollaPAUL achieves state-of-the-art performances on SFUOD benchmarks, and extensive experiments validate its effectiveness.
Keon-Hee Park, Seun-An Choe, Gyeong-Moon Park
ICCV3
2025 Do Not Mimic My Voice : Speaker Identity Unlearning for Zero-Shot Text-to-Speech
abstract
The rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to selectively remove the knowledge to replicate unwanted individual voices from pre-trained model parameters has not been explored. In this paper, we address the new challenge of speaker identity unlearning for ZS-TTS systems. To meet this goal, we propose the first machine unlearning frameworks for ZS-TTS, especially Teacher-Guided Unlearning (TGU), designed to ensure the model forgets designated speaker identities while retaining its ability to generate accurate speech for other speakers. Our proposed methods incorporate randomness to prevent consistent replication of forget speakers' voices, assuring unlearned identities remain untraceable. Additionally, we propose a new evaluation metric, speaker-Zero Retrain Forgetting (spk-ZRF). This assesses the model's ability to disregard prompts associated with forgotten speakers, effectively neutralizing its knowledge of these voices. The experiments conducted on the state-of-the-art model demonstrate that TGU prevents the model from replicating forget speakers' voices while maintaining high quality for other speakers. The demo is available at https://speechunlearn.github.io/ .
Taesoo Kim, Jinju Kim, Jong Hwan Ko, Gyeong-Moon Park
ICML5
2025 When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series
abstract
Recently, forecasting future abnormal events has emerged as an important scenario to tackle realworld necessities. However, the solution of predicting specific future time points when anomalies will occur, known as Anomaly Prediction (AP), remains under-explored. Existing methods dealing with time series data fail in AP, focusing only on immediate anomalies or failing to provide precise predictions for future anomalies. To address AP, we propose a novel framework called Anomaly to Prompt (A2P), comprised of Anomaly-Aware Forecasting (AAF) and Synthetic Anomaly Prompting (SAP). To enable the forecasting model to forecast abnormal time points, we adopt a strategy to learn the relationships of anomalies. For the robust detection of anomalies, our proposed SAP introduces a learnable Anomaly Prompt Pool (APP) that simulates diverse anomaly patterns using signal-adaptive prompt. Comprehensive experiments on multiple real-world datasets demonstrate the superiority of A2P over state-of-the-art methods, showcasing its ability to predict future anomalies.
Min-Yeong Park, Won-Jeong Lee, Seong Tae Kim 0001, Gyeong-Moon Park
ICML4
2025 Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
abstract
Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods—based on saliency—produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Language-based approaches offer structure but often fail to explain motions due to their tacit nature—intuitively understood but difficult to verbalize. To address these challenges, we propose Disentangled Action aNd Context concept-based Explainable (DANCE) video action recognition, a framework that predicts actions through disentangled concept types: motion dynamics, objects, and scenes. We define motion dynamics concepts as human pose sequences. We employ a large language model to automatically extract object and scene concepts. Built on an ante-hoc concept bottleneck design, DANCE enforces prediction through these concepts. Experiments on four datasets—KTH, Penn Action, HAA500, and UCF101—demonstrate that DANCE significantly improves explanation clarity with competitive performance. Through a user study, we validate the superior interpretability of DANCE. Experimental results also show that DANCE is beneficial for model debugging, editing, and failure analysis.
Jongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim 0001, Jinwoo Choi 0001
NeurIPS3
2025 Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion Models
abstract
Recent advances in diffusion models have enabled high-quality synthesis of specific subjects, such as identities or objects. This capability, while unlocking new possibilities in content creation, also introduces significant privacy risks, as personalization techniques can be misused by malicious users to generate unauthorized images. Although several studies have attempted to counter this by generating adversarially perturbed samples designed to disrupt personalization, they rely on unrealistic assumptions and become ineffective in the presence of even a few clean images or under simple image transformations. To address these challenges, we shift the protection target from the images to the diffusion model itself to hinder the personalization of specific subjects, through our novel framework called $\textbf{A}$nti-$\textbf{P}$ersonalized $\textbf{D}$iffusion $\textbf{M}$odels ($\textbf{APDM}$). We first provide a theoretical analysis demonstrating that a naive approach of existing loss functions to diffusion models is inherently incapable of ensuring convergence for robust anti-personalization. Motivated by this finding, we introduce Direct Protective Optimization (DPO), a novel loss function that effectively disrupts subject personalization in the target model without compromising generative quality. Moreover, we propose a new dual-path optimization strategy, coined Learning to Protect (L2P). By alternating between personalization and protection paths, L2P simulates future personalization trajectories and adaptively reinforces protection at each step. Experimental results demonstrate that our framework outperforms existing methods, achieving state-of-the-art performance in preventing unauthorized personalization. The code is available at https://github.com/KU-VGI/APDM.
Tae-Young Lee, Juwon Seo, Jong Hwan Ko, Gyeong-Moon Park
NeurIPS4
2025 WINE: Wavelet-Guided GAN Inversion and Editing for High-Fidelity Refinement
abstract
Recent advanced GAN inversion models aim to convey high-fidelity information from original images to generators through methods using generator tuning or high-dimensional feature learning. Despite these efforts, accurately reconstructing image-specific details remains as a challenge due to the inherent limitations both in terms of training and structural aspects, leading to a bias towards low-frequency information. In this paper, we look into the widely used pixel loss in GAN inversion, revealing its predominant focus on the reconstruction of low-frequency features. We then propose WINE, a Wavelet-guided GAN Inversion aNd Editing model, which transfers the high-frequency information through wavelet coefficients via newly proposed wavelet loss and wavelet fusion scheme. Notably, WINE is the first attempt to interpret GAN inversion in the frequency domain. Our experimental results showcase the precision of WINE in preserving high-frequency details and enhancing image quality. Even in editing scenarios, WINE outperforms existing state-of-the-art GAN inversion models with a fine balance between editability and reconstruction quality. Pre-trained model and codes will be publicized after the review process.
Seung Jun Moon, Gyeong-Moon Park
WACV3
2025 Towards High-fidelity Head Blending with Chroma Keying for Industrial Applications
abstract
We introduce an industrial Head Blending pipeline for the task of seamlessly integrating an actor's head onto a target body in digital content creation. The key challenge stems from discrepancies in head shape and hair structure, which lead to unnatural boundaries and blending artifacts. Existing methods treat foreground and back-ground as a single task, resulting in suboptimal blending quality. To address this problem, we propose CHANGER, a novel pipeline that decouples background integration from foreground blending. By utilizing chroma keying for artifact-free background generation and introducing Head shape and long Hair augmentation (H2augmen-tation) to simulate a wide range of head shapes and hair styles, CHANGER improves generalization on innumerable various real-world cases. Furthermore, our Fore-ground Predictive Attention Transformer (FPAT) module enhances foreground blending by predicting and focusing on key head and body regions. Quantitative and qualitative evaluations on benchmark datasets demonstrate that our CHANGER outperforms state-of-the-art methods, de-livering high-fidelity, industrial-grade results.
Hah Min Lew, Sahng-Min Yoo, Hyunwoo Kang, Gyeong-Moon Park
WACV4
2024 Open-Set Domain Adaptation for Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) for semantic segmentation aims to transfer the pixel-wise knowledge from the labeled source domain to the unlabeled target do-main. However, current UDA methods typically assume a shared label space between source and target, limiting their applicability in real-world scenarios where novel cat-egories may emerge in the target domain. In this paper, we introduce Open-Set Domain Adaptation for Semantic Segmentation (OSDA -SS) for the first time, where the target domain includes unknown classes. We identify two major problems in the OSDA -SS scenario as follows: 1) the existing UDA methods struggle to predict the exact boundary of the unknown classes, and 2) they fail to accurately predict the shape of the unknown classes. To address these issues, we propose Boundary and Unknown Shape-Aware open-set domain adaptation, coined BUS. Our BUS can accu-rately discern the boundaries between known and unknown classes in a contrastive manner using a novel dilation-erosion-based contrastive loss. In addition, we propose OpenReMix, a new domain mixing augmentation method that guides our model to effectively learn domain and size-invariant features for improving the shape detection of the known and unknown classes. Through extensive experiments, we demonstrate that our proposed BUS effectively detects unknown classes in the challenging OSDA-SS sce-nario compared to the previous methods by a large margin. The code is available at https://github.com/KHUAGI/BUS.
Seun-An Choe, Ah-Hyung Shin, Keon-Hee Park, Jinwoo Choi 0001, Gyeong-Moon Park
CVPR5
2024 Pre-trained Vision and Language Transformers are Few-Shot Incremental Learners
abstract
Few-Shot Class Incremental Learning (FSCIL) is a task that requires a model to learn new classes incrementally without forgetting when only a few samples for each class are given. FSCIL encounters two significant challenges: catastrophic forgetting and overfitting, and these challenges have driven prior studies to primarily rely on shallow models, such as ResNet-18. Even though their limited capacity can mitigate both forgetting and overfitting issues, it leads to inadequate knowledge transfer during few-shot in-cremental sessions. In this paper, we argue that large models such as vision and language transformers pre-trained on large datasets can be excellent few-shot incremental learn-ers. To this end, we propose a novel FSCIL framework called PriViLege, Pre-trained Vision and Language trans-formers with prompting functions and knowledge distillation. Our framework effectively addresses the challenges of catastrophic forgetting and overfitting in large models through new pre-trained knowledge tuning (PKT) and two losses: entropy-based divergence loss and semantic knowl-edge distillation loss. Experimental results show that the proposed PriViLege significantly outperforms the existing state-of-the-art methods with a large margin, e.g., +9.38% in CUB200, +20.58% in CIFAR-100, and +13.36% in miniImageNet. Our implementation code is available at https://github.com/KHU-AGI/PriViLege.
Keon-Hee Park, Kyungwoo Song, Gyeong-Moon Park
CVPR3
2024 Generative Unlearning for Any Identity
abstract
Recent advances in generative models trained on large-scale datasets have made it possible to synthesize highquality samples across various domains. Moreover, the emergence of strong inversion networks enables not only a reconstruction of real-world images but also the modification of attributes through various editing methods. However, in certain domains related to privacy issues, e.g., human faces, advanced generative models along with strong inversion methods can lead to potential misuses. In this paper, we propose an essential yet under-explored task called generative identity unlearning, which steers the model not to generate an image of specific identity. In the generative identity unlearning, we target the following objectives: (i) preventing the generation of images with a certain identity, and (ii) preserving the overall quality of the generative model. To satisfy these goals, we propose a novel framework, Generative Unlearning for Any IDEntity (GUIDE), which prevents the reconstruction of a specific identity by unlearning the generator with only a single image. GUIDE consists of two parts: (i) finding a target point for optimization that unidentifies the source latent code and (ii) novel loss functions that facilitate the unlearning procedure while less affecting the learned distribution. Our extensive experiments demonstrate that our proposed method achieves state-of-the-art performance in the generative machine unlearning task. The code is available at https://github.com/KHU-AGI/GUIDE.
Juwon Seo, Sung-Hoon Lee, Tae-Young Lee, Seungjun Moon, Gyeong-Moon Park
CVPR5
2024 Towards Model-Agnostic Dataset Condensation by Heterogeneous Models
Jun-Yeong Moon, Jung Uk Kim, Gyeong-Moon Park
ECCV (29)3
2024 Versatile Incremental Learning: Towards Class and Domain-Agnostic Incremental Learning
Min-Yeong Park, Gyeong-Moon Park
ECCV (31)3
2024 Online Continuous Generalized Category Discovery
Keon-Hee Park, Hakyung Lee, Kyungwoo Song, Gyeong-Moon Park
ECCV (81)4
2024 GLAD: Global-Local View Alignment and Background Debiasing for Unsupervised Video Domain Adaptation with Large Domain Gap
abstract
In this work, we tackle the challenging problem of unsupervised video domain adaptation (UVDA) for action recognition. We specifically focus on scenarios with a substantial domain gap, in contrast to existing works primarily deal with small domain gaps between labeled source domains and unlabeled target domains. To establish a more realistic setting, we introduce a novel UVDA scenario, denoted as Kinetics→BABEL, with a more considerable domain gap in terms of both temporal dynamics and background shifts. To tackle the temporal shift, i.e., action duration difference between the source and target domains, we propose a global-local view alignment approach. To mitigate the background shift, we propose to learn temporal order sensitive representations by temporal order learning and background invariant representations by background augmentation. We empirically validate that the proposed method shows significant improvement over the existing methods on the Kinetics→BABEL dataset with a large domain gap. The code is available at https://github.com/KHU-VLL/GLAD.
Hyogun Lee, Kyungho Bae, Seong Jong Ha, Yumin Ko, Gyeong-Moon Park, Jinwoo Choi 0001
WACV5
2024 RADIO: Reference-Agnostic Dubbing Video Synthesis
abstract
One of the most challenging problems in audio-driven talking head generation is achieving high-fidelity detail while ensuring precise synchronization. Given only a single reference image, extracting meaningful identity attributes becomes even more challenging, often causing the network to mirror the facial and lip structures too closely. To address these issues, we introduce RADIO, a framework engineered to yield high-quality dubbed videos regardless of the pose or expression in reference images. The key is to modulate the decoder layers using latent space composed of audio and reference features. Additionally, we incorporate ViT blocks into the decoder to emphasize high-fidelity details, especially in the lip region. Our experimental results demonstrate that RADIO displays high synchronization without the loss of fidelity. Especially in harsh scenarios where the reference frame deviates significantly from the ground truth, our method outperforms state-of-the-art methods, highlighting its robustness.
Dongyeun Lee, Sangjoon Yu, Jaejun Yoo 0001, Gyeong-Moon Park
WACV5
2024 Wav2NeRF: Audio-driven realistic talking head generation via wavelet-based NeRF
abstract
Talking head generation is an essential task in various real-world applications such as film making and virtual reality. To this end, recent works focus on the NeRF-based methods that can capture the 3D structural information of faces and generate more natural and vivid talking videos. However, the existing NeRF-based methods fail to accurately generate the audio-synced videos. In this paper, we point out that the previous methods do not consider the audio-visual representations explicitly, which is crucial for precise lip synchronization. Moreover, the existing methods struggle to generate high-frequency details, making the generation results unnatural. To overcome these problems, we propose a novel audio-synced and high-fidelity NeRF-based talking head generation framework, named Wav2NeRF, which learns audio-visual cross-modality representations and employs the wavelet transform for better visual quality. In precise, we adopt a 2D CNN-based neural rendering decoder to a NeRF-based encoder for fast generation of the whole image to employ a new multi-level SyncNet loss for accurate lip synchronization. We also propose a novel cross-attention module to effectively fuse the image and the audio representation. In addition, we integrate the wavelet transform into our framework by proposing the wavelet loss function to enhance high-frequency details. We demonstrate that the proposed method renders realistic and audio-synced talking head videos and shows outstanding performances on average in 4 representative metrics, including PSNR (+ 4.7%), SSIM (+ 2.2%), LMD (+ 51.3%), and SyncNet Confidence (+ 154.7%) compared to the NeRF-based current state-of-the-art methods.
Ah-Hyung Shin, Jiwon Hwang, Yoonhyung Kim, Gyeong-Moon Park
Image Vis. Comput.5
2023 LINe: Out-of-Distribution Detection by Leveraging Important Neurons
abstract
It is important to quantify the uncertainty of input samples, especially in mission-critical domains such as autonomous driving and healthcare, where failure predictions on out-of-distribution (OOD) data are likely to cause big problems. OOD detection problem fundamentally begins in that the model cannot express what it is not aware of. Post-hoc OOD detection approaches are widely explored because they do not require an additional re-training process which might degrade the model's performance and increase the training cost. In this study, from the perspective of neurons in the deep layer of the model representing high-level features, we introduce a new aspect for analyzing the difference in model outputs between in-distribution data and OOD data. We propose a novel method, Leveraging Important Neurons (LINe), for post-hoc Out of distribution detection. Shapley value-based pruning reduces the effects of noisy outputs by selecting only high-contribution neurons for predicting specific classes of input data and masking the rest. Activation clipping fixes all values above a certain threshold into the same value, allowing LINe to treat all the class-specific features equally and just consider the difference between the number of activated feature differences between in-distribution and OOD data. Comprehensive experiments verify the effectiveness of the proposed method by outperforming state-of-the-art post-hoc OOD detection methods on CIFAR-10, CIFAR-100, and ImageNet datasets. Code is available on https://github.com/LINe-OOD
Yong Hyun Ahn, Gyeong-Moon Park, Seong Tae Kim 0001
CVPR2
2023 Online Class Incremental Learning on Stochastic Blurry Task Boundary via Mask and Visual Prompt Tuning
abstract
Continual learning aims to learn a model from a continuous stream of data, but it mainly assumes a fixed number of data and tasks with clear task boundaries. However, in real-world scenarios, the number of input data and tasks is constantly changing in a statistical way, not a static way. Although recently introduced incremental learning scenarios having blurry task boundaries somewhat address the above issues, they still do not fully reflect the statistical properties of real-world situations because of the fixed ratio of disjoint and blurry samples. In this paper, we propose a new Stochastic incremental Blurry task boundary scenario, called Si-Blurry, which reflects the stochastic properties of the real-world. We find that there are two major challenges in the Si-Blurry scenario: (1) intra- and inter-task forget-tings and (2) class imbalance problem. To alleviate them, we introduce Mask and Visual Prompt tuning (MVP). In MVP, to address the intra- and inter-task forgetting issues, we propose a novel instance-wise logit masking and contrastive visual prompt tuning loss. Both of them help our model discern the classes to be learned in the current batch. It results in consolidating the previous knowledge. In addition, to alleviate the class imbalance problem, we introduce a new gradient similarity-based focal loss and adaptive feature scaling to ease overfitting to the major classes and underfitting to the minor classes. Extensive experiments show that our proposed MVP significantly outperforms the existing state-of-the-art methods in our challenging Si-Blurry scenario. The code is available at https://github.com/moonjunyyy/Si-Blurry
Jun-Yeong Moon, Keon-Hee Park, Jung Uk Kim, Gyeong-Moon Park
ICCV4
2023 LFS-GAN: Lifelong Few-Shot Image Generation
abstract
We address a challenging lifelong few-shot image generation task for the first time. In this situation, a generative model learns a sequence of tasks using only a few samples per task. Consequently, the learned model encounters both catastrophic forgetting and overfitting problems at a time. Existing studies on lifelong GANs have proposed modulation-based methods to prevent catastrophic forgetting. However, they require considerable additional parameters and cannot generate high-fidelity and diverse images from limited data. On the other hand, the existing few-shot GANs suffer from severe catastrophic forgetting when learning multiple tasks. To alleviate these issues, we propose a framework called Lifelong Few-Shot GAN (LFS-GAN) that can generate high-quality and diverse images in lifelong few-shot image generation task. Our proposed framework learns each task using an efficient task-specific modulator - Learnable Factorized Tensor (LeFT). LeFT is rank-constrained and has a rich representation ability due to its unique reconstruction technique. Furthermore, we propose a novel mode seeking loss to improve the diversity of our model in low-data circumstances. Extensive experiments demonstrate that the proposed LFS-GAN can generate high-fidelity and diverse images without any forgetting and mode collapse in various domains, achieving state-of-the-art in lifelong few-shot image generation task. Surprisingly, we find that our LFS-GAN even outperforms the existing few-shot GANs in the few-shot image generation task. The code is available at Github.
Juwon Seo, Ji-Su Kang, Gyeong-Moon Park
ICCV3
2022 IntereStyle: Encoding an Interest Region for Robust StyleGAN Inversion
Seung Jun Moon, Gyeong-Moon Park
ECCV (15)2
2022 SR-EM: Episodic Memory Aware of Semantic Relations Based on Hierarchical Clustering Resonance Network
abstract
An intelligent robot requires episodic memory that can retrieve a sequence of events for a service task learned from past experiences to provide a proper service to a user. Various episodic memories, which can learn new tasks incrementally without forgetting the tasks learned previously, have been designed based on adaptive resonance theory (ART) networks. The conventional ART-based episodic memories, however, do not have the adaptability to the changing environments. They cannot utilize the retrieved task episode adaptively in the working environment. Moreover, if a user wants to receive multiple services of the same kind in a given situation, the user should repeatedly command multiple times. To tackle these limitations, in this article, a novel hierarchical clustering resonance network (HCRN) is proposed, which has a high clustering performance on multimodal data and can compute the semantic relations between learned clusters. Using HCRN, a semantic relation-aware episodic memory (SR-EM) is designed, which can adapt the retrieved task episode to the current working environment to carry out the task intelligently. Experimental simulations demonstrate that HCRN outperforms the conventional ART in terms of clustering performance on multimodal data. Besides, the effectiveness of the proposed SR-EM is verified through robot simulations for two scenarios.
Jaewoo Choi 0001, Gyeong-Moon Park, Jong-Hwan Kim 0001
IEEE Trans. Cybern.2
2021 Adaptive Developmental Resonance Network
abstract
Adaptive resonance theory (ART) networks, including developmental resonance network (DRN), basically use a vigilance parameter as a hyperparameter to determine whether a current input can belong to any existing categories or not. The problem here is that the clustering quality of those networks is sensitive to the vigilance parameter so that the users are required to fine-tune the parameter delicately beforehand. Another problem is that those networks only deal with a hyperrectangular decision boundary, which means they cannot learn categories of arbitrary shape. In addition, the order of data processing is a critical factor to categorize clusters correctly because each category can expand its boundary into the areas of other categories erroneously. To deal with these problems, we propose an advanced version of DRN, Adaptive DRN (A-DRN), which learns the vigilance parameters assigned for individual category nodes as well as category weights. The proposed A-DRN combines close categories to construct a cluster that contains the categories identifying a cluster boundary of arbitrary shape. Our A-DRN also employs a sliding window. The sliding window buffers sequential data points to presume the data distribution roughly, which helps our network to have a robust and consistent performance to a random order of input data. Through the experiments, we empirically demonstrate the effectiveness of A-DRN in both synthetic and real-world benchmark data sets.
Gyeong-Moon Park, Jong-Hwan Kim 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Convolutional Neural Network With Developmental Memory for Continual Learning
abstract
Convolutional neural networks (CNNs) are one of the most successful deep neural networks. Indeed, most of the recent applications related to computer vision are based on CNNs. However, when learning new tasks in a sequential manner, CNNs face catastrophic forgetting: they forget a considerable amount of previously learned tasks while adapting to novel tasks. To overcome this main barrier to continual learning with CNNs, we introduce developmental memory (DM) into a CNN, continually generating submemory networks to learn important features of individual tasks. A novel training method, referred to here as guided learning (GL), guides the newly generated submemory to become an expert on the new task, eventually improving the performance of the overall network. At the same time, the existing submemories attempt to preserve the knowledge of old tasks. Experiments on image classification tasks show that compared with the state-of-the-art algorithms, the proposed CNN with DM not only improves the classification performance on the new image task but also leads to less forgetting of previous image tasks to facilitate continual learning.
Gyeong-Moon Park, Sahng-Min Yoo, Jong-Hwan Kim 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Non-Probabilistic Cosine Similarity Loss for Few-Shot Image Classification
Joonhyuk Kim, Inug Yoon, Gyeong-Moon Park, Jong-Hwan Kim 0001
BMVC3
2019 Developmental Resonance Network
abstract
Adaptive resonance theory (ART) networks deal with normalized input data only, which means that they need the normalization process for the raw input data, under the assumption that the upper and lower bounds of the input data are known in advance. Without such an assumption, ART networks cannot be utilized. To solve this problem and improve the learning performance, inspired by the ART networks, we propose a developmental resonance network (DRN) by employing new techniques of a global weight and node connection and grouping processes. The proposed DRN learns the global weight converging to the unknown range of the input data and properly clusters by grouping similar nodes into one. These techniques enable DRN to learn the raw input data without the normalization process while retaining the stability, plasticity, and memory usage efficiency without node proliferation. Simulation results verify that our DRN, applied to the unsupervised clustering problem, can cluster raw data properly without a prior normalization process.
Gyeong-Moon Park, Jaewoo Choi 0001, Jong-Hwan Kim 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Deep ART Neural Model for Biologically Inspired Episodic Memory and Its Application to Task Performance of Robots
abstract
Robots are expected to perform smart services and to undertake various troublesome or difficult tasks in the place of humans. Since these human-scale tasks consist of a temporal sequence of events, robots need episodic memory to store and retrieve the sequences to perform the tasks autonomously in similar situations. As episodic memory, in this paper we propose a novel Deep adaptive resonance theory (ART) neural model and apply it to the task performance of the humanoid robot, Mybot, developed in the Robot Intelligence Technology Laboratory at KAIST. Deep ART has a deep structure to learn events, episodes, and even more like daily episodes. Moreover, it can retrieve the correct episode from partial input cues robustly. To demonstrate the effectiveness and applicability of the proposed Deep ART, experiments are conducted with the humanoid robot, Mybot, for performing the three tasks of arranging toys, making cereal, and disposing of garbage.
Gyeong-Moon Park, Yong-Ho Yoo, Deok-Hwa Kim, Jong-Hwan Kim 0001
IEEE Trans. Cybern.1
2017 Context preference-based deep adaptive resonance theory: Integrating user preferences into episodic memory encoding and retrieval
abstract
Episodic memory which can store and recall episodes has been modeled by various research. Those models focus on encoding and retrieving the same sequence of events of episodes. In this paper, we propose context preference-based deep adaptive resonance theory (CPD-ART). CPD-ART uses a new approach in encoding and retrieving a temporal sequence of events considering subjects, preference criteria such as weather, and object contexts such as beverage. A new layer, context preference field, is added to the encoding and retrieval processes for decision making. Context preference field encodes and stores the knowledge of criteria and object contexts, along with their relations in probability weight vectors. Simulation results demonstrate that CPD-ART is able to conduct decision making analysis and retrieve the sequence of events of an episode correctly through decision making analysis based on subjects, preference criteria, and the object contexts.
Dick Sigmund, Gyeong-Moon Park, Jong-Hwan Kim 0001
IJCNN2
2016 Deep Adaptive Resonance Theory for learning biologically inspired episodic memory
abstract
Biologically inspired episodic memory is able to store time sequential events, and to recall all of them from partial information. Because of the advantages of episodic memory, the biological concepts of episodic memory have been utilized to many applications. In this research, we propose a new memory model, called Deep ART (Adaptive Resonance Theory), to make a robust memory system for learning episodic memory. Deep ART has an attribute field in the bottom layer, which is newly designed to get semantic information of inputs. After encoding all inputs with their features, events are categorized in the event field using specified inputs. Since an episode is made of a temporal sequence of events, Deep ART makes event sequences with proposed sequence encoding and decoding processes. They can encode any temporal sequence of events, even if there are duplicated events in the episode. Moreover, based on the result of the analysis of retrieval error, Deep ART does not use the complement coding for partial inputs to enhance the accuracy of episode retrieval from partial cues. The simulation results demonstrate the effectiveness of Deep ART as the long term memory.
Gyeong-Moon Park, Jong-Hwan Kim 0001
IJCNN1
2016 Biologically-inspired episodic memory model considering the context information
abstract
Episodic memory can store time sequential events and retrieve them anytime with specific cues. However, if the episodic memory only stores events comprised of actions and objects, execution of episodes may fail if current situation is different from the settings it learned in. As a solution, we propose Deep C-ART (Context-Adaptive Resonance Theory) which considers not only time sequential events but also their contexts. In addition to the learning process of Deep ART, Deep C-ART stores context information such as situation of objects, states of robots, place, and time of episodes. Since context changes over each event in an episode, Deep C-ART forms an episode with an event sequence and a context sequence. During retrieval and execution of episode, it compares the current situation with the learned one to verify that it is executable or in an anomaly situation. The effectiveness of Deep C-ART is demonstrated through computer simulations.
Gyeong-Moon Park, Sanghyun Cho, Jong-Hwan Kim 0001
SMC1