VLDB 2026 Research / reviewers in the wild / expert
Kyungwoo Song
dblp:155/4867
· DBLP profile ↗
57ranked-venue papers
7as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 5 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing LLMs for Manufacturing Information Extraction
Subeen Park, Hakyung Lee, Ryunyi Lee, Hyo-won Suh, Kyungwoo Song |
PAKDD (4) | 5 |
| 2026 | DART: Diffusion-based approach for predicting trajectories of ballistic missiles
Joonseong Kang, Seonggyun Lee, Jeyoon Yeom, Dongwg Hong, Kyungwoo Song |
Expert Syst. Appl. | 6 |
| 2026 | ReCoD: Enhancing image description for cross-modal understanding via retrieval and comparison feedback mechanism
Geunyoung Jung, Jun Park, Hankyeol Lee, Kyungwoo Song, Jiyoung Jung |
Neurocomputing | 4 |
| 2026 | Robust Adaptation of Foundation Models With Black-Box Visual PromptingabstractWith a surge of large-scale pre-trained models, parameter-efficient transfer learning (PETL) of large models has garnered significant attention. While promising, they commonly rely on two optimistic assumptions: 1) full access to the parameters of a PTM and 2) sufficient memory capacity to cache all intermediate activations for gradient computation. However, in most real-world applications, PTMs serve as black-box APIs or proprietary software without full parameter accessibility. Besides, it is hard to meet a large memory requirement for modern PTMs. This work proposes black-box visual prompting (BlackVIP), which efficiently adapts the PTMs without knowledge of their architectures or parameters. BlackVIP has two components: 1) Coordinator and 2) simultaneous perturbation stochastic approximation with gradient correction (SPSA-GC). The Coordinator designs input-dependent visual prompts, which allow the target PTM to adapt in the wild. SPSA-GC efficiently estimates the gradient of PTM to update Coordinator. Besides, we introduce a variant, BlackVIP-SE, which significantly reduces the runtime and computational cost of BlackVIP. Extensive experiments on 19 datasets demonstrate that BlackVIPs enable robust adaptation to diverse domains and tasks with minimal memory requirements. We further provide a theoretical analysis on the generalization of visual prompting methods by presenting their connection to the certified robustness of randomized smoothing, and presenting an empirical support for improved robustness. Changdae Oh, Gyeongdeok Seo, Geunyoung Jung, Zhi-Qi Cheng, Hosik Choi, Jiyoung Jung, Kyungwoo Song |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Graph Perceiver IO: A general architecture for graph-structured data
Seyun Bae, Hoyoon Byun, Changdae Oh, Yoon-Sik Cho, Kyungwoo Song |
Pattern Recognit. | 5 |
| 2026 | Data Adaptive Stochastic Ensemble Net: Optimizing Infection Predictions for COVID-19 Cluster AnalysisabstractMachine learning has garnered significant interest and is extensively utilized in the medical field due to its direct impact on human life. Two components are necessary to develop an AI-based infection prediction assistance system: a training dataset and machine learning prediction model. For AI-based infection prediction model, we first gathered a real-world COVID-19 cluster dataset, consisting of 8,844 confirmed cases across 519 clusters, which includes individual properties and contact relationships between confirmed cases. Second, we introduce the Data Adaptive Stochastic Ensemble Network (DASEN) to enhance prediction robustness. DASEN dynamically adjusts the weight of each component by optimizing the Dirichlet distribution concentration parameter based on the data distribution. We demonstrate the validity of DASEN, showing that different models focus on distinct features and perform well on data with varying characteristics, thus preventing overfitting to majority labels. Notably, DASEN provides superior robustness across all settings with minimal overhead for parameter optimization. Sungjun Lim 0002, YongTaek Lim, Hojun Park, Junggu Lee, Jaehun Jung, Kyungwoo Song |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | RAILL: Retrieval-Augmented Instruction Tuning for Low-Resource Language Model Training
Youngjun Choi 0001, Sungjun Lim 0002, Minhoi Park, Jaekyeong Jung, Eunsik Kim, Chul-su Kim, Kyongjae Lee, Hosik Choi, Kyungwoo Song |
IEEE Big Data | 10 |
| 2025 | Causal Effect Variational Transformer for Public Health Measures and COVID-19 Infection Cluster AnalysisabstractRecent research increasingly integrates causal inference into deep learning models to enhance the explainability and robustness of medical applications. However, data scarcity remains a fundamental challenge due to privacy constraints and the high cost of data collection. This issue, compounded by complex variable dependencies and unobserved latent confounders, hinders the reliable estimation of causal effects. To address these challenges, we collect two real-world COVID-19 infection cluster datasets, including public health measures, from distinct distributions in collaboration with local governments, a medical university, and a hospital. We also propose a cut-off augmentation method that generates diverse feature-label pairs by slicing time-series sequences at different observation windows, effectively simulating partial observations common in real-world settings. We further introduce the Causal Effect Variational Transformer (CEVT), a Transformer-based model that captures temporal structure and addresses the difficulty of causal estimation under scarce data, complex dependencies, and latent confounding by modeling multiple treatments through an iterative conditioning mechanism. We validate the causal modeling capability of CEVT on synthetic datasets and demonstrate that, on two distinct COVID-19 datasets, it consistently outperforms baselines in infection prediction. Notably, the causal effects estimated by CEVT converge with findings from medical studies on infection control, reinforcing its reliability and underscoring its potential to inform public health decision-making. Jinho Kang, Sungjun Lim 0002, Hojun Park, Jiyoung Jung, Jaehun Jung, Kyungwoo Song |
CIKM | 6 |
| 2025 | Sufficient Invariant Learning for Distribution ShiftabstractLearning robust models under distribution shifts between training and test datasets is a fundamental challenge in machine learning. While learning invariant features across environments is a popular approach, it often assumes that these features are fully observed in both training and test sets—a condition frequently violated in practice. When models rely on invariant features absent in the test set, their robustness in new environments can deteriorate. To tackle this problem, we introduce a novel learning principle called the Sufficient Invariant Learning (SIL) framework, which focuses on learning a sufficient subset of invariant features rather than relying on a single feature. After demonstrating the limitation of existing invariant learning methods, we propose a new algorithm, Adaptive Sharpness-aware Group Distributionally Robust Optimization (ASGDRO), to learn diverse invariant features by seeking common flat minima across the environments. We theoretically demonstrate that finding a common flat minima enables robust predictions based on diverse invariant features. Empirical evaluations on multiple datasets, including our new benchmark, confirm ASGDRO’s robustness against distribution shifts, highlighting the limitations of existing methods. Code: https://github.com/MLAI-Yonsei/SIL-ASGDRO. Taero Kim, Subeen Park, Sungjun Lim 0002, Yonghan Jung, Krikamol Muandet, Kyungwoo Song |
CVPR | 6 |
| 2025 | TIDES: Technical Information Discovery and Extraction SystemabstractAddressing the challenges in QA for specific technical domains requires identifying relevant portions of extensive documents and generating answers based on this focused content.Traditional pre-trained LLMs often struggle with domain-specific terminology, while fine-tuned LLMs demand substantial computational resources.To overcome these limitations, we propose TIDES, Technical Information Distillation and Extraction System.TIDES is a trainingfree approach that combines traditional TF-IDF techniques with prompt-based LLMs in a hybrid process, effectively addressing complex technical questions.It uses TF-IDF to identify and prioritize domain-specific words that are rare in other documents and LLMs to refine the candidate pool by focusing on the most relevant segments in documents through multiple stages.Our approach improves the precision and efficiency of QA systems in technical contexts without LLM retraining. Subeen Park, Hakyung Lee, YongTaek Lim, Hyo-won Suh, Kyungwoo Song |
EMNLP | 6 |
| 2025 | Brain-inspired Lp-Convolution benefits large kernels and aligns better with visual cortexabstractConvolutional Neural Networks (CNNs) have profoundly influenced the field of computer vision, drawing significant inspiration from the visual processing mechanisms inherent in the brain. Despite sharing fundamental structural and representational similarities with the biological visual system, differences in local connectivity patterns within CNNs open up an interesting area to explore. In this work, we explore whether integrating biologically observed receptive fields (RFs) can enhance model performance and foster alignment with brain representations. We introduce a novel methodology, termed $L_p$-convolution, which employs the multivariate $L_p$-generalized normal distribution as an adaptable $L_p$-masks, to reconcile disparities between artificial and biological RFs. $L_p$-masks finds the optimal RFs through task-dependent adaptation of conformation such as distortion, scale, and rotation. This allows $L_p$-convolution to excel in tasks that require flexible RF shapes, including not only square-shaped regular RFs but also horizontal and vertical ones. Furthermore, we demonstrate that $L_p$-convolution with biological RFs significantly enhances the performance of large kernel CNNs possibly by introducing structured sparsity inspired by $L_p$-generalized normal distribution in convolution. Lastly, we present that neural representations of CNNs align more closely with the visual cortex when -convolution is close to biological RFs. Jea Kwon, Sungjun Lim 0002, Kyungwoo Song, C. Justin Lee |
ICLR | 3 |
| 2025 | DaWin: Training-free Dynamic Weight Interpolation for Robust AdaptationabstractAdapting a pre-trained foundation model on downstream tasks should ensure robustness against distribution shifts without the need to retrain the whole model. Although existing weight interpolation methods are simple yet effective, we argue their static nature limits downstream performance while achieving efficiency. In this work, we propose DaWin, a training-free dynamic weight interpolation method that leverages the entropy of individual models over each unlabeled test sample to assess model expertise, and compute per-sample interpolation coefficients dynamically. Unlike previous works that typically rely on additional training to learn such coefficients, our approach requires no training. Then, we propose a mixture modeling approach that greatly reduces inference overhead raised by dynamic interpolation. We validate DaWin on the large-scale visual recognition benchmarks, spanning 14 tasks across robust fine-tuning -- ImageNet and derived five distribution shift benchmarks -- and multi-task learning with eight classification tasks. Results demonstrate that DaWin achieves significant performance gain in considered settings, with minimal computational overhead. We further discuss DaWin's analytic behavior to explain its empirical success. Changdae Oh, Yixuan Li 0001, Kyungwoo Song, Sangdoo Yun, Dongyoon Han |
ICLR | 3 |
| 2025 | Exploring the Potential of Foundation Models as Reliable AI Contact Centers
Hoyoon Byun, Minhoi Park, Seolah Kim, EunBi Kim, Kyungwoo Song |
KDD (2) | 5 |
| 2025 | LBC: Language-Based-Classifier for Out-Of-Variable GeneralizationabstractKangjun Noh, Baekryun Seong, Hoyoon Byun, Youngjun Choi, Sungjin Song, Kyungwoo Song. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kangjun Noh, Baekryun Seong, Hoyoon Byun, Youngjun Choi 0001, Sungjin Song, Kyungwoo Song |
NAACL (Long Papers) | 6 |
| 2025 | CCL: Causal-aware In-context Learning for Out-of-Distribution GeneralizationabstractIn-context learning (ICL), a nonparametric learning method based on the knowledge of demonstration sets, has become a de facto standard for large language models (LLMs). The primary goal of ICL is to select valuable demonstration sets to enhance the performance of LLMs. Traditional ICL methods choose demonstration sets that share similar features with a given query. However, we have found that the performance of these traditional ICL approaches is limited on out-of-distribution (OOD) datasets, where the demonstration set and the query originate from different distributions. To ensure robust performance in OOD datasets, it is essential to learn causal representations that remain invariant between the source and target datasets. Inspired by causal representation learning, we propose causal-aware in-context learning (CCL). CCL captures the causal representations of a given dataset and selects demonstration sets that share similar causal features with the query. To achieve this, CCL employs a novel VAE-based causal representation learning technique. We demonstrate that CCL improves the OOD generalization performance of LLMs both theoretically and empirically. Code is available at: \url{https://github.com/MLAI-Yonsei/causal-context-learning} Hoyoon Byun, Gyeongdeok Seo, Joonseong Kang, Taero Kim, Kyungwoo Song |
NeurIPS | 6 |
| 2025 | Flat Posterior Does Matter For Bayesian Model AveragingabstractBayesian neural networks (BNNs) estimate the posterior distribution of model parameters and utilize posterior samples for Bayesian Model Averaging (BMA) in prediction. However, despite the crucial role of flatness in the loss landscape in improving the generalization of neural networks, its impact on BMA has been largely overlooked. In this work, we explore how posterior flatness influences BMA generalization and empirically demonstrate that \emph{(1) most approximate Bayesian inference methods fail to yield a flat posterior} and \emph{(2) BMA predictions, without considering posterior flatness, are less effective at improving generalization}. To address this, we propose Flat Posterior-aware Bayesian Model Averaging (FP-BMA), a novel training objective that explicitly encourages flat posteriors in a principled Bayesian manner. We also introduce a Flat Posterior-aware Bayesian Transfer Learning scheme that enhances generalization in downstream tasks. Empirically, we show that FP-BMA successfully captures flat posteriors, improving generalization performance. Sungjun Lim 0002, Jeyoon Yeom, Sooyon Kim, Hoyoon Byun, Jinho Kang, Yohan Jung, Jiyoung Jung, Kyungwoo Song |
UAI | 8 |
| 2025 | Enhancing Visual Classification Using Comparative DescriptorsabstractThe performance of vision-language models (VLMs), such as CLIP, in visual classification tasks, has been enhanced by leveraging semantic knowledge from large language models (LLMs), including GPT. Recent studies have shown that in zero-shot classification tasks, descriptors incorporating additional cues, high-level concepts, or even random characters often outperform those using only category names. In many classification tasks, while the top-1 accuracy may be relatively low, the top-5 accuracy is often significantly higher. This gap implies that most misclassifications occur among a few similar classes, highlighting the model's difficulty in distinguishing between classes with subtle differences. To address this challenge, we introduce a novel concept of comparative descriptors. These descriptors emphasize the unique features of a target class against its most similar classes, enhancing differentiation. By generating and integrating these comparative descriptors into the classification framework, we refine the semantic focus and improve classification accuracy. An additional filtering process ensures that these descriptors are closer to the image embeddings in the CLIP space, further enhancing performance. Our approach demonstrates improved accuracy and robustness in visual classification tasks by addressing the specific challenge of subtle inter-class differences. Code is available at ht tps: //github.com/hklee / Comparative-CLIP Hankyeol Lee, Gawon Seo, Wonseok Choi 0014, Geunyoung Jung, Kyungwoo Song, Jiyoung Jung |
WACV | 5 |
| 2025 | Correction to: Dirichlet stochastic weights averaging for graph neural networks
Minhoi Park, Rakwoo Chang, Kyungwoo Song |
Appl. Intell. | 3 |
| 2025 | Correction To: Dirichlet stochastic weights averaging for graph neural networks
Minhoi Park, Rakwoo Chang, Kyungwoo Song |
Appl. Intell. | 3 |
| 2025 | GFML: Gravity function for metric learning
Hoyoon Byun, Sungjun Lim 0002, Kyungwoo Song |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | COVID-19 prediction with doubly multi-task Gaussian Process
Sooyon Kim, Yongtaek Lim, Sungjun Lim 0002, Gyeongdeok Seo, Hojun Park, Jaehun Jung, Kyungwoo Song |
J. Biomed. Informatics | 8 |
| 2025 | Multi-query frequency prompting for physiological signal domain adaptation
Jinho Kang, Hoyoon Byun, Taero Kim, Jiyoung Jung, Kyungwoo Song |
Knowl. Based Syst. | 5 |
| 2025 | Fog-free training for foggy scene understanding
Kyungwoo Song, Junsuk Choe |
Pattern Recognit. Lett. | 2 |
| 2025 | TC-BERT: large-scale language model for Korean technology commercialization documents
Taero Kim, Changdae Oh, Hyeji Hwang, Eunkyeong Lee 0001, Yunjeong Choi, Hosik Choi, Kyungwoo Song |
J. Supercomput. | 9 |
| 2025 | Multi-purpose technology commercialization recommender system with large-scale Korean language model
Kangjun Noh, Hyeji Hwang, YongTaek Lim, Changdae Oh, Eunkyeong Lee 0001, Yunjeong Choi, Hosik Choi, Kyungwoo Song |
J. Supercomput. | 10 |
| 2024 | Pre-trained Vision and Language Transformers are Few-Shot Incremental LearnersabstractFew-Shot Class Incremental Learning (FSCIL) is a task that requires a model to learn new classes incrementally without forgetting when only a few samples for each class are given. FSCIL encounters two significant challenges: catastrophic forgetting and overfitting, and these challenges have driven prior studies to primarily rely on shallow models, such as ResNet-18. Even though their limited capacity can mitigate both forgetting and overfitting issues, it leads to inadequate knowledge transfer during few-shot in-cremental sessions. In this paper, we argue that large models such as vision and language transformers pre-trained on large datasets can be excellent few-shot incremental learn-ers. To this end, we propose a novel FSCIL framework called PriViLege, Pre-trained Vision and Language trans-formers with prompting functions and knowledge distillation. Our framework effectively addresses the challenges of catastrophic forgetting and overfitting in large models through new pre-trained knowledge tuning (PKT) and two losses: entropy-based divergence loss and semantic knowl-edge distillation loss. Experimental results show that the proposed PriViLege significantly outperforms the existing state-of-the-art methods with a large margin, e.g., +9.38% in CUB200, +20.58% in CIFAR-100, and +13.36% in miniImageNet. Our implementation code is available at https://github.com/KHU-AGI/PriViLege. Keon-Hee Park, Kyungwoo Song, Gyeong-Moon Park |
CVPR | 2 |
| 2024 | Perturb-and-Compare Approach for Detecting Out-of-Distribution Samples in Constrained Access EnvironmentsabstractAccessing machine learning models through remote APIs has been gaining prevalence following the recent trend of scaling up model parameters for increased performance. Even though these models exhibit remarkable ability, detecting out-of-distribution (OOD) samples remains a crucial safety concern for end users as these samples may induce unreliable outputs from the model. In this work, we propose an OOD detection framework, MixDiff, that is applicable even when the model’s parameters or its activations are not accessible to the end user. To bypass the access restriction, MixDiff applies an identical input-level perturbation to a given target sample and a similar in-distribution (ID) sample, then compares the relative difference in the model outputs of these two samples. MixDiff is model-agnostic and compatible with existing output-based OOD detection methods. We provide theoretical analysis to illustrate MixDiff’s effectiveness in discerning OOD samples that induce overconfident outputs from the model and empirically demonstrate that MixDiff consistently enhances the OOD detection performance on various datasets in vision and text domains. Hoyoon Byun, Changdae Oh, JinYeong Bak, Kyungwoo Song |
ECAI | 5 |
| 2024 | Online Continuous Generalized Category Discovery
Keon-Hee Park, Hakyung Lee, Kyungwoo Song, Gyeong-Moon Park |
ECCV (81) | 3 |
| 2024 | Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsabstractImproving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation. Changdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han, Sangdoo Yun, Jaegul Choo, Alex Hauptmann 0001, Zhi-Qi Cheng, Kyungwoo Song |
NeurIPS | 9 |
| 2024 | Dirichlet stochastic weights averaging for graph neural networks
Minhoi Park, Rakwoo Chang, Kyungwoo Song |
Appl. Intell. | 3 |
| 2024 | Language model-guided student performance prediction with multimodal auxiliary information
Changdae Oh, Minhoi Park, Sungjun Lim 0002, Kyungwoo Song |
Expert Syst. Appl. | 4 |
| 2024 | Bibimbap : Pre-trained models ensemble for Domain Generalization
Jinho Kang, Taero Kim, Changdae Oh, Jiyoung Jung, Rakwoo Chang, Kyungwoo Song |
Pattern Recognit. | 7 |
| 2024 | Structural and positional ensembled encoding for Graph Transformer
Jeyoon Yeom, Taero Kim, Rakwoo Chang, Kyungwoo Song |
Pattern Recognit. Lett. | 4 |
| 2023 | BlackVIP: Black-Box Visual Prompting for Robust Transfer LearningabstractWith the surge of large-scale pre-trained models (PTMs), fine-tuning these models to numerous downstream tasks becomes a crucial problem. Consequently, parameter efficient transfer learning (PETL) of large models has grasped huge attention. While recent PETL methods showcase impressive performance, they rely on optimistic assumptions: 1) the entire parameter set of a PTM is available, and 2) a sufficiently large memory capacity for the fine-tuning is equipped. However, in most real-world applications, PTMs are served as a black-box API or proprietary software without explicit parameter accessibility. Besides, it is hard to meet a large memory requirement for modern PTMs. In this work, we propose black-box visual prompting (Black-VIP), which efficiently adapts the PTMs without knowledge about model architectures and parameters. Black-VIP has two components; 1) Coordinator and 2) simultaneous perturbation stochastic approximation with gradient correction (SPSA-GC). The Coordinator designs inputdependent image-shaped visual prompts, which improves few-shot adaptation and robustness on distribution/location shift. SPSA-GC efficiently estimates the gradient of a target model to update Coordinator. Extensive experiments on 16 datasets demonstrate that BlackVIP enables robust adaptation to diverse domains without accessing PTMs' parameters, with minimal memory requirements. Code: https://github.com/changdaeoh/BlackVIP Changdae Oh, Hyeji Hwang, YongTaek Lim, Geunyoung Jung, Jiyoung Jung, Hosik Choi, Kyungwoo Song |
CVPR | 8 |
| 2023 | Causally Disentangled Generative Variational AutoEncoderabstractWe present a new supervised learning technique for the Variational AutoEncoder (VAE) that allows it to learn a causally disentangled representation and generate causally disentangled outcomes simultaneously. We call this approach Causally Disentangled Generation (CDG). CDG is a generative model that accurately decodes an output based on a causally disentangled representation. Our research demonstrates that adding supervised regularization to the encoder alone is insufficient for achieving a generative model with CDG, even for a simple task. Therefore, we explore the necessary and sufficient conditions for achieving CDG within a specific model. Additionally, we introduce a universal metric for evaluating the causal disentanglement of a generative model. Empirical results from both image and tabular datasets support our findings. Seunghwan An, Kyungwoo Song, Jong-June Jeon |
ECAI | 2 |
| 2023 | SAAL: Sharpness-Aware Active LearningabstractWhile deep neural networks play significant roles in many research areas, they are also prone to overfitting problems under limited data instances. To overcome overfitting, this paper introduces the first active learning method to incorporate the sharpness of loss space into the acquisition function. Specifically, our proposed method, Sharpness-Aware Active Learning (SAAL), constructs its acquisition function by selecting unlabeled instances whose perturbed loss becomes maximum. Unlike the Sharpness-Aware learning with fully-labeled datasets, we design a pseudo-labeling mechanism to anticipate the perturbed loss w.r.t. the ground-truth label, which we provide the theoretical bound for the optimization. We conduct experiments on various benchmark datasets for vision-based tasks in image classification, object detection, and domain adaptive semantic segmentation. The experimental results confirm that SAAL outperforms the baselines by selecting instances that have the potentially maximal perturbation on the loss. The code is available at https://github.com/YoonyeongKim/SAAL. Yoon-Yeong Kim, Youngjae Cho 0002, JoonHo Jang, Byeonghu Na, Yeongmin Kim, Kyungwoo Song, Wanmo Kang, Il-Chul Moon |
ICML | 6 |
| 2023 | Geodesic Multi-Modal Mixup for Robust Fine-TuningabstractPre-trained multi-modal models, such as CLIP, provide transferable embeddings and show promising results in diverse applications. However, the analysis of learned multi-modal embeddings is relatively unexplored, and the embedding transferability can be improved. In this work, we observe that CLIP holds separated embedding subspaces for two different modalities, and then we investigate it through the lens of \textit{uniformity-alignment} to measure the quality of learned representation. Both theoretically and empirically, we show that CLIP retains poor uniformity and alignment even after fine-tuning. Such a lack of alignment and uniformity might restrict the transferability and robustness of embeddings. To this end, we devise a new fine-tuning method for robust representation equipping better alignment and uniformity. First, we propose a \textit{Geodesic Multi-Modal Mixup} that mixes the embeddings of image and text to generate hard negative samples on the hypersphere. Then, we fine-tune the model on hard negatives as well as original negatives and positives with contrastive loss. Based on the theoretical analysis about hardness guarantee and limiting behavior, we justify the use of our method. Extensive experiments on retrieval, calibration, few- or zero-shot classification (under distribution shift), embedding arithmetic, and image captioning further show that our method provides transferable representations, enabling robust model adaptation on diverse tasks. Changdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim, Minchul Shin, Jong-June Jeon, Kyungwoo Song |
NeurIPS | 7 |
| 2023 | Sequential Likelihood-Free Inference with Neural Proposal
Kyungwoo Song, Yoon-Yeong Kim, Yongjin Shin, Wanmo Kang, Il-Chul Moon, Weonyoung Joo |
Pattern Recognit. Lett. | 2 |
| 2022 | From Noisy Prediction to True Label: Noisy Prediction Calibration via Generative ModelabstractNoisy labels are inevitable yet problematic in machine learning society. It ruins the generalization of a classifier by making the classifier over-fitted to noisy labels. Existing methods on noisy label have focused on modifying the classifier during the training procedure. It has two potential problems. First, these methods are not applicable to a pre-trained classifier without further access to training. Second, it is not easy to train a classifier and regularize all negative effects from noisy labels, simultaneously. We suggest a new branch of method, Noisy Prediction Calibration (NPC) in learning with noisy labels. Through the introduction and estimation of a new type of transition matrix via generative model, NPC corrects the noisy prediction from the pre-trained classifier to the true label as a post-processing scheme. We prove that NPC theoretically aligns with the transition matrix based methods. Yet, NPC empirically provides more accurate pathway to estimate true label, even without involvement in classifier learning. Also, NPC is applicable to any classifier trained with noisy label methods, if training instances and its predictions are available. Our method, NPC, boosts the classification performances of all baseline models on both synthetic and real-world datasets. The implemented code is available at https://github.com/BaeHeeSun/NPC. HeeSun Bae, Byeonghu Na, JoonHo Jang, Kyungwoo Song, Il-Chul Moon |
ICML | 5 |
| 2022 | Efficient Approximate Inference for Stationary Kernel on Frequency DomainabstractBased on the Fourier duality between a stationary kernel and its spectral density, modeling the spectral density using a Gaussian mixture density enables one to construct a flexible kernel, known as a Spectral Mixture kernel, that can model any stationary kernel. However, despite its expressive power, training this kernel is typically difficult because scalability and overfitting issues often arise due to a large number of training parameters. To resolve these issues, we propose an approximate inference method for estimating the Spectral mixture kernel hyperparameters. Specifically, we approximate this kernel by using the finite random spectral points based on Random Fourier Feature and optimize the parameters for the distribution of spectral points by sampling-based variational inference. To improve this inference procedure, we analyze the training loss and propose two special methods: a sampling method of spectral points to reduce the error of the approximate kernel in training, and an approximate natural gradient to accelerate the convergence of parameter inference. Yohan Jung, Kyungwoo Song, Jinkyoo Park |
ICML | 2 |
| 2022 | Soft Truncation: A Universal Training Technique of Score-based Diffusion Model for High Precision Score EstimationabstractRecent advances in diffusion models bring state-of-the-art performance on image generation tasks. However, empirical results from previous research in diffusion models imply an inverse correlation between density estimation and sample generation performances. This paper investigates with sufficient empirical evidence that such inverse correlation happens because density estimation is significantly contributed by small diffusion time, whereas sample generation mainly depends on large diffusion time. However, training a score network well across the entire diffusion time is demanding because the loss scale is significantly imbalanced at each diffusion time. For successful training, therefore, we introduce Soft Truncation, a universally applicable training technique for diffusion models, that softens the fixed and static truncation hyperparameter into a random variable. In experiments, Soft Truncation achieves state-of-the-art performance on CIFAR-10, CelebA, CelebA-HQ $256\times 256$, and STL-10 datasets. Kyungwoo Song, Wanmo Kang, Il-Chul Moon |
ICML | 3 |
| 2022 | Learning Fair Representation via Distributional Contrastive DisentanglementabstractLearning fair representation is crucial for achieving fairness or debiasing sensitive information. Most existing works rely on adversarial representation learning to inject some invariance into representation. However, adversarial learning methods are known to suffer from relatively unstable training, and this might harm the balance between fairness and predictiveness of representation. We propose a new approach, learningFAir Representation via distributional CONtrastive Variational AutoEncoder (FarconVAE), which induces the latent space to be disentangled into sensitive and non-sensitive parts. We first construct the pair of observations with different sensitive attributes but with the same labels. Then, FarconVAE enforces each non-sensitive latent to be closer, while sensitive latents to be far from each other and also far from the non-sensitive latent by contrasting their distributions. We provide a new type of contrastive loss motivated by Gaussian and Student-t kernels for distributional contrastive learning with theoretical analysis. Besides, we adopt a new swap-reconstruction loss to boost the disentanglement further. FarconVAE shows superior performance on fairness, pretrained model debiasing, and domain generalization tasks from various modalities, including tabular, image, and text. Changdae Oh, Heeji Won, Junhyuk So, Taero Kim, Hosik Choi, Kyungwoo Song |
KDD | 7 |
| 2022 | Unknown-Aware Domain Adversarial Learning for Open-Set Domain AdaptationabstractOpen-Set Domain Adaptation (OSDA) assumes that a target domain contains unknown classes, which are not discovered in a source domain. Existing domain adversarial learning methods are not suitable for OSDA because distribution matching with $\textit{unknown}$ classes leads to negative transfer. Previous OSDA methods have focused on matching the source and the target distribution by only utilizing $\textit{known}$ classes. However, this $\textit{known}$-only matching may fail to learn the target-$\textit{unknown}$ feature space. Therefore, we propose Unknown-Aware Domain Adversarial Learning (UADAL), which $\textit{aligns}$ the source and the target-$\textit{known}$ distribution while simultaneously $\textit{segregating}$ the target-$\textit{unknown}$ distribution in the feature alignment procedure. We provide theoretical analyses on the optimized state of the proposed $\textit{unknown-aware}$ feature alignment, so we can guarantee both $\textit{alignment}$ and $\textit{segregation}$ theoretically. Empirically, we evaluate UADAL on the benchmark datasets, which shows that UADAL outperforms other methods with better feature alignments by reporting state-of-the-art performances. JoonHo Jang, Byeonghu Na, Donghyeok Shin, Mingi Ji, Kyungwoo Song, Il-Chul Moon |
NeurIPS | 5 |
| 2021 | Counterfactual Fairness with Disentangled Causal Effect Variational AutoencoderabstractThe problem of fair classification can be mollified if we develop a method to remove the embedded sensitive information from the classification features. This line of separating the sensitive information is developed through the causal inference, and the causal inference enables the counterfactual generations to contrast the what-if case of the opposite sensitive attribute. Along with this separation with the causality, a frequent assumption in the deep latent causal model defines a single latent variable to absorb the entire exogenous uncertainty of the causal graph. However, we claim that such structure cannot distinguish the 1) information caused by the intervention (i.e., sensitive variable) and 2) information correlated with the intervention from the data. Therefore, this paper proposes Disentangled Causal Effect Variational Autoencoder (DCEVAE) to resolve this limitation by disentangling the exogenous uncertainty into two latent variables: either 1) independent to interventions or 2) correlated to interventions without causality. Particularly, our disentangling approach preserves the latent variable correlated to interventions in generating counterfactual examples. We show that our method estimates the total effect and the counterfactual effect without a complete causal graph. By adding a fairness regularization, DCEVAE generates a counterfactual fair dataset while losing less original information. Also, DCEVAE generates natural counterfactual images by only flipping sensitive information. Additionally, we theoretically show the differences in the covariance structures of DCEVAE and prior works from the perspective of the latent disentanglement. Hyemi Kim, JoonHo Jang, Kyungwoo Song, Weonyoung Joo, Wanmo Kang, Il-Chul Moon |
AAAI | 4 |
| 2021 | Implicit Kernel AttentionabstractAttention computes the dependency between representations, and it encourages the model to focus on the important selective features. Attention-based models, such as Transformer and graph attention network (GAT), are widely utilized for sequential data and graph-structured data. This paper suggests a new interpretation and generalized structure of the attention in Transformer and GAT. For the attention in Transformer and GAT, we derive that the attention is a product of two parts: 1) the RBF kernel to measure the similarity of two instances and 2) the exponential of L2 norm to compute the importance of individual instances. From this decomposition, we generalize the attention in three ways. First, we propose implicit kernel attention with an implicit kernel function instead of manual kernel selection. Second, we generalize L2 norm as the Lp norm. Third, we extend our attention to structured multi-head attention. Our generalized attention shows better performance on classification, translation, and regression tasks. Kyungwoo Song, Yohan Jung, Il-Chul Moon |
AAAI | 1 |
| 2021 | LADA: Look-Ahead Data Acquisition via Augmentation for Deep Active LearningabstractActive learning effectively collects data instances for training deep learning models when the labeled dataset is limited and the annotation cost is high. Data augmentation is another effective technique to enlarge the limited amount of labeled instances. The scarcity of labeled dataset leads us to consider the integration of data augmentation and active learning. One possible approach is a pipelined combination, which selects informative instances via the acquisition function and generates virtual instances from the selected instances via augmentation. However, this pipelined approach would not guarantee the informativeness of the virtual instances. This paper proposes Look-Ahead Data Acquisition via augmentation, or LADA framework, that looks ahead the effect of data augmentation in the process of acquisition. LADA jointly considers both 1) unlabeled data instance to be selected and 2) virtual data instance to be generated by data augmentation, to construct the acquisition function. Moreover, to generate maximally informative virtual instances, LADA optimizes the data augmentation policy to maximize the predictive acquisition score, resulting in the proposal of InfoSTN and InfoMixup. The experimental results of LADA show a significant improvement over the recent augmentation and acquisition baselines that were independently applied. Yoon-Yeong Kim, Kyungwoo Song, JoonHo Jang, Il-Chul Moon |
NeurIPS | 2 |
| 2020 | Sequential Recommendation with Relation-Aware Kernelized Self-AttentionabstractRecent studies identified that sequential Recommendation is improved by the attention mechanism. By following this development, we propose Relation-Aware Kernelized Self-Attention (RKSA) adopting a self-attention mechanism of the Transformer with augmentation of a probabilistic model. The original self-attention of Transformer is a deterministic measure without relation-awareness. Therefore, we introduce a latent space to the self-attention, and the latent space models the recommendation context from relation as a multivariate skew-normal distribution with a kernelized covariance matrix from co-occurrences, item characteristics, and user information. This work merges the self-attention of the Transformer and the sequential recommendation by adding a probabilistic model of the recommendation task specifics. We experimented RKSA over the benchmark datasets, and RKSA shows significant improvements compared to the recent baseline models. Also, RKSA were able to produce a latent space model that answers the reasons for recommendation. Mingi Ji, Weonyoung Joo, Kyungwoo Song, Yoon-Yeong Kim, Il-Chul Moon |
AAAI | 3 |
| 2020 | Hierarchically Clustered Representation LearningabstractThe joint optimization of representation learning and clustering in the embedding space has experienced a breakthrough in recent years. In spite of the advance, clustering with representation learning has been limited to flat-level categories, which often involves cohesive clustering with a focus on instance relations. To overcome the limitations of flat clustering, we introduce hierarchically-clustered representation learning (HCRL), which simultaneously optimizes representation learning and hierarchical clustering in the embedding space. Compared with a few prior works, HCRL firstly attempts to consider a generation of deep embeddings from every component of the hierarchy, not just leaf components. In addition to obtaining hierarchically clustered embeddings, we can reconstruct data by the various abstraction levels, infer the intrinsic hierarchical structure, and learn the level-proportion features. We conducted evaluations with image and text domains, and our quantitative analyses showed competent likelihoods and the best accuracies compared with the baselines. Su-Jin Shin, Kyungwoo Song, Il-Chul Moon |
AAAI | 2 |
| 2020 | Bivariate Beta-LSTM
Kyungwoo Song, JoonHo Jang, Il-Chul Moon |
AAAI | 1 |
| 2020 | Deep Generative Positive-Unlabeled Learning under Selection BiasabstractLearning in the positive-unlabeled (PU) setting is prevalent in real world applications. Many previous works depend upon theSelected Completely At Random (SCAR) assumption to utilize unlabeled data, but the SCAR assumption is not often applicable to the real world due to selection bias in label observations. This paper is the first generative PU learning model without the SCAR assumption. Specifically, we derive the PU risk function without the SCAR assumption, and we generate a set of virtual PU examples to train the classifier. Although our PU risk function is more generalizable, the function requires PU instances that do not exist in the observations. Therefore, we introduce the VAE-PU, which is a variant of variational autoencoders to separate two latent variables that generate either features or observation indicators. The separated latent information enables the model to generate virtual PU instances. We test the VAE-PU on benchmark datasets with and without the SCAR assumption. The results indicate that the VAE-PU is superior when selection bias exists, and the VAE-PU is also competent under the SCAR assumption. The results also emphasize that the VAE-PU is effective when there are few positive-labeled instances due to modeling on selection bias. Byeonghu Na, Hyemi Kim, Kyungwoo Song, Weonyoung Joo, Yoon-Yeong Kim, Il-Chul Moon |
CIKM | 3 |
| 2020 | Context Aware Sequence ModelingabstractContext modeling helps understand the data, such as sentence or user behavior. Contextual information captures the important underlying feature, and it enhances the relationship between data instances or hidden representations. As the importance of the sequential model grows, so does the importance of the sequential contextual modeling. Under the sequential data, we need to consider the context change over time. In this paper, we present our research works on context modeling and its dynamics modeling over time. Furthermore, we extend our research to handle the multi-granularity of sequential context modeling to consider rich context representations. Kyungwoo Song |
IJCAI | 1 |
| 2019 | Adversarial Dropout for Recurrent Neural Networks
Sungrae Park, Kyungwoo Song, Mingi Ji, Wonsung Lee, Il-Chul Moon |
AAAI | 2 |
| 2019 | Hierarchical Context Enabled Recurrent Neural Network for RecommendationabstractA long user history inevitably reflects the transitions of personal interests over time. The analyses on the user history require the robust sequential model to anticipate the transitions and the decays of user interests. The user history is often modeled by various RNN structures, but the RNN structures in the recommendation system still suffer from the long-term dependency and the interest drifts. To resolve these challenges, we suggest HCRNN with three hierarchical contexts of the global, the local, and the temporary interests. This structure is designed to withhold the global long-term interest of users, to reflect the local sub-sequence interests, and to attend the temporary interests of each transition. Besides, we propose a hierarchical context-based gate structure to incorporate our interest drift assumption. As we suggest a new RNN structure, we support HCRNN with a complementary bi-channel attention structure to utilize hierarchical context. We experimented the suggested structure on the sequential recommendation tasks with CiteULike, MovieLens, and LastFM, and our model showed the best performances in the sequential recommendations. Kyungwoo Song, Mingi Ji, Sungrae Park, Il-Chul Moon |
AAAI | 1 |
| 2018 | Neural Ideal Point Estimation Network
Kyungwoo Song, Wonsung Lee, Il-Chul Moon |
AAAI | 1 |
| 2017 | Augmented Variational Autoencoders for Collaborative Filtering with Auxiliary InformationabstractRecommender systems offer critical services in the age of mass information. A good recommender system selects a certain item for a specific user by recognizing why the user might like the item. This awareness implies that the system should model the background of the items and the users. This background modeling for recommendation is tackled through the various models of collaborative filtering with auxiliary information. This paper presents variational approaches for collaborative filtering to deal with auxiliary information. The proposed methods encompass variational autoencoders through augmenting structures to model the auxiliary information and to model the implicit user feedback. This augmentation includes the ladder network and the generative adversarial network to extract the low-dimensional representations influenced by the auxiliary information. These two augmentations are the first trial in the venue of the variational autoencoders, and we demonstrate their significant improvement on the performances in the applications of the collaborative filtering. Wonsung Lee, Kyungwoo Song, Il-Chul Moon |
CIKM | 2 |
| 2016 | Data-driven ballistic coefficient learning for future state prediction of high-speed vehicles
Kyungwoo Song, Jinhyung Tak, Han-Lim Choi, Il-Chul Moon |
FUSION | 1 |
| 2014 | Identifying the evolution of disasters and responses with network-text analysisabstractDisasters and responses have evolved over-time, and the evolution has been affected by various factors, such as societal change, climate change, and technological advance. To better prepare the future disasters, we need to estimate the evolution trend of the past disasters and the responses. This paper analyzes the academic articles of the field with network-text analyses. The analyses captured the word level and the topic level evolution over-time with statistical significance tests. Further, we turn the text mining results into the network analysis data to identify the key words and topics in the evolution paths. The proposed method suggests the swift of interests, i.e. the new ways of organizational interoperation, the evolution of logistic issues, in the disaster and response field. Kyungwoo Song, Do-Hyeong Kim, Su-Jin Shin, Il-Chul Moon |
SMC | 1 |