Kyungwoo Song

dblp:155/4867 · DBLP profile ↗
← Back
8ranked-venue papers in the field
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 3Information Retrieval & Web Search · 3Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1 (1 first)
YearPublicationVenuePosition
2026 Enhancing LLMs for Manufacturing Information Extraction
Subeen Park, Hakyung Lee, Ryunyi Lee, Hyo-won Suh, Kyungwoo Song
PAKDD (4)5
2025 RAILL: Retrieval-Augmented Instruction Tuning for Low-Resource Language Model Training
Youngjun Choi 0001, Sungjun Lim 0002, Minhoi Park, Jaekyeong Jung, Eunsik Kim, Chul-su Kim, Kyongjae Lee, Hosik Choi, Kyungwoo Song
IEEE Big Data10
2025 Causal Effect Variational Transformer for Public Health Measures and COVID-19 Infection Cluster Analysis
abstract
Recent research increasingly integrates causal inference into deep learning models to enhance the explainability and robustness of medical applications. However, data scarcity remains a fundamental challenge due to privacy constraints and the high cost of data collection. This issue, compounded by complex variable dependencies and unobserved latent confounders, hinders the reliable estimation of causal effects. To address these challenges, we collect two real-world COVID-19 infection cluster datasets, including public health measures, from distinct distributions in collaboration with local governments, a medical university, and a hospital. We also propose a cut-off augmentation method that generates diverse feature-label pairs by slicing time-series sequences at different observation windows, effectively simulating partial observations common in real-world settings. We further introduce the Causal Effect Variational Transformer (CEVT), a Transformer-based model that captures temporal structure and addresses the difficulty of causal estimation under scarce data, complex dependencies, and latent confounding by modeling multiple treatments through an iterative conditioning mechanism. We validate the causal modeling capability of CEVT on synthetic datasets and demonstrate that, on two distinct COVID-19 datasets, it consistently outperforms baselines in infection prediction. Notably, the causal effects estimated by CEVT converge with findings from medical studies on infection control, reinforcing its reliability and underscoring its potential to inform public health decision-making.
Jinho Kang, Sungjun Lim 0002, Hojun Park, Jiyoung Jung, Jaehun Jung, Kyungwoo Song
CIKM6
2025 Exploring the Potential of Foundation Models as Reliable AI Contact Centers
Hoyoon Byun, Minhoi Park, Seolah Kim, EunBi Kim, Kyungwoo Song
KDD (2)5
2022 Learning Fair Representation via Distributional Contrastive Disentanglement
abstract
Learning fair representation is crucial for achieving fairness or debiasing sensitive information. Most existing works rely on adversarial representation learning to inject some invariance into representation. However, adversarial learning methods are known to suffer from relatively unstable training, and this might harm the balance between fairness and predictiveness of representation. We propose a new approach, learningFAir Representation via distributional CONtrastive Variational AutoEncoder (FarconVAE), which induces the latent space to be disentangled into sensitive and non-sensitive parts. We first construct the pair of observations with different sensitive attributes but with the same labels. Then, FarconVAE enforces each non-sensitive latent to be closer, while sensitive latents to be far from each other and also far from the non-sensitive latent by contrasting their distributions. We provide a new type of contrastive loss motivated by Gaussian and Student-t kernels for distributional contrastive learning with theoretical analysis. Besides, we adopt a new swap-reconstruction loss to boost the disentanglement further. FarconVAE shows superior performance on fairness, pretrained model debiasing, and domain generalization tasks from various modalities, including tabular, image, and text.
Changdae Oh, Heeji Won, Junhyuk So, Taero Kim, Hosik Choi, Kyungwoo Song
KDD7
2020 Deep Generative Positive-Unlabeled Learning under Selection Bias
abstract
Learning in the positive-unlabeled (PU) setting is prevalent in real world applications. Many previous works depend upon theSelected Completely At Random (SCAR) assumption to utilize unlabeled data, but the SCAR assumption is not often applicable to the real world due to selection bias in label observations. This paper is the first generative PU learning model without the SCAR assumption. Specifically, we derive the PU risk function without the SCAR assumption, and we generate a set of virtual PU examples to train the classifier. Although our PU risk function is more generalizable, the function requires PU instances that do not exist in the observations. Therefore, we introduce the VAE-PU, which is a variant of variational autoencoders to separate two latent variables that generate either features or observation indicators. The separated latent information enables the model to generate virtual PU instances. We test the VAE-PU on benchmark datasets with and without the SCAR assumption. The results indicate that the VAE-PU is superior when selection bias exists, and the VAE-PU is also competent under the SCAR assumption. The results also emphasize that the VAE-PU is effective when there are few positive-labeled instances due to modeling on selection bias.
Byeonghu Na, Hyemi Kim, Kyungwoo Song, Weonyoung Joo, Yoon-Yeong Kim, Il-Chul Moon
CIKM3
2017 Augmented Variational Autoencoders for Collaborative Filtering with Auxiliary Information
abstract
Recommender systems offer critical services in the age of mass information. A good recommender system selects a certain item for a specific user by recognizing why the user might like the item. This awareness implies that the system should model the background of the items and the users. This background modeling for recommendation is tackled through the various models of collaborative filtering with auxiliary information. This paper presents variational approaches for collaborative filtering to deal with auxiliary information. The proposed methods encompass variational autoencoders through augmenting structures to model the auxiliary information and to model the implicit user feedback. This augmentation includes the ladder network and the generative adversarial network to extract the low-dimensional representations influenced by the auxiliary information. These two augmentations are the first trial in the venue of the variational autoencoders, and we demonstrate their significant improvement on the performances in the applications of the collaborative filtering.
Wonsung Lee, Kyungwoo Song, Il-Chul Moon
CIKM2
2016 Data-driven ballistic coefficient learning for future state prediction of high-speed vehicles
Kyungwoo Song, Jinhyung Tak, Han-Lim Choi, Il-Chul Moon
FUSION1