VLDB 2026 Research / reviewers in the wild / expert
Seunghwan An
dblp:293/9384
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-1891-1174ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Impute Missing Entries with UncertaintyabstractMissing data presents a widespread challenge in real-world data collection. In this paper, our goal is to impute missing entries while accurately reflecting the uncertainty associated with them. We introduce U-VAE, a method that employs a non-parametric distributional learning strategy to parameterize the likelihood of missing values. To address the infeasibility of directly estimating the underlying conditional distributions due to data incompleteness, we incorporate stochastic re-masking and un-masking techniques during training. Specifically, we replace the conventional reconstruction loss with the continuous ranked probability score (CRPS), a strictly proper scoring rule, and theoretically demonstrate that the discrepancy between the underlying conditional distribution and our imputer is upper-bounded. We evaluate the performance of U-VAE on 11 real-world datasets, showing its effectiveness in both single and multiple imputations, while also enhancing post-imputation performance and supporting valid statistical inference. Jaesung Lim 0002, Seunghwan An, Jong-June Jeon |
AAAI | 2 |
| 2026 | Estimating Subgraph Importance with Structural Prior Domain Knowledge
Seunghwan An, Jong-June Jeon |
PAKDD (1) | 2 |
| 2026 | DrIM: Context-Driven Nearest Neighbor Imputation Using Language Representation
Jaesung Lim 0002, Seunghwan An, Jong-June Jeon |
PAKDD (2) | 2 |
| 2025 | Masked Language Modeling Becomes Conditional Density Estimation for Tabular Data SynthesisabstractIn this paper, our goal is to generate synthetic data for heterogeneous (mixed-type) tabular datasets with high machine learning utility (MLu). Since the MLu performance depends on accurately approximating the conditional distributions, we focus on devising a synthetic data generation method based on conditional distribution estimation. We introduce MaCoDE by redefining the consecutive multi-class classification task of Masked Language Modeling (MLM) as histogram-based non-parametric conditional density estimation. Our approach enables the estimation of conditional densities across arbitrary combinations of target and conditional variables. We bridge the theoretical gap between distributional learning and MLM by demonstrating that minimizing the orderless multi-class classification loss leads to minimizing the total variation distance between conditional distributions. To validate our proposed model, we evaluate its performance in synthetic data generation across 10 real-world datasets, demonstrating its ability to adjust data privacy levels easily without re-training. Additionally, since masked input tokens in MLM are analogous to missing data, we further assess its effectiveness in handling training datasets with missing values, including multiple imputations of the missing entries. Seunghwan An, Gyeongdong Woo, Jaesung Lim 0002, Sungchul Hong, Jong-June Jeon |
AAAI | 1 |
| 2025 | Improving SMOTE via fusing conditional VAE for data-adaptive noise filteringabstractRecent advances in a generative neural network model extend the development of data augmentation methods. However, the augmentation methods based on the modern generative models fail to achieve notable improvement in class imbalance data compared to the conventional model, Synthetic Minority Oversampling Technique (SMOTE). We investigate the problem of the generative model for imbalanced classification and introduce a framework to enhance the SMOTE algorithm using Variational Autoencoders (VAE s). Our approach systematically quantifies the density of data points in a low-dimensional latent space using the VAE, simultaneously incorporating information on class labels and classification difficulty. Then, the data points potentially degrading the augmentation are systematically excluded, and the neighboring observations are directly augmented on the data space. Empirical studies on several imbalanced datasets represent that this simple process innovatively improves the conventional SMOTE algorithm over the deep learning models. Consequently, we conclude that the selection of minority data and the interpolation in the data space are beneficial for imbalanced classification problems with a relatively small number of data points. Sungchul Hong, Seunghwan An, Jong-June Jeon |
Appl. Intell. | 2 |
| 2025 | Variational autoencoder for distributional learning via quantile function estimationabstractThe Gaussianity assumption in Variational AutoEncoders (VAEs) enhances computational efficiency and provides a solid theoretical basis for estimating probability distributions. However, we have empirically found that approximating distributions with non-smooth densities using the Gaussian VAE is challenging. Therefore, we propose an approach for distributional learning in VAEs that extends to estimating the quantile function while accommodating both smooth and non-smooth densities. This is achieved by utilizing the continuous ranked probability score, a strictly proper scoring rule, as our reconstruction loss. Our method can be seen as a specialized form of a nonparametric M-estimator for estimating general quantile functions, and we establish a theoretical connection between our model and quantile estimation. Furthermore, we demonstrate that our reconstruction loss functions as the lower bound of an infinite mixture of asymmetric Laplace distributions, which allows our synthetic data generation mechanism to maintain differential privacy. We validate the effectiveness of our model in capturing the underlying distribution through experiments involving synthetic data generation on real-world tabular datasets, showing that the level of data privacy can be easily adjusted. Seunghwan An, Sungchul Hong, Jong-June Jeon |
Neural Networks | 1 |
| 2024 | Cryptocurrency Price Forecasting using Variational Autoencoder with Versatile Quantile ModelingabstractIn recent years, there has been a growing interest in probabilistic forecasting methods that offer more comprehensive insights by considering prediction uncertainties rather than point estimates. This paper introduces a novel variational autoencoder learning framework for multivariate distributional forecasting. Our approach employs distributional learning to directly estimate the cumulative distribution function of future time series conditional distributions using the continuous ranked probability score. By incorporating a temporal structure within the latent space and utilizing versatile quantile models, such as the generalized lambda distribution, we enable distributional forecasting by generating synthetic time series data for future time points. To assess the effectiveness of our method, we conduct experiments using a multivariate dataset of real cryptocurrency prices, demonstrating its superiority in forecasting high-volatility scenarios. Sungchul Hong, Seunghwan An, Jong-June Jeon |
CIKM | 2 |
| 2024 | Customization of latent space in semi-supervised Variational AutoEncoderabstractWe propose a novel semi-supervised learning method of Variational AutoEncoder (VAE), which yields a customized latent space through our EXplainable encoder Network (EXoN). The customization involves a manual design of the interpolation and structural constraint, such as proximity, which enhances the interpretability of the latent space. To improve the classification performance, we introduce a new semi-supervised classification method called SCI (Soft-label Consistency Interpolation). Combining the classification loss and the Kullback–Leibler divergence is crucial in constructing an explainable latent space. Additionally, the variability of the generated samples is determined by an active latent subspace, which effectively captures distinctive characteristics. We conduct experiments using the MNIST, SVHN, and CIFAR-10 datasets, and the results demonstrate that our approach yields an explainable latent space while significantly reducing the effort required to analyze representation patterns within the latent space. Seunghwan An, Jong-June Jeon |
Pattern Recognit. Lett. | 1 |
| 2023 | Causally Disentangled Generative Variational AutoEncoderabstractWe present a new supervised learning technique for the Variational AutoEncoder (VAE) that allows it to learn a causally disentangled representation and generate causally disentangled outcomes simultaneously. We call this approach Causally Disentangled Generation (CDG). CDG is a generative model that accurately decodes an output based on a causally disentangled representation. Our research demonstrates that adding supervised regularization to the encoder alone is insufficient for achieving a generative model with CDG, even for a simple task. Therefore, we explore the necessary and sufficient conditions for achieving CDG within a specific model. Additionally, we introduce a universal metric for evaluating the causal disentanglement of a generative model. Empirical results from both image and tabular datasets support our findings. Seunghwan An, Kyungwoo Song, Jong-June Jeon |
ECAI | 1 |
| 2023 | Distributional Learning of Variational AutoEncoder: Application to Synthetic Data GenerationabstractThe Gaussianity assumption has been consistently criticized as a main limitation of the Variational Autoencoder (VAE) despite its efficiency in computational modeling. In this paper, we propose a new approach that expands the model capacity (i.e., expressive power of distributional family) without sacrificing the computational advantages of the VAE framework. Our VAE model's decoder is composed of an infinite mixture of asymmetric Laplace distribution, which possesses general distribution fitting capabilities for continuous variables. Our model is represented by a special form of a nonparametric M-estimator for estimating general quantile functions, and we theoretically establish the relevance between the proposed model and quantile estimation. We apply the proposed model to synthetic data generation, and particularly, our model demonstrates superiority in easily adjusting the level of data privacy. Seunghwan An, Jong-June Jeon |
NeurIPS | 1 |