EDBT 2026 Demo / reviewers in the wild / expert
Jong-June Jeon
dblp:203/0387
· DBLP profile ↗
19ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0002-1423-4292ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 16 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Impute Missing Entries with UncertaintyabstractMissing data presents a widespread challenge in real-world data collection. In this paper, our goal is to impute missing entries while accurately reflecting the uncertainty associated with them. We introduce U-VAE, a method that employs a non-parametric distributional learning strategy to parameterize the likelihood of missing values. To address the infeasibility of directly estimating the underlying conditional distributions due to data incompleteness, we incorporate stochastic re-masking and un-masking techniques during training. Specifically, we replace the conventional reconstruction loss with the continuous ranked probability score (CRPS), a strictly proper scoring rule, and theoretically demonstrate that the discrepancy between the underlying conditional distribution and our imputer is upper-bounded. We evaluate the performance of U-VAE on 11 real-world datasets, showing its effectiveness in both single and multiple imputations, while also enhancing post-imputation performance and supporting valid statistical inference. Jaesung Lim 0002, Seunghwan An, Jong-June Jeon |
AAAI | 3 |
| 2026 | Estimating Subgraph Importance with Structural Prior Domain Knowledge
Seunghwan An, Jong-June Jeon |
PAKDD (1) | 3 |
| 2026 | DrIM: Context-Driven Nearest Neighbor Imputation Using Language Representation
Jaesung Lim 0002, Seunghwan An, Jong-June Jeon |
PAKDD (2) | 3 |
| 2026 | Revisiting IPS in Recommendation Models: Unveiling Its Impact on Model PerformanceabstractIn neural network-based recommendation models, numerous studies have adopted inverse propensity scoring (IPS) as a principal approach, namely, minimizing unbiased risk. IPS is commonly employed to address selection bias inherent in user-item interactions and to improve model generalization performance. In this study, we critically examine whether minimizing the unbiased risk obtained via IPS genuinely contributes to improvements in recommendation performance. To this end, we first analyze the mechanism through which IPS influences model behavior, particularly focusing on its role beyond bias correction. Under the data generation assumptions commonly adopted in prior work, we demonstrate that the effectiveness of IPS is more closely associated with enriching embedding space by multi-task learning, rather than the unbiasedness of the risk function. This observation suggests that IPS may serve as a regularization mechanism that indirectly enhances model expressiveness. Furthermore, our experimental results reveal that the reweighting process of IPS can, in certain scenarios, lead to a degradation in recommendation performance. These findings highlight the limitations of IPS when used solely for unbiased risk minimization. In light of this, we show that multi-task learning proves beneficial in effectively managing this complexity, thereby offering a more robust approach to improving recommendation quality. Wonhyung Shin, Jong-June Jeon |
WSDM | 2 |
| 2025 | Masked Language Modeling Becomes Conditional Density Estimation for Tabular Data SynthesisabstractIn this paper, our goal is to generate synthetic data for heterogeneous (mixed-type) tabular datasets with high machine learning utility (MLu). Since the MLu performance depends on accurately approximating the conditional distributions, we focus on devising a synthetic data generation method based on conditional distribution estimation. We introduce MaCoDE by redefining the consecutive multi-class classification task of Masked Language Modeling (MLM) as histogram-based non-parametric conditional density estimation. Our approach enables the estimation of conditional densities across arbitrary combinations of target and conditional variables. We bridge the theoretical gap between distributional learning and MLM by demonstrating that minimizing the orderless multi-class classification loss leads to minimizing the total variation distance between conditional distributions. To validate our proposed model, we evaluate its performance in synthetic data generation across 10 real-world datasets, demonstrating its ability to adjust data privacy levels easily without re-training. Additionally, since masked input tokens in MLM are analogous to missing data, we further assess its effectiveness in handling training datasets with missing values, including multiple imputations of the missing entries. Seunghwan An, Gyeongdong Woo, Jaesung Lim 0002, Sungchul Hong, Jong-June Jeon |
AAAI | 6 |
| 2025 | Generalizing Query Performance Prediction under Retriever and Concept Shifts via Data-driven CorrectionabstractQuery Performance Prediction (QPP) aims to estimate the effectiveness of an information retrieval (IR) system without access to ground-truth relevance judgments. Existing supervised QPP methods typically follow a regression model framework that maps query-document representations to target metrics such as RR@10 or nDCG@10. However, these approaches often suffer from degraded performance under concept shift, where the distribution of relevance given a query-document pair changes between training and test datasets. This paper proposes a novel classification-based framework, QPP-MLC (QPP Multi-Label Classification), which formulates QPP as a multi-label classification task. QPP-MLC infers the relevance of each document among the top-k retrieved results and aggregates these document-level relevance predictions to predict the overall query performance. As a result, QPP-MLC provides a diagnosis tool for the concept shift and a correction method under the concept shift by modulating a threshold level of classification tasks. Experiments on MS MARCO and TREC DL benchmarks show that QPP-MLC achieves strong prediction accuracy and outperforms traditional regression-based QPP methods. Jaehwan Jung 0001, Jong-June Jeon |
CIKM | 2 |
| 2025 | CHEM: Causally and Hierarchically Explaining MoleculesabstractGraph Neural Networks (GNNs) have significantly advanced in analyzing graph-structured data; however, their explainability remains challenging, affecting their applicability in critical domains such as medicine and pharmacology. In particular, violating the subgraph structure can degrade model interpretability and generalization performance. To address this problem, we propose a hierarchical and explainable causal inference-based GNN. Our model selects features based on explainable subgraph units informed by prior knowledge. Our method begins by clustering molecules into functional groups via the BRICS algorithm, then constructing a hierarchical structure at both the node and motif levels. The proposed model employs a gate module that distills causal features on the motif level and a loss function that disconnects information flow from non-causal features to the target level. The classification results on real-world molecular graphs demonstrate that our model outperforms other causal inference-based GNN models. In addition, it is confirmed that leveraging molecular docking data effectively identifies true causal substructures in the proposed model. Gyeongdong Woo, Soyoung Cho, Kimoon Na, Jinhee Choi, Jong-June Jeon |
CIKM | 7 |
| 2025 | Dynamic Higher-Order Relations and Event-Driven Temporal Modeling for Stock Price ForecastingabstractIn stock price forecasting, modeling the probabilistic dependence between stock prices within a time-series framework has remained a persistent and highly challenging area of research. We propose a novel model to explain the extreme co-movement in multivariate data with time-series dependencies. Our model incorporates a Hawkes process layer to capture abrupt co-movements, thereby enhancing the temporal representation of market dynamics. We introduce dynamic hypergraphs into our model adapting to higher-order (groupwise rather than pairwise) relationships within the stock market. Extensive experiments on real-world benchmarks demonstrate the robustness of our approach in predictive performance and portfolio stability. Kijeong Park, Sungchul Hong, Jong-June Jeon |
IJCAI | 3 |
| 2025 | Improving SMOTE via fusing conditional VAE for data-adaptive noise filteringabstractRecent advances in a generative neural network model extend the development of data augmentation methods. However, the augmentation methods based on the modern generative models fail to achieve notable improvement in class imbalance data compared to the conventional model, Synthetic Minority Oversampling Technique (SMOTE). We investigate the problem of the generative model for imbalanced classification and introduce a framework to enhance the SMOTE algorithm using Variational Autoencoders (VAE s). Our approach systematically quantifies the density of data points in a low-dimensional latent space using the VAE, simultaneously incorporating information on class labels and classification difficulty. Then, the data points potentially degrading the augmentation are systematically excluded, and the neighboring observations are directly augmented on the data space. Empirical studies on several imbalanced datasets represent that this simple process innovatively improves the conventional SMOTE algorithm over the deep learning models. Consequently, we conclude that the selection of minority data and the interpolation in the data space are beneficial for imbalanced classification problems with a relatively small number of data points. Sungchul Hong, Seunghwan An, Jong-June Jeon |
Appl. Intell. | 3 |
| 2025 | Adaptive adversarial augmentation for molecular property prediction
Soyoung Cho, Sungchul Hong, Jong-June Jeon |
Expert Syst. Appl. | 3 |
| 2025 | Variational autoencoder for distributional learning via quantile function estimationabstractThe Gaussianity assumption in Variational AutoEncoders (VAEs) enhances computational efficiency and provides a solid theoretical basis for estimating probability distributions. However, we have empirically found that approximating distributions with non-smooth densities using the Gaussian VAE is challenging. Therefore, we propose an approach for distributional learning in VAEs that extends to estimating the quantile function while accommodating both smooth and non-smooth densities. This is achieved by utilizing the continuous ranked probability score, a strictly proper scoring rule, as our reconstruction loss. Our method can be seen as a specialized form of a nonparametric M-estimator for estimating general quantile functions, and we establish a theoretical connection between our model and quantile estimation. Furthermore, we demonstrate that our reconstruction loss functions as the lower bound of an infinite mixture of asymmetric Laplace distributions, which allows our synthetic data generation mechanism to maintain differential privacy. We validate the effectiveness of our model in capturing the underlying distribution through experiments involving synthetic data generation on real-world tabular datasets, showing that the level of data privacy can be easily adjusted. Seunghwan An, Sungchul Hong, Jong-June Jeon |
Neural Networks | 3 |
| 2024 | Cryptocurrency Price Forecasting using Variational Autoencoder with Versatile Quantile ModelingabstractIn recent years, there has been a growing interest in probabilistic forecasting methods that offer more comprehensive insights by considering prediction uncertainties rather than point estimates. This paper introduces a novel variational autoencoder learning framework for multivariate distributional forecasting. Our approach employs distributional learning to directly estimate the cumulative distribution function of future time series conditional distributions using the continuous ranked probability score. By incorporating a temporal structure within the latent space and utilizing versatile quantile models, such as the generalized lambda distribution, we enable distributional forecasting by generating synthetic time series data for future time points. To assess the effectiveness of our method, we conduct experiments using a multivariate dataset of real cryptocurrency prices, demonstrating its superiority in forecasting high-volatility scenarios. Sungchul Hong, Seunghwan An, Jong-June Jeon |
CIKM | 3 |
| 2024 | Customization of latent space in semi-supervised Variational AutoEncoderabstractWe propose a novel semi-supervised learning method of Variational AutoEncoder (VAE), which yields a customized latent space through our EXplainable encoder Network (EXoN). The customization involves a manual design of the interpolation and structural constraint, such as proximity, which enhances the interpretability of the latent space. To improve the classification performance, we introduce a new semi-supervised classification method called SCI (Soft-label Consistency Interpolation). Combining the classification loss and the Kullback–Leibler divergence is crucial in constructing an explainable latent space. Additionally, the variability of the generated samples is determined by an active latent subspace, which effectively captures distinctive characteristics. We conduct experiments using the MNIST, SVHN, and CIFAR-10 datasets, and the results demonstrate that our approach yields an explainable latent space while significantly reducing the effort required to analyze representation patterns within the latent space. Seunghwan An, Jong-June Jeon |
Pattern Recognit. Lett. | 2 |
| 2023 | Causally Disentangled Generative Variational AutoEncoderabstractWe present a new supervised learning technique for the Variational AutoEncoder (VAE) that allows it to learn a causally disentangled representation and generate causally disentangled outcomes simultaneously. We call this approach Causally Disentangled Generation (CDG). CDG is a generative model that accurately decodes an output based on a causally disentangled representation. Our research demonstrates that adding supervised regularization to the encoder alone is insufficient for achieving a generative model with CDG, even for a simple task. Therefore, we explore the necessary and sufficient conditions for achieving CDG within a specific model. Additionally, we introduce a universal metric for evaluating the causal disentanglement of a generative model. Empirical results from both image and tabular datasets support our findings. Seunghwan An, Kyungwoo Song, Jong-June Jeon |
ECAI | 3 |
| 2023 | Distributional Learning of Variational AutoEncoder: Application to Synthetic Data GenerationabstractThe Gaussianity assumption has been consistently criticized as a main limitation of the Variational Autoencoder (VAE) despite its efficiency in computational modeling. In this paper, we propose a new approach that expands the model capacity (i.e., expressive power of distributional family) without sacrificing the computational advantages of the VAE framework. Our VAE model's decoder is composed of an infinite mixture of asymmetric Laplace distribution, which possesses general distribution fitting capabilities for continuous variables. Our model is represented by a special form of a nonparametric M-estimator for estimating general quantile functions, and we theoretically establish the relevance between the proposed model and quantile estimation. We apply the proposed model to synthetic data generation, and particularly, our model demonstrates superiority in easily adjusting the level of data privacy. Seunghwan An, Jong-June Jeon |
NeurIPS | 2 |
| 2023 | Geodesic Multi-Modal Mixup for Robust Fine-TuningabstractPre-trained multi-modal models, such as CLIP, provide transferable embeddings and show promising results in diverse applications. However, the analysis of learned multi-modal embeddings is relatively unexplored, and the embedding transferability can be improved. In this work, we observe that CLIP holds separated embedding subspaces for two different modalities, and then we investigate it through the lens of \textit{uniformity-alignment} to measure the quality of learned representation. Both theoretically and empirically, we show that CLIP retains poor uniformity and alignment even after fine-tuning. Such a lack of alignment and uniformity might restrict the transferability and robustness of embeddings. To this end, we devise a new fine-tuning method for robust representation equipping better alignment and uniformity. First, we propose a \textit{Geodesic Multi-Modal Mixup} that mixes the embeddings of image and text to generate hard negative samples on the hypersphere. Then, we fine-tune the model on hard negatives as well as original negatives and positives with contrastive loss. Based on the theoretical analysis about hardness guarantee and limiting behavior, we justify the use of our method. Extensive experiments on retrieval, calibration, few- or zero-shot classification (under distribution shift), embedding arithmetic, and image captioning further show that our method provides transferable representations, enabling robust model adaptation on diverse tasks. Changdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim, Minchul Shin, Jong-June Jeon, Kyungwoo Song |
NeurIPS | 6 |
| 2023 | Quantile estimation for encrypted data
Jaeseon Kim, Sungchul Shin, Cheolwoo Park, Jong-June Jeon, Soon-Sun Kwon, Hosik Choi |
Appl. Intell. | 5 |
| 2021 | Learning a High-dimensional Linear Structural Equation Model via l1-Regularized RegressionabstractThis paper develops a new approach to learning high-dimensional linear structural equation models (SEMs) without the commonly assumed faithfulness, Gaussian error distribution, and equal error distribution conditions. A key component of the algorithm is component-wise ordering and parent estimations, where both problems can be efficiently addressed using l1-regularized regression. This paper proves that sample sizes n = Omega( d^{2} \log p) and n = \Omega( d^2 p^{2/m} ) are sufficient for the proposed algorithm to recover linear SEMs with sub-Gaussian and (4m)-th bounded-moment error distributions, respectively, where p is the number of nodes and d is the maximum degree of the moralized graph. Further shown is the worst-case computational complexity O(n (p^3 + p^2 d^2 ) ), and hence, the proposed algorithm is statistically consistent and computationally feasible for learning a high-dimensional linear SEM when its moralized graph is sparse. Through simulations, we verify that the proposed algorithm is statistically consistent and computationally feasible, and it performs well compared to the state-of-the-art US, GDS, LISTEN and TD algorithms with our settings. We also demonstrate through real COVID-19 data that the proposed algorithm is well-suited to estimating a virus-spread map in China. Gunwoong Park, Sang Jun Moon, Sion Park, Jong-June Jeon |
J. Mach. Learn. Res. | 4 |
| 2018 | The sparse Luce model
Jong-June Jeon, Hosik Choi |
Appl. Intell. | 1 |