VLDB 2026 Research / reviewers in the wild / expert
Yoonhyuk Choi
dblp:304/8407
· DBLP profile ↗
17ranked-venue papers
12as first author
17since 2021 · last 2026
0000-0003-4359-5596ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 11 first-author · 14 since 2021Databases, data management, data science and information retrieval · 9 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sheaf Graph Neural Networks via PAC-Bayes Spectral OptimizationabstractOver-smoothing in Graph Neural Networks (GNNs) causes collapse in distinct node features, particularly on heterophilic graphs where adjacent nodes often have dissimilar labels. Although sheaf neural networks partially mitigate this problem, they typically rely on static or heavily parameterized sheaf structures that hinder generalization and scalability. Existing sheaf-based models either predefine restriction maps or introduce excessive complexity, yet fail to provide rigorous stability guarantees. In this paper, we introduce a novel scheme called SGPC (Sheaf GNNs with PAC-Bayes Calibration), a unified architecture that combines cellular-sheaf message passing with several mechanisms, including optimal transport-based lifting, variance-reduced diffusion, and PAC-Bayes spectral regularization for robust semi-supervised node classification. We establish performance bounds theoretically and demonstrate that end-to-end training in linear computational complexity can achieve the resulting bound-aware objective. Experiments on nine homophilic and heterophilic benchmarks show that SGPC outperforms state-of-the-art spectral and sheaf-based GNNs while providing certified confidence intervals on unseen nodes. Yoonhyuk Choi, Jiho Choi, Taewook Ko, JongWook Kim, Chong-Kwon Kim |
AAAI | 1 |
| 2026 | Delay-Aware Sequential Recommendation with Dynamic Graphs and Variate-Temporal Decomposition
Yoonhyuk Choi |
SIGIR | 1 |
| 2026 | Identifying heterophilic neighbors via confidence-based subgraph matching for graph neural networksabstractGraph Neural Networks (GNNs) often struggle with heterophilic graphs, where neighboring nodes tend to have dissimilar labels-a common scenario in real-world networks. This paper addresses this limitation through a two-phase framework called ConSM (Confidence-based Subgraph Matching). First, we introduce a confidence-aware subgraph matching module that estimates edge coefficients by comparing the structural similarity of 2-hop neighborhoods using optimal transport. This process identifies task-irrelevant or misleading edges based on a tunable confidence ratio. Second, we integrate these edge coefficients into a sign-aware label propagation mechanism that adaptively encourages or discourages message passing based on edge confidence, thereby enhancing GNN robustness under heterophily. Compared to our earlier conference version [1], this manuscript provides (i) a clearer justification of key design choices such as subgraph-based reasoning and the use of 2-hop neighborhoods, (ii) an adaptive strategy to tune the confidence ratio without manual search, and (iii) extensive new experiments covering recent heterophily-oriented baselines and larger, leakage-free datasets. Empirical results show that ConSM improves classification accuracy, mitigates over-smoothing, and remains effective across both homophilic and heterophilic regimes. Yoonhyuk Choi, Chong-Kwon Kim |
Artif. Intell. | 1 |
| 2026 | Edge-conditioned Markov kernels for stable label refinement on heterophilous graphs
Taewook Ko, Yoonhyuk Choi |
Inf. Sci. | 2 |
| 2025 | Selective Blocking for Message-Passing Neural Networks on Heterophilic GraphsabstractGraph Neural Networks (GNNs) thrive on message passing (MP) but are vulnerable when the graph carries many heterophilic or misclassified edges. Prior analyses suggest that signed propagation can mitigate over-smoothing under low edge-error rates, yet they implicitly assume perfect edge labels and the presence of self-loops. We revisit this setting and show that, under high edge uncertainty, propagating any information may harm node separability even with signed weights. Our key insight is to decide not to propagate along uncertain edges adaptively. Concretely, we intentionally omit self-loops to isolate pure neighbor influence for a clearer theoretical analysis, adopt a row-stochastic (asymmetric) operator that matches the Markov-chain view of MP and simplifies spectral-radius proofs, and dynamically estimate the local homophily $b_i$ and edge-classification error $e_t$ during training via an EM procedure. We prove that our selective blocking yields a sub-stochastic propagation matrix whose joint spectral radius exceeds that of signed GNNs under high $e_t$, guaranteeing reduced over-smoothing, and we supply a lemma showing that class-discriminative signals survive even when the operator is rank-deficient. Extensive experiments on seven homophilic and heterophilic benchmarks confirm that the proposed adaptive blocking outperforms strong baselines. Yoonhyuk Choi, Taewook Ko, Jiho Choi, Chong-Kwon Kim |
UAI | 1 |
| 2025 | Mitigating Overfitting in Graph Neural Networks via Feature and Hyperplane PerturbationabstractMessage-passing neural networks are widely employed in various graph mining applications. However, these methods are susceptible to the scarcity of labeled data, which often leads to overfitting. Our observations suggest that sparse initial vectors further exacerbate this issue by failing to fully represent the range of learnable parameters. This sparsity can hinder the optimization of specific dimensions in the initial projection matrix, as the training samples may not adequately span these parameters. To overcome this challenge, we propose a novel perturbation technique that introduces variability to the initial features and the projection hyperplane. Notably, even without employing grid search, we demonstrate that shifting with a small estimated value mitigates this problem more effectively than other perturbation methods. Experimental results on real-world datasets reveal that our technique significantly enhances node classification accuracy in semi-supervised scenarios. Yoonhyuk Choi, Jiho Choi, Taewook Ko, Chong-Kwon Kim |
WSDM | 1 |
| 2025 | Review-Based Hyperbolic Cross-Domain RecommendationabstractThe issue of data sparsity poses a significant challenge to recommender systems. In response to this, algorithms that leverage side information such as review texts have been proposed. Furthermore, Cross-Domain Recommendation (CDR), which captures domain-shareable knowledge and transfers it from a richer domain (source) to a sparser one (target) has emerged recently. Nevertheless, existing methodologies assume an Euclidean embedding space, encountering difficulties in accurately representing richer text information and managing complex user-item interactions. This paper advocates a hyperbolic CDR approach for modeling review-based user-item relationships. We first emphasize that conventional distance-based domain alignment techniques may cause problems because small modifications in hyperbolic geometry result in magnified perturbations, ultimately leading to the collapse of hierarchical structures. To address this challenge, we propose hierarchy-aware embedding and domain alignment schemes that adjust the scale to extract domain-shareable information without disrupting structural forms. Extensive experiments substantiate the efficiency, robustness, and scalability of the proposed model. The source code is given here https://github.com/ChoiYoonHyuk/HEAD. Yoonhyuk Choi, Jiho Choi, Taewook Ko, Chong-Kwon Kim |
WSDM | 1 |
| 2025 | Beyond Binary: Improving Signed Message Passing in Graph Neural Networks for Multi-Class GraphsabstractGraph Neural Networks (GNNs) exhibit satisfactory performance on homophilic networks, where most edges connect two nodes with the same label. However, their effectiveness diminishes as the graphs become heterophilic (low homophily), prompting the exploration of various message-passing schemes. In particular, assigning negative weights to heterophilic edges (signed propagation) for message-passing has gained significant attention, and some studies theoretically confirm its effectiveness. Nevertheless, prior theorems assume binary classification scenarios, which may not hold well for graphs with multiple classes. To solve this limitation, we offer new theoretical insights into GNNs in multi-class environments and identify the drawbacks of employing signed propagation from two perspectives: message-passing and parameter update. We found that signed propagation without considering feature distribution can degrade the separability of dissimilar neighbors, which also increases prediction uncertainty (e.g., conflicting evidence) that can cause instability. To address these limitations, we introduce two novel calibration strategies aiming to improve discrimination power while reducing entropy in predictions. Through theoretical and extensive experimental analysis, we demonstrate that the proposed schemes enhance the performance of both signed and general message-passing neural networks (Choi et al. 2023). Yoonhyuk Choi, Taewook Ko, Jiho Choi, Chong-Kwon Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Beyond Message-Passing: Generalization of Graph Neural Networks via Feature Perturbation for Semi-Supervised Node ClassificationabstractGraph neural networks (GNNs) that collect information from neighbors are commonly utilized in semi-supervised learning contexts. In particular, a significant body of research has been dedicated to developing effective graph filters and aggregation methods to filter the information from adjacent nodes. Despite their efficacy, these approaches may encounter challenges due to the sparsity of training nodes, especially when their features are represented as sparse vectors (e.g., bag-of-words). This condition can lead to the overfitting of certain dimensions within the first projection matrix (hyperplane), as the training samples may not adequately represent the full spectrum of learnable parameters. To solve this limitation, we propose an innovative perturbation technique. Specifically, we introduce additional training variability by modifying both the initial features and the hyperplane, which contributes to the reduction of prediction variance by updating the entire dimensions. To the best of our knowledge, our approach is the first to address the overfitting issue in GNNs precipitated by sparse node features. Comprehensive experiments on real-world datasets and ablation studies affirm that our proposed method significantly enhances node classification performance, with improvements of up to 46.5% in GNN algorithms. Yoonhyuk Choi, Jiho Choi, Taewook Ko, Chong-Kwon Kim |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Improving the Text Convolution Mechanism with Large Language Model for Review-Based RecommendationabstractRecent studies in recommender systems focus on addressing data sparsity and cold-start problems by utilizing side information, such as tags, images, and testimonials. Among these, user-written testimonials (purchase reviews) are precious for analyzing personal preferences, and many methods have been developed based on this context. Generally, existing methods apply 2D text convolution followed by selecting important words using the attention mechanism. However, the text convolution scheme inevitably suffers from information loss since the number of words in reviews commonly exceeds hundreds. To address this limitation, we focus on the Large Language Model (LLM), which has shown promising results in various fields, including search engines, natural language processing, and healthcare. In particular, LLM has demonstrated excellent performance in text summarization and QA tasks, leading to the development of text-based recommender systems. Nevertheless, LLM alone struggles to perform collaborative filtering, which is essential in a recommender system. Thus, we propose LLM-based text summarization before applying 2D convolution, followed by the widely used collaborative filtering mechanism. This approach can improve recommendation quality by removing unnecessary words in advance, reducing the smoothing effect while capturing the rich user-item interactions. Our method is integrated with recent text-based recommendation algorithms, which have proven to improve the quality of all baselines by about 16.9 % on average. We conduct experiments and ablation studies using benchmark datasets, demonstrating that our method is scalable and efficient. Yoonhyuk Choi, Fahim Tasneema Azad |
IEEE Big Data | 1 |
| 2024 | Prioritizing Potential Wetland Areas via Region-to-Region Knowledge Transfer and Adaptive PropagationabstractWetlands are important to communities, offering benefits ranging from water purification, and flood protection to recreation and tourism. Therefore, identifying and prioritizing potential wetland areas is a critical decision problem. While data-driven solutions are feasible, this is complicated by significant data sparsity due to the low proportion of wetlands (3-6%) in many areas of interest in the southwestern US. This makes it hard to develop data-driven models that can help guide the identification of additional wetland areas. To solve this limitation, we propose two strategies: (1) knowledge transfer from regions with rich wetlands (such as the Eastern US) to regions with sparser wetlands (such as the Southwestern area). , and (2) spatial data enrichment strategy that relies on an adaptive propagation mechanism. This mechanism differentiates between node pairs that have positive and negative impacts on each other for Graph Neural Networks (GNNs). We conduct rigorous experiments to substantiate our proposed method's effectiveness, robustness, and scalability compared to state-of-the-art baselines. Additionally, an ablation study demonstrates that each module is essential in prioritizing potential wetlands. Yoonhyuk Choi, Reepal Shah, John Sabo, Huan Liu 0001, K. Selçuk Candan |
IEEE Big Data | 1 |
| 2024 | Introducing CausalBench: A Flexible Benchmark Framework for Causal Analysis and Machine Learning
Ahmet Kapkiç, Pratanu Mandal, Shu Wan 0002, Paras Sheth, Abhinav Gorantla, Yoonhyuk Choi, Huan Liu 0001, K. Selçuk Candan |
CIKM | 6 |
| 2023 | Universal Graph Contrastive Learning with a Novel Laplacian PerturbationabstractGraph Contrastive Learning (GCL) is an effective method for discovering meaningful patterns in graph data. By evaluating diverse augmentations of the graph, GCL learns discriminative representations and provides a flexible and scalable mechanism for various graph mining tasks. This paper proposes a novel contrastive learning framework by introducing Laplacian perturbation. The proposed framework offers a distinct advantage by employing an indirect perturbation method, which provides a more stable approach while maintaining the perturbation effects. Moreover, it exhibits a wide range of applicability by not being restricted to specific graph types. We demonstrate that a spectral graph convolution based on the Laplacian successfully extracts representations from diverse graph types. Our extensive experiments on a variety of real-world datasets, covering multiple graph types, show that the proposed model outperforms state-of-the-art baselines in both node classification and link sign prediction tasks. Taewook Ko, Yoonhyuk Choi, Chong-Kwon Kim |
UAI | 2 |
| 2023 | A spectral graph convolution for signed directed graphs via magnetic LaplacianabstractSigned directed graphs contain both sign and direction information on their edges, providing richer information about real-world phenomena compared to unsigned or undirected graphs. However, analyzing such graphs is more challenging due to their complexity, and the limited availability of existing methods. Consequently, despite their potential uses, signed directed graphs have received less research attention. In this paper, we propose a novel spectral graph convolution model that effectively captures the underlying patterns in signed directed graphs. To this end, we introduce a complex Hermitian adjacency matrix that can represent both sign and direction of edges using complex numbers. We then define a magnetic Laplacian matrix based on the adjacency matrix, which we use to perform spectral convolution. We demonstrate that the magnetic Laplacian matrix is positive semi-definite (PSD), which guarantees its applicability to spectral methods. Compared to traditional Laplacians, the magnetic Laplacian captures additional edge information, which makes it a more informative tool for graph analysis. By leveraging the information of signed directed edges, our method generates embeddings that are more representative of the underlying graph structure. Furthermore, we showed that the proposed method has wide applicability for various graph types and is the most generalized Laplacian form. We evaluate the effectiveness of the proposed model through extensive experiments on several real-world datasets. The results demonstrate that our method outperforms state-of-the-art techniques in signed directed graph embedding. Taewook Ko, Yoonhyuk Choi, Chong-Kwon Kim |
Neural Networks | 2 |
| 2022 | Finding Heterophilic Neighbors via Confidence-based Subgraph Matching for Semi-supervised Node ClassificationabstractGraph Neural Networks (GNNs) have proven to be powerful in many graph-based applications. However, they fail to generalize well under heterophilic setups, where neighbor nodes have different labels. To address this challenge, we employ a confidence ratio as a hyper-parameter, assuming that some of the edges are disassortative (heterophilic). Here, we propose a two-phased algorithm. Firstly, we determine edge coefficients through subgraph matching using a supplementary module. Then, we apply GNNs with a modified label propagation mechanism to utilize the edge coefficients effectively. Specifically, our supplementary module identifies a certain proportion of task-irrelevant edges based on a given confidence ratio. Using the remaining edges, we employ the widely used optimal transport to measure the similarity between two nodes with their subgraphs. Finally, using the coefficients as supplementary information on GNNs, we improve the label propagation mechanism which can prevent two nodes with smaller weights from being closer. The experiments on benchmark datasets show that our model alleviates over-smoothing and improves performance. Yoonhyuk Choi, Jiho Choi, Taewook Ko, Hyungho Byun, Chong-Kwon Kim |
CIKM | 1 |
| 2022 | Review-Based Domain Disentanglement without Duplicate Users or Contexts for Cross-Domain RecommendationabstractA cross-domain recommendation has shown promising results in solving data-sparsity and cold-start problems. Despite such progress, existing methods focus on domain-shareable information (overlapped users or same contexts) for a knowledge transfer, and they fail to generalize well without such requirements. To deal with these problems, we suggest utilizing review texts that are general to most e-commerce systems. Our model (named SER) uses three text analysis modules, guided by a single domain discriminator for disentangled representation learning. Here, we suggest a novel optimization strategy that can enhance the quality of domain disentanglement, and also debilitates detrimental information of a source domain. Also, we extend the encoding network from a single to multiple domains, which has proven to be powerful for review-based recommender systems. Extensive experiments and ablation studies demonstrate that our method is efficient, robust, and scalable compared to the state-of-the-art single and cross-domain recommendation methods. Yoonhyuk Choi, Jiho Choi, Taewook Ko, Hyungho Byun, Chong-Kwon Kim |
CIKM | 1 |
| 2022 | DiVa: An Accelerator for Differentially Private Machine LearningabstractThe widespread deployment of machine learning (ML) is raising serious concerns on protecting the privacy of users who contributed to the collection of training data. Differential privacy (DP) is rapidly gaining momentum in the industry as a practical standard for privacy protection. Despite DP’s importance, however, little has been explored within the computer systems community regarding the implication of this emerging ML algorithm on system designs. In this work, we conduct a detailed workload characterization on a state-of-the-art differentially private ML training algorithm named DPSGD. We uncover several unique properties of DP-SGD (e.g., its high memory capacity and computation requirements vs. non-private ML), root-causing its key bottlenecks. Based on our analysis, we propose an accelerator for differentially private ML named DiVa, which provides a significant improvement in compute utilization, leading to 2.6× higher energy-efficiency vs. conventional systolic arrays. Beomsik Park, Ranggi Hwang, Dongho Yoon, Yoonhyuk Choi, Minsoo Rhu |
MICRO | 4 |