EDBT 2026 Demo / reviewers in the wild / expert
Yuankai Luo
dblp:299/6707
· DBLP profile ↗
13ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0003-3844-7214ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Generative Recommendation via Simple Categorical User Sequence CompressionabstractAlthough generative recommenders demonstrate improved performance with longer sequences, their real-time deployment is hindered by substantial computational costs. To address this challenge, we propose a simple yet effective method for compressing long-term user histories by leveraging inherent item categorical features, thereby preserving user interests while enhancing efficiency. Experiments on two large-scale datasets demonstrate that, compared to the influential HSTU model, our approach achieves up to a 6× reduction in computational cost and up to 39% higher accuracy at comparable cost (i.e., similar sequence length). The source code will be available at https://github.com/Genemmender/CAUSE. Qijiong Liu, Zhongzhou Liu, Yuankai Luo, Guoyuan An, Nuo Chen 0004, Wei Guo 0006, Yong Liu 0020, Xiao-Ming Wu 0003 |
WSDM | 5 |
| 2025 | Beyond Random Masking: When Dropout meets Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) have emerged as powerful tools for learning on graph-structured data, yet the behavior of dropout in these models remains poorly understood. This paper presents a comprehensive theoretical analysis of dropout in GCNs, revealing that its primary role differs fundamentally from standard neural networks - preventing oversmoothing rather than co-adaptation. We demonstrate that dropout in GCNs creates dimension-specific stochastic sub-graphs, leading to a form of structural regularization not present in standard neural networks. Our analysis shows that dropout effects are inherently degree-dependent, resulting in adaptive regularization that considers the topological importance of nodes. We provide new insights into dropout's role in mitigating oversmoothing and derive novel generalization bounds that account for graph-specific dropout effects. Furthermore, we analyze the synergistic interaction between dropout and batch normalization in GCNs, uncovering a mechanism that enhances overall regularization. Our theoretical findings are validated through extensive experiments on both node-level and graph-level tasks across 14 datasets. Notably, GCN with dropout and batch normalization outperforms state-of-the-art methods on several benchmarks, demonstrating the practical impact of our theoretical insights. Yuankai Luo, Xiao-Ming Wu 0003 |
ICLR | 1 |
| 2025 | Node Identifiers: Compact, Discrete Representations for Efficient Graph LearningabstractWe present a novel end-to-end framework that generates highly compact (typically 6-15 dimensions), discrete (int4 type), and interpretable node representations—termed node identifiers (node IDs)—to tackle inference challenges on large-scale graphs. By employing vector quantization, we compress continuous node embeddings from multiple layers of a Graph Neural Network (GNN) into discrete codes, applicable under both self-supervised and supervised learning paradigms. These node IDs capture high-level abstractions of graph data and offer interpretability that traditional GNN embeddings lack. Extensive experiments on 34 datasets, encompassing node classification, graph classification, link prediction, and attributed graph clustering tasks, demonstrate that the generated node IDs significantly enhance speed and memory efficiency while achieving competitive performance compared to current state-of-the-art methods. Our source code is available at https://github.com/LUOyk1999/NodeID. Yuankai Luo, Hongkang Li, Qijiong Liu, Lei Shi 0002, Xiao-Ming Wu 0003 |
ICLR | 1 |
| 2025 | Can Classic GNNs Be Strong Baselines for Graph-level Tasks? Simple Architectures Meet ExcellenceabstractMessage-passing Graph Neural Networks (GNNs) are often criticized for their limited expressiveness, issues like over-smoothing and over-squashing, and challenges in capturing long-range dependencies. Conversely, Graph Transformers (GTs) are regarded as superior due to their employment of global attention mechanisms, which potentially mitigate these challenges. Literature frequently suggests that GTs outperform GNNs in graph-level tasks, especially for graph classification and regression on small molecular graphs. In this study, we explore the untapped potential of GNNs through an enhanced framework, GNN+, which integrates six widely used techniques: edge feature integration, normalization, dropout, residual connections, feed-forward networks, and positional encoding, to effectively tackle graph-level tasks. We conduct a systematic re-evaluation of three classic GNNs—GCN, GIN, and GatedGCN—enhanced by the GNN+ framework across 14 well-known graph-level datasets. Our results reveal that, contrary to prevailing beliefs, these classic GNNs consistently match or surpass the performance of GTs, securing top-three rankings across all datasets and achieving first place in eight. Furthermore, they demonstrate greater efficiency, running several times faster than GTs on many datasets. This highlights the potential of simple GNN architectures, challenging the notion that complex mechanisms in GTs are essential for superior graph-level performance. Our source code is available at https://github.com/LUOyk1999/GNNPlus. Yuankai Luo, Lei Shi 0002, Xiao-Ming Wu 0003 |
ICML | 1 |
| 2025 | Decoy Effect in Search Interaction: Understanding User Behavior and Measuring System VulnerabilityabstractThis study addresses (1) the influence of the decoy effect, a cognitive bias where the presence of an inferior item alters preferences between two options, on users’ search interactions and (2) the measurement of information retrieval systems’ vulnerability to the decoy effect. 1 From the perspective of user behavior, this study investigates the influence of the decoy effect in information retrieval (IR) by examining how decoy results affect users’ interaction on search engine result pages (SERPs), particularly in terms of click-through likelihood, browsing dwell time, and perceived document usefulness. We conducted an experiment based upon regression analysis on user interaction logs from three user study datasets which in total encompass 24 topics, 841 unique search sessions, and 2,685 queries. The findings indicate that decoys significantly increase the likelihood of document clicks and perceived usefulness. To investigate whether the influence of the decoy varies across different levels of task difficulty and user knowledge, we ran an additional experiment on one of the three datasets, which encompasses 6 topics, 166 search sessions and 652 queries. The results indicate that when the task is less challenging, users are more likely to click on a document with a decoy. Additionally, they spend more time on the target document and assign it a higher usefulness score. Furthermore, users with lower knowledge levels about the topic tend to give higher usefulness ratings to the target document. Regarding IR system evaluation, this study provides empirical insights into measuring the vulnerability of text retrieval models to potential decoy effect. An evaluation metric, namely DEcoy Judgement and Assessment VUlnerability (DEJA-VU), is proposed to evaluate the possibility of a retrieval model ranking results in a way that could trigger decoy biases. The experiments on the Text REtrieval Conference (TREC) 19 Deep Learning (DL) passage retrieval task and the TREC 20 DL passage retrieval task demonstrate that ColBERT and SPLADE show higher relevance-oriented retrieval effectiveness while also displaying lower vulnerability to decoy effect. Overall, this work advances the understanding of decoy effect, a well-established concept in cognitive psychology and behavioral economics, in a novel application field (i.e., Information Retrieval). It contributes to modeling users’ search behavior in the context of cognitive biases, as well as assessment of the vulnerability of systems and ranking algorithms to the decoy effect. Nuo Chen 0004, Jiqun Liu, Hanpei Fang, Yuankai Luo, Tetsuya Sakai, Xiao-Ming Wu 0003 |
ACM Trans. Inf. Syst. | 4 |
| 2025 | GeneticPrism: Multifaceted Visualization of Citation-Based Scholarly Research EvolutionabstractUnderstanding the evolution of scholarly research is essential for many real-life decision-making processes in academia, such as research planning, frontier exploration, and award selection. Popular platforms like Google Scholar and Web of Science rely on numerical indicators that are too abstract to convey the context and content of scientific research, while most existing visualization approaches on mapping science do not consider the presentation of individual scholars' research evolution using curated self-citation data. This paper builds on our previous work and proposes an integrated pipeline to visualize a scholar's research evolution from multiple topic facets. A novel 3D prism-shaped visual metaphor is introduced as the overview of a scholar's research profile, whilst their scientific evolution on each topic is displayed in a more structured manner. Additional designs by topic chord diagram, streamgraph visualization, and inter-topic flow map, optimized by an elaborate layout algorithm, assist in perceiving the scholar's scientific evolution across topics. A new six-degree-impact glyph metaphor highlights key interdisciplinary works driving the evolution. The proposed visualization methods are evaluated through case studies analyzing the careers of prestigious Turing award laureates, one major visualization venue, and a focused user study. Ye Sun 0009, Zipeng Liu, Yuankai Luo, Lei Shi 0002 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Classic GNNs are Strong Baselines: Reassessing GNNs for Node ClassificationabstractGraph Transformers (GTs) have recently emerged as popular alternatives to traditional message-passing Graph Neural Networks (GNNs), due to their theoretically superior expressiveness and impressive performance reported on standard node classification benchmarks, often significantly outperforming GNNs. In this paper, we conduct a thorough empirical analysis to reevaluate the performance of three classic GNN models (GCN, GAT, and GraphSAGE) against GTs. Our findings suggest that the previously reported superiority of GTs may have been overstated due to suboptimal hyperparameter configurations in GNNs. Remarkably, with slight hyperparameter tuning, these classic GNN models achieve state-of-the-art performance, matching or even exceeding that of recent GTs across 17 out of the 18 diverse datasets examined. Additionally, we conduct detailed ablation studies to investigate the influence of various GNN configurations—such as normalization, dropout, residual connections, and network depth—on node classification performance. Our study aims to promote a higher standard of empirical rigor in the field of graph machine learning, encouraging more accurate comparisons and evaluations of model capabilities. Our implementation is available at https://github.com/LUOyk1999/tunedGNN. Yuankai Luo, Lei Shi 0002, Xiao-Ming Wu 0003 |
NeurIPS | 1 |
| 2024 | Enhancing Graph Transformers with Hierarchical Distance Structural EncodingabstractGraph transformers need strong inductive biases to derive meaningful attention scores. Yet, current methods often fall short in capturing longer ranges, hierarchical structures, or community structures, which are common in various graphs such as molecules, social networks, and citation networks. This paper presents a Hierarchical Distance Structural Encoding (HDSE) method to model node distances in a graph, focusing on its multi-level, hierarchical nature. We introduce a novel framework to seamlessly integrate HDSE into the attention mechanism of existing graph transformers, allowing for simultaneous application with other positional encodings. To apply graph transformers with HDSE to large-scale graphs, we further propose a high-level HDSE that effectively biases the linear transformers towards graph hierarchies. We theoretically prove the superiority of HDSE in terms of expressivity and generalization. Empirically, we demonstrate that graph transformers with HDSE excel in graph classification, regression on 7 graph-level datasets, and node classification on 11 large-scale graphs. Yuankai Luo, Hongkang Li, Lei Shi 0002, Xiao-Ming Wu 0003 |
NeurIPS | 1 |
| 2023 | Impact-Oriented Contextual Scholar Profiling using Self-Citation GraphsabstractQuantitatively profiling a scholar's scientific impact is important to modern research society. Current practices with bibliometric indicators (e.g., h-index), lists, and networks perform well at scholar ranking, but do not provide structured context for scholar-centric, analytical tasks such as profile reasoning and understanding. This work presents GeneticFlow (GF), a suite of novel graph-based scholar profiles that fulfill three essential requirements: structured-context, scholar-centric, and evolution-rich. We propose a framework to compute GF over large-scale academic data sources with millions of scholars. The framework encompasses a new unsupervised advisor-advisee detection algorithm, a well-engineered citation type classifier using interpretable features, and a fine-tuned graph neural network (GNN) model. Evaluations are conducted on the real-world task of scientific award inference. Experiment outcomes show that the F1 score of best GF profile significantly outperforms alternative methods of impact indicators and bibliometric networks in all the 6 computer science fields considered. Moreover, the core GF profiles, with 63.6%\sim66.5% nodes and 12.5%\sim29.9% edges of the full profile, still significantly outrun existing methods in 5 out of 6 fields studied. Visualization of GF profiling result also reveals human explainable patterns for high-impact scholars. Yuankai Luo, Lei Shi 0002, Mufan Xu, Yuwen Ji, Fengli Xiao, Chunming Hu, Zhiguang Shan |
KDD | 1 |
| 2023 | Improving Self-supervised Molecular Representation Learning using Persistent HomologyabstractSelf-supervised learning (SSL) has great potential for molecular representation learning given the complexity of molecular graphs, the large amounts of unlabelled data available, the considerable cost of obtaining labels experimentally, and the hence often only small training datasets. The importance of the topic is reflected in the variety of paradigms and architectures that have been investigated recently, most focus on designing views for contrastive learning.
In this paper, we study SSL based on persistent homology (PH), a mathematical tool for modeling topological features of data that persist across multiple scales. It has several unique features which particularly suit SSL, naturally offering: different views of the data, stability in terms of distance preservation, and the opportunity to flexibly incorporate domain knowledge.
We (1) investigate an autoencoder, which shows the general representational power of PH, and (2) propose a contrastive loss that complements existing approaches.
We rigorously evaluate our approach for molecular property prediction and demonstrate its particular features in improving the embedding space:
after SSL, the representations are better and offer considerably more predictive power than the baselines over different probing tasks; our loss
increases baseline performance, sometimes largely; and we often obtain substantial improvements over very small datasets, a common scenario in practice. Yuankai Luo, Lei Shi 0002, Veronika Thost |
NeurIPS | 1 |
| 2023 | Transformers over Directed Acyclic GraphsabstractTransformer models have recently gained popularity in graph representation learning as they have the potential to learn complex relationships beyond the ones captured by regular graph neural networks.
The main research question is how to inject the structural bias of graphs into the transformer architecture,
and several proposals have been made for undirected molecular graphs and, recently, also for larger network graphs.
In this paper, we study transformers over directed acyclic graphs (DAGs) and propose architecture adaptations tailored to DAGs:
(1) An attention mechanism that is considerably more efficient than the regular quadratic complexity of transformers and at the same time faithfully captures the DAG structure, and (2) a positional encoding of the DAG's partial order, complementing the former.
We rigorously evaluate our approach over various types of tasks, ranging from classifying source code graphs to nodes in citation networks, and show that it is effective in two important aspects: in making graph transformers generally outperform graph neural networks tailored to DAGs and in improving SOTA graph transformer performance in terms of both quality and efficiency. Yuankai Luo, Veronika Thost, Lei Shi 0002 |
NeurIPS | 1 |
| 2023 | Mobility Inference on Long-Tailed Sparse TrajectoryabstractAnalyzing the urban trajectory in cities has become an important topic in data mining. How can we model the human mobility consisting of stay and travel states from the raw trajectory data? How can we infer these mobility states from a single user’s trajectory information? How can we further generalize the mobility inference to the real-world trajectory data that span multiple users and are sparsely sampled over time? In this article, based on formal and rigid definitions of the stay/travel mobility, we propose a single trajectory inference algorithm that utilizes a generic long-tailed sparsity pattern in the large-scale trajectory data. The algorithm guarantees a 100% precision in the stay/travel inference with a provable lower bound in the recall metric. Furthermore, we design a transformer-like deep learning architecture on the problem of mobility inference from multiple sparse trajectories. Several adaptations from the standard transformer network structure are introduced, including the singleton design to avoid the negative effect of sparse labels in the decoder side, the customized space-time embedding on features of location records, and the mask apparatus at the output side for loss function correction. Evaluations on three trajectory datasets of 40 million urban users validate the performance guarantees of the proposed inference algorithm and demonstrate the superiority of our deep learning model, in comparison to sequence learning methods in the literature. On extremely sparse trajectories, the deep learning model improves from the single trajectory inference algorithm with more than two times of overall and F1 accuracy. The model also generalizes to large-scale trajectory data from different sources with good scalability. Lei Shi 0002, Yuankai Luo, Shuai Ma 0001, Hanghang Tong, Zhetao Li, Zhiguang Shan |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | Attention Based Short-Term Metro Passenger Flow Prediction
Linjiang Zheng, Xuanxuan Luo, Congjun Xie, Yuankai Luo |
KSEM | 6 |