EDBT 2026 Demo / reviewers in the wild / expert
Yurui Lai
dblp:307/3251
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0000-4402-3798ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Simple yet Effective Graph Distillation via Clustering
Yurui Lai, Taiyan Zhang, Renchi Yang |
KDD (2) | 1 |
| 2025 | Leveraging Large Language Models for Effective Label-free Node Classification in Text-Attributed GraphsabstractGraph neural networks (GNNs) have become the preferred models for node classification in graph data due to their robust capabilities in integrating graph structures and attributes. However, these models heavily depend on a substantial amount of high-quality labeled data for training, which is often costly to obtain. With the rise of large language models (LLMs), a promising approach is to utilize their exceptional zero-shot capabilities and extensive knowledge for node labeling. Despite encouraging results, this approach either requires numerous queries to LLMs or suffers from reduced performance due to noisy labels generated by LLMs. To address these challenges, we introduce Locle, an active self-training framework that does Label-free nOde Classification with LLMs cost-Effectively. Locle iteratively identifies small sets of ''critical'' samples using GNNs and extracts informative pseudo-labels for them with both LLMs and GNNs, serving as additional supervision signals to enhance model training. Specifically, Locle comprises three key components: (i) an effective active node selection strategy for initial annotations; (ii) a careful sample selection scheme to identify ''critical'' nodes based on label disharmonicity and entropy; and (iii) a label refinement module that combines LLMs and GNNs with a rewired topology. Extensive experiments on five benchmark text-attributed graph datasets demonstrate that Locle significantly outperforms state-of-the-art methods under the same query budget to LLMs in terms of label-free node classification. Notably, on the DBLP dataset with 14.3k nodes, Locle achieves an 8.08% improvement in accuracy over the state-of-the-art at a cost of less than one cent. Our code is available at https://github.com/HKBU-LAGAS/Locle. Taiyan Zhang, Renchi Yang, Yurui Lai, Mingyu Yan, Xiaochun Ye, Dongrui Fan |
SIGIR | 3 |
| 2024 | Efficient Topology-aware Data Augmentation for High-Degree Graph Neural NetworksabstractIn recent years, graph neural networks (GNNs) have emerged as a potent tool for learning on graph-structured data and won fruitful successes in varied fields. The majority of GNNs follow the message-passing paradigm, where representations of each node are learned by recursively aggregating features of its neighbors. However, this mechanism brings severe over-smoothing and efficiency issues over high-degree graphs (HDGs), wherein most nodes have dozens (or even hundreds) of neighbors, such as social networks, transaction graphs, power grids, etc. Additionally, such graphs usually encompass rich and complex structure semantics, which are hard to capture merely by feature aggregations in GNNs.Motivated by the above limitations, we propose TADA, an efficient and effective front-mounted data augmentation framework for GNNs on HDGs. Under the hood, TADA includes two key modules: (i) feature expansion with structure embeddings, and (ii) topology- and attribute-aware graph sparsification. The former obtains augmented node features and enhanced model capacity by encoding the graph structure into high-quality structure embeddings with our highly-efficient sketching method. Further, by exploiting task-relevant features extracted from graph structures and attributes, the second module enables the accurate identification and reduction of numerous redundant/noisy edges from the input graph, thereby alleviating over-smoothing and facilitating faster feature aggregations over HDGs. Empirically, \algo considerably improves the predictive performance of mainstream GNN models on 8 real homophilic/heterophilic HDGs in terms of node classification, while achieving efficient training and inference processes. Yurui Lai, Xiaoyang Lin, Renchi Yang |
KDD | 1 |
| 2024 | Improved Topology Features for Node Classification on Heterophilic Graphs
Yurui Lai, Taiyan Zhang, Rui Fan 0004 |
ECML/PKDD (7) | 1 |
| 2022 | NCCR: Neighbor and Cluster Consistency Regularization for Improving Graph Node ClassificationabstractSemi-supervised node classification in graphs is a key problem in machine learning, and graph neural networks (GNNs) currently achieve state-of-the-art performance. However, traditional GNNs fail to make use of a substantial amount of information available in a graph. For example, the training loss is often defined only with respect to labeled training nodes, which usually make up a small proportion of the graph. Also, most graphs exhibit a certain degree of homophily, in which neighboring nodes are likely to belong to the same class, but GNNs typically do not make use of this property in an explicit way. In this work, we introduce a new type of consistency regularization which is able to make use of data from unlabeled nodes and also exploits graph homophily in a novel and more accurate way. Additionally, we observe that nodes in a graph may exhibit different amounts of homophily, so that uniformly enforcing neighbor consistency regularization across all nodes can reduce accuracy. We thus introduce a second clustering based regularization targeting low homophily nodes which lack reliable information from their neighbors. We show that we can flexibly combine the two regularizations with existing GNN backbones, and then demonstrate the effectiveness of the combined method by achieving state-of-the-art accuracy on a number of datasets. Feiming Yang, Yurui Lai, Leshan Wang, Rui Fan 0004 |
ICTAI | 2 |