EDBT 2026 Demo / reviewers in the wild / expert
Yilun Liu 0005
dblp:51/8650-5
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0001-5448-5806ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning OnlyabstractZero-shot reasoning on text-rich networks (TRNs) remains a challenging frontier, as models must integrate textual semantics with relational structure without task-specific supervision.While graph neural networks rely on fixed label spaces and supervised objectives, recent large language model (LLM)-based approaches often overlook graph context or depend on distillation from larger models, limiting generalisation.We propose TRN-R1-Zero, a post-training framework for TRN reasoning trained solely via reinforcement learning.TRN-R1-Zero directly optimises base LLMs using a Neighbour-aware Group Relative Policy Optimisation objective that dynamically adjusts rewards based on a novel margin gain metric for the informativeness of neighbouring signals, effectively guiding the model toward relational reasoning.Unlike prior methods, TRN-R1-Zero requires no supervised fine-tuning or chain-of-thought data generated from large reasoning models.Extensive experiments across citation, hyperlink, social and co-purchase TRN benchmarks demonstrate the superiority and robustness of TRN-R1-Zero.Moreover, relying strictly on node-level training, TRN-R1-Zero achieves zero-shot inference on edgeand graph-level tasks, extending beyond crossdomain transfer. Yilun Liu 0005, Ruihong Qiu, Zi Huang |
ACL (1) | 1 |
| 2026 | LEXA: Legal case retrieval via graph contrastive learning with contextualised LLM embeddingsabstractAbstract Legal case retrieval (LCR) is a specialised information retrieval task aimed at identifying relevant cases given a query case. LCR holds pivotal significance in facilitating legal practitioners to locate legal precedents. Existing LCR methods predominantly rely on traditional lexical models or language models; however, they typically overlook the domain-specific structural information embedded in legal documents. Our previous work CaseGNN (Tang et al., In: ECIR, 2024) successfully harnesses text-attributed graphs and graph neural networks to incorporate structural legal information. Nonetheless, three key challenges remain in enhancing the representational capacity of CaseGNN: (1) The under-utilisation of rich edge information in text-attributed case graph (TACG). (2) The insufficiency of training signals for graph contrastive learning. (3) The lack of contextualised legal information in node and edge features. In this paper, the LEXA model, an extension of CaseGNN, is proposed to overcome these limitations by jointly leveraging rich edge information, enhanced training signals, and contextualised embeddings derived from large language models (LLMs). Specifically, an edge-updated graph attention layer (EUGAT) is proposed to comprehensively update node and edge features during graph modelling, resulting in a full utilisation of structural information of legal cases. Moreover, LEXA incorporates a novel graph contrastive learning objective with graph augmentation to provide additional training signals, thereby strengthening the model’s legal comprehension capabilities. What’s more, given the remarkable contextualised understanding capabilities of LLMs for text encoding, LLMs are employed to generate node and edge features for the text-attributed case graph (TACG). Extensive experiments on two benchmark datasets from COLIEE 2022 and COLIEE 2023 demonstrate that LEXA not only significantly improves CaseGNN but also achieves supreme performance compared to state-of-the-art LCR methods. Code has been released on https://github.com/yanran-tang/CaseGNN . Yanran Tang, Ruihong Qiu, Yilun Liu 0005, Xue Li 0001, Zi Huang |
World Wide Web (WWW) | 3 |
| 2025 | GCondenser: Benchmarking Graph CondensationabstractLarge-scale graphs are valuable for graph representation learning, but the vast volume of data often hinders model building efficiency. Graph condensation (GC) addresses this challenge by compressing a large graph into a significantly smaller one that still supports effective model training. While recent studies have proposed various techniques to enhance condensation effectiveness, comprehensive and practical evaluations of these methods remain limited. In this paper, we introduce GCondenser, a large-scale graph condensation toolkit designed to facilitate flexible development, holistic evaluation and comparison of mainstream GC approaches. GCondenser provides a standardised GC pipeline with condensation, validation, and evaluation stages, and offers straightforward extensibility to accommodate new methods and datasets. Additionally, we conduct a thorough empirical study of existing GC methods, offering insights into multiple facets of condensation performance. The toolkit is available at https://github.com/superallen13/GCondenser. Yilun Liu 0005, Ruihong Qiu, Zi Huang |
CIKM | 1 |
| 2025 | PUMA: Efficient Continual Graph Learning for Node Classification With Graph CondensationabstractWhen handling streaming graphs, existing graph representation learning models encounter a catastrophic forgetting problem, where previously learned knowledge of these models is easily overwritten when learning with newly incoming graphs. In response, Continual Graph Learning (CGL) emerges as a novel paradigm enabling graph representation learning from static to streaming graphs. Our prior work, Condense and Train (CaT) (Liu et al. 2023) is a replay-based CGL framework with a balanced continual learning procedure, which designs a small yet effective memory bank for replaying data by condensing incoming graphs. Although the CaT alleviates the catastrophic forgetting problem, there exist three issues: (1) The graph condensation algorithm derived in CaT only focuses on labelled nodes while neglecting abundant information carried by unlabelled nodes; (2) The continual training scheme of the CaT overemphasises on the previously learned knowledge, limiting the model capacity to learn from newly added memories; (3) Both the condensation process and replaying process of the CaT are time-consuming. In this paper, we propose aPsUdo-label guidedMemory bAnk (PUMA) CGL framework, extending from the CaT to enhance its efficiency and effectiveness by overcoming the above-mentioned weaknesses and limits. To fully exploit the information in a graph, PUMA expands the coverage of nodes during graph condensation with both labelled and unlabelled nodes. Furthermore, a training-from-scratch strategy is proposed to upgrade the previous continual learning scheme for a balanced training between the historical and the new graphs. Besides, PUMA uses a one-time prorogation and wide graph encoders to accelerate the graph condensation and the graph encoding process in the training stage to improve the efficiency of the whole framework. Extensive experiments on seven datasets for the node classification task demonstrate the state-of-the-art performance and efficiency over existing methods. Yilun Liu 0005, Ruihong Qiu, Yanran Tang, Hongzhi Yin, Zi Huang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | CaseGNN: Graph Neural Networks for Legal Case Retrieval with Text-Attributed Graphs
Yanran Tang, Ruihong Qiu, Yilun Liu 0005, Xue Li 0001, Zi Huang |
ECIR (2) | 3 |
| 2023 | CaT: Balanced Continual Graph Learning with Graph CondensationabstractContinual graph learning (CGL) is purposed to continuously update a graph model with graph data being fed in a streaming manner. Since the model easily forgets previously learned knowledge when training with new-coming data, the catastrophic forgetting problem has been the major focus in CGL. Recent replay-based methods intend to solve this problem by updating the model using both (1) the entire new-coming data and (2) a sampling-based memory bank that stores replayed graphs to approximate the distribution of historical data. After updating the model, a new replayed graph sampled from the incoming graph will be added to the existing memory bank. Despite these methods are intuitive and effective for the CGL, two issues are identified in this paper. Firstly, most samplingbased methods struggle to fully capture the historical distribution when the storage budget is tight. Secondly, a significant data imbalance exists in terms of the scales of the complex newcoming graph data and the lightweight memory bank, resulting in unbalanced training. To solve these issues, a Condense and Train (CaT) framework is proposed in this paper. Prior to each model update, the new-coming graph is condensed to a small yet informative synthesised replayed graph, which is then stored in a Condensed Graph Memory with historical replay graphs. In the continual learning phase, a Training in Memory scheme is used to update the model directly with the Condensed Graph Memory rather than the whole new-coming graph, which alleviates the data imbalance problem. Extensive experiments conducted on four benchmark datasets successfully demonstrate superior performances of the proposed CaT framework in terms of effectiveness and efficiency. The code has been released on https://github.com/superallen13/CaT-CGL. Yilun Liu 0005, Ruihong Qiu, Zi Huang |
ICDM | 1 |