Xingrui Zhuo

dblp:311/8749 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-8349-3597ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Relink: Constructing Query-Driven Evidence Graph On-the-Fly for GraphRAG
abstract
Graph-based Retrieval-Augmented Generation (GraphRAG) mitigates hallucinations in Large Language Models (LLMs) by grounding them in structured knowledge. However, current GraphRAG methods are constrained by a prevailing build-then-reason paradigm, which relies on a static, pre-constructed Knowledge Graph (KG). This paradigm faces two critical challenges. First, the KG's inherent incompleteness often breaks reasoning paths. Second, the graph’s low signal-to-noise ratio introduces distractor facts, presenting query-relevant but misleading knowledge that disrupts the reasoning process. To address these challenges, we argue for a reason-and-construct paradigm and propose Relink, a framework that dynamically builds a query-specific evidence graph. To tackle incompleteness, Relink instantiates required facts from a latent relation pool derived from the original text corpus, repairing broken paths on the fly. To handle misleading or distractor facts, Relink employs a unified, query-aware evaluation strategy that jointly considers candidates from both the KG and latent relations, selecting those most useful for answering the query rather than relying on their pre-existence. This empowers Relink to actively discard distractor facts and construct the most faithful and precise evidence path for each query. Extensive experiments on five Open-Domain Question Answering benchmarks show that Relink achieves significant average improvements of 5.4% in EM and 5.2% in F1 over leading GraphRAG baselines, demonstrating the superiority of our proposed framework.
Manzong Huang, Chenyang Bu, Yi He 0007, Xingrui Zhuo, Xindong Wu 0001
AAAI4
2025 Progressive Prefix-Memory Tuning for Complex Logical Query Answering on Knowledge Graphs
abstract
Conducting complex logical queries over knowledge graphs remains a significant challenge. Recent research has successfully leveraged Pre-trained Language Models (PLMs) to tackle Knowledge Graph Complex Query Answering (KGCQA) tasks, which is attributed to PLMs' ability to comprehend logical semantics of queries through context learning. However, existing PLM-based KGCQA methods usually overlook the harm of disordered syntax or fragmented contexts within a serialized query, posing the problem of “impossible language” to limit PLMs in grasping the logical semantics. To address this problem, we propose a Progressive Prefix-Memory Tuning (PPMT) framework for KGCQA tasks, which effectively rectifies erroneous segments in serialized queries to assist PLMs in query answering. First, we propose a prefix-memory rectification mechanism embedded in a PLM module. This mechanism assigns rectification parameters in memory stores to polish the language segments of entities, relations, and queries through specific prefixes. To further capture the logical semantics in queries, we design a progressive fine-tuning strategy, which optimizes our model through a conditional gradient update process guided by knowledge translation constraints. Extensive experiments on widely used KGCQA benchmarks demonstrate the significant superiority of PPMT in terms of HR@3 and MRR. Our codes are available at https://github.com/lazyloafer/PPMT.
Xingrui Zhuo, Shirui Pan, Jiapu Wang, Gong-Qing Wu, Zan Zhang 0002, Zizhong Wei, Xindong Wu 0001
IJCAI1
2025 Effective Instruction Parsing Plugin for Complex Logical Query Answering on Knowledge Graphs
abstract
Knowledge Graph Query Embedding (KGQE) aims to embed First-Order Logic (FOL) queries in a low-dimensional KG space for complex reasoning over incomplete KGs. To enhance the generalization of KGQE models, recent studies integrate various external information (such as entity types and relation context) to better capture the logical semantics of FOL queries. The whole process is commonly referred to as Query Pattern Learning (QPL). However, current QPL methods typically suffer from the pattern-entity alignment bias problem, leading to the learned defective query patterns limiting KGQE models' performance. To address this problem, we propose an effective Query Instruction Parsing Plugin (QIPP) that leverages the context awareness of Pre-trained Language Models (PLMs) to capture latent query patterns from code-like query instructions. Unlike the external information introduced by previous QPL methods, we first propose code-like instructions to express FOL queries in an alternative format. This format utilizes textual variables and nested tuples to convey the logical semantics within FOL queries, serving as raw materials for a PLM-based instruction encoder to obtain complete query patterns. Building on this, we design a query-guided instruction decoder to adapt query patterns to KGQE models. To further enhance QIPP's effectiveness across various KGQE models, we propose a query pattern injection mechanism based on compressed optimization boundaries and an adaptive normalization component, allowing KGQE models to utilize query patterns more efficiently. Extensive experiments demonstrate that our plug-and-play method improves the performance of eight basic KGQE models and outperforms two state-of-the-art QPL methods.
Xingrui Zhuo, Jiapu Wang, Gong-Qing Wu, Shirui Pan, Xindong Wu 0001
WWW1
2024 A Lightweight, Effective, and Efficient Model for Label Aggregation in Crowdsourcing
abstract
Due to the presence of noise in crowdsourced labels, label aggregation (LA) has become a standard procedure for post-processing these labels. LA methods estimate true labels from crowdsourced labels by modeling worker quality. However, most existing LA methods are iterative in nature. They require multiple passes through all crowdsourced labels, jointly and iteratively updating true labels and worker qualities until a termination condition is met. As a result, these methods are burdened with high space and time complexities, which restrict their applicability in scenarios where scalability and online aggregation are essential. Furthermore, defining a suitable termination condition for iterative algorithms can be challenging. In this article, we view LA as a dynamic system and represent it as a Dynamic Bayesian Network. From this dynamic model, we derive two lightweight and scalable algorithms: LAonepassand LAtwopass. These algorithms can efficiently and effectively estimate worker qualities and true labels by traversing all labels at most twice, thereby eliminating the need for explicit termination conditions and multiple traversals over the crowdsourced labels. Due to their dynamic nature, the proposed algorithms are also capable of performing label aggregation online. We provide theoretical proof of the convergence property of the proposed algorithms and bound the error of the estimated worker qualities. Furthermore, we analyze the space and time complexities of our proposed algorithms, demonstrating their equivalence to those of majority voting. Through experiments conducted on 20 real-world datasets, we demonstrate that our proposed algorithms can effectively and efficiently aggregate labels in both offline and online settings, even though they traverse all labels at most twice. The code is on https://github.com/yyang318/LA_onepass .
Yi Yang 0036, Zhong-Qiu Zhao, Gong-Qing Wu, Xingrui Zhuo, Qing Liu 0001, Quan Bai 0001, Weihua Li 0007
ACM Trans. Knowl. Discov. Data4
2024 Geometric-Contextual Mutual Infomax Path Aggregation for Relation Reasoning on Knowledge Graph
abstract
Relation reasoning inKnowledgeGraphCompletion (KGC) aims at predicting missing relations between entities. Recently, effective KGC methods have usually focused on exploring the path pattern between entities, such as reward-based path walking and path context mining, to complete target relations. However, these methods typically suffer from two challenges: 1) They have difficulty in handling the individual representation limitation of candidate paths when there are no paths that directly represent latent relations between entities; 2) They overlook the biases of path context induction, which leads to unreasonable information interfering with the model's reasoning. To manage these challenges, aGeometric-ContextualMutualInfomax (GCMI) path aggregator is proposed for relation reasoning. First, we design an attentive path aggregator with a shared Transformer encoder to capture the contexts from several candidate paths parallelly and integrate these contexts to sufficiently represent the latent relations of each entity pair for reasoning. Then, the GCMI modules are proposed to constrain the local and global biases of path context induction in the Transformer encoder and the path aggregator, respectively, by a straightforward geometric rule. Extensive experiments on 32 real-world relation reasoning tasks demonstrate that our method significantly outperforms 8 state-of-the-art baselines in terms of AP and AUC.
Xingrui Zhuo, Gong-Qing Wu, Zan Zhang 0002, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.1
2024 Multi-Hop Multi-View Memory Transformer for Session-Based Recommendation
abstract
A Session-Based Recommendation (SBR) seeks to predict users’ future item preferences by analyzing their interactions with previously clicked items. In recent approaches, Graph Neural Networks (GNNs) have been commonly applied to capture item relations within a session to infer user intentions. However, these GNN-based methods typically struggle with feature ambiguity between the sequential session information and the item conversion within an item graph, which may impede the model’s ability to accurately infer user intentions. In this article, we propose a novel Multi-hop Multi-view Memory Transformer (M 3 T) to effectively integrate the sequence-view information and relation conversion (graph-view information) of items in a session. First, we propose a Multi-view Memory Transformer (M 2 T) module to concurrently obtain multi-view information of items. Then, a set of trainable memory matrices are employed to store sharable item features, which mitigates cross-view item feature ambiguity. To comprehensively capture latent user intentions, an M 3 T framework is designed to integrate user intentions across different hops of an item graph. Specifically, a k-order power method is proposed to manage the item graph to alleviate the over-smoothing problem when obtaining high-order relations of items. Extensive experiments conducted on three real-world datasets demonstrate the superiority of our method.
Xingrui Zhuo, Shengsheng Qian, Jun Hu 0016, Fuxin Dai, Kangyi Lin, Gong-Qing Wu
ACM Trans. Inf. Syst.1
2023 Label Enhanced Graph Attention Network for Truth Inference
Ningjing Zhao, Xingrui Zhuo, Gong-Qing Wu, Zan Zhang 0002
ICANN (4)2
2023 Semantic-Reconstructed Graph Transformer Network for Event Detection
abstract
Event detection (ED) is a key subtask of information extraction to extract key events, such as stock rise and fall and social public opinion, from news or social media. Although current GCN-based event detection methods achieve remarkable success via building graphs with dependency trees, they typically suffer from two challenges: 1) They use sequence models to learn contextual information of sentences, ignoring the longterm dependencies problem of sequence models might learn ineffective information and make it propagate in GCN layers. 2) Most methods do not exploit global dependency label information and grammatical structure information that convey rich linguistic knowledge directly, and only consider local dependency label information. To cope with these challenges, we propose a novel event detection model via semantic-reconstructed graph transformer networks (SRGTNED), which incorporates semantic reconstruction and path information collection methods. Using the semantic reconstruction method, we assign a pruned sequence to each word based on the path information to capture contextual information consistent with sentence semantics. Moreover, to better utilize global dependency label information and grammatical structure information, a Graph Transformer Network (GTN)-based heterogeneous graph embedding framework is introduced to automatically learn path information between important words by converting sentences as heterogeneous graphs. We conduct experiments on the ACE2005 dataset and the Commodity News dataset, and the experimental results demonstrate that our method significantly outperforms 11 state-of-the-art baselines in terms of the F1-score.
Zhuochun Miao, Xingrui Zhuo, Gong-Qing Wu, Chenyang Bu
IJCNN2
2023 Multivariate Time Series Classification via Hierarchical Graph Embedding
abstract
Multivariate time series classification aims to determine the labels for multivariate time series samples. Although variable interaction relationships and sample similarity relationships exist in multivariate time series, the available related methods usually ignore the rich relationships and are ineffective in exploiting these. To solve this problem, we propose a Hierarchical Graph Embedding for Multivariate Time Series Classification (MTSC-HGE), which consists of a variable-wise attentive graph pooling module and a sample-wise graph convolutional module to obtain the relationships of variables and samples. Specifically, we design an attentive graph pooling module based on self-attention, which can obtain sample features fusing temporal patterns and variable interaction relationships in samples. Furthermore, we propose a graph mapping criterion that converts the MTS dataset into a graph based on dynamic time warping to explicitly reflect the similarity relationships between samples. To capture latent sample relationships, a GCN module is utilized on the sample graph to integrate sample features obtained from the attentive graph pool module. In addition, a classifier takes the rich representation output by the model to get the final predicted class. Extensive experiments on 14 public datasets show that MTSC-HGE significantly outperforms state-of-the-art baselines.
Wenhao Niu, Xingrui Zhuo, Gong-Qing Wu, Junwei Lv, Zan Zhang 0002, Chenyang Bu
IJCNN2
2023 Crowdsourcing Truth Inference via Reliability-Driven Multi-View Graph Embedding
abstract
Crowdsourcing truth inference aims to assign a correct answer to each task from candidate answers that are provided by crowdsourced workers. A common approach is to generate workers’ reliabilities to represent the quality of answers. Although crowdsourced triples can be converted into various crowdsourced relationships, the available related methods are not effective in capturing these relationships to alleviate the harm to inference that is caused by conflicting answers. In this research, we propose aReliability-drivenMulti-viewGraphEmbedding framework forTruthinference (TiReMGE), which explores multiple crowdsourced relationships by organically integrating worker reliabilities into a graph space that is constructed from crowdsourced triples. Specifically, to create an interactive environment, we propose a reliability-driven initialization criterion for initializing vectors of tasks and workers as interactive carriers of reliabilities. From the perspective of multiple crowdsourced relationships, a multi-view graph embedding framework is proposed for reliability information interaction on a task-worker graph, which encodes latent crowdsourced relationships into vectors of workers and tasks for reliability update and truth inference. A heritable reliability updating method based on the Lagrange multiplier method is proposed to obtain reliabilities that match the quality of workers for interaction by a novel constraint law. Our ultimate goal is to minimize the Euclidean distance between the encoded task vector and the answer that is provided by a worker with high reliability. Extensive experimental results on nine real-world datasets demonstrate that TiReMGE significantly outperforms the nine state-of-the-art baselines.
Gong-Qing Wu, Xingrui Zhuo, Xianyu Bao, Xuegang Hu, Richang Hong, Xindong Wu 0001
ACM Trans. Knowl. Discov. Data2
2023 TIRA: Truth Inference via Reliability Aggregation on Object-Source Graph
abstract
Crowdsourcing platforms collect massive dirty claims that are provided by sources for crowdsourced objects, which prompts truth inference to be proposed for crowdsourcing data denoising. Although current graph-based truth-inference methods achieve remarkable success by capturing complex crowdsourcing relationships, they typically suffer from two challenges: 1) They fail to obtain complete crowdsourcing relationships because of the structural limitations of crowdsourcing relationship graphs; 2) Their vector initialization methods for objects and sources are disturbed by claim noise, which limits them from obtaining correct object and source semantics. To cope with these challenges, we propose a novelTruth-Inference method viaReliabilityAggregation (TIRA) on an object-source graph. Specifically, we propose a hierarchical graph auto-encoder to adapt to a reasonable object-source graph, which enables TIRA to capture complete crowdsourcing relationships from multiple perspectives. To better guide TIRA, we design a vector initialization method based on source reliabilities to map the denoised claims to a representation space of objects and sources. Finally, TIRA aggregates the reliability information on an object-source graph to generate object embeddings for truth inference. We conducted extensive experiments on 12 real-world datasets. The experimental results demonstrate that our method significantly outperforms 12 state-of-the-art baselines in terms of the$accuracy$and$weighted\_{F}1$.
Gong-Qing Wu, Xingrui Zhuo, Liangzhu Zhou, Xianyu Bao, Richang Hong, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.2