Xiaohua Li 0004

dblp:33/4610-4 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Adapting Language Models to Text Matching Based Recommendation Systems
Haidong Xin, Sen Mei, Zhenghao Liu 0001, Xiaohua Li 0004, Minghe Yu 0001, Yu Gu 0002, Ge Yu 0001
WISA4
2025 Secure Sharing of Health Data Based on Crosschain: Secure and Available System
abstract
With the rapid growth of multi-source and heterogeneous health data, achieving efficient and trustworthy crosschain sharing while ensuring privacy and security has become a critical challenge in blockchain-based healthcare applications. To address this, this paper proposes a secure health data sharing method based on a cross-chain blockchain architecture. A fourlayer system framework is constructed, and two original core modules, D-FBAC and HAEPP, are designed. The D-FBAC module introduces a reputation-weighted attribute mechanism to achieve resilient and fine-grained access control. The HAEPP module combines hybrid encryption strategies with a key lifecycle management mechanism to enhance privacy protection during both on-chain and cross-chain data flows. To improve system interpretability and user trust, an off-chain explainable access rejection module is also integrated, enabling rule-level traceability and feedback for access denials. The experimental evaluation is conducted on real-world healthcare datasets using a comprehensive set of metrics, results show that the proposed system significantly outperforms existing approaches across multiple performance dimensions, offering valuable insights for the interoperability and trusted governance of healthcare information systems.
Hanyu Mao, Tiezheng Nie, Minghe Yu 0001, Xiaomei Dong, Xiaohua Li 0004, Ge Yu 0001
BIBM5
2025 LCHGNN: Towards Distributed Hypergraph Neural Network Training Based on Communication Graphs with Lightweight Communication Optimization
abstract
Hypergraph Neural Networks (HGNNs) build on Graph Neural Networks (GNNs) by using hyperedges to capture complex, high-order relationships in data. However, training HGNNs on large hypergraphs is limited by computational and memory bottlenecks on a single machine. To overcome this, we propose LCHGNN, a distributed training method based on a new data structure called the communication graph, which simplifies hypergraph communication by representing cut hyperedges as vertices for structured message passing. LCHGNN employs a vertex-centric, hyperedge-replication-based storage scheme and introduces specialized forward and backward propagation mechanisms tailored for distributed execution. To mitigate communication overhead, We propose a lightweight optimization strategy that employs full synchronization in the initial round, followed by lightweight synchronization in subsequent rounds. Additionally, we present a learnable semi-supervised synchronization (LSS) aggregation mechanism for adaptive hyperedge selection. Extensive experiments on benchmark datasets demonstrate that LCHGNN preserves training accuracy while substantially reducing communication costs and enhancing scalability. This work addresses a critical gap in distributed HGNN research by delivering a communication-efficient and scalable training method, thereby facilitating the application of hypergraph learning to large-scale problems.
Taibo Wang, Yu Gu 0002, Xinning Cui, Zhen Song 0004, Xiaohua Li 0004, Fangfang Li 0002
CIKM5
2025 MassBFT: Fast and Scalable Geo-Distributed Byzantine Fault-Tolerant Consensus
abstract
Geo-distributed consensus protocols provide high availability and resilience for distributed database services. These protocols group nodes by their data centers to leverage the network topology that spans across multiple data centers, thereby reducing costly cross-datacenter communication. However, they still face performance and scalability challenges due to inefficient log replication mechanisms. 1) These protocols rely on the leader node in each group to perform cross-datacenter log replication, creating a single-node performance bottleneck. 2) Byzantine receivers can behave arbitrarily, forcing the group leader to send multiple log copies during replication to prevent loss, thus causing redundant transmissions. 3) Since all groups must execute these logs in the same order, synchronizations across groups are necessary to maintain consistency when multiple groups are proposing concurrently, which also slow down log replication. This paper presents MassBFT, a Byzantine fault-tolerant geo-consensus protocol that achieves high performance and scalability. We design an encoded bijective log replication to eliminate the leader bottleneck and reduce the cross-datacenter network consumption. We also propose asynchronous log ordering to eliminate synchronization across groups. Experimental results show that MassBFT is scalable, fault-tolerant, and outperforms state-of-the-art protocols with 5.49-29.96 times higher throughput under YCSB, SmallBank, and TPC-C workloads.
Zeshun Peng, Yanfeng Zhang 0001, Tinghao Feng, Weixing Zhou, Xiaohua Li 0004, Ge Yu 0001
ICDE5
2025 Enhancing the Patent Matching Capability of Large Language Models via the Memory Graph
abstract
Intellectual Property (IP) management involves strategically protecting and utilizing intellectual assets to enhance organizational innovation, competitiveness, and value creation. Patent matching is a crucial task in intellectual property management, which facilitates the organization and utilization of patents. Existing models often rely on the emergent capabilities of Large Language Models (LLMs) and leverage them to identify related patents directly. However, these methods usually depend on matching keywords and overlook the hierarchical classification and categorical relationships of patents. In this paper, we propose MemGraph, a method that augments the patent matching capabilities of LLMs by incorporating a memory graph derived from their parametric memory. Specifically, MemGraph prompts LLMs to traverse their memory to identify relevant entities within patents, followed by attributing these entities to corresponding ontologies. After traversing the memory graph, we utilize extracted entities and ontologies to improve the capability of LLM in comprehending the semantics of patents. Experimental results on the PatentMatch dataset demonstrate the effectiveness of MemGraph, achieving a 17.68% performance improvement over baseline LLMs. The further analysis highlights the generalization ability of MemGraph across various LLMs, both in-domain and out-of-domain, and its capacity to enhance the internal reasoning processes of LLMs during patent matching. All data and codes are available at https://github.com/NEUIR/MemGraph.
Qiushi Xiong, Zhenghao Liu 0001, Mengjia Wang, Zulong Chen, Yu Gu 0002, Xiaohua Li 0004, Ge Yu 0001
SIGIR8
2025 Self-Guided Graph Refinement With Progressive Fusion for Multiplex Graph Contrastive Representation Learning
abstract
Multiplex Graph Contrastive Learning (MGCL) has attracted significant attention. However, existing MGCL methods often struggle with suboptimal graph structures and fail to fully capture intricate interdependencies across multiplex views. To address these issues, we propose a novel self-supervised framework, Multiplex Graph Refinement with progressive fusion (MGRefine), for multiplex graph contrastive representation learning. Specifically, MGRefine introduces a multi-view learning module to extract a structural guidance matrix by exploring the underlying relationships between nodes. Then, a progressive fusion module is employed to progressively enhance and fuse representations from different views, capturing and leveraging nuanced interdependencies and comprehensive information across the multiplex graphs. The fused representation is then used to construct a consensus guidance matrix. A self-enhanced refinement module continuously refines the multiplex graphs using these guidance matrices while providing effective supervision signals. MGRefine achieves mutual reinforcement between graph structures and representations, ensuring continuous optimization of the model throughout the learning process in a self-enhanced manner. Extensive experiments demonstrate that MGRefine outperforms state-of-the-art methods and also verify the effectiveness of MGRefine across various downstream tasks on several benchmark datasets.
Yu Gu 0002, Xiaofeng Zhu 0001, Xiaohua Li 0004, Fangfang Li 0002, Ge Yu 0001
IEEE Trans. Big Data4
2025 Multi-Evidence Based Fact Verification via A Confidential Graph Neural Network
abstract
Fact verification tasks aim to identify the integrity of textual contents according to the truthful corpus. Existing fact verification models usually build a fully connected reasoning graph, which regards claim-evidence pairs as nodes and connects them with edges. They employ the graph to propagate the semantics of the nodes. Nevertheless, the noisy nodes usually propagate their semantics via the edges of the reasoning graph, which misleads the semantic representations of other nodes and amplifies the noise signals. To mitigate the propagation of noisy semantic information, we introduce a Confidential Graph Attention Network (CO-GAT), which proposes a node masking mechanism for modeling the nodes. Specifically, CO-GAT calculates the node confidence score by estimating the relevance between the claim and evidence pieces. Then, the node masking mechanism uses the node confidence scores to control the noise information flow from the vanilla node to the other graph nodes. CO-GAT achieves a 73.59% FEVER score on the FEVER dataset and shows the generalization ability by broadening the effectiveness to the science-specific domain.
Yuqing Lan, Zhenghao Liu 0001, Yu Gu 0002, Xiaoyuan Yi, Xiaohua Li 0004, Liner Yang, Ge Yu 0001
IEEE Trans. Big Data5
2024 Diffusion Model-Enhanced Contrastive Learning for Graph Representation
Yumeng Song, Yu Gu 0002, Fangfang Li 0002, Xiaohua Li 0004
DASFAA (6)5
2024 Hammer: A General Blockchain Evaluation Framework
abstract
With the rising proliferation of blockchain systems and applications, choosing the appropriate blockchains to deploy applications is critical to achieving optimal performance. Evaluation frameworks provide a systematic approach to assessing and comparing different blockchain systems, guiding application developers to choose the most suitable one. However, existing evaluation frameworks still have limitations that affect their accuracy. First, most frameworks utilize workloads initially de-signed for traditional databases, which fail to capture the unique characteristics and requirements of blockchain systems. Second, these frameworks fail to generate correct results under heavy workloads due to their imbalanced task processing algorithms. Third, existing frameworks are tailored only for non-sharding blockchain architectures, limiting their ability to evaluate diverse blockchains. This paper introduces Hammer, a general blockchain evaluation framework that addresses the above limitations. It consists of two key components: workload prediction and asynchronous task processing. Workload prediction accurately predicts real-world workload trends by expanding the scope of temporal control sequences, providing a more realistic evaluation of blockchain performance. Asynchronous task processing handles heavy-load situations, enabling accurate evaluation of blockchain performance. Extensive experiments on various blockchains under Smallbank workload empower application developers to make informed decisions about blockchain selection and optimization.
Gang Wang 0012, Yanfeng Zhang 0001, Chenhao Ying 0001, Xiaohua Li 0004, Ge Yu 0001
ICDCS4
2024 HIChain: A Hierarchical IoT Permissioned Blockchain with Edge Cloud Architecture
Tinghao Feng, Zeshun Peng, Yanfeng Zhang 0001, Xiaohua Li 0004, Xiaomei Dong, Ge Yu 0001
WISE (3)4
2023 Text Matching Improves Sequential Recommendation by Reducing Popularity Biases
abstract
This paper proposes Text mAtching based SequenTial rEcommenda-tion model (TASTE), which maps items and users in an embedding space and recommends items by matching their text representations. TASTE verbalizes items and user-item interactions using identifiers and attributes of items. To better characterize user behaviors, TASTE additionally proposes an attention sparsity method, which enables TASTE to model longer user-item interactions by reducing the self-attention computations during encoding. Our experiments show that TASTE outperforms the state-of-the-art methods on widely used sequential recommendation datasets. TASTE alleviates the cold start problem by representing long-tail items using full-text modeling and bringing the benefits of pretrained language models to recommendation systems. Our further analyses illustrate that TASTE significantly improves the recommendation accuracy by reducing the popularity bias of previous item id based recommendation models and returning more appropriate and text-relevant items to satisfy users. All codes are available at https://github.com/OpenMatch/TASTE.
Zhenghao Liu 0001, Sen Mei, Chenyan Xiong, Xiaohua Li 0004, Shi Yu 0001, Zhiyuan Liu 0001, Yu Gu 0002, Ge Yu 0001
CIKM4
2023 CLNIE: A Contrastive Learning Based Node Importance Evaluation Method for Knowledge Graphs with Few Labels
Yumeng Song, Yu Gu 0002, Xiaohua Li 0004, Fangfang Li 0002
DASFAA (2)4
2022 Efficient Subhypergraph Containment Queries on Hypergraph Databases
Yang Song 0022, Xiaohua Li 0004, Fangfang Li 0002, Yu Gu 0002
WISA3
2022 CSGNN: Improving Graph Neural Networks with Contrastive Semi-supervised Learning
Yumeng Song, Yu Gu 0002, Xiaohua Li 0004, Chuanwen Li, Ge Yu 0001
DASFAA (1)3
2022 Dimension Reduction for Efficient Dense Retrieval via Conditional Autoencoder
abstract
Dense retrievers encode queries and documents and map them in an embedding space using pre-trained language models.These embeddings need to be high-dimensional to fit training signals and guarantee the retrieval effectiveness of dense retrievers.However, these highdimensional embeddings lead to larger index storage and higher retrieval latency.To reduce the embedding dimensions of dense retrieval, this paper proposes a Conditional Autoencoder (ConAE) to compress the high-dimensional embeddings to maintain the same embedding distribution and better recover the ranking features.Our experiments show that ConAE is effective in compressing embeddings by achieving comparable ranking performance with its teacher model and making the retrieval system more efficient.Our further analyses show that ConAE can alleviate the redundancy of the embeddings of dense retrieval with only one linear layer.All codes of this work are available at https://github.com/NEUIR/ConAE.
Zhenghao Liu 0001, Chenyan Xiong, Zhiyuan Liu 0001, Yu Gu 0002, Xiaohua Li 0004
EMNLP6
2022 NeuChain: A Fast Permissioned Blockchain System with Deterministic Ordering
abstract
Blockchain serves as a replicated transactional processing system in a trustless distributed environment. Existing blockchain systems all rely on an explicit ordering step to determine the global order of transactions that are collected from multiple peers. The ordering consensus can be the bottleneck since it must be Byzantine-fault tolerant and can scarcely benefit from parallel execution. In this paper, we propose an ordering-free architecture that makes ordering implicit through deterministic execution. Based on this novel architecture, we develop a permissioned blockchain system NeuChain. A number of key optimizations such as asynchronous block generation and pipelining are leveraged for high throughput and low latency. Several security mechanisms are also designed to make our system robust to malicious attacks. Our geo-distributed experimental results show that NeuChain can achieve 47.2--64.1X throughput improvement over HyperLedger Fabric and 1.6--12.2X throughput improvement over the state-of-the-art high performance blockchains.
Zeshun Peng, Yanfeng Zhang 0001, Haixu Liu, Yuxiao Gao, Xiaohua Li 0004, Ge Yu 0001
Proc. VLDB Endow.6
2019 A Blockchain Based Secure E-Commerce Transaction System
Yun Zhang 0020, Xiaohua Li 0004, Jili Fan, Tiezheng Nie, Ge Yu 0001
WISA2
2017 Refreshment of the shortest path cache with change of single edge
Xiaohua Li 0004, Tao Qiu, Ning Wang 0003, Xiaochun Yang 0001, Bin Wang 0015, Ge Yu 0001
Expert Syst. Appl.1
2016 An Update Method for Shortest Path Caching with Burst Paths Based on Sliding Windows
Xiaohua Li 0004, Ning Wang 0003, Kanggui Peng, Xiaochun Yang 0001, Ge Yu 0001
WAIM (2)1
2014 Refreshment Strategies for the Shortest Path Caching Problem with Changing Edge Weight
Xiaohua Li 0004, Tao Qiu, Xiaochun Yang 0001, Bin Wang 0015, Ge Yu 0001
APWeb1