EDBT 2026 Demo / reviewers in the wild / expert
Lei Zhou 0016
dblp:72/5749-16
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiffKGR: Diffusion-Based Virtual Edge Generation for Knowledge Graph Recommendation
Lyuwen Wu, Xiaoying Gan, Luoyi Fu, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou |
DASFAA (1) | 4 |
| 2026 | Graph Out-of-Distribution Generalization Based on Structural-Entropy-Guided Information BottleneckabstractOut-of-Distribution (OOD) generalization is a promising yet challenging goal that guarantees the test performance of Graph Neural Networks (GNNs) in open-world settings. However, due to the intricate internal topology of graph-structured data, redundant information from the spurious topologies severely confuses GNNs to deviate from the labels. Extracting concise and label-relevant subgraphs from the original graphs can alleviate this problem. Unfortunately, existing methods either overlook the global structural distribution or rely heavily on manually predefined assumptions. As a result, they fall short of well capturing the structural distribution changes between input graph and extracted subgraph, thus compromising adaptability of extracted invariant subgraphs to diverse OOD scenarios. This motivates us to propose a framework called S tructural- E ntropy-guided I nformation B ottleneck (OOD-SEIB) that aims to more traceably measure the inherent information changes for better and more flexible OOD generalization. The core of OOD-SEIB lies in concise topology extraction module, where we measure the mutual information flow between input graph and extracted subgraph based on structural entropy, termed Compression Index (CI). Specifically, the CI is a quantifiable metric that calculates the codeword length required to describe entire graph structure via a biased random walk. Under this guidance, OOD-SEIB then launches a structural information bottleneck compression module that jointly optimizes both CI and label-relevance of the subgraph topology by iteratively balancing between informativeness and compression. To further improve GNN’s invariant subgraph identification capability, OOD-SEIB generates multiple augmented environments and distill the invariant subgraphs into GNN as knowledge in an inside-out manner. When iteratively optimizing in above prescribed way, OOD-SEIB progressively reinforce the invariant subgraph extraction, thereby enhancing its generalization capability. Extensive experiments on synthetic and three real-world graph-level OOD benchmarks demonstrate that our proposed OOD-SEIB improves classification accuracy by 4.85%–38.03% on average compared to state-of-the-art baselines. Additionally, we extend OOD-SEIB to two node-level benchmarks, achieving average classification accuracy improvements of 14.52% and 13.15%. Zijun Di, Bin Lu 0005, Luoyi Fu, Ningdi Jin, Xiaoying Gan, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou |
ACM Trans. Knowl. Discov. Data | 9 |
| 2025 | DeepReport: An AI-assisted Idea Generation System for Scientific ResearchabstractNowadays, the explosive growth of academic literature has been going far beyond scientists' limited capability to read through, making it increasingly difficult for them to absorb disciplinary insights and extract intellectual essences critical for generating novel research ideas in interdisciplinary studies. To address this, we develop DeepReport, an AI-assisted scientific idea generation system to alleviate the research burden. Technically, DeepReport maintains evolving concept co-occurrence graphs to extract core insights from over 260 million publications across all disciplines. These concepts are periodically collected and updated, enabling the automatic extraction of hidden cross-domain connections. Combining temporal link prediction and analysis techniques with large language models, DeepReport is able to further transform these patterns of insights into actionable ideas. With the function of integrating up-to-date academic databases, visualizing dynamic relationships of concepts, and automatically generating new ideas, DeepReport empowers researchers to navigate complex knowledge landscapes, reduce cognitive burdens, and accelerate the generation of groundbreaking concepts. This work provides an in-depth exploration of DeepReport's architecture, functionalities, and applications, highlighting its transformative potential for advancing interdisciplinary research and fostering innovation. DeepReport is available at https://idea.acemap.cn/. Yi Xu 0004, Luoyi Fu, Shuqian Sheng, Jiaxin Ding 0001, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou |
SIGIR | 6 |
| 2025 | Connectivity maintenance against link uncertainty and heterogeneity in adversarial networksabstractThis paper delves into the challenge of maintaining connectivity in adversarial networks, focusing on the preservation of essential links to prevent the disintegration of network components under attack. Unlike previous approaches that assume a stable and homogeneous network topology, this study introduces a more realistic model that incorporates both link uncertainty and heterogeneity. Link uncertainty necessitates additional probing to confirm link existence, while heterogeneity reflects the varying resilience of links against attacks. We model the network as a random graph where each link is defined by its existence probability, probing cost, and resilience. The primary objective is to devise a defensive strategy that maximizes the expected size of the largest connected component at the end of an adversarial process while minimizing the probing cost, irrespective of the attack patterns employed. We begin by establishing the NP-hardness of the problem and then introduce an optimal defensive strategy based on dynamic programming. Due to the high computational cost of achieving optimality, we also develop two approximate strategies that offer efficient solutions within polynomial time. The first is a heuristic method that assesses link importance across three heterogeneous subnetworks, and the second is an adaptive minimax policy designed to minimize the defender’s potential worst-case loss, with guaranteed performance. Through extensive testing on both synthetic and real-world datasets across various attack scenarios, our strategies demonstrate significant advantages over existing methods. Jianzhi Tang, Luoyi Fu, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou |
High Confid. Comput. | 3 |
| 2024 | RepEval: Effective Text Evaluation with LLM RepresentationabstractShuqian Sheng, Yi Xu, Tianhang Zhang, Zanwei Shen, Luoyi Fu, Jiaxin Ding, Lei Zhou, Xiaoying Gan, Xinbing Wang, Chenghu Zhou. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shuqian Sheng, Yi Xu 0004, Tianhang Zhang, Zanwei Shen, Luoyi Fu, Jiaxin Ding 0001, Lei Zhou 0016, Xiaoying Gan, Xinbing Wang, Chenghu Zhou |
EMNLP | 7 |
| 2024 | OxyGenerator: Reconstructing Global Ocean Deoxygenation Over a Century with Deep LearningabstractAccurately reconstructing the global ocean deoxygenation over a century is crucial for assessing and protecting marine ecosystem. Existing expert-dominated numerical simulations fail to catch up with the dynamic variation caused by global warming and human activities. Besides, due to the high-cost data collection, the historical observations are severely sparse, leading to big challenge for precise reconstruction. In this work, we propose OxyGenerator, the first deep learning based model, to reconstruct the global ocean deoxygenation from 1920 to 2023. Specifically, to address the heterogeneity across large temporal and spatial scales, we propose zoning-varying graph message-passing to capture the complex oceanographic correlations between missing values and sparse observations. Additionally, to further calibrate the uncertainty, we incorporate inductive bias from dissolved oxygen (DO) variations and chemical effects. Compared with in-situ DO observations, OxyGenerator significantly outperforms CMIP6 numerical simulations, reducing MAPE by 38.77%, demonstrating a promising potential to understand the “breathless ocean” in data-driven manner. Bin Lu 0005, Ze Zhao, Luyu Han, Xiaoying Gan, Yuntao Zhou, Lei Zhou 0016, Luoyi Fu, Xinbing Wang, Chenghu Zhou |
ICML | 6 |
| 2024 | Is Reference Necessary in the Evaluation of NLG Systems? When and Where?abstractShuqian Sheng, Yi Xu, Luoyi Fu, Jiaxin Ding, Lei Zhou, Xinbing Wang, Chenghu Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shuqian Sheng, Yi Xu 0004, Luoyi Fu, Jiaxin Ding 0001, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou |
NAACL-HLT | 5 |
| 2024 | Community Deception in Large Networks: Through the Lens of Laplacian SpectrumabstractMany complex networks in the real world have community structures. Typical examples include online social networks and ecology networks. While the identification of communities bears numerous practical applications, with the increasing awareness of data security and privacy concerns, the need to protect the community affiliations of individuals from disclosing by attackers emerges. This raises the community deception (CD) problem, that is, the opposite of community detection, which asks for ways to minimally perturb the network structures by rewiring nodes so that the target communities maximally hide themself from community detection algorithms. To this end, we investigate the CD problem through a Laplacian spectrum lens and propose a method named$\mathtt {ComDeceptor}$to hide a flexible target set of communities, which is more universal than most existing methods that either focus on hiding the entire communities or a single community. The key idea of$\mathtt {ComDeceptor}$is to first allocate the resources of perturbations fairly and effectively. By proving that hiding communities through intercommunity edge addition and intracommunity edge deletion correspond to maximizing the second smallest eigenvalue$\lambda _{2}$and minimizing the largest eigenvalue$\lambda _{n}$of the graph Laplacian, respectively,$\mathtt {ComDeceptor}$then incorporates efficient heuristics for approximately solving the problems, thus selecting the appropriate edge to perturb. Experimental results over nine real-world networks and six community detection algorithms not only demonstrate the efficiency of$\mathtt {ComDeceptor}$, but also the superior performance on obfuscating community structures over the baselines. Luoyi Fu, Jiaxin Ding 0001, Xinde Cao, Xinbing Wang, Lei Zhou 0016, Chenghu Zhou |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2024 | Hi-PART: Going Beyond Graph Pooling with Hierarchical Partition Tree for Graph-Level Representation LearningabstractGraph pooling refers to the operation that maps a set of node representations into a compact form for graph-level representation learning. However, existing graph pooling methods are limited by the power of the Weisfeiler–Lehman (WL) test in the performance of graph discrimination. In addition, these methods often suffer from hard adaptability to hyper-parameters and training instability. To address these issues, we propose Hi-PART, a simple yet effective graph neural network (GNN) framework with Hi erarchical Par tition T ree (HPT). In HPT, each layer is a partition of the graph with different levels of granularities that are going toward a finer grain from top to bottom. Such an exquisite structure allows us to quantify the graph structure information contained in HPT with the aid of structural information theory. Algorithmically, by employing GNNs to summarize node features into the graph feature based on HPT’s hierarchical structure, Hi-PART is able to adequately leverage the graph structure information and provably goes beyond the power of the WL test. Due to the separation of HPT optimization from graph representation learning, Hi-PART involves the height of HPT as the only extra hyper-parameter and enjoys higher training stability. Empirical results on graph classification benchmarks validate the superior expressive power and generalization ability of Hi-PART compared with state-of-the-art graph pooling approaches. Yuyang Ren, Haonan Zhang 0004, Luoyi Fu, Shiyu Liang, Lei Zhou 0016, Xinbing Wang, Xinde Cao, Chenghu Zhou |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Analyzing Information Cascading in Large Scale Networks: A Fixed Point ApproachabstractInformation cascading, referred as the phenomenon of an individual following the behavior of the preceding individual after observing its actions, is prevalent in real social networks and triggers intense research interests for the purpose of monitoring and controlling network epidemics. One of the typical lines of information cascading study belongs to the Influence Maximization Problem, which aims to algorithmically find the optimal seeds that can spread the information to the maximum number of nodes. Regardless of the tremendous efforts made in various algorithm design of finding such optimal seeds, it has not yet been well understood how the absolute influence power of the “optimal” source set affects the ultimate cascading, i.e., under which conditions the seeds are able or unable to influence an substantial fraction of the entire network. Most existing works have investigated the conditions of network scale influence under linear threshold model, where the activation of a node requires a large number of infected neighbors. Instead, in this paper we focus on the case of single source cascading, which is only possible to occur under the independent cascading model. We launch information cascading analysis from two aspects, i.e., the influence scale and network-scale cascading probability. Firstly, percolation analysis of the cascading outcome shows that estimating influence scale is equivalent to solving fixed point equations. Then, we investigate the speed and stability of information cascading based on fixed point analysis, which shows that the information cascading process almost surely terminates within logarithmic time complexity. Furthermore, the results are generalized to the stochastic block model, where we find that network-scale cascading is determined by the spectral radius of the community matrix. The analysis presented in this paper could help us better understand the conditions for different information cascading outcomes. Luoyi Fu, Jiasheng Xu, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | FlowerCast: Efficient Time-sensitive Multicast in Wireless Sensor Networks with Link UncertaintyabstractThis article studies time-sensitive multicast in wireless sensor networks (WSNs) with link uncertainty, where information from the source needs to be delivered to multiple receivers within an imposed delay constraint. Prior art on static WSNs minimizes the multicast delay via the construction of a multicast tree that approximates the Steiner tree in length, which, however, may be invalidated by the time-varying network topology of WSNs with uncertain link states. Moreover, for multicast in WSNs with link uncertainty, the possible link failure necessitates a suitable measurement of the uncertain communication distance and calls for the performance guarantee in both delay and delivery ratio. In this work, by modeling a WSN as a random graph with each link associated with a transmission probability, we propose FlowerCast, an efficient multicast scheme, to jointly minimize the expected multicast delay and to maximize the expected delivery ratio of multicast under delay constraint. The core of FlowerCast is to quantify the uncertain communication distance by the expected transmission delay of a time-varying path, based on which a delay-optimal multicast tree is constructed in accordance with the directionality of delay. Candidate paths with a high expected delivery ratio and low expected delay are then selected in a distributed manner to conditionally connect adjacent multicast members and thus transform the multicast tree into a multicast flower. Despite the NP-hardness of optimal candidate paths’ addition, the transformation with the highest expected delivery ratio of multicast under delay constraint can be guaranteed through a pseudo-polynomial time derandomization-based greedy approach. We further demonstrate the time and energy efficiency of FlowerCast through asymptotic analysis. To make full use of the possible overlapping links in a multicast flower, a hybrid routing strategy is presented to wisely switch between sequential routing and synchronous routing for extra enhancement of the multicast performance. Extensive experiments on various datasets verify the superiority of FlowerCast and hybrid routing over baselines and indicate their wide applicability to practical scenarios. Jianzhi Tang, Luoyi Fu, Shiyu Liang, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou |
ACM Trans. Sens. Networks | 5 |