EDBT 2026 Demo / reviewers in the wild / expert
Pengpeng Zhou
dblp:205/7633
· DBLP profile ↗
12ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Push and Pull: Defending against Retrieval Poisoning Attacks via Embedding Space ReshapingabstractRetrieval-Augmented Generation (RAG) improves the performance of Large Language Models (LLMs) by retrieving and integrating relevant information from external knowledge bases, which helps generate more accurate responses. However, RAG is vulnerable to retrieval poisoning attacks , where attackers can induce LLM to produce inaccurate responses by injecting malicious documents into the retrieval process. In this article, we propose ShieldRAG , a novel defense framework designed to counteract retrieval poisoning attacks by reshaping the retrieval embedding space. ShieldRAG leverages a dual-strategy effect realized via a majority-consensus mechanism: ① Push : Implicitly forces the embedding of a user query away from malicious documents by filtering out their minority signals, reducing their influence. ② Pull : Aligns the embedding of a user query closer to that of benign documents, reinforcing accurate retrieval. These strategies work synergistically to preserve retrieval integrity and enhance the quality of LLM-generated responses. Specifically, ShieldRAG operates through three key steps: Sliding Retrieval Explanation Generation , Keyword Aggregation , and Query Targeting Optimization . These three steps collectively ensure the effective integration of information from benign sources while filtering out malicious interference, thereby significantly enhancing the robustness of RAG systems against retrieval poisoning attacks. We evaluate ShieldRAG on four open-domain Question Answering (QA) datasets: Natural Questions, MS-MARCO, HotpotQA, and 2WikiMultiHopQA, using seven representative LLMs. Extensive experiments demonstrate that ShieldRAG significantly improves response accuracy while mitigating adversarial effects, showcasing strong generalization across multiple datasets and LLM architectures. Longzhu He, Chaozhuo Li, Zheng Liu 0011, Pengpeng Zhou, Sen Su |
ACM Trans. Inf. Syst. | 6 |
| 2026 | High-Performance RoCE-Capable Multicast for Commodity RDMA DatacentersabstractModern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g.,$5.2\times $faster multicast communication and$2.7\times $higher replication throughput for distributed storage. Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005 |
IEEE Trans. Netw. | 7 |
| 2025 | An improved hierarchical neural network model with local and global feature matching for script event prediction
Pengpeng Zhou, Bin Wu 0001, Caiyong Wang, Longzhu He |
Expert Syst. Appl. | 1 |
| 2025 | Mitigating privacy risks in Retrieval-Augmented Generation via locally private entity perturbation
Longzhu He, Peng Tang 0002, Yuanhe Zhang, Pengpeng Zhou, Sen Su |
Inf. Process. Manag. | 4 |
| 2025 | P4KVS: A Role-Replica Separation Offloading Method to Achieve In-Network Consistency for KV Stores Based on P4 SwitchesabstractStrong consistency, particularly linearizability, is essential for distributed DBMSs deployed in correctness-critical domains such as finance and defense. In general, an optimal linearizability DBMS system focus on two key principles: (1) matching single-node (no-consistency cost) Read/Write performance under strong consistency, and (2) practical deployability via general database compatibility. Unfortunately, existing solutions fall short on both fronts. %However, achieving strong consistency often comes with steep performance penalties. For example, etcd-a widely-used Raft-based system-achieves only ~5.9% of the throughput of LevelDB, a single-node store without consistency overhead. To achieve higher performance, software approaches adopt weaker consistency models (e.g., ZAB), rely on narrow network assumptions (e.g., NOPaxos), or expose protocol internals to clients (e.g., CURP), yet still fail to close the performance gap. Recent programmable networking hardware offers promising advances, yet current hardware solutions face practical limitations, including minimal storage and incompatibility with general-purpose databases. We propose P4KVS, the first practical Raft-based in-network consensus offloading solution leveraging programmable switches (P4) for distributed key-value stores. P4KVS offloads only the Leader role to the switch while retaining Followers on servers. Under linearizability, it achieves 74% of single-node LevelDB's throughput for write-heavy workloads, and up to 222.4% for read-heavy workloads by distributing reads across three replicas. This demonstrates that, even under strong consistency, P4KVS can match or exceed the performance of a single-node system. Compared to etcd (which also uses Raft), P4KVS delivers 37.5× higher read throughput and 3520× lower write latency. These results validate our hardware role-replica separation design in eliminating software Raft bottlenecks, while preserving compatibility via standard database interfaces (e.g., LevelDB, etcd) and scaling beyond typical switch memory constraints. Haojuan Li, Zongpu Zhang, Chenzhen Ye, Ruohan Tang, Jian Li 0021, Haibing Guan, Qiaoling Wang, Pengpeng Zhou |
Proc. ACM Manag. Data | 8 |
| 2025 | A language-guided cross-modal semantic fusion retrieval method
Ligu Zhu, Suping Wang, Lei Shi 0030, Feifei Kou, Pengpeng Zhou |
Signal Process. | 7 |
| 2024 | Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable MulticastabstractModern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g., 5.2 × faster multicast communication and 2.7 × higher replication throughput for distributed storage. Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005 |
HPCA | 7 |
| 2022 | What happens next? Combining enhanced multilevel script learning and dual fusion strategies for script event predictionabstractScript event prediction (SEP), aiming at predicting next event from context event sequences (i.e., scripts), has played an important role in many real-world applications such as government decision-making. While most of the existing research only depend on the top-level event prediction, they ignore the influence of other bottom levels or other relationship modeling manners. In this paper, we focus on the problem of SEP via multilevel script learning where the goal of is to explore a multistage, multiprediction and multilevel information fusion model for SEP. This is challenging in (1) simultaneously modeling of the multilevel event relationship semantic information and (2) effectively designing multilevel information fusion strategies. In this paper, we propose a new script event prediction model based on Enhanced Multilevel script learning and Dual Fusion strategies, named EMDF-Net. Specifically, EMDF-Net designs the multilevel (event/chain/segment level) script learning to model both temporal and casual information as well as the rich structural relevance via neural stacking of self-attention mechanism and graph neural networks. Then it proposes dual fusion strategies to fully integrate different-level information by nonlinear feature composition and weighted score fusion. Finally, a deep supervision strategy is utilized to end-to-end train the whole model and provide a good initialization for information fusion. Experimental results on the popular NYT corpus demonstrate the effectiveness and superiority of EMDF-Net. Pengpeng Zhou, Bin Wu 0001, Caiyong Wang, Hao Peng 0001, Juwei Yue, Song Xiao 0004 |
Int. J. Intell. Syst. | 1 |
| 2020 | LogSayer: Log Pattern-driven Cloud Component Anomaly Diagnosis with Machine LearningabstractAnomaly diagnosis is a critical task for building a reliable cloud system and speeding up the system recovery form failures. With the increase of scales and applications of clouds, they are more vulnerable to various anomalies, and it is more challenging for anomaly troubleshooting. System logs that record significant events at critical time points become excellent sources of information to perform anomaly diagnosis. Never-theless, existing log-based anomaly diagnosis approaches fail to achieve high precision in highly concurrent environments due to interleaved unstructured logs. Besides, transient anomalies that have no obvious features are hard to detect by these approaches. To address this gap, this paper proposes LogSayer, a log pattern-driven anomaly detection model. LogSayer represents the system state by identifying suitable statistical features (e.g. frequency, surge), which are not sensitive to the exact log sequence. It then measures changes in the log pattern when a transient anomaly occurs. LogSayer uses Long Short-Term Memory (LSTM) neural networks to learn the historical correlation of log patterns and applies a BP neural network for adaptive anomaly decisions. Our experimental evaluations over the HDFS and OpenStack data sets show that LogSayer outperforms the state-of-the-art log-based approaches with precision over 98%. Pengpeng Zhou, Yang Wang 0147, Zhenyu Li 0001, Xin Wang 0001, Gareth Tyson, Gaogang Xie |
IWQoS | 1 |
| 2020 | Logchain: Cloud workflow reconstruction & troubleshooting with unstructured logs
Pengpeng Zhou, Yang Wang 0147, Zhenyu Li 0001, Gareth Tyson, Hongtao Guan, Gaogang Xie |
Comput. Networks | 1 |
| 2019 | An Adaptive Cross-Layer Sampling-Based Node Embedding for Multiplex NetworksabstractNetwork embedding aims to learn a latent representation of each node which preserves the structure information. Many real-world networks have multiple dimensions of nodes and multiple types of relations. Therefore, it is more appropriate to represent such kind of networks as multiplex networks. A multiplex network is formed by a set of nodes connected in different layers by links indicating interactions of different types. However, existing random walk based multiplex networks embedding algorithms have problems with sampling bias and imbalanced relation types, thus leading the poor performance in the downstream tasks. In this paper, we propose a node embedding method based on adaptive cross-layer forest fire sampling (FFS) for multiplex networks (FFME). We first focus on the sampling strategies of FFS to address the bias issue of random walk. We utilize a fixed-length queue to record previously visited layers, which can balance the edge distribution over different layers in sampled node sequences. In addition, to adaptively sample node's context, we also propose a metric for node called Neighbors Partition Coefficient (N P C ). The generation process of node sequence is supervised by NPC for adaptive cross-layer sampling. Experiments on real-world networks in diverse fields show that our method outperforms the state-of-the-art methods in application tasks such as cross-domain link prediction and shared community structure detection. Nianwen Ning, Chenguang Song, Pengpeng Zhou, Bin Wu 0001 |
ICTAI | 3 |
| 2017 | An entity disambiguation method based on LeaderRankabstractEntity Disambiguation is commonly faced in semantic search and knowledge base population. However, it is a challenging task because of the diversity of mentions. Previous methods can be classified into two main groups. One focuses on disambiguating mentions in a document independently and mainly relies on the local context similarity. The other collectively disambiguates mentions only taking into account link information. These are not appropriate when the context and the link information are poor or misleading. In this paper, we propose a new method to collectively disambiguate mentions in documents. Our proposed framework considers three features, including text similarity, entity popularity, and entity relationship. First we adopt LeaderRank algorithm on the graph model to rank entities according to the link information among entities. Then we combine with global text similarity between entity and document to disambiguate mentions. Our detailed experimental evaluation on two benchmark datasets demonstrates our methods is effective. Bingjing Jia, Bin Wu 0001, Jinna Lv, Pengpeng Zhou, Yao Bu |
IEEE BigData | 4 |