Pengjie Cui

dblp:239/4418 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0008-3753-490XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Nezha: An Efficient Distributed Graph Processing System on Heterogeneous Hardware
abstract
The growing scale of graph data across various applications demands efficient distributed graph processing systems. Despite the widespread use of the Scatter-Gather model for large-scale graph processing across distributed machines, the performance still can be significantly improved as the computation ability of each machine is not fully utilized and the communication costs during graph processing are expensive in the distributed environment. In this work, we propose a novel and efficient distributed graph processing system Nezha on heterogeneous hardware, where each machine is equipped with both CPU and GPU processors and all these machines in the distributed cluster are interconnected via Remote Direct Memory Access (RDMA).To reduce the communication costs, we devise an effective communication mode with a graph-friendly communication protocol in the graph-based RDMA communication adapter of Nezha. To improve the computation efficiency, we propose a multi-device cooperative execution mechanism in Nezha, which fully utilizes the CPU and GPU processors of each machine in the distributed cluster. We also alleviate the workload imbalance issue at inter-machine and intra-machine levels via the proposed workload balancer in Nezha. We conduct extensive experiments by running 4 widely-used graph algorithms on 5 graph datasets to demonstrate the superiority of Nezha over existing systems.
Pengjie Cui, Dong Jiang 0004, Bo Tang 0016, Ye Yuan 0001
Proc. ACM Manag. Data1
2025 Errata for "CGgraph: An Ultra-fast Graph Processing System on Modern Commodity CPU-GPU Co-processor"
Pengjie Cui, Bo Tang 0016, Ye Yuan 0001
Proc. VLDB Endow.1
2024 CGgraph: An Ultra-fast Graph Processing System on Modern Commodity CPU-GPU Co-processor
abstract
In recent years, many CPU-GPU heterogeneous graph processing systems have been developed in both academic and industrial to facilitate large-scale graph processing in various applications, e.g., social networks and biological networks. However, the performance of existing systems can be significantly improved by addressing two prevailing challenges: GPU memory over-subscription and efficient CPU-GPU cooperative processing. In this work, we propose CGgraph, an ultra-fast CPU-GPU graph processing system to address these challenges. In particular, CGgraph overcomes GPU-memory over-subscription by extracting a subgraph which only needs to be loaded into GPU memory once, but its vertices and edges can be used in multiple iterations during the graph processing procedure. To support efficient CPU-GPU co-processing, we design a CPU-GPU cooperative processing scheme, which balances the workloads between CPU and GPU by on-demand task allocation. To evaluate the efficiency of CG-graph, we conduct extensive experiments, comparing it with 7 state-of-the-art systems using 4 well-known graph algorithms on 6 real-world graphs. Our prototype system CGgraph outperforms all existing systems, delivering up to an order of magnitude improvement. Moreover, CGgraph on a modern commodity machine with a CPU-GPU co-processor yields superior (or at the very least, comparable) performance compared to existing systems on a high-end CPU-GPU server.
Pengjie Cui, Bo Tang 0016, Ye Yuan 0001
Proc. VLDB Endow.1
2019 Local Experts Finding Across Multiple Social Networks
Yuliang Ma 0001, Ye Yuan 0001, Guoren Wang, Yishu Wang 0001, Delong Ma, Pengjie Cui
DASFAA (2)6