EDBT 2026 Demo / reviewers in the wild / expert
Jie Zhang 0130
dblp:84/6889-130
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-4008-9703ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCSR: A Fast Data Structure with Leaf-Oriented Locks for Streaming Graph Processing
Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Xuejun An |
EDBT | 2 |
| 2026 | B-Graphless: Batch-based serverless graph processing for embodied AI backends
Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Xuejun An, Xiaochun Ye |
Future Gener. Comput. Syst. | 3 |
| 2025 | CGP-Graphless: Towards Efficient Serverless Graph Processing via CPU-GPU Pipelined Collaboration
Jie Zhang 0130, Huawei Cao, Xuejun An, Xiaochun Ye |
Euro-Par (1) | 3 |
| 2025 | A Co-Design Framework for Graph Processing on CPU-GPU Heterogeneous PlatformsabstractRecently, large-scale graph processing on CPU-GPU heterogeneous platforms has attracted considerable attention. However, disparities in memory bandwidth and parallel computational capabilities between CPUs and GPUs, coupled with the irregular structure of graphs and the inherent unpredictability of graph algorithms, often lead to inefficient utilization of CPU-GPU hardware resources, ultimately degrading graph processing performance. To address this, we propose and implement CoDgraph, a co-design framework for high-performance graph processing on CPU-GPU heterogeneous platforms. Specifically, we introduce a fine-grained partitioning strategy to balance workloads, minimize communication overhead, and enhance data locality. Next, we develop an adaptive co-scheduling computing scheme, leveraging a cost model that accounts for CPU and GPU hardware resources to improve system utilization. Finally, to further optimize largescale graph processing, we design and implement an efficient overlapping pipeline execution mode that employs asynchronous parallel execution. Extensive evaluations demonstrate that CoDgraph outperforms state-of-the-art CPU and CPU-GPU graph processing systems, including Ligra (CoDgraph is$14.59 \times$faster on average) and Subway (CoDgraph is$4.17 \times$faster on average). In addition, CoDgraph also has comparable performance to the advanced in-memory graph processing Tigr on GPU and shows good scalability for different graph scales and CPU-GPU heterogeneous platforms. Yuan Zhang 0031, Huawei Cao, Ming Dun, Jie Zhang 0130, Xiaochun Ye |
ICCD | 5 |
| 2025 | CGCGraph: Efficient CPU-GPU Co-execution for Concurrent Dynamic Graph ProcessingabstractWith the continuous growth of user scale and application data, the demand for large-scale concurrent graph processing is increasing. Typically, large-scale concurrent graph processing jobs need to process corresponding snapshots of dynamically changing graph data to obtain information at different time points. To enhance the throughput of such applications, current solutions concurrently process multiple graph snapshots on the GPU. However, when dealing with rapidly changing graph data, transferring multiple snapshots of concurrent jobs to the GPU results in high data transfer overhead between CPU and GPU. Additionally, the execution mode of existing work suffers from underutilization of GPU computational resources. In this work, we introduce CGCGraph, which can be integrated into existing GPU graph processing systems like Subway, to enable efficient concurrent graph snapshot processing jobs and enhance overall system resource utilization. The key idea is to offload unshared graph data of multiple concurrent snapshots to the CPU, reducing CPU-GPU transfer overhead. By implementing CPU-GPU co-execution, there is potential for enhanced utilization of GPU computing resources. Specifically, CGCGraph leverages kernel fusion to process shared graph data concurrently on the GPU, while executing all snapshots in parallel on the CPU, with each snapshot assigned a dedicated thread. This approach enables efficient concurrent processing within a novel CPU-GPU co-execution model, incorporating three optimization strategies targeting storage, computation, and synchronization. We integrate CGCGraph with Subway, an existing system designed for out-of-GPU-memory static graph processing. Experimental results show that the integration of CGCGraph with current GPU-based systems obtains performance improvements ranging from 1.7 to 4.5 times. Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Xuejun An, Junying Huang, Xiaochun Ye |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | ArkGPU: enabling applications' high-goodput co-location execution on multitasking GPUs
Jie Lou, Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Ninghui Sun |
CCF Trans. High Perform. Comput. | 3 |
| 2023 | FSGraph: fast and scalable implementation of graph traversal on GPUs
Yuan Zhang 0031, Huawei Cao, Jie Zhang 0130, Junying Huang, Xiaochun Ye, Xuejun An |
CCF Trans. High Perform. Comput. | 4 |