Cheng Zhao 0001

dblp:93/3598-1 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 6 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Efficient Graph Data Access for Out-of-Memory GPU Streaming Graph Processing
abstract
Leveraging GPUs' high parallelism can significantly improve the real-time computation efficiency of streaming graph processing. However, when a large-scale graph exceeds GPU memory capacity, CPU-GPU cooperative processing often results in substantial and irregular CPU-to-GPU data transfer overhead. This stems from the extensive redundant graph accesses during continuous computation, which can hardly be addressed by existing solutions. In this work, we present Grapin, an out-of-memory GPU streaming graph processing system designed to minimize graph data transfer via two effective techniques for eliminating redundant accesses: (1) Extending advanced incremental processing algorithms to GPUs by converting their heavyweight data dependency processing into GPU-friendly forms, eliminating redundant graph accesses from the computation side; and (2) providing a lightweight yet efficient GPU hot subgraph management framework that finely caches the frequently accessed dynamic subgraphs in a vertex-centric manner. Experimental results demonstrate that Grapin can efficiently process large-scale streaming graphs with billions of edges on a single NVIDIA A5000 GPU. Enabling incremental computation reduces data transfer by 61%, and the integration of GPU hot subgraph reuse further reduces the remaining transfer by 72%, resulting in a total reduction of 89%. Compared with CPU-based solutions, Grapin achieves speedups ranging from 1.8x to 96.9x (17.9x on average).
Qiange Wang, Yongze Yan, Hongshi Tan, Cheng Chen 0008, Cheng Zhao 0001, Jiaming Tian, Xiaoliang Cong, Yanfeng Zhang 0001, Ge Yu 0001, Weng-Fai Wong, Bingsheng He
Proc. VLDB Endow.5
2024 Exploring Performance and Cost Optimization with ASIC-Based CXL Memory
abstract
As memory-intensive applications continue to drive the need for advanced architectural solutions, Compute Express Link (CXL) has risen as a promising interconnect technology that enables seamless high-speed, low-latency communication between host processors and various peripheral devices. In this study, we explore the application performance of ASIC CXL memory in various data-center scenarios. We then further explore multiple potential impacts (e.g., throughput, latency, and cost reduction) of employing CXL memory via carefully designed policies and strategies. Our empirical results show the high potential of CXL memory, reveal multiple intriguing observations of CXL memory and contribute to the wide adoption of CXL memory in real-world deployment environments. Based on our benchmarks, we also develop an Abstract Cost Model that can estimate the cost benefit from using CXL memory.
Yupeng Tang, Henry Hu, Tongping Liu, Jiaxin Shan, Ruoyun Huang, Cheng Zhao 0001, Cheng Chen 0008, Xiaoning Ding, Jianjun Chen 0001
EuroSys10
2024 FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
abstract
Dynamic graph random walk (DGRW) emerges as a practical tool for capturing structural relations within a graph. Effectively executing DGRW on GPU presents certain challenges. First, existing sampling methods demand a pre-processing buffer, causing substantial space complexity. Moreover, the power-law distribution of graph vertex degrees introduces workload imbalance issues, rendering DGRW embarrassed to parallelize. In this paper, we propose FlowWalker, a GPU-based dynamic graph random walk framework. FlowWalker implements an efficient parallel sampling method to fully exploit the GPU parallelism and reduce space complexity. Moreover, it employs a sampler-centric paradigm alongside a dynamic scheduling strategy to handle the huge amounts of walking queries. FlowWalker stands as a memory-efficient framework that requires no auxiliary data structures in GPU global memory. We examine the performance of FlowWalker extensively on ten datasets, and experiment results show that FlowWalker achieves up to 752.2×, 72.1×, and 16.4× speedup compared with existing CPU, GPU, and FPGA random walk frameworks, respectively. Case study shows that FlowWalker diminishes random walk time from 35% to 3% in a pipeline of ByteDance friend recommendation GNN training.
Junyi Mei, Shixuan Sun, Chao Li 0009, Cheng Chen 0008, Jing Wang 0055, Cheng Zhao 0001, Xiaofeng Hou, Minyi Guo, Bingsheng He, Xiaoliang Cong
Proc. VLDB Endow.8
2022 On the maxima of motzkin-straus programs and cliques of graphs
Qingsong Tang, Xiangde Zhang, Cheng Zhao 0001
J. Glob. Optim.3
2016 An extension of the Motzkin-Straus theorem to non-uniform hypergraphs and its applications
Yuejian Peng, Qingsong Tang, Cheng Zhao 0001
Discret. Appl. Math.4
2014 Some results on Lagrangians of hypergraphs
Qingsong Tang, Yuejian Peng, Xiangde Zhang, Cheng Zhao 0001
Discret. Appl. Math.4
2013 A novel iris and chaos-based random number generator
Hegui Zhu, Cheng Zhao 0001, Xiangde Zhang, Lianping Yang
Comput. Secur.2
2013 A novel image encryption-compression scheme using hyper-chaos and Chinese remainder theorem
Hegui Zhu, Cheng Zhao 0001, Xiangde Zhang
Signal Process. Image Commun.2
2008 Generating non-jumping numbers recursively
Yuejian Peng, Cheng Zhao 0001
Discret. Appl. Math.2
2007 Characterization of P6-free graphs
Jiping Liu, Yuejian Peng, Cheng Zhao 0001
Discret. Appl. Math.3
1996 Graphs That Admit 3-to-1 or 2-to-1 Maps onto the Circle
Anthony J. W. Hilton, Jiping Liu, Cheng Zhao 0001
Discret. Appl. Math.3