Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yongze Yan

dblp:160/0970 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0009-3166-0072ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
GPUs and heterogeneous computing · 71% Storage systems · 24% Parallel and multicore computing · 5%
Databases, data mining, and information retrieval
1 paper
Graph data management · 100%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU graph processing
1.722025
Towards Communication-Efficient Out-of-Core Graph Processing on the GPU · IEEE Trans. Parallel Distributed Syst. 2025
Efficient Graph Data Access for Out-of-Memory GPU Streaming Graph Processing · Proc. VLDB Endow. 2025
GPUs and heterogeneous computing
GPU memory management
0.912025
Towards Communication-Efficient Out-of-Core Graph Processing on the GPU · IEEE Trans. Parallel Distributed Syst. 2025
Storage systems › out-of-core computation
out-of-core graph processing
0.912025
Towards Communication-Efficient Out-of-Core Graph Processing on the GPU · IEEE Trans. Parallel Distributed Syst. 2025
Graph data management › dynamic graph processing
incremental graph processing
0.712023
Layph: Making Change Propagation Constraint in Incremental Graph Processing by Layering Graph · ICDE 2023
Graph algorithms and graph theory
graph processing
0.312025
Towards Communication-Efficient Out-of-Core Graph Processing on the GPU · IEEE Trans. Parallel Distributed Syst. 2025
Parallel and multicore computing › graph processing
iterative graph processing
0.212023
Layph: Making Change Propagation Constraint in Incremental Graph Processing by Layering Graph · ICDE 2023

Methods — techniques the papers use, named apart from their topics

task scheduling optimization · 1.7hybrid transfer management · 1.7skeleton graph · 1.3graph layering · 1.3incremental processing · 0.9hot subgraph caching · 0.9
YearPublicationVenuePosition
2025 Efficient Graph Data Access for Out-of-Memory GPU Streaming Graph Processing
abstract
Leveraging GPUs' high parallelism can significantly improve the real-time computation efficiency of streaming graph processing. However, when a large-scale graph exceeds GPU memory capacity, CPU-GPU cooperative processing often results in substantial and irregular CPU-to-GPU data transfer overhead. This stems from the extensive redundant graph accesses during continuous computation, which can hardly be addressed by existing solutions. In this work, we present Grapin, an out-of-memory GPU streaming graph processing system designed to minimize graph data transfer via two effective techniques for eliminating redundant accesses: (1) Extending advanced incremental processing algorithms to GPUs by converting their heavyweight data dependency processing into GPU-friendly forms, eliminating redundant graph accesses from the computation side; and (2) providing a lightweight yet efficient GPU hot subgraph management framework that finely caches the frequently accessed dynamic subgraphs in a vertex-centric manner. Experimental results demonstrate that Grapin can efficiently process large-scale streaming graphs with billions of edges on a single NVIDIA A5000 GPU. Enabling incremental computation reduces data transfer by 61%, and the integration of GPU hot subgraph reuse further reduces the remaining transfer by 72%, resulting in a total reduction of 89%. Compared with CPU-based solutions, Grapin achieves speedups ranging from 1.8x to 96.9x (17.9x on average).
Qiange Wang, Yongze Yan, Hongshi Tan, Cheng Chen 0008, Cheng Zhao 0001, Jiaming Tian, Xiaoliang Cong, Yanfeng Zhang 0001, Ge Yu 0001, Weng-Fai Wong, Bingsheng He
Proc. VLDB Endow.2
2025 Towards Communication-Efficient Out-of-Core Graph Processing on the GPU
abstract
The key performance bottleneck of large-scale graph processing on memory-limited GPUs is the host-GPU graph data transfer. Existing GPU-accelerated graph processing frameworks address this issue by managing the active subgraph transfer at runtime. Some frameworks adopt explicit transfer management approaches based on explicit memory copy with filter or compaction. In contrast, others adopt implicit transfer management approaches based on on-demand accesses with the zero-copy mechanism or unified virtual memory. Having made intensive analysis, we find that as the active vertices evolve, the performance of the two approaches varies in different workloads. Due to heavy redundant data transfers, high CPU compaction overhead, or low bandwidth utilization, adopting a single approach often results in suboptimal performance. Moreover, these methods lack effective cache management methods to address the irregular and sparse memory access pattern of graph processing. In this work, we propose a hybrid transfer management approach that takes the merits of both two transfer approaches at runtime. Moreover, we present an efficient vertex-centric graph caching framework that minimizes CPU-GPU communication by caching frequently accessed graph data at runtime. Based on these techniques, we present HytGraph, a GPU-accelerated graph processing framework, which is empowered by a set of effective task-scheduling optimizations to improve performance. Experiments on real-world and synthetic graphs show that HytGraph achieves average speedups of 2.5 ×, 5.0 ×, and 2.0 × compared to the state-of-the-art GPU-accelerated graph processing systems, Grus, Subway, and EMOGI, respectively.
Qiange Wang, Xin Ai 0006, Yongze Yan, Shufeng Gong 0001, Yanfeng Zhang 0001, Jing Chen 0037, Ge Yu 0001
IEEE Trans. Parallel Distributed Syst.3
2023 Layph: Making Change Propagation Constraint in Incremental Graph Processing by Layering Graph
abstract
Real-world graphs are constantly evolving, which demands updates of the previous analysis results to accommodate graph changes. By using the memoized previous computation state, incremental graph computation can reduce unnecessary recomputation. However, a small change may propagate over the whole graph and lead to large-scale iterative computations. To address this problem, we propose Layph, a two-layered graph framework. The upper layer is a skeleton of the graph which is much smaller than the original graph, and the lower layer has some disjoint subgraphs. Layph limits costly global iterative computations on the original graph to the small graph skeleton and a few subgraphs updated with the input graph changes. In this way, many vertices and edges are not involved in iterative computations, which significantly reduces the computation overhead and improves the performance of incremental graph processing. Our experimental results show that Layph outperforms current state-of-the-art incremental graph systems by 9.08× on average (up to 36.66×) in response time.
Song Yu 0004, Shufeng Gong 0001, Yanfeng Zhang 0001, Wenyuan Yu, Qiang Yin 0002, Chao Tian 0001, Yongze Yan, Ge Yu 0001, Jingren Zhou 0001
ICDE8