Zengyuan Zhang

dblp:350/1323 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-1643-4191ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 40% Memory systems · 26% High-performance computing · 24%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing
0.922025
DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUs · IEEE Trans. Computers 2023
Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache Hierarchy · IEEE Trans. Computers 2025
Memory systems
cache
0.912025
Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache Hierarchy · IEEE Trans. Computers 2025
Memory systems › cache management
cache interference
0.912025
Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache Hierarchy · IEEE Trans. Computers 2025
Parallel and multicore computing › task scheduling
contention-aware scheduling
0.912025
Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache Hierarchy · IEEE Trans. Computers 2025
Parallel and multicore computing
task scheduling
0.912025
Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache Hierarchy · IEEE Trans. Computers 2025
GPUs and heterogeneous computing
GPU computing
0.712023
DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUs · IEEE Trans. Computers 2023
High-performance computing
stencil computation
0.712023
DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUs · IEEE Trans. Computers 2023
Parallel and multicore computing › parallel program transformation
tiling
0.712023
DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUs · IEEE Trans. Computers 2023
Parallel and multicore computing
parallel programming models
0.212023
DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUs · IEEE Trans. Computers 2023

Methods — techniques the papers use, named apart from their topics

two-level scheduling · 0.9task grouping · 0.9heuristic · 0.9static tiling · 0.7dynamic rectangular tiling · 0.7dynamic hybrid tiling · 0.7
YearPublicationVenuePosition
2025 Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache Hierarchy
abstract
For scientific computing applications that consist of many loosely coupled tasks, efficient scheduling is critical to achieve high performance and good quality of service (QoS). One of the challenges for co-running tasks is the frequent contention for shared cache hierarchy of multi-core processors. Such contention significantly increases cache miss rate and therefore, results in performance deterioration for computational tasks. This paper presents Scalpel, a contention-aware task grouping and co-scheduling approach for efficient task scheduling on shared cache hierarchy. Scalpel utilizes the shared cache access features of tasks to group them in a heuristic way, which reduces the contention within groups by achieving equal shared cache locality, while maintaining load balancing between groups. Based thereon, it proposes a two-level scheduling strategy to schedule groups to processors and assign tasks to available cores in a timely manner, while considering the impact of task scheduling on shared cache locality to minimize task execution time. Experiments show that Scalpel reduces the shared cache miss rate by up to 2.14× and optimizes the execution time by up to 1.53× for scientific computing benchmarks, compared to several baseline approaches.
Song Liu 0007, Zengyuan Zhang, Xinhe Wan, Bo Zhao 0019, Weiguo Wu
IEEE Trans. Computers3
2023 TurboStencil: You only compute once for stencil computation
Song Liu 0007, Xinhe Wan, Zengyuan Zhang, Bo Zhao 0019, Weiguo Wu
Future Gener. Comput. Syst.3
2023 DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUs
abstract
Stencil computation is an important class of computational modes in scientific computing applications. Loop tiling techniques have been widely studied to accelerate stencil computations on different architectures by exploiting parallelism and data locality. Recent advanced tiling methods enable the tile-wise concurrent start-up to improve the execution performance. However, such methods statically partition all dimensions of iteration space into tiles with predetermined complex shapes and sizes, and thus lead to low thread utilization and memory access efficiency on GPUs. In this paper, we present DHTS, a novel dynamic hybrid tiling strategy for stencil computations. DHTS employs static tiling on the outer dimensions to achieve concurrent start-up parallelism, while proposes a dynamic rectangular tiling method on the inner dimensions to improve thread utilization and memory access efficiency. By deriving tile size constraints, DHTS adaptively achieves equal-size workload of tiles, and therefore reducing idle threads and increasing coalesced memory accesses within tiles. We implement the proposed strategy with different complex tile shapes. Experimental results on Titan V and Tesla V100 GPUs show that DHTS effectively improves the execution performance of 2D/3D stencils compared to state-of-the-art tiling methods, and achieves the best improvement of 28×.
Song Liu 0007, Zengyuan Zhang, Weiguo Wu
IEEE Trans. Computers2