Peng Qu 0001

dblp:20/3008-1 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-1786-5372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Look Before You Leap : Precision Instruction Supply via SmartScout
abstract
Modern high-performance processors extensively employ Fetch-Directed Instruction Prefetching (FDIP) to mitigate instruction supply bottlenecks. However, the efficacy of FDIP is fundamentally constrained by the accuracy of the Branch Prediction Unit (BPU). As the critical component within the BPU, the Branch Target Buffer (BTB) faces severe capacity bottlenecks. While pre-decoding-based prefetching offers a remedy, existing approaches suffer from two critical impediments: (1) The Noise Dilemma: Suboptimal trade-off between coverage and accuracy. (2) Inefficient Miss Resolution: Current designs rely on reactive recovery or stalls, failing to leverage available front-end slack for proactive correction.
Peng Qu 0001, Tingji Zhang, Fang Su, Zhe Pan 0001, Youhui Zhang
ICS2
2026 VDHA: Vector-Driven Hash Aggregation for Sparse Matrix-Sparse Vector Multiplication on GPUs
abstract
Sparse matrix-sparse vector multiplication (SpMSpV) is a core primitive in graph analytics and scientific computing, also arising in spiking neural networks for event-driven spike propagation. On GPUs, the performance of the prevalent and efficient SpMSpV paradigm is often bottlenecked by the write-back phase of accumulating non-zero multiply–accumulate results; its many-to-one index scatter pattern causes severe conflicts and poor bandwidth utilization on GPUs. We present VDHA, a GPU-based weighted SpMSpV kernel that leverages block-private hash tables for local aggregation, substantially reducing write conflicts and improving memory coalescing. To further amplify this benefit, we incorporate column splitting with lightweight reordering to expose more locality, and employ a fetch–compute–writeback pipeline to overlap hash computation with memory accesses. Extensive evaluation on over 300 matrices with more than 5 million nonzeros, including web-scale graphs (Konect/LAW) and scientific workloads (SuiteSparse), shows that VDHA consistently outperforms state-of-the-art baselines. On web graphs, it achieves a 1.41× geometric-mean speedup (up to 3.42×), while on SuiteSparse it delivers 1.13× (up to 2.55×). We also provide a lightweight predictive model that identifies matrices favorable to VDHA with 91.3% accuracy.
Zhe Pan 0001, Peng Qu 0001, Youhui Zhang
PPoPP3
2026 Root-Down Exposure for Maximal Clique Enumeration on GPUs
abstract
Maximal clique enumeration (MCE) in large-scale graphs is critical across various application domains, including social network analysis, bioinformatics, and computer vision. However, existing GPU-based MCE solutions suffer from inefficient load-balancing mechanisms. These mechanisms force busy workers to pause and hand over workloads to idle workers, which introduces significant synchronization overhead and increases memory usage. To address these limitations, we introduce a root-down exposure mechanism where busy workers dynamically expose their current root, enabling idle workers to pull workloads from the exposed node directly without synchronization. We then propose a bitmap-centric MCE scheme and an aggressive node generation rule to further simplify the memory layout and accelerate enumeration. We combine them into RDMCE, a Root-Down MCE solution on GPUs. Across large real-world graphs with up to 146 billion maximal cliques, RDMCE is 1.25-5.38× faster than any next-best state-of-the-art GPU-based solutions and is the only one that completes enumeration on every test dataset, offering a more efficient and scalable MCE solution.
Zhe Pan 0001, Peng Qu 0001, Youhui Zhang
PPoPP2
2025 Hierarchical Prefetching: A Software-Hardware Instruction Prefetcher for Server Applications
abstract
The large working set of instructions in server-side applications causes a significant bottleneck in the front-end, even for high-performance processors equipped with fetch-directed instruction prefetching (FDIP). Prefetchers specifically designed for server scenarios typically rely on a record-and-replay mechanism that exploits the repetitiveness of instruction sequences. However, the efficacy of these techniques is compromised by discrepancies between actual and predicted control flows, resulting in loss of coverage and timeliness. This paper proposes Hierarchical Prefetching, a novel approach that tackles the limitations of existing prefetchers. It identifies common coarse-grained functionality blocks (called Bundles) within the server code and prefetches them as a whole. Bundles are significantly larger than typical prefetch targets, encompassing tens to hundreds of kilobytes of code. The approach combines simple software analysis of code for bundle formation and light-weight hardware for record-and-replay prefetching. The prefetcher requires under 2KB of on-chip storage by keeping most of the metadata in main memory. Experiments with 11 popular server workloads reveal that Hierarchical Prefetching significantly improves miss coverage and timeliness over prior techniques, achieving a 6.6% average performance gain over FDIP.
Tingji Zhang, Boris Grot, Wenjian He, Yashuai Lv, Peng Qu 0001, Fang Su, Guowei Zhang 0002, Youhui Zhang
ASPLOS (2)5
2025 SoftGuide: A Hardware-Software Co-design Predictor for Data-Dependent Branches
abstract
Branch predictors based on historical information exhibit exemplary performance. However, a subset of data-dependent branches still pose a significant challenge, often resulting in severe mispredictions. Such branches are commonly encountered during the processing of various data structures, and increasing the capacity of predictors has shown limited effective-ness. Thus, a dedicated branch predictor is necessary to enhance performance for varying data structures while maintaining lower hardware complexity. This paper proposes SoftGuide, a novel hardware-software cooperative branch predictor designed to tackle two main chal-lenges inherent in data-dependent branch prediction: (1) the detection and identification of branch dependencies(software-friendly) and (2) data prefetching and dependency chain pre-execution(hardware-friendly). SoftGuide leverages software to convey the memory access patterns and the dependency chains associated with the branch, thereby circumventing the overhead of hardware-based detection. Utilizing the information provided by the software, the enhanced hardware prefetches data and triggers pre-execution in advance. SoftGuide can perfectly unify prefetching and prediction tasks. For SPEC2006 and GAP benchmarks with the method of SimPoint, SoftGuide realizes a decrease in branch mispredictions per 1K instructions (MPKI) by 46.4% and an increase in Instructions Per Cycle (IPC) by 1.25x average on the processor equipped with the state-of-the-art branch predictor. Moreover, the storage overhead is just 3.88KB.
Peng Qu 0001, Tingji Zhang, Youhui Zhang
CCGrid2
2024 A Row Decomposition-based Approach for Sparse Matrix Multiplication on GPUs
abstract
Sparse-Matrix Dense-Matrix Multiplication (SpMM) and Sampled Dense Dense Matrix Multiplication (SDDMM) are important sparse kernels in various computation domains. The uneven distribution of nonzeros in the sparse matrix and the tight data dependence between sparse and dense matrices make it a challenge to run sparse matrix multiplication efficiently on GPUs. By analyzing the aforementioned problems, we propose a row decomposition (RoDe)-based approach to optimize the two kernels on GPUs, using the standard Compressed Sparse Row (CSR) format. Specifically, RoDe divides the sparse matrix rows into regular parts and residual parts, to fully optimize their computations separately. We also devise the corresponding load balancing and finegrained pipelining technologies. Profiling results show that RoDe can achieve more efficient memory access and reduce warp stall cycles significantly. Compared to the state-of-the-art (SOTA) alternatives, RoDe achieves a speedup of up to 7.86× with a geometric mean of 1.45× for SpMM, and a speedup of up to 8.99× with a geometric mean of 1.49× for SDDMM; the dataset is SuiteSparse. RoDe also outperforms its counterpart in the deep learning dataset. Furthermore, its preprocessing overhead is significantly smaller, averaging only 16% of the SOTA.
Peng Qu 0001, Youhui Zhang, Zhaolin Li
PPoPP3
2023 ENLARGE: An Efficient SNN Simulation Framework on GPU Clusters
abstract
Spiking Neural Networks (SNNs) are currently the most widely used computing model for neuroscience communities. There is also an increasing research interest in exploring the potential of SNN in brain-inspired computing, artificial intelligence, and other areas. As SNNs possess distinguished characteristics that originate from biological authenticity, they require dedicated simulation frameworks to achieve usability and efficiency. However, there is no widely-used, easily accessible, high performance SNN simulation framework for GPU clusters. In this paper, we propose ENLARGE, an efficient SNN simulation framework on GPU clusters. ENLARGE provides a multi-level architecture that deals with computation, communication, and synchronization hierarchically. We also propose an efficient communication method with an all-to-all communication pattern. To deal with the delay of spike delivery, which is the most distinguished SNN characteristic, several delay-aware optimization methods are also proposed. We further propose a multilevel workload management method. Various experiments are carried out to demonstrate the performance and scalability of the framework, as well as the effects of the optimization methods. Test results show that ENLARGE can achieve$3.17\times \sim 28.12\times$speedup compared with the most widely used NEST simulator and$3.26\times \sim 13.57\times$speedup compared with the widely used NEST GPU simulator for GPU clusters.
Peng Qu 0001, Youhui Zhang
IEEE Trans. Parallel Distributed Syst.1
2022 A review of basic software for brain-inspired computing
Peng Qu 0001, Youhui Zhang
CCF Trans. High Perform. Comput.1
2020 High Performance Simulation of Spiking Neural Network on GPGPUs
abstract
Spiking neural network (SNN) is the most commonly used computational model for neuroscience and neuromorphic computing communities. It provides more biological reality and possesses the potential to achieve high computational power and energy efficiency. Because existing SNN simulation frameworks on general-purpose graphics processing units (GPGPUs) do not fully consider the biological oriented properties of SNNs, like spike-driven, activity sparsity, etc., they suffer from insufficient parallelism exploration, irregular memory access, and load imbalance. In this article, we propose specific optimization methods to speed up the SNN simulation on GPGPU. First, we propose a fine-grained network representation as a flexible and compact intermediate representation (IR) for SNNs. Second, we propose the cross-population/-projection parallelism exploration to make full use of GPGPU resources. Third, sparsity aware load balance is proposed to deal with the activity sparsity. Finally, we further provide dedicated optimization to support multiple GPGPUs. Accordingly, BSim, a code generation framework for high-performance simulation of SNN on GPGPUs is also proposed. Tests show that, compared to a state-of-the-art GPU-based SNN simulator GeNN, BSim achieves 1.41× - 9.33× speedup for SNNs with different configurations; it outperforms other simulators much more.
Peng Qu 0001, Youhui Zhang
IEEE Trans. Parallel Distributed Syst.1