Runzhen Xue

dblp:316/6303 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-1956-1284ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FZKP: Alleviating Dataflow Complexity to Exploit Fine-Grained Parallelism for ZKP Acceleration
abstract
Zero-knowledge proof (ZKP) is a promising cryptographic protocol, but its practical deployment is hindered by the time-consuming proof generation. The proof generation inherently exhibits high-degree parallelism, yet challenges persist in exploiting fine-grained parallelism due to the dataflow complexity, impeding previous work to achieve optimal acceleration. In this work, we propose FZKP, a ZKP accelerator that utilizes two novel fine-grained dataflows coupled with two forward-flow microarchitectures to alleviate dataflow complexity, efficiently exploiting fine-grained parallelism. The proposed dataflows simplify the dataflow pattern for parallel execution, disclosing fine-grained parallelism at a low cost. The microarchitectures employ a base design to handle large bit-width intermediate results for timely consumption. They then replicate and combine the base design following the proposed dataflow to facilitate parallel execution. When evaluated in 12 nm, FZKP achieves an average speedup of 10.3× and 2.2× over the state-of-the-art GPU-based solution and ZKP accelerator on real-world workloads, respectively.
Ziheng Xiao, Mingyu Yan, Mingyu Gao 0001, Runzhen Xue, Xiaochun Ye, Dongrui Fan
IEEE Trans. Parallel Distributed Syst.4
2025 MetaDSE: A Few-shot Meta-learning Framework for Cross-workload CPU Design Space Exploration
abstract
Cross-workload design space exploration (DSE) is crucial in CPU architecture design. Existing DSE methods typically employ the transfer learning technique to leverage knowledge from source workloads, aiming to minimize the requirement of target workload simulation. However, these methods struggle with overfitting, data ambiguity, and workload dissimilarity. To address these challenges, we reframe the cross-workload CPU DSE task as a few-shot meta-learning problem and further introduce MetaDSE. By leveraging model agnostic meta-learning, MetaDSE swiftly adapts to new target workloads, greatly enhancing the efficiency of cross-workload CPU DSE. Additionally, MetaDSE introduces a novel knowledge transfer method called the workload-adaptive architectural mask algorithm, which uncovers the inherent properties of the architecture. Experiments on SPEC CPU 2017 demonstrate that MetaDSE significantly reduces prediction error by 44.3% compared to the state-of-theart. MetaDSE is open-sourced and available at this anonymous GitHub.
Runzhen Xue, Hao Wu 0070, Mingyu Yan, Ziheng Xiao, Xiaochun Ye, Dongrui Fan
DAC1
2025 LiGNN: Accelerating GNN Training Through Locality-Aware Dropout
abstract
Graph Neural Networks (GNNs) have demonstrated significant success in graph learning and are widely adopted across various critical domains. However, the irregular connectivity between vertices leads to inefficient neighbor aggregation, resulting in substantial irregular and coarse-grained DRAM accesses. This lack of data locality presents significant challenges for execution platforms, ultimately degrading performance. While previous accelerator designs have leveraged on-chip memory and data access scheduling strategies to address this issue, they still inevitably access features at irregular addresses from DRAM. In this work, we propose LiGNN, a hardware-based solution that enhances locality and applies dropout to aggregation to accelerate GNN training. Unlike algorithmic dropout approaches that primarily focus on improving accuracy and neglects hardware costs, LiGNN is specifically designed to drop nodes' features with data locality awareness, directly targeting the reduction of irregular DRAM accesses, meanwhile maintaining accuracy. LiGNN introduces locality-aware ordering and a DRAM row integrity policy, enabling configurable burst and row-granularity dropout at the DRAM level. This approach improves data locality and ensures more efficient DRAM access. Compared to state-of-the-art methods, under classic 0.5 droprate, LiGNN achieves a 1.62~2.2× speedup, reduces DRAM accesses by 44~50% and DRAM row activation by 41~82%, all without losing accuracy.
Gongjian Sun, Mingyu Yan, Dengke Han, Runzhen Xue, Xiaochun Ye, Dongrui Fan
DATE4
2025 DropNaE: Alleviating irregularity for large-scale graph representation learning
Xin Liu 0073, Xunbin Xiong, Mingyu Yan, Runzhen Xue, Shirui Pan, Songwen Pei, Lei Deng 0003, Xiaochun Ye, Dongrui Fan
Neural Networks4
2025 SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN Acceleration
abstract
Heterogeneous graph neural networks (HGNNs) have expanded graph representation learning to heterogeneous graph fields. Recent studies have demonstrated their superior performance across various applications, including circuit representation, chip design automation, and placement optimization, often surpassing existing methods. However, GPUs often experience inefficiencies when executing HGNNs due to their unique and complex execution patterns. Compared to traditional graph neural networks (GNNs), these patterns further exacerbate irregularities in memory access. To tackle these challenges, recent studies have focused on developing domain-specific accelerators for HGNNs. Nonetheless, most of these efforts have concentrated on optimizing the datapath or scheduling data accesses, while largely overlooking the potential benefits that could be gained from leveraging the inherent properties of the semantic graph, such as its topology, layout, and generation. In this work, we focus on leveraging the properties of semantic graphs to enhance HGNN performance. First, we analyze the semantic graph build (SGB) stage and identify significant opportunities for data reuse during semantic graph generation. Next, we uncover the phenomenon of buffer thrashing during the graph feature processing (GFP) stage, revealing potential optimization opportunities in semantic graph layout. Furthermore, we propose a lightweight hardware accelerator frontend for HGNNs, called SiHGNN. This accelerator frontend incorporates a tree-based SGB for efficient semantic graph generation and features a novel Graph Restructurer for optimizing semantic graph layouts. Experimental results show that SiHGNN enables the state-of-the-art HGNN accelerator to achieve an average performance improvement of$2.95\times $.
Runzhen Xue, Mingyu Yan, Dengke Han, Ziheng Xiao, Xiaochun Ye, Dongrui Fan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
abstract
Heterogeneous Graph Neural Networks (HGNNs) have broadened the applicability of graph representation learning to heterogeneous graphs. However, the irregular memory access pattern of HGNNs leads to the buffer thrashing issue in HGNN accelerators.
Runzhen Xue, Mingyu Yan, Dengke Han, Yihan Teng, Xiaochun Ye, Dongrui Fan
DAC1
2024 ADE-HGNN: Accelerating HGNNs Through Attention Disparity Exploitation
Dengke Han, Meng Wu 0006, Runzhen Xue, Mingyu Yan, Xiaochun Ye, Dongrui Fan
Euro-Par (2)3
2024 HiHGNN: Accelerating HGNNs Through Parallelism and Data Reusability Exploitation
abstract
Heterogeneous graph neural networks (HGNNs) have emerged as powerful algorithms for processing heterogeneous graphs (HetGs), widely used in many critical fields. To capture both structural and semantic information in HetGs, HGNNs first aggregate the neighboring feature vectors for each vertex in each semantic graph and then fuse the aggregated results across all semantic graphs for each vertex. Unfortunately, existing graph neural network accelerators are ill-suited to accelerate HGNNs. This is because they fail to efficiently tackle the specific execution patterns and exploit the high-degree parallelism as well as data reusability inside and across the processing of semantic graphs in HGNNs. In this work, we first quantitatively characterize a set of representative HGNN models on GPU to disclose the execution bound of each stage, inter-semantic-graph parallelism, and inter-semantic-graph data reusability in HGNNs. Guided by our findings, we propose a high-performance HGNN accelerator, HiHGNN, to alleviate the execution bound and exploit the newfound parallelism and data reusability in HGNNs. Specifically, we first propose a bound-aware stage-fusion methodology that tailors to HGNN acceleration, to fuse and pipeline the execution stages being aware of their execution bounds. Second, we design an independency-aware parallel execution design to exploit the inter-semantic-graph parallelism. Finally, we present a similarity-aware execution scheduling to exploit the inter-semantic-graph data reusability. Compared to the state-of-the-art software framework running on NVIDIA GPU T4 and GPU A100, HiHGNN respectively achieves an average 40.0× and 8.3× speedup as well as 99.59% and 99.74% energy reduction with quintile the memory bandwidth of GPU A100.
Runzhen Xue, Dengke Han, Mingyu Yan, Mo Zou, Xiaocheng Yang, John Kim 0001, Xiaochun Ye, Dongrui Fan
IEEE Trans. Parallel Distributed Syst.1
2022 A Practical Highly Paralleled ReRAM-Based DNN Accelerator by Reusing Weight Pattern Repetitions
abstract
Resistive random access memory (ReRAM)-based processing-in-memory (PIM) architecture has been designed to accelerate deep neural networks (DNNs) by concurring computation and memory barriers. To further improve memory and computation efficiency, the weight sparsity characteristic has been explored to optimize the ReRAM-based DNN accelerators. However, these designs only focus on compressing zero weights to eliminate ineffectual computation. In this article, we thoroughly analyze the weight distribution characteristics of several typical DNN models and observe many nonzero weight pattern repetitions (WPRs). Therefore, there is an opportunity to further improve the performance and energy efficiency by reusing these WPR. We propose a novel ReRAM-based accelerator—PattPIM, to achieve space compression and computation reuse by exploring DNN WPR based on practical ReRAM crossbars. In PattPIM, we propose a configurable WPR-aware DNN engine and a WPR-to-OU mapping scheme to save both space and computation resources. An intraprocessing engine (PE) pipeline is designed to improve the parallelism of the computation process. Furthermore, we adopt an approximate weight pattern transform algorithm to improve the DNN WPR ratio to enhance the reuse efficiency with negligible accuracy loss. Our evaluation with 6 DNN models shows that the proposed PattPIM delivers significant performance improvement, ReRAM resource efficiency and energy saving.
Yuhao Zhang 0006, Zhiping Jia, Hongchao Du, Runzhen Xue, Zhaoyan Shen, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4