Runzhou Zhang

dblp:152/4444 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Libra: A Hybrid-Sparse Attention Accelerator Featuring Multi-Level Workload Balance
abstract
Transformers have delivered exceptional performance and are widely used across various natural language processing (NLP) tasks, owing to their powerful attention mechanism. However, the high computational complexity and substantial memory usage pose significant challenges to inference efficiency. Numerous quantization and value-level sparsification methods have been proposed to overcome these challenges. Since higher sparsity leads to greater acceleration efficiency, leveraging both value-level and bit-level sparsity (hybrid sparsity) can effectively exploit the acceleration potential of the attention mechanism. However, increased sparsity exacerbates load imbalance across compute units, potentially limiting the extent of acceleration benefits. To fully exploit the acceleration potential of hybrid sparsity, we propose Libra, an attention accelerator developed through algorithm-hardware co-design. At the algorithm level, we design the bit-group-based algorithm consisting of filtered bit-group sparsification (FBS) and dynamic bit-group quantization (DBQ) to maximize the utilization of sparsity in attention. FBS imposes structured sparsity on weights, while DBQ introduces dynamic sparsification during the computation of activations. At the hardware level, we design task pool to achieve multi-level workload balance, effectively mitigating the load imbalance among compute units induced by hybrid sparsity. Additionally, different stages in DBQ can be executed in parallel, with each stage operating at distinct bit-widths. To support this, we design an adaptive bit-width architecture that enables simultaneous computations at varying bitwidths. Our experiments demonstrate that, compared to state-of-the-art (SOTA) attention accelerators, Libra achieves up to $1.49 \times \sim 5.89 \times$ speedup and $2.65 \times \sim 10.82 \times$ enhancement in energy efficiency.
Faxian Sun, Runzhou Zhang, Heng Liao, Zhinan Qin, Jianli Chen, Jun Yu 0010, Kun Wang 0005
DAC2
2025 Blaze: An Efficient Bit-Sparse Attention Architecture With Workload Orchestration Optimization
abstract
The attention mechanism is a core neural network primitive widely utilized in state-of-the-art models of Natural Language Processing (NLP) applications. However, the high computational complexity and substantial power consumption hinder its deployment and efficient inference. To address these challenges, various methods leveraging sparsity and quantization have been proposed. Compared to these methods, the exploitation of abundant bit-level sparsity in attention-based models presents great potential for the performance enhancement of attention inference. Existing bit-sparsity optimization methods primarily focus on Convolutional Neural Networks (CNNs), which are not ideally suitable for the attention mechanism, and they have not effectively solved the workload imbalance and hardware under-utilization issues caused by the irregular distribution of non-zero bits in tensor data. In this work, we introduce Blaze, an efficient attention architecture that leverages both value and bit-level sparsity in tensor data along with workload orchestration optimization. To mitigate the workload imbalance issues often encountered by sparse bit-serial architecture, we propose an Approximate-Computing-Based (ACB) workload orchestration mechanism. Additionally, to fully exploit the redundancy in the attention mechanism, we propose a Leading-Booth mechanism to further enhance the performance of attention computation. We also design a reconfigurable computing engine to support both mechanisms. Experimental results indicate that, compared to state-of-the-art (SOTA) attention accelerators, our Blaze can achieve $2.37 \times \sim 6.18 \times$ improvement in performance and $9.69 \times \sim 43.96 \times$ enhancement in energy efficiency. Our accelerator can reach up to $1.58 \times$ speedup in attention computing performance compared with the SOTA bit-sparse accelerator.
Runzhou Zhang, Faxian Sun, Kunchen Zou, Zhinan Qin, Jianli Chen, Jun Yu 0010, Kun Wang 0005
DAC1
2024 The Audience Effect: Do Observations Change Outcomes in HCI Studies?
abstract
Observational studies are widely used in Human-Computer Interaction (HCI) research to evaluate usability and user experience with technologies. However, the act of observation may influence participant behaviour and performance, threatening the validity of study findings. This paper investigates the impact of three observation types on participant outcomes in a simulated HCI study context. Participants completed Sudoku puzzles under baseline (no observation), human observation, sensor-based observation, and combined human/sensor conditions. Performance was assessed by puzzle completion rates. The mental workload was measured via NASA-TLX surveys, heart rate, galvanic skin response, and infrared thermal imaging. Results showed observations negatively impacted performance versus baseline, with human observers inducing the greatest distraction. Experienced participants were more influenced than novices. Task medium also affected engagement and observation reactivity. Findings demonstrate observations introduce bias in HCI research, emphasising careful consideration of observation methods to improve result validity.
Boon-Giin Lee, Dave Towey, Kaiyi Chen, Yichu Fang, Runzhou Zhang, Matthew Pike
COMPSAC6
2024 FLOP: A Flexible Memory-Optimized Processor for Parallel Graph Mining on FPGA
abstract
Graph mining is an important and complex emerging algorithmic model with extensive applications in fields including social sciences, chemoinformatics, and bioinformatics. However, contemporary graph mining accelerators still face challenges related to excessive on-chip resource utilization, and low set processing efficiency. To address these issues, we propose FLOP, a memory-optimized processor that leverages a new on-chip and off-chip memory partitioning design scheme. First, FLOP's memory design can accommodate the varying memory requirements of graph vertex sets. Second, we devise input-size aware processing engines (PEs) to optimize resource utilization and maximize computation efficiency. Third, FLOP adopts a pattern-aware instruction set architecture and a two-stage compiler to satisfy the mining needs of different patterns. We evaluated FLOP using five commonly used datasets and different pattern mining tasks. Experiment results show that, FLOP outperforms the state-of-the-art FPGA-based accelerator Gramer by 3.09× ~ 15.93× and compared to the CPU-based design GraphPi, FLOP achieves an average of 6.55× speedup. Additionally, FLOP also has competitive performance compared to the ASIC-based design FINGERS.
Runzhou Zhang, Jun Yu 0010, Kun Wang 0005
ICCAD2
2024 Feature extraction of trajectories for mobility modeling in 5G NB-IoT networks
Runzhou Zhang, Lei Ning, Mengkun Li, Chengcai Wang
Wirel. Networks1
2023 Edge FPGA-based Onsite Neural Network Training
abstract
Conjugate gradient (CG) is widely used in training sparse neural networks. However, CG, involving a large amount of sparse matrix and vector operations, cannot be efficiently implemented on resource-limited edge devices. In this paper, a high-performance and energy-efficient CG accelerator implemented on edge Field Programmable Gate Array is proposed for fast onsite neural networks training. According to the profiling, we propose a unified matrix multiplier that is compatible with the sparse and dense matrix. We also design a novel T-engine to handle transpose operation with the compressed sparse format. Experimental results show that our proposal outperforms the state-of-the-art FPGA work with a resource reduction of up to 41.3%. In addition, we achieve on average$10.2\times$and$2.0\times$speedup, while$10.1\times$and$3.5\times$better energy efficiency than implementations on CPU and GPU, respectively.
Ruiqi Chen 0001, Yu Li 0003, Runzhou Zhang, Jun Yu 0010, Kun Wang 0005
ISCAS4
2021 Trajectory Mining-Based City-Level Mobility Model for 5G NB-IoT Networks
abstract
Due to the large coverage of 5G NB‐IoT networks, a more realistic mobility model for a macroscopic scene will greatly facilitate the development of optimal radio resource management algorithms. However, models devised for a random motion scene are no longer applicable in circumstances. Therefore, in this paper, a city‐level mobility model is proposed based on the feature mining of the real trajectory of vehicles in the city of Shenzhen. The proposed model is separately designed in the motion trajectory to reduce the mutual influence between the time and spatial sequence. Simulation results show that it can better present specific node motions with the physical constraints of the city layout, which are motivated with a high degree of fit in terms of self‐similarity, hotspots, and long‐tail features.
Runzhou Zhang, Tongyi Zheng, Lei Ning
Wirel. Commun. Mob. Comput.1
2020 Fundamental System-Degrading Effects in THz Communications Using Multiple OAM beams With Turbulence
abstract
We explore and find the fundamental systemdegrading effects when using multiple orbital-angular-momentum (OAM) beams in a THz communications link under atmospheric turbulence in simulation. Unlike optical links with relatively small divergence effects, the crosstalk performance of THz OAM links is dependent on divergence-related parameters, including OAM mode order, frequency, and beam waist. Simulation results show: (i) for the cases with the same ratio of beam diameter to the Fried parameter (D/r0), the signal power increases and the crosstalk (XT) decreases when increasing the divergence-related parameters; and (ii) for the cases with the same atmospheric structure constant Cn2, the signal power decreases and the XT increases when increasing the divergence-related parameters. Moreover, for building a link where OAM +4 is transmitted with the parameters: (i) beam waist of 0.1 m and link distance of 200 m, and (ii) beam waist of 1 m and link distance of 1 km, the XT from neighbouring mode remains less than -15 dB when carrier wave frequency is <; 1 THz and 0.1 THz, respectively. In addition, simulation results also show that: (i) limited aperture size of the system has high influence on the XT performance under both weak and strong turbulence; and (ii) displacement of the system has high influence on the XT performance under no and weak turbulence.
Zhe Zhao 0003, Runzhou Zhang, Hao Song 0006, Kai Pang, Ahmed Almaiman, Huibin Zhou, Haoqian Song, Cong Liu 0010, Nanzhe Hu, Xinzhou Su, Amir Minoofar, Shlomo Zach, Moshe Tur, Andreas F. Molisch, Alan E. Willner
ICC2
2017 Performance of Using Antenna Arrays to Generate and Receive mm-Wave Orbital-Angular-Momentum Beams
abstract
Generation and detection of millimeter-wave carrying orbital angular momentum (OAM) have been of growing interest. In this paper, we evaluate patch antenna arrays with different arrangements as OAM generators and receivers by simulation. We compare beam evolution processes and steering performance for circular and ring-antenna-array based OAM links. Mode purity of the generated OAM +1 beam fluctuates between 10% and 99% for ring antenna arrays with 10 cm diameters while it remains >99% for circular antenna arrays. Compared to a ring-antenna-array based link, the circular-antenna-array based link could have an ~10 dB lower power loss at the distance up to 0.5m. We also show that a 5cm diameter circular antenna array could steer an OAM +1 beam up to 80° with a mode purity degradation of <;1%, while ring antenna arrays have ~10% mode purity degradation. OAM spectrum analysis shows both the lattice shape and the boundary shape of an antenna array could cause power leakage to harmonic OAM orders. Such power leakages would increase as the designed OAM order or the lattice period d increases, while it would decrease as the array diameter D increases.
Zhe Zhao 0003, Guodong Xie, Long Li 0001, Haoqian Song, Cong Liu 0010, Kai Pang, Runzhou Zhang, Changjing Bao, Zhe Wang 0020, Soji Sajuyigbe, Shilpa Talwar, Hosein Nikopour, Alan E. Willner
GLOBECOM7