Qizhe Wu

dblp:307/9939 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-4977-5363ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Efficient Top-k Sorter through Runtime Lane Selection of Priority Queues on a Directed Ring
Huawen Liang, Wei Yuan 0006, Qizhe Wu, Xi Jin 0002
ISCAS3
2025 Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
abstract
General matrix-matrix multiplication (GEMM), serving as a cornerstone of AI computations, has positioned tensor processing engines (TPEs) as increasingly critical components within existing GPUs and domain-specific architectures (DSA). Our analysis identifies that the prevailing architectures primarily focus on dataflow or operand reuse strategies, when considering the combination of matrix multiplication with multiply-accumulator (MAC) itself, it provides greater optimization space for the design of TPEs. This work introduces a novel perspective on matrix multiplication from a hardware standpoint, focusing on the bit-weight dimension of MACs. Through this lens, we propose a finer-grained TPE notation, using matrix triple loops as an example, introducing new methods and ideas for designing and optimizing PE microarchitecture. Based on the new notation and transformations, we propose four optimization techniques that achieve varying degrees of improvement in timing, area, and power consumption. We implement our design in RTL using the SMIC-28nm process. Applying our methods to four classic TPE architectures (include systolic array [20], 3D-Cube [27], multiplier-adder tree [48], and 2D-Matrix [30]), we achieved area efficiency improvements of $1.27 \times, 1.28 \times, 1.56 \times$, and $1.44 \times$, and $1.04 \times, 1.56 \times, 1.49 \times$, and $1.20 \times$ for energy efficiency respectively. When applied to a bit-slice architecture, we achieved a $12.10 \times$ improvement in energy efficiency and $2.85 \times$ in area efficiency compared to Laconic [38]. Our Verilog HDL code, along with timing, area, and power reports for circuit synthesis in URL: https://github.com/wqzustc/High-Performance-Tensor-Processing-Engines.
Qizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 0002, Zerong He, Linfeng Tao, Letian Zhao, Zhaoxi Zeng, Wei Yuan 0006, Xi Jin 0002
HPCA1
2025 SageSC: Accelerating GraphSAGE Minibatch Inference on Memory-Intensive Graphs
abstract
Graph neural networks demonstrate excellent performance on node classification tasks in graph datasets. For inference tasks on memory-intensive graphs, the storage burden, memory access bottlenecks, and load imbalance issues arise. The minibatch inference proposed in GraphSAGE is an effective method for minimizing these problems. However, minibatch inference introduces new challenges: while it facilitates subsequent computation, the irregular random memory access pressure shifts to the minibatch construction phase, creating performance bottlenecks in the system. In this work, to address the aforementioned challenges, we propose a novel scattered minibatch construction and aggregation (SMCA) algorithm to optimize sampling, batch construction, and aggregation computations for minimizing their latency. This method distributes memoryintensive workloads and exploits the parallelism between memory groups. Evaluation results show that the proposed accelerator SageSC achieves speedups ranging from 3x to 96x compared to CPU/GPU baselines, especially on memory-intensive graphs, while outperforming existing state-of-the-art designs.
Yuchen Gui, Wei Yuan 0006, Qizhe Wu, Huawen Liang, Letian Zhao, Linfeng Tao, Zhongguang Xu, Xi Jin 0002
ICCD3
2025 MHE-TPE: Multi-Operand High-Radix Encoder for Mixed-Precision Fixed-Point Tensor Processing Engines
Qizhe Wu, Jinyi Zhou, Zhanhe Hu, Zhichen Zeng 0002, Huawen Liang, Jiuru Zhu, Linfeng Tao, Xin Zhang 0176, Zekang Cheng, Letian Zhao, Wei Yuan 0006, Xi Jin 0002
MICRO1
2024 A FPGA-HBM-Based Hardware Streaming Accelerator for GNN Sampling
abstract
Sampling takes a long time during GNN training and inference, especially in large-scale graph datasets, so accelerating the sampling process is of great value. Present GNN samplers are mainly focused on the CPU or GPU side, with fewer sampling accelerator implementations on hardware platforms such as FPGAs, and most of the existing accelerators focus on the entire GNN computation process rather than on sampling. Sampling GNNs on the FPGA side is difficult because of the intensive random memory access. In this work, we propose a FPGA-HBM-based streaming sampler working in the time domain to accelerate traditional FPGA sampling, which realizes GNN sampling while reading data in bus bursts and can satisfy both playback and non-playback sampling modes. In addition, the fast loading of node feature vectors is also realized by combining the features of HBM. The proposed hardware achieves a speedup of$2\times$to$20\times$relative to traditional FPGA-based node index sampling and$5\times$to$18\times$relative to feature vector loading at the CPU side.
Yuchen Gui, Qizhe Wu, Wei Yuan 0006, Huawen Liang, Xi Jin 0002
ASAP2
2024 RingTK: A Ring, Parallel and High Performance Top-K Sorter on FPGA
abstract
Getting the K largest/smallest elements from$N$inputs is one of the essential operations in many applications. In this article, we propose RingTK, a ring, parallel, and high performance Top-K sorter implemented on FPGA. We use a priority queue as the basic processing unit and design a Top-K sorter based on a ring topology with a global maximum module(GMM) to obtain good scalability and high performance. Based on the ring topology, we design Ring Multiplexers (RMUX) and modify the GMM to enable RingTK to efficiently handle different K-sizes and concurrent tasks. We design a encoder to flexibly deal with different data formats and max/min Top-K tasks. Finally, we implement the proposed architecture on the Xilinx XCVU37P FPGA. The results show that the proposed architecture has good scalability, and the throughput and parallelism are approximately linear. We can achieve a throughput of 35.38GB/s with 32 PQs that exceed existing literature with K=8160.
Huawen Liang, Qizhe Wu, Wei Yuan 0006, Teng Tian, Xi Jin 0002
FCCM2
2024 Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
abstract
Graph Convolutional Networks (GCNs) are state-of-the-art deep learning models for representation learning on graphs. However, the efficient training of GCNs is hampered by constraints in memory capacity and bandwidth, compounded by the irregular data flow that results in communication bottlenecks. To address these challenges, we propose a message-passing architecture that leverages NUMA-based memory access properties and employs a parallel multicast routing algorithm based on a 4-D hypercube network within the accelerator for efficient message passing in graphs. Additionally, we have re-engineered the backpropagation algorithm specific to GCNs within our proposed accelerator. This redesign strategically mitigates the memory demands prevalent during the training phase and diminishes the computational overhead associated with the transposition of extensive matrices. Compared to the state-of-the-art HP-GNN architecture we achieved a performance improvement of 1.03×~1.81×.
Qizhe Wu, Letian Zhao, Yuchen Gui, Huawen Liang, Xi Jin 0002
FPGA1
2024 EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
abstract
Tensor computations, with matrix multiplication being the primary operation, serve as the fundamental basis for data analysis, physics, machine learning, and deep learning. As the scale and complexity of data continue to grow rapidly, the demand for tensor computations has also increased significantly. To meet this demand, several research institutions have started developing dedicated hardware for tensor computations. To further improve the computational performance of tensor process units, we have reexamined the issue of computation reuse that was previously overlooked in existing architectures. As a result, we propose a novel EN-T architecture that can reduce chip area and power consumption. Furthermore, our method is compatible with existing tensor processing units. We evaluated our method on prevalent microarchitectures, the results demonstrate an average improvement in area efficiency of 8.7 %, 12.2 %, and 11.0 % for tensor computing units at computational scales of 256 GOPS, 1 TOPS, and 4 TOPS, respectively. Similarly, there were energy efficiency enhancements of 13.0 %, 17.5 %, and 15.5 %.
Qizhe Wu, Yuchen Gui, Zhichen Zeng 0002, Huawen Liang, Xi Jin 0002
ICCD1
2022 FP-GNN: Adaptive FPGA accelerator for Graph Neural Networks
Teng Tian, Letian Zhao, Qizhe Wu, Wei Yuan 0006, Xi Jin 0002
Future Gener. Comput. Syst.4
2022 QEGCN: An FPGA-based accelerator for quantized GCNs with edge-level parallelism
Wei Yuan 0006, Teng Tian, Qizhe Wu, Xi Jin 0002
J. Syst. Archit.3