EDBT 2026 Demo / reviewers in the wild / expert
Qi Wang 0051
dblp:19/1924-51
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-2644-2873ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RV-WINO: A RISC-V Neural Network Accelerator Based on Winograd Algorithm Fabricated in 55-nm CMOS ProcessabstractThe rapid evolution of artificial intelligence (AI) in IoT applications necessitates the execution of inference tasks on edge devices. However, the deployment of computation-intensive neural networks on resource-constrained edge systems presents a significant challenge. This brief presents the RV-WINO processor, the first silicon implementation of a RISC-V processor based on the Winograd algorithm for convolution and general matrix multiplication (GEMM) acceleration. The processor incorporates a Winograd module, which significantly reduces multiplication operations during convolutions, leading to a substantial decrease in energy consumption. In addition, the processor includes a matrix multiplication module that reuses the multipliers of the Winograd module, accelerating fully connected and dot product operations in neural networks. The RV-WINO processor fabricated in a 55-nm CMOS process achieves the peak computational performance of 0.95 and 2.39 GOPS in INT32 and INT8 modes, with its peak energy efficiency reaching 112 and 237 GOPS/W. In convolutional neural network (CNN) inference tasks, the execution time is reduced by over 80% compared with the baseline processor. Yucong Huang, Qu Lu, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | NNia-8: An 8-Core RISC-V Neural Network Inference Accelerator with Efficient Processing Elements and Memory Utilization
Yucong Huang, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye |
NPC (2) | 5 |
| 2025 | RV-SCNN: A RISC-V Processor With Customized Instruction Set for SNN and CNN Inference Acceleration on Edge PlatformsabstractThe rapid advancement of artificial intelligence (AI) applications has driven an increasing demand for conducting inference tasks on edge devices. However, implementing computation-intensive neural networks on resource-constrained edge systems remains a significant challenge. In this article, we propose a novel processor architecture called RV-SCNN to address this challenge. The architecture is based on the RISC-V generic instruction set and incorporates various single instruction multiple data (SIMD) custom instruction extensions to accelerate the computation of spike neural networks (SNNs) and convolutional neural networks (CNNs), enabling efficient execution of complex neural network models. The core operators of the processor are shared by both SNN and CNN operations, thus supporting both computation modes. Other acceleration implementations include an internal hardware loop control unit that reduces the instruction overhead, an address calculation unit and an interlayer fusion unit that minimize the memory access overhead, as well as an image to column (IM2COL) unit that improves the computational efficiency of the$3 \times 3$convolutions in SNNs and CNNs. The custom instructions are called through inline assembly in the C program, providing higher flexibility compared to traditional ASICs and supporting custom complex SNN/CNN network structures. Compared to traditional instruction sets, the RV-SCNN processor reduces the execution time of CNNs and SNNs by over 90%. We validate the processor on FPGA platform and evaluate its performance under CMOS 55-nm process. The processor achieves an operational efficiency of 9.88 pJ/SOP in SNN network inference tasks, while the peak energy efficiency reaches 679 GOPS/W in CNN network inference. Chenxi Feng, Xinyu Kang, Qi Wang 0051, Yucong Huang, Terry Tao Ye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | RV-GEMM: Neural Network Inference Acceleration with Near-Memory GEMM Instructions on RISC-VabstractGeneral Matrix Multiply (GEMM), as a fundamental operation in neural network, plays an important role in artificial intelligence and signal processing applications. In this paper, we proposed three SMID RISC-V custom instructions to accelerate GEMM computations, supporting multiple precisions including 32-bit, 16-bit and 8-bit fixed. Furthermore, we implemented address calculation and loop control units along with the GEMM acceleration module to reduce the memory access overhead. These three GEMM custom instructions, along with the near-memory optimization units, were incorporated in the RV-GEMM processor and implemented on the FPGA platform for speedup evaluation. It was also compiled in Synopsys Design Compiler with CMOS 55nm process for hardware overhead estimation. Compared to the baseline RISC-V processor, for GEMM computations under precisions of 32-bit, 16-bit and 8-bit fixed, the RV-GEMM processor achieved speedup ratios of 15.8×, 28.7× and 42.5×. The peak energy efficiency also reached 260 GOPS/W, 420 GOPS/W and 609 GOPS/W, respectively. Chenxi Feng, Bingzhen Chen, Qi Wang 0051, Yucong Huang, Terry Tao Ye |
CF | 4 |
| 2024 | Optimizing CNN Computation Using RISC-V Custom Instruction Sets for Edge PlatformsabstractBenefit from the custom instruction extension capabilities, RISC-V architecture can be optimized for many domain-specific applications. In this paper, we propose seven RISC-V SIMD (single instruction multiple data) custom instructions that can significantly optimize the convolution, activation and pool operations in CNN inference computation. More specifically, instruction CONV23 can greatly speed up the operation ofF(2 × 2, 3 × 3). With the adoption of Winograd algorithm, the number of multiplications can be reduced from 36 to 16, and the execution time is also reduced from 140 to 21 clock cycles. These custom instructions can be executed in batch mode within the acceleration module where the immediate data can be reused, so the latency and energy overhead associated with excess memory accesses can be eliminated. Using inline assembler in C language, the custom instructions can be called and compiled together with C source code. A revised RISC-V processor, RI5CY-Accel is constructed on FPGA to accommodate these custom instructions. Revised LeNet-5, VGG16 and ResNet18 model; called LeNet-Accel, VGG16-Accel and ResNet18-Accel are also optimized based on RI5CY-Accel architecture. Benchmark experiments demonstrated that the inference of LeNet-Accel, VGG16-Accel and ResNet18-Accel based on RI5CY-Accel can greatly reduce the execution latency by over 76.6%, 88.8% and 87.1%, with the total energy consumption saving of 74.8%, 87.8% and 85.1% respectively. Bingzhen Chen, Chenxi Feng, Qi Wang 0051, Terry Tao Ye |
IEEE Trans. Computers | 6 |
| 2022 | Multiplication Through a Single Look-Up-Table (LUT) in CNN Inference ComputationabstractParameter quantization with lower bit width is the common approach to reduce the computation loads in CNN inference. With the parameters being replaced by fixed-width binaries, multiplication operations can be replaced by the look-up-table (LUT), where the multiplier-multiplicand operands serve as the table index, and the precalculated products serve as table elements. Because the histogram profiles of the parameters in different layers/channels differ significantly in CNN, previous LUT-based computation methods have to use different LUTs for each layer/channel, and consequently demand larger memory space along with extra access time and power consumption. In this work, we first normalize the parameters Gaussian profiles of different layers/channels to have similar means and variances, and further quantize the normalized parameters into fixed width through nonlinear quantization. Because of the normalized parameters profile, we can use one single compact LUT ($16\times 16$entries) to replace all multiplication operations in the whole network. Furthermore, the normalization procedure also reduces the errors induced from quantization. Experiments demonstrate that with a compact 256-entry LUT, we can achieve the accuracy comparable to the results from 32-bit floating-point calculation; while significantly reducing the computation loads and memory spaces, along with power consumption and hardware resources. Qi Wang 0051, Terry Tao Ye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Customized Instruction on RISC-V for Winograd-Based Convolution AccelerationabstractConvolution operation accounts for the major work-load in convolutional neural networks (CNN). However, standard instruction set for RISC-V processor cannot efficiently perform the matrix convolution between kernel and input matrices. In this paper, we construct a custom instruction under the RISC-V ISA that can perform the F(2×2,3×3) convolution within one single execution. Particularly, optimized by the Winograd algorithm, the operation only needs 16 multiplications instead of 36 multiplications as needed by standard ISA. Benefit from this cycles, as compared to 140 cycles using standard instructions. Thenew instruction, F(2×2,3×3) can be calculated within 19 clock power consumed during convolution operation is also reduced significantly. Jianghan Zhu, Qi Wang 0051, Can He, Terry Tao Ye |
ASAP | 3 |