Chuanning Wang

dblp:345/2732 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2027 STREAM: Spatial-tuned refinement with enhanced adaptive masking for few-shot surface defect segmentation
Chuanning Wang
Expert Syst. Appl.2
2026 Top-V: A Flexible and Programmable Top-K Acceleration Framework Based on the RISC-V ISA
Qixiang Chen, Chuanning Wang
ISCAS2
2026 Feature disentanglement and adaptive patch routing: A unified framework for zero-shot anomaly detection with vision foundation models
Chuanning Wang, Yanzhao Zhou, Zhenxiong Gu, Lining Sun
Neurocomputing1
2026 Prototypical rectification with interpretable fusion for enhanced industrial defect segmentation
Chuanning Wang, Yanzhao Zhou, Zhenxiong Gu, Mei Lin
J. Supercomput.1
2025 A RISC-V Domain-Specific Processor for Deep Learning-Based Channel Estimation
abstract
Channel estimation (CE) is a critical component in the massive multi-input multi-output (MIMO) communication systems. Compared with conventional CE algorithms, deep learning (DL)-based approach becomes a promising alternative, due to its capability of offering enhanced performance and robustness across diverse scenarios. However, efficient DL-based CE algorithms have two key properties that make them challenging for implementation in existing architectures at the edge side: the diversity of deep neural networks (DNNs) and CE strategies, and the involvements of multiple computation-intensive tasks that compass conventional signal processing, artificial intelligence (AI) inference, and online learning. To address these challenges, a domain-specific processor based on an extended RISC-V instruction set architecture (ISA) is proposed to perform these DL-based CE algorithms. First, a dedicated RISC-V ISA extension is developed to support all essential operations required by a DL-based CE algorithm, such as matrix inversion, in a flexible manner. Building on the customized ISA extension, a highly adaptable and scalable RISC-V processor is developed, featuring scalar and vector posit arithmetic units to alleviate high computational and memory demands of DNNs during both inference and training phase. Additionally, a coarse-grained matrix accelerator is integrated to expedite various matrix operations ensuring high throughput. In this way, both high flexibility and computational efficiency are achieved. Finally, our processor is implemented on a TSMC 28-nm technology. Implementation results show that the processor achieves a speedup of$5.16\sim 6.80\times $for all matrix operations compared with the state-of-the-art work. Moreover, the proposed processor provides an area efficiency improvement of$1.61\times $and an energy efficiency enhancement of$6.6\sim 15.4\times $compared to the open-source vector processor Ara. Notably, this work is the first RISC-V domain-specific processor tailored for diverse DL-based CE algorithms.
Chuanning Wang, Yangcan Zhou, Shaowei Wang 0001, Chuan Zhang 0001, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 SPEED: A Scalable RISC-V Vector Processor Enabling Efficient Multiprecision DNN Inference
abstract
Deploying deep neural networks (DNNs) on those resource-constrained edge platforms is hindered by their substantial computation and storage demands. Quantized multiprecision DNNs (MP-DNNs), denoted as MP-DNNs, offer a promising solution for these limitations but pose challenges for the existing RISC-V processors due to complex instructions, suboptimal parallel processing, and inefficient dataflow mapping. To tackle the challenges mentioned above, SPEED, a scalable RISC-V vector (RVV) processor, is proposed to enable efficient MP-DNN inference, incorporating innovations in customized instructions, hardware architecture, and dataflow mapping. First, some dedicated customized RISC-V instructions are introduced based on RVV extensions to reduce the instruction complexity, allowing SPEED to support processing precision ranging from 4- to 16-bit with minimized hardware overhead. Second, a parameterized multiprecision tensor unit (MPTU) is developed and integrated within the scalable module to enhance parallel processing capability by providing reconfigurable parallelism that matches the computation patterns of diverse MP-DNNs. Finally, a flexible mixed dataflow method is adopted to improve computational and energy efficiency according to the computing patterns of different DNN operators. The synthesis of SPEED is conducted on TSMC 28-nm technology. Experimental results show that SPEED achieves a peak throughput of 737.9 GOPS and an energy efficiency of 1383.4 GOPS/W for 4-bit operators. Furthermore, SPEED exhibits superior area efficiency compared with prior RVV processors, with the enhancements of$5.9\sim 26.9\times $and$8.2\sim 18.5\times $for 8-bit operator and best integer performance, respectively, which highlights SPEED’s significant potential for efficient MP-DNN inference.
Chuanning Wang, Chao Fang 0005, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2024 A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
abstract
RISC-V processors encounter substantial challenges in deploying multi-precision deep neural networks (DNNs) due to their restricted precision support, constrained throughput, and suboptimal dataflow design. To tackle these challenges, a scalable RISC-V vector (RVV) processor, namely SPEED, is proposed to enable efficient multi-precision DNN inference by innovations from customized instructions, hardware architecture, and dataflow mapping. Firstly, dedicated customized RISC-V instructions are proposed based on RVV extensions, providing SPEED with fine-grained control over processing precision ranging from 4 to 16 bits. Secondly, a parameterized multi-precision systolic array unit is incorporated within the scalable module to enhance parallel processing capability and data reuse opportunities. Finally, a mixed multi-precision dataflow strategy, compatible with different convolution kernels and data precision, is proposed to effectively improve data utilization and computational efficiency. We perform synthesis of SPEED in TSMC 28nm technology. The experimental results demonstrate that SPEED achieves a peak throughput of 287.41 GOPS and an energy efficiency of 1335.79 GOPS/W at 4-bit precision condition, respectively. Moreover, when compared to the pioneer open-source vector processor Ara, SPEED provides an area efficiency improvement of 2.04× and 1.63× under 16-bit and 8-bit precision conditions, respectively, which shows SPEED’s significant potential for efficient multi-precision DNN inference.
Chuanning Wang, Chao Fang 0005, Zhongfeng Wang 0001, Jun Lin 0001
ISCAS1