Zhen Wang 0019

dblp:78/6727-19 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-7821-2252ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TNNSAT: A Tree-NN Based Small-Footprint ASR Processor with Scene-Adaptive On-Chip Training
Ningyuan Li 0004, Zhichen Zhang, Zhixuan Gong, Zhen Wang 0019, Bo Liu 0019
ISCAS7
2024 Optimizing the Micro-Architectural Performance of the Current and Emerging Edge Infrastructure
abstract
The Network Function Virtualization (NFV) is the essential technology proposed to tackle the next-generation mobile system’s various flexibility features. In this article, we implement a thorough micro-architectural performance investigation on the NFV-enabled edge virtual Radio Access Network (vRAN) and the emerging 5G new-radio (nr) platform to unveil the main micro-architectural bottlenecks of the next-generation network’s vRAN system. Based on our experimental results, we find that the high core bound hinders the processing speed of the vRAN and 5G nr platforms. Several solutions alleviating the vRAN’s core bound are proposed to accelerate the vRAN system’s processing speed. Besides, we observe that the current co-location strategy cannot maximize the COTS servers’ CPU utilization and meanwhile eliminate the system hang-up caused by CPU resource contention. We fill this gap by proposing an optimized co-location strategy based on our observed vRAN co-location characterization. Finally, we detect that on the modern hyper-threading-enabled COTS servers, the current pin core policy of 5G nr will cause L3 cache contention, which will lead to severe system hang-up. A novel threads management mechanism is proposed to eliminate this system hang-up on the hyper-threading-enabled COTS servers.
Zhen Wang 0019, Weili Wu 0001, Yang Hu 0001
IEEE Trans. Cloud Comput.2
2023 Towards an Efficient SIMD Virtual Radio Access Network (vRAN) and Edge Cloud System
abstract
Nowadays, virtual Radio Access Network (vRAN) plays a vital role in today's mobile edge system for its better support for latency-sensitive applications. However, our characterization of vRAN on modern processors depicts a frustrating picture of Single-Instruction Multi-Data (SIMD) acceleration. Specifically, the existing data arrangement processes cannot efficiently utilize the ports in modern processors, which leads to high backend bound and fails to saturate the memory bandwidth between registers and the L1 cache. To tackle the issue, we thoroughly examine the state-of-the-art CPU architecture and observe the idle ports which could be utilized by the process. Motivated by this observation, we propose an “Arithmetic Ports Consciousness Mechanism” (APCM) utilizing these idle ports to eliminate the backend bound and saturate the memory bandwidth. The APCM decreases the data arrangement's backend bound from 45$\%$to 3$\%$and promotes its memory bandwidth utilization by 4X-16X. Moreover, we illustrate that the APCM can be utilized to promote the performance of typical mobile edge applications such as network routing, image processing, and AI applications. The CPU time of the data arrangement process time of the selected typical mobile edge applications can be reduced by 55$\%$- 95$\%$when utilizing the proposed mechanism.
Zhen Wang 0019, Yang Hu 0001
IEEE Trans. Cloud Comput.2
2022 A Target-Separable BWN Inspired Speech Recognition Processor with Low-power Precision-adaptive Approximate Computing
abstract
This paper proposes a speech recognition processor based on a target-separable binarized weight network (BWN), capable of performing both speaker verification (SV) and keyword spotting (KWS). In traditional speech recognition system, the SV based on traditional model and the KWS based on neural networks (NN) model are two independent hardware modules. In this work, both SV and KWS are processed by the proposed BWN with unified training and optimization framework which can be performed for various application scenarios. By the system-architecture co-design, SV and KWS share most of the network parameters, and the classification part is calculated separately according to different targets. An energy-efficient NN accelerator which can be dynamically reconfigured to process different layers of the BWN with splitting calculation of frequency domain convolution is proposed. SV and KWS can be achieved with only one time calculation of each input speech frame, which greatly improves the computing energy efficiency. The computing units of the NN accelerator are optimized using precision-adaptive approximate computing method with Dual-VDD to further reduce the energy cost. Compared to state-of-the-arts, this work can achieve about 4 × reduction in power consumption while maintaining high system adaptability and accuracy.
Bo Liu 0019, Hao Cai 0001, Haige Wu, Anfeng Xue, Zhen Wang 0019, Jun Yang 0006
DATE7
2022 Self-compensation tensor multiplication unit for adaptive approximate computing in low-power CNN processing
Bo Liu 0019, Hao Cai 0001, Reyuan Zhang, Zhen Wang 0019, Jun Yang 0006
Sci. China Inf. Sci.5
2022 An Efficient BCNN Deployment Method Using Quality-Aware Approximate Computing
abstract
As the artificial intelligence and Internet of Things (AIoT) develop rapidly, the deployment of artificial neural networks in edge computing is becoming significant with great challenge. The binarized convolutional neural network (BCNN) is one of the most widely adopted light-weight ANNs in AIoT, which can achieve the balance of system accuracy and hardware resource consumption, compared to others. To achieve high power and area efficiency in BCNN deployment, many approximate computing (AxC) techniques are integrated to make full use of the resilience of BCNN. As the research focused on the integration of AxC in circuit design, the design of AxC itself is not fully considered when applied to specific applications or domains. Based on circuit-architecture-system co-design, this article proposes an efficient BCNN deployment method, including a quality-circuit co-design method for approximate adder generation, a quality-aware intercompensation approach for addition tree, and a computing quality involved retraining approach for BCNN deployment. Experimental results show that the proposed quality model can achieve 86.43% in average accuracy while evaluating nine types of typical approximate adders. The proposed method is conducted on the applications of keyword spotting of GSCD, MNIST, and CIFAR-10, and we can further rise the approximation degree by 50%–75%, while reducing the accuracy by less than 1%.
Bo Liu 0019, Xuetao Wang, Anfeng Xue, Qiao Shen 0001, Na Xie, Yu Gong 0002, Zhen Wang 0019, Jun Yang 0006, Hao Cai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2022 Quality Driven Systematic Approximation for Binary-Weight Neural Network Deployment
abstract
Neural networks (NNs) with large scales of artificial neurons are increasingly used in recognition and classification tasks. In power-constrained scenarios, the tradeoff between performance and hardware consumptions must be carefully evaluated before silicon tape-out. In this paper, we proposed a systematic approach to design ultra-low power NN system. This work is motivated by the facts that NNs are resilient to approximation in many of the computations and NNs are outputting statistical tensors which are acceptable to less-than-perfect results. We resort to the front-back end approach with a twofold aim: (1) a fast and accurate design approach is proposed by estimating the computing quality of low-power approximate adder arrays, and it is adopted to evaluate the neural network system; (2) a quality configurable engine with different approximation degrees while processing NNs is implemented. The proposed work is demonstrated with a comprehensive keyword spotting (KWS) system as an ultra-low power NN engine. The experimental environment is setup with ten keywords from the google speech command dataset (GSCD) using an industrial 22-nm ultra-low-leakage (ULL) process. Comparing to the state-of-the-art KWS processors, the proposed approximate NN engine can demonstrate over 60% improvement in power efficiency and$1.1\times $area efficiency while achieving similar recognition accuracy.
Yu Gong 0002, Hao Cai 0001, Haige Wu, Hao Yan 0002, Zhen Wang 0019, Longxing Shi, Bo Liu 0019
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 More is Less: Domain-Specific Speech Recognition Microprocessor Using One-Dimensional Convolutional Recurrent Neural Network
abstract
Low-power keywords recognition has been a focus of acoustic signal processing for several decades. This work investigates the domain-specific speech recognition microprocessor based on optimized one-dimensional convolutional recurrent neural network (1D-CRNN). Compared to previous DNN based frameworks, the proposed 1D-CRNN can process both the feature extraction and keywords classification, and achieve high recognition accuracy with reduced computation operations under wide range background noise SNRs. An energy-efficient 1D-CRNN accelerator is implemented to dynamically reconfigure and process the different layers. This accelerator has the characteristics of “More is Less” in three aspects: 1) the hybrid network with more complex layers is much more compact and requires less computation; 2) although the weight width quantized to 8 bits requires more memory size and multiplication energy cost, the required network neurons can be reduced and hardware utilization can be improved; 3) an energy-aware self-compensation tensor multiplication unit with dual power supply based on approximation design method can be utilized for 1D-CRNN computing. Compared to the state-of-the-art architectures, the novel more-is-less architecture can achieve a much lower power consumption of$1.4~\mu \text{W}\sim 2.1~\mu \text{W}$(over 80% reduced) under an industry 22nm technology, while maintaining higher system adaptability (support SNRs: −5dB~Clean) for 1~5 real-time keywords recognition.
Bo Liu 0019, Hao Cai 0001, Xiaoling Ding, Yu Gong 0002, Weiqiang Liu 0001, Jinjiang Yang, Zhen Wang 0019, Jun Yang 0006
IEEE Trans. Circuits Syst. I Regul. Pap.9
2021 A survey of in-spin transfer torque MRAM computing
Hao Cai 0001, Bo Liu 0019, Juntong Chen, Lirida A. B. Naviner, Yongliang Zhou, Zhen Wang 0019, Jun Yang 0006
Sci. China Inf. Sci.6
2020 An Ultra-low Power Keyword-Spotting Accelerator Using Circuit-Architecture-System Co-design and Self-adaptive Approximate Computing Based BWN
abstract
This paper proposed an ultra-low power keyword-spotting (KWS) accelerator using circuit-architecture-system co-design and precision self-adaptive approximate computing based binarized weight network (BWN). To reduce the power consumption while maintaining the system recognition accuracy for different background noise, we first proposed a bit-by-bit layer-by-layer quantization method to quantize the deep neural network (DNN) to BWN. Then, we proposed a precision self-adaptive approximate addition unit to further reduce the BWN energy consumption. Evaluated under TSMC22nm ULL process technology, this work can support up to 10 keywords real time recognition under different background noise types and SNRs (from 5dB to near microphone) with power consumption of 13.6uW.
Bo Liu 0019, Hao Cai 0001, Zeyu Shen 0003, Yu Gong 0002, Lepeng Huang, Zhen Wang 0019
ACM Great Lakes Symposium on VLSI7
2020 A Background Noise Self-adaptive VAD Using SNR Prediction Based Precision Dynamic Reconfigurable Approximate Computing
abstract
This paper proposed a background-noise self-adaptive voice activity detection (VAD) accelerator using SNR prediction based precision dynamic reconfigurable approximate computing. To improve the energy efficiency while maintaining high recognition accuracy for different background noises, two optimization techniques are proposed. Firstly, we proposed a SNR prediction module to analyze and pre-classify the back-ground noise into different levels, and a binarized weight network (BWN) accelerator with reconfigurable data bit width to implement the feature classification of VAD. Then, we proposed an approximate computing architecture with precision self-adaptive approximate addition unit to further reduce the energy consumption of BWN accelerator. Evaluated under 28nm process technology, this work can achieve high recognition accuracy (speech/none-speech hit rate: 95%/92% @10dB, 90%/87% @5dB, and 85%/80% @-5dB) under different background noise (SNR-5dB) with a low power consumption of 2 ~ 8uW.
Bo Liu 0019, Yan Li 0056, Lepeng Huang, Hao Cai 0001, Shisheng Guo, Yu Gong 0002, Zhen Wang 0019
ACM Great Lakes Symposium on VLSI8
2020 Binarized Weight Neural-Network Inspired Ultra-Low Power Speech Recognition Processor with Time-Domain Based Digital-Analog Mixed Approximate Computing
abstract
In this paper, an ultra-low power speech recognition processor is implemented based on an optimized binarized weight neural-network (BWN). To accelerate the BWN and make it energy efficient, we proposed an approximate computing architecture for the quantized BWN based on time-domain digital-analog mixed addition unit and precision optimization with fault-tolerant training method. Experimental results show that the proposed digital-analog mixed approximate computing architecture can significantly reduce the power consumption while maintaining the recognition accuracy. Implemented under TSMC 28nm, the proposed processor can support 10 keywords real time recognition under different noise types and SNRs, while the power consumption is 56μW.
Bo Liu 0019, Hao Cai 0001, Yu Gong 0002, Yan Li 0056, Zhen Wang 0019
ISCAS7
2020 Enabling Latency-Aware Data Initialization for Integrated CPU/GPU Heterogeneous Platform
abstract
Nowadays, driven by the needs of autonomous driving and edge intelligence, integrated CPU/GPU heterogeneous platform has gained significant attention from both academia and industry. As the representative series, NVIDIA Jetson family perform well in terms of computation capability, power consumption, and mobile size. Even so, the integrated heterogeneous platform only contains one limited physical memory, which is shared by the CPU and GPU cores and can be the performance bottleneck of the mobile/edge applications. On the other hand, with the unified memory (UM) model introduced in GPU programming, not only the memory allocation is significantly reduced, which mitigates the memory bottleneck of the integrated platforms but also the memory management and programming are simplified. However, as a programming legacy, the UM model still follows the conventional copy-then-execute model, initializing data on the CPU side after allocating memory. This legacy programming mode not only causes significant initialization latency but also slows the execution of the following kernel. In this article, we propose a framework to enable the latency-aware data initialization on the integrated heterogeneous platform. The framework not only includes three data initialization modes, the CPU initialization, GPU initialization, and hybrid initialization, but also utilizes an affinity estimation model to wisely decide the best initialization mode for an application such that the initialization latency performance of the application can be optimized. We evaluate our design on NVIDIA TX2 and AGX platforms. The results demonstrate that the framework can accurately select a data initialization mode for a given application to significantly reduce the initialization latency. We envision this latency-aware data initialization framework being adopted in a full-version of autonomous solution (e.g., Autoware) in the future.
Zihang Jiang, Zhen Wang 0019, Xulong Tang, Cong Liu 0005, Shouyi Yin, Yang Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3