VLDB 2026 Research / reviewers in the wild / expert
Wei Mao 0002
dblp:51/4914-2
· DBLP profile ↗
20ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0003-2527-6778ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 7 first-author · 16 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 442.42 TOPS/W RRAM-based Digital Computing-in-Memory Accelerator for BF16×1-bit Vision Transformer
Jingyun Gu, Jiaqi Yang 0009, Jingyao Dong, Qilong Chen, Dingbang Liu, Wei Mao 0002, Hao Yu 0001 |
ISCAS | 8 |
| 2026 | FeRAM-Based Reconfigurable Strong PUF with Ultra-Low Power and Enhanced Attack Resilience
Xiguang Wu, Bo Li 0155, Jiuren Zhou, Wei Mao 0002, Yan Liu 0016, Genquan Han |
ISCAS | 6 |
| 2026 | Collaborative Design of FeRAM via a Joint Ferroelectric Device and Circuit AnalysisabstractFerroelectric random access memory (FeRAM) is a promising candidate to further dynamic random access memory (DRAM) scaling. However, the design of the FeRAM bit cell is nontrivial as the ferroelectric device model is not well supported by EDA tools. Modern integrated circuit design heavily depends on circuit-level SPICE simulators that integrate compact device models through modified nodal analysis (MNA) representation. This paper presents a novel MNA-based SPICE simulation method for ferroelectric device models, targeted at the design space exploration of FeRAM bitcells. Furthermore, this paper provides a co-design procedure for FeRAM bitcells and sense amplifiers via a comprehensive case study. Bo Li 0056, Junfeng Tan, Tingjie Yang, Huanning Zhang, Xueyang Bai, Wei Mao 0002, Jiuren Zhou, Guoyong Shi, Yan Liu 0016, Genquan Han |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2025 | A Layer-wised Mixed-Precision CIM Accelerator with Bit-level Sparsity-aware ADCs for NAS-Optimized CNNsabstractExploring multiple precisions as well as sparsities for a computingin-memory (CIM) based convolutional accelerators is challenging. To further improve energy efficiency with minimal accuracy loss, this paper develops a neural architecture search (NAS) method to identify precision for each layer of the CNN and further leverages bit-level sparsity. The results indicate that following this approach, ResNet-18 and VGG-16 not only maintain their accuracy but also implement layer-wised mixed-precision effectively. Furthermore, there is a substantial enhancement in the bit-level sparsity of weights within each layer, with an average bit-level sparsity exceeding 90% per bit, thus providing broader possibilities for hardware-level sparsity optimization. In terms of hardware design, a mixed-precision (2/4/8-bit) readout circuit as well as a bit-level sparsity-aware Analog-to-Digital Converter (ADC) are both proposed to reduce system power consumption. Based on bit-level sparsity mixed-precision CNNs benchmarks, post-layout simulation results in 28nm reveal that the proposed accelerator achieves up to 245.72 TOPS/W energy efficiency, which shows about 2.52 -- 6.57× improvement compared to the state-of-the-art SRAM-based CIM accelerators. Haoxiang Zhou, Zikun Wei, Dingbang Liu, Liuyang Zhang, Chenchen Ding, Jiaqi Yang 0009, Wei Mao 0002, Hao Yu 0001 |
ASP-DAC | 7 |
| 2025 | A 20.98TOPS/W Energy-Efficient Binary BERT Model on Group Vector Systolic CIM AcceleratorabstractTransformer-based large language models (LLMs) impose significant bandwidth and compute challenges when deployed on edge devices. SRAM-based compute-in-memory (CIM) accelerators offer a promising solution to reduce data movement but are still limited by model size. This work develops a ternary weight splitting (TWS) binarization to obtain Brain-Floating-Point-16×INT1 (BF16×1-b) and INT8×INT1 (8-b×1-b) based transformers that exhibit competitive accuracy while significantly reducing model size compared to full precision counterparts. Then, a fully digital SRAM-based CIM accelerator is designed incorporating a bit-parallel SRAM macro within a highly efficient group vector systolic architecture, which can store one column of BERT-Tiny model with stationary systolic data reuse. The design in a 28nm technology only requires 2KB SRAM with an area of 2mm2. It achieves a throughput of 6.55TOPS and consumes a total power of 312.5mW and 221mW at 400MHz, resulting in a state-of-the-art area efficiency of 3.3TOPS/mm2and normalized energy efficiency of 20.98TOPS/W and 34.35TOPS/W for BF16×1-b and 8-b×1-b respectively on BERT-Tiny model, demonstrating a 10.25× improvement in area efficiency and a 2.23× improvement in energy efficiency compared to other state-of-the-art counterparts. Additionally, our proposed configuration compresses the model size by 32% with only a 0.5% accuracy loss on SST-2. Dingbang Liu, Qilong Chen, Jingyun Gu, Jiaqi Yang 0009, Kai Li 0024, Wei Mao 0002, Ngai Wong 0001, Chang Wen Chen, Hao Yu 0001 |
ISLPED | 7 |
| 2025 | A parallel computing-in-memory accelerator utilizing FeRAM array with retention loss correction
Wei Mao 0002, Bo Li 0155, Xiaomeng Lv, Fuyi Li, Haiqiao Hong, Shirui Zhao, Siying Zheng, Jiuren Zhou, Yan Liu 0016, Genquan Han |
Sci. China Inf. Sci. | 3 |
| 2025 | A 28-nm 135.19 TOPS/W Bootstrapped-SRAM Compute-in-Memory Accelerator With Layer-Wise Precision and SparsityabstractArtificial intelligence (AI) edge devices demand high energy efficiency as well as inference accuracy. SRAM-based compute-in-memory (CIM) accelerators have great potential for power reduction but still need to exploit higher throughput and better linearity performance. To meet edge-AI computing demands by CIM works, it is crucial to optimize algorithms and parameters for specific circuit systems to achieve hardware acceleration. This work firstly employs neural network search (NAS) method to find out the layer-wise optimized precisions and sparsities for convolutional neural networks (CNNs). Then, a 144-Kb charge-domain signed mixed-precision (2/4/8-bit) CIM accelerator employing bootstrapped SRAM cells with 9-transistors and 1-capacitor (9T1C) structure is proposed that incorporates a bit-level sparsity-aware analog-to-digital converter (ADC). This work not only achieves highly linear parallel accumulation operations to meet AI computing demands but also implements a hardware and software co-optimization system tailored to specific data characteristics. The design is verified on NAS-optimized networks VGG-16 and ResNet-18 using Cifar-10 dataset, which could achieve an equivalent accuracy at 4-bit of 68.68% while maintaining a high energy efficiency at 2-bit of 135.19TOPS/W by measurements. Wei Mao 0002, Dingbang Liu, Haoxiang Zhou, Fuyi Li, Kai Li 0024, Qiuping Wu, Jiaqi Yang 0009, Liuyang Zhang, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | APCCAS 2022 Guest Editorial Special Issue Based on the 18th Asia Pacific Conference on Circuits and SystemsabstractThe IEEE Asia Pacific Conference on Circuits and Systems (APCCAS) is the regional flagship conference of the IEEE Circuits and Systems Society (CASS) in Asia. This conference is a major international forum established by the IEEE Circuits and Systems Society for researchers to exchange their latest findings in circuits and systems. It covers a wide range of topics, including analog, mixed-signal, digital, communication, sensory, biomedical, power/energy, nonlinear, and artificial intelligence circuits and systems. Xiaojin Zhao, Hailong Jiao, Wei Mao 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | An Energy-Efficient Bit-Split-and-Combination Systolic Accelerator for NAS-Based Multi-Precision Convolution Neural NetworksabstractOptimized convolutional neural network (CNN) models and energy-efficient hardware design are of great importance in edge-computing applications. The neural architecture search (NAS) methods are employed for CNN model optimization with multi-precision networks. To satisfy the computation requirements, multi-precision convolution accelerators are highly desired. The existing high-precision-split (HPS) designs reduce the additional logics for reconfiguration while resulting in low throughput for low precisions. The low-precision-combination (LPC) designs improve the low-precision throughput with large hardware cost. In this work, a bit-split-and-combination (BSC) systolic accelerator is proposed to overcome the bottlenecks. Firstly, BSC-based multiply-accumulate (MAC) unit is designed to support multi-precision computation operations. Secondly, multi-precision systolic dataflow is developed with improved data-reuse and transmission efficiency. The proposed work is designed by Chisel and synthesized in 28-nm process. The BSC MAC unit achieves maximum 2.40× and 1.64× energy efficiency than HPS and LPC units, respectively. Compared with published accelerator designs Gemmini, Bit-fusion and Bit-serial, the proposed accelerator achieves up to 2.94 × area efficiency and 6.38 × energy-saving performance on the multi-precision VGG-16, ResNet-18 and LeNet-5 benchmarks. Liuyao Dai, Gengbin Huang, Junzhuo Zhou, Kai Li 0024, Wei Mao 0002, Hao Yu 0001 |
ASP-DAC | 7 |
| 2022 | A Precision-Scalable Energy-Efficient Bit-Split-and-Combination Vector Systolic Accelerator for NAS-Optimized DNNs on EdgeabstractOptimized model and energy-efficient hardware are both required for deep neural networks (DNNs) in edge-computing area. Neural architecture search (NAS) methods are employed for DNN model optimization with resulted multi-precision networks. Previous works have proposed low-precision-combination (LPC) and high-precision-split (HPS) methods for multi-precision networks, which are not energy-efficient for precision-scalable vector implementation. In this paper, a bit-split-and-combination (BSC) based vector systolic accelerator is developed for a precision-scalable energy-efficient convolution on edge. The maximum energy efficiency of the proposed BSC vector processing element (PE) is up to 1.95× higher in 2-bit, 4-bit and 8-bit operations when compared with LPC and HPS PEs. Further with NAS optimized multi-precision CNN networks, the averaged energy efficiency of the proposed vector systolic BSC PE array achieves up to 2.18× higher in 2-bit, 4-bit and 8-bit operations than that of LPC and HPS PE arrays. Kai Li 0024, Junzhuo Zhou, Junyi Luo, Zhengke Yang, Shuxin Yang, Wei Mao 0002, Mingqiang Huang, Hao Yu 0001 |
DATE | 7 |
| 2022 | ANT-UNet: Accurate and Noise-Tolerant Segmentation for Pathology Image ProcessingabstractPathology image segmentation is an essential step in early detection and diagnosis for various diseases. Due to its complex nature, precise segmentation is not a trivial task. Recently, deep learning has been proved as an effective option for pathology image processing. However, its efficiency is highly restricted by inconsistent annotation quality. In this article, we propose an accurate and noise-tolerant segmentation approach to overcome the aforementioned issues. This approach consists of two main parts: a preprocessing module for data augmentation and a new neural network architecture, ANT-UNet. Experimental results demonstrate that, even on a noisy dataset, the proposed approach can achieve more accurate segmentation with 6% to 35% accuracy improvement versus other commonly used segmentation methods. In addition, the proposed architecture is hardware friendly, which can reduce the amount of parameters to one-tenth of the original and achieve 1.7× speed-up. Yufei Chen 0007, Tingtao Li, Qinming Zhang, Wei Mao 0002, Nan Guan, Hao Yu 0001, Cheng Zhuo |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | A High Performance Multi-Bit-Width Booth Vector Systolic Accelerator for NAS Optimized Deep Learning Neural NetworksabstractMulti-bit-width convolutional neural network (CNN) maintains the balance between network accuracy and hardware efficiency, thus enlightening a promising method for accurate yet energy-efficient edge computing. In this work, we develop state-of-the-art multi-bit-width accelerator for NAS Optimized deep learning neural networks. To efficiently process the multi-bit-width network inferencing, multi-level optimizations have been proposed. Firstly, differential Neural Architecture Search (NAS) method is adopted for the high accuracy multi-bit-width network generation. Secondly, hybrid Booth based multi-bit-width multiply-add-accumulation (MAC) unit is developed for data processing. Thirdly, vector systolic array is proposed for effectively accelerating the matrix multiplications. With vector-style systolic dataflow, both the processing time and logic resources consumption can be reduced when compared with the classical systolic array. Finally, The proposed multi-bit-width CNN acceleration scheme has been practically deployed on FPGA platform of Xilinx ZCU102. Average performance on accelerating the full NAS optimized VGG16 network is 784.2 GOPS, and peek performance of the convolutional layer can reach as high as 871.26 GOPS for INT8, 1676.96 GOPS for INT4, and 2863.29 GOPS for INT2 respectively, which is among the best results in previous CNN accelerator benchmarks. Mingqiang Huang, Yucen Liu, Changhai Man, Kai Li 0024, Wei Mao 0002, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | A Fall Detection Network by 2D/3D Spatio-temporal Joint Models with Tensor Compression on EdgeabstractFalling is ranked highly among the threats in elderly healthcare, which promotes the development of automatic fall detection systems with extensive concern. With the fast development of the Internet of Things (IoT) and Artificial Intelligence (AI), camera vision-based solutions have drawn much attention for single-frame prediction and video understanding on fall detection in the elderly by using Convolutional Neural Network (CNN) and 3D-CNN, respectively. However, these methods hardly supervise the intermediate features with good accurate and efficient performance on edge devices, which makes the system difficult to be applied in practice. This work introduces a fast and lightweight video fall detection network based on a spatio-temporal joint-point model to overcome these hurdles. Instead of detecting fall motion by the traditional CNNs, we propose a Long Short-Term Memory (LSTM) model based on time-series joint-point features extracted from a pose extractor . We also introduce the increasingly mature RGB-D camera and propose 3D pose estimation network to further improve the accuracy of the system. We propose to apply tensor train decomposition on the model to reduce storage and computational consumption so the deployment on edge devices can to realized. Experiments are conducted to verify the proposed framework. For fall detection task, the proposed video fall detection framework achieves a high sensitivity of 98.46% on Multiple Cameras Fall, 100% on UR Fall, and 98.01% on NTU RGB-D 120. For pose estimation task, our 2D model attains 73.3 mAP in the COCO keypoint challenge, which outperforms the OpenPose by 8%. Our 3D model attains 78.6% mAP on NTU RGB-D dataset with 3.6× faster speed than OpenPose. Shuwei Li, Changhai Man, Wei Mao 0002, Shaobo Luo, Rumin Zhang, Hao Yu 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2022 | An Energy-Efficient Mixed-Bitwidth Systolic Accelerator for NAS-Optimized Deep Neural NetworksabstractOptimized deep neural network (DNN) models and energy-efficient hardware designs are of great importance in edge-computing applications. The neural architecture search (NAS) methods are employed for DNN model optimization with mixed-bitwidth networks. To satisfy the computation requirements, mixed-bitwidth convolution accelerators are highly desired for low-power and high-throughput performance. There exist several methods to support mixed-bitwidth multiply-accumulate (MAC) operations in DNN accelerator designs. The low-bitwidth-combination (LBC) method improves the low-bitwidth throughput with a large hardware cost. The high-bitwidth-split (HBS) method minimizes the additional logic gates for configuration. However, the throughput performance in the low-bitwidth mode is poor. In this work, a bit-split-and-combination (BSC) systolic accelerator is proposed. The BSC-based MAC unit is designed to support mixed-bitwidth operations with the best overall performance. Besides, interprocessing element (PE) systolic and intra-PE paralleled dataflow not only improves throughput performance in mixed-bitwidth modes, but also saves power performance for data transmission. The proposed work is designed and synthesized in a 28-nm process. The BSC MAC unit achieves a maximum $2.08\times $ and $1.75\times $ energy efficiency improvement than the HBS and LBC unit, respectively. Compared with the state-of-the-art accelerators, the proposed work also achieves excellent energy-efficient performance with 20.02, 23.55, and 30.17 TOPS/W on mixed-bitwidth VGG-16, ResNet-18, and LeNet-5 benchmarks at 0.6 V, respectively. Wei Mao 0002, Liuyao Dai, Kai Li 0024, Laimin Du, Shaobo Luo, Mingqiang Huang, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | A Configurable Floating-Point Multiple-Precision Processing Element for HPC and AI Converged ComputingabstractThere is an emerging need to design configurable accelerators for the high-performance computing (HPC) and artificial intelligence (AI) applications in different precisions. Thus, the floating-point (FP) processing element (PE), which is the key basic unit of the accelerators, is necessary to meet multiple-precision requirements with energy-efficient operations. However, the existing structures by using high-precision-split (HPS) and low-precision-combination (LPC) methods result in low utilization rate of the multiplication array and long multiterm processing period, respectively. In this article, a configurable FP multiple-precision PE design is proposed with the LPC structure. Half precision, single precision, and double precision are supported. The 100% multiplier utilization rate of the multiplication array for all precisions is achieved with improved speed in the comparison and summation process. The proposed design is realized in a 28-nm process with 1.429-GHz clock frequency. Compared with the existing multiple-precision FP methods, the proposed structure achieves 63% and 88% area-saving performance for FP16 and FP32 operations, respectively. The$4\times $and$20\times $maximum throughput rates are obtained when compared with fixed FP32 and FP64 operations. Compared with the previous multiple-precision PEs, the proposed one achieves the best energy-efficiency performance with 975.13 GFLOPS/W. Wei Mao 0002, Kai Li 0024, Liuyao Dai, Xinang Xie, He Li 0008, Longyang Lin, Hao Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | A Video-based Fall Detection Network by Spatio-temporal Joint-point Model on Edge DevicesabstractTripping or falling is among the top threats in elderly healthcare, and the development of automatic fall detection systems are of considerable importance. With the fast development of the Internet of Things (IoT), camera vision-based solutions have drawn much attention in recent years. The traditional fall video analysis on the cloud has significant communication overhead. This work introduces a fast and lightweight video fall detection network based on a spatio-temporal joint-point model to overcome these hurdles. Instead of detecting falling motion by the traditional Convolutional Neural Networks (CNNs), we propose a Long Short-Term Memory (LSTM) model based on time-series joint-point features, extracted from a pose extractor and then filtered from a geometric joint-point filter. Experiments are conducted to verify the proposed framework, which shows a high sensitivity of 98.46% on Multiple Cameras Fall Dataset and 100% on UR Fall Dataset. Furthermore, our model can achieve pose estimation tasks simultaneously, attaining 73.3 mAP in the COCO keypoint challenge dataset, which outperforms the OpenPose work by 8%. Shuwei Li, Changhai Man, Wei Mao 0002, Ngai Wong 0001, Hao Yu 0001 |
DATE | 5 |
| 2021 | A Reconfigurable Multiple-Precision Floating-Point Dot Product Unit for High-Performance ComputingabstractThere is an emerging need to optimize floating-point (FP) dot product units (DPU) for high-performance scientific computing as well as training deep learning models. Due to different precision requirements of applications, a reconfigurable multiple-precision DPU operation can largely reduce the cost of area and power. However, the existing methods could result in redundant bits for unit multipliers, but also leave idle hardware resources for the operations in different precisions. In this paper, a reconfigurable multiple-precision FP DPU design is proposed for high-performance computing (HPC) applications. The FP DPU can be reconfigured as follows. A bit-partitioning method is provided to minimize the redundant bits with a configurable mixed-precision multiplier for three-mode operations: 20 half-precision Dot Product (DP), 5 single-precision DP, and 1 double-precision DP operations. Any of the modes can be executed in two successive clock cycles without idle hardware resources. The proposed design is realized by using the UMC 55-nm process with simulation results. Compared with the existing multiple-precision FP methods, the proposed DPU achieves 88.9% and 35.8% area-saving performance for FP16 and FP32 operations, respectively. Moreover, when using benchmarked HPC applications where multiple precisions can be used, the proposed reconfigurable DPU can accelerate up to 4× and 20× maximum throughput rates when compared with fixed FP32 and FP64 operations, respectively. Wei Mao 0002, Kai Li 0024, Xinang Xie, Shirui Zhao, He Li 0008, Hao Yu 0001 |
DATE | 1 |
| 2020 | Energy-Efficient Machine Learning Accelerator for Binary Neural NetworksabstractBinary neural network (BNN) has shown great potential to be implemented with power efficiency and high throughput. Compared with its counterpart, the convolutional neural network (CNN), BNN is trained with binary constrained weights and activations, which are more suitable for edge devices with less computing and storage resource requirements. In this paper, we introduce the BNN characteristics, basic operations and the binarized-network optimization methods. Then we summarize several accelerator designs for BNN hardware implementation by using three mainstream structures, i.e., ReRAM-based crossbar, FPGA and ASIC. Based on the BNN characteristics and hardware custom designs, all these methods achieve massively parallelized computations and highly pipelined data flow to enhance its latency and throughput performance. In addition, the intermediate data with the binary format are stored and processed on chip by constructing the computing-in-memory (CIM) architecture to reduce the off-chip communication costs, including power and latency. Wei Mao 0002, Zhihua Xiao, Peng Xu 0035, Dingbang Liu, Shirui Zhao, Fengwei An, Hao Yu 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | High Dynamic Performance Current-Steering DAC Design With Nested-Segment Structure
Wei Mao 0002, Yongfu Li 0002, Chun-Huat Heng, Yong Lian 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Zero-bias true random number generator using LFSR-based scramblerabstractIn this paper, we proposed an improved true random number generator (TRNG), which comprises a low-bias hardware random number generator (HRNG) and a scrambler based on linear-feedback shift register (LFSR). The HRNG reduces both DC offset from the noise sources and offset voltage from the comparator to generate low-bias bitstream. The LFSR-based scrambler further reduces the bias to zero without sacrificing the throughput rate. Randomness quality is verified by Monte Carlo simulations using the randomness test suite. Wei Mao 0002, Yongfu Li 0002, Chun-Huat Heng, Yong Lian 0001 |
ISCAS | 1 |