Aifei Zhang

dblp:333/4170 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Efficient Weight Mapping and Resource Scheduling on Crossbar-based Multi-core CIM Systems
abstract
Crossbar-based computing-in-memory (CIM) systems facilitate large-scale parallel multiply-and-accumulate (MAC) operations, while a domain-specific compiler (DSC) plays a pivotal role in optimizing the deployment of neural network algorithms on such systems. With the development of multi-core and large-core architectures, some key compiler problems such as high parallel processing, resource utilization, and crossbar array assignment methods have not been solved. For low-latency application scenarios, we have designed a resource scheduling strategy for our hardware system based on stream data processing to reduce the latency caused by intra-core and intercore communication. Additionally, a weight mapping strategy has been developed to maximize the potential of crossbar arrays in convolutional neural networks (CNNs) deployment. Experimental results on our multi-core eFlash-based CIM system-on-chip (SoC) demonstrate that these two technologies help CNNs achieve a 76% reduction in latency, a 30% improvement in resource utilization, and the use rate of crossbar array that can reach up to 94.7%.
Sifan Sun, Aifei Zhang, Haiyan Qin, Minhao Gu, Shihang Fu, Shuaikai Liu, Baosen Liu, Wang Kang 0001
DAC3
2025 HyIMC: Analog-Digital Hybrid In-Memory Computing SoC for High-Quality Low-Latency Speech Enhancement
abstract
In-memory computing (IMC) holds significant promise for accelerating deep learning-based speech enhancement (DL-SE). However, existing IMC architectures face challenges in simultaneously achieving high precision, energy efficiency, and the necessary parallelism for DL-SE's inherent temporal dependencies. This paper introduces HyIMC, a novel hybrid analog-digital IMC architecture designed to address these limitations. HyIMC features: 1) a hybrid analog-digital design optimized for DL-SE algorithms; 2) a schedule controller that efficiently manages recurrent dataflow within skip connections; and 3) non-key dimension shrinkage, a model compression technique that preserves accuracy. Implemented on a 40nm eFlash-based IMC SoC prototype, HyIMC achieves 160 TOPS/W energy efficiency, compresses the DL-SE model size by ~600%, improves the feature of merit by ~1200%, and enhances perceptual evaluation of speech quality by ~120%.
Wanru Mao, Guangyao Wang, Tianshuo Bai, Jingcheng Gu, Xitong Yang, Aifei Zhang, Xiaohang Wei, Wang Kang 0001
DATE8
2024 An End-to-End In-Memory Computing System Based on a 40-nm eFlash-Based IMC SoC: Circuits, Toolchains, and Systems Co-Design Framework
abstract
Despite its promising potential for Artificial Intelligence (AI) applications, current In-Memory Computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance including (a) an 8-bit hardware-friendly Quantization-Aware Training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data, (b) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips, and (c) an efficient mapping strategy based on the Integer Linear Programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40nm eFlash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 hours for voice recognition, a 21.53% improvement for the Perceptual Evaluation of Speech Quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection.
Tianshuo Bai, Wanru Mao, Guangyao Wang, Aifei Zhang, Shihang Fu, Shuaikai Liu, Jianchao Hu, Xitong Yang, Biao Pan, Wei W. Xing, Wang Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 A 4-bit Integer-Only Neural Network Quantization Method Based on Shift Batch Normalization
abstract
Neural networks are powerful, but at the cost of huge amounts of computation. Deploying neural networks on edge devices is especially challenging. Quantization is a possible solution to alleviate the huge cost, while most quantization methods are not sufficiently hardware-friendly. In this paper, we proposed an integer-only quantization method. With no division or big integer multiplication, this quantization method is suitable to be deployed on co-designed hardware platforms. We applied 4-bit quantization on some classical networks and corresponding datasets. On MNIST, CIFAR10 and CFAR100, quantization networks perform as well as original networks. On SpeechCommands, accuracy error induced by quantization is 0.16%. We also deployed quantized networks under OpenCL framework and on a flash-based in-memory-computing chip to verify this method’s feasibility.
Qingyu Guo, Xiaoxin Cui, Aifei Zhang, Xinjie Guo, Yuan Wang 0001
ISCAS4