Ao Shi

dblp:380/5630 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0001-8200-5360ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 High-accuracy and Energy-efficient RRAM-based Real-time Object Detection System for Thermal Infrared Imaging Applications
Kexun Li, Ao Shi, Hairuo Lu, Lianliang Wu, Yulin Feng, Lifeng Liu
ISCAS3
2026 A 28nm 143.4-322.5TOPS/W INT8 time-domain CIM macro featuring zero-weight skipping and shift-and-add embedded TDC for Deep Neural Network
Ao Shi, Lianliang Wu, Haobin Shang, Kexun Li, Lifeng Liu, Jinfeng Kang, Peng Huang 0004
ISCAS1
2026 RRAM based CIM, PUF and True Random Number Generator for High-Security AES Encryption System
Kefan Tao, Shiyue Song, Yading Yi, Ao Shi, Haokai Guan, Lianliang Wu, Hao Ai, Lifeng Liu, Yulin Feng, Peng Huang 0004
ISCAS4
2026 RRAM-based CAM for Energy-Efficient In-Memory Text Compression System
Lianliang Wu, Hao Ai, Ao Shi, Haokai Guan, Kexun Li, Kefan Tao, Yulin Feng, Zongwei Wang 0001, Yimao Cai, Peng Huang 0004
ISCAS4
2025 CIMUS: 3D-Stacked Computing-in-Memory Under Image Sensor Architecture for Efficient Machine Vision
abstract
Computational image sensors with CNN processing capabilities are emerging to alleviate the energy-intensive and time-consuming data movement between sensors and external processors. However, deploying CNN models onto these computational image sensors faces challenges from the limited on-chip memory resources and insufficient image processing throughput. This work proposes a 3D-stacked NAND flash-based computing-in-memory under image sensor architecture (CIMUS) to facilitate the complete deployment of CNN model. To fully leverage the potential of high bandwidth from the 3D-stacked integration, we design a novel distributed CNN mapping and dataflow to process the full focal plane image in parallel, which senses and recognizes ImageNet tasks with >1000fps. To tackle the computational error of inputs “0” in 3D NAND flash-based CIM, we propose an input-independent offset compensation method, which reduces the average vector-matrix multiplication (VMM) error by 48%. Evaluation results indicate that CIMUS architecture achieves a 9.8× improvement in CNN inference speed and a 33× boost in energy efficiency compared to the state-of-the-art computational image sensor in the ImageNet recognition task.
Lixia Han, Haozhang Yang, Ao Shi, Guihai Yu, Yijiao Wang, Yanzhi Wang 0001, Jinfeng Kang, Peng Huang 0004
IEEE Trans. Computers5
2024 Low Quantization Error Readout Circuit with Fully Charge-Domain Calculation for Computation-in-Memory Deep Neural Network
abstract
This work presents a low quantization error readout circuit with fully-charge-domain calculation for quantization and post-process of computation-in-memory (CIM)-based neural network. The contributions include: (1) A novel residual charge accumulation function is designed to achieve charge-domain summation of quantized partial sum, and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize <1 LSB INL at ±7 bits and speed of 285MHz/LSB; (3) Sample & hold, current subtraction and bidirectional counter are designed to improve 3.95× energy efficiency and 2.48× area efficiency.
Ao Shi, Lixia Han, Lifeng Liu, Linxiao Shen, Peng Huang 0004, Jinfeng Kang
ISCAS1
2024 Specific ADC of NVM-Based Computation-in-Memory for Deep Neural Networks
abstract
Non-volatile memory (NVM)-based Computation-in-memory has demonstrated a significant advantage in high-efficiency neural networks. However, the requirement of analog-to-digital converter (ADC) and post-processing circuits not only cost high energy and area but also results in high computation errors, which tradeoffs the performance boost brought by CIM. Here, we present a specific ADC and post-processing circuit of the NVM-based CIM neural network to address these issues. The main contributions include: (1) A novel residual charge accumulation function (RCA) is designed to achieve charge-domain summation of quantized partial sum and reduces 38% quantization error; (2) Charge reset is introduced in the integrate & fire circuit to realize$3.95\times $energy efficiency and$2.48\times $area efficiency. Evaluation based on the measured results of the fabricated chip shows that the VGG-11 neural network with the proposed ADC circuit can achieve a 3.28-time improvement in energy efficiency while maintaining the same network recognition rate.
Ao Shi, Lixia Han, Haozhang Yang, Lifeng Liu, Linxiao Shen, Jinfeng Kang, Peng Huang 0004
IEEE Trans. Circuits Syst. I Regul. Pap.1