Chih-Cheng Lu

dblp:50/7360 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author
YearPublicationVenuePosition
2026 A High-Area-Efficiency Computing-in-Memory Deep Learning Accelerator Based on Weight-Sparsity Activation Compression and Sparsity Reciprocity
abstract
The large number of parameters in DNNs poses significant challenges for resource-constrained hardware. To address this, researchers have focused on model sparsity to skip redundant computations and improve inference speed and computing-in-memory (CIM) techniques to reduce data transfer delays and lower energy consumption. However, leveraging activation sparsity within CIM macros remains challenging due to its architectural limitations, leading to bottlenecks in this research direction. This study proposes a activation compression method based on the CIM structure and weight sparsity, along with a new approach called sparse reciprocity, to enhance model sparsity. A high-area-efficiency CIM-based accelerator chip was developed using these techniques and prior expertise. Experiments on various models and data precisions demonstrate a activation compression ratio of up to 54%, with the proposed accelerator achieving an area efficiency of 243 GOPS/mm² for the VGG16 model and 363 GOPS/mm² for the ResNet18 model.
Chu-Yao Lee, Chih-Cheng Lu, Meng-Fan Chang, Kea-Tiong Tang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 SUN: Dynamic Hybrid-Precision SRAM-Based CIM Accelerator With High Macro Utilization Using Structured Pruning Mixed-Precision Networks
abstract
Convolutional neural networks (CNNs) play a key role in many deep learning applications; however, these networks are resource-intensive. The parallel computing ability of computing-in-memory (CIM) enables high energy efficiency in artificial intelligence accelerators. When implementing a CNN in CIM, quantization and pruning are indispensable for reducing the calculation complexity and improving the efficiency of hardware calculations. Mixed-precision quantization with flexible bit widths provides a better efficiency-accuracy trade-off than fixed-precision quantization. However, CIM calculations for mixed-precision models are inefficient because the fixed capacity of CIM macros is redundant for hybrid precision distributions. To address this, we propose a software and hardware co-design SRAM-based CIM architecture called SUN, including a CIM-adaptive mixed precision joint pruning quantization algorithm and dynamic hybrid precision CNN accelerator. Three techniques are implemented in this architecture: (1) a mixed precision joint pruning algorithm for reducing the memory access and removing the redundant computing, (2) a CIM-adaptive filter-wise and paired mixed-precision quantization for improving CIM macro utilization, and (3) an SRAM-based CIM CNN accelerator in which the SRAM CIM macro is used as the processing element to support sparse and mixed-precision CNN computation with high CIM macro utilization. This architecture achieves a system area efficiency of 428.2 TOPS/mm2 and throughput of 792.2 GOPS on the CIFAR-10 dataset.
Yen-Wen Chen, Rui-Hsuan Wang, Yu-Hsiang Cheng, Chih-Cheng Lu, Meng-Fan Chang, Kea-Tiong Tang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 MARS: Multimacro Architecture SRAM CIM-Based Accelerator With Co-Designed Compressed Neural Networks
abstract
Convolutional neural networks (CNNs) play a key role in deep learning applications. However, the large storage overheads and the substantial computational cost of CNNs are problematic in hardware accelerators. Computing-in-memory (CIM) architecture has demonstrated great potential to effectively compute large-scale matrix–vector multiplication. However, the intensive multiply and accumulation (MAC) operations executed on CIM macros remain bottlenecks for further improvement of energy efficiency and throughput. To reduce computational costs, model compression is a widely studied method to shrink the model size. For implementation in a static random access memory (SRAM) CIM–based accelerator, the model compression algorithm must consider the hardware limitations of CIM macros. In this study, a software and hardware co-design approach is proposed to design MARS, a SRAM-based CIM (SRAM CIM)-based CNN accelerator that can utilize multiple SRAM CIM macros as processing units and support a sparse CNN, and an SRAM CIM-aware model compression algorithm that considers a CIM architecture to reduce the number of network parameters. With the proposed hardware software co-designed method, MARS can reach over 700 and 400 FPS for CIFAR-10 and CIFAR-100, respectively. In addition, MARS achieves 52.3 and 88.2 TOPs/W in VGG16 and ResNet18, respectively.
Syuan-Hao Sie, Jye-Luen Lee, Yi-Ren Chen, Zuo-Wei Yeh, Zhaofang Li, Chih-Cheng Lu, Chih-Cheng Hsieh, Meng-Fan Chang, Kea-Tiong Tang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2012 Learning from biological neurons to compute with electronic noise special
abstract
Biological neurons seem able to compute with noise reliably, or even to use noise to achieve probabilistic inference. This paper introduces two neuro-inspired algorithms and their implementation in the Very Large Scale Integration (VLSI). By generalising data variability with noise, the algorithms are able to classify noisy data more reliably. The VLSI implementation further demonstrates the feasibility of utilising electronic noise for stochastic computation. To exploit the intrinsic noise of transistors for computation, two transistors with enhanced and adaptable noise are further developed and modelled. These technologies would allow us to compute with noisy devices just like how the brain computes with noisy neurons.
Hsin Chen, Chih-Cheng Lu, Yi-Da Wu, Tang-Jung Chiu
ICCAD2
2010 Mapping the Diffusion Network into a stochastic system in Very Large Scale Integration
abstract
The Diffusion Network (DN) is a probabilistic model capable of recognising continuous-time, continuous-valued biomedical data. As the stochastic process of the DN is described by stochastic differential equations, realising the DN with analogue circuits is important to facilitate real-time simulation of a large network. This paper presents the translation of the DN into analogue Very Large Scale Integration (VLSI). With extensive simulation, the dynamic ranges of parameters and their representation in VLSI are identified. The VLSI circuits realising the stochastic unit of the DN are further designed and interconnected to form a stochastic system using noise to induce stochastic dynamics in VLSI. The circuit simulation demonstrate that the VLSI translation of the DN is satisfactory and the DN system is capable of using noise-induced stochastic dynamics to regenerate various types of continuous-time sequences.
Chen-Han Chien, Chih-Cheng Lu, Hsin Chen
IJCNN2
2009 Current-Mode Computation with Noise in a Scalable and Programmable Probabilistic Neural VLSI System
Chih-Cheng Lu, Hsin Chen
ICANN (1)1
2009 Minimising Contrastive Divergence with Dynamic Current Mirrors
Chih-Cheng Lu, Hsin Chen
ICANN (1)1
2007 A Scalable and Programmable Architecture for the Continuous Restricted Boltzmann Machine in VLSI
abstract
The continuous restricted Boltzmann machine (CRBM) has been attractive as a probabilistic model both useful in biomedical applications and hardware-amenable. To implement a large-scale CRBM system for real applications, the performance of the CRBM under hardware constraints must be investigated. This paper examines the effects of parameter precision on the performance of the CRBM, identifies the required precision for modelling real biomedical data, and finally proposes a scalable and programmable architecture based on which a CRBM-embedded intelligent system can be formed
Chih-Cheng Lu, C. Y. Hong, Hsin Chen
ISCAS1