VLDB 2026 Research / reviewers in the wild / expert
Meng-Fan Chang
dblp:89/6934 · also Meng-Fan Marvin Chang
· DBLP profile ↗
48ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0001-6905-6350ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 6 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SATA: Sparsity-Aware Scheduling for Selective Token AttentionabstractTransformers have become the foundation of numerous state-of-the-art AI models across diverse domains, thanks to their powerful attention mechanism for modeling long-range dependencies. However, the quadratic scaling complexity of attention poses significant challenges for efficient hardware implementation. While techniques such as quantization and pruning help mitigate this issue, selective token attention offers a promising alternative by narrowing the attention scope to only the most relevant tokens, reducing computation and filtering out noise.In this work, we propose SATA, a locality-centric dynamic scheduling scheme that proactively manages sparsely distributed access patterns from selective Query-Key operations. By reordering operand flow and exploiting data locality, our approach enables early fetch and retirement of intermediate Query/Key vectors, improving system utilization. We implement and evaluate our token management strategy in a control and compute system, using runtime traces from selective-attention-based models. Experimental results show that our method improves system throughput by up to 1.76× and boosts energy efficiency by 2.94×, while incurring minimal scheduling overhead. Zhenkun Fan, Zishen Wan, Che-Kai Liu, Ashwin Sanjay Lele, Win-San Khwa, Meng-Fan Chang, Arijit Raychowdhury |
DATE | 7 |
| 2026 | A High-Area-Efficiency Computing-in-Memory Deep Learning Accelerator Based on Weight-Sparsity Activation Compression and Sparsity ReciprocityabstractThe large number of parameters in DNNs poses significant challenges for resource-constrained hardware. To address this, researchers have focused on model sparsity to skip redundant computations and improve inference speed and computing-in-memory (CIM) techniques to reduce data transfer delays and lower energy consumption. However, leveraging activation sparsity within CIM macros remains challenging due to its architectural limitations, leading to bottlenecks in this research direction. This study proposes a activation compression method based on the CIM structure and weight sparsity, along with a new approach called sparse reciprocity, to enhance model sparsity. A high-area-efficiency CIM-based accelerator chip was developed using these techniques and prior expertise. Experiments on various models and data precisions demonstrate a activation compression ratio of up to 54%, with the proposed accelerator achieving an area efficiency of 243 GOPS/mm² for the VGG16 model and 363 GOPS/mm² for the ResNet18 model. Chu-Yao Lee, Chih-Cheng Lu, Meng-Fan Chang, Kea-Tiong Tang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | In-3-D nand Flash Computing for Vector Similarity Search Acceleration on Edge DevicesabstractVector Similarity Search (VSS) on edge devices is increasingly essential for data privacy but faces substantial latency and energy overhead due to frequent data transfers between storage and DRAM. To address these challenges, we propose the Intelligent Cognition Engine (ICE), a fully digital non-volatile in-memory computing (nvIMC) framework designed for integration with commercial 3DnandFlash. ICE avoids the use of analog-digital conversion (ADC/DAC) and reduces data movement by performing vector similarity computation directly withinnandstorage. Key architectural components include a digital page multiplier, a two’s-complement accumulator that supports signed computation with minimal circuit modification, and a hierarchical Top-N search strategy to reduce unnecessary data accesses. The proposed framework is evaluated through a combination of post-layout circuit simulations, representative silicon measurements, and system-level modeling on edge platforms. Experimental results indicate that ICE can achieve 17.8–$122.6\times $speedup and 11.5–$162\times $improvement in energy efficiency compared to conventional von Neumann-based approaches, demonstrating the feasibility and scalability of fully digital 3Dnand-based nvIMC for edge AI workloads. Han-Wen Hu, Yuan-Hao Chang 0001, Bo-Rong Lin, Huai-Mu Wang, Yung-Chun Lee, Hsiang-Pang Li, Chung Kuang Chen, Tei-Wei Kuo, Meng-Fan Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2026 | 3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge RenderingabstractDynamic 3D Gaussian splatting (3DGS) extends static 3DGS to render dynamic scenes, enabling AR/VR applications with moving objects. However, implementing dynamic 3DGS on edge devices faces challenges: (1) Loading all Gaussian parameters from DRAM for frustum culling incurs high energy costs. (2) Increased parameters for dynamic scenes elevate sorting latency and energy consumption. (3) Limited on-chip buffer capacity with higher parameters reduces buffer reuse, causing frequent DRAM access. (4) Dynamic 3DGS operations are not readily compatible with digital compute-in-memory (DCIM). These challenges hinder real-time performance and power efficiency on edge devices, leading to reduced battery life or requiring bulky batteries. To tackle these challenges, we propose algorithm-hardware co-design techniques. At the algorithmic level, we introduce three optimizations: (1) DRAM-access reduction frustum culling to lower DRAM access overhead, (2) Adaptive tile grouping to enhance on-chip buffer reuse, and (3) Adaptive interval initialization Bucket-Bitonic sort to reduce sorting latency. At the hardware level, we present a DCIM-friendly computation flow that is evaluated using the measured data from a 16 nm DCIM prototype chip. Our experimental results on Large-Scale Real-World Static/Dynamic Datasets demonstrate the ability to achieve high frame rate real-time rendering exceeding 200 frames per second (FPS) with minimal power consumption—merely 0.28 W for static Large-Scale Real-World scenes and 0.63 W for dynamic Large-Scale Real-World scenes. This work successfully addresses the significant challenges of implementing static/dynamic 3DGS technology on resource-constrained edge devices. Wei-Hsing Huang, Cheng-Jhih Shih, Jian-Wei Su, Samuel Wade Wang, Vaidehi Garg, Yuyao Kong, Jen-Chun Tien, Nealson Li, Arijit Raychowdhury, Meng-Fan Chang, Yingyan (Celine) Lin, Shimeng Yu |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2026 | Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale SystemsabstractRecent developments have introduced Kolmogorov– Arnold networks (KANs), an innovative architectural paradigm capable of replicating conventional deep neural network (DNN) capabilities while utilizing significantly reduced parameter counts through the employment of parameterized B-spline functions incorporating trainable coefficients. Nevertheless, the B-spline functional components inherent to KAN architectures introduce distinct hardware acceleration complexities. While B-spline function evaluation can be accomplished through lookup table (LUT) implementations that directly encode functional mappings, thus minimizing computational overhead, such approaches continue to demand considerable circuit infrastructure, including LUTs, multiplexers, decoders, and associated components. This work presents an algorithm-hardware co-design approach for KAN acceleration. At the algorithmic level, techniques include alignment–symmetry and PowerGap KAN hardware-aware quantization, KAN sparsity-aware mapping strategy, and circuit-level techniques include N:1 time modulation dynamic voltage input generator with analog-compute-in-memory (ACIM) circuits. Furthermore, this work conducts comprehensive evaluations on large-scale KAN networks to validate the proposed methodologies. Nonideality factors, including partial sum deviations arising from process variations, have been evaluated with the statistics measured from the TSMC 22-nm RRAM-ACIM prototype chips. Utilizing optimally determined KAN hyperparameters in conjunction with circuit optimizations implemented and evaluated at the 22-nm technology node, despite the model sizes for large-scale tasks in this work increasing by 435 K$\times $to 756 K$\times $compared to tiny-scale tasks in previous work, the area overhead increases by only 26 K$\times $to 40 K$\times $, with power consumption rising by merely$48\times $to$93\times $, while accuracy degradation remains minimal at 0.11%–0.22%, thereby demonstrating the scaling potential of our proposed architecture. Wei-Hsing Huang, Jianwei Jia, Yuyao Kong, Faaiq G. Waqar, Tai-Hao Wen, Meng-Fan Chang, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge InferenceabstractRecently, a novel model named Kolmogorov-Arnold Networks (KAN) has been proposed with the potential to achieve the functionality of traditional deep neural networks (DNNs) using orders of magnitude fewer parameters by parameterized B-spline functions with trainable coefficients. However, the B-spline functions in KAN present new challenges for hardware acceleration. Evaluating the B-spline functions can be performed by using lookup tables (LUTs) to directly map the B-spline functions, thereby reducing computational resource requirements. However, this method still requires substantial circuit resources (LUTs, MUXs, decoders, etc.). For the first time, this paper employs an algorithm-hardware co-design methodology to accelerate KAN. The proposed algorithm-level techniques include Alignment-Symmetry and PowerGap KAN hardware aware quantization, KAN sparsity aware mapping strategy, and circuit-level techniques include N:1 Time Modulation Dynamic Voltage input generator with analog-CIM (ACIM) circuits. The impact of non-ideal effects, such as partial sum errors caused by the process variations, has been evaluated with the statistics measured from the TSMC 22nm RRAM-ACIM prototype chips. With the best searched hyperparameters of KAN and the optimized circuits implemented in 22 nm node, we can reduce hardware area by 41.78x, energy by 77.97x with 3.03% accuracy boost compared to the traditional DNN hardware. Wei-Hsing Huang, Jianwei Jia, Yuyao Kong, Faaiq G. Waqar, Tai-Hao Wen, Meng-Fan Chang, Shimeng Yu |
ASP-DAC | 6 |
| 2025 | Clo-HDnn: Continual On-Device Learning Accelerator with Hyperdimensional Computing via Progressive SearchabstractClo-HDnn is an on-device learning (ODL) accelerator designed for emerging continual learning (CL) tasks. Clo-HDnn integrates hyperdimensional computing (HDC) along with low-cost Kronecker HD Encoder and weight clustering feature extraction (WCFE) to optimize accuracy and efficiency. Clo-HDnn adopts gradient-free CL to efficiently update and store the learned knowledge in the form of class hypervectors. Its dual-mode operation enables bypassing costly feature ex- traction for simpler datasets, while progressive search reduces complexity by up to $61 \%$ by encoding and comparing only partial query hypervectors. Achieving 4.66 TFLOPS/W (FE) and 3.78 TOPS/W (classifier), Clo-HDnn delivers $7.77 \times$ and $4.85 \times$ higher energy efficiency compared to SOTA ODL accelerators. Chang Eun Song, Keming Fan, Soumil Jain, Gopabandhu Hota, Haichao Yang, Leo Liu, Meng-Fan Chang, Carlos H. Diaz, Gert Cauwenberghs, Tajana Rosing, Mingu Kang |
HCS | 8 |
| 2025 | A 16.42 TOPS/mm2 Reverse Rotating Charge Sharing DRAM Compute-in-Memory Design in 28nm ProcessabstractIn recent years, the complexity of neural network (NN) models has increase continuously, resulting in a substantial increase in multiply-accumulate (MAC) operations. Therefore, the need for area- and power-efficient DRAM compute-in-memory (CIM) designs has become a pressing issue. The proposed CIM design uses the reverse rotating charge sharing (RRCS) technique to save 79% of silicon area compared to conventional designs. In addition, the use of the proposed low-threshold prediction technique can reduce analog-to-digital converter (ADC) calculations by 25% and ADC power consumption by 17%. Energy efficiency and area reduction are 10x and 1900x, respectively, higher than previous CIM designs. Tsung-Han Wu, Yu-Tse Shih, Hsin-Yung Fu, Chia Ling Ho, Wai-Chi Fang, Ke-Horng Chen, Kuo-Lin Zheng, Ying-Hsi Lin, Shian-Ru Lin, Tsung-Yen Tsai, Jui-Jen Wu, Meng-Fan Chang |
ISCAS | 12 |
| 2024 | Non-volatile Memory Technologies for Edge AI ApplicationsabstractEmbedded non-volatile memories (NVMs) hold significant promise for advancing edge AI applications by offering unique advantages in near/in-memory computing-based accelerators. This invited paper delivers an in-depth review of the use cases for NVMs, emphasizing their potential to enhance area- and energy-efficiency in edge devices. We begin by outlining the latest advancements in TSMC NVM technologies and examining several NVM-based accelerator test chips enabled by the TSMC University Shuttle Program. Additionally, we delve into the tradeoffs involved in optimizing NVM devices, explore their potential for approximate computing applications, and assess the impact of NVM non-idealities on inference accuracy. Xiaoyu Sun 0001, Win-San Khwa, Xiaochen Peng, Meng-Fan Chang, Kerem Akarvardar |
ICCAD | 4 |
| 2024 | SUN: Dynamic Hybrid-Precision SRAM-Based CIM Accelerator With High Macro Utilization Using Structured Pruning Mixed-Precision NetworksabstractConvolutional neural networks (CNNs) play a key role in many deep learning applications; however, these networks are resource-intensive. The parallel computing ability of computing-in-memory (CIM) enables high energy efficiency in artificial intelligence accelerators. When implementing a CNN in CIM, quantization and pruning are indispensable for reducing the calculation complexity and improving the efficiency of hardware calculations. Mixed-precision quantization with flexible bit widths provides a better efficiency-accuracy trade-off than fixed-precision quantization. However, CIM calculations for mixed-precision models are inefficient because the fixed capacity of CIM macros is redundant for hybrid precision distributions. To address this, we propose a software and hardware co-design SRAM-based CIM architecture called SUN, including a CIM-adaptive mixed precision joint pruning quantization algorithm and dynamic hybrid precision CNN accelerator. Three techniques are implemented in this architecture: (1) a mixed precision joint pruning algorithm for reducing the memory access and removing the redundant computing, (2) a CIM-adaptive filter-wise and paired mixed-precision quantization for improving CIM macro utilization, and (3) an SRAM-based CIM CNN accelerator in which the SRAM CIM macro is used as the processing element to support sparse and mixed-precision CNN computation with high CIM macro utilization. This architecture achieves a system area efficiency of 428.2 TOPS/mm2 and throughput of 792.2 GOPS on the CIFAR-10 dataset. Yen-Wen Chen, Rui-Hsuan Wang, Yu-Hsiang Cheng, Chih-Cheng Lu, Meng-Fan Chang, Kea-Tiong Tang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Efficient Processing of MLPerf Mobile Workloads Using Digital Compute-In-Memory MacrosabstractCompute-in-memory (CIM) has recently emerged as a promising design paradigm to accelerate deep neural network (DNN) processing. Continuously better energy and area efficiency at the macrolevel had been reported through many testchips over the last few years. However, in those macro design-oriented studies, accelerator-level considerations, such as memory accesses and processing of entire DNN workloads have not been investigated in-depth. In this article, we aim to fill this gap starting with the characteristics of our latest CIM macro fabricated with cutting-edge FinFET CMOS technology at 4-nm node. We then study, through an accelerator simulator developed in-house, three key items that would determine the efficiency of our CIM macro in the accelerator context while running MLPerf Mobile suite: 1) dataflow optimization; 2) optimal selection of CIM macro dimensions to further improve macro utilization; and 3) optimal combination of multiple CIM macros. Although there is typically a stark contrast between macro-level peak and accelerator-level average throughput and energy efficiency, the aforementioned optimizations are shown to improve the macro utilization by$3.04\times $and reduce the energy-delay product (EDP) to$0.34\times $compared to the original macro on MLPerf Mobile inference workloads. While we exploit a digital CIM macro in this study, the findings and proposed methods remain valid for other types of CIM (such as analog CIM and analog–digital–hybrid CIM) as well. Xiaoyu Sun 0001, Weidong Cao 0001, Brian Crafton, Kerem Akarvardar, Haruki Mori, Hidehiro Fujiwara, Hiroki Noguchi, Yu-Der Chih, Meng-Fan Chang, Yih Wang, Tsung-Yung Jonathan Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2024 | Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical FrameworkabstractAugmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration. Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 11 |
| 2023 | A Fully-Integrated Energy-Scalable Transformer Accelerator Supporting Adaptive Model Configuration and Word Elimination for Language Understanding on Edge DevicesabstractEfficient natural language processing on the edge is needed to interpret voice commands, which have become a standard way to interact with devices around us. Due to the tight power and compute constraints of edge devices, it is important to adapt the computation to the hardware conditions. We present a Transformer accelerator with a variable-depth adder tree to support different model dimensions, a SuperTransformer model from which Sub Transformers of various sizes can be sampled enabling adaptive model configuration, and a dedicated word elimination unit to prune redundant tokens. We achieve up to 6.9× scalability in network latency and energy between the largest and smallest Sub Transformers, under the same operating conditions. Word elimination can reduce network energy by 16%, with a 14.5% drop in F1 score. At 0.68V and 80MHz, processing a 32-length input with our custom 2-layer Transformer model for intent detection and slot filling takes 0.61ms and 1.6μJ. Zexi Ji, Hanrui Wang 0002, Miaorong Wang, Win-San Khwa, Meng-Fan Chang, Song Han 0003, Anantha P. Chandrakasan |
ISLPED | 5 |
| 2022 | Enabling High-Quality Uncertainty Quantification in a PIM Designed for Bayesian Neural NetworkabstractUncertainty quantification measures the prediction uncertainty of a neural network facing out-of-training-distribution samples. Bayesian Neural Networks (BNNs) can provide high-quality uncertainty quantification by introducing specific noise to the weights during inference. To accelerate BNN inference, ReRAM processing-in-memory (PIM) architecture is a competitive solution to provide both high-efficient computing and in-situ noise generation at the same time. However, there normally exists a huge gap between the generated noise in PIM hardware and that required by a BNN model. We demonstrate that the quality of uncertainty quantification is substantially degraded due to this gap. To solve this problem, we propose a holistic framework called W2W-PIM. We first introduce an efficient method to generate noise in ReRAM PIM design according to the demand of a BNN model. In addition, the PIM architecture is carefully modified to enable the noise generation and evaluate uncertainty quality. Moreover, a calibration unit is further introduced to reduce the noise gap caused by imperfection of the noise model. Comprehensive evaluation results demonstrate that W2W-PIM framework can achieve high-quality uncertainty quantification and high energy-efficiency at the same time. Bingzhe Wu, Guangyu Sun 0003, Zhe Zhang 0006, Zhihang Yuan, Runsheng Wang, Ru Huang 0001, Dimin Niu, Hongzhong Zheng, Zhichao Lu, Meng-Fan Chang, Tianchan Guan, Xin Si |
HPCA | 12 |
| 2022 | Reliable Computing of ReRAM Based Compute-in-Memory Circuits for AI Edge DevicesabstractCompute-in-memory macros based on non-volatile memory (nvCIM) are a promising approach to break through the memory bottleneck for artificial intelligence (AI) edge devices; however, the development of these devices involves unavoidable tradeoffs between reliability, energy efficiency, computing latency, and readout accuracy. This paper outlines the background of ReRAM-based nvCIM as well as the major challenges in its further development, including process variation in ReRAM devices and transistors and the small signal margins associated with variation in input-weight patterns. This paper also investigates the error model of a nvCIM macro, and the correspondent degradation of inference accuracy as a function of error model when using nvCIM macros. Finally, we summarize recent trends and advances in the development of reliable ReRAM-based nvCIM macro. Meng-Fan Chang, Je-Ming Hung, Ping-Cheng Chen, Tai-Hao Wen |
ICCAD | 1 |
| 2022 | ICE: An Intelligent Cognition Engine with 3D NAND-based In-Memory Computing for Vector Similarity Search AccelerationabstractVector similarity search (VSS) for unstructured vectors generated via machine learning methods is a promising solution for many applications, such as face search. With increasing awareness and concern about data security requirements, there is a compelling need to store data and process VSS applications locally on edge devices rather than send data to servers for computation. However, the explosive amount of data movement from NAND storage to DRAM across memory hierarchy and data processing of the entire dataset consume enormous energy and require long latency for VSS applications. Specifically, edge devices with insufficient DRAM capacity will trigger data swap and deteriorate the execution performance. To overcome this crucial hurdle, we propose an intelligent cognition engine (ICE) with cognitive 3D NAND, featuring non-volatile in-memory computing (nvIMC) to accelerate the processing, suppress the data movement, and reduce data swap between the processor and storage. This cognitive 3D NAND features digital nvIMC techniques (i. e., ADClDAC-free approach), high-density 3D NAND, and compatibility with standard 3D NAND products with minor modifications. To facilitate parallel INT8/INT4 vector-vector multiplication (VVM) and mitigate the reliability issue of 3D NAND, we develop a bit-error-tolerance data encoding and a two’s complement-based digital accumulator. VVM can support similarity computations (e.g., cosine similarity and Euclidean distance), which are required to search “the most similar data” right where they are stored. In addition, the proposed solution can be realized on edge storage products, e.g., embedded Multi-Media Card (eMMC). The measured and simulated results on real 3D NAND chips show that ICE enhances the system execution time by $17\times to 95\times$ and energy efficiency by $11\times to 140\times$, compared to traditional von Neumann approaches using state-of-the-art edge systems with MobileFaceNet on CASIA-WebFace dataset. To the best of our knowledge, this work demonstrates the first 3D NAND-based digital nvIMC technique with measured silicon data. Han-Wen Hu, Wei-Chen Wang 0002, Yuan-Hao Chang 0001, Yung-Chun Lee, Bo-Rong Lin, Huai-Mu Wang, Yen-Po Lin, Chong-Ying Lee, Tzu-Hsiang Su, Chih-Chang Hsieh, Chia-Ming Hu, Yi-Ting Lai, Chung Kuang Chen, Han-Sung Chen, Hsiang-Pang Li, Tei-Wei Kuo, Meng-Fan Chang, Keh-Chung Wang, Chun-Hsiung Hung, Chih-Yuan Lu |
MICRO | 18 |
| 2022 | MARS: Multimacro Architecture SRAM CIM-Based Accelerator With Co-Designed Compressed Neural NetworksabstractConvolutional neural networks (CNNs) play a key role in deep learning applications. However, the large storage overheads and the substantial computational cost of CNNs are problematic in hardware accelerators. Computing-in-memory (CIM) architecture has demonstrated great potential to effectively compute large-scale matrix–vector multiplication. However, the intensive multiply and accumulation (MAC) operations executed on CIM macros remain bottlenecks for further improvement of energy efficiency and throughput. To reduce computational costs, model compression is a widely studied method to shrink the model size. For implementation in a static random access memory (SRAM) CIM–based accelerator, the model compression algorithm must consider the hardware limitations of CIM macros. In this study, a software and hardware co-design approach is proposed to design MARS, a SRAM-based CIM (SRAM CIM)-based CNN accelerator that can utilize multiple SRAM CIM macros as processing units and support a sparse CNN, and an SRAM CIM-aware model compression algorithm that considers a CIM architecture to reduce the number of network parameters. With the proposed hardware software co-designed method, MARS can reach over 700 and 400 FPS for CIFAR-10 and CIFAR-100, respectively. In addition, MARS achieves 52.3 and 88.2 TOPs/W in VGG16 and ResNet18, respectively. Syuan-Hao Sie, Jye-Luen Lee, Yi-Ren Chen, Zuo-Wei Yeh, Zhaofang Li, Chih-Cheng Lu, Chih-Cheng Hsieh, Meng-Fan Chang, Kea-Tiong Tang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | DL-RSIM: A Reliability and Deployment Strategy Simulation Framework for ReRAM-based CNN AcceleratorsabstractMemristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. In addition, due to the hardware constraints, the way to deploy neural network models on memristor crossbar arrays affects the computation parallelism and communication overheads. To enable reliable and energy-efficient memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit/device properties on the inference accuracy and the influence of different deployment strategies on performance and energy consumption. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. A rich set of reliability impact factors and deployment strategies are explored by DL-RSIM, and it can be incorporated with any deep learning neural networks implemented by TensorFlow. Using several representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and energy-efficient deployment strategies and develop optimization techniques accordingly. Hsiang-Yun Cheng, Chia-Lin Yang, Meng-Yao Lin, Kai Lien, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang, Yen-Ting Tsou, Chin-Fu Nien |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2021 | Sparsity-Aware Clamping Readout Scheme for High Parallelism and Low Power Nonvolatile Computing-in-Memory Based on Resistive MemoryabstractThe input parallelism of resistive memory (RRAM) based nonvolatile computing-in-memory (nvCIM) structure is limited by the signal margin as well as the readout precision. In this work, we propose a sparsity-aware clamping (SAC) scheme and its circuit implementation for nvCIM by co-design of circuit and algorithm. It can adaptively tune the quantized range and resolution of the readout circuit according to the degree of sparsity in neural network models. As a result, the SAC scheme can effectively increase the input parallelism of nvCIMs without incurring degradation on the signal margin or increasing the hardware cost for analogue readout. A case study on processing a multi-layer perceptron (MLP) model with the proposed nvCIM structure shows that the SAC scheme can improve the throughput by 2 times and increase the energy efficiency by 25.35% with negligible inference accuracy loss. Linfang Wang, Wang Ye, Junjie An, Chunmeng Dou, Qi Liu 0010, Meng-Fan Chang, Ming Liu 0022 |
ISCAS | 6 |
| 2021 | A 40nm 1Mb 35.6 TOPS/W MLC NOR-Flash Based Computation-in-Memory Structure for Machine LearningabstractComputation-in-memory (CIM) is a feasible method to overcome "Von-Neumann bottleneck" with high throughput and energy efficiency. In this paper, we proposed a 1Mb Multi-Level (MLC) NOR Flash based CIM (MLFlash- CIM) structure with 40nm technology node. A multi-bit readout circuit was proposed to realize adaptive quantization, which comprises a current interface circuit, a multi-level analog shift amplifier (AS-Amp) and an 8-bit SAR-ADC. When applied to a modified VGG-16 Network with 16 layers, the proposed MLFlash-CIM can achieve 92.73% inference accuracy under CIFAR-10 dataset. This CIM structure also achieved a peak throughput of 3.277 TOPS and an energy efficiency of 35.6 TOPS/W with 4-bit multiplication and accumulation (MAC) operations. Sitao Zeng, Zhiguo Zhu, Zhaolong Qin, Chen Wang 0130, Jingjing Li 0001, Sanfeng Zhang 0001, Yajuan He, Chunmeng Dou, Xin Si, Meng-Fan Chang, Qiang Li 0021 |
ISCAS | 11 |
| 2021 | In-memory Learning with Analog Resistive Switching Memory: A Review and PerspectiveabstractIn this article, we review the existing analog resistive switching memory (RSM) devices and their hardware technologies for in-memory learning, as well as their challenges and prospects. Since the characteristics of the devices are different for in-memory learning and digital memory applications, it is important to have an in-depth understanding across different layers from devices and circuits to architectures and algorithms. First, based on a top-down view from architecture to devices for analog computing, we define the main figures of merit (FoMs) and perform a comprehensive analysis of analog RSM hardware including the basic device characteristics, hardware algorithms, and the corresponding mapping methods for device arrays, as well as the architecture and circuit design considerations for neural networks. Second, we classify the FoMs of analog RSM devices into two levels. Level 1 FoMs are essential for achieving the functionality of a system (e.g., linearity, symmetry, dynamic range, level numbers, fluctuation, variability, and yield). Level 2 FoMs are those that make a functional system more efficient and reliable (e.g., area, operational voltage, energy consumption, speed, endurance, retention, and compatibility with back-end-of-line processing). By constructing a device-to-application simulation framework, we perform an in-depth analysis of how these FoMs influence in-memory learning and give a target list of the device requirements. Lastly, we evaluate the main FoMs of most existing devices with analog characteristics and review optimization methods from programming schemes to materials and device structures. The key challenges and prospects from the device to system level for analog RSM devices are discussed. Bin Gao 0006, Jianshi Tang, Meng-Fan Chang, Xiaobo Sharon Hu, Jan Van der Spiegel, He Qian, Huaqiang Wu |
Proc. IEEE | 5 |
| 2021 | Challenges and Trends of SRAM-Based Computing-In-Memory for AI Edge DevicesabstractWhen applied to artificial intelligence edge devices, the conventionally von Neumann computing architecture imposes numerous challenges (e.g., improving the energy efficiency), due to the memory-wall bottleneck involving the frequent movement of data between the memory and the processing elements (PE). Computing-in-memory (CIM) is a promising candidate approach to breaking through this so-called memory wall bottleneck. SRAM cells provide unlimited endurance and compatibility with state-of-the-art logic processes. This paper outlines the background, trends, and challenges involved in the further development of SRAM-CIM macros. This paper also reviews recent silicon-verified SRAM-CIM macros designed for logic and multiplication-accumulation (MAC) operations. Chuan-Jia Jhang, Cheng-Xin Xue, Je-Min Hung, Fu-Chun Chang, Meng-Fan Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | A Two-way SRAM Array based Accelerator for Deep Neural Network On-chip TrainingabstractOn-chip training of large-scale deep neural networks (DNNs) is challenging due to computational complexity and resource limitation. Compute-in-memory (CIM) architecture exploits the analog computation inside the memory array to speed up the vectormatrix multiplication (VMM) and alleviate the memory bottleneck. However, existing CIM prototype chips, in particular, SRAM-based accelerators target at implementing low-precision inference engine only. In this work, we propose a two-way SRAM array design that could perform bi-directional in-memory VMM with minimum hardware overhead. A novel solution of signed number multiplication is also proposed to handle the negative input in backpropagation. We taped-out and validated proposed two-way SRAM array design in TSMC 28nm process. Based on the silicon measurement data on CIM macro, we explore the hardware performance for the entire architecture for DNN on-chip training. The experimental data shows that proposed accelerator can achieve energy efficiency of ~3.2 TOPS/W, >1000 FPS and >300 FPS for ResNet and DenseNet training on ImageNet, respectively. Hongwu Jiang, Shanshi Huang, Xiaochen Peng, Jian-Wei Su, Yen-Chi Chou, Wei-Hsing Huang, Ta-Wei Liu, Ruhui Liu, Meng-Fan Chang, Shimeng Yu |
DAC | 9 |
| 2019 | Monolithic-3D Integration Augmented Design Techniques for Computing in SRAMsabstractIn-memory computing has emerged as a promising solution to address the logic-memory performance gap. We propose design techniques using monolithic-3D integration to achieve reliable multirow activation which in turn help in computation as part of data readout. Our design is 1.8x faster than the existing techniques for Boolean computations. We quantitatively show no impact to cell stability when multiple rows are activated and thereby requiring no extra hardware for maintaining the cell stability during computations. In-memory digital to analog conversion technique is proposed using a 3D-CAM primitive. The design utilizes relatively low strength layer-2 transistors effectively and provides 7x power savings when compared with a specialized converter in-memory. Lastly, we present a linear classifier system by making use the above-mentioned techniques which is 47x faster while computing vector matrix multiplication using a dedicated hardware engine. Srivatsa Rangachar Srinivasa, Wei-Hao Chen, Yung-Ning Tu, Meng-Fan Chang, Jack Sampson, Narayanan Vijaykrishnan |
ISCAS | 4 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2018 | Parallelizing SRAM arrays with customized bit-cell for binary neural networksabstractRecent advances in deep neural networks (DNNs) have shown Binary Neural Networks (BNNs) are able to provide a reasonable accuracy on various image datasets with a significant reduction in computation and memory cost. In this paper, we explore two BNNs: hybrid BNN (HBNN) and XNOR-BNN, where the weights are binarized to +1/-1 while the neuron activations are binarized to 1/0 and +1/-1 respectively. Two SRAM bit cell designs are proposed, namely, 6T SRAM for HBNN and customized 8T SRAM for XNOR-BNN. In our design, the high-precision multiply-and-accumulate (MAC) is replaced by bitwise multiplication for HBNN or XNOR for XNOR-BNN plus bit-counting operations. To parallelize the weighted sum operation, we activate multiple word lines in the SRAM array simultaneously and digitize the analog voltage developed along the bit line by a multi-level sense amplifier (MLSA). In order to partition the large matrices in DNNs, we investigate the impact of sensing bit-levels of MLSA on the accuracy degradation for different sub-array sizes and propose using the nonlinear quantization technique to mitigate the accuracy degradation. With 64×64 sub-array size and 3-bit MLSA, HBNN and XNOR-BNN architectures can minimize the accuracy degradation to 2.37% and 0.88%, respectively, for an inspired VGG-16 network on the CIFAR-10 dataset. Design space exploration of SRAM based synaptic architectures with the conventional row-by-row access scheme and our proposed parallel access scheme are also performed, showing significant benefits in the area, latency and energy-efficiency. Finally, we have successfully taped-out and validated the proposed HBNN and XNOR-BNN designs in TSMC 65 nm process with measured silicon data, achieving energy-efficiency >100 TOPS/W for HBNN and >50 TOPS/W for XNOR-BNN. Rui Liu 0005, Xiaochen Peng, Xiaoyu Sun 0001, Win-San Khwa, Xin Si, Jia-Jing Chen, Jia-Fang Li, Meng-Fan Chang, Shimeng Yu |
DAC | 8 |
| 2018 | DL-RSIM: a simulation framework to enable reliable ReRAM-based accelerators for deep learningabstractMemristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. To enable reliable memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit and device properties on the inference accuracy. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. DL-RSIM simulates the error rates of every sum-of-products computation in the memristor-based accelerator and injects the errors in the targeted TensorFlow-based neural network model. A rich set of reliability impact factors are explored by DL-RSIM, and it can be incorporated with any deep learning neural network implemented by TensorFlow. Using three representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and develop reliability optimization techniques. Meng-Yao Lin, Hsiang-Yun Cheng, Tzu-Hsien Yang, I-Ching Tseng, Chia-Lin Yang, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang |
ICCAD | 10 |
| 2018 | A 2-GHz Direct Digital Frequency Synthesizer Based on LUT and RotationabstractThis paper proposes a direct digital frequency synthesizer (DDFS) based on Lookup-Table-Rotation (LUT-ROT) architecture. The DDFS takes the advantages of coarse-fine LUT and pipelined rotation units to achieve both high-speed and high-resolution. Based on the analysis of approximation error, the trade-off between accuracy and memory usage is achieved, which leads to the lowest amplitude of noise and the smallest size of LUT. Experimental results show that the DDFS achieves 11.7mW/GHz power consumption and 96dBc SFDR in 2-GHz clock frequency. Yixiong Yang, Zhibo Wang 0004, Meng-Fan Chang, Mon-Shu Ho, Huazhong Yang, Yongpan Liu |
ISCAS | 4 |
| 2018 | A Monolithic-3D SRAM Design with Enhanced Robustness and In-Memory Computation SupportabstractWe present a novel 3D-SRAM cell using a Monolithic 3D integration (M3D-IC) technology for realizing both robustness and In-memory Boolean logic compute support. The proposed two-layer design makes use of additional transistors over the SRAM layer to enable assist techniques as well as provide logic functions (such as AND/NAND, OR/NOR, XNOR/XOR) without degrading cell density. Through analysis, we provide insights into the benefits provided by three memory assist and two logic modes and evaluate the energy efficiency of our proposed design. Assist techniques improve SRAM read stability by 2.2x and increase the write margin by 17.6%, while staying within the SRAM footprint. By virtue of increased robustness, the cell enables seamless operation at lower supply voltages and thereby ensures energy efficiency. Energy Delay Product (EDP) reduces by 1.6x over standard 6T SRAM with a faster data access. Transistor placement and their biasing technique in layer-2 enables In-memory bitwise Boolean computation. When computing bulk In-memory operations, 6.5x energy savings is achieved as compared to computing outside the memory system. Srivatsa Rangachar Srinivasa, Akshay Krishna Ramanathan, Xueqing Li 0002, Wei-Hao Chen, Fu-Kuo Hsueh, Chih-Chao Yang, Chang-Hong Shen, Jia-Min Shieh, Sumeet Kumar Gupta, Meng-Fan Chang, Swaroop Ghosh, Jack Sampson, Narayanan Vijaykrishnan |
ISLPED | 10 |
| 2018 | A Dual-Data Line Read Scheme for High-Speed Low-Energy Resistive Nonvolatile MemoriesabstractResistance-based memory devices are considered as a strong candidate for next-generation nonvolatile memories as well as potential application in high density embedded cache. These devices can be programmed to different resistance states by applying electrical bias. Read operation then senses the programmed state by discharging a shared data line through the memory cell. However, as the dimensions of the device scale down, its resistance increases and the distribution widens, which leads to reduced sensing margins and degradation in read performance. In the proposed dual-data line (DDL) read scheme, we recycle current flowing through the memory cell during the read operation to create an additional voltage swing on a secondary data line, and combine it with the signal on the original data line to reduce sensing time and energy. Performance comparison with the conventional read scheme is derived theoretically by using circuit analysis and verified through simulation on an array critical path constructed in 65-nm technology. Results show that the DDL scheme can improve sensing margins by an average of 86%, which translates to a sensing time reduction of 47%, across various device conditions. For the same read performance, the sensing energy is decreased by 48%. Hochul Lee, Farbod Ebrahimi, Bonnie Lam, Wei-Hao Chen, Meng-Fan Chang, Pedram Khalili Amiri, Kang-Lung Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | Compact 3-D-SRAM Memory With Concurrent Row and Column Data Access Capability Using Sequential Monolithic 3-D IntegrationabstractThis paper proposes the use of monolithic 3-D integration technology in designing a novel two-layer 3-D-static random access memory (3-D-SRAM) cell in standard 6T-SRAM footprint. The proposed 3-D-SRAM cell is capable of data access from both the layers. The cell is designed to retrieve row-wise and column-wise data concurrently from the memory array. This memory design can cater to applications and workloads requiring multidimensional data access for enhancing system performance. The novel 3-D layout technique ensures the same footprint as a 6T-SRAM cell despite enhancing the functionality. The design ensures no degradation in the cell stability and performance. Voltage reduction in layer-2 provides 5.4× power savings during column-wise data access. We analyze the implications of employing the proposed SRAM to achieve efficient data access for integral image algorithm. We obtain 2.15× savings in access time and 7.81% access energy savings while accessing data from a 32-kB memory array to compute integral image for a region of 32 rows and 16 columns. Srivatsa Rangachar Srinivasa, Xueqing Li 0002, Meng-Fan Chang, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Data Backup Optimization for Nonvolatile SRAM in Energy Harvesting Sensor NodesabstractNonvolatile static random access memory (nvSRAM) has been widely investigated as a promising on-chip memory architecture in energy harvesting sensor nodes, due to zero standby power, resilience to power failures, and fast read/write operations. However, conventional approaches back up all data from static random access memory into nonvolatile memory when power failures happen. It leads to significant energy overhead and peak inrush current, which has a negative impact on the system performance and circuit reliability. This paper proposes a holistic data backup optimization to mitigate these problems in nvSRAM, consisting of a partial backup algorithm and a run-time adaptive write policy. A statistic dead-block predictor is employed to achieve dead block identification with trivial hardware overhead. An adaptive policy is used to switch between write-back and write-through strategy to reduce the rollback induced by backup failures. Experimental results show that the proposed scheme improves the performance by 4.6% on average while the backup power consumption and the inrush current are reduced by 38.1% and 54% on average compared to the full backup scheme. What is more, the backup capacitor size for energy buffer can be reduced by 40% on average under the same performance constraint. Yongpan Liu, Jinshan Yue, Hehe Li, Qinghang Zhao, Mengying Zhao, Chun Jason Xue, Guangyu Sun 0003, Meng-Fan Chang, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2017 | Low-VDD Operation of SRAM Synaptic Array for Implementing Ternary Neural NetworkabstractFor Internet of Things (IoT) edge devices, it is very attractive to have the local sensemaking capability instead of sending all the data back to the cloud for information processing. For image pattern recognition, neuro-inspired machine learning algorithms have demonstrated enormous powerfulness. To effectively implement learning algorithms on-chip for IoT edge devices, on-chip synaptic memory architectures have been proposed to implement the key operations such as weighted-sum or matrix-vector multiplication. In this paper, we proposed a low-power design of static random access memory (SRAM) synaptic array for implementing a low-precision ternary neural network. We experimentally demonstrated that the supply voltage (VDD) of the SRAM array could be aggressively reduced to a level, where the SRAM cell is susceptible to bit failures. The testing results from 65-nm SRAM chips indicate that VDD could be reduced from the nominal 1-0.55 V (or 0.5 V) with a bit error rate ~0.23% (or ~1.56%), which only introduced ~0.08% (or ~1.68%) degradation in the classification accuracy. As a result, the power consumption could be reduced by more than 8× (or 10×). Xiaoyu Sun 0001, Rui Liu 0005, Yi-Ju Chen, Hsiao-Yun Chiu, Wei-Hao Chen, Meng-Fan Chang, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2017 | A Flexible Wildcard-Pattern Matching Accelerator via Simultaneous Discrete Finite AutomataabstractRegular expression matching becomes indispensable elements of Internet of Things network security. However, traditional ternary content addressable memory (TCAM) search engine is unable to handle patterns with wildcards, as it precisely tracks only one active state with single transition. This paper proposes a promising simultaneous pattern matching methodology for wildcard patterns by two separated engines to represent discrete finite automata. A key preprocessing to encode possible postfix pattern by a unique key ensures that follow-up patterns can accurately traverse all possible matches with limited hardware resources. This approach is practical and scalable for achieving good performance and low space consumption in network security, and it can be applicable to any regular expressions even with multiwildcard patterns. The experimental results demonstrate that this scheme can efficiently and accurately recognize wildcard patterns by simultaneously tracking only two active states. By adopting SRAM TCAM in the proposed architecture, the energy consumption is reduced to around 39%, compared with the energy consumption using a computing system that contains a large memory lookup and comparison overhead. Hsiang-Jen Tsai, Chien-Chih Chen, Yin-Chi Peng, Ya-Han Tsao, Yen-Ning Chiang, Wei-Cheng Zhao, Meng-Fan Chang, Tien-Fu Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2017 | Energy-Efficient TCAM Search Engine Design Using Priority-Decision in Memory TechnologyabstractTernary content-addressable memory (TCAM)-based search engines generally need a priority encoder (PE) to select the highest priority match entry for resolving the multiple match problem due to the don't care (X) features of TCAM. In contemporary network security, TCAM-based search engines are widely used in regular expression matching across multiple packets to protect against attacks, such as by viruses and spam. However, the use of PE results in increased energy consumption for pattern updates and search operations. Instead of using PEs to determine the match, our solution is a three-phase search operation that utilizes the length information of the matched patterns to decide the longest pattern match data. This paper proposes a promising memory technology called priority-decision in memory (PDM), which eliminates the need for PEs and removes restrictions on ordering, implying that patterns can be stored in an arbitrary order without sorting their lengths. Moreover, we present a sequential input-state (SIS) scheme to disable the mass of redundant search operations in state segments on the basis of an analysis distribution of hex signatures in a virus database. Experimental results demonstrate that the PDM-based technology can improve update energy consumption of nonvolatile TCAM (nvTCAM) search engines by 36%-67%, because most of the energy in these search engines is used to reorder. By adopting the SIS-based method to avoid unnecessary search operations in a TCAM array, the search energy reduction is around 64% of nvTCAM search engines. Hsiang-Jen Tsai, Keng-Hao Yang, Yin-Chi Peng, Chien-Chen Lin, Ya-Han Tsao, Meng-Fan Chang, Tien-Fu Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2016 | Nonvolatile memory design based on ferroelectric FETsabstractFerroelectric FETs (FEFETs) offer intriguing possibilities for the design of low power nonvolatile memories by virtue of their three-terminal structure coupled with the ability of the ferroelectric (FE) material to retain its polarization in the absence of an electric field. Utilizing the distinct features of FEFETs, we propose a 2-transistor (2T) FEFET-based nonvolatile memory with separate read and write paths. With proper co-design at the device, cell and array levels, the proposed design achieves non-destructive read and lower write power at iso-write speed compared to standard FERAM. In addition, the FEFET-based memory exhibits high distinguishability with six orders of magnitude difference in the read currents corresponding to the two states. Comparative analysis based on experimentally calibrated models shows significant improvement of access energy-delay. For example, at a fixed write time of 550ps, the write voltage and energy are 58.5% and 67.7% lower than FERAM, respectively. These benefits are achieved with 2.4 times the area overhead. Further exploration of the proposed FEFET memory in energy harvesting nonvolatile processors shows an average improvement of 27% in forward progress over FERAM. Sumitha George, Kaisheng Ma, Ahmedullah Aziz, Xueqing Li 0002, Asif Islam Khan, Sayeef S. Salahuddin, Meng-Fan Chang, Suman Datta, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
DAC | 7 |
| 2016 | Designs of emerging memory based non-volatile TCAM for Internet-of-Things (IoT) and big-data processing: A 5T2R universal cellabstractMany search engines or filters for the internet-of-things and big-data employ ternary content-addressable-memory (TCAM) to suppress power consumption in the transmission of data between end-devices and servers. Nonvolatile TCAMs (nvTCAM) are designed to achieve zero standby power with smaller area overhead and faster power off/on operations than those found in conventional TCAM+NVM 2-macro schemes. In this paper, we discuss the challenges involved in the design of nvTCAMs and propose a universal 5T2R nvTCAM cell with tolerance for the various R-ratios and write parameters associated with emerging memory devices. A 128×64b-nvTCAM macro was fabricated using HfO ReRAM and a 90nm-CMOS process for concept verification. Meng-Fan Chang, Ching-Hao Chuang, Yen-Ning Chiang, Shyh-Shyuan Sheu, Chia-Chen Kuo, Hsiang-Yun Cheng, Jack Sampson, Mary Jane Irwin |
ISCAS | 1 |
| 2016 | Design of nonvolatile processors and applicationsabstractEnergy harvesting is under intense investigation as a promising substitute for batteries. However, given the erratic nature of ambient energy sources, temporary status in conventional CMOS circuits can be lost upon a sudden power outage. Taking advantage of emerging nonvolatile memories (NVMs), nonvolatile processor (NVP) backs up system contexts when power failure occurs, and recalls pre-stored data on resumption. It has become a hot topic for the capability to survive power variations and to guarantee forward progress on computation tasks. This paper acts as a brief guide introducing the concepts, the current status, the challenges and opportunities, as well as emerging applications of NVPs. Through this paper, we expect to help researchers who are new in this area better understand the development trends and future research prospects. We also hope to attract researchers to join and explore more innovative applications of NVP. Fang Su, Zhibo Wang 0004, Jinyang Li 0002, Meng-Fan Chang, Yongpan Liu |
VLSI-SoC | 4 |
| 2015 | Read circuits for resistive memory (ReRAM) and memristor-based nonvolatile LogicsabstractResistive memory device (Memristor) is one of the candidates for energy-efficient nonvolatile memory and nonvolatile logics (nvLogics) in the applications of wearable devices, Internet of Things (IoT), cloud computing, and big-data processing. However, resistive RAM (ReRAM) and memristor-based nvLogics suffer limited performance and low yield due to process variations in transistors and resistance of memristor. This presentation discusses the design challenges in read circuits for high-speed, area-efficient, and low-voltage ReRAM and nvLogics. Memristor-based nvLogics, such as nonvolatile-SRAM (nvSRAM), nonvolatile flip-flops (nvFF), and nonvolatile TCAM (nvTCAM) are included in this presentation. Several silicon-verified solutions on read scheme and sense amplifiers are also discussed in this presentation. Meng-Fan Chang, Chien-Chen Lin, Mon-Shu Ho, Ping-Cheng Chen, Chia-Chen Kuo, Ming-Pin Chen, Pei-Ling Tseng, Tzu-Kun Ku, Chien-Fu Chen, Kai-Shin Li, Jia-Min Shieh |
ASP-DAC | 1 |
| 2015 | Ambient energy harvesting nonvolatile processors: from circuit to systemabstractEnergy harvesting is gaining more and more attentions due to its characteristics of ultra-long operation time without maintenance. However, frequent unpredictable power failures from energy harvesters bring performance and reliability challenges to traditional processors. Nonvolatile processors are promising to solve such a problem due to their advantage of zero leakage and efficient backup and restore operations. To optimize the nonvolatile processor design, this paper proposes new metrics of nonvolatile processors to consider energy harvesting factors for the first time. Furthermore, we explore the nonvolatile processor design from circuit to system level. A prototype of energy harvesting nonvolatile processor is set up and experimental results show that the proposed performance metric meets the measured results by less than 6.27% average errors. Finally, the energy consumption of nonvolatile processor is analyzed under different benchmarks. Yongpan Liu, Hehe Li, Xueqing Li 0002, Kaisheng Ma, Shuangchen Li, Meng-Fan Chang, Jack Sampson, Yuan Xie 0001, Jiwu Shu, Huazhong Yang |
DAC | 8 |
| 2015 | Energy-efficient non-volatile TCAM search engine design using priority-decision in memory technology for DPIabstractTCAM-based search engines are widely used in regular expression matching across multiple packets. However, the use of priority encoder results in increased energy consumption of pattern updates and search operations. This work, proposes a promising memory technology, called Priority-Decision in Memory (PDM), which eliminates the need for priority encoders and removes restrictions on ordering, meaning that patterns can be stored in an arbitrary order without sorting their lengths. Moreover, we present a Sequential Input-State Search (SIS) scheme to disable the mass of redundant search operations in state segments, based on the analysis distribution of hex signatures in a virus database. Experimental results demonstrate that PDM-based technology can improve update energy consumption of nvTCAM search engines by 36%~67% because most of the energy in the latter is used to reorder. By adopting the SIS-based method to avoid unnecessarily search operations in a TCAM array, the search energy reduction is around 64% of nvTCAM search engines. Hsiang-Jen Tsai, Keng-Hao Yang, Yin-Chi Peng, Chien-Chen Lin, Ya-Han Tsao, Meng-Fan Chang, Tien-Fu Chen |
DAC | 6 |
| 2015 | An energy efficient backup scheme with low inrush current for nonvolatile SRAM in energy harvesting sensor nodes
Hehe Li, Yongpan Liu, Qinghang Zhao, Yizi Gu, Xiao Sheng, Guangyu Sun 0003, Chao Zhang 0007, Meng-Fan Chang, Huazhong Yang |
DATE | 8 |
| 2014 | Leveraging Data Lifetime for Energy-Aware Last Level Non-Volatile SRAM Caches using Redundant Store EliminationabstractNVM has commonly been used to address increasingly large last-level caches (LLCs) requirements by reducing leakage. However, frequent data-writing operations result in increased energy consumption. In this context, a promising memory technology, Non-volatile SRAM (nvSRAM), enables normal and standby operation modes which can be used to store various types of data. However, nvSRAM suffers from high dynamic energy usage due to frequent switching between operation modes. In this paper, we propose a redundant store elimination (RSE) scheme which, on average, discards 94% of needless bit-write operations. Moreover, we present a retention-aware cache management policy to reduce data updates of cache blocks, based on the correlation between data lifetime and cache types. Experimental results demonstrate that our proposal can improve energy consumption of SRAM-based and RRAM-based LLCs by 57% and 31%, respectively. Hsiang-Jen Tsai, Chien-Chih Chen, Keng-Hao Yang, Ting-Chin Yang, Li-Yue Huang, Ching-Hao Chuang, Meng-Fan Chang, Tien-Fu Chen |
DAC | 7 |
| 2013 | Special session 4C: Hot topic 3D-IC design and testabstractThree-dimensional (3D) integration using through silicon via (TSV) is a promising approach to coping with the challenges faced by the current 2D technology. A TSV-based 3D IC is implemented by stacking multiple dies which are vertically connected by TSVs. This may shorten the global interconnects of a 3D IC and greatly improve its performance and power consumption. High bandwidth is achieved by the increase of IO channels provided by the TSVs, which also reduce the unnecessary waste of energy during data movement. In addition, the 3D integration technology shows other advantages over 2D technology, such as high functionality, heterogeneous integration, small form factor, etc. However, there are still challenges that need to be tackled before volume production of 3D ICs using TSV becomes possible, including technology scalability, quality and reliability, yield, thermal management, equipment and infrastructure, and costs. To demonstrate the feasibility of 3D-IC technologies, many academic and industrial institutes have been working on various test vehicles, especially in recent years. An increasing attention also has been attracted around the world in the semiconductor industry by the development of related technologies. In this special session, we will discuss the evolutionary efforts toward the realization of 3-D ICs. As memory dies need to be integrated in most 3D-IC system, we will also address the challenges in the design and test of 3D memories against cross-layer process, voltage and temperature variations while suppressing thermal effect and power consumption, etc. We will share our experiences and show results from some of the test vehicles we have worked on, including processor/memory stacks, analog/logic stacks, logic/logic stacks, etc. We will show a 1,024-bit wide bus chip-to-chip interconnection using 40×40 fine-pitch TSVs to demonstrate ultra-low-power operation. For memories, we will present a 3D-RAM structure using small voltage-swing TSVs, vertical-device-stacking nonvolatile-SRAM and ReRAM, and a 3D vertical-gate NAND flash. To address the challenge of reliability and yield, we will also discuss important test techniques. Demonstration of benefits provided by 3D integration technology will also be shown. Last but not least, we will describe our development plan regarding various types of die stacking using heterogeneous process integration, especially processor/memory stacking that is widely believed to be a key technology in future generations of smart handheld devices. Jin-Fu Li 0001, Cheng-Wen Wu, Masahiro Aoyagi, Meng-Fan Chang, Ding-Ming Kwai |
VTS | 4 |
| 2012 | Endurance-aware circuit designs of nonvolatile logic and nonvolatile sram using resistive memory (memristor) deviceabstractThe use of low voltage circuits and power-off mode help to reduce the power consumption of chips. Non-volatile logic (nvLogic) and nonvolatile SRAM (nvSRAM) enable a chip to preserve its key local states and data, while providing faster power-on/off speeds than those available with conventional two-macro schemes. Resistive memory (memristor) devices feature fast write speed and low write power. Applying memristors to nvLogic and nvSRAMs not only enables chips to achieve low power consumption for store operations, but also achieve fast power-on/off processes and reliable operation even in the event of sudden power failure. However, current memristor devices suffer from limited endurance, which influences the design of the circuit structure for memristor-based nvLogic and nvSRAM. Moreover, previous nvLogic/nvSRAM circuits cannot achieve low voltage operation. This paper explores various circuit structures for nvLogic and nvSRAM, taking into account memristor endurance, especially for low-voltage applications. Meng-Fan Chang, Ching-Hao Chuang, Min-Ping Chen, Lai-Fu Chen, Hiroyuki Yamauchi, Pi-Feng Chiu, Shyh-Shyuan Sheu |
ASP-DAC | 1 |
| 2011 | Circuit design challenges in embedded memory and resistive RAM (RRAM) for mobile SoC and 3D-ICabstractMobile systems require high-performance and low-power SoC or 3D-IC chips to perform complex operations, ensure a small form-factor and ensure a long battery life time. A low supply voltage (VDD) is frequently utilized to suppress dynamic power consumption, standby current, and thermal effects in SoC and 3D-IC. Furthermore, lowering the VDD reduces the voltage stress of the devices and slows the aging of chips. However, a low VDD for embedded memories can cause functional failure and low yield. This paper reviews various challenges in the design of low-voltage circuits for embedded memory (SRAM and ROM). It also discusses emerging embedded memory solutions. Alternative memory interfaces and architectures for mobile SoC and 3D-IC are also explored. Meng-Fan Chang, Pi-Feng Chiu, Shyh-Shyuan Sheu |
ASP-DAC | 1 |
| 2009 | Analysis and Reduction of Supply Noise Fluctuations Induced by Embedded Via-Programming ROMabstractVarious input addresses and accessed code-patterns of a via-programming read only memory (ROM) cause substantial fluctuations in peak current and supply noise across cycles. This work analyzes the fluctuations in the supply noise that are associated with the pattern-dependent current profile of embedded via-programming ROM on a QFN package with various decoupling capacitances. A pattern-insensitive (PI) technique is developed for via-programming ROM to reduce both fluctuations of peak current and cycle current across various input addresses and accessed code-patterns. The PI technique involves the arranging of the data patterns of a ROM-code and the adjustment of the structures of row decoders and peripheral circuits. Experiments based on the designed test-setup on fabricated 0.25 mum 256 kb ROM macros demonstrate the fluctuation in peak current of conventional ROM and its reduction by the PI technique. The fluctuations of measured peak and cycle currents of PI-ROM are only 0.7% and 13.1% of those of conventional ROM. The PI-ROM also has a 94.5% lower standby current than conventional ROM. Meng-Fan Chang, Shu-Meng Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |