VLDB 2026 Research / reviewers in the wild / expert
Zheyu Yan
dblp:248/8154
· DBLP profile ↗
38ranked-venue papers
9as first author
36since 2021 · last 2026
0000-0003-1830-606XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 8 first-author · 34 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CDACiM: A Charge-Domain Compute-in-Memory Macro for FP/INT MAC Operations with Reconfigurable Capacitor Digital-Analog-ConverterabstractAdvanced edge artificial intelligence (AI) chips need to balance flexible computation, high energy efficiency, and sufficient inference accuracy across diverse workloads. Many compute-in-memory (CiM) designs enable efficient neural network acceleration but focus solely on integer (INT) multiply-and-accumulate (MAC) operations, limiting precision. Some CiM macros add extra circuitry to support floating point (FP) MACs, but these dedicated exponent-handling blocks often waste area when running INT workloads. In this paper, we propose CDACiM, a charge-domain CiM macro that supports both FP and INT MAC operations with minimal overhead. CDACiM introduces a reconfigurable capacitor digital-to-analog converter (RCDAC) that performs both exponent summation and bitwise AND for mantissa multiplication. To calculate exponent offsets, we develop a shared single-slope ADC (SS-ADC) that finds the maximum exponent and computes differences in time domain simultaneously. Our design includes a sparsity-aware computation scheme with tunable thresholds that skips low-importance input-weight pairs, boosting energy efficiency through higher input sparsity. We also introduce a multi-bit input accumulation method that leverages ADC redundancy during quantization and normalization to improve performance. Implemented in a 40nm CMOS process, CDACiM demonstrates an excellent flexibility and trade-off between accuracy and resource usage. Notably, it is the first CiM design to reconfigure capacitor-based INT macro for parallel exponent computation. CDACiM achieves 16.2 TOPS/W for INT MACs and $\mathbf{1 5. 9}$ TFLOPS/W for FP MACs. It delivers a $\mathbf{1. 3 6 - 1. 4 8 \times}$ improvement in energy efficiency with minimal accuracy loss compared to recent FP CiM macros. Jinting Yao, Yuxiao Jiang, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
ASP-DAC | 5 |
| 2026 | FeFET-Based Analog In-Memory Computing With Inherent Shift-and-Add CapabilityabstractIn-memory computing (IMC) architecture has emerged as a highly promising approach, enhancing the energy efficiency of multiply-and-accumulate (MAC) operations in deep neural networks (DNNs) by embedding parallel computations directly into memory arrays. However, existing ferroelectric FET (FeFET)-based analog IMC designs are often constrained to cell-level optimizations and struggle to achieve high-precision MAC operations. In contrast, high-precision analog IMC architectures typically perform MAC operations for partial inputs and weights within the array in a single cycle and then accumulate partial results over multiple cycles. During this procedure, circuits that handle weight shift-and-add process, whether in digital or analog form, incur significant overhead. This paper presents energy-efficient high-precision analog IMC designs leveraging FeFET technology, which inherently support a shift-and-add mechanism for weights. Initially, we introduce an IMC array paradigm that performs partial MAC operations within each column, and seamlessly incorporates the shift-and-add process for weights by utilizing the analog storage properties of FeFET-based cells. Building upon this paradigm, we propose single-level cell (SLC) FeFET-based designs, namely CurFe and ChgFe, operating in the current and charge modes, respectively. Additionally, to leverage FeFET’s multi-level cell (MLC) properties, we propose a novel hybrid SLC-MLC FeFET-based design, MulFe, which offers higher storage density and energy efficiency. Comprehensive evaluations are conducted at both the circuit and system levels, and the results indicate that the average energy efficiency of the proposed FeFET-based analog IMC designs is 1.32× to 2.71× higher compared to state-of-the-art (SOTA) IMC designs. Qingrong Huang, Yu Qian 0002, Jiahao Cai, Kai Ni 0004, Thomas Kämpfe, Zheyu Yan, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Computers | 8 |
| 2026 | QUNF+: A Quadratic Approximation Framework With Hardware Co-Design for Universal Nonlinear Function Acceleration in Neural NetworksabstractModern neural networks have undergone extensive hardware optimization to address the increasing computational demands. While most existing acceleration strategies concentrate on linear operations, the relative cost of these nonlinear operations has become a critical efficiency bottleneck. This work presents QUNF+, a hardware-centric, quadratic-based approximation framework that offers a universal, scalable, and accurate approach for accelerating a wide range of nonlinear functions. Unlike conventional piecewise linear methods, QUNF+ segments functions uniformly and applies second-order Taylor expansions, yielding superior accuracy with fewer segments. QUNF+ also introduces a hardware-efficient approximation scheme with adjustable precision and an optional remainder compensation mechanism. In addition, we propose a set of novel function remapping techniques further reducing approximation error with minimal overhead. We design a fully integrated hardware architecture incorporating Canonical Signed Digit encoding and logic pruning to minimize resource usage without accuracy loss. Experimental results demonstrate that QUNF+ achieves up to 4.1x, 1.5x and 1.7x improvements in power, area, and latency, respectively, over state-of-the-art PWL methods. Application to real-world Transformer models results in an accuracy degradation of less than 0.72% and 0.10% on NLP and CV tasks, with system-level evaluations showing up to 24.38% reduction in total energy consumption and 92.39% in nonlinear operation energy. These results establish QUNF+ as a robust and scalable solution for nonlinear acceleration in modern AI hardware. Haonan Du, Chenyi Wen, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Zheyu Yan, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | NeFT: Negative Feedback Training to Improve Robustness of Compute-in-Memory DNN AcceleratorsabstractCompute-in-memory accelerators built upon non-volatile memory devices excel in energy efficiency and latency when performing deep neural network (DNN) inference, thanks to their in-situ data processing capability. However, the stochastic nature and intrinsic variations of non-volatile memory devices often result in performance degradation during DNN inference. Introducing these non-ideal device behaviors in DNN training enhances robustness, but drawbacks include limited accuracy improvement, reduced prediction confidence, and convergence issues. This arises from a mismatch between the deterministic training and non-deterministic device variations, as such training, though considering variations, relies solely on the model’s final output. In this work, inspired by control theory, we propose Negative Feedback Training (NeFT)—a novel concept supported by theoretical analysis—to more effectively capture the multi-scale noisy information throughout the network. We instantiate this concept with two specific instances, oriented variational forward (OVF) and intermediate representation snapshot (IRS). Based on device variation models extracted from measured data, extensive experiments show that our NeFT outperforms existing state-of-the-art methods with up to a 45.08% improvement in inference accuracy while reducing epistemic uncertainty, boosting output confidence, and improving convergence probability. These results underline the generality and practicality of our NeFT framework for increasing the robustness of DNNs against device variations. The source code for these two instances is available at https://github.com/YifanQin-ND/NeFT_CIM. Zheyu Yan, Dailin Gan, Jun Xia 0003, Zixuan Pan, Wujie Wen, Xiaobo Sharon Hu, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | ZlibBoost: An Efficient and Flexible Open-Source Framework for Standard Cell CharacterizationabstractAs VLSI designs grow increasingly complex and transition to smaller process nodes, accurate and efficient library characterization has become essential for modern design workflows. Existing open-source tools are often constrained by limited functionality, efficiency, and accuracy, making them insufficient for today’s design challenges. This article reviews the shortcomings of current open-source tools and introduces ZlibBoost, a novel open-source framework designed to provide both flexibility and high performance. Its modular, front-end and back-end separated architecture, along with user-friendly interfaces, enables seamless customization, integration of machine learning models, and expanded simulator compatibility. A variety of key features are introduced to significantly enhance both accuracy and efficiency of library characterization. Experimental results demonstrate ZlibBoost’s capability to meet the demands of both academic research and practical applications, establishing it as a robust solution for advancing semiconductor design. Zhengrui Chen, Chengjun Guo, Shizhang Wang, Guozhu Feng, Zixuan Song, Xunzhao Yin, Weiquan Song, Li Zhang 0021, Zheyu Yan, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2026 | A High-Parallelism Softmax Hardware-Software Co-Design for Fast and Efficient LLM InferenceabstractLarge language models (LLMs) have been the driving force behind significant advancements in artificial intelligence. However, their unique self-attention mechanism leads to difficulties in accelerating the inference. Softmax, with its complex nonlinear operations and low parallelism, significantly limits the efficiency of LLMs for long sequences. This work proposes a novel high-parallelism hardware/software co-design Softmax solution. By incorporating a sum estimation algorithm, we eliminate the need for complex exponential on all elements. A high-speed low-energy hardware architecture is introduced by applying high-parallelism statistical module and simplified division module. Experimental results demonstrate that our approach achieves a 9.90–44.75% reduction in latency and up to a 17.36% reduction in energy consumption compared with state-of-the-art Softmax hardware, with minimal impact on model inference accuracy. Chenyi Wen, Haonan Du, Xuyang He, Zheyu Yan, Qi Sun 0002, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | Invited Paper: Boosting Standard Cell Library Characterization with Machine LearningabstractAs VLSI designs grow more complex and transition to smaller process nodes, accurate and efficient library characterization has become increasingly crucial within DTCO and STCO flows. Current open-source tools, however, are constrained to basic library characterization functions and fail to adequately meet modern design demands. In this paper, we review the existing open-source standard cell characterization tools, summarize their limitations, and introduce ZlibBoost---a new open-source framework designed to offer both flexibility and efficiency. We leverage ZlibBoost for LUT index optimization, dynamic power supply noise modeling, and machine learning-based prediction to enhance efficiency and accuracy in library characterization. Experimental results show that such a tool is helpful for both academia and industry to effectively navigate DTCO and STCO challenges. Zhengrui Chen, Chengjun Guo, Zixuan Song, Guozhu Feng, Shizhang Wang, Li Zhang 0021, Xunzhao Yin, Zheyu Yan, Cheng Zhuo |
ASP-DAC | 9 |
| 2025 | A 10.60 μW 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia DetectionabstractThis paper proposes an ultra-low power, mixed-bit-width sparse convolutional neural network (CNN) accelerator to accelerate ventricular arrhythmia (VA) detection. The chip achieves 50% sparsity in a quantized 1D CNN using a sparse processing element (SPE) architecture. Measurement on the prototype chip TSMC 40nm CMOS low-power (LP) process for the VA classification task demonstrates that it consumes 10.60 μW of power while achieving a performance of 150 GOPS and a diagnostic accuracy of 99.95%. The computation power density is only 0.57 μW/mm2, which is 14.23× smaller than state-of-the-art works, making it highly suitable for implantable and wearable medical devices. Zhenge Jia, Zheyu Yan, Jay Mok, Manto Yung, Yu Liu 0007, Wujie Wen, Luhong Liang, Kwang-Ting Cheng, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ASP-DAC | 3 |
| 2025 | Rethinking Medical Anomaly Detection in Brain MRI: An Image Quality Assessment PerspectiveabstractReconstruction-based methods, particularly those leveraging autoencoders, have been widely adopted for anomaly detection task in brain MRI. Unlike most existing works try to improve the task accuracy through architectural or algorithmic innovations, we tackle this task from image quality assessment (IQA) perspective, an under-explored direction in the field. Due to the limitations of conventional metrics such as £1 in capturing the nuanced differences in reconstructed images for medical anomaly detection, we propose fusion quality, a novel metric that wisely integrates the structure-level sensitivity of Structural Similarity Index Measure (SSIM) with the pixel-level precision of £1. The metric offers a more comprehensive assessment of reconstruction quality, considering intensity (subtractive property of l1and divisive property of SSIM), contrast, and structural similarity. Furthermore, the proposed metric makes subtle regional variations more impactful in the final assessment. Thus, considering the inherent divisive properties of SSIM, we design an average intensity ratio (AIR)-based data transformation that amplifies the divisive discrepancies between normal and abnormal regions, thereby enhancing anomaly detection. By fusing the aforementioned two components, we devise the IQA approach. Experimental results on two distinct brain MRI datasets show that our IQA approach significantly enhances medical anomaly detection performance when integrated with state-of-the-art baselines. Code is provided here. Zixuan Pan, Jun Xia 0003, Zheyu Yan, Guoyue Xu, Yawen Wu, Zhenge Jia, Jianxu Chen 0001, Yiyu Shi 0001 |
BIBM | 3 |
| 2025 | Hardware-Aware Compilation and Simulation for In-Memory ComputingabstractThis brief presents an overview of recent tools and research efforts aimed at enhancing the programmability and reliability of In-Memory Computing (IMC)-based systems. We discuss hardware-aware training techniques that improve model resilience to analog device imperfections, and explore mapping strategies that balance accuracy and performance for heterogeneous IMC-based accelerators. Additionally, we examine a compiler framework that abstracts hardware complexities and enables seamless integration of these accelerators into existing deployment pipelines. By combining these approaches with advanced simulation tools, we propose an end-to-end workflow that facilitates the practical deployment and optimization of IMC technologies across diverse memory types and architectural designs. Asif Ali Khan, Hadjer Benmeziane, Hamid Farzaneh, João Paulo C. de Lima, William Andrew Simon, Yiyu Shi 0001, Zheyu Yan, Abu Sebastian, Xiaobo Sharon Hu, Jerónimo Castrillón, Corey Lammie |
CASES | 7 |
| 2025 | VQT-CiM: Accelerating Vector Quantization Enhanced Transformer with Ferroelectric Compute-in-MemoryabstractTransformer models have achieved state-of-the-art performance in various natural language processing (NLP) and computer vision (CV) tasks. To meet their substantial computational demands, the compute-in-memory (CiM) architectures, which alleviate the memory wall problem and enable efficient vector-matrix multiplication (VMM), have been adopted for transformer accelerators. However, the dynamic VMM involved in the attention mechanism, which necessitates runtime write operations, presents significant challenges for non-volatile memory (NVM)-based CiM designs. High write overhead, complex compute-write-compute (CWC) dependencies, and limited endurance reduce their effectiveness. In this paper, we propose VQT-CiM, a ferroelectric FET (FeFET)-based CiM design that accelerates vector quantization (VQ) enhanced transformers by eliminating the runtime write operations. VQT-CiM quantizes keys and values in self-attention to convert dynamic VMMs in inner-product and weighted-sum into static VMMs with the codebooks, enabling efficient calculations with CiM crossbars. However, directly applying VQ hinders the accuracy of transformer model due to its limited representation capability. To address this, we introduce a vector quantization scheme that integrates residual VQ (RVQ) and product VQ (PVQ) for enhanced representation space. We present an efficient hardware implementation for the proposed VQT-CiM with optimized dataflow in RVQ, which incorporates the FeFET-based CiM crossbars and peripheral digital circuits. Evaluation results suggest that VQT-CiM achieves the $3.54 \times$ and $4.53 \times$ improvements in energy efficiency and throughput, respectively, compared to state-of-the-art NVM-based CiM transformer designs. Xuchu Huang, Haonan Du, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
DAC | 4 |
| 2025 | FactorHD: A Hyperdimensional Computing Model for Multi-Object Multi-Class Representation and FactorizationabstractNeuro-symbolic artificial intelligence (neurosymbolic AI) excels in logical analysis and reasoning. Hyperdimensional Computing (HDC), a promising braininspired computational model, is integral to neuro-symbolic AI. Various HDC models have been proposed to represent class-instance and class-class relations, but when representing the more complex class-subclass relation, where multiple objects associate different levels of classes and subclasses, they face challenges for factorization, a crucial task for neuro-symbolic AI systems. In this article, we propose FactorHD, a novel HDC model capable of representing and factorizing the complex class-subclass relation efficiently. FactorHD features a symbolic encoding method that embeds an extra memorization clause, preserving more information for multiple objects. In addition, it employs an efficient factorization algorithm that selectively eliminates redundant classes by identifying the memorization clause of the target class. Such model significantly enhances computing efficiency and accuracy in representing and factorizing multiple objects with class-subclass relation, overcoming limitations of existing HDC models such as “superposition catastrophe” and “the problem of 2 “. Evaluations show that FactorHD achieves approximately $5667 \times$ speedup at a representation size of $10^{9}$ compared to existing HDC models. When integrated with the ResNet-18 neural network, FactorHD achieves $92.48 \%$ factorization accuracy on the Cifar-10 dataset. Xuchu Huang, Chenyu Ni, Zheyu Yan, Xunzhao Yin, Cheng Zhuo |
DAC | 5 |
| 2025 | Algorithm-Hardware Co-Design of a Unified Accelerator for Non-Linear Functions in TransformersabstractNonlinear functions (NFs) in Transformers require high-precision computation consuming significant time and energy, despite the aggressive quantization schemes for other components. Piece-wise Linear (PWL) approximation-based methods offer more efficient processing schemes for NFs but fall short in dealing with functions with high nonlinearities. Moreover, PWL-based methods still suffer from inevitably high latency introduced by the Multiply-And-Add (MADD) unit. To address these issues, this paper proposes a novel quadratic approximation scheme and a highly integrated, multiplier-less hardware structure, as a unified method to accelerate any unary nonlinear function. We also demonstrate implementation examples for GELU, Softmax, and LayerNorm. The experimental results show that the proposed method achieves up to 5.41% higher inference accuracy and 60.12% lower area-delay product. Haonan Du, Chenyi Wen, Zhengrui Chen, Li Zhang 0021, Qi Sun 0002, Zheyu Yan, Cheng Zhuo |
DATE | 6 |
| 2025 | NVCiM-PT: An NVCiM-Assisted Prompt Tuning Framework for Edge LLMsabstractLarge Language Models (LLMs) deployed on edge devices, known as edge LLMs, need to continuously fine-tune their model parameters from user-generated data under limited resource constraints. However, most existing learning methods are not applicable for edge LLMs because of their reliance on high resources and low learning capacity. Prompt tuning (PT) has recently emerged as an effective fine-tuning method for edge LLMs by only modifying a small portion of LLM parameters, but it suffers from user domain shifts, resulting in repetitive training and losing resource efficiency. Conventional techniques to address domain shift issues often involve complex neural networks and sophisticated training, which are incompatible for PT for edge LLMs. Therefore, an open research question is how to address domain shift issues for edge LLMs with limited resources. In this paper, we propose a prompt tuning framework for edge LLMs, exploiting the benefits offered by non-volatile computing-in-memory (NVCiM) architectures. We introduce a novel NVCiM-assisted PT framework, where we narrow down the core operations to matrix-matrix multiplication, which can then be accelerated by performing in-situ computation on NVCiM. To the best of our knowledge, this is the first work employing NVCiM to improve the edge LLM PT performance. Ruiyang Qin, Zheyu Yan, Liu Liu 0023, Dancheng Liu, Amir Nassereldine, Jinjun Xiong, Kai Ni 0004, Xiaobo Sharon Hu, Yiyu Shi 0001 |
DATE | 3 |
| 2025 | GenFIQA: Generative Face Image Quality Assessment via Identity-conditioned Diffusion ModelabstractFace recognition (FR) systems are widely deployed but often struggle due to unconstrained image-capturing conditions. Face image quality assessment (FIQA), applied before recognition, mitigates these challenges by filtering out unreliable samples. Current leading FIQA methods evaluate image quality based on the characteristics observed within the FR model pipeline. However, they leave out the inherent differences in identity embeddings between high-and low-quality face images. To this end, we propose Gen-FIQA, which utilizes a generative model to probe and amplify this difference. Specifically, we extract the identity embedding from an input image using a pre-trained FR model, and then use it as a conditioning signal to generate several face images of the same identity. This generation process leverages the inherent prior in the generative model to translate the difference in identity embedding space back to pixel space. To quantify these differences, the quality score is computed as the average cosine similarity between embeddings from the original and generated images. To improve computational efficiency, we further distill GenFIQA into a lightweight regression-based variant, GenFIQA(R). Extensive experiments across five benchmark datasets and four FR models demonstrate the superiority of our methods over thirteen state-of-the-art FIQA methods. Zheyu Yan, Weisong Zhao, Kai Pang, Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IJCB | 1 |
| 2025 | FACAM: Design and Optimization of A Compact Energy Efficient FeFET-Based Analog Content Addressable MemoryabstractContent Addressable Memory (CAM) is known for highly parallel pattern matching capability, which is widely used for data-centric applications and advanced machine learning models that involve associative search tasks. However, most state-of-the-art CAM designs focus on binary/multi-bit CAMs (B/MCAMs) based on CMOS or emerging nonvolatile memories (NVMs), which struggle in scenarios where analog values, rather than discrete levels, need to be stored and searched. Therefore, analog CAMs (ACAMs) offer a promising solution to further increase memory density, improve energy efficiency and extend practical scenarios. Among NVMs, ferroelectric field effect transistors (FeFETs) have emerged as a strong candidate for efficient CAM designs due to the three-terminal structure, high on-off ratio, high OFF resistance and voltage-driven write/read mechanisms. In this paper, we propose FACAM, a compact and energy efficient single-input 2FeFET-1T ACAM cell design, with a two-phase search scheme, that sets the location and width of the matching range through two FeFETs, respectively. We further present a FACAM array which reduces the matchline (ML) voltage swing by shifting ML precharging into the in-cell search operations. We also propose an adaptive scheme to selectively early-terminate second search phase for further search energy optimization. Evaluation results suggest that our proposed FACAM achieves 8.39× and 2.94× energy efficiency compared with the state-of-the-art better 6T-2R ACAM and 2FeFET ACAM. Benchmarking results in deep random forest accelerator show that our approach is 2.14× faster and 7.82× energy efficient than 2FeFET ACAM. Jiahao Cai, Ann Franchesca Laguna, Thomas Kämpfe, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
ICCAD | 6 |
| 2025 | Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on EdgeabstractThe combination of Large Language Models (LLM) and Automatic Speech Recognition (ASR), when deployed on edge devices (called edge ASR-LLM), can serve as a powerful personalized assistant to enable audio-based interaction for users. Compared to text-based interaction, edge ASR-LLM allows accessible and natural audio interactions. Unfortunately, existing ASR-LLM models are mainly trained in high-performance computing environments and produce substantial model weights, making them difficult to deploy on edge devices. More importantly, to better serve users’ personalized needs, the ASR-LLM must be able to learn from each distinct user, given that audio input often contains highly personalized characteristics that necessitate personalized on-device training. Since individually fine-tuning the ASR or LLM often leads to suboptimal results due to modality-specific limitations, end-to-end training ensures seamless integration of audio features and language understanding (cross-modal alignment), ultimately enabling a more personalized and efficient adaptation on edge devices. However, due to the complex training requirements and substantial computational demands of existing approaches, cross-modal alignment between ASR audio and LLM can be challenging on edge devices. In this work, we propose a resource-efficient cross-modal alignment framework that bridges ASR and LLMs on edge devices to handle personalized audio input. Our framework enables efficient ASR-LLM alignment on resource-constrained devices like Raspberry Pi 5 (8GB RAM), achieving 50x training time speedup while improving the alignment quality by more than 50%. To the best of our knowledge, this is the first work to study efficient ASR-LLM alignment on resource-constrained edge devices. Ruiyang Qin, Dancheng Liu, Gelei Xu, Amir Nassereldine, Zheyu Yan, Chenhui Xu, Xiaobo Sharon Hu, Jinjun Xiong, Yiyu Shi 0001 |
ICCAD | 5 |
| 2025 | ANAS: Software-hardware co-design of approximate neural network accelerators via neural architecture search
Zheyu Yan, Xunzhao Yin, Lenian He, Cheng Zhuo |
Integr. | 2 |
| 2025 | CSA-CiM: Enhancing Multifunctional Computing-in-Memory With Configurable Sense AmplifiersabstractComputing-in-memory (CiM) effectively alleviates the memory wall problem faced by traditional von Neumann architectures when handling data-intensive applications. Most CiM arrays employ dedicated sense amplifiers (SAs) to perform specific functions, and prior configurable CiM arrays achieve multifunctionality by stacking multiple SAs with corresponding functions. However, the independent nature of these SAs, particularly the analog-to-digital converter (ADC), results in excessive energy and area consumption. In this article, we propose a configurable multifunctional ferroelectric field effect transistor (FeFET)-based CiM array design, including configurable peripheral circuit with corresponding multifunctionalities and reusable SA components, to reduce energy consumption and latency. The array cells perform logical AND and XNOR operations, and the proposed SA can be configured to operate in either ADC or winner-take-all (WTA) modes, thereby enabling the array to implement both multiplication-accumulation (MAC) and associative search operations. Instead of operating independently, the WTA component within the SA participates as a flash stage in successive approximation register (SAR) conversions in ADC mode, thus enhancing the WTA utilization, energy efficiency and compactness. By integrating the multifunctional CiM array and the configurable SA, our design supports MAC, Hamming-distance computation (HDC), and nearest neighbor search (NNS) operations within the same structure. Compared to existing works, our design achieves energy efficiency improvements of$7.2\times $for MAC,$2.9\times $for HDC, and EDP improvement of$6.4\times $for NNS, respectively. Yuxiao Jiang, Kai Ni 0004, Thomas Kämpfe, Cheng Zhuo, Zheyu Yan, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge DevicesabstractThe scaling laws have become the de facto guidelines for designing large language models (LLMs), but they were studied under the assumption of unlimited computing resources for both training and inference. As LLMs are increasingly used as personalized intelligent assistants, their customization (i.e., learning through fine-tuning) and deployment onto resource-constrained edge devices will become more and more prevalent. An urgent but open question is how a resource-constrained computing environment would affect the design choices for a personalized LLM. We study this problem empirically in this work. In particular, we consider the tradeoffs among a number of key design factors and their intertwined impacts on learning efficiency and accuracy. The factors include the learning methods for LLM customization, the amount of personalized data used for learning customization, the types and sizes of LLMs, the compression methods of LLMs, the amount of time afforded to learn, and the difficulty levels of the target use cases. Through extensive experimentation and benchmarking, we draw a number of surprisingly insightful guidelines for deploying LLMs onto resource-constrained devices. For example, an optimal choice between parameter learning and RAG may vary depending on the difficulty of the downstream task, the longer fine-tuning time does not necessarily help the model, and a compressed LLM may be a better choice than an uncompressed LLM to learn from limited personalized data. Ruiyang Qin, Dancheng Liu, Chenhui Xu, Zheyu Yan, Zhaoxuan Tan, Zhenge Jia, Amir Nassereldine, Jiajie Li 0002, Meng Jiang 0001, Ahmed Abbasi, Jinjun Xiong, Yiyu Shi 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | SenHDC: A 3-D NAND Flash-Based Processing-in-Sensor Hyperdimensional Computing ArchitectureabstractThe rapid growth of the Internet of Things (IoT) promotes vision application deployments in embedded devices. To alleviate the data conversion and transmission overheads in CMOS image sensor (CIS), processing-in-sensor (PIS) has been proposed, enabling computations to occur directly within sensory systems. Meanwhile, brain-inspired hyperdimensional computing (HDC) emerges as a promising computing paradigm well-suited for PIS due to its high accuracy and efficiency in various cognitive tasks. However, HDC requires transforming the sensed signals into long hypervectors (HVs), which imposes significant overhead for PIS systems. Thus, in this article, we propose SenHDC, a hardware-software co-design framework for HDC-based PIS. We introduce a novel hardware-friendly encoding paradigm that eliminates complex HV transformations, coupled with a noise-resilient position HV generation method that mitigates the impact of capacitance load imbalance. We further present an efficient 3-D NAND flash-based compute-in-memory (CIM) hardware design, comprising an encoding module that performs encoding in charge domain and an associative search module that checks the similarity between the input and prestored classes for classification. Experimental results show that SenHDC improves energy efficiency by$40.6\times $and$1.6\times $in the encoding module and associative search module, respectively, compared to the state-of-the-art HDC CIM implementations. Xuchu Huang, Qingrong Huang, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | FL-NAS: Towards Fairness of NAS for Resource Constrained Devices via Large Language Models : (Invited Paper)abstractNeural Architecture Search (NAS) has become the de fecto tools in the industry in automating the design of deep neural networks for various applications, especially those driven by mobile and edge devices with limited computing resources. The emerging large language models (LLMs), due to their prowess, have also been incorporated into NAS recently and show some promising results. This paper conducts further exploration in this direction by considering three important design metrics simultaneously, i.e., model accuracy, fairness, and hardware deployment efficiency. We propose a novel LLM-based NAS framework, FL-NAS, in this paper, and show experimentally that FL-NAS can indeed find high-performing DNNs, beating state-of-the-art DNN models by orders-of-magnitude across almost all design considerations. Ruiyang Qin, Zheyu Yan, Jinjun Xiong, Ahmed Abbasi, Yiyu Shi 0001 |
ASPDAC | 3 |
| 2024 | Special Session: Sustainable Deployment of Deep Neural Networks on Non-Volatile Compute-in-Memory AcceleratorsabstractNon-volatile memory (NVM) based compute-in-memory (CIM) accelerators have emerged as a sustainable solution to significantly boost energy efficiency and minimize latency for Deep Neural Networks (DNNs) inference due to their in-situ data processing capabilities. However, the performance of NVCIM accelerators degrades because of the stochastic nature and intrinsic variations of NVM devices. Conventional write-verify operations, which enhance inference accuracy through iterative writing and verification during deployment, are costly in terms of energy and time. Inspired by negative feedback theory, we present a novel negative optimization training mechanism to achieve robust DNN deployment for NVCIM. We develop an Oriented Variational Forward (OVF) training method to implement this mechanism. Experiments show that OVF outperforms existing state-of-the-art techniques with up to a 46.71% improvement in inference accuracy while reducing epistemic uncertainty. This mechanism reduces the reliance on write-verify operations and thus contributes to the sustainable and practical deployment of NVCIM accelerators, addressing performance degradation while maintaining the benefits of sustainable computing with NVCIM accelerators. Zheyu Yan, Wujie Wen, Xiaobo Sharon Hu, Yiyu Shi 0001 |
CODES+ISSS | 2 |
| 2024 | TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM AcceleratorsabstractCompute-in-memory (CIM) accelerators using non-volatile memory (NVM) devices offer promising solutions for energy-efficient and low-latency Deep Neural Network (DNN) inference execution. However, practical deployment is often hindered by the challenge of dealing with the massive amount of model weight parameters impacted by the inherent device variations within non-volatile computing-in-memory (NVCIM) accelerators. This issue significantly offsets their advantages by increasing training overhead, the time and energy needed for mapping weights to device states, and diminishing inference accuracy. To mitigate these challenges, we propose the "Tiny Shared Block (TSB)" method, which integrates a small shared 1 × 1 convolution block into the DNN architecture. This block is designed to stabilize feature processing across the network, effectively reducing the impact of device variation. Extensive experimental results show that TSB achieves over 20× inference accuracy gap improvement, over 5× training speedup, and weights-to-device mapping cost reduction while requiring less than 0.4% of the original weights to be write-verified during programming, when compared with state-of-the-art baseline solutions. Our approach provides a practical and efficient solution for deploying robust DNN models on NVCIM accelerators, making it a valuable contribution to the field of energy-efficient AI hardware. Zheyu Yan, Zixuan Pan, Wujie Wen, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ICCAD | 2 |
| 2024 | Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory ArchitecturesabstractLarge Language Models (LLMs) deployed on edge devices learn through fine-tuning and updating a certain portion of their parameters. Although such learning methods can be optimized to reduce resource utilization, the overall required resources remain a heavy burden on edge devices. Instead, Retrieval-Augmented Generation (RAG), a resource-efficient LLM learning method, can improve the quality of the LLM-generated content without updating model parameters. However, the RAG-based LLM may involve repetitive searches on the profile data in every user-LLM interaction. This search can lead to significant latency along with the accumulation of user data. Conventional efforts to decrease latency result in restricting the size of saved user data, thus reducing the scalability of RAG as user data continuously grows. It remains an open question: how to free RAG from the constraints of latency and scalability on edge devices? In this paper, we propose a novel framework to accelerate RAG via Computing-in-Memory (CiM) architectures. It accelerates matrix multiplications by performing in-situ computation inside the memory while avoiding the expensive data transfer between the computing unit and memory. Our framework, Robust CiM-backed RAG (RoCR), utilizing a novel contrastive learning-based training method and noise-aware training, can enable RAG to efficiently search profile data with CiM. To the best of our knowledge, this is the first work utilizing CiM to accelerate RAG. Ruiyang Qin, Zheyu Yan, Dewen Zeng, Zhenge Jia, Dancheng Liu, Ahmed Abbasi, Zhi Zheng 0002, Ningyuan Cao, Kai Ni 0004, Jinjun Xiong, Yiyu Shi 0001 |
ICCAD | 2 |
| 2024 | Personalized Meta-Federated Learning for IoT-Enabled Health MonitoringabstractFederated learning (FL) has been widely adopted in IoT-enabled health monitoring on biosignals thanks to its advantages in data privacy preservation. However, the global model trained from FL generally performs unevenly across subjects since biosignal data is inherent with complex temporal dynamics. The morphological characteristics of biosignals with the same label can vary significantly among different subjects (i.e., inter-subject variability) while biosignals with varied temporal patterns can be collected on the same subject (i.e., intra-subject variability). To address the challenges, we present the Personalized Meta-Federated learning (PMFed) framework for personalized IoT-enabled health monitoring. Specifically, in the federated learning stage, a novel momentum-based model aggregating strategy is introduced to aggregate clients' models based on domain similarity in the meta-federated learning paradigm to obtain a well-generalized global model while speeding up the convergence. In the model personalizing stage, an adaptive model personalization mechanism is devised to adaptively tailor the global model based on the subject-specific biosignal features while preserving the learned cross-subject representations. We develop an IoT-enabled computing framework to evaluate the effectiveness of PMFed over three real-world health monitoring tasks. Experimental results show that the PMFed excels at detection performances in terms of F1 and accuracy by up to 9.4% and 8.7%, and reduces training overhead and throughput by up to 56.3% and 63.4% when compared with the SOTA federated learning algorithms. Zhenge Jia, Tianren Zhou, Zheyu Yan, Jingtong Hu, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | U-SWIM: Universal Selective Write-Verify for Computing-in-Memory Neural AcceleratorsabstractArchitectures that incorporate computing-in-memory (CiM) using emerging nonvolatile memory (NVM) devices have become strong contenders for deep neural network (DNN) acceleration due to their impressive energy efficiency. Yet, a significant challenge arises when using these emerging devices: they can show substantial variations during the weight-mapping process. This can severely impact DNN accuracy if not mitigated. A widely accepted remedy for imperfect weight mapping is the iterative write-verify approach, which involves verifying conductance values and adjusting devices if needed. In all existing publications, this procedure is applied to every individual device, resulting in a significant programming time overhead. In our research, we illustrate that only a small fraction of weights need this write-verify treatment for the corresponding devices and the DNN accuracy can be preserved, yielding a notable programming acceleration. Building on this, we introduce U-SWIM, a novel method based on the second derivative. It leverages a single iteration of forward and backpropagation to pinpoint the weights demanding write-verify. Through extensive tests on diverse DNN designs and datasets, U-SWIM manifests up to a$10\times $programming acceleration against the traditional exhaustive write-verify method, all while maintaining a similar accuracy level. Furthermore, compared to our earlier SWIM technique, U-SWIM excels, showing a$7\times $speedup when dealing with devices exhibiting nonuniform variations. Zheyu Yan, Xiaobo Sharon Hu, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Compute-in-Memory-Based Neural Network Accelerators for Safety-Critical Systems: Worst-Case Scenarios and ProtectionsabstractEmerging non-volatile memory (NVM)-based Computing-in-Memory (CiM) architectures show substantial promise in accelerating deep neural networks (DNNs) due to their exceptional energy efficiency. However, NVM devices are prone to device variations. Consequently, the actual DNN weights mapped to NVM devices can differ considerably from their targeted values, inducing significant performance degradation. Many existing solutions aim to optimize average performance amidst device variations, which is a suitable strategy for general-purpose conditions. However, the worst-case performance that is crucial for safety-critical applications is largely overlooked in current research. In this study, we define the problem of pinpointing the worst-case performance of CiM DNN accelerators affected by device variations. Additionally, we introduce a strategy to identify a specific pattern of the device value deviations in the complex, high-dimensional value deviation space, responsible for this worst-case outcome. Our findings reveal that even subtle device variations can precipitate a dramatic decline in DNN accuracy, posing risks for CiM-based platforms in supporting safety-critical applications. Notably, we observe that prevailing techniques to bolster average DNN performance in CiM accelerators fall short in enhancing worst-case scenarios. In light of this issue, we propose a novel worst-case-aware training technique named A-TRICE that efficiently combines adversarial training and noise-injection training with right-censored Gaussian noise to improve the DNN accuracy in the worst-case scenarios. Our experimental results demonstrate that A-TRICE improves the worst-case accuracy under device variations by up to 33%. Zheyu Yan, Xiaobo Sharon Hu, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | DASALS: Differentiable Architecture Search-Driven Approximate Logic SynthesisabstractApproximate computing is a promising computing paradigm for designing energy-efficient systems. To automatically generate approximate circuits, many local iterative approximate logic synthesis (ALS) methods have been proposed. They need to specify a particular local approximation change and apply it to modify the local structure of a circuit in each round. This will lose some global optimization opportunities, thus, degrading circuit quality. In this paper, we propose DASALS, a differentiable archltecture search-driven ALS method, to directly search the whole circuit structure to obtain the approximate circuits with better circuit quality-accuracy trade-off. DASALS is based on a proper continuous relaxation of the discrete search space of ALS and an efficient gradient descent-based search algorithm. The experimental results show that compared with a state-of-the-art method, DASALS on average reduces the area-delay product by 10.82% and mean square error by 10.93%. Xuan Wang 0027, Zheyu Yan, Chang Meng, Yiyu Shi 0001, Weikang Qian |
ICCAD | 2 |
| 2023 | Improving Realistic Worst-Case Performance of NVCiM DNN Accelerators Through Training with Right-Censored Gaussian NoiseabstractCompute-in-Memory (CiM), built upon non-volatile memory (NVM) devices, is promising for accelerating deep neural networks (DNNs) owing to its in-situ data processing capability and superior energy efficiency. To battle device variations, noise injection training is commonly used, which perturbs weights with Gaussian noise during training to make the model more robust to weight variations. Despite its prevalence, however, existing successes are mostly empirical, and very little theoretical support is available. Even the most fundamental questions such as why Gaussian but not other types of noises should be used is not answered. In this work, through formally analyzing the effect of injecting Gaussian noise in training to improve the k-th percentile performance (KPP), a realistic worst-case performance metric, for the first time we provide a theoretical justification of the effectiveness of the approach. We further show that surprisingly Gaussian noise is not the best option, contrary to what has been taken for granted in the literature. Instead, a right-censored Gaussian noise significantly improves the KPP of DNNs. We further propose an automated method to determine the optimal hyperparameters for injecting this right-censored Gaussian noise during the training process. Our method achieves up to a 26% improvement in KPP compared to the state-of-the-art methods employed to enhance DNN robustness under the impact of device variations. Zheyu Yan, Wujie Wen, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ICCAD | 1 |
| 2022 | RADARS: Memory Efficient Reinforcement Learning Aided Differentiable Neural Architecture SearchabstractDifferentiable neural architecture search (DNAS) is known for its capacity in the automatic generation of superior neural networks. However, DNAS based methods suffer from memory usage explosion when the search space expands, which may prevent them from running successfully on even advanced GPU platforms. On the other hand, reinforcement learning (RL) based methods, while being memory efficient, are extremely time-consuming. Combining the advantages of both types of methods, this paper presents RADARS, a scalable RL aided DNAS framework that can explore large search spaces in a fast and memory-efficient manner. RADARS iteratively applies RL to prune undesired architecture candidates and identifies a promising subspace to carry out DNAS. Experiments using a workstation with 12 GB GPU memory show that on CIFAR-10 and ImageNet datasets, RADARS can achieve up to 3.41% higher accuracy with 2.5X search time reduction compared with a state-of-the-art RL-based method, while the two DNAS baselines cannot complete due to excessive memory usage or search time. To the best of the authors’ knowledge, this is the first DNAS framework that can handle large search spaces with bounded memory usage. Zheyu Yan, Weiwen Jiang, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ASP-DAC | 1 |
| 2022 | SWIM: selective write-verify for computing-in-memory neural acceleratorsabstractComputing-in-Memory architectures based on non-volatile emerging memories have demonstrated great potential for deep neural network (DNN) acceleration thanks to their high energy efficiency. However, these emerging devices can suffer from significant variations during the mapping process (i.e., programming weights to the devices), and if left undealt with, can cause significant accuracy degradation. The non-ideality of weight mapping can be compensated by iterative programming with a write-verify scheme, i.e., reading the conductance and rewriting if necessary. In all existing works, such a practice is applied to every single weight of a DNN as it is being mapped, which requires extensive programming time. In this work, we show that it is only necessary to select a small portion of the weights for write-verify to maintain the DNN accuracy, thus achieving significant speedup. We further introduce a second derivative based technique SWIM, which only requires a single pass of forward and backpropagation, to efficiently select the weights that need write-verify. Experimental results on various DNN architectures for different datasets show that SWIM can achieve up to 10x programming speedup compared with conventional full-blown write-verify while attaining a comparable accuracy. Zheyu Yan, Xiaobo Sharon Hu, Yiyu Shi 0001 |
DAC | 1 |
| 2022 | Computing-In-Memory Neural Network Accelerators for Safety-Critical Systems: Can Small Device Variations Be Disastrous?abstractComputing-in-Memory (CiM) architectures based on emerging nonvolatile memory (NVM) devices have demonstrated great potential for deep neural network (DNN) acceleration thanks to their high energy efficiency. However, NVM devices suffer from various non-idealities, especially device-to-device variations due to fabrication defects and cycle-to-cycle variations due to the stochastic behavior of devices. As such, the DNN weights actually mapped to NVM devices could deviate significantly from the expected values, leading to large performance degradation. To address this issue, most existing works focus on maximizing average performance under device variations. This objective would work well for general-purpose scenarios. But for safety-critical applications, the worst-case performance must also be considered. Unfortunately, this has been rarely explored in the literature. In this work, we formulate the problem of determining the worst-case performance of CiM DNN accelerators under the impact of device variations. We further propose a method to effectively find the specific combination of device variation in the high-dimensional space that leads to the worst-case performance. We find that even with very small device variations, the accuracy of a DNN can drop drastically, causing concerns when deploying CiM accelerators in safety-critical applications. Finally, we show that surprisingly none of the existing methods used to enhance average DNN performance in CiM accelerators are very effective when extended to enhance the worst-case performance, and further research down the road is needed to address this problem. Zheyu Yan, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ICCAD | 1 |
| 2022 | VisualNet: An End-to-End Human Visual System Inspired Framework to Reduce Inference Latency of Deep Neural NetworksabstractAcceleration of deep neural network (DNN) inference has gained increasing attention recently with the wide adoption of DNNs for practical applications. For computer vision tasks where inputs are images, existing works mostly focus on improving the throughput of inference for multiple images. However, in many real-time applications, it is critical to reduce the latency of a single image inference, which is more complicated than improving the throughput because of the inherent data dependencies. On the other hand, from human brain's perspective, the complexity in our visual surroundings is first encoded as a pattern of light on a two dimensional array of photoreceptors, with little direct resemblance to the original input or the ultimate percept. Within just a few hundred microns of retinal thickness, this initial signal encoded by our photoreceptors must be transformed into an adequate representation of the entire visual scene. Inspired by how the retina helps human brain incept new information efficiently, we present an end-to-end structured framework built using any existing convolutional neural network (CNN) as the backbone. The proposed framework, called VisualNet, can create task parallelism for the backbone during the inference of a single image. Experiments using a number of neural networks for the ImageNet classification task and the CIFAR-10 classification task on GPUs and CPUs show that the proposed VisualNet reduces the latency of the regular network it builds on by up to 80.6% when both are fully parallelized with state-of-the-art acceleration libraries. At the same time, VisualNet can achieve similar or slightly higher accuracy. Jinjun Xiong, Song Bian 0001, Zheyu Yan, Meiping Huang, Jian Zhuang, Takashi Sato 0001, Xiaowei Xu 0004, Yiyu Shi 0001 |
IEEE Trans. Computers | 5 |
| 2021 | Uncertainty Modeling of Emerging Device based Computing-in-Memory Neural Accelerators with Application to Neural Architecture Searchabstractemerging device based Computing-in-memory (CiM) has been proved to be a promising candidate for high energy efficiency deep neural network (DNN) computations. However, most emerging devices suffer uncertainty issues, resulting in a difference between actual data stored and the weight value it is design to be. This leads to an accuracy drop from trained models to actually deployed platforms. In this work, we offer a thorough analysis on the effect of such uncertainties induced changes in DNN models. To reduce the impact of device uncertainties, we propose UAE, a uncertainty-aware Neural Architecture Search scheme to identify a DNN model that is both accurate and robust against device uncertainties. Zheyu Yan, Da-Cheng Juan, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ASP-DAC | 1 |
| 2021 | Device-Circuit-Architecture Co-Exploration for Computing-in-Memory Neural AcceleratorsabstractCo-exploration of neural architectures and hardware design is promising due to its capability to simultaneously optimize network accuracy and hardware efficiency. However, state-of-the-art neural architecture search algorithms for the co-exploration are dedicated for the conventional von-Neumann computing architecture, whose performance is heavily limited by the well-known memory wall. In this article, we are the first to bring the computing-in-memory architecture, which can easily transcend the memory wall, to interplay with the neural architecture search, aiming to find the most efficient neural architectures with high network accuracy and maximized hardware efficiency. Such a novel combination makes opportunities to boost performance, but also brings a bunch of challenges: The optimization space spans across multiple design layers from device type and circuit topology to neural architecture; and the presence of device variation may drastically degrade the neural network performance. To address these challenges, we propose a cross-layer exploration framework, namely NACIM, which jointly explores device, circuit and architecture design space and takes device variation into consideration to find the most robust neural architectures, coupled with the most efficient hardware design. Experimental results demonstrate that NACIM can find the robust neural network with 0.45 percent accuracy loss in the presence of device variation, compared with a 76.44 percent loss from the state-of-the-art NAS without consideration of variation; in addition, NACIM achieves an energy efficiency up to 16.3 TOPs/W, 3.17x higher than the state-of-the-art NAS. Weiwen Jiang, Qiuwen Lou, Zheyu Yan, Lei Yang 0018, Jingtong Hu, Xiaobo Sharon Hu, Yiyu Shi 0001 |
IEEE Trans. Computers | 3 |
| 2020 | When Single Event Upset Meets Deep Neural Networks: Observations, Explorations, and RemediesabstractDeep Neural Network has proved its potential in various perception tasks and hence become an appealing option for interpretation and data processing in security sensitive systems. However, security-sensitive systems demand not only high perception performance, but also design robustness under various circumstances. Unlike prior works that study network robustness from software level, we investigate from hardware perspective about the impact of Single Event Upset (SEU) induced parameter perturbation (SIPP) on neural networks. We systematically define the fault models of SEU and then provide the definition of sensitivity to SIPP as the robustness measure for the network. We are then able to analytically explore the weakness of a network and summarize the key findings for the impact of SIPP on different types of bits in a floating point parameter, layer-wise robustness within the same network and impact of network depth. Based on those findings, we propose two remedy solutions to protect DNNs from SIPPs, which can mitigate accuracy degradation from 28% to 0.27% for ResNet with merely 0.24-bit SRAM area overhead per parameter. Zheyu Yan, Yiyu Shi 0001, Wang Liao 0001, Masanori Hashimoto, Xichuan Zhou, Cheng Zhuo |
ASP-DAC | 1 |
| 2020 | Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple TasksabstractNeural Architecture Search (NAS) has demonstrated its power on various AI accelerating platforms such as Field Programmable Gate Arrays (FPGAs) and Graphic Processing Units (GPUs). However, it remains an open problem how to integrate NAS with Application-Specific Integrated Circuits (ASICs), despite them being the most powerful AI accelerating platforms. The major bottleneck comes from the large design freedom associated with ASIC designs. Moreover, with the consideration that multiple DNNs will run in parallel for different workloads with diverse layer operations and sizes, integrating heterogeneous ASIC sub-accelerators for distinct DNNs in one design can significantly boost performance, and at the same time further complicate the design space. To address these challenges, in this paper we build ASIC template set based on existing successful designs, described by their unique dataflows, so that the design space is significantly reduced. Based on the templates, we further propose a framework, namely ASICNAS, which can simultaneously identify multiple DNN architectures and the associated heterogeneous ASIC accelerator design, such that the design specifications (specs) can be satisfied, while the accuracy can be maximized. Experimental results show that compared with successive NAS and ASIC design optimizations which lead to design spec violations, ASICNAS can guarantee the results to meet the design specs with 17.77%, 2.49×, and 2.32× reductions on latency, energy, and area and less than 1.6% accuracy loss. To the best of the authors’ knowledge, this is the first work on neural architecture and ASIC accelerator design co-exploration. Lei Yang 0018, Zheyu Yan, Meng Li 0004, Hyoukjun Kwon, Liangzhen Lai, Tushar Krishna, Vikas Chandra, Weiwen Jiang, Yiyu Shi 0001 |
DAC | 2 |