Yu Cao 0001

dblp:68/6563-1 · also Yu Kevin Cao · DBLP profile ↗
← Back
145ranked-venue papers
9as first author
41since 2021 · last 2026
0000-0001-6968-1180ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 124 · 9 first-author · 32 since 2021Artificial intelligence and machine learning · 14 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 since 2021Software engineering, systems software and programming languages · 10 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Chiplet-NAS: Chiplet-aware Neural Architecture Search for Efficient AI Inference on 2.5D Integration
abstract
The co-design of neural network architectures and their target chiplet-based hardware systems presents a significant challenge due to the vast and combinatorial design space. Identifying solutions that are Pareto-optimal across competing objectives of task accuracy, system latency, and power consumption requires solutions beyond manual design and brute-force methods. This paper proposes a closed-loop chiplet-aware neural architecture search (Chiplet-NAS) framework to automate the exploration and discover hardware-optimized models for efficient AI inference on 2.5 D chiplet-based systems. The framework integrates a Tree-structured Parzen Estimator (TPE) for sampleefficient search with CLAIRE, a chiplet-based library and fast performance benchmarking tool, to provide direct hardware feedback on latency and energy consumption, along with accuracy optimization. The framework is evaluated by co-designing ResNet-based model architectures with chiplet based hardware systems. Compared to a baseline NAS that optimizes only on the task accuracy, our Chiplet-NAS achieves significant power and performance benefits at the iso-accuracy.
Pragnya Sudershan Nalla, Nikhil K. Cherukuri, Sachin S. Sapatnekar, Chaitali Chakrabarti, Yu Cao 0001, Jeff Zhang 0001
ASP-DAC7
2026 HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases
abstract
Retrieval Augmented Generation (RAG) is an essential agent for Large Language Model (LLM) aided Description Language (HDL) tasks, addressing the challenges of limited training data and prohibitively long prompts. However, its performance in handling ambiguous queries and real-world, repository-level HDL projects containing thousands or even tens of thousands of code lines remains limited. Our analysis demonstrates two fundamental mismatches, structural and vocabulary, between conventional semantic similarity-based RAGs and HDL codes. To this end, we propose HDLxGraph, the first framework that integrates the inherent graph characteristics of HDLs with RAGs for LLM-assisted tasks. Specifically, HDLxGraph incorporates Syntax Trees (ASTs) to capture HDLs’ hierarchical structures and Data Flow Graphs (DFGs) to address the vocabulary mismatch. In addition, to overcome the lack of comprehensive HDL search benchmarks, we introduce HDLSearch, an LLMgenerated dataset derived from real-world, repository-level HDL projects. Evaluations show that HDLxGraph improves search, debugging, and completion accuracy by $\mathbf{1 2. 0 4 \%} \boldsymbol{/} \mathbf{1 2. 2 2 \%} \boldsymbol{/} \mathbf{5. 0 4 \%}$ and by $\mathbf{1 1. 5 9 \%} \boldsymbol{/} \mathbf{8. 1 8 \%} \boldsymbol{/} \mathbf{4. 0 7 \%}$ over state-of-the-art similarity-based RAG and software-code Graph RAG baselines, respectively. The code of HDLxGraph and HDLSearch benchmark are available at https://github.com/UMN-ZhaoLab/HDLxGraph.
Pingqing Zheng, Jiayin Qin, Fuqi Zhang, Niraj Chitla, Zishen Wan, Shang Wu 0003, Yu Cao 0001, Caiwen Ding, Yang Zhao 0013
ASP-DAC7
2026 EGO: Efficient Compression of Unstructured Sparse DNNs for Compute-in-Memory based on Graph Minimum-Cost Matching Optimization
abstract
Compute-in-memory (CiM) for edge AI inference operates under strict memory and energy constraints. While unstructured pruning reduces model size and computation, efficiently deploying the sparse weights on CiM’s dense, regular arrays remains challenging. Existing studies either incur high indexing overhead by storing per-element indexing metadata, or achieve limited compression by relying on scarce structural patterns within unstructured weights. The column packing method, which avoids the high overhead of per-element indexing and offers rich compression potential, shows promise to reconcile unstructured sparsity with CiM’s regular compute pattern, but its direct application to CiM is hindered by heuristic grouping algorithms that either yield suboptimal compression or sacrifice model accuracy.To bridge this gap and unlock the potential of column packing for CiM, this study presents EGO, an algorithm-hardware co-designed framework. EGO overcomes the inefficiency of heuristic grouping by introducing a combinatorially optimized grouping algorithm, which formulates column packing as minimum-cost graph matching. A digital CiM architecture is co-designed with the EGO column grouping formulation, which features a custom Sparsity Processing Unit (SPU) to enable efficient activation routing while preserving CiM’s dense and regular dataflow. Circuit-level simulations show that EGO achieves 1.4–3.7x average improvement in energy efficiency and 1.2–1.8x average improvement in area efficiency compared to previous state-of-the-art methods.
Teng Wan, Yu Cao 0001, Huazhong Yang, Xueqing Li 0002
DATE2
2026 vFPGA: Towards Sub-µs Reconfiguration via 3D FPGA and Packaging Co-Design
Nikhil K. Cherukuri, Sharad Nag, Pragnya Sudershan Nalla, Ashish K. Kola, Chetan S. Gadireddi, Kevin Dai, Jae-sun Seo, Zhenman Fang, Jeff Zhang 0001, Yu Cao 0001
FPGA10
2026 A 22nm Reconfigurable Systolic Array for FFT and AI Inference
John Stolzberg-Schray, Sharad Nag, Jacob Johnson, Nikhil K. Cherukuri, Ashish K. Kola, Gopikrishnan Raveendran Nair, Jeff Zhang 0001, Jae-sun Seo, Yu Cao 0001
ISCAS9
2026 2.5D/3D Chiplet-based Integration: New Dimensions in Design and Testing
Ganap A. Tewary, Partho Bhoumik, Pragnya S. Nalla, Yu Cao 0001, Krishnendu Chakrabarty, Jeff Zhang 0001
VTS4
2025 CLAIRE: Composable Chiplet Libraries for AI Inference
abstract
Artificial intelligence has made a significant impact on fields like computer vision, Natural Language Processing (NLP), healthcare, and robotics. However, recent AI models, such as GPT-4 and LLaMAv3, demand significant number of computational resources, pushing monolithic chips to their technological and practical limits. 2.5D chiplet-based heterogeneous architectures have been proposed to address these technological and practical limits. While chiplet optimization for models like Convolutional Neural Networks (CNNs) is well-established, scaling this approach to accommodate diverse AI inference models with different computing primitives, data volumes, and different chiplet sizes is very challenging. A set of hardened IPs and chiplet libraries optimized for a broad range of AI applications is proposed in this work. We derive the set of chiplet configurations that are composable, scalable and reusable by employing an analytical framework trained on a diverse set of AI algorithms. Testing these set of library synthesized configurations on a different set of algorithms, we achieve a$1.99\times-3.99\times$improvement in non-recurring engineering (NRE) chiplet design costs, with minimal performance overhead compared to custom chiplet-based ASIC designs. Similar to soft IPs for SoC development, the library of chiplets improves flexibility, reusability, and efficiency for AI hardware designs.
Pragnya Sudershan Nalla, Emad Haque, Yaotian Liu, Sachin S. Sapatnekar, Jeff Zhang 0001, Chaitali Chakrabarti, Yu Cao 0001
DATE7
2025 MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging
abstract
As program workloads (e.g., AI) increase in size and algorithmic complexity, the primary challenge lies in their high dimensionality, encompassing computing cores, array sizes, and memory hierarchies. To overcome these obstacles, innovative approaches are required. Agile chip design has already benefited from machine learning integration at various stages, including logic synthesis, placement, and routing. With Large Language Models (LLMs) recently demonstrating impressive proficiency in Hardware Description Language (HDL) generation, it is promising to extend their abilities to 2.5D integration, an advanced technique that saves area overhead and development costs. However, LLM-driven chiplet design faces challenges such as flatten design, high validation cost and imprecise parameter optimization, which limit its chiplet design capability. To address this, we propose MAHL, a hierarchical LLM-based chiplet design generation framework that features six agents which collaboratively enable AI algorithm-hardware mapping, including hierarchical description generation, retrieval-augmented code generation, diverseflow-based validation, and multi-granularity design space exploration. These components together enhance the efficient generation of chiplet design with optimized Power, Performance and Area (PPA). Experiments show that MAHL not only significantly improves the generation accuracy of simple RTL design, but also increases the generation accuracy of real-world chiplet design, evaluated by Pass@5, from 0 to 0.72 compared to conventional LLMs under the best-case scenario. Compared to state-of-the-art CLARIE (expert-based), MAHL achieves comparable or even superior PPA results under certain optimization objectives.
Jinwei Tang, Jiayin Qin, Nuo Xu 0013, Pragnya Sudershan Nalla, Yu Cao 0001, Yang Zhao 0013, Caiwen Ding
ICCAD5
2025 Adaptive Graph Learning for Efficient Thermal Analysis of Multi-Stacking Chiplet Systems under Interface Variations
abstract
Efficient thermal analysis is critical for ensuring the reliability and performance of modern integrated circuits, particularly in multi-stacking technologies. Traditional thermal analysis methods rely on numerical solutions of partial differential equations (PDEs), which are computationally expensive. This paper introduces a fast thermal framework that synergizes a graph neural network (GNN), hybrid with finite element methods (FEM), to accelerate thermal predictions for a wide range of 2.5D/3D design configurations. Our graph neural network architecture is inherently scalable, accommodating various design sizes. Furthermore, it is able to incorporate additional trainable nodes to be adaptive to new temperature profiles under realistic defect conditions during the assembly process. Validated against fine-grained numerical solutions and real post-silicon thermal imaging data, the proposed framework achieves an average mean absolute percentage error (MAPE) of 0.05%. It completes thermal simulation within a few hundred milliseconds, yielding a speedup of over 1000× compared to conventional steady-state finite difference method (FDM) solvers. Moreover, our method exhibits robust adaptability to previously unseen process and material variations without the need for retraining.
Ziyao Yang, Jingbo Sun 0003, Vidya A. Chhabria, Yu Cao 0001
ICCAD4
2025 Occult: Optimizing Collaborative Communications across Experts for Accelerated Parallel MoE Training and Inference
abstract
Mixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-to-all communication across devices. Unfortunately, such communication overhead typically constitutes a significant portion of the total runtime, hampering the scalability of distributed training and inference for modern MoE models (consuming over 40% runtime in large-scale training). In this paper, we first define $\textit{collaborative communication}$ to illustrate this intrinsic limitation, and then propose system- and algorithm-level innovations to reduce communication costs. Specifically, given a pair of experts co-activated by one token, we call them as $\textit{collaborated}$, which comprises $2$ cases as $\textit{intra-}$ and $\textit{inter-collaboration}$, depending on whether they are kept on the same device. Our pilot investigations reveal that augmenting the proportion of intra-collaboration can accelerate expert parallel at scale. It motivates us to strategically $\underline{\texttt{o}}$ptimize $\underline{\texttt{c}}$ollaborative $\underline{\texttt{c}}$omm$\underline{\texttt{u}}$nication for acce$\underline{\texttt{l}}$era$\underline{\texttt{t}}$ed MoE training and inference, dubbed $\textbf{\texttt{Occult}}$. Our designs are capable of $\underline{either}$ delivering exact results with reduced communication cost, $\underline{or}$ controllably minimizing the cost with collaboration pruning, materialized by modified fine-tuning. Comprehensive experiments on various MoE-LLMs demonstrate that $\texttt{Occult}$ can be faster than popular state-of-the-art inference or training frameworks (over 50% speed up across multiple tasks and models) with comparable or superior quality compared to the standard fine-tuning. Codes will be available upon acceptance.
Shuqing Luo, Pingzhi Li, Jie Peng 0002, Yang Zhao 0013, Yu Cao 0001, Yu Cheng 0001, Tianlong Chen 0001
ICML5
2025 Uncertainty-Based Extensible Codebook for Discrete Federated Learning in Heterogeneous Data Silos
abstract
Federated learning (FL), aimed at leveraging vast distributed datasets, confronts a crucial challenge: the heterogeneity of data across different silos. While previous studies have explored discrete representations to enhance model generalization across minor distributional shifts, these approaches often struggle to adapt to new data silos with significantly divergent distributions. In response, we have identified that models derived from FL exhibit markedly increased uncertainty when applied to data silos with unfamiliar distributions. Consequently, we propose an innovative yet straightforward iterative framework, termed \emph{Uncertainty-Based Extensible-Codebook Federated Learning (UEFL)}. This framework dynamically maps latent features to trainable discrete vectors, assesses the uncertainty, and specifically extends the discretization dictionary or codebook for silos exhibiting high uncertainty. Our approach aims to simultaneously enhance accuracy and reduce uncertainty by explicitly addressing the diversity of data distributions, all while maintaining minimal computational overhead in environments characterized by heterogeneous data silos. Extensive experiments across multiple datasets demonstrate that UEFL outperforms state-of-the-art methods, achieving significant improvements in accuracy (by 3\%--22.1\%) and uncertainty reduction (by 38.83\%--96.24\%). The source code is available at https://github.com/destiny301/uefl.
Yu Cao 0001, Dianbo Liu
ICML2
2025 RTGS: Real-Time 3D Gaussian Splatting SLAM via Multi-Level Redundancy Reduction
Leshu Li, Jiayin Qin, Jie Peng 0002, Zishen Wan, Huaizhi Qu, Pingqing Zheng, Hongsen Zhang, Yu Cao 0001, Tianlong Chen 0001, Yang Zhao 0013
MICRO9
2025 Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
abstract
Mixture-of-Experts (MoE) architecture offers enhanced efficiency for Large Language Models (LLMs) with modularized computation, yet its inherent sparsity poses significant hardware deployment challenges, including memory locality issues, communication overhead, and inefficient computing resource utilization. Inspired by the modular organization of the human brain, we propose $\texttt{Mozart}$, a novel algorithm-hardware co-design framework tailored for efficient training of MoE-based LLMs on 3.5D wafer-scale chiplet architectures. On the algorithm side, $\texttt{Mozart}$ exploits the inherent modularity of chiplets and introduces: ($1$) an expert allocation strategy that enables efficient on-package all-to-all communication, and ($2$) a fine-grained scheduling mechanism that improves communication-computation overlap through streaming tokens and experts. On the architecture side, $\texttt{Mozart}$ adaptively co-locates heterogeneous modules on specialized chiplets with a 2.5D NoP-Tree topology and hierarchical memory structure. Evaluation across three popular MoE models demonstrates significant efficiency gains, enabling more effective parallelization and resource utilization for large-scale modularized MoE-LLMs.
Shuqing Luo, Pingzhi Li, Jiayin Qin, Jie Peng 0002, Yang Zhao 0013, Yu Cao 0001, Tianlong Chen 0001
NeurIPS7
2025 HISIM: Analytical Performance Modeling and Design Space Exploration of 2.5D/3D Integration for AI Computing
abstract
Monolithic designs face significant fabrication cost and data movement challenges, especially when executing complex and diverse AI models. Advanced 2.5D/3D packaging promises high bandwidth and connection density to overcome these challenges, yet it also introduces new electro-thermal constraints. This article develops a suite of analytical performance models to enable efficient benchmarking of a 2.5D/3D heterogeneous system for energy-efficient AI computing. These models encompass various performance metrics related to computing units, network-on-chip (NoC), and network-on-package (NoP). The results are summarized into a new tool, HISIM, which is$10^{4} \times $–$10^{6} \times $faster than state-of-the-art AI benchmark tools. Furthermore, HISIM integrates rapid thermal simulation for the 2.5D/3D system, helping shed light on both the potential and limitations of 2.5D/3D heterogeneous integration (HI) on representative AI algorithms. The code of HISIM is available athttps://github.com/mec-UMN/HISIM.
Zhenyu Wang 0016, Pragnya Sudershan Nalla, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Jae-sun Seo, Vidya A. Chhabria, Jeff Zhang 0001, Chaitali Chakrabarti, Ümit Y. Ogras, Yu Cao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.11
2024 Transformer-Based Selective Super-resolution for Efficient Image Refinement
abstract
Conventional super-resolution methods suffer from two drawbacks: substantial computational cost in upscaling an entire large image, and the introduction of extraneous or potentially detrimental information for downstream computer vision tasks during the refinement of the background. To solve these issues, we propose a novel transformer-based algorithm, Selective Super-Resolution (SSR), which partitions images into non-overlapping tiles, selects tiles of interest at various scales with a pyramid architecture, and exclusively reconstructs these selected tiles with deep features. Experimental results on three datasets demonstrate the efficiency and robust performance of our approach for super-resolution. Compared to the state-of-the-art methods, the FID score is reduced from 26.78 to 10.41 with 40% reduction in computation cost for the BDD100K dataset.
Kishore Kasichainula, Yaoxin Zhuo, Baoxin Li, Jae-sun Seo, Yu Cao 0001
AAAI6
2024 Exploiting 2.5D/3D Heterogeneous Integration for AI Computing
abstract
The evolution of AI algorithms has not only revolutionized many application domains, but also posed tremendous challenges on the hardware platform. Advanced packaging technology today, such as 2.5D and 3D interconnection, provides a promising solution to meet the ever-increasing demands of bandwidth, data movement, and system scale in AI computing. This work presents HISIM, a modeling and benchmarking tool for chiplet-based heterogeneous integration. HISIM emphasizes the hierarchical interconnection that connects various chiplets through network-on-package. It further integrates technology roadmap, power/latency prediction, and thermal analysis together to support electro-thermal co-design. Leveraging HISIM with in-memory computing chiplets, we explore the advantages and limitations of 2.5D and 3D heterogenous integration on representative AI algorithms, such as DNNs, transformers, and graph neural networks.
Zhenyu Wang 0016, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Yaotian Liu, Jae-sun Seo, Chaitali Chakrabarti, Ümit Y. Ogras, Vidya A. Chhabria, Jeff Zhang 0001, Yu Cao 0001
ASPDAC11
2024 A 16nm Heterogeneous Accelerator for Energy-Efficient Sparse and Dense AI Computing
abstract
Artificial intelligence (AI) has evolved from dense Deep Neural Networks (DNNs) toward a diverse set of models, such as sparse graph convolutional neural networks (GCNs). These new models differ in model size, processing flow, memory access patterns, and data/model sparsity. Hardware platforms optimized for dense DNNs with a regular data structure are inefficient to manage new unstructured, sparse workloads, such as GCNs. For instance, in-memory computing (IMC) units that is suitable for dense matrix/vector computation, but significantly underutilized for sparse data.
Gopikrishnan Raveendran Nair, Fengyang Jiang, Jeff Zhang 0001, Yu Cao 0001
ISLPED4
2024 Cooling the Chaos: Mitigating the Effect of Threshold Voltage Variation in Cryogenic CMOS Memories
abstract
Cryogenic CMOS is a promising technology for high performance computing due to its improvement in subthreshold slope, carrier mobilities and reduced wire resistance. The threshold voltage (Vth) increase at 77K can be mitigated by metal gate work function (PHIG) engineering to achieve matched off current (Ioff) further enhancing the device performance allowing us to operate at very low supply voltage thereby reducing the Energy Delay Product (EDP). However, the effect of variation on noise margins of static random access memories (SRAM) deploying these matched Ioff devices is very prominent especially at low supply voltages (Vdd) limiting its scaling. In this work, we propose a framework to perform Vth retargeting for cryogenic SRAM for improving noise margins in high performance cryogenic SRAM cells under variation. The proposed framework comprises of a Monte-Carlo engine which performs statistical analysis and DC characterization and a backend processing engine to analyze noise margins and tune the PHIG. To demonstrate the framework, we use calibrated 14nm FinFET models at 300K and 77K. First, we analyze the logic blocks using iso-Ioff devices, which yield up to 3x improvement in delay at iso-energy and a 4.5x reduction in energy at iso-delay. Next, we study the effect of Vth variation on the device currents. Finally, the framework is deployed to tune PHIG, and results show that it can enhance the noise margins by 23%, 31% and 19% for hold, read and write operations respectively at 77K compared to iso-Ioff devices. Further, a 1kb SRAM array has been simulated using iso-Ioff tuned peripherals and framework tuned SRAM cells, and it shows 5.4x reduction in read/write energies along with 1.2x delay reduction and better noise margins at 77K compared to 300K.
Rakshith Saligram, Amol D. Gaidhane, Yu Cao 0001, Suman Datta, Arijit Raychowdhury
ISLPED3
2024 Cryogenic Operation of Computing-In-Memory based Spiking Neural Network
abstract
This paper introduces a Computing-In-Memory based Spiking Neural Network (SNN) architecture for cryogenic operation of CMOS (Cryo-SNN). The paper demonstrates design strategies to improve energy efficiency of Cryo-SNN by coupling low-voltage operation at cryogenic temperature with innovative design of neuron circuits optimized for cryogenic conditions. By exploiting the enhanced device characteristics of 14 nm FinFET transistors at cryogenic temperatures, our architecture outlines critical adaptations to SNN components for optimal functionality in extreme environments. The circuit simulation using measurement calibrated 14nm FinFET models shows that a Cryo-SNN designed for MNIST classification operates with 4.54X improved energy-delay-product (EDP) over room temperature operation while maintaining similar accuracy. Further, the paper designs an optimized SNN architecture for autonomous health monitoring of miniaturized satellites at cryogenic temperature consuming less than 1mW of power.
Laith A. Shamieh, Wei-Chun Wang 0001, Shida Zhang, Rakshith Saligram, Amol D. Gaidhane, Yu Cao 0001, Arijit Raychowdhury, Suman Datta, Saibal Mukhopadhyay
ISLPED6
2024 Patch-based Selection and Refinement for Early Object Detection
abstract
Early object detection (OD) is a crucial task for the safety of many dynamic systems. Current OD algorithms have limited success for small objects at a long distance. To improve the accuracy and efficiency of such a task, we propose a novel set of algorithms that divide the image into patches, select patches with objects at various scales, elaborate the details of a small object, and detect it as early as possible. Our approach is built upon a transformer-based network and integrates the diffusion model to improve the detection accuracy. As demonstrated on BDD100K, our algorithms enhance the mAP for small objects from 1.03 to 8.93, and reduce the data volume in computation by more than 77%.
Kishore Kasichainula, Yaoxin Zhuo, Baoxin Li, Jae-sun Seo, Yu Cao 0001
WACV6
2024 A Progressive Subnetwork Searching Framework for Dynamic Inference
abstract
Deep neural network (DNN) model compression is a popular and important optimization method for efficient and fast hardware acceleration. However, the compressed model is usually fixed, without the capability to tune the computing complexity (i.e., latency in hardware) on-the-fly, depending on dynamic latency requirements, workloads, and computing hardware resource allocation. To address this challenge, dynamic DNN with run-time adaption of computing structures has been constructed through training with a cross-entropy objective function consisting of multiple subnets sampled from the supernet. Our investigations in this work show that the performance of dynamic inference highly relies on the quality of subnet sampling. To construct a dynamic DNN with multiple high-quality subnets, we propose a progressive subnetwork searching framework, which is embedded with several proposed new techniques, including trainable noise ranking, channel-group sampling, selective fine-tuning, and subnet filtering. Our proposed framework empowers the target dynamic DNN with higher accuracy for all the subnets compared with prior works on both the Canadian Institute for Advanced Research dataset with 10 classes (CIFAR-10) and ImageNet datasets. Specifically, compared with United States-Neural Network (US-NN), our method achieves 0.9% average accuracy gain for Alexnet, 2.5% for ResNet18, 1.1% for Visual Geometry Group (VGG)11, and 0.58% for MobileNetv1, on the ImageNet dataset, respectively. Moreover, to demonstrate run-time tuning of computing latency of dynamic DNN in real computing system, we have deployed our constructed dynamic networks into Nvidia Titan graphics processing unit (GPU) and Intel Xeon central processing unit (CPU), showing great improvement over prior works. The code is available at https://github.com/ASU-ESIC-FAN-Lab/Dynamic-inference.
Li Yang 0009, Zhezhi He, Yu Cao 0001, Deliang Fan
IEEE Trans. Neural Networks Learn. Syst.3
2024 High Throughput FPGA-Based Object Detection via Algorithm-Hardware Co-Design
abstract
Object detection and classification is a key task in many computer vision applications such as smart surveillance and autonomous vehicles. Recent advances in deep learning have significantly improved the quality of results achieved by these systems, making them more accurate and reliable in complex environments. Modern object detection systems make use of lightweight convolutional neural networks (CNNs) for feature extraction, coupled with single-shot multi-box detectors (SSDs) that generate bounding boxes around the identified objects along with their classification confidence scores. Subsequently, a non-maximum suppression (NMS) module removes any redundant detection boxes from the final output. Typical NMS algorithms must wait for all box predictions to be generated by the SSD-based feature extractor before processing them. This sequential dependency between box predictions and NMS results in a significant latency overhead and degrades the overall system throughput, even if a high-performance CNN accelerator is used for the SSD feature extraction component. In this paper, we present a novel pipelined NMS algorithm that eliminates this sequential dependency and associated NMS latency overhead. We then use our novel NMS algorithm to implement an end-to-end fully pipelined FPGA system for low-latency SSD-MobileNet-V1 object detection. Our system, implemented on an Intel Stratix 10 FPGA, runs at 400 MHz and achieves a throughput of 2,167 frames per second with an end-to-end batch-1 latency of 2.13 ms. Our system achieves 5.3× higher throughput and 5× lower latency compared to the best prior FPGA-based solution with comparable accuracy.
Anupreetham Anupreetham, Mohamed Ibrahim 0005, Mathew Hall, Andrew Boutros, Ajay Kuzhively, Abinash Mohanty, Eriko Nurvitadhi, Vaughn Betz, Yu Cao 0001, Jae-sun Seo
ACM Trans. Reconfigurable Technol. Syst.9
2023 FPGA Acceleration of GCN in Light of the Symmetry of Graph Adjacency Matrix
abstract
Graph Convolutional Neural Networks (GCNs) are widely used to process large-scale graph data. Different from deep neural networks (DNNs), GCNs are sparse, irregular, and unstructured, posing unique challenges to hardware acceleration with regular processing elements (PEs). In particular, the adja-cency matrix of a GCN is extremely sparse, leading to frequent but irregular memory access, low spatial/temporal data locality and poor data reuse. Furthermore, a realistic graph usually consists of unstructured data (e.g., unbalanced distributions), creating significantly different processing times and imbalanced workload for each node in GCN acceleration. To overcome these challenges, we propose an end-to-end hardware-software co-design to accelerate GCNs on resource-constrained FPGAs with the features including: (1) A custom dataflow that leverages symmetry along the diagonal of the adjacency matrix to accelerate feature aggregation for undirected graphs. We utilize either the upper or the lower triangular matrix of the adjacency matrix to perform aggregation in GCN to improve data reuse. (2) Unified compute cores for both aggregation and transform phases, with full support to the symmetry-based dataflow. These cores can be dynamically reconfigured to the systolic mode for transformation or as individual accumulators for aggregation in GCN processing. (3) Preprocessing of the graph in software to rearrange the edges and features to match the custom dataflow. This step improves the regularity in memory access and data reuse in the aggregation phase. Moreover, we quantize the GCN precision from FP32 to INT8 to reduce the memory footprint without losing the inference accuracy. We implement our accelerator design in Intel Stratix10 MX FPGA board with HBM2, and demonstrate$1.3\times-110.5\times$improvement in end-to-end GCN latency as compared to the state-of the-art FPGA implementations, on the graph datasets of Cora, Pubmed, Citeseer and Reddit.
Gopikrishnan Raveendran Nair, Han-Sok Suh, Mahantesh Halappanavar, Frank Liu 0001, Jae-sun Seo, Yu Cao 0001
DATE6
2023 Improving the Efficiency of CMOS Image Sensors through In-Sensor Selective Attention
abstract
Inspired by the selective attention mechanism in human vision, we propose to introduce a saliency-based processing step in the CMOS image sensor, to continuously select pixels corresponding to salient objects and feedback such information to the sensor, instead of blindly passing all pixels to the sensor output. To minimize the overhead of saliency detection in this feedback loop, we propose two techniques: (1) saliency detection with low-precision, down-sampled grayscale images, and (2) Optimization of the loss function and model structure. Finally, we pad the minimum number of pixels around the selected pixels to maintain the accuracy of object detection (OD). Our method is experimented with two types of OD algorithms on three representative datasets. At the similar OD accuracy with the full image, our proposed selective feedback method successfully achieves 70.5% reduction in the volume of output pixels for BDD100K, which translates to 4.3× and 3.4× reduction in power consumption and latency, respectively.
Kishore Kasichainula, Dong-Woo Jee, Injune Yeo, Yaoxin Zhuo, Baoxin Li, Jae-sun Seo, Yu Cao 0001
ISCAS8
2023 SpikeSim: An End-to-End Compute-in-Memory Hardware Evaluation Tool for Benchmarking Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are an active research domain toward energy-efficient machine intelligence. Compared to conventional artificial neural networks (ANNs), SNNs use temporal spike data and bio-plausible neuronal activation functions such as leaky-integrate fire/integrate fire (LIF/IF) for data processing. However, SNNs incur significant dot-product operations causing high memory and computation overhead in standard von-Neumann computing platforms. To this end, in-memory computing (IMC) architectures have been proposed to alleviate the “memory-wall bottleneck” prevalent in von-Neumann architectures. Although recent works have proposed IMC-based SNN hardware accelerators, the following key implementation aspects have been overlooked: 1) the adverse effects of crossbar nonideality on SNN performance due to repeated analog dot-product operations over multiple time-steps and 2) hardware overheads of essential SNN-specific components, such as the LIF/IF and data communication modules. To this end, we propose SpikeSim, a tool that can perform realistic performance, energy, latency and area evaluation of IMC-mapped SNNs. SpikeSim consists of a practical monolithic IMC architecture called SpikeFlow for mapping SNNs. Additionally, the nonideality computation engine (NICE) and energy–latency–area (ELA) engine performs hardware-realistic evaluation of SpikeFlow-mapped SNNs. Based on 65nm CMOS implementation and experiments on CIFAR10, CIFAR100 and TinyImagenet datasets, we find that the LIF/IF neuronal module has significant area contribution$(>11\%$of the total hardware area). To this end, we propose SNN topological modifications that leads to$1.24\times $and$10\times $reduction in the neuronal module’s area and the overall energy-delay-product value, respectively. Furthermore, in this work, we perform a holistic comparison between IMC implemented ANN and SNNs and conclude that lower number of time-steps are the key to achieve higher throughput and energy-efficiency for SNNs compared to 4-bit ANNs. The code repository for the SpikeSim tool is available at Github link.
Abhishek Moitra, Abhiroop Bhattacharjee, Runcong Kuang, Yu Cao 0001, Priyadarshini Panda
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 MNSIM 2.0: A Behavior-Level Modeling Tool for Processing-In-Memory Architectures
abstract
In the age of Artificial Intelligence (AI), the huge data movements between memory and computing units become the bottleneck of von Neumann architectures, i.e., the “memory wall” problem. In order to tackle this challenge, Processing-In-Memory (PIM) architectures are proposed, which perform in-situ computations in memory and give alternative solutions to boost the computing energy efficiency and performance. Because of the large-scale Neural Network (NN) algorithm models and the huge hardware design space, various factors affect computing accuracy and performance, bringing the need for efficient PIM modeling and evaluation tools. In this work, we propose a behavior-level modeling tool, MNSIM 2.0, to model the performance of PIM architectures efficiently. At the hardware level, MNSIM 2.0 provides a hierarchical PIM modeling structure with flexible architecture configurability and components extensibility. Moreover, the first unified PIM memory array model is proposed for describing both digital and analog PIM. At the algorithm level, MNSIM 2.0 supports the PIM-based NN computing accuracy simulation considering various architecture and device parameters. A PIM-oriented NN model training and quantization flow is also integrated to improve the performance gain brought by PIM. At the scheduling level, MNSIM 2.0 adopts a universal scheduling description compatible with different scheduling strategies. Validation using fabricated PIM macros shows the relative modeling error rate of MNSIM 2.0 is 3:8 5:5%. Case studies show that MNSIM 2.0 enables PIM design space explorations, influences analysis of device parameters, and architecture design insight discoveries.
Zhenhua Zhu 0002, Hanbo Sun, Tongxin Xie, Guohao Dai 0001, Lixue Xia, Dimin Niu, Xiaoming Chen 0003, Xiaobo Sharon Hu, Yu Cao 0001, Yuan Xie 0001, Huazhong Yang, Yu Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.10
2023 Algorithm-hardware Co-optimization for Energy-efficient Drone Detection on Resource-constrained FPGA
abstract
Convolutional neural network (CNN)-based object detection has achieved very high accuracy; e.g., single-shot multi-box detectors (SSDs) can efficiently detect and localize various objects in an input image. However, they require a high amount of computation and memory storage, which makes it difficult to perform efficient inference on resource-constrained hardware devices such as drones or unmanned aerial vehicles (UAVs). Drone/UAV detection is an important task for applications including surveillance, defense, and multi-drone self-localization and formation control. In this article, we designed and co-optimized an algorithm and hardware for energy-efficient drone detection on resource-constrained FPGA devices. We trained an SSD object detection algorithm with a custom drone dataset. For inference, we employed low-precision quantization and adapted the width of the SSD CNN model. To improve throughput, we use dual-data rate operations for DSPs to effectively double the throughput with limited DSP counts. For different SSD algorithm models, we analyze accuracy or mean average precision (mAP) and evaluate the corresponding FPGA hardware utilization, DRAM communication, and throughput optimization. We evaluated the FPGA hardware for a custom drone dataset, Pascal VOC, and COCO2017. Our proposed design achieves a high mAP of 88.42% on the multi-drone dataset, with a high energy efficiency of 79 GOPS/W and throughput of 158 GOPS using the Xilinx Zynq ZU3EG FPGA device on the Open Vision Computer version 3 (OVC3) platform. Our design achieves 1.1 to 8.7× higher energy efficiency than prior works that used the same Pascal VOC dataset, using the same FPGA device, but at a low-power consumption of 2.54 W. For the COCO dataset, our MobileNet-V1 implementation achieved an mAP of 16.8, and 4.9 FPS/W for energy-efficiency, which is ∼ 1.9× higher than prior FPGA works or other commercial hardware platforms.
Han-Sok Suh, Jian Meng, Ty Nguyen, Vijay Kumar 0001, Yu Cao 0001, Jae-sun Seo
ACM Trans. Reconfigurable Technol. Syst.5
2022 Gradient-Based Novelty Detection Boosted by Self-Supervised Binary Classification
abstract
Novelty detection aims to automatically identify out-of-distribution (OOD) data, without any prior knowledge of them. It is a critical step in data monitoring, behavior analysis and other applications, helping enable continual learning in the field. Conventional methods of OOD detection perform multi-variate analysis on an ensemble of data or features, and usually resort to the supervision with OOD data to improve the accuracy. In reality, such supervision is impractical as one cannot anticipate the anomalous data. In this paper, we propose a novel, self-supervised approach that does not rely on any pre-defined OOD data: (1) The new method evaluates the Mahalanobis distance of the gradients between the in-distribution and OOD data. (2) It is assisted by a self-supervised binary classifier to guide the label selection to generate the gradients, and maximize the Mahalanobis distance. In the evaluation with multiple datasets, such as CIFAR-10, CIFAR-100, SVHN and TinyImageNet, the proposed approach consistently outperforms state-of-the-art supervised and unsupervised methods in the area under the receiver operating characteristic (AUROC) and area under the precision-recall curve (AUPR) metrics. We further demonstrate that this detector is able to accurately learn one OOD class in continual learning.
Jingbo Sun 0003, Li Yang 0009, Jiaxin Zhang 0005, Frank Liu 0001, Mahantesh Halappanavar, Deliang Fan, Yu Cao 0001
AAAI7
2022 XBM: A Crossbar Column-wise Binary Mask Learning Method for Efficient Multiple Task Adaption
abstract
Recently, utilizing ReRAM crossbar array to accelerate DNN inference on single task has been widely studied. However, using the crossbar array for multiple task adaption has not been well explored. In this paper, for the first time, we propose XBM, a novel crossbar column-wise binary mask learning method for multiple task adaption in ReRAM crossbar DNN accelerator. XBM leverages the mask-based learning algorithm's benefit to avoid catastrophic forgetting to learn a task-specific mask for each new task. With our hardware-aware design innovation, the required masking operation to adapt for a new task could be easily implemented in existing crossbar based convolution engine with minimal hardware/ memory overhead and, more importantly, no need of power hungry cell re-programming, unlike prior works. The extensive experimental results show that compared with state-of-the-art multiple task adaption methods, XBM keeps the similar accuracy on new tasks while only requires 1.4% mask memory size compared with popular piggyback. Moreover, the elimination of cell re-programming or tuning saves up to 40% energy during new task adaption.
Fan Zhang 0069, Li Yang 0009, Jian Meng, Yu Cao 0001, Jae-sun Seo, Deliang Fan
ASP-DAC4
2022 XMA: a crossbar-aware multi-task adaption framework via shift-based mask learning method
abstract
ReRAM crossbar array as a high-parallel fast and energy-efficient structure attracts much attention, especially on the acceleration of Deep Neural Network (DNN) inference on one specific task. However, due to the high energy consumption of weight re-programming and the ReRAM cells' low endurance problem, adapting the crossbar array for multiple tasks has not been well explored. In this paper, we propose XMA, a novel crossbar-aware shift-based mask learning method for multiple task adaption in the ReRAM crossbar DNN accelerator for the first time. XMA leverages the popular mask-based learning algorithm's benefit to mitigate catastrophic forgetting and learn a task-specific, crossbar column-wise, and shift-based multi-level mask, rather than the most commonly used element-wise binary mask, for each new task based on a frozen backbone model. With our crossbar-aware design innovation, the required masking operation to adapt for a new task could be implemented in an existing crossbar-based convolution engine with minimal hardware/memory overhead and, more importantly, no need for power-hungry cell re-programming, unlike prior works. The extensive experimental results show that, compared with state-of-the-art multiple task adaption Piggyback method [1], XMA achieves 3.19% higher accuracy on average, while saving 96.6% memory overhead. Moreover, by eliminating cell re-programming, XMA achieves ~4.3x higher energy efficiency than Piggyback.
Fan Zhang 0069, Li Yang 0009, Jian Meng, Jae-sun Seo, Yu Cao 0001, Deliang Fan
DAC5
2022 XST: A Crossbar Column-wise Sparse Training for Efficient Continual Learning
abstract
Leveraging the ReRAM crossbar-based In-Memory-Computing (IMC) to accelerate single task DNN inference has been widely studied. However, using the ReRAM crossbar for continual learning has not been explored yet. In this work, we propose XST, a novel crossbar column-wise sparse training framework for continual learning. XST significantly reduces the training cost and saves inference energy. More importantly, it is friendly to existing crossbar-based convolution engine with almost no hardware overhead. Compared with the state-of-the-art CPG method, the experiments show that XST's accuracy achieves 4.95 % higher accuracy. Furthermore, XST demonstrates ~5.59 × training speedup and 1.5 × inference energy-saving.
Fan Zhang 0069, Li Yang 0009, Jian Meng, Jae-sun Seo, Yu Cao 0001, Deliang Fan
DATE5
2022 Big-Little Chiplets for In-Memory Acceleration of DNNs: A Scalable Heterogeneous Architecture
abstract
Monolithic in-memory computing (IMC) architectures face significant yield and fabrication cost challenges as the complexity of DNNs increases. Chiplet-based IMCs that integrate multiple dies with advanced 2.5D/3D packaging offers a low-cost and scalable solution. They enable heterogeneous architectures where the chiplets and their associated interconnection can be tailored to the non-uniform algorithmic structures to maximize IMC utilization and reduce energy consumption. This paper proposes a heterogeneous IMC architecture with big-little chiplets and a hybrid network-on-package (NoP) to optimize the utilization, interconnect bandwidth, and energy efficiency. For a given DNN, we develop a custom methodology to map the model onto the big-little architecture such that the early layers in the DNN are mapped to the little chiplets with higher NoP bandwidth and the subsequent layers are mapped to the big chiplets with lower NoP bandwidth. Furthermore, we achieve a scalable solution by incorporating a DRAM into each chiplet to support a wide range of DNNs beyond the area limit. Compared to a homogeneous chiplet-based IMC architecture, the proposed big-little architecture achieves up to 329× improvement in the energy-delay-area product (EDAP) and up to 2× higher IMC utilization. Experimental evaluation of the proposed big-little chiplet-based RRAM IMC architecture for ResNet-50 on ImageNet shows 259×, 139×, and 48× improvement in energy-efficiency at lower area compared to Nvidia V100 GPU, Nvidia T4 GPU, and SIMBA architecture, respectively.
A. Alper Goksoy, Sumit K. Mandal, Zhenyu Wang 0016, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001
ICCAD8
2022 Impact of On-chip Interconnect on In-memory Acceleration of Deep Neural Networks
abstract
With the widespread use of Deep Neural Networks (DNNs), machine learning algorithms have evolved in two diverse directions—one with ever-increasing connection density for better accuracy and the other with more compact sizing for energy efficiency. The increase in connection density increases on-chip data movement, which makes efficient on-chip communication a critical function of the DNN accelerator. The contribution of this work is threefold. First, we illustrate that the point-to-point (P2P)-based interconnect is incapable of handling a high volume of on-chip data movement for DNNs. Second, we evaluate P2P and network-on-chip (NoC) interconnect (with a regular topology such as a mesh) for SRAM- and ReRAM-based in-memory computing (IMC) architectures for a range of DNNs. This analysis shows the necessity for the optimal interconnect choice for an IMC DNN accelerator. Finally, we perform an experimental evaluation for different DNNs to empirically obtain the performance of the IMC architecture with both NoC-tree and NoC-mesh. We conclude that, at the tile level, NoC-tree is appropriate for compact DNNs employed at the edge, and NoC-mesh is necessary to accelerate DNNs with high connection density. Furthermore, we propose a technique to determine the optimal choice of interconnect for any given DNN. In this technique, we use analytical models of NoC to evaluate end-to-end communication latency of any given DNN. We demonstrate that the interconnect optimization in the IMC architecture results in up to 6 × improvement in energy-delay-area product for VGG-19 inference compared to the state-of-the-art ReRAM-based IMC architectures.
Sumit K. Mandal, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001
ACM J. Emerg. Technol. Comput. Syst.6
2022 Exploring Model Stability of Deep Neural Networks for Reliable RRAM-Based In-Memory Acceleration
abstract
RRAM-based in-memory computing (IMC) effectively accelerates deep neural networks (DNNs). Furthermore, model compression techniques, such as quantization and pruning, are necessary to improve algorithm mapping and hardware performance. However, in the presence of RRAM device variations, low-precision and sparse DNNs suffer from severe post-mapping accuracy loss. To address this, in this work, we investigate a new metric,model stability, from the loss landscape to help shed light on accuracy loss under variations and model compression, which guides an algorithmic solution to maximize model stability and mitigate accuracy loss. Based on statistical data from a CMOS/RRAM 1T1R test chip at 65nm, we characterize wafer-level RRAM variations and develop a cross-layer benchmark tool that incorporates quantization, pruning, device variations, model stability, and IMC architecture parameters to assess post-mapping accuracy and hardware performance. Leveraging this tool, we show that a loss-landscape-based DNN model selection for stability effectively tolerates device variations and achieves a post-mapping accuracy higher than that with 50% lower RRAM variations. Moreover, we quantitatively interpret why model pruning increases the sensitivity to variations, while a lower-precision model has better tolerance to variations. Finally, we propose a novel variation-aware training method to improve model stability, in which there exists the most stable model for the best post-mapping accuracy of compressed DNNs. Experimental evaluation of the method shows up to 19%, 21%, and 11% post-mapping accuracy improvement for our 65nm RRAM device, across various precision and sparsity, on CIFAR-10, CIFAR-100, and SVHN datasets, respectively.
Li Yang 0009, Jingbo Sun 0003, Jubin Hazra, Xiaocong Du, Maximilian Liehr, Zheng Li 0020, Karsten Beckmann, Rajiv V. Joshi, Nathaniel C. Cady, Deliang Fan, Yu Cao 0001
IEEE Trans. Computers12
2022 Hybrid RRAM/SRAM in-Memory Computing for Robust DNN Acceleration
abstract
RRAM-based in-memory computing (IMC) effectively accelerates deep neural networks (DNNs) and other machine learning algorithms. On the other hand, in the presence of RRAM device variations and lower precision, the mapping of DNNs to RRAM-based IMC suffers from severe accuracy loss. In this work, we propose a novel hybrid IMC architecture that integrates an RRAM-based IMC macro with a digital SRAM macro using a programmable shifter to compensate for the RRAM variations and recover the accuracy. The digital SRAM macro consists of a small SRAM memory array and an array of multiply-and-accumulate (MAC) units. The nonideal output from the RRAM macro, due to device and circuit nonidealities, is compensated by adding the precise output from the SRAM macro. In addition, the programmable shifter allows for different scales of compensation by shifting the SRAM macro output relative to the RRAM macro output. On the algorithm side, we develop a framework for the training of DNNs to support the hybrid IMC architecture through ensemble learning. The proposed framework performs quantization (weights and activations), pruning, RRAM IMC-aware training, and employs ensemble learning through different compensation scales by utilizing the programmable shifter. Finally, we design a silicon prototype of the proposed hybrid IMC architecture in the 65-nm SUNY process to demonstrate its efficacy. Experimental evaluation of the hybrid IMC architecture shows that the SRAM compensation allows for a realistic IMC architecture with multilevel RRAM cells (MLCs) even though they suffer from high variations. The hybrid IMC architecture achieves up to 21.9%, 12.65%, and 6.52% improvement in post-mapping accuracy over state-of-the-art techniques, at minimal overhead, for ResNet-20 on CIFAR-10, VGG-16 on CIFAR-10, and ResNet-18 on ImageNet, respectively.
Zhenyu Wang 0016, Injune Yeo, Li Yang 0009, Jian Meng, Maximilian Liehr, Rajiv V. Joshi, Nathaniel C. Cady, Deliang Fan, Jae-sun Seo, Yu Cao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.11
2021 SWIFT: Small-World-based Structural Pruning to Accelerate DNN Inference on FPGA
abstract
State-of-the-art DNN pruning approaches achieved high sparsity. However, these methods usually do not consider the intrinsic graph property of DNNs, leading to an irregular pruned network. Consequently, hardware accelerators cannot directly benefit from such pruning, suffering additional cost on indexing, control and data paths. Inspired by the observation that the brain and real-world networks follow a Small-World model, we propose a graph-based progressive structural pruning technique, SWIFT, that integrates local clusters and global sparsity in DNNs to benefit the dataflow and workload balance of the accelerators. In particular, we propose an output stationary FPGA architecture to accelerate DNN inference and integrate it with the structural sparsity by SWIFT, so that the communication and computation of clustered zero weights are eliminated. In addition, a full mesh data router is designed to adaptively direct inputs into corresponding processing elements (PEs) for different layer configurations and skipping zero operations. The proposed SWIFT is evaluated with multiple DNNs on different datasets. It achieves sparsity ratio up to 76% for CIFAR-10, 83% for CIFAR-100, 76% for the SVHN datasets. Moreover, our proposed SWIFT FPGA accelerator achieves up to 4.4× improvement in throughput for different dense networks with a marginal hardware overhead.
Yufei Ma 0002, Yu Cao 0001, Le Ye, Ru Huang 0001
FPGA3
2021 End-to-End FPGA-based Object Detection Using Pipelined CNN and Non-Maximum Suppression
abstract
Object detection is an important computer vision task, with many applications in autonomous driving, smart surveillance, robotics, and other domains. Single-shot detectors (SSD) coupled with a convolutional neural network (CNN) for feature extraction can efficiently detect, classify and localize various objects in an input image with very high accuracy. In such systems, the convolution layers extract features and predict the bounding box locations for the detected objects as well as their confidence scores. Then, a non-maximum suppression (NMS) algorithm eliminates partially overlapping boxes and selects the bounding box with the highest score per class. However, these two components are strictly sequential; a conventional NMS algorithm needs to wait for all box predictions to be produced before processing them. This prohibits any overlap between the execution of the convolutional layers and NMS, resulting in significant latency overhead and throughput degradation. In this paper, we present a novel NMS algorithm that alleviates this bottleneck and enables a fully-pipelined hardware implementation. We also implement an end-to-end system for low-latency SSD-MobileNet-V1 object detection, which combines a state-of-the-art deeply-pipelined CNN accelerator with a custom hardware implementation of our novel NMS algorithm. As a result of our new algorithm, the NMS module adds a minimal latency overhead of only 0.13μ s to the SSD-MobileNet-V1 convolution layers. Our end-to-end object detection system implemented on an Intel Stratix 10 FPGA runs at a maximum operating frequency of 350 MHz, with a throughput of 609 frames-per-second and an end-to-end batch-1 latency of 2.4 ms. Our system achieves 1.5× higher throughput and 4.4× lower latency compared to the current state-of-the-art SSD-based object detection systems on FPGAs.
Anupreetham Anupreetham, Mohamed Ibrahim 0005, Mathew Hall, Andrew Boutros, Ajay Kuzhively, Abinash Mohanty, Eriko Nurvitadhi, Vaughn Betz, Yu Cao 0001, Jae-sun Seo
FPL9
2021 Algorithm-Hardware Co-Optimization for Energy-Efficient Drone Detection on Resource-Constrained FPGA
abstract
Convolutional neural network (CNN) based object detection has achieved very high accuracy, e.g. single-shot multi-box detectors (SSD) can efficiently detect and localize various objects in an input image. However, they require a high amount of computation and memory storage, which makes it difficult to perform efficient inference on resource-constrained hardware devices such as drones or unmanned aerial vehicles (UAVs). Drone/UAV detection is an important task for applications including surveillance, defense, and multi-drone self-localization and formation control. In this paper, we designed and co-optimized algorithm and hardware for energy-efficient drone detection on resource-constrained FPGA devices. We trained SSD object detection algorithm with a custom drone dataset. For inference, we employed low-precision quantization and adapted the width of the SSD CNN model. To improve throughput, we use dual-data rate operations for DSPs to effectively double the throughput with limited DSP counts. For different SSD algorithm models, we analyze accuracy or mean average precision (mAP) and evaluate the corresponding FPGA hardware utilization, DRAM communication, throughput optimization. Our proposed design achieves a high mAP of 88.42% on the multi-drone dataset, with a high energy-efficiency of 79 GOPS/W and throughput of 158 GOPS using Xilinx Zynq ZU3EG FPGA device on the Open Vision Computer version 3 (OVC3) platform. Our design achieves 2.7X higher energy efficiency than prior works using the same FPGA device, at a low-power consumption of 1.98 W.
Han-Sok Suh, Jian Meng, Ty Nguyen, Shreyas K. Venkataramanaiah, Vijay Kumar 0001, Yu Cao 0001, Jae-sun Seo
FPT6
2021 Alternate Model Growth and Pruning for Efficient Training of Recommendation Systems
abstract
Deep learning recommendation systems at scale have provided remarkable gains through increasing model capacity (i.e. wider and deeper neural networks), but it comes at significant training cost and infrastructure cost. Model pruning is an effective technique to reduce computation overhead for deep neural networks by removing redundant parameters. However, modern recommendation systems are still thirsty for model capacity due to the demand for handling big data. Thus, pruning a recommendation model at scale results in a smaller model capacity and consequently lower accuracy. To reduce computation cost without sacrificing model capacity, we propose a dynamic training scheme, namely alternate model growth and pruning, to alternatively construct and prune weights in the course of training. Our method leverages structured sparsification to reduce computational cost without hurting the model capacity at the end of offline training so that a full-size model is available in the recurring training stage to learn new data in real time. To the best of our knowledge, this is the first work to provide in-depth experiments and discussion of applying structural dynamics to recommendation systems at scale to reduce training cost. The proposed method is validated with an open-source deep learning recommendation model (DLRM) and state-of-the-art industrial-scale production models.
Xiaocong Du, Bhargav Bhushanam, Jiecao Yu, Dhruv Choudhary, Tianxiang Gao, Sherman Wong, Louis Feng, Jongsoo Park, Yu Cao 0001, Arun Kejariwal
ICMLA9
2021 Evolutionary NAS in Light of Model Stability for Accurate Continual Learning
abstract
Continual learning, the capability to learn new knowledge from streaming data without forgetting the previous knowledge, is a critical requirement for dynamic learning systems, especially for emerging edge devices such as self-driving cars and drones. However, continual learning is still facing the catastrophic forgetting problem. Previous work illustrate that model performance on continual learning is not only related to the learning algorithms but also strongly dependent on the inherited model, i.e., the model where continual learning starts. The better stability of the inherited model, the less catastrophic forgetting and thus, the inherited model should be elaborately selected. Inspired by this finding, we develop an evolutionary neural architecture search (ENAS) algorithm that emphasizes the Stability of the inherited model, namely ENAS-S. ENAS-S aims to find optimal architectures for accurate continual learning on edge devices. On CIFAR-10 and CIFAR-100, we present that ENAS-S achieves competitive architectures with lower catastrophic forgetting and smaller model size when learning from a data stream, as compared with handcrafted DNNs.
Xiaocong Du, Zheng Li 0020, Jingbo Sun 0003, Frank Liu 0001, Yu Cao 0001
IJCNN5
2021 SIAM: Chiplet-based Scalable In-Memory Acceleration with Mesh for Deep Neural Networks
abstract
In-memory computing (IMC) on a monolithic chip for deep learning faces dramatic challenges on area, yield, and on-chip interconnection cost due to the ever-increasing model sizes. 2.5D integration or chiplet-based architectures interconnect multiple small chips (i.e., chiplets) to form a large computing system, presenting a feasible solution beyond a monolithic IMC architecture to accelerate large deep learning models. This paper presents a new benchmarking simulator, SIAM, to evaluate the performance of chiplet-based IMC architectures and explore the potential of such a paradigm shift in IMC architecture design. SIAM integrates device, circuit, architecture, network-on-chip (NoC), network-on-package (NoP), and DRAM access models to realize an end-to-end system. SIAM is scalable in its support of a wide range of deep neural networks (DNNs), customizable to various network structures and configurations, and capable of efficient design space exploration. We demonstrate the flexibility, scalability, and simulation speed of SIAM by benchmarking different state-of-the-art DNNs with CIFAR-10, CIFAR-100, and ImageNet datasets. We further calibrate the simulation results with a published silicon result, SIMBA. The chiplet-based IMC architecture obtained through SIAM shows 130 and 72 improvement in energy-efficiency for ResNet-50 on the ImageNet dataset compared to Nvidia V100 and T4 GPUs.
Sumit K. Mandal, Manvitha Pannala, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001
ACM Trans. Embed. Comput. Syst.7
2020 Accurate Inference with Inaccurate RRAM Devices: Statistical Data, Model Transfer, and On-line Adaptation
abstract
Resistive random-access memory (RRAM) is a promising technology for in-memory computing with high storage density, fast inference, and good compatibility with CMOS. However, the mapping of a pre-trained deep neural network (DNN) model on RRAM suffers from realistic device issues, especially the variation and quantization error, resulting in a significant reduction in inference accuracy. In this work, we first extract these statistical properties from 65 nm RRAM data on 300mm wafers. The RRAM data present 10-levels in quantization and 50% variance, resulting in an accuracy drop to 31.76% and 10.49% for MNIST and CIFAR-10 datasets, respectively. Based on the experimental data, we propose a combination of machine learning algorithms and on-line adaptation to recover the accuracy with the minimum overhead. The recipe first applies Knowledge Distillation (KD) to transfer an ideal model into a student model with statistical variations and 10 levels. Furthermore, an on-line sparse adaptation (OSA) method is applied to the DNN model mapped on to the RRAM array. Using importance sampling, OSA adds a small SRAM array that is sparsely connected to the main RRAM array; only this SRAM array is updated to recover the accuracy. As demonstrated on MNIST and CIFAR-10 datasets, a 7.86% area cost is sufficient to achieve baseline accuracy for the 65 nm RRAM devices.
Gouranga Charan, Jubin Hazra, Karsten Beckmann, Xiaocong Du, Rajiv V. Joshi, Nathaniel C. Cady, Yu Cao 0001
DAC8
2020 Non-uniform DNN Structured Subnets Sampling for Dynamic Inference
abstract
With the success of Deep Neural Networks (DNN), many recent works have been focusing on developing hardware accelerator for power and resource-limited system via model compression techniques, such as quantization, pruning, low-rank approximation and etc. However, almost all existing compressed DNNs are fixed after deployment, which lacks run-time adaptive structure to adapt to its dynamic hardware resource allocation, power budget, throughput requirement, as well as dynamic workload. As the countermeasure, to construct a novel run-time dynamic DNN structure, we propose a novel DNN sub-network sampling method via non-uniform channel selection for subnets generation. Thus, user can trade off between power, speed, computing load and accuracy on-the-fly after the deployment, depending on the dynamic requirements or specifications of the given system. We verify the proposed model on both CIFAR-10 and ImageNet dataset using ResNets, which outperforms the same sub-nets trained individually and other related works. It shows that, our method can achieve latency trade-off among 13.4, 24.6, 41.3, 62.1(ms) and 30.5, 38.7, 51, 65.4(ms) for GPU with 128 batch-size and CPU respectively on ImageNet using ResNet18.
Li Yang 0009, Zhezhi He, Yu Cao 0001, Deliang Fan
DAC3
2020 MNSIM 2.0: A Behavior-Level Modeling Tool for Memristor-based Neuromorphic Computing Systems
abstract
Memristor based neuromorphic computing systems give alternative solutions to boost the computing energy efficiency of Neural Network (NN) algorithms. Because of the large-scale applications and the large architecture design space, many factors will affect the computing accuracy and system's performance. In this work, we propose a behavior-level modeling tool for memristor-based neuromorphic computing systems, MNSIM 2.0, to model the performance and help researchers to realize an early-stage design space exploration. Compared with the former version and other benchmarks, MNSIM 2.0 has the following new features: 1. In the algorithm level, MNSIM 2.0 supports the inference accuracy simulation for mixed-precision NNs considering non-ideal factors. 2. In the architecture level, a hierarchical modeling structure for PIM systems is proposed. Users can customize their designs from the aspects of devices, interfaces, processing units, buffer designs, and interconnections. 3. Two hardware-aware algorithm optimization methods are integrated in MNSIM 2.0 to realize software-hardware co-optimization.
Zhenhua Zhu 0002, Hanbo Sun, Kaizhong Qiu, Lixue Xia, Guohao Dai 0001, Dimin Niu, Xiaoming Chen 0003, Xiaobo Sharon Hu, Yu Cao 0001, Yuan Xie 0001, Yu Wang 0002, Huazhong Yang
ACM Great Lakes Symposium on VLSI10
2020 FPGA-based Low-Batch Training Accelerator for Modern CNNs Featuring High Bandwidth Memory
abstract
Training convolutional neural networks (CNNs) requires intensive computations as well as a large amount of storage and memory access. While low bandwidth off-chip memories in prior FPGA works have hindered the system-level performance, modern FPGAs offer high bandwidth memory (HBM2) that unlocks opportunities to improve the throughput/energy of FPGA-based CNN training. This paper presents a FPGA accelerator for CNN training which (1) uses HBM2 for efficient off-chip communication, and (2) supports various training operations (e.g. residual connections, stride-2 convolutions) for modern CNNs. We analyze the impact of HBM2 on CNN training workloads, provide a comprehensive comparison with DDR3, and present the strategies to efficiently use HBM2 features for enhanced CNN training performance. For training ResNet-20/VGG-like CNNs for CIFAR-10 dataset with low batch size of 2, the proposed CNN training accelerator on Intel Stratix-10 MX FPGA demonstrates 1.4/1.7X energy-efficiency improvement compared to Stratix-10 GX FPGA with DDR3 memory, and 4.5/9.7 X energy-efficiency improvement compared to Tesla V100 GPU.
Shreyas K. Venkataramanaiah, Han-Sok Suh, Shihui Yin, Eriko Nurvitadhi, Aravind Dasu, Yu Cao 0001, Jae-sun Seo
ICCAD6
2020 DAT-RNN: Trajectory Prediction with Diverse Attention
abstract
Trajectory prediction, an emerging application of spatial-temporal graph, is extremely critical in dynamic applications such as autonomous vehicles and robots. However, the diversity of trajectories and the modeling of mutual relations make it difficult to predict trajectories precisely and efficiently. In this work, we propose a novel approach, diverse attention RNN (DAT-RNN), to handle the diversity of trajectories and the accurate modeling of neighboring relations with two novel and well-designed modules: DAT-RNN first uses a diversity-aware memory (DAM) module, which is based on the detour integral of each individual, to capture the temporal behavior of each person; then DAT-RNN employs an anomaly attention module (AAM), which integrates a weighted sum of spatial relations from multiple neighbors to assist the prediction. With the well-elaborated modules, DAT-RNN integrates both temporal and spatial relations to improve the prediction under various circumstances. Comprehensive experiments on ETH and UCY datasets demonstrate the efficacy of the proposed approach.
Zheng Li 0020, Xiaocong Du, Yu Cao 0001
ICMLA3
2020 Efficient and Modularized Training on FPGA for Real-time Applications
abstract
Training of deep Convolution Neural Networks (CNNs) requires a tremendous amount of computation and memory and thus, GPUs are widely used to meet the computation demands of these complex training tasks. However, lacking the flexibility to exploit architectural optimizations, GPUs have poor energy efficiency of GPUs and are hard to be deployed on energy-constrained platforms. FPGAs are highly suitable for training, such as real-time learning at the edge, as they provide higher energy efficiency and better flexibility to support algorithmic evolution. This paper first develops a training accelerator on FPGA, with 16-bit fixed-point computing and various training modules. Furthermore, leveraging model segmentation techniques from Progressive Segmented Training, the newly developed FPGA accelerator is applied to online learning, achieving much lower computation cost. We demonstrate the performance of representative CNNs trained for CIFAR-10 on Intel Stratix-10 MX FPGA, evaluating both the conventional training procedure and the online learning algorithm.
Shreyas K. Venkataramanaiah, Xiaocong Du, Zheng Li 0020, Shihui Yin, Yu Cao 0001, Jae-sun Seo
IJCAI5
2020 Online Knowledge Acquisition with the Selective Inherited Model
abstract
Continual learning, which updates machine learning models according to streaming data, is increasingly needed in the dynamic systems. Such a scenario requires both the preservation of previous knowledge, as well as the adaptation to new observations, with high computational and memory efficiency at the edge. Previous approaches attempt to learn the knowledge class by class from scratch, using either regularization based or memory replay-based methods. However, they still suffer from severe accuracy drop, a.k.a catastrophic forgetting, during this incremental process. Moreover, as the entire model is involved in each updating, their computation cost is too expensive for edge computing. In this work, we propose a novel brain- inspired paradigm named acquisitive learning (AL). Different from previous approaches that focus only on model adaptation, AL emphasizes the importance of both knowledge inheritance and acquisition: the model is first pre-trained and selected in the cloud (the selective inherited model) and then adapted to new knowledge (the acquisition). The quality of the inherited model is monitored by the landscape of the loss function, while the acquisition is realized by segmented training. The combination of both steps reduces accuracy drop by >10× on the CIFAR- 100 dataset. Furthermore, AL benefits edge computing with 5× reduction in latency per training image on FPGA prototype and 150× reduction in training FLOPs.
Xiaocong Du, Shreyas K. Venkataramanaiah, Zheng Li 0020, Jae-sun Seo, Frank Liu 0001, Yu Cao 0001
IJCNN6
2020 GAR: Graph Assisted Reasoning for Object Detection
abstract
It is well believed that object-object relations and object-scene relations inherently improve the accuracy of object detection. However, the way to efficiently model relations remains a problem. Graph Convolutional Network (GCN), an effective method to handle structured data with relations, inspires us to leverage graphs in modeling relations for objection detection tasks. In this work, we propose a novel approach, Graph Assisted Reasoning (GAR), to utilize a heterogeneous graph in modeling object-object relations and object-scene relations. GAR fuses the features from neigh-boring object nodes as well as scene nodes and produces better recognition than that produced from individual object nodes. Moreover, compared to previous approaches using Recurrent Neural Network (RNN), the light-weight and low-coupling architecture of GAR further facilitates its integration into the object detection module. Comprehensive experiments on PASCAL VOC and MS COCO datasets demonstrate the efficacy of GAR.
Zheng Li 0020, Xiaocong Du, Yu Cao 0001
WACV3
2020 Automatic Compilation of Diverse CNNs Onto High-Performance FPGA Accelerators
abstract
A broad range of applications are increasingly benefiting from the rapid and flourishing development of convolutional neural networks (CNNs). The FPGA-based CNN inference accelerator is gaining popularity due to its high-performance and low-power as well as FPGA's conventional advantage of reconfigurability and flexibility. Without a general compiler to automate the implementation, however, significant efforts and expertise are still required to customize the design for each CNN model. In this paper, we present an register-transfer level (RTL)-level CNN compiler that automatically generates customized FPGA hardware for the inference tasks of various CNNs, in order to enable high-level fast prototyping of CNNs from software to FPGA and still keep the benefits of low-level hardware optimization. First, a general-purpose library of RTL modules is developed to model different operations at each layer. The integration and dataflow of physical modules are predefined in the top-level system template and reconfigured during compilation for a given CNN algorithm. The runtime control of layer-by-layer sequential computation is managed by the proposed execution schedule so that even highly irregular and complex network topology, e.g., GoogLeNet and ResNet, can be compiled. The proposed methodology is demonstrated with various CNN algorithms, e.g., NiN, VGG, GoogLeNet, and ResNet, on two standalone Intel FPGAs, Arria 10, and Stratix 10, achieving end-to-end inference throughputs of 969 GOPS and 1604 GOPS, respectively, with batch size of one.
Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Performance Modeling for CNN Inference Accelerators on FPGA
abstract
The recently reported successes of convolutional neural networks (CNNs) in many areas have generated wide interest in the development of field-programmable gate array (FPGA)-based accelerators. To achieve high performance and energy efficiency, an FPGA-based accelerator must fully utilize the limited computation resources and minimize the data communication and memory access, both of which are impacted and constrained by a variety of design parameters, e.g., the degree and dimension of parallelism, the size of on-chip buffers, the bandwidth of the external memory, and many more. The large design space of the accelerator makes it impractical to search for the optimal design in the implementation phase. To address this problem, a performance model is described to estimate the performance and resource utilization of an FPGA implementation. By this means, the performance bottleneck and design bound can be identified and the optimal design option can be explored early in the design phase. The proposed performance model is validated using a variety of CNN algorithms comparing the results with on-board test results on two different FPGAs.
Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Automatic Compiler Based FPGA Accelerator for CNN Training
abstract
Training of convolutional neural networks (CNNs) on embedded platforms to support on-device learning is earning vital importance in recent days. Designing flexible training hardware is much more challenging than inference hardware, due to design complexity and large computation/memory requirement. In this work, we present an automatic compiler based FPGA accelerator with 16-bit fixed-point precision for complete CNN training, including Forward Pass (FP), Backward Pass (BP) and Weight Update (WU). We implemented an optimized RTL library to perform training-specific tasks and developed an RTL compiler to automatically generate FPGA-synthesizable RTL based on user-defined constraints. We present a new cyclic weight storage/access scheme for on-chip BRAM and off-chip DRAM to efficiently implement non-transpose and transpose operations during FP and BP phases, respectively. Representative CNNs for CIFAR-10 dataset are implemented and trained on Intel Stratix 10 GX FPGA using proposed hardware architecture, demonstrating up to 479 GOPS performance.
Shreyas K. Venkataramanaiah, Yufei Ma 0002, Shihui Yin, Eriko Nurvitadhi, Aravind Dasu, Yu Cao 0001, Jae-sun Seo
FPL6
2019 Single-Net Continual Learning with Progressive Segmented Training
abstract
There is an increasing need of continual learning in dynamic systems, such as the self-driving vehicle, the surveillance drone, and the robotic system. Such a system requires learning from the data stream, training the model to preserve previous information and adapt to a new task, and generating a single-headed vector for future inference. Different from previous approaches with dynamic structures, this work focuses on a single network and model segmentation to prevent catastrophic forgetting. Leveraging the redundant capacity of a single network, model parameters for each task are separated into two groups: one important group which is frozen to preserve current knowledge, and secondary group to be saved (not pruned) for a future learning. A fixed-size memory containing a small amount of previously seen data is further adopted to assist the training. Without additional regularization, the simple yet effective approach of Progressive Segmented Training (PST) successfully incorporates multiple tasks and achieves the state-of-the-art accuracy in the single-head evaluation on CIFAR-10 and CIFAR-100 datasets. Moreover, the segmented training significantly improves computation efficiency in continual learning at the edge.
Xiaocong Du, Gouranga Charan, Frank Liu 0001, Yu Cao 0001
ICMLA4
2019 Guest Editors' Introduction to the Special Section on Hardware and Algorithms for Energy-Constrained On-chip Machine Learning
abstract
No abstract available.
Jae-sun Seo, Yu Cao 0001, Xin Li 0001, Paul N. Whatmough
ACM J. Emerg. Technol. Comput. Syst.2
2019 Guest Editors' Introduction: Hardware and Algorithms for Energy-Constrained On-Chip Machine Learning (Part 2)
abstract
No abstract available.
Jae-sun Seo, Yu Cao 0001, Xin Li 0001, Paul N. Whatmough
ACM J. Emerg. Technol. Comput. Syst.2
2018 Towards a Wearable Cough Detector Based on Neural Networks
abstract
Persistent cough is a symptom common to a number of respiratory disorders; however, reliable monitoring of cough frequency and cough severity over an extended period of time can be a challenge. Traditional methods involve subjective evaluation by care providers or patient self-reports. As an alternative, we propose an objective method for monitoring cough using a wearable microphone. We collected 24-hour audio recordings from 9 patients suffering from chronic obstructive pulmonary disease, asthma, and lung cancer using the VitaloJAK wearable microphone. Trained professionals carefully listened to each audio stream and manually labeled each cough event. Using this data, we propose a new neural-network-based cough detection scheme. A pre-processing algorithm is used to estimate the start and end of each cough and the deep neural network is trained using each cough instance. Experiments demonstrate an average leave-one-participant-out cross-validation specificity and sensitivity of 93.7% and 97.6% respectively.
Prad Kadambi, Abinash Mohanty, Jaclyn Smith, Kevin McGuinnes, Kimberly Holt, Armin Furtwaengler, Roberto Slepetys, Jae-sun Seo, Junseok Chae, Yu Cao 0001, Visar Berisha
ICASSP12
2018 Algorithm-hardware co-design of single shot detector for fast object detection on FPGAs
abstract
The rapid improvement in computation capability has made convolutional neural networks (CNNs) a great success in recent years on image classification tasks, which has also prospered the development of objection detection algorithms with significantly improved accuracy. However, during the deployment phase, many applications demand low latency processing of one image with strict power consumption requirement, which reduces the efficiency of GPU and other general-purpose platform, bringing opportunities for specific acceleration hardware, e.g. FPGA, by customizing the digital circuit specific for the inference algorithm. Therefore, this work proposes to customize the detection algorithm, e.g. SSD, to benefit its hardware implementation with low data precision at the cost of marginal accuracy degradation. The proposed FPGA-based deep learning inference accelerator is demonstrated on two Intel FPGAs for SSD algorithm achieving up to 2.18 TOPS throughput and up to 3.3× superior energy-efficiency compared to GPU.
Yufei Ma 0002, Tu Zheng, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
ICCAD3
2018 ALAMO: FPGA acceleration of deep learning algorithms with a modularized RTL compiler
Yufei Ma 0002, Naveen Suda, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
Integr.3
2018 Guest Editors' Introduction: Frontiers of Hardware and Algorithms for On-chip Learning
abstract
No abstract available.
Yu Cao 0001, Xin Li 0001, Jae-sun Seo, Ganesh Dasika
ACM J. Emerg. Technol. Comput. Syst.1
2018 Power, Performance, and Area Benefit of Monolithic 3D ICs for On-Chip Deep Neural Networks Targeting Speech Recognition
abstract
In recent years, deep learning has become widespread for various real-world recognition tasks. In addition to recognition accuracy, energy efficiency and speed (i.e., performance) are other grand challenges to enable local intelligence in edge devices. In this article, we investigate the adoption of monolithic three-dimensional (3D) IC (M3D) technology for deep learning hardware design, using speech recognition as a test vehicle. M3D has recently proven to be one of the leading contenders to address the power, performance, and area (PPA) scaling challenges in advanced technology nodes. Our study encompasses the influence of key parameters in DNN hardware implementations towards their performance and energy efficiency, including DNN architectural choices, underlying workloads, and tier partitioning choices in M3D designs. Our post-layout M3D designs, together with hardware-efficient sparse algorithms, produce power savings and performance improvement beyond what can be achieved using conventional 2D ICs. Experimental results show that M3D offers 22.3% iso-performance power saving and 6.2% performance improvement, convincingly demonstrating its entitlement as a solution for DNN ASICs. We further present architectural and physical design guidelines for M3D DNNs to maximize the benefits.
Kyungwook Chang, Deepak Kadetotad, Yu Cao 0001, Jae-sun Seo, Sung Kyu Lim
ACM J. Emerg. Technol. Comput. Syst.3
2018 MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System
abstract
Memristor-based computation provides a promising solution to boost the power efficiency of the neuromorphic computing system. However, a behavior-level memristor-based neuromorphic computing simulator, which can model the performance and realize an early stage design space exploration, is still missing. In this paper, we propose a simulation platform for the memristor-based neuromorphic system, called MNSIM. A hierarchical structure for memristor-based neuromorphic computing accelerator is proposed to provides flexible interfaces for customization. A detailed reference design is provided for large-scale applications. A behavior-level computing accuracy model is incorporated to evaluate the computing error rate affected by interconnect lines and nonideal device factors. Experimental results show that MNSIM achieves over 7000 times speed-up than SPICE simulation. MNSIM can optimize the design and estimate the tradeoff relationships among different performance metrics for users.
Lixue Xia, Boxun Li, Tianqi Tang 0001, Peng Gu 0008, Pai-Yu Chen, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Yuan Xie 0001, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2018 Optimizing the Convolution Operation to Accelerate Deep Neural Networks on FPGA
Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
IEEE Trans. Very Large Scale Integr. Syst.2
2017 A real-time 17-scale object detection accelerator with adaptive 2000-stage classification in 65nm CMOS
abstract
This paper presents an object detection accelerator that features many-scale (17), many-object (up to 50), multi-class (e.g., face, traffic sign), and high accuracy (average precision (AP) of 0.81/0.72 for AFW/BTSD datasets) detection. Employing 10 gradient/color channels, integral features are extracted and 2,000 simple classifiers for rigid boosted templates are adaptively combined to make a strong classification. The prototype chip implemented in 65nm CMOS demonstrates 16–40 frames per second and 22–160 mW power at 0.6–1.0V supply.
Minkyu Kim 0001, Abinash Mohanty, Deepak Kadetotad, Naveen Suda, Luning Wei, Pooja Saseendran, Xiaofei He 0001, Yu Cao 0001, Jae-sun Seo
ASP-DAC8
2017 Optimizing Loop Operation and Dataflow in FPGA Acceleration of Deep Convolutional Neural Networks
Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
FPGA2
2017 An automatic RTL compiler for high-throughput FPGA implementation of diverse deep convolutional neural networks
abstract
Convolutional neural networks (CNNs) are rapidly evolving and being applied to a broad range of applications. Given a specific application, an increasing challenge is to search the appropriate CNN algorithm and efficiently map it to the target hardware. The FPGA-based accelerator has the advantage of reconfigurability and flexibility, and has achieved high-performance and low-power. Without a general compiler to automate the implementation, however, significant efforts and expertise are still required to customize the design for each CNN model. In this work, we present an RTL-level CNN compiler that automatically generates customized FPGA hardware for the inference tasks of various CNNs, in order to enable high-level fast prototyping of CNNs from software to FPGA and still keep the benefits of low-level hardware optimization. First, a general-purpose library of RTL modules is developed to model different operations at each layer. The implementation of each module is optimized at the RTL level. Given a CNN algorithm, its structure is abstracted to a directed acyclic graph (DAG) and then complied with RTL modules in the library. The integration and dataflow of physical modules are predefined in the top-level system template and reconfigured during compilation. The runtime control of layer-by-layer sequential computation is managed by the proposed execution schedule so that even highly irregular and complex network topology, e.g. ResNet, can be compiled. The proposed methodology is demonstrated with end-to-end FPGA implementations of various CNN algorithms (e.g. NiN, VGG-16, ResNet-50, and ResNet-152) on two standalone Intel FPGAs, Stratix V and Arria 10. The performance and overhead of the automated compilation are evaluated. The compiled FPGA accelerators exhibit superior performance compared to state-of-the-art automation-based works by >2× for various CNNs.
Yufei Ma 0002, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
FPL2
2017 A real-time 17-scale object detection accelerator with adaptive 2000-stage classification in 65nm CMOS
abstract
This paper presents an object detection accelerator that features many-scale (17), many-object (up to 50), multi-class (e.g., face, traffic sign), and high accuracy (average precision of 0.79/0.65 for AFW/BTSD datasets). Employing 10 gradient/color channels, integral features are extracted, and the results of 2,000 simple classifiers for rigid boosted templates are adaptively combined to make a strong classification. By jointly optimizing the algorithm and the hardware architecture, the prototype chip implemented in 65nm CMOS demonstrates real-time object detection of 13-35 frames per second with low power consumption of 22-160mW at 0.58-1.0V supply.
Minkyu Kim 0001, Abinash Mohanty, Deepak Kadetotad, Naveen Suda, Luning Wei, Pooja Saseendran, Xiaofei He 0001, Yu Cao 0001, Jae-sun Seo
ISCAS8
2017 End-to-end scalable FPGA accelerator for deep residual networks
abstract
This work presents an efficient hardware accelerator design of deep residual learning algorithms, which have shown superior image recognition accuracy (>90% top-5 accuracy on ImageNet database). Two key objectives of the acceleration strategy are to (1) maximize resource utilization and minimize data movements, and (2) employ scalable and reusable computing primitives to optimize physical design under hardware constraints. Furthermore, we present techniques for efficient integration and communication of these primitives in deep residual convolutional neural networks (CNNs) that exhibit complex, non-uniform layer connections. The proposed hardware accelerator efficiently implements state-of-the-art ResNet-50/152 algorithms on Arria-10 FPGA, demonstrating 285.1/315.5 GOPS of throughput and 27.2/71.7 ms of latency, respectively.
Yufei Ma 0002, Minkyu Kim 0001, Yu Cao 0001, Sarma B. K. Vrudhula, Jae-sun Seo
ISCAS3
2017 Monolithic 3D IC designs for low-power deep neural networks targeting speech recognition
abstract
In recent years, deep learning has become widespread for various real-world recognition tasks. In addition to recognition accuracy, energy efficiency is another grand challenge to enable local intelligence in edge devices. In this paper, we investigate the adoption of monolithic 3D IC (M3D) technology for deep learning hardware design, using speech recognition as a test vehicle. M3D has recently proven to be one of the leading contenders to address the power, performance and area (PPA) scaling challenges in advanced technology nodes. Our study encompasses the influence of key parameters in DNN hardware implementations towards energy efficiency, including DNN architectural choices, underlying workloads, and tier partitioning choices in M3D. Our post-layout M3D designs, together with hardware-efficient sparse algorithms, produce power savings beyond what can be achieved using conventional 2D ICs. Experimental results show that M3D offers 22.3% iso-performance power saving, convincingly demonstrating its entitlement as a solution for DNN ASICs. We further present architectural guidelines for M3D DNNs to maximize the power saving.
Kyungwook Chang, Deepak Kadetotad, Yu Cao 0001, Jae-sun Seo, Sung Kyu Lim
ISLPED3
2017 Improving efficiency in sparse learning with the feedforward inhibitory motif
Steven Skorheim, Visar Berisha, Shimeng Yu, Jae-sun Seo, Maxim Bazhenov, Yu Cao 0001
Neurocomputing8
2017 Guest Editors' Introduction: Hardware and Algorithms for On-Chip Learning
abstract
editorial Free Access Share on Guest Editors’ Introduction: Hardware and Algorithms for On-Chip Learning Authors: Yu Cao Arizona State University, Tempe, Arizona Arizona State University, Tempe, ArizonaView Profile , Xin Li Carnegie Mellon University, Pittsburgh, PA Carnegie Mellon University, Pittsburgh, PAView Profile , Taemin Kim Intel, Hillsboro, OR Intel, Hillsboro, ORView Profile , Suyog Gupta Google, Mountain View, CA Google, Mountain View, CAView Profile Authors Info & Claims ACM Journal on Emerging Technologies in Computing SystemsVolume 13Issue 3July 2017 Article No.: 30pp 1–3https://doi.org/10.1145/3022193Published:09 February 2017Publication History 0citation335DownloadsMetricsTotal Citations0Total Downloads335Last 12 Months17Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Yu Cao 0001, Xin Li 0001, Suyog Gupta
ACM J. Emerg. Technol. Comput. Syst.1
2017 RTN in Scaled Transistors for On-Chip Random Seed Generation
abstract
Random numbers play a vital role in cryptography, where they are used to generate keys, nonce, one-time pads, and initialization vectors for symmetric encryption. The quality of random number generator (RNG) has significant implications on vulnerability and performance of these algorithms. A pseudo-RNG uses a deterministic algorithm to produce numbers with a distribution very similar to uniform. True RNGs (TRNGs), on the other hand, use some natural phenomenon/process to generate random bits. They are nondeterministic, because the next number to be generated cannot be determined in advance. In this paper, a novel on-chip noise source, random telegraph noise (RTN), is exploited for simple and reliable TRNG. RTN, a microscopic process of stochastic trapping/detrapping of charges, is usually considered as a noise and mitigated in design. Through physical modeling and silicon measurement, we demonstrate that RTN is appropriate for TRNG, especially in highly scaled MOSFETs. Due to the slow speed of RTN, we purpose the system for on-chip seed generation for random number. Our contributions are: 1) physical model calibration of RTN with comprehensive 65- and 180-nm transistor measurements; 2) the scaling trend of RTN, validated with silicon data down to 28 nm; 3) design principles to achieve 50% signal probability by using intrinsic RTN physical properties, without traditional postprocessing algorithms, the generated sequence passes the National Institute of Standards and Technology (NIST) tests; and 4) solutions to manage realistic issues in practice, including multilevel RTN signal, robustness to voltage and temperature fluctuations and the operation speed.
Abinash Mohanty, Ketul Sutaria, Hiromitsu Awano, Takashi Sato 0001, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2016 MNSIM: Simulation platform for memristor-based neuromorphic computing system
Lixue Xia, Boxun Li, Tianqi Tang 0001, Peng Gu 0008, Xiling Yin, Wenqin Huangfu, Pai-Yu Chen, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Yuan Xie 0001, Huazhong Yang
DATE9
2016 Throughput-Optimized OpenCL-based FPGA Accelerator for Large-Scale Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have gained popularity in many computer vision applications such as image classification, face detection, and video analysis, because of their ability to train and classify with high accuracy. Due to multiple convolution and fully-connected layers that are compute-/memory-intensive, it is difficult to perform real-time classification with low power consumption on today?s computing systems. FPGAs have been widely explored as hardware accelerators for CNNs because of their reconfigurability and energy efficiency, as well as fast turn-around-time, especially with high-level synthesis methodologies. Previous FPGA-based CNN accelerators, however, typically implemented generic accelerators agnostic to the CNN configuration, where the reconfigurable capabilities of FPGAs are not fully leveraged to maximize the overall system throughput. In this work, we present a systematic design space exploration methodology to maximize the throughput of an OpenCL-based FPGA accelerator for a given CNN model, considering the FPGA resource constraints such as on-chip memory, registers, computational resources and external memory bandwidth. The proposed methodology is demonstrated by optimizing two representative large-scale CNNs, AlexNet and VGG, on two Altera Stratix-V FPGA platforms, DE5-Net and P395-D8 boards, which have different hardware resources. We achieve a peak performance of 136.5 GOPS for convolution operation, and 117.8 GOPS for the entire VGG network that performs ImageNet classification on P395-D8 board.
Naveen Suda, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma 0002, Sarma B. K. Vrudhula, Jae-sun Seo, Yu Cao 0001
FPGA8
2016 Scalable and modularized RTL compilation of Convolutional Neural Networks onto FPGA
abstract
Despite its popularity, deploying Convolutional Neural Networks (CNNs) on a portable system is still challenging due to large data volume, intensive computation and frequent memory access. Although previous FPGA acceleration schemes generated by high-level synthesis tools (i.e., HLS, OpenCL) have allowed for fast design optimization, hardware inefficiency still exists when allocating FPGA resources to maximize parallelism and throughput. A direct hardware-level design (i.e., RTL) can improve the efficiency and achieve greater acceleration. However, this requires an in-depth understanding of both the algorithm structure and the FPGA system architecture. In this work, we present a scalable solution that integrates the flexibility of high-level synthesis and the finer level optimization of an RTL implementation. The cornerstone is a compiler that analyzes the CNN structure and parameters, and automatically generates a set of modular and scalable computing primitives that can accelerate various deep learning algorithms. Integrating these modules together for end-to-end CNN implementations, this work quantitatively analyzes the complier's design strategy to optimize the throughput of a given CNN model with the FPGA resource constraints. The proposed methodology is demonstrated on Altera Stratix-V GXA7 FPGA for AlexNet and NIN CNN models, achieving 114.5 GOPS and 117.3 GOPS, respectively. This represents a 1.9× improvement in throughput when compared to the OpenCL-based design. The results illustrate the promise of the automatic compiler solution for modularized and scalable hardware acceleration of deep learning.
Yufei Ma 0002, Naveen Suda, Yu Cao 0001, Jae-sun Seo, Sarma B. K. Vrudhula
FPL3
2016 Ranking the parameters of deep neural networks using the fisher information
abstract
The large number of parameters in deep neural networks (DNNs) often makes them prohibitive for low-power devices, such as field-programmable gate arrays (FPGA). In this paper, we propose a method to determine the relative importance of all network parameters by measuring the amount of information that the network output carries about each of the parameters - the Fisher Information. Based on the importance ranking, we design a complexity reduction scheme that discards unimportant parameters and assigns more quantization bits to more important parameters. For evaluation, we construct a deep autoencoder and learn a non-linear dimensionality reduction scheme for accelerometer data measuring the gait of individuals with Parkinson's disease. Experimental results confirm that the proposed ranking method can help reduce the complexity of the network with minimal impact on performance.
Visar Berisha, Martin Woolf, Jae-sun Seo, Yu Cao 0001
ICASSP5
2016 Compact oscillation neuron exploiting metal-insulator-transition for neuromorphic computing
abstract
The phenomenon of metal-insulator-transition (MIT) in strongly correlated oxides, such as NbO2, have shown the oscillation behavior in recent experiments. In this work, the MIT based two-terminal device is proposed as a compact oscillation neuron for the parallel read operation from the resistive synaptic array. The weighted sum is represented by the frequency of the oscillation neuron. Compared to the complex CMOS integrate-and-fire neuron with tens of transistors, the oscillation neuron achieves significant area reduction, thereby alleviating the column pitch matching problem of the peripheral circuitry in resistive memories. Firstly, the impact of MIT device characteristics on the weighted sum accuracy is investigated when the oscillation neuron is connected to a single resistive synaptic device. Secondly, the array-level performance is explored when the oscillation neurons are connected to the resistive synaptic array. To address the interference of oscillation between columns in simple cross-point arrays, a 2-transistor-1-resistor (2T1R) array architecture is proposed at negligible increase in array area. Finally, the circuit-level benchmark of the proposed oscillation neuron with the CMOS neuron is performed. At single neuron node level, oscillation neuron shows >12.5× reduction of area. At 128×128 array level, oscillation neuron shows a reduction of ∼4% total area, >30% latency, ∼5× energy and ∼40× leakage power, demonstrating its advantage of being integrated into the resistive synaptic array for neuro-inspired computing.
Pai-Yu Chen, Jae-sun Seo, Yu Cao 0001, Shimeng Yu
ICCAD3
2016 Bi-Level Rare Temporal Pattern Detection
abstract
Nowadays, temporal data is generated at an unprecedented speed from a variety of applications, such as wearable devices, sensor networks, wireless networks and etc. In contrast to such large amount of temporal data, it is usually the case that only a small portion of them contains information of interest. For example, for the ECG signals collected by wearable devices, most of them collected from healthy people are normal, and only a small number of them collected from people with certain heart diseases are abnormal. Furthermore, even for the abnormal temporal sequences, the abnormal patterns may only be present in a few time segments and are similar among themselves, forming a rare category of temporal patterns. For example, the ECG signal collected from an individual with a certain heart disease may be normal in most time segments, and abnormal in only a few time segments, exhibiting similar patterns. What is even more challenging is that such rare temporal patterns are often non-separable from the normal ones. Existing works on outlier detection for temporal data focus on detecting either the abnormal sequences as a whole, or the abnormal time segments directly, ignoring the relationship between abnormal sequences and abnormal time segments. Moreover, the abnormal patterns are typically treated as isolated outliers instead of a rare category with self-similarity. In this paper, for the first time, we propose a bi-level (sequence-level/ segment-level) model for rare temporal pattern detection. It is based on an optimization framework that fully exploits the bi-level structure in the data, i.e., the relationship between abnormal sequences and abnormal time segments. Furthermore, it uses sequence-specific simple hidden Markov models to obtain segment-level labels, and leverages the similarity among abnormal time segments to estimate the model parameters. To solve the optimization framework, we propose the unsupervised algorithm BIRAD, and also the semi-supervised version BIRAD-K which learns from a single labeled example. Experimental results on both synthetic and real data sets demonstrate the performance of the proposed algorithms from multiple aspects, outperforming state-of-the-art techniques on both temporal outlier detection and rare category analysis.
Dawei Zhou 0003, Jingrui He, Yu Cao 0001, Jae-sun Seo
ICDM3
2016 High-performance face detection with CPU-FPGA acceleration
abstract
Face detection is a critical function in many embedded applications, such as computer vision and security. Although face detection has been well studied, detecting a large number of faces with different scales and excessive variations (pose, expression, or illumination) usually involves computationally expensive classification algorithms. These algorithms may divide an image into sub-windows at different scales, evaluate a large set of features for each sub-window, and determine the presence and location of a face. Even with state-of-the-art CPUs, it is still challenging to perform real-time face detection with sufficiently high energy efficiency and accuracy. In this paper, we propose a suite of acceleration techniques to enable such a capability on the CPU-FPGA platform, based on a state-of-the-art face detection algorithm that employs a large number of simple classifiers. We first map the algorithm using the integrated OpenCL environment for FPGA. Matching the structure of the algorithm, a nested architecture is proposed to speed up both memory access and the computing iterations. This multi-layer architecture distributes parallel computing cores with the memory. The physical aspects of the nested architecture, such as the core size and the number of cores, are further optimized to achieve real-time face detection, under realistic hardware constraints.
Abinash Mohanty, Naveen Suda, Minkyu Kim 0001, Sarma B. K. Vrudhula, Jae-sun Seo, Yu Cao 0001
ISCAS6
2016 Design of a reliable RRAM-based PUF for compact hardware security primitives
abstract
Physical Unclonable Functions (PUF) have to be highly reliable especially when it is being used along with cryptographic hash modules for key generation. To achieve ultrahigh reliability, the conventional approach employs error correction codes (ECC) based on helper data input. Such an approach not only increases the hardware overhead of the PUF but also reduces the entropy of the system, resulting in both hardware and software security issues. In this paper we design a compact and highly reliable PUF architecture based on resistive random access memory (RRAM). We propose a new design where the sum of the read-out currents of multiple RRAM cells is used for generating one response bit. This method statistically minimizes any early-lifetime failure due to RRAM retention degradation at high temperature or under voltage stress. We employ a device model that is calibrated with IMEC HfOx RRAM experimental data and show that with 8 cells per bit, we can ensure99.9999% reliability) for a lifetime >10 years at 125°C. We embed the RRAM PUF into SHA-256 and show that the hardware overhead of the proposed RRAM PUF based architecture is significantly lower than one that uses a traditional RRAM PUF with ECC.
Ayush Shrivastava, Pai-Yu Chen, Yu Cao 0001, Shimeng Yu, Chaitali Chakrabarti
ISCAS3
2016 Technological Exploration of RRAM Crossbar Array for Matrix-Vector Multiplication
Lixue Xia, Peng Gu 0008, Boxun Li, Tianqi Tang 0001, Xiling Yin, Wenqin Huangfu, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Huazhong Yang
J. Comput. Sci. Technol.8
2015 Technological exploration of RRAM crossbar array for matrix-vector multiplication
abstract
The matrix-vector multiplication is the key operation for many computationally intensive algorithms. In recent years, the emerging metal oxide resistive switching random access memory (RRAM) device and RRAM crossbar array have demonstrated a promising hardware realization of the analog matrix-vector multiplication with ultra-high energy efficiency. In this paper, we analyze the impact of nonlinear voltage-current relationship of RRAM devices and the interconnect resistance as well as other crossbar array parameters on the circuit performance and present a design guide. On top of that, we propose a technological exploration flow for device parameter configuration to overcome the impact of nonideal factors and achieve a better trade-off among performance, energy and reliability for each specific application. The simulation results of a support vector machine (SVM) and MNIST pattern recognition dataset show that the RRAM crossbar array-based SVM is robust to the input signal fluctuation but sensitive to the tunneling gap deviation. A further resistance resolution test presents that a 4-bit RRAM device is able to realize a recognition accuracy of ∼ 90%, indicating the physical feasibility of RRAM crossbar array-based SVM. In addition, the proposed technological exploration flow is able to achieve 10.98% improvement of recognition accuracy on the MNIST dataset and 26.4% energy savings compared with previous work.
Peng Gu 0008, Boxun Li, Tianqi Tang 0001, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Huazhong Yang
ASP-DAC5
2015 Technology-design co-optimization of resistive cross-point array for accelerating learning algorithms on chip
Pai-Yu Chen, Deepak Kadetotad, Abinash Mohanty, Jieping Ye, Sarma B. K. Vrudhula, Jae-sun Seo, Yu Cao 0001, Shimeng Yu
DATE9
2015 On-chip Sparse Learning with Resistive Cross-point Array Architecture
abstract
Unsupervised learning with sparse coding is widely adopted in applications of feature extraction, pattern classification, and compressive sensing. However, even with the state-of-the-art hardware platform of CPUs/GPUs, solving a sparse coding problem is still expensive in computation. In this paper, the resistive cross-point array architecture (CPA) is proposed to achieve on-chip acceleration of sparse coding, especially the matrix/vector operations that are intensively used in the algorithm. Learning and recognition experiments are conducted with the MNIST handwriting dataset. By co-optimizing the algorithm, architecture, circuit, and resistive synaptic devices, SPICE simulation at 65nm demonstrates that the CPA is able to accelerate sparse coding computation by more than 3800X, compared to software running on an 8-core CPU. Furthermore, this work investigates the technological limitations of a realistic resistive CPA, including reduced ON/OFF range of synaptic devices, nonlinearity in programming, spatial and temporal variations, and interconnect parasitics. The results illustrate both enormous opportunities and practical barriers of resistive CPA in real-time learning on a chip.
Shimeng Yu, Yu Cao 0001
ACM Great Lakes Symposium on VLSI2
2015 Mitigating Effects of Non-ideal Synaptic Device Characteristics for On-chip Learning
abstract
The cross-point array architecture with resistive synaptic devices has been proposed for on-chip implementation of weighted sum and weight update in the training process of learning algorithms. However, the non-ideal properties of the synaptic devices available today, such as the nonlinearity in weight update, limited ON/OFF range and device variations, can potentially hamper the learning accuracy. This paper focuses on the impact of these realistic properties on the learning accuracy and proposes the mitigation strategies. Unsupervised sparse coding is selected as a case study algorithm. With the calibration of the realistic synaptic behavior from the measured experimental data, our study shows that the recognition accuracy of MNIST handwriting digits degrades from ∜97 % to ∜65 %. To mitigate this accuracy loss, the proposed strategies include 1) the smart programming schemes for achieving linear weight update; 2) a dummy column to eliminate the off-state current; 3) the use of multiple cells for each weight element to alleviate the impact of device variations. With the improved synaptic behavior by these strategies, the accuracy increases back to ∜95 %, enabling the reliable integration of realistic synaptic devices in the neuromorphic systems.
Pai-Yu Chen, I-Ting Wang, Tuo-Hung Hou, Jieping Ye, Sarma B. K. Vrudhula, Jae-sun Seo, Yu Cao 0001, Shimeng Yu
ICCAD8
2015 Energy-efficient reconstruction of compressively sensed bioelectrical signals with stochastic computing circuits
abstract
Compressive sensing (CS) allows acquiring sparse signals at sub-Nyquist rate, offering an energy-efficient solution to data acquisition. This is especially important to reduce communication data for mobile medical applications. However, reconstructing the signal from CS is usually left off-line due to the complex computations. In this paper, we integrate two key technologies to enable on-line energy-efficient CS signal reconstruction. These are (1) the use of Bayesian CS Belief Propagation (CS-BP) as the algorithm basis and (2) the novel design of stochastic computing (SC) circuits to efficiently map CS-BP algorithm. The overall signal reconstruction system is implemented with digital SC circuits in 65nm CMOS and recovers compressively sensed electrocardiography (ECG) and electromyography (EMG) signals with 11X to 8X data compression factor. Compared to a conventional binary design, post-layout simulation results show that the proposed stochastic design performs reconstruction with 5X energy-delay product improvement and 2X area reduction.
Yufei Ma 0002, Minkyu Kim 0001, Yu Cao 0001, Jae-sun Seo, Sarma B. K. Vrudhula
ICCD3
2015 Optimizing latency, energy, and reliability of 1T1R ReRAM through appropriate voltage settings
abstract
Resistive RAM (ReRAM) has fast access time, ultra-low stand-by power and high reliability, making it a viable memory technology to replace DRAM for main memory. The 1-transistor-1-resistor (1T1R) ReRAM array has density comparable to that of a DRAM array and the advantages of lower programming energy and higher reliability compared to the ultrahigh density ReRAM cross-point array. In this paper, we show how circuit operation parameters, such as the pulse amplitude and pulse widths of word-line (WL) voltage, bit-line (BL) voltage, and source-line (SL) voltage can be used to lower latency, lower power and improve reliability. SPICE simulation results demonstrate that appropriate choice of voltage settings can be used to reduce the write latency of the 1T1R cell by 29.4% and reduce write energy by 46.7% over the DRAM cell. Next, we show how the endurance of ReRAM cell can be improved by increasing the ratio between OFF and ON resistances and reducing SL voltage. We find that of these, reducing the SL voltage results in significant improvement in endurance with smaller energy overhead. Next, we evaluate the system-level performance of a 1GB ReRAM and DRAM memory system using CACTI and GEM5. Simulation results using SPEC CPU INT 2006 and DaCapo-9.12 benchmarks show that the ReRAM based main memory can improve IPC by 4.2% and energy by up to 77.8% compared to a DRAM system.
Manqing Mao, Yu Cao 0001, Shimeng Yu, Chaitali Chakrabarti
ICCD2
2015 Impact of temporal transistor variations on circuit reliability
abstract
With the ever-increasing importance of temporal transistor variations during circuit run time and aging, this paper focuses on impacts of the two major temporal effects: the Bias Temperature Instability (BTI) and Random Telegraph Noise (RTN), illustrating their scaling trend, challenges, and potential solutions for future design robustness.
Runsheng Wang, Yu Cao 0001
ISCAS2
2015 Finite-point method for efficient timing characterization of sequential elements
Anupama R. Subramaniam, Janet Roveda, Yu Cao 0001
Integr.3
2015 A Finite-Point Method for Efficient Gate Characterization Under Multiple Input Switching
abstract
Timing characterization of standard cells is one of the essential steps in VLSI design. The traditional static timing analysis (STA) tool assumes single input switching models for the characterization of multiple input gates. However, due to technology scaling, increasing operating frequency, and process variation, the probability of the occurrence of multiple input switching (MIS) is increasing. On the other hand, considering all possible MIS scenarios for the characterization of multiple input logic gates, is computationally intensive. To improve the efficiency, this work proposes a finite-point-based characterization methodology for multiple input gates with the effects of MIS. Furthermore, delay variation due to MIS is integrated into the STA flow through propagation of switching windows. The proposed modeling methodology is validated using benchmark circuits at the 45nm technology node for various operating conditions. Experimental results demonstrate significant reduction in computation cost and data volume with less than ∼10% error compared to that of traditional SPICE simulation.
Anupama R. Subramaniam, Janet Roveda, Yu Cao 0001
ACM Trans. Design Autom. Electr. Syst.3
2014 Statistical analysis of random telegraph noise in digital circuits
abstract
Random telegraph noise (RTN) has become an important reliability issue at the sub-65nm technology node. Existing RTN simulation approaches mainly focus on single trap induced RTN and transient response of RTN, which are usually time-consuming for circuit-level simulation. This paper proposes a statistical algorithm to study multiple traps induced RTN in digital circuits, to show the temporal distribution of circuit delay under RTN. Based on the simulation results we show how to protect circuit from RTN. Bias dependence of RTN is also discussed.
Xiaoming Chen 0003, Yu Wang 0002, Yu Cao 0001, Huazhong Yang
ASP-DAC3
2014 BTI-Induced Aging under Random Stress Waveforms: Modeling, Simulation and Silicon Validation
abstract
The BTI effect, which consists of both stress and recovery phases, poses a unique challenge to long-term aging prediction, because the degradation rate strongly depends on the stress pattern. Previous approaches usually resort to an average, constant stress waveform to simplify the situation. They are efficient, but fail to capture the reality, especially under dynamic voltage scaling (DVS) or in analog/mixed signal designs where the stress waveform is much more random. This paper presents a suite of solutions that enable aging simulation under all possible stress conditions. Key contributions include: (1) Compact modeling of BTI when the stress voltage is varying. The results to both reaction-diffusion (RD) and trapping/detrapping (TD) mechanisms are derived. (2) Efficient simulation under DVS, leveraging the new BTI models; (3) Silicon validation at 45nm and 65nm, at both device and circuit levels. As the results illustrate, it is necessary to combine both RD and TD mechanisms to accurately predict aging under changing stress voltages. Our proposed work provides a general and comprehensive solution to aging analysis under random stress patterns.
Ketul Sutaria, Athul Ramkumar, Rongjun Zhu, Renju Rajveev, Yu Cao 0001
DAC6
2014 The Stochastic Loss of Spikes in Spiking Neural P Systems: Design and Implementation of Reliable Arithmetic Circuits
abstract
Spiking neural P systems (in short, SN P systems) have been introduced as computing devices inspired by the structure and functioning of neural cells. The presence of unreliable components in SN P systems can be considered in many different aspects.
Matteo Cavaliere, Pei An, Sarma B. K. Vrudhula, Yu Cao 0001
Fundam. Informaticae5
2014 Cross-Layer Modeling and Simulation of Circuit Reliability
abstract
Integrated circuit design in the late CMOS era is challenged by the ever-increasing variability and reliability issues. The situation is further compounded by real-time uncertainties in workload and ambient conditions, which dynamically influence the degradation rate. To improve design predictability and guarantee system lifetime, accurate modeling, and simulation tools for reliability are essential to both digital and analog circuits. This paper presents cross-layer solutions for emerging reliability threats, including: 1) device-level modeling of reliability mechanisms, such as transistor aging and its statistical behavior; 2) circuit-level long-term aging models that capture unique operation patterns in digital and analog design, and directly predict the degradation; and 3) simulation methods for very-large-scale designs. Built on the long-term model, the new methods significantly enhance the accuracy and efficiency of reliability analysis. As validated by silicon data, these solutions close the gap between the underlying reliability physics and circuit/system design for resilience.
Yu Cao 0001, Jyothi Velamala, Ketul Sutaria, Mike Shuo-Wei Chen, Jonathan Ahlbin, Ivan Sanchez Esqueda, Michael Bajura, Michael Fritze
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2013 NBTI-aware circuit node criticality computation
abstract
For sub-65nm technology nodes, Negative Bias Temperature Instability (NBTI) has become a primary limiting factor of circuit lifetime. During the past few years, researchers have spent considerable effort on accurate modeling and characterization of circuit delay degradation caused by NBTI at different design levels. The search for techniques and methodologies which can aid in effectively minimizing the NBTI effect on circuit delay is still underway. In this work, we present the usage of node criticality computation to drive NBTI-aware timing analysis and optimization. Circuits that have undergone this optimization flow show strong resistance to NBTI delay degradation. For the first time, this work proposes a node criticality computation algorithm under an NBTI-aware timing analysis and optimization framework. Our work provides answers to the following yet unaddressed questions: (a) what is the definition of node criticality in a circuit under the NBTI effect? (b) how do we identify the critical nodes that, once protected, will be immune to NBTI timing degradation? and (c) what are the NBTI effect attenuation approaches? Experimental results indicate that by protecting the critical nodes found by our proposed methodology, circuit delay degradation can be reduced by up to 50%. Combined with peak temperature reduction, the delay degradation can be be further improved.
Shengqi Yang, Wenping Wang 0004, Mark Hagan, Wei Zhang 0012, Pallav Gupta, Yu Cao 0001
ACM J. Emerg. Technol. Comput. Syst.6
2012 Exploring sub-20nm FinFET design with predictive technology models
abstract
Predictive MOSFET models are critical for early stage design-technology co-optimization and circuit design research. In this work, Predictive Technology Model files for sub-20nm multi-gate transistors have been developed (PTM-MG). Based on MOSFET scaling theory, the 2011 ITRS roadmap and early stage silicon data from published results, PTM for FinFET devices are generated for 5 technology nodes corresponding to the years 2012-2020 on the ITRS roadmap.
Saurabh Sinha 0001, Greg Yeric, Vikas Chandra, Brian Cline, Yu Cao 0001
DAC5
2012 Physics matters: statistical aging prediction under trapping/detrapping
abstract
Randomness in Negative Bias Temperature Instability (NBTI) process poses a dramatic challenge on reliability prediction of digital circuits. Accurate statistical aging prediction is essential in order to develop robust guard banding and protection strategies during the design stage. Variations in device level and supply voltage due to Dynamic Voltage Scaling (DVS) need to be considered in aging analysis. The statistical device data collected from 65nm test chip shows that degradation behavior derived from trapping/detrapping mechanism is accurate under statistical variations compared to conventional Reaction Diffusion (RD) theory. The unique features of this work include (1) Aging model development as a function of technology parameters based on trapping/detrapping theory (2) Reliability prediction under device variations and DVS with solid validation with using 65nm statistical silicon data (3) Asymmetric aged timing analysis under NBTI and comprehensive evaluation of our framework in ISCAS89 sequential circuits. Further, we show that RD based NBTI model significantly overestimates the degradation and TD model correctly captures aging variability. These results provide design insights under statistical NBTI aging and enhance the prediction efficiency.
Jyothi Velamala, Ketul Sutaria, Takashi Sato 0001, Yu Cao 0001
DAC4
2012 Hierarchical modeling of Phase Change memory for reliable design
abstract
As CMOS based memory devices near their end, memory technologies, such as Phase Change Random Access Memory (PRAM), have emerged as viable alternatives. This work develops a hierarchical modeling framework that connects the unique device physics of PRAM with its circuit and state transition properties. Such an approach enables design exploration at various levels in order to optimize the performance and yield. By providing a complete set of compact models, it supports SPICE simulation of PRAM in the presence of process variations and temporal degradation. Furthermore, this work proposes a new metric, State Transition Curve (STC) that supports the assessment of other performance metrics (e.g., power, speed, yield, etc.), helping gain valuable insights on PRAM reliability.
Ketul Sutaria, Chengen Yang, Chaitali Chakrabarti, Yu Cao 0001
ICCD5
2012 Design benchmarking to 7nm with FinFET predictive technology models
abstract
The coming ten years promise great changes in silicon technology, with the end of planar bulk CMOS and the rise of interconnect parasitics to true significance. With such shifts in the underlying technology, the simple extrapolation of performance metrics may lead to pronounced prediction errors in design pathfinding. In this work, we utilize newly developed Predictive Technology Models for FinFETs aligned to the 2011 ITRS. Together with predictive interconnect models, we project performance and power landscape for the technology nodes from 20nm to 7nm. We present an overview of models, assess the advantage of FinFET over bulk CMOS devices, benchmark the scaling of critical design metrics, and illustrate major design barriers toward the 7nm node.
Saurabh Sinha 0001, Brian Cline, Greg Yeric, Vikas Chandra, Yu Cao 0001
ISLPED5
2012 A self-tuning design methodology for power-efficient multi-core systems
abstract
This article aims to achieve computational reliability and energy efficiency through codevelopment of algorithms, device, and circuit designs for application-specific, reconfigurable architectures. The new methodology characterizes aging-switching activity and aging-supply voltage relationships that are applicable for minimizing power consumption and task execution efficiency in order to achieve low bit energy ratio (BER). In addition, a new dynamic management algorithm (DMA) is proposed to alleviate device degradation and to extend system lifespan. In contrast to traditional workload balancing schemes in which cores are regarded as homogeneous, the new algorithm ranks cores as “highly competitive,” “less competitive,” and “not competitive” according to their various competitiveness. Core competitiveness is evaluated based upon their reliability, temperature, and timing requirements. Consequently, “competitive” cores will take charge of the majority of the tasks at relatively high voltage/frequency without violating power and timing budgets, while “not competitive” cores will have light workloads to ensure their reliability. The new approach combines intrinsic device characteristics (aging-switching activity and aging-supply voltage curves) into an integrated framework to achieve high reliability and low energy level with graceful degradation of system performance. Experimental results show that the proposed method has achieved up to 20% power reduction, with about 4% performance degradation (in terms of accomplished workload and system throughput), compared with traditional workload balancing methods. The new method also improves system mean-time-to-failure (MTTF) by up to 25%.
Jin Sun 0006, Jyothi Velamala, Yu Cao 0001, Roman L. Lysecky, Karthik Shankar, Janet Roveda
ACM Trans. Design Autom. Electr. Syst.4
2012 Variation-Aware Supply Voltage Assignment for Simultaneous Power and Aging Optimization
abstract
As technology scales, negative bias temperature instability (NBTI) has become a major reliability concern for circuit designers. And the growing process variations can no longer be ignored. Meanwhile, reducing power consumption remains to be one of the design goals. In this paper, a variation-aware supply voltage assignment (SVA) technique combining dual$V_{dd}$assignment and dynamic$V_{dd}$scaling is proposed on a statistical platform, to minimize circuit power under an aging-aware timing constraint. The experimental results show that our SVA technique can mitigate on average 62% of the NBTI-induced circuit delay degradation. Compared with guard-banding and single$V_{dd}$scaling approaches, our approach saves more energy.
Xiaoming Chen 0003, Yu Wang 0002, Yu Cao 0001, Yuchun Ma, Huazhong Yang
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Design sensitivity of single event transients in scaled logic circuits
abstract
Single Event Transients (SET) in digital logic pose an ever increasing reliability challenge as device dimensions shrink in modern technologies. Projection of SET sensitivity with scaling is essential to assess the logic failure and error probability in modern technology generations. This paper discusses the effects of device scaling from 45nm to 12nm processes and circuit parameter tuning on SETs. The failure due to particle strikes i.e., Single Event upsets (SEU) as well as its behavior with process variations and reliability mechanisms such as NBTI is evaluated in this work. The critical supply voltage required to avoid SET propagation with circuit parameters is investigated. This work also proposes a probability model which examines the propagation of SET at any node to the output of a circuit. The proposed methodology can be extended to any complex digital circuit to investigate its vulnerability to SET.
Jyothi Velamala, Robert LiVolsi, Myra Torres, Yu Cao 0001
DAC4
2011 Programmable analog device array (PANDA): a platform for transistor-level analog reconfigurability
abstract
The design and development of analog/mixed-signal (AMS) ICs is becoming increasingly expensive, complex, and lengthy. Lacking a reconfigurable platform, analog designers are denied the benefits of rapid prototyping, hardware emulation, and smooth migration to advanced technology nodes. To overcome these limitations, this work proposes a new approach that maps any AMS design problem to a transistor-level reconfigurable vehicle, thus enabling fast validation and a reduction in post-Silicon bugs, and minimizing design risk and costs. The unique features of the approach include: (1) transistor-level programmability that emulates each transistor behavior in an analog design, reproducing the system and achieving very fine granularity of reconfiguration; (2) programmable switches that are treated as a design component during analog transistor mapping, and optimized with the reconfiguration matrix; (3) parasitics reduction that leverages the aggressive scaling of CMOS technology. Based on these principles, a digitally controlled PANDA platform is designed at a 32nm node. Several 90nm analog blocks are successfully emulated with the 32nm platform, including a folded-cascode operational amplifier, a sample-and-hold module (S/H), and a voltage-controlled oscillator (VCO). A solid basis to future efforts on the architecture, hierarchical optimization, and related design automation tools is demonstrated.
Jounghyuk Suh, Nagib Hakim, Bertan Bakkaloglu, Yu Cao 0001
DAC6
2011 Failure diagnosis of asymmetric aging under NBTI
abstract
Design for reliability is becoming an important step in the design cycle with CMOS technology scaling, demanding need for efficient and accurate reliability simulation methods in the design stage. Traditional aging analysis does not differentiate NBTI induced delay shift in rising and falling edges, thereby assuming averaging effect due to recovery. It is essential to identify the critical operation conditions that are more susceptible to timing violations under aging. In this paper, by identifying the critical moments in circuit operation and considering the asymmetric aging effects, timing violations under NBTI effect are correctly predicted. The unique features of this work include: (1) delay modeling of a digital gate due to threshold voltage (Vth) shift using delay dependence on supply voltage from cell library; (2) asymmetric aging analysis is conducted by recognizing the critical points in circuit operation; and (3) setup and hold timing violations due to NBTI induced path delay shift in logic and clock buffer are investigated. This failure assessment method is further demonstrated in ISCAS89 benchmark circuits using 45nm Nangate standard cell library to extract aging information in critical paths. The proposed failure diagnosis enables resilient design techniques to mitigate circuit aging under NBTI.
Jyothi Velamala, Venkatesa Ravi, Yu Cao 0001
ICCAD3
2011 Self-Tuning for Maximized Lifetime Energy-Efficiency in the Presence of Circuit Aging
abstract
This paper presents an integrated framework, together with control policies, for optimizing dynamic control of self-tuning parameters of a digital system over its lifetime in the presence of circuit aging. A variety of self-tuning parameters such as supply voltage, operating clock frequency, and dynamic cooling are considered, and jointly optimized using efficient algorithms described in this paper. Our optimized self-tuning approach satisfies performance constraints at all times, and maximizes a lifetime computational power efficiency (LCPE) metric, which is defined as the total number of clock cycles achieved over lifetime divided by the total energy consumed over lifetime. We present three control policies: 1) progressive-worst-case-aging (PWCA), which assumes worst-case aging at all times; 2) progressive-on-state-aging (POSA), which estimates aging by tracking active/sleep modes, and then assumes worst-case aging in active mode and long recovery effects in sleep mode; and 3) progressive-real-time-aging-assisted (PRTA), which acquires real-time information and initiates optimized control actions. Various flavors of these control policies for systems with dynamic voltage and frequency scaling (DVFS) are also analyzed. Simulation results on benchmark circuits, using aging models validated by 45 nm measurements, demonstrate the effectiveness and practicality of our approach in significantly improving LCPE and/or lifetime compared to traditional one-time worst-case guardbanding. We also derive system design guidelines to maximize self-tuning benefits.
Evelyn Mintarno, Joëlle Skaf, Jyothi Velamala, Yu Cao 0001, Stephen P. Boyd, Robert W. Dutton, Subhasish Mitra
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2011 Leakage Power and Circuit Aging Cooptimization by Gate Replacement Techniques
abstract
As technology scales, the aging effect caused by negative bias temperature instability (NBTI) has become a major reliability concern. In the mean time, reducing leakage power remains to be one of the key design goals. Because both NBTI-induced circuit degradation and standby leakage power have a strong dependency on the input vectors, input vector control (IVC) technique could be adopted to reduce the leakage power and mitigate NBTI-induced degradation. The IVC technique, however, is ineffective for larger circuits. Consequently, in this paper, we propose two gate replacement algorithms [direct gate replacement (DGR) algorithm and divide and conquer-based gate replacement (DCBGR) algorithm], together with optimal input vector selection, to simultaneously reduce the leakage power and mitigate NBTI-induced degradation. Our experimental results on 23 benchmark circuits reveal the following. 1) Both DGR and DCBGR algorithms outperform pure IVC technique by 15%–30% with 5% delay relaxation for three different design goals: leakage power reduction only, NBTI mitigation only, and leakage/NBTI cooptimization. 2) The DCBGR algorithm leads to better optimization results and save on average more than 10$\times$runtime compared to the DGR algorithm. 3) The area overhead for leakage reduction is much more than that for NBTI mitigation.
Yu Wang 0002, Xiaoming Chen 0003, Wenping Wang 0004, Yu Cao 0001, Yuan Xie 0001, Huazhong Yang
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Statistical Modeling and Simulation of Threshold Variation Under Random Dopant Fluctuations and Line-Edge Roughness
abstract
The threshold voltage (Vth) of a nanoscale transistor is severely affected by random dopant fluctuations and line-edge roughness. The analysis of these effects usually requires atomistic simulations which are too expensive in computation for statistical design. In this work, we develop an efficient SPICE simulation method and statistical variation model that accurately predict threshold variation as a function of dopant fluctuations and gate length change caused by lithography and the etching process. By understanding the physical principles of atomistic simulations, we: 1) identify the appropriate method to divide a nonuniform gate into slices in order to map those fluctuations into the device model; 2) extract the variation ofVthfrom the strong-inversion region instead of the leakage current, benefiting from the linearity of the saturation current with respect toVth; 3) propose a compact model ofVthvariation that is scalable with gate size and the amount of dopant and gate length fluctuations; and 4) investigate the interaction with non-rectangular gate (NRG) and reverse narrow width effect (RNWE). The proposed SPICE simulation method is validated with atomistic simulation results. Given the post-lithography gate geometry, this approach correctly models the variation of device output current in all operating regions. Based on the new results, we further project the amount ofVthvariation at advanced technology nodes, helping shed light on the challenges of future robust circuit design.
Yun Ye 0001, Frank Liu 0001, Min Chen 0024, Sani R. Nassif, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2010 In-situ characterization and extraction of SRAM variability
abstract
Measurement and extraction of as fabricated SRAM cell variability is essential to process improvement and robust design. This is challenging in practice, due to the complexity in the test procedure and requisite numerical analysis. This work proposes a new singleended test procedure for SRAM cell write margin measurement. Moreover, an efficient decomposition method is developed to extract transistor threshold voltage (VTH) variations from the measurements, allowing accurate determination of SRAM cell stability. The entire approach is demonstrated in a 90nm test chip with 32K cells. The advantages of the proposed method include: (1) a single-ended SRAM test structure with no disturbance to SRAM operations; (2) a convenient test procedure that only requires quasistatic control of external voltages; and (3) a non-iterative method that extracts the VTH variation of each transistor from eight measurements. The new procedure enables accurate predictions of SRAM performance variability. As validated with 90nm data of write margin and data retention voltage, the prediction error from extracted VTH variations is < 4% at all corners.
Srivatsan Chellappa, Jia Ni, Xiaoyin Yao, Nathan D. Hindman, Jyothi Velamala, Min Chen 0024, Yu Cao 0001, Lawrence T. Clark
DAC7
2010 Optimized self-tuning for circuit aging
abstract
We present a framework and control policies for optimizing dynamic control of various self-tuning parameters over lifetime in the presence of circuit aging. Our framework introduces dynamic cooling as one of the self-tuning parameters, in addition to supply voltage and clock frequency. Our optimized self-tuning satisfies performance constraints at all times and maximizes a lifetime computational power efficiency (LCPE) metric, which is defined as the total number of clock cycles achieved over lifetime divided by the total energy consumed over lifetime. Our framework features three control policies: 1. Progressive-worst-case-aging (PWCA), which assumes worst-case aging at all times; 2. Progressive-on-state-aging (POSA), which estimates aging by tracking active/sleep mode, and then assumes worst-case aging in active mode and long recovery effects in sleep mode; 3. Progressive-real-time-aging-assisted (PRTA), which estimates the actual amount of aging and initiates optimized control action. Simulation results on benchmark circuits, using aging models validated by 45nm CMOS stress measurements, demonstrate the practicality and effectiveness of our approach. We also analyze design constraints and derive system design guidelines to maximize self-tuning benefits.
Evelyn Mintarno, Joëlle Skaf, Jyothi Velamala, Yu Cao 0001, Stephen P. Boyd, Robert W. Dutton, Subhasish Mitra
DATE5
2010 A resilience roadmap
abstract
Technology scaling has an increasing impact on the resilience of CMOS circuits. This outcome is the result of (a) increasing sensitivity to various intrinsic and extrinsic noise sources as circuits shrink, and (b) a corresponding increase in parametric variability causing behavior similar to what would be expected with hard (topological) faults. This paper examines the issue of circuit resilience, then proposes and demonstrates a roadmap for evaluating fault rates starting at the 45 nm and going down to the 12 nm nodes. The complete infrastructure necessary to make these predictions is placed in the open source domain, with the hope that it will invigorate research in this area.
Sani R. Nassif, Nikil Mehta, Yu Cao 0001
DATE3
2010 A self-evolving design methodology for power efficient multi-core systems
abstract
This paper introduces a new methodology that characterizes aging-duty cycle and aging-supply voltage relationships that are applicable to minimizing power consumption and task execution time to achieve low Bit-Energy-Ratio (BER). In contrast to the traditional workload balancing scheme where cores are regarded as homogeneous, we proposed a new task scheduler that ranks cores according to their various competitiveness evaluated based upon their reliability, temperature and timing requirements. Consequently, the new approach combines internal characteristics (aging-duty cycle and aging-supply voltage curves) into an integrated framework to achieve system performance improvement or graceful degradation with high reliability and low power. Experimental results show that the proposed method has achieved 18% power reduction with about 4% performance degradation (in terms of accomplished workload) compared with traditional workload balancing methods.
Jin Sun 0006, Jyothi Velamala, Yu Cao 0001, Roman L. Lysecky, Karthik Shankar, Janet Roveda
ICCAD4
2010 Simulation of random telegraph Noise with 2-stage equivalent circuit
abstract
With the continuous reduction of CMOS device dimension, the importance of Random Telegraph Noise (RTN) keeps growing. To determine its impact on circuit performance and optimize the design, it is essential to physically model RTN effect and embed it into the standard simulation environment. In this paper, a new simulation method of time domain RTN effect is proposed to benchmark important digital circuits: (1) A two-stage L-shaped circuit is proposed to generate RTN signal by integrating a white noise source. An L-shaped circuit is a RC filter connected with an ideal comparator, where RC values are calibrated with the physical property of RTN; (2) This sub-circuit is fully compatible with SPICE, enabling the time domain analysis in nanometer scale digital design; (3) The importance of discrete RTN is demonstrated on a 32nm SRAM design and a 22nm low power ring oscillator (RO), using the proposed method. As compared to traditional 1/f noise, the impact of RTN is more significant under low voltages, leading to tremendous differences in the prediction of Vccminand failure probability in SRAM, as well as jitter noise in RO.
Yun Ye 0001, Chi-Chao Wang, Yu Cao 0001
ICCAD3
2010 Workload-adaptive process tuning strategy for power-efficient multi-core processors
abstract
As more devices are integrated with technology scaling, reducing the power consumption of both high-performance and low-power processors has become the first-class design constraint. Reducing power consumption while satisfying required performance is critical for increasing the operating time of mobile devices and lowering the operating cost of offices and data centers. Meanwhile, dynamic voltage and frequency scaling (DVFS) and clock-gating (CG) techniques have been widely used for two of the most powerful techniques to reduce the power consumption of such processors. Depending on performance and power demands, a processor runs at various performance and power states to trade power with performance. In this paper, we propose process tuning strategy to minimize the average power consumption of multi-core processors that use the DVFS and CG techniques, while providing the same maximum performance. The proposed optimization method incorporates with workload characteristics of commercial high-performance and low-power multi-core processors. The experimental results show that our optimized 32nm technologies for workstation, mobile, and server multi-core processors minimize the average power by up to 13, 18, and 9%, respectively.
Jungseob Lee, Chi-Chao Wang, Hamid Reza Ghasemi, Lloyd Bircher, Yu Cao 0001, Nam Sung Kim
ISLPED5
2010 Workload-aware neuromorphic design of low-power supply voltage controller
abstract
A workload-aware low-power neuromorphic controller for dynamic voltage scaling in VLSI systems is presented. The neuromorphic controller predicts future workload values and preemptively regulates supply voltage based on past workload profile. Our specific contributions include: (1) implementation of a digital and analog version of the controller in 45nm CMOS technology, resulting in 3% performance hit with a power overhead in the range of 10-150 microwatts, (2) higher prediction accuracy compared to a software based OS-governed DVS scheme by 50%, reducing wasted power and improving error margins, (3) digital design has minimal power overhead and is more reconfigurable, while analog design is better suited for nonlinear and complex computational tasks.
Saurabh Sinha 0001, Jounghyuk Suh, Bertan Bakkaloglu, Yu Cao 0001
ISLPED4
2010 Modeling and Analysis of the Nonrectangular Gate Effect for Postlithography Circuit Simulation
abstract
For nanoscale CMOS devices, gate roughness has severe impact on the deviceI-Vcharacteristics, particularly in the subthreshold region. In particular, the nonrectangular gate (NRG) geometries are caused by subwavelength lithography and have relatively low spatial frequency. In this paper, we present an analytical approach to model NRG effects onI-Vcharacteristics. To predict the change ofI-Vcharacteristics due to the NRG effect, the proposed model converts the postlithography gate profile into an equivalent gate length (Le) , which is a function of the gate bias voltage but independent of the drain bias voltage. We demonstrate the accuracy of this approach by comparing it to TCAD simulation results for 65-nm technology. The newLemodel is readily integrated into standard transistor models in traditional circuit simulation tools, such as SPICE, for both dc and transient analyses. We further develop a generic procedure to systematically extract theLevalue from the postlithography gate profile. The interaction with the narrow-width effect is also efficiently incorporated into the proposed algorithm. TCAD verification demonstrates that the proposedLemodel is simple for implementation, scalable with both transistor geometries and bias conditions, and also continuous across all the operation regions.
Ritu Singhal, Asha Balijepalli, Anupama R. Subramaniam, Chi-Chao Wang, Frank Liu 0001, Sani R. Nassif, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.7
2010 The Impact of NBTI Effect on Combinational Circuit: Modeling, Simulation, and Analysis
abstract
Negative-bias-temperature instability (NBTI) has become the primary limiting factor of circuit life time. In this paper, we develop a hierarchical framework for analyzing the impact of NBTI on the performance of logic circuits under various operation conditions, such as the supply voltage, temperature, and node switching activity. Given a circuit topology and input switching activity, we propose an efficient method to predict the degradation of circuit speed over a long period of time. The effectiveness of our method is comprehensively demonstrated with the International Symposium on Circuits and Systems (ISCAS) benchmarks and a 65-nm industrial design. Furthermore, we extract the following key design insights for reliable circuit design under NBTI effect, including: 1) During dynamic operation, NBTI-induced degradation is relatively insensitive to supply voltage, but strongly dependent on temperature; 2) There is an optimum supply voltage that leads to the minimum of circuit performance degradation; circuit degradation rate actually goes up if supply voltage is lower than the optimum value; 3) Circuit performance degradation due to NBTI is highly sensitive to input vectors. The difference in delay degradation is up to 5× for various static and dynamic operations. Finally, we examine the interaction between NBTI effect, and process and design uncertainty in realistic conditions.
Wenping Wang 0004, Shengqi Yang, Sarvesh Bhardwaj, Sarma B. K. Vrudhula, Frank Liu 0001, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2009 A framework for estimating NBTI degradation of microarchitectural components
abstract
Degradation of device parameters over the lifetime of a system is emerging as a significant threat to system reliability. Among the aging mechanisms, wearout resulting from NBTI is of particular concern in deep submicron technology generations. To facilitate architectural level aging analysis, a tool capable of evaluating NBTI vulnerabilities early in the design cycle has been developed. The tool includes workload-based temperature and performance degradation analysis across a variety of technologies and operating conditions, revealing a complex interplay between factors influencing NBTI timing degradation.
Michael DeBole, Krishnan Ramakrishnan, Varsha Balakrishnan, Wenping Wang 0004, Yu Wang 0002, Yuan Xie 0001, Yu Cao 0001, Narayanan Vijaykrishnan
ASP-DAC8
2009 Variability analysis under layout pattern-dependent rapid-thermal annealing process
abstract
Rapid-Thermal Annealing (RTA) with radiation heating is recently adopted in nanoscale CMOS fabrication in order to achieve ultra-shallow junction with maximum dopant activation rate. However, recent results report the systematic shift of threshold voltage (Vth) and increased Vth variation due to RTA process [1--2]. The exact amount of variations depends on layout pattern density, RTA heating temperature (T) and effective annealing time. In this work, we develop joint thermal/TCAD simulation and compact modeling tools to analyze performance variability under various layout pattern densities and RTA conditions. With the new simulation capability, we recognize two major variation mechanisms under RTA: the change of effective channel length (Leff) induced by lateral dopant diffusion, and the fluctuation of equivalent oxide thickness (EOT) due to incomplete dopant activation. We perform device simulations to quantify transistor performance shift due to Leff and EOT variations. Moreover, we propose a suite of compact models that bridge the underlying RTA process with device parameter change for efficient design optimization. The new tools are validated with published silicon data at 45nm and 65nm nodes. They will facilitate physical designers to predict and mitigate circuit performance variability due to the layout-dependent RTA process.
Yun Ye 0001, Frank Liu 0001, Min Chen 0024, Yu Cao 0001
DAC4
2009 Gate replacement techniques for simultaneous leakage and aging optimization
abstract
As technology scales, the aging effect caused by Negative Bias Temperature Instability (NBTI) has become a major reliability concern for circuit designers. On the other hand, reducing leakage power remains to be one of the design goals. Because both NBTI-induced circuit degradation and standby leakage power have a strong dependency on the input vectors, Input Vector Control (IVC) technique may be adopted to mitigate leakage and NBTI. However, IVC technique is in-effective for larger circuits. Therefore, in this paper, we propose two fast gate replacement algorithms together with optimal input vector selection to simultaneously mitigate leakage power and NBTI induced circuit degradation: Direct Gate Replacement (DGR) algorithm and Divide and Conquer Based Gate Replacement (DCBGR) algorithm. Our experimental results on 20 benchmark circuits at 65nm technology node reveal that: 1) Both DGR and DCBGR algorithms outperform pure IVC about on average 20% for three different object functions: leakage power reduction only, NBTI mitigation only, and leakage/NBTI co-optimization. 2) The DCBGR algorithm leads to better optimization results and save on average 100X runtime compared with the DGR algorithm.
Yu Wang 0002, Xiaoming Chen 0003, Wenping Wang 0004, Yu Cao 0001, Yuan Xie 0001, Huazhong Yang
DATE4
2009 Modeling of layout-dependent stress effect in CMOS design
abstract
Strain technology has been successfully integrated into CMOS fabrication to improve carrier transport properties since 90 nm node. Due to the non-uniform stress distribution in the channel, the enhancement in carrier mobility, velocity, and threshold voltage shift strongly depend on circuit layout, leading to systematic performance variations among transistors. A compact stress model that physically captures this behavior is essential to bridge the process technology with design optimization. In this paper, starting from the first principle, a new layout-dependent stress model is proposed as a function of layout, temperature, and other device parameters. Furthermore, a method of layout decomposition is developed to partition the layout into a set of simple patterns for efficient model extraction. These solutions significantly reduce the complexity in stress modeling and simulation. They are comprehensively validated by TCAD simulation and published Si-data, including the state-of-the-art strain technologies and the STI stress effect. By embedding them into circuit analysis, the interaction between layout and circuit performance is well benchmarked at 45 nm node.
Chi-Chao Wang, Frank Liu 0001, Min Chen 0024, Yu Cao 0001
ICCAD5
2009 Enabling resonant clock distribution with scaled on-chip magnetic inductors
abstract
Resonant clock distribution with distributed LC oscillators is promising to reducing clock power and jitter noise. Yet the difficulty in the integration of on-chip inductors still limits its application in practice. This paper resolves such a key issue with sub-50 ¿m magnetic inductors, which are fully compatible with the CMOS process. These inductors leverage soft magnetic coils to achieve inductances up to 4nH, Q-factor of 3 at 1 GHz with a device diameter of only 30-50 ¿m, resulting in area savings of nearly 100X as compared to conventional design. The latency and noise performance of the resonant clock network is demonstrated to be comparable to those using conventional inductors without soft magnetic materials. In addition, inductors with integrated magnetic materials significantly reduce mutual coupling and eddy current loss in the power grid below the clock network. These design advantages enable high density of on-chip distributed oscillators, providing better phase averaging, lower power and superior noise characteristics as compared to traditional buffer-tree based clock network.
Saurabh Sinha 0001, Jyothi Velamala, Tawab Dastagir, Bertan Bakkaloglu, Yu Cao 0001
ICCD7
2009 Variation-aware supply voltage assignment for minimizing circuit degradation and leakage
abstract
Abstract—As technology scales, negative bias temperature instability (NBTI) has become a major reliability concern for circuit designers. And the growing process variations can no longer be ignored. Meanwhile, reducing power consumption remains to be one of the design goals. In this paper, a variation-aware supply voltage assignment (SVA) technique combining dual assignment and dynamic scaling is proposed on a statistical platform, to minimize circuit power under an aging-aware timing constraint. The experimental results show that our SVA technique can mitigate on average 62 % of the NBTI-induced circuit delay degrada-tion. Compared with guard-banding and single scaling approaches, our approach saves more energy. Index Terms—Dynamic power, leakage power, negative bias temperature instability (NBTI), supply voltage assignment (SVA). I.
Xiaoming Chen 0003, Yu Wang 0002, Yu Cao 0001, Yuchun Ma, Huazhong Yang
ISLPED3
2009 Finite-Point-Based Transistor Model: A New Approach to Fast Circuit Simulation
abstract
In this paper, a new approach of transistor modeling is developed for fast statistical circuit simulation in the presence of variations. For both the I-V and C-V characteristics of a transistor, finite data points are identified based on their physical meanings and their importance in circuit operation. The impact of process and design variations is embedded into these key points using analytical expressions. During the simulation, the entire I -V and C -V curves are interpolated from these points with simple polynomial formulas. This novel approach significantly enhances the simulation speed with sufficient accuracy. The model is implemented in Verilog-A to support generic circuit simulators. The accuracy and convergence of the proposed model are comprehensively evaluated through a set of benchmark circuits, including nand, a pass-gate, latches, AOI, ring oscillators, and an adder. Compared to SPICE simulations with the BSIM models, the simulation time can be reduced by 7 times in transient analysis and more than 9 times in Monte-Carlo simulations.
Min Chen 0024, Frank Liu 0001, Yu Cao 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Statistical modeling and simulation of threshold variation under dopant fluctuations and line-edge roughness
abstract
The threshold voltage (Vth) of a nanoscale transistor is severely affected by random dopant fluctuations and line-edge roughness. The analysis of these effects usually requires atomistic simulations that are too expensive in computation for statistical circuit design. In this work, we develop an efficient SPICE simulation method and statistical transistor model that accurately predict threshold variation as a function of dopant fluctuations and gate length change caused by sub-wavelength lithography and the gate etching process. By understanding the physical principles of atomistic simulations, we (a) identify the appropriate method to divide a non-uniform gate into slices in order to map those fluctuations into the device model; (b) extract the variation of V th from the strong-inversion region instead of the leakage current, benefiting from the linearity of the saturation current with respect to Vth; and (c) propose a compact model of Vth variation that is scalable with gate size and the amount of dopant and gate length fluctuations. The proposed SPICE simulation method is fully validated against atomistic simulation results. Given the post-lithography gate geometry, this approach correctly models the variation of device output current in all operating regions. Based on the new results, we further project the amount of V th variation at advanced technology nodes, helping shed light on the challenges of future robust circuit design.
Yun Ye 0001, Frank Liu 0001, Sani R. Nassif, Yu Cao 0001
DAC4
2008 Optimized Circuit Failure Prediction for Aging: Practicality and Promise
abstract
Circuit failure prediction is used to predict occurrences of circuit failures, during system operation, before errors appear in system data and states. This technique is applicable for overcoming major scaled-CMOS reliability challenges posed by aging mechanisms such as Negative-Bias-Temperature-Instability (NBTI). This is possible because of the gradual nature of degradation associated with such aging mechanisms. Circuit failure prediction uses special on-chip circuits called aging sensors. In this paper, we experimentally demonstrate correct functionality and practicality of two flavors of flip-flop designs with built-in aging sensors using 90 nm test chips. We also present an aging-aware timing analysis technique to strategically place such flip-flops with built-in aging sensors at selective locations inside a chip for effective circuit failure prediction. This aging-aware timing analysis approach also minimizes the chip-level area impact of such aging sensors. Results from two 90 nm designs demonstrate the practicality and effectiveness of optimized circuit failure prediction with overall chip-level area impact of 2.5% and 0.6%.
Mridul Agarwal, Varsha Balakrishnan, Anshuman Bhuyan, Kyunglok Kim, Bipul Chandra Paul, Wenping Wang 0004, Yu Cao 0001, Subhasish Mitra
ITC8
2008 Digital Circuit Design Challenges and Opportunities in the Era of Nanoscale CMOS
abstract
Well-designed circuits are one key ldquoinsulatingrdquo layer between the increasingly unruly behavior of scaled complementary metal-oxide-semiconductor devices and the systems we seek to construct from them. As we move forward into the nanoscale regime, circuit design is burdened to ldquohiderdquo more of the problems intrinsic to deeply scaled devices. How this is being accomplished is the subject of this paper. We discuss new techniques for logic circuits and interconnect, for memory, and for clock and power distribution. We survey work to build accurate simulation models for nanoscale devices. We discuss the unique problems posed by nanoscale lithography and the role of geometrically regular circuits as one promising solution. Finally, we look at recent computer-aided design efforts in modeling, analysis, and optimization for nanoscale designs with ever increasing amounts of statistical variation.
Benton H. Calhoun, Yu Cao 0001, Xin Li 0001, Ken Mai, Lawrence T. Pileggi, Rob A. Rutenbar, Kenneth L. Shepard
Proc. IEEE2
2007 Modeling and Analysis of Non-Rectangular Gate for Post-Lithography Circuit Simulation
abstract
In the nano regime it has become increasingly important to consider the impact of non-rectangular gate (NRG) shape caused due to sub-wavelength lithography. NRG dramatically increases the leakage current and requires geometry dependent transistor models for post-litho circuit simulation. In this paper, we propose a coherent modeling approach for non-rectangular gates based on equivalent gate length (Le). A gate-voltage dependent model of Le is developed which is scalable with design conditions, continuous across weak and strong inversion regions, accurate for both leakage and saturation current, and compatible with standard circuit analysis tools. We systematically verify this approach with 65nm TCAD simulations. A generic CAD algorithm is further proposed to predict the value of Le under various non-rectangular geometries. The interaction with the narrow-width effect is efficiently convolved in this method. Depending on the gate geometry, the leakage current can vary more than 15X at 65nm technology node. Our analytical method well captures this effect. Finally, we extrapolate the impact of NRG effect on future technology generations. The proposed model can be easily extracted from TCAD tools or direct silicon data. It bridges the gap between lithography, simulation, and circuit analysis for measuring transistor performance under increasingly severe NRG effect.
Ritu Singhal, Asha Balijepalli, Anupama R. Subramaniam, Frank Liu 0001, Sani R. Nassif, Yu Cao 0001
DAC6
2007 The Impact of NBTI on the Performance of Combinational and Sequential Circuits
abstract
Negative-bias-temperature-instability (NBTI) has become the primary limiting factor of circuit lifetime. In this work, we develop a general framework for analyzing the impact of NBTI on the performance of a circuit, based on various circuit parameters such as the supply voltage, temperature, and node switching activity of the signals etc. We propose an efficient method to predict the degradation of circuit performance based on circuit topology and the switching activity of the signals over long periods of time. We demonstrate our results on ISCAS benchmarks and a 65nm industrial design. The framework is used to provide key design insights for designing reliable circuits. The key design insights that we obtain are: (1) degradation due to NBTI is most sensitive on the input patterns and the duty cycle; the difference in the delay degradation can be up to 5X for various static and dynamic conditions, (2) during dynamic operation, NBTI-induced degradation is relatively insensitive to supply voltage, but strongly dependent on temperature; (3) NBTI has marginal impact on the clock signal.
Wenping Wang 0004, Shengqi Yang, Sarvesh Bhardwaj, Rakesh Vattikonda, Sarma B. K. Vrudhula, Frank Liu 0001, Yu Cao 0001
DAC7
2007 Fast statistical circuit analysis with finite-point based transistor model
Min Chen 0024, Frank Liu 0001, Yu Cao 0001
DATE4
2007 Optimizing finfet technology for high-speed and low-power design
abstract
The threshold voltage (Vth) of a FinFET device varies between the on and off mode: Vth is lower when the transistor is on and it is higher when the transistor is off. Such a property is ideal for low-power designs with low supply voltage (Vdd). The low Vth provides high circuit speed even when Vdd is low, while the high standby Vth effectively controls the leakage. In this work, we exploit this property to achieve both high-speed and low-power operations. First, we develop an equivalent sub-circuit model of a FinFET transistor for design explorations. The accuracy of this model is verified with TCAD simulations. Then, we optimize key device parameters to obtain a large range of Vth between dynamic and standby mode: a thicker gate oxide and a thinner silicon body are desirable for this low-power design. Using the optimized FinFET device at 32nm node, we demonstrate that more than 35% reduction in total energy can be achieved without sacrificing the speed. We further benchmark the performance of representative logic and memory units under process variations.
Tarun Sairam, Yu Cao 0001
ACM Great Lakes Symposium on VLSI3
2007 A robust finite-point based gate model considering process variations
abstract
This paper proposes a robust gate model based on a finite-point modeling scheme. With current source model (CSM) framework, a robust, finite-point gate model is constructed. The new model depends on the selective points of I-V curves of gates. Thus, it implicitly incorporates the variation related parameters into finite points. In addition, to provide good accuracy on output waveform, the new model creates the input and output capacitance elements as nonlinear dependency on input/output waveform and process variation parameters. Experimental results show that the generated gate model has less than 3.7% error at mean, less than 6.2% error at variance and less than 5.8% at 90% percentile for cumulative density functions (CDFs).
Alexander V. Mitev, Dinesh Ganesan, Dheepan Shanmugasundaram, Yu Cao 0001, Janet Roveda
ICCAD4
2007 An efficient method to identify critical gates under circuit aging
abstract
Negative bias temperature instability (NBTI) is the leading factor of circuit performance degradation. Due to its complex dependence on operating conditions, especially signal probability, it is a tremendous challenge to accurately predict the degradation rate in reality. On the other hand, we demonstrate in this work that it is feasible to reliably predict the relative importance of gates under NBTI. By identifying critical gates that are the most important ones for timing degradation, we will be able to effectively protect the circuit from aging, with the minimum design overhead. The proposed method is based on a new timing analysis framework that integrates a NBTI-aware library. For each potential critical path, we prove that there exists a particular signal probability, which leads to the worst case of timing degradation. The search of such worst case signal probability provides a safe guardband for the degradation, yet avoiding overly pessimistic analysis. By applying this method to ISCAS and ITC benchmark circuits at the 65 nm node, we demonstrate that in average only 1% of total gates need to be protected in order to control the timing degradation within 10% in ten years. Since this method only requires one-time analysis of each critical path, it is very efficient in computation. With the information of critical gates available, it further enables other resilient design techniques to mitigate circuit aging under NBTI.
Wenping Wang 0004, Zile Wei, Shengqi Yang, Yu Cao 0001
ICCAD4
2007 Compact modeling of carbon nanotube transistor for early stage process-design exploration
abstract
Carbon nanotube transistor (CNT) is promising to be the technology of choice for nanoscale integration. In this work, we develop the first compact model of CNT, with the objective to explore the optimal process and design space for robust low-power applications. Based on the concept of the surface potential, the new model accurately predicts the characteristics of a CNT device under various process and design conditions, such as diameter, chirality, gate dielectrics, and bias voltages. With the physical modeling of the contact, this model covers both the Schottky-barrier CNT (SB-CNT) and MOS-type CNT. The proposed model does not require any iteration and thus, significantly enhances the simulation efficiency to support large-scale design research. Using this model, we benchmark the performance of a FO4 inverter with CNT and 22nm CMOS technology. The following key insights are extracted: (1) even with the SB-CNT and realistic layout parasitics, the circuit speed can be more than 10X that of 22nm CMOS; (2) The diameter range of 1-1.5nm exhibits the maximum tolerance to contact materials and process variations; (3) a CNT circuit allows better scaling of the supply voltage (Vdd) for power reduction. For a fixed energy consumption and Vdd, the CNT speed is 4X that of 22nm CMOS. Overall, the new model enables efficient design research with CNT, revealing tremendous opportunities for both high-speed and low-power applications.
Asha Balijepalli, Saurabh Sinha 0001, Yu Cao 0001
ISLPED3
2007 Predictive technology model for nano-CMOS design exploration
abstract
A predictive MOSFET model is critical for early circuit design research. In this work, a new generation of Predictive Technology Model (PTM) is developed, covering emerging physical effects and alternative structures, such as the double-gate device (i.e., FinFET). Based on physical models and early stage silicon data, PTM of bulk and double-gate devices are successfully generated from 130nm to 32nm technology nodes, with effective channel length down to 13nm. By tuning only ten primary parameters, PTM can be easily customized to cover a wide range of process uncertainties. The accuracy of PTM predictions is comprehensively verified with published silicon data: the error of the current is below 10% for both NMOS and PMOS. Furthermore, the new PTM correctly captures process sensitivities in the nanometer regime. PTM is available online at http://www.eas.asu.edu/~ptm.
Yu Cao 0001
ACM J. Emerg. Technol. Comput. Syst.2
2006 Statistical leakage minimization through joint selection of gate sizes, gate lengths and threshold voltage
abstract
This paper proposes a novel methodology for statistical leakage minimization of digital circuits. A function of mean and variance of the circuit leakage is minimized with constraint on a-percentile of the delay using physical delay models. Since the leakage is a strong function of the threshold voltage and gate length, considering them as design variables can provide significant amount of power savings. The leakage minimization problem is formulated as a multivariable convex optimization problem. We demonstrate that statistical optimization can lead to more than 37% savings in nominal leakage compared to worst-case techniques that perform only gate sizing.
Sarvesh Bhardwaj, Yu Cao 0001, Sarma B. K. Vrudhula
ASP-DAC2
2006 Modeling of intra-die process variations for accurate analysis and optimization of nano-scale circuits
abstract
This paper proposes the use of Karhunen-Loève Expansion (KLE) for accurate and efficient modeling of intra-die correlations in the semiconductor manufacturing process. We demonstrate that the KLE provides a significantly more accurate representation of the underlying stochastic process compared to the traditional approach of dividing the layout into grids and applying Principal Component Analysis (PCA). By comparing the results of leakage analysis using both KLE and the existing approaches, we show that using KLE can provide up to 4-5x reduction in the variability space (number of random variables) while maintaining the same accuracy. We also propose an efficient leakage minimization algorithm that maximizes the leakage yield while satisfying probabilistic constraints on the delay.
Sarvesh Bhardwaj, Sarma B. K. Vrudhula, Praveen Ghanta, Yu Cao 0001
DAC4
2006 Modeling and minimization of PMOS NBTI effect for robust nanometer design
abstract
Negative bias temperature instability (NBTI) has become the dominant reliability concern for nanoscale PMOS transistors. In this paper, a predictive model is developed for the degradation of NBTI in both static and dynamic operations. Model scalability and generality are comprehensively verified with experimental data over a wide range of process and bias conditions. By implementing the new model into SPICE for an industrial 90nm technology, key insights are obtained for the development of robust design solutions: (1) the most effective techniques to mitigate the NBTI degradation are VDD tuning, PMOS sizing, and reducing the duty cycle; (2) an optimal VDD exists to minimize the degradation of circuit performance; (3) tuning gate length or the switching frequency has little impact on the NBTI effect; (4) a new switching scenario is identified for worst case timing analysis during NBTI stress.
Rakesh Vattikonda, Wenping Wang 0004, Yu Cao 0001
DAC3
2006 Analysis and modeling of CD variation for statistical static timing
abstract
Statistical static timing analysis (SSTA) has become a key method for analyzing the effect of process variation in aggressively scaled CMOS technologies. Much research has focused on the modeling of spatial correlation in SSTA. However, the vast majority of these works used artificially generated process data to test the proposed models. Hence, it is difficult to determine the actual effectiveness of these methods, the conditions under which they are necessary, and whether they lead to a significant increase in accuracy that warrants their increased runtime and complexity. In this paper, we study 5 different correlation models and their associated SSTA methods using 35420 critical dimension (CD) measurements that were extracted from 23 reticles on 5 wafers in a 130nm CMOS process. Based on the measured CD data, we analyze the correlation as a function of distance and generate 5 distinct correlation models, ranging from simple models which incorporate one or two variation components to more complex models that utilize principle component analysis and Quad-trees. We then study the accuracy of the different models and compare their SSTA results with the result of running STA directly on the extracted data. We also examine the trade-off between model accuracy and run time, as well as the impact of die size on model accuracy. We show that, especially for small dies (< 6.6mm x 5.7mm), the simple models provide comparable accuracy to that of the more complex ones, while incurring significantly less runtime and implementation difficulty. The results of this study demonstrate that correlation models for SSTA must be carefully tested on actual process data and must be used judiciously.
Brian Cline, Kaviraj Chopra, David T. Blaauw, Yu Cao 0001
ICCAD4
2005 Impact of on-chip interconnect frequency-dependent R(f)L(f) on digital and RF design
abstract
On-chip global interconnect exhibits clear frequency dependence in both resistance (R) and inductance ( L). In this paper, its impact on modern digital and radio frequency (RF) circuit design is examined. First, a physical and compact ladder circuit model is developed to capture this behavior, which only employs frequency independent R and L elements, and thus, supports transient analysis. Using this new model we demonstrate that the use of dc values for R and L is sufficient for timing analysis (i.e., 50% delay and slew rate) in digital designs. However, RL frequency dependence is critical for the analysis of signal integrity, shield line insertion, power supply stability, and RF inductor performance.
Yu Cao 0001, Xuejue Huang, Dennis Sylvester, Tsu-Jae King Liu, Chenming Hu
IEEE Trans. Very Large Scale Integr. Syst.1
2005 Switch-factor based loop RLC modeling for efficient timing analysis
abstract
Timing uncertainty caused by inductive and capacitive coupling is one of the major bottlenecks in timing analysis. In this paper, we propose an effective loop RLC modeling technique to efficiently decouple lines with both inductive and capacitive coupling. We generalize the RLC decoupling problem based on the theory of distributed RLC lines and a switch-factor, which is the voltage ratio between two nets. This switch-factor is also known as the Miller factor, and is widely used to model capacitive coupling. The proposed modeling technique can be directly applied to partial RLC netlists extracted using existing parasitic extraction tools without advance knowledge of the return path. The new model captures the impact of neighboring switching activity as it significantly affects the current return path. As demonstrated in our experiments, the new model accurately predicts both upper and lower delay bounds as a function of neighboring switching patterns. Therefore, this approach can be easily implemented into existing timing analysis flows such as max-timing and min-timing analysis. Finally, we apply the new modeling approach to a range of activities across the design process including timing optimization, static timing analysis, high frequency clock design, and data-bus wire planning.
Yu Cao 0001, Xuejue Huang, Dennis Sylvester
IEEE Trans. Very Large Scale Integr. Syst.1
2003 Switch-Factor Based Loop RLC Modeling for Efficient Timing Analysis
Yu Cao 0001, Xuejue Huang, Dennis Sylvester
ICCAD1
2003 Bidirectional closed-form transformation between on-chip coupling noise waveforms and interconnect delay-change curves
abstract
A novel concept of bidirectional transformation between on-chip coupling noise waveform and delay-change curve (DCC) using closed-form equations is described in this paper. These equations are targeted for use in: 1) the efficient generation of DCCs and 2) accurate experimental determination of subnanosecond coupling noise. In particular, we explore the concept of using analytical models to efficiently generate DCCs that can then be used to characterize the impact of noise on any victim/aggressor configuration. The concept is model independent, although we investigate several common noise modeling choices and perform a sensitivity analysis to optimize the generation of DCCs. By extending existing noise models, arbitrary configurations can be considered including multiple aggressors in the timing-analysis framework. Simulation using the analytical approach closely matches time-consuming SPICE simulations, making noise-aware timing analysis using DCCs both efficient and accurate. A test chip using a 0.25-/spl mu/m CMOS process was designed and its measurement results also show good agreement with SPICE simulations.
Takashi Sato 0001, Yu Cao 0001, Kanak Agarwal 0001, Dennis Sylvester, Chenming Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Improved a priori interconnect predictions and technology extrapolation in the GTX system
abstract
A priori interconnect prediction and technology extrapolation are closely intertwined. Interconnect predictions are at the core of technology extrapolation models of achievable system power, area density, and speed. Technology extrapolation, in turn, informs a priori interconnect prediction via models of interconnect technology and interconnect optimizations. In this paper, we address the linkage between a priori interconnect prediction and technology extrapolation in two ways. First, we describe how rapid changes in technology, as well as rapid evolution of prediction methods, require a dynamic and flexible framework for technology extrapolation. We then develop a new tool, the GSRC technology extrapolation system (GTX), which allows capture of such knowledge and rapid development of new studies. Second, we identify several "nontraditional" facets of interconnect prediction and quantify their impact on key technology extrapolations. In particular, we explore the effects of interconnect design optimizations such as shield insertion, repeater sizing and repeater staggering, as well as modeling choices for RLC interconnects.
Yu Cao 0001, Chenming Hu, Xuejue Huang, Andrew B. Kahng, Igor L. Markov, Michael Oliver, Dirk Stroobandt, Dennis Sylvester
IEEE Trans. Very Large Scale Integr. Syst.1
2002 Effective on-chip inductance modeling for multiple signal lines and application to repeater insertion
abstract
/sup A/ new approach to handle inductance effects for multiple signal lines is presented. The worst-case switching pattern is first identified. Then a numerical approach is used to model the effective loop inductance (L/sub eff/) for multiple lines. Based on a look-up table for L/sub eff/, an equivalent single line model can be generated to decouple a specific signal line from the others to perform static timing analysis. Compared to the use of full RLC netlists for multiple lines, this approach greatly improves the computational efficiency and maintains accuracy for timing and signal integrity analysis. We apply these models to repeater insertion in critical paths and find that, for a single line, the RLC model minimizes delay with fewer number of repeaters than RC model. However, for multiple lines, we find that same number of repeaters is inserted for optimal delay according to both the RC and RLC models.
Yu Cao 0001, Xuejue Huang, N. H. Chang, O. Sam Nakagawa, Weize Xie, Dennis Sylvester, Chenming Hu
IEEE Trans. Very Large Scale Integr. Syst.1
2000 GTX: the MARCO GSRC technology extrapolation system
abstract
Technology extrapolation — the calibration and prediction of achievable design in future technology generations — drives the evolution of VLSI system architectures, design methodologies, and design tools. This paper describes initial experiences with development and use of GTX, the MARCO GSRC Technology Extrapolation system. GTX provides a robust, portable framework for interactive specification and comparison of modeling choices, e.g., for predicting system cycle time, die size and power dissipation. We use GTX to reveal surprising levels of uncertainty (modeling and parameter sensitivity) in widely-cited cycle-time models that drive recent roadmaps. We also describe new SOI and bulk device models that have been developed for GTX, as well as studies of power dissipation and delay uncertainty under various implementation assumptions for global interconnects.
Andrew E. Caldwell, Yu Cao 0001, Andrew B. Kahng, Farinaz Koushanfar, Hua Lu 0004, Igor L. Markov, Michael Oliver, Dirk Stroobandt, Dennis Sylvester
DAC2
2000 Effects of Global Interconnect Optimizations on Performance Estimation of Deep Submicron Design
abstract
In this paper, we quantify the impact of global interconnect optimization techniques that address such design objectives as delay, peak noise, delay uncertainty due to noise, power, and cost. In doing so, we develop a new system-performance simulation model as a set of studies within the MARCO GSRC Technology Extrapolation (GTX) system. We model a typical point-to-point global interconnect and focus on accurate assessment of both circuit and design technology with respect to such issues as inductance, signal line shielding, dynamic delay, buffer placement uncertainty and repeater staggering. We demonstrate, for example, that optimal wire sizing models need to consider inductive effects-and that use of more accurate (-1,3) worst-case capacitive coupling noise switch factors substantially increases peak noise estimates compared to traditional (0,2) bounds. We also find that optimal repeater sizes are significantly smaller than conventional models would suggest, especially when considering energy-delay issues.
Yu Cao 0001, Chenming Hu, Xuejue Huang, Andrew B. Kahng, Sudhakar Muddu, Dirk Stroobandt, Dennis Sylvester
ICCAD1