Peng Li 0001

dblp:83/6353-1 · DBLP profile ↗
← Back
186ranked-venue papers
18as first author
39since 2021 · last 2026
0000-0003-3548-4589ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 152 · 18 first-author · 24 since 2021Artificial intelligence and machine learning · 32 · 14 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021
YearPublicationVenuePosition
2026 Khan-GCL: Kolmogorov-Arnold Network Based Graph Contrastive Learning with Hard Negatives
abstract
Graph contrastive learning (GCL) has demonstrated great promise for learning generalizable graph representations from unlabeled data. However, conventional GCL approaches face two critical limitations: (1) the restricted expressive capacity of multilayer perceptron (MLP) based encoders, and (2) suboptimal negative samples that either from random augmentations—failing to provide effective 'hard negatives'—or generated hard negatives without addressing the semantic distinctions crucial for discriminating graph data. To this end, we propose Khan-GCL, a novel framework that integrates the Kolmogorov–Arnold Network (KAN) into the GCL encoder architecture, substantially enhancing its representational capacity. Furthermore, we exploit the rich information embedded within KAN coefficient parameters to develop two novel critical feature identification techniques that enable the generation of semantically meaningful hard negative samples for each graph representation. These strategically constructed hard negatives guide the encoder to learn more discriminative features by emphasizing critical semantic differences between graphs. Extensive experiments demonstrate that our approach achieves state-of-the-art performance compared to existing GCL methods across a variety of datasets and tasks.
Zihu Wang, Boxun Xu, Hejia Geng, Peng Li 0001
AAAI4
2026 AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
abstract
Visual autoregressive modeling (VAR) via next-scale prediction has emerged as a scalable image generation paradigm. While Key and Value (KV) caching in large language models (LLMs) has been extensively studied, next-scale prediction presents unique challenges, and KV caching design for next-scale based VAR transformers remains largely unexplored. A major bottleneck is the excessive KV memory growth with the increasing number of scales—severely limiting scalability. Our systematic investigation reveals that: (1) Attending to tokens from local scales significantly contributes to generation quality (2) Allocating a small amount of memory for the coarsest scales, termed as condensed scales, stabilizes multi-scale image generation (3) Strong KV similarity across finer scales is predominantly observed in cache-efficient layers, whereas cache-demanding layers exhibit weaker inter-scale similarity. Based on the observations, we introduce AMS-KV, a scale-adaptive KV caching policy for next-scale prediction in VAR models. AMS-KV prioritizes storing KVs from condensed and local scales, preserving the most relevant tokens to maintain generation quality. It further optimizes KV cache utilization and computational efficiency identifying cache-demanding layers through inter-scale similarity analysis. Compared to the vanilla next-scale prediction-based VAR models, AMS-KV reduces KV cache usage by up to 84.83% and self-attention latency by 60.48%. Moreover, when the baseline VAR-d30 model encounters out-of-memory failures at a batch size of 128, AMS-KV enables stable scaling to a batch size of 256 with improved throughput.
Boxun Xu, Yu Wang 0167, Zihu Wang, Peng Li 0001
AAAI4
2026 LLM-USO: Large Language Model-Based Universal Sizing Optimizer
abstract
The design of analog circuits is a cornerstone of integrated circuit (IC) development, requiring the optimization of complex, interconnected sub-structures such as amplifiers, comparators, and buffers. Traditionally, this process relies heavily on expert human knowledge to refine design objectives by carefully tuning sub-components while accounting for their interdependencies. Existing methods, such as Bayesian Optimization (BO), offer a mathematically driven approach for efficiently navigating large design spaces. However, these methods fall short in two critical areas compared to human expertise: (i) they lack the semantic understanding of the sizing solution space and its direct correlation with design objectives before optimization, and (ii) they fail to reuse knowledge gained from optimizing similar sub-structures across different circuits. To overcome these limitations, we propose the Large Language Model-based Universal Sizing Optimizer (LLM-USO), which introduces a novel method for knowledge representation to encode circuit design knowledge in a structured text format. This representation enables the systematic reuse of optimization insights for circuits with similar sub-structures. LLM-USO employs a hybrid framework that integrates BO with large language models (LLMs) and a learning summary module. This approach serves to: (i) infuse domain-specific knowledge into the BO process and (ii) facilitate knowledge transfer across circuits, mirroring the cognitive strategies of expert designers. Specifically, LLM-USO constructs a knowledge summary mechanism to distill and apply design insights from one circuit to related ones. It also incorporates a knowledge summary critiquing mechanism to ensure the accuracy and quality of the summaries and employs BO-guided suggestion filtering to identify optimal design points efficiently. We evaluate the LLM-USO framework through transfer learning experiments on various analog circuits, demonstrating its ability to improve the quality of circuit design optimization.
Karthik Somayaji Nanjangud Suryanarayana, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are promising biologically plausible models of computation which utilize a spiking binary activation function similar to that of biological neurons. SNNs are well positioned to process spatiotemporal data, and are advantageous in ultra-low power and real-time processing. Despite a large body of work on conventional artificial neural network accelerators, much less attention has been given to efficient SNN hardware accelerator design. In particular, SNNs exhibit inherent unstructured spatial and temporal firing sparsity, an opportunity yet to be fully explored for great hardware processing efficiency. In this work, we propose a novel systolic-array SNN accelerator architecture, called SpikeX, to take on the challenges and opportunities stemming from unstructured sparsity while taking into account the unique characteristics of spike-based computation. By developing an efficient dataflow targeting expensive multi-bit weight data movements, SpikeX reduces memory access and increases data sharing and hardware utilization for computations spanning across both time and space, thereby significantly improving energy efficiency and inference latency. Furthermore, recognizing the importance of SNN network and hardware co-design, we develop a co-optimization methodology facilitating not only hardware-aware SNN training but also hardware accelerator architecture search, allowing joint network weight parameter optimization and accelerator architectural reconfiguration. This end-to-end network/accelerator co-design approach offers a significant reduction of 15.1×−150.87× in energy-delay-product(EDP) without comprising model accuracy.
Boxun Xu, Richard Boone, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Trimming Down Large Spiking Vision Transformers Via Heterogeneous Quantization Search
abstract
Spiking neural networks (SNNs) are amenable to deployment on edge devices and neuromorphic hardware due to their lower dissipation. Recently, SNN-based transformers have garnered significant interest, incorporating attention mechanisms akin to their counterparts in artificial neural networks (ANNs) while demonstrating excellent performance. However, deploying large-scale spiking transformer models on resourceconstrained edge devices such as mobile devices still poses significant challenges resulting from the high computational demands of large uncompressed high-precision models. In this work, we introduce a novel heterogeneous quantization method to compress spiking transformers through layer-wise quantization. Our approach optimizes the quantization of each layer using one of two distinct quantization schemes, that is, uniform or power-of-two quantization, with mixed precision. Our heterogeneous quantization demonstrates the feasibility of maintaining high performance for spiking transformers while utilizing an average effective precision of 3.14-3.67 bits with less than a$\mathbf{1 \%}$accuracy drop on neuromorphic DVS Gesture and CIFAR10-DVS datasets. It achieves$8.71 \times-10.19 \times$model compression rate for standard floating-point spiking transformers. Furthermore, the proposed approach achieves substantial energy reductions of$5.69 \times, 8.72 \times$, and$10.2 \times$on the N-Caltech101, DVS-Gesture, and CIFAR10DVS datasets, respectively, while maintaining high accuracy levels of$85.3 \%, 97.57 \%$, and 80.4 %.
Boxun Xu, Peng Li 0001
ASAP3
2025 3D Acceleration for Mixture-of-Experts and Multi-Head Attention Spiking Transformers with Dynamic Head Pruning
abstract
Spiking Neural Networks (SNNs) provide a brain-inspired and event-driven mechanism that is believed to be critical to unlock energy-efficient deep learning. On the other hand, mixture-of-experts (MoE) models mirror the parallel distributed processing of the nervous system, and expand model capacity without scaling up the number of computational operations. However, there is currently a lack of hardware support for highly parallel distributed processing in spiking based MoE models. This paper introduces the first 3D hardware architecture and design methodology for Mixture-of-Experts and Multi-Head Attention spiking transformers. By leveraging 3D integration with memory-on-logic and logic-on-logic stacking and exploring energy-efficient dynamic head pruning, we explore such brain-inspired accelerators with spatially stackable circuitry, demonstrating significant improvements of energy efficiency and latency over conventional 2D CMOS integration.
Boxun Xu, Junyoung Hwang, Pruek Vanna-Iampikul, Yuxuan Yin, Sung Kyu Lim, Peng Li 0001
ICCAD6
2025 Bishop: Sparsified Bundling Spiking Transformers on Heterogeneous Cores with Error-constrained Pruning
abstract
Spiking neural networks(SNNs) have emerged as a promising solution for deployment on resource-constrained edge devices and neuromorphic hardware due to their low power consumption.Spiking transformers, which integrate attention mechanisms similar to those found in artificial neural networks (ANNs), have recently exhibited impressive performance.However, these models are large in size and involve high-volume computation in both time and space, posing significant challenges for efficient hardware acceleration.We present Bishop, the first dedicated hardware accelerator architecture and HW/SW co-design framework for spiking transformers that optimally represents, manages, and processes spike-based workloads while exploring spatiotemporal sparsity and data reuse.Specifically, we introduce the concept of Token-Time Bundle (TTB), a container that bundles spiking data of a set of tokens over multiple time points.Our heterogeneous accelerator architecture Bishop concurrently processes workload packed in TTBs and explores intra-and inter-bundle multiple-bit weight reuse to significantly reduce memory access.Bishop utilizes a stratifier, a dense core array, and a sparse core array to process MLP blocks and projection layers.The stratifier routes high-density spiking activation workload to the dense core and low-density counterpart to the sparse core, ensuring optimized processing tailored to the given spatiotemporal sparsity level.To further reduce data access and computation, we introduce a novel Bundle Sparsity-Aware (BSA) training pipeline that enhances not only the overall but also structured TTB-level firing sparsity.Moreover, the processing efficiency of self-attention layers is boosted by the proposed Error-Constrained TTB Pruning (ECP), which trims activities in spiking queries, keys, and values both before and after the computation of spiking attention maps with a well-defined error bound.Finally, we design a reconfigurable TTB spiking attention core to efficiently compute spiking attention maps by executing highly simplified "AND" and "Accumulate" operations.On average, Bishop achieves a 5.91× speedup and 6.11×
Boxun Xu, Yuxuan Yin, Vikram Iyer, Peng Li 0001
ISCA4
2025 Transfer Learning for Minimum Operating Voltage Prediction in Advanced Technology Nodes: Leveraging Legacy Data and Silicon Odometer Sensing
abstract
Accurate prediction of chip performance is critical for ensuring energy efficiency and reliability in semiconductor manufacturing. However, developing minimum operating voltage (Vmin) prediction models at advanced technology nodes is challenging due to limited training data and the complex relationship between process variations and Vmin. To address these issues, we propose a novel transfer learning framework that leverages abundant legacy data from the 16nm technology node to enable accurate Vminprediction at the advanced 5nm node. A key innovation of our approach is the integration of input features derived from on-chip silicon odometer sensor data, which provide fine-grained characterization of localized process variations—an essential factor at the 5nm node—resulting in significantly improved prediction accuracy.
Yuxuan Yin, Rebecca Chen, Boxun Xu, Peng Li 0001
ITC5
2025 Data-Efficient Prediction of Minimum Operating Voltage via Inter- and Intra-Wafer Variation Alignment
abstract
Predicting the minimum operating voltage (Vmin) of chips stands as a crucial technique in enhancing the speed and reliability of manufacturing testing flow. However, existing Vminprediction methods often overlook various sources of variations in both training and deployment phases. Notably, overlooking wafer zone-to-zone (intra-wafer) variations and wafer-to-wafer (inter-wafer) variations diminishes the accuracy, data efficiency, and reliability of Vminpredictors. To address this challenge, we propose Restricted Bias Alignment (RBA), a novel data-efficient Vminprediction framework that introduces a variation alignment technique to simultaneously estimate inter- and intra-wafer variations. Furthermore, we propose utilizing class probe data to model inter-wafer variations for the first time.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
VTS4
2025 Reliable Board-Level Degradation Prediction with Monotonic Segmented Regression under Noisy Measurement
abstract
The increasing complexity of electronic systems in autonomous electric vehicles necessitates robust methods for forecasting the degradation of critical components such as printed circuit boards (PCBs). Various time series forecasting methods have been investigated to predict in-situ resistance degradation under vibration loads. However, these methods failed to capture the degradation trend under strong measurement noise. This paper introduces Monotonic Segmented Linear Regression (MSLR), a novel approach designed to capture monotonic degradation trends in time series data under significant measurement noise. By incorporating monotonic constraints, MSLR effectively models the non-decreasing behavior characteristic of degradation processes. To further enhance reliability of the prediction, we integrate Adaptive Conformal Inference (ACI) with MSLR, enabling the estimation of statistically valid upper bounds for resistance degradation with high confidence. Extensive experiments demonstrate that MSLR outperforms state-of-the-art time series forecasting baselines on real-world PCB degradation datasets.
Yuxuan Yin, Rebecca Chen, Varun Thukral, Peng Li 0001
VTS5
2025 Backpropagation-based learning with local derivative approximation and memory replay in biologically plausible neural systems
Richard Boone, Peng Li 0001
Neurocomputing2
2025 Rare Event Detection by Acquisition-Guided Sampling
abstract
Motivated by the challenges in detecting extremely rare failures for sophisticated specifications in circuit design, we consider the problem of detecting regions of interest (ROIs) that consist of specifications with the value of a complex target function for the system performance being below or above a certain pre-specified threshold. Though Bayesian optimization (BO) has been applied to this problem, it is not effective in identifying multiple ROIs as it was originally designed for global optimization and tends to focus on searching the area where the global optimum is most likely to be. In this work, we propose a sampling strategy for fast ROI detection within a limited number of target function evaluations. The sampling distribution is designed so that the probability of a specification being sampled is proportional to the corresponding value of the acquisition function. Such an acquisition-guided sampling algorithm promotes a wider search of the sample space and a simpler incorporation of different criteria to determine the specifications to be evaluated next. To further improve the performance, we propose a new design of the acquisition function and two modifications of existing acquisition functions. Numerical studies on synthetic functions and a real-world circuit design application demonstrate that the proposed method can enjoy a stronger exploration ability provided by sampling and achieve faster ROI detection with higher coverage. Note to Practitioners—This study considers the extremely rare failure detection in automated circuit design, fabrication, packaging, and verification. Obtaining enough observations of interest within a given budget of evaluations is challenging due to the scarcity of extremely rare failures. Bayesian optimization (BO) has been adopted to tackle this problem, but it may not achieve satisfactory coverage of multiple failure regions, since its goal is to find the global optimum of the target function. In this paper, we propose a sampling-based rare event detection strategy tailored to efficiently detect regions of interest (ROIs) with high coverage, along with newly designed acquisition functions incorporating the pre-specified threshold and ideas of experimental design. This sampling strategy has greater robustness to the choice of acquisition function. Also, since multiple queries can be easily obtained through sampling, various criteria can be easily incorporated for determining the next batch of evaluation specifications without much increase in computational complexity.
Huiling Liao, Xiaoning Qian, Jianhua Z. Huang, Peng Li 0001
IEEE Trans Autom. Sci. Eng.4
2024 Semi-supervised Learning of Dynamical Systems with Neural Ordinary Differential Equations: A Teacher-Student Model Approach
abstract
Modeling dynamical systems is crucial for a wide range of tasks, but it remains challenging due to complex nonlinear dynamics, limited observations, or lack of prior knowledge. Recently, data-driven approaches such as Neural Ordinary Differential Equations (NODE) have shown promising results by leveraging the expressive power of neural networks to model unknown dynamics. However, these approaches often suffer from limited labeled training data, leading to poor generalization and suboptimal predictions. On the other hand, semi-supervised algorithms can utilize abundant unlabeled data and have demonstrated good performance in classification and regression tasks. We propose TS-NODE, the first semi-supervised approach to modeling dynamical systems with NODE. TS-NODE explores cheaply generated synthetic pseudo rollouts to broaden exploration in the state space and to tackle the challenges brought by lack of ground-truth system data under a teacher-student model. TS-NODE employs an unified optimization framework that corrects the teacher model based on the student's feedback while mitigating the potential false system dynamics present in pseudo rollouts. TS-NODE demonstrates significant performance improvements over a baseline Neural ODE model on multiple dynamical system modeling tasks.
Yu Wang 0167, Yuxuan Yin, Karthik Somayaji Nanjangud Suryanarayana, Ján Drgona, Malachi Schram, Mahantesh Halappanavar, Frank Liu 0001, Peng Li 0001
AAAI8
2024 Learn-by-Compare: Analog Performance Prediction using Contrastive Regression with Design Knowledge
abstract
This paper introduces Learn-by-Compare (LbC), a novel approach for analog performance modeling by employing semi-supervised contrastive regression. LbC employs a deep neural network encoder to come up with latent representations of sizing solutions by comparing similarity/dissimilarity of the underlying performance. Leveraging two levels of transistor level sizing data augmentation (DA), namely LS-DA and GS-DA, LbC produces new data samples by employing design knowledge. Experimental results highlight LbC's superior predictive accuracy compared to traditional regression methods. Offering a streamlined semi-supervised learning methodology, LbC effectively incorporates simple design knowledge and representation learning for efficient analog performance modeling.
Zihu Wang, Karthik Somayaji Nanjangud Suryanarayana, Peng Li 0001
DAC3
2024 Data-Efficient Conformalized Interval Prediction of Minimum Operating Voltage Capturing Process Variations
abstract
Accurate minimum operating voltage (Vmin) prediction is a critical element in manufacturing tests. Conventional methods lack coverage guarantees in interval predictions. Conformal Prediction (CP), a distribution-free machine learning approach, excels in providing rigorous coverage guarantees for interval predictions. However, standard CP predictors may fail due to a lack of knowledge of process variations. We address this challenge by providing principled conformalized interval prediction in the presence of process variations with high data efficiency, where the data from a few additional chips is utilized for calibration. We demonstrate the superiority of the proposed method on industrial 16nm chip data.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
DAC4
2024 Reliable Interval Prediction of Minimum Operating Voltage Based on On-Chip Monitors via Conformalized Quantile Regression
abstract
Predicting the minimum operating voltage$V_{min}$of chips is one of the important techniques for improving the manufacturing testing flow, as well as ensuring the long-term reliability and safety of in-field systems. Current$V_{min}$prediction methods often provide only point estimates, necessitating additional techniques for constructing prediction confidence intervals to cover uncertainties caused by different sources of variations. While some existing techniques offer region predictions, but they rely on certain distributional assumptions and/or provide no coverage guarantees. In response to these limitations, we propose a novel distribution-free$V_{min}$interval estimation methodology possessing a theoretical guarantee of coverage. Our approach leverages conformalized quantile regression and on-chip monitors to generate reliable prediction intervals. We demonstrate the effectiveness of the proposed method on an industrial 5nm automotive chip dataset. Moreover, we show that the use of on-chip monitors can reduce the interval length significantly for$V_{min}$prediction.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
DATE5
2024 On Learning Discriminative Features from Synthesized Data for Self-supervised Fine-Grained Visual Recognition
Zihu Wang, Lingqiao Liu, Scott Ricardo Figueroa Weston, Samuel Tian, Peng Li 0001
ECCV (89)5
2024 Spiking Transformer Hardware Accelerators in 3D Integration
abstract
Spiking neural networks (SNNs) are powerful models of spatiotemporal computation and are well suited for deployment on resource-constrained edge devices and neuromorphic hardware due to their low power consumption. Leveraging attention mechanisms similar to those found in their artificial neural network counterparts, recently emerged spiking transformers have showcased promising performance and efficiency by capitalizing on the binary nature of spiking operations. Recognizing the current lack of dedicated hardware support for spiking transformers, this paper presents the first work on 3D spiking transformer hardware architecture and design methodology. We present an architecture and physical design co-optimization approach tailored specifically for spiking transformers. Through memory-on-logic and logic-on-logic stacking enabled by 3D integration, we demonstrate significant energy and delay improvements compared to conventional 2D CMOS integration.
Boxun Xu, Junyoung Hwang, Pruek Vanna-Iampikul, Sung Kyu Lim, Peng Li 0001
ICCAD5
2024 ADO-LLM: Analog Design Bayesian Optimization with In-Context Learning of Large Language Models
abstract
Analog circuit design requires substantial human expertise and involvement, which is a significant roadblock to design productivity. Bayesian Optimization (BO), a popular machine-learning-based optimization strategy, has been leveraged to automate analog design given its applicability across various circuit topologies and technologies. Traditional BO methods employ black-box Gaussian Process surrogate models and optimized labeled data queries to find optimization solutions by trading off between exploration and exploitation. However, the search for the optimal design solution in BO can be expensive from both a computational and data usage point of view, particularly for high-dimensional optimization problems. This paper presents ADO-LLM, the first work integrating large language models (LLMs) with Bayesian Optimization for analog design optimization. ADO-LLM leverages the LLM's ability to infuse domain knowledge to rapidly generate viable design points to remedy BO's inefficiency in finding high-value design areas specifically under the limited design space coverage of the BO's probabilistic surrogate model. In the meantime, sampling of design points evaluated in the iterative BO process provides quality demonstrations for the LLM to generate high-quality design points while leveraging infused broad design knowledge. Furthermore, the diversity brought by BO's exploration enriches the contextual understanding of the LLM and allows it to more broadly search in the design space and prevent repetitive and redundant suggestions. We evaluate the proposed framework on two different types of analog circuits and demonstrate notable improvements in design efficiency and effectiveness.
Yuxuan Yin, Yu Wang 0167, Boxun Xu, Peng Li 0001
ICCAD4
2024 High-Dimensional Bayesian Optimization via Semi-Supervised Learning with Optimized Unlabeled Data Sampling
abstract
We introduce a novel semi-supervised learning approach, named Teacher-Student Bayesian Optimization ($\texttt{TSBO}$), integrating the teacher-student paradigm into BO to minimize expensive labeled data queries for the first time. $\texttt{TSBO}$ incorporates a teacher model, an unlabeled data sampler, and a student model. The student is trained on unlabeled data locations generated by the sampler, with pseudo labels predicted by the teacher. The interplay between these three components implements a unique selective regularization to the teacher in the form of student feedback. This scheme enables the teacher to predict high-quality pseudo labels, enhancing the generalization of the GP surrogate model in the search space. To fully exploit $\texttt{TSBO}$, we propose two optimized unlabeled data samplers to construct effective student feedback that well aligns with the objective of Bayesian optimization. Furthermore, we quantify and leverage the uncertainty of the teacher-student model for the provision of reliable feedback to the teacher in the presence of risky pseudo-label predictions. $\texttt{TSBO}$ demonstrates significantly improved sample-efficiency in several global optimization tasks under tight labeled data budgets. The implementation is available at https://github.com/reminiscenty/TSBO-Official.
Yuxuan Yin, Yu Wang 0167, Peng Li 0001
ICML3
2024 Systolic Array Acceleration of Spiking Neural Networks with Application-Independent Split-Time Temporal Coding
abstract
Spiking Neural Networks (SNNs) are brain-inspired computing models with event-driven based low-power operations and unique temporal dynamics. However, temporal dynamics in SNNs pose a significant overhead in accelerating neural computations and limit the computing capabilities of neuromorphic accelerators. Especially, unstructured sparsity emergent in both space and time, i.e., across neurons and time points, and iterative computations across time points cause a primary bottleneck in data movement.
Jeongjun Lee, Peng Li 0001
ISLPED2
2024 Pareto Optimization of Analog Circuits Using Reinforcement Learning
abstract
Analog circuit optimization and design presents a unique set of challenges in the IC design process. Many applications require the designer to optimize for multiple competing objectives, which poses a crucial challenge. Motivated by these practical aspects, we propose a novel method to tackle multi-objective optimization for analog circuit design in continuous action spaces. In particular, we propose to (i) extrapolate current techniques in Multi-Objective Reinforcement Learning to continuous state and action spaces and (ii) provide for a dynamically tunable trained model to query user defined preferences in multi-objective optimization in the analog circuit design context.
Karthik Somayaji Nanjangud Suryanarayana, Peng Li 0001
ACM Trans. Design Autom. Electr. Syst.2
2023 AutoNF: Automated Architecture Optimization of Normalizing Flows with Unconstrained Continuous Relaxation Admitting Optimal Discrete Solution
abstract
Normalizing flows (NF) build upon invertible neural networks and have wide applications in probabilistic modeling. Currently, building a powerful yet computationally efficient flow model relies on empirical fine-tuning over a large design space. While introducing neural architecture search (NAS) to NF is desirable, the invertibility constraint of NF brings new challenges to existing NAS methods whose application is limited to unstructured neural networks. Developing efficient NAS methods specifically for NF remains an open problem. We present AutoNF, the first automated NF architectural optimization framework. First, we present a new mixture distribution formulation that allows efficient differentiable architecture search of flow models without violating the invertibility constraint. Second, under the new formulation, we convert the original NP-hard combinatorial NF architectural optimization problem to an unconstrained continuous relaxation admitting the discrete optimal architectural solution, circumventing the loss of optimality due to binarization in architectural optimization. We evaluate AutoNF with various density estimation datasets and show its superior performance-cost trade-offs over a set of existing hand-crafted baselines.
Yu Wang 0167, Ján Drgona, Jiaxin Zhang 0005, Karthik Somayaji Nanjangud Suryanarayana, Malachi Schram, Frank Liu 0001, Peng Li 0001
AAAI7
2023 Recognizing Wafer Map Patterns Using Semi-Supervised Contrastive Learning with Optimized Latent Representation Learning and Data Augmentation
abstract
Wafer map analysis is essential for process issue detection and yield improvement in semiconductor manufacturing. Accurate wafer map pattern recognition facilitates root-causing of abnormal chip fabrication conditions. However, manually annotating wafer map data is expensive and time-consuming, which drives up the demand for exploring label-efficient methods for wafer analysis. This paper proposes a novel contrastive learning framework for wafer map pattern feature extraction and classification. Under the semi-supervised learning setting, the proposed approach aims at learning from a large amount of unlabeled data while efficiently exploiting a small amount of expensive labeled data. Our method utilizes supervised contrastive learning on a small amount of labeled data to learn a better latent space representation with well-separated wafer pattern classes. Furthermore, a dual-encoder latent-space model is incorporated to best optimize the simultaneous use of labeled, unlabeled data, and varying types of data augmentations for representation learning. Finally, we enrich the semantics of the learned latent representation space by introducing a novel interwafer data augmentation to synthesize data which are not present in the given dataset. Experiments show that our method leads existing wafer pattern recognition techniques including recent contrastive learning based approaches by a large performance gain, and suggest that superior accuracy may be achieved simply by semi-supervised learning without resorting to labeling-intensive supervised learning.
Zihu Wang, Hanbin Hu, Peng Li 0001
ITC4
2023 Domain-Specific Machine Learning Based Minimum Operating Voltage Prediction Using On-Chip Monitor Data
abstract
Determining the minimum operating voltage ($V_{min}$) of chip designs is critical for low power dissipation and assurance of quality and functional safety during manufacturing tests and in-field monitoring. We demonstrate how on-chip monitor data can be leveraged to provide accurate minimum operating voltage prediction using a domain-specific machine learning approach. Given limited measured chip data, the key challenge in developing a machine learning approach is to provide an accurate prediction while addressing overfitting and selecting a subset of optimal features. To this end, we propose to utilize a novel monotonic lattice neural network architecture that is geared towards accurate prediction by imposing domain-specific monotonic relationships between the input sensor data and$V_{min}$. Furthermore, we perform an effective feature selection by considering both the correlation between each feature and$V_{min}$as well as the co-linearity between the features. Experiments demonstrate superior performance in comparison with linear regression and conventional neural networks.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
ITC4
2023 Comprehensive SNN Compression Using ADMM Optimization and Activity Regularization
abstract
As well known, the huge memory and compute costs of both artificial neural networks (ANNs) and spiking neural networks (SNNs) greatly hinder their deployment on edge devices with high efficiency. Model compression has been proposed as a promising technique to improve the running efficiency via parameter and operation reduction, whereas this technique is mainly practiced in ANNs rather than SNNs. It is interesting to answer how much an SNN model can be compressed without compromising its functionality, where two challenges should be addressed: 1) the accuracy of SNNs is usually sensitive to model compression, which requires an accurate compression methodology and 2) the computation of SNNs is event-driven rather than static, which produces an extra compression dimension on dynamic spikes. To this end, we realize a comprehensive SNN compression through three steps. First, we formulate the connection pruning and weight quantization as a constrained optimization problem. Second, we combine spatiotemporal backpropagation (STBP) and alternating direction method of multipliers (ADMMs) to solve the problem with minimum accuracy loss. Third, we further propose activity regularization to reduce the spike events for fewer active operations. These methods can be applied in either a single way for moderate compression or a joint way for aggressive compression. We define several quantitative metrics to evaluate the compression performance for SNNs. Our methodology is validated in pattern recognition tasks over MNIST, N-MNIST, CIFAR10, and CIFAR100 datasets, where extensive comparisons, analyses, and insights are provided. To the best of our knowledge, this is the first work that studies SNN compression in a comprehensive manner by exploiting all compressible components and achieves better results.
Lei Deng 0003, Yujie Wu 0002, Yifan Hu 0013, Ling Liang 0003, Guoqi Li 0002, Xing Hu 0001, Yufei Ding 0001, Peng Li 0001, Yuan Xie 0001
IEEE Trans. Neural Networks Learn. Syst.8
2023 Exploring Adversarial Attack in Spiking Neural Networks With Spike-Compatible Gradient
abstract
Spiking neural network (SNN) is broadly deployed in neuromorphic devices to emulate brain function. In this context, SNN security becomes important while lacking in-depth investigation. To this end, we target the adversarial attack against SNNs and identify several challenges distinct from the artificial neural network (ANN) attack: 1) current adversarial attack is mainly based on gradient information that presents in a spatiotemporal pattern in SNNs, hard to obtain with conventional backpropagation algorithms; 2) the continuous gradient of the input is incompatible with the binary spiking input during gradient accumulation, hindering the generation of spike-based adversarial examples; and 3) the input gradient can be all-zeros (i.e., vanishing) sometimes due to the zero-dominant derivative of the firing function. Recently, backpropagation through time (BPTT)-inspired learning algorithms are widely introduced into SNNs to improve the performance, which brings the possibility to attack the models accurately given spatiotemporal gradient maps. We propose two approaches to address the above challenges of gradient-input incompatibility and gradient vanishing. Specifically, we design a gradient-to-spike (G2S) converter to convert continuous gradients to ternary ones compatible with spike inputs. Then, we design a restricted spike flipper (RSF) to construct ternary gradients that can randomly flip the spike inputs with a controllable turnover rate, when meeting all-zero gradients. Putting these methods together, we build an adversarial attack methodology for SNNs. Moreover, we analyze the influence of the training loss function and the firing threshold of the penultimate layer on the attack effectiveness. Extensive experiments are conducted to validate our solution. Besides the quantitative analysis of the influence factors, we also compare SNNs and ANNs against adversarial attacks under different attack methods. This work can help reveal what happens in SNN attacks and might stimulate more research on the security of SNN models and neuromorphic devices.
Ling Liang 0003, Xing Hu 0001, Lei Deng 0003, Yujie Wu 0002, Guoqi Li 0002, Yufei Ding 0001, Peng Li 0001, Yuan Xie 0001
IEEE Trans. Neural Networks Learn. Syst.7
2022 Parallel Time Batching: Systolic-Array Acceleration of Sparse Spiking Neural Computation
abstract
Spiking Neural Networks (SNNs) are brain- inspired computing models incorporating unique temporal dynamics and event-driven processing. Rich dynamics in both space and time offer great challenges and opportunities for efficient processing of sparse spatiotemporal data compared with conventional artificial neural networks (ANNs). Specifically, the additional overheads for handling the added temporal dimension limit the computational capabilities of neuromorphic accelerators. Iterative processing at every time-point with sparse inputs in a temporally sequential manner not only degrades the utilization of the systolic array but also intensifies data movement.In this work, we propose a novel technique and architecture that significantly improve utilization and data movement while efficiently handling temporal sparsity of SNNs on systolic arrays. Unlike time-sequential processing in conventional SNN accelerators, we pack multiple time points into a single time window (TW) and process the computations induced by active synaptic inputs falling under several TWs in parallel, leading to the proposed parallel time batching. It allows weight reuse across multiple time points and enhances the utilization of the systolic array with reduced idling of processing elements, overcoming the irregularity of sparse firing activities. We optimize the granularity of time-domain processing, i.e., the TW size, which significantly impacts the data reuse and utilization. We further boost the utilization efficiency by simultaneously scheduling non-overlapping sparse spiking activities onto the array. The proposed architectures offer a unifying solution for general spiking neural networks with commonly exhibited temporal sparsity, a key challenge in hardware acceleration, delivering 248X energy-delay product (EDP) improvement on average compared to an SNN baseline for accelerating various networks. Compared to ANN based accelerators, our approach improves EDP by 47X on the CIFAR10 dataset.
Jeong-Jun Lee, Peng Li 0001
HPCA3
2022 SaARSP: An Architecture for Systolic-Array Acceleration of Recurrent Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are brain-inspired event-driven models of computation with promising ultra-low energy dissipation. Rich network dynamics emergent in recurrent spiking neural networks (R-SNNs) can form temporally based memory, offering great potential in processing complex spatiotemporal data. However, recurrence in network connectivity produces tightly coupled data dependency in both space and time, rendering hardware acceleration of R-SNNs challenging. We present the first work to exploit spatiotemporal parallelisms to accelerate the R-SNN-based inference on systolic arrays using an architecture called SaARSP. We decouple the processing of feedforward synaptic connections from that of recurrent connections to allow for the exploitation of parallelisms across multiple time points. We propose a novel time window size optimization (TWSO) technique, to further explore the temporal granularity of the proposed decoupling in terms of optimal time window size and reconfiguration of the systolic array considering layer-dependent connectivity to boost performance. Stationary dataflow and time window size are jointly optimized to trade off between weight data reuse and movements of partial sums, the two bottlenecks in latency and energy dissipation of the accelerator. The proposed systolic-array architecture offers a unifying solution to an acceleration of both feedforward and recurrent SNNs, and delivers 4,000X EDP improvement on average for different R-SNN benchmarks over a conventional baseline.
Jeong-Jun Lee, Yuan Xie 0001, Peng Li 0001
ACM J. Emerg. Technol. Comput. Syst.4
2022 H2Learn: High-Efficiency Learning Accelerator for High-Accuracy Spiking Neural Networks
abstract
Although spiking neural networks (SNNs) take benefits from the bioplausible neural modeling, the low accuracy under the common local synaptic plasticity learning rules limits their application in many practical tasks. Recently, an emerging SNN supervised learning algorithm inspired by backpropagation through time (BPTT) from the domain of artificial neural networks (ANNs) has successfully boosted the accuracy of SNNs, and helped improve the practicability of SNNs. However, current general-purpose processors suffer from low efficiency when performing BPTT for SNNs due to the ANN-tailored optimization. On the other hand, current neuromorphic chips cannot support BPTT because they mainly adopt local synaptic plasticity rules for simplified implementation. In this work, we propose H2Learn, a novel architecture that can achieve high efficiency for BPTT-based SNN learning, which ensures high accuracy of SNNs. At the beginning, we characterized the behaviors of BPTT-based SNN learning. Benefited from the binary spike-based computation in the forward pass and weight update, we first design look-up table (LUT)-based processing elements in the forward engine and weight update engine to make accumulations implicit and to fuse the computations of multiple input points. Second, benefited from the rich sparsity in the backward pass, we design a dual-sparsity-aware backward engine, which exploits both input and output sparsity. Finally, we apply a pipeline optimization between different engines to build an end-to-end solution for the BPTT-based SNN learning. Compared with the modern NVIDIA V100 GPU, H2Learn achieves$7.38\times $area saving,$5.74-10.20\times $speedup, and$5.25-7.12\times $energy saving on several benchmark datasets.
Ling Liang 0003, Zheng Qu 0002, Zhaodong Chen 0001, Fengbin Tu, Yujie Wu 0002, Lei Deng 0003, Guoqi Li 0002, Peng Li 0001, Yuan Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2021 Algorithm and Hardware Co-Design for FPGA Acceleration of Hamiltonian Monte Carlo Based No-U-Turn Sampler
abstract
Monte Carlo (MC) methods are widely used in many research areas such as physical simulation, statistical analysis, and machine learning. Application of MC methods requires drawing fast mixing samples from a given probability distribution. Among existing sampling methods, the Hamiltonian Monte Carlo (HMC) utilizes gradient information during Hamiltonian simulation and can produce fast mixing samples at the highest efficiency. However, without carefully chosen simulation parameters for a specific problem, HMC generally suffers from simulation locality and computation waste. As a result, the No-U-Turn Sampler (NUTS) has been proposed to automatically tune these parameters during simulation and is the current state-of-the-art sampling algorithm. However, application of NUTS requires frequent gradient calculation of a given distribution and high-volume vector processing, especially for large-scale problems, leading to drawing an expensively large number of samples and a desire of hardware acceleration. While some hardware acceleration works have been proposed for traditional Markov Chain Monte Carlo (MCMC) and HMC methods, there is no existing work targeting hardware acceleration of the NUTS algorithm. In this paper, we present the first NUTS accelerator on FPGA while addressing the high complexity of this state-of-the-art algorithm. Our hardware and algorithm co-optimizations include an incremental resampling technique which leads to a more memory efficient architecture and pipeline optimization for multi-chain sampling to maximize the throughput. We also explore three levels of parallelism in the NUTS accelerator to further boost performance. Compared with optimized C++ NUTS package: RSTAN, our NUTS accelerator can reach a maximum speedup of 50.6X and an energy improvement of 189.7X.
Yu Wang 0167, Peng Li 0001
ASAP2
2021 Reversible Gating Architecture for Rare Failure Detection of Analog and Mixed-Signal Circuits
abstract
Due to the growing complexity and numerous manufacturing variation in safety-critical analog and mixed-signal (AMS) circuit design, rare failure detection in the high-dimensional variational space is one of the major challenges in AMS verification. Efficient AMS failure detection is very demanding with limited samples on account of high simulation and manufacturing cost. In this work, we combine a reversible network and a gating architecture to identify essential features from datasets and reduce feature dimension for fast failure detection. While reversible residual networks (RevNets) have been actively studied for its restoration ability from output to input without the loss of information, the gating network facilitates the RevNet to aim at effective dimension reduction. We incorporate the proposed reversible gating architecture into Bayesian optimization (BO) framework to reduce the dimensionality of BO embedding important features clarified by gating fusion weights so that the failure points can be efficiently located. Furthermore, we propose a conditional density estimation of important and non-important features to extract high-dimensional original input features from the low-dimension important features, improving the efficiency of the proposed methods. The improvements of our proposed approach on rare failure detection is demonstrated in AMS data under the high-dimensional process variations.
Myung Seok Shim, Hanbin Hu, Peng Li 0001
DAC3
2021 Prioritized Reinforcement Learning for Analog Circuit Optimization With Design Knowledge
abstract
Analog circuit design and optimization manifests as a critical phase in IC design, which still heavily relies on extensive and time-consuming manual designing by experienced experts. In recent years, the development of reinforcement learning (RL) algorithms draws attention with related techniques being introduced into the analog design field for circuit optimization. However, for robust and efficient analog circuit design, a smart and rapid search for high-quality design points is more desired than finding a globally optimal agent as in traditional RL applications, which was a point not fully considered in some previous works. In this work, we propose three techniques within the RL framework aiming at fast high-quality design point search in a data efficient manner. In particular, we (i) incorporate design knowledge from experienced designers into the critic network design to achieve a better reward evaluation with less data; (ii) guide the RL training with non-uniform sampling techniques prioritizing exploitation over high quality designs and exploration for poorly-trained space; (iii) leverage the trained critic network and limited additional circuit simulation for smart and efficient sampling to get high-quality design points. The experimental results demonstrate the effectiveness and efficiency of our proposed techniques.
Karthik Somayaji Nanjangud Suryanarayana, Hanbin Hu, Peng Li 0001
DAC3
2021 Backpropagated Neighborhood Aggregation for Accurate Training of Spiking Neural Networks
abstract
While Backpropagation (BP) has been applied to spiking neural networks (SNNs) achieving encouraging results, a key challenge involved is to backpropagate a differentiable continuous-valued loss over layers of spiking neurons exhibiting discontinuous all-or-none firing activities. Existing methods deal with this difficulty by introducing compromises that come with their own limitations, leading to potential performance degradation. We propose a novel BP-like method, called neighborhood aggregation (NA), which computes accurate error gradients guiding weight updates that may lead to discontinuous modifications of firing activities. NA achieves this goal by aggregating the error gradient over multiple spike trains in the neighborhood of the present spike train of each neuron. The employed aggregation is based on a generalized finite difference approximation with a proposed distance metric quantifying the similarity between a given pair of spike trains. Our experiments show that the proposed NA algorithm delivers state-of-the-art performance for SNN training on several datasets including CIFAR10.
Yukun Yang 0001, Peng Li 0001
ICML3
2021 Systolic-Array Spiking Neural Accelerators with Dynamic Heterogeneous Voltage Regulation
abstract
Spiking neural networks (SNNs) have emerged as a new generation of neural networks, presenting a brain-inspired event-driven model with advantages in spatiotemporal information processing. Due to the need for high power consumption of compute-intensive neural accelerators, adequate power delivery network (PDN) design is a key requirement to ensure power efficiency and integrity. However, PDN design for SNN accelerators has not been extensively studied despite its great potential benefit in energy efficiency. In this paper, we present the first study on dynamic heterogeneous voltage regulation (HVR) for spiking neural accelerators to maximize system energy efficiency while ensuring power integrity. We propose a novel sparse-workload-aware dynamic PDN control policy, which enables high energy efficiency of sparse spiking computation on a systolic array. By exploring sparse inputs and all-or-none nature of spiking computations for PDN control, we explore different types of PDNs to accelerate spiking convolutional neural networks (S-CNNs) trained with the dynamic vision sensor (DVS) gesture dataset. Furthermore, we demonstrate various power gating schemes to further optimize the proposed PDN architecture, which leads to a more than a three-fold reduction in total energy overhead for spiking neural computations on systolic array-based accelerators.
Jeong-Jun Lee, Peng Li 0001
IJCNN4
2021 Spiking Neural Networks with Laterally-Inhibited Self-Recurrent Units
abstract
In biological brains, recurrent connections play a crucial role in cortical computation, modulation of network dynamics, and communication. However, in recurrent spiking neural networks (SNNs), recurrence is mostly constructed by random connections. How excitatory and inhibitory recurrent connections affect network responses and what kinds of connectivity benefit learning performance is still obscure. In this work, we propose a novel recurrent structure called the Laterally-Inhibited Self-Recurrent Unit (LISR), which consists of one excitatory neuron with a self-recurrent connection wired together with an inhibitory neuron through excitatory and inhibitory synapses. The self-recurrent connection of the excitatory neuron mitigates the information loss caused by the firing-and-resetting mechanism and maintains the long-term neuronal memory. The lateral inhibition from the inhibitory neuron to the corresponding excitatory neuron, on the one hand, adjusts the firing activity of the latter. On the other hand, it plays as a forget gate to clear the memory of the excitatory neuron. Based on speech and image datasets commonly used in neuromorphic computing, RSNNs based on the proposed LISR improve performance significantly by up to 9.26% over feedforward SNNs trained by a state-of-the-art backpropagation method with similar computational costs.
Peng Li 0001
IJCNN2
2021 Semi-supervised Wafer Map Pattern Recognition using Domain-Specific Data Augmentation and Contrastive Learning
abstract
Wafer map pattern recognition is instrumental for detecting systemic manufacturing process issues. However, high cost in labeling wafer patterns renders it impossible to leverage large amounts of valuable unlabeled data in conventional machine learning based wafer map pattern prediction. We proposed a contrastive learning framework for semi-supervised learning and prediction of wafer map patterns. Our framework incorporates an encoder to learn good representation for wafer maps in an unsupervised manner, and a supervised head to recognize wafer map patterns. In particular, contrastive learning is applied for the unsupervised encoder representation learning supported by augmented data generated by different transformations (views) of wafer maps. We identified a set of transformations to effectively generate similar variants of each original pattern. We further proposed a novel rotation-twist transformation to augment wafer map data by rotating each given wafer map for which the angle of rotation is a smooth function of the radius. Experimental results demonstrate that the proposed semi-supervised learning framework greatly improves recognition accuracy compared to traditional supervised methods, and the rotation-twist transformation further enhances the recognition accuracy in both semi-supervised and supervised tasks.
Hanbin Hu, Peng Li 0001
ITC3
2021 Skip-Connected Self-Recurrent Spiking Neural Networks With Joint Intrinsic Parameter and Synaptic Weight Training
abstract
As an important class of spiking neural networks (SNNs), recurrent spiking neural networks (RSNNs) possess great computational power and have been widely used for processing sequential data like audio and text. However, most RSNNs suffer from two problems. First, due to the lack of architectural guidance, random recurrent connectivity is often adopted, which does not guarantee good performance. Second, training of RSNNs is in general challenging, bottlenecking achievable model accuracy. To address these problems, we propose a new type of RSNN, skip-connected self-recurrent SNNs (ScSr-SNNs). Recurrence in ScSr-SNNs is introduced by adding self-recurrent connections to spiking neurons. The SNNs with self-recurrent connections can realize recurrent behaviors similar to those of more complex RSNNs, while the error gradients can be more straightforwardly calculated due to the mostly feedforward nature of the network. The network dynamics is enriched by skip connections between nonadjacent layers. Moreover, we propose a new backpropagation (BP) method, backpropagated intrinsic plasticity (BIP), to boost the performance of ScSr-SNNs further by training intrinsic model parameters. Unlike standard intrinsic plasticity rules that adjust the neuron's intrinsic parameters according to neuronal activity, the proposed BIP method optimizes intrinsic parameters based on the backpropagated error gradient of a well-defined global loss function in addition to synaptic weight training. Based on challenging speech, neuromorphic speech, and neuromorphic image data sets, the proposed ScSr-SNNs can boost performance by up to 2.85% compared with other types of RSNNs trained by state-of-the-art BP methods.
Peng Li 0001
Neural Comput.2
2021 Exponential Stabilization of Inertial Memristive Neural Networks With Multiple Time Delays
abstract
This article investigates the global exponential stabilization (GES) of inertial memristive neural networks with discrete and distributed time-varying delays (DIMNNs). By introducing the inertial term into memristive neural networks (MNNs), DIMNNs are formulated as the second-order differential equations with discontinuous right-hand sides. Via a variable transformation, the initial DIMNNs are rewritten as the first-order differential equations. By exploiting the theories of differential inclusion, inequality techniques, and the comparison strategy, the p th moment GES ( p ≥ 1 ) of the addressed DIMNNs is presented in terms of algebraic inequalities within the sense of Filippov, which enriches and extends some published results. In addition, the global exponential stability of MNNs is also performed in the form of an M-matrix, which contains some existing ones as special cases. Finally, two simulations are carried out to validate the correctness of the theories, and an application is developed in pseudorandom number generation.
Yin Sheng, Tingwen Huang, Zhigang Zeng, Peng Li 0001
IEEE Trans. Cybern.4
2020 Dynamic Heterogeneous Voltage Regulation for Systolic Array-Based DNN Accelerators
abstract
With the growing performance and wide application of deep neural networks (DNNs), recent years have seen enormous efforts on DNN accelerator hardware design for platforms from mobile devices to data centers. The systolic array has been a popular architectural choice for many proposed DNN accelerators with hundreds to thousands of processing elements (PEs) for parallel computing. Systolic array-based DNN accelerators for datacenter applications have high power consumption and nonuniform workload distribution, which makes power delivery network (PDN) design challenging. Server-class multicore processors have benefited from distributed on-chip voltage regulation and heterogeneous voltage regulation (HVR) for improving energy efficiency while guaranteeing power delivery integrity. This paper presents the first work on HVR-based PDN architecture and control for systolic array-based DNN accelerators. We propose to employ a PDN architecture comprising heterogeneous on-chip and off-chip voltage regulators and multiple power domains. By analyzing patterns of typical DNN workloads via a modeling framework, we propose a DNN workload-aware dynamic PDN control policy to maximize system energy efficiency while ensuring power integrity. We demonstrate significant energy efficiency improvements brought by the proposed PDN architecture, dynamic control, and power gating, which lead to a more than five-fold reduction of leakage energy and PDN energy overhead for systolic array DNN accelerators.
Joseph Riad, Edgar Sánchez-Sinencio, Peng Li 0001
ICCD4
2020 Reconfigurable Dataflow Optimization for Spatiotemporal Spiking Neural Computation on Systolic Array Accelerators
abstract
Spiking neural networks (SNNs) offer a promising biologically-plausible computing model and lend themselves to ultra-low-power event-driven processing on neuromorphic processors. Compared with the conventional artificial neural networks, SNNs are well-suited for processing complex spatiotemporal data. Despite its significance, dataflow optimization of spiking neural accelerator architectures has not been extensively studied. Recognizing the need for efficient processing of complex spatiotemporal data while considering the all-or-none nature of spiking activities, we propose holistic reconfigurable dataflow optimization for systolic array acceleration of spiking convolutional networks (S-CNNs). A novel scheme is introduced for parallel acceleration of computation across multiple time points, which further allows for systemic optimization of variable tiling for a large performance and efficiency gains. We show how variable tiling, in particular, the positioning of the temporal dimension, can be targeted to optimize data movement, throughput, and energy efficiency. Furthermore, we explore joint layer-dependent dataflow and accelerator hardware optimization to further boost performance and energy efficiency. To support systemic design space exploration, we develop an SNN dataflow simulator capable of analyzing the throughput and energy dissipation of systolic array accelerators for any targeted S-CNN while considering the inherent spatiotemporal characteristics of spiking neural computation. The proposed techniques deliver orders of magnitude of improvements on throughput, energy efficiency, and delay-energy product for accelerating deep Alexnet and VGG-16 SNNs.
Jeong-Jun Lee, Peng Li 0001
ICCD2
2020 Advanced Outlier Detection Using Unsupervised Learning for Screening Potential Customer Returns
abstract
Due to the extreme scarcity of customer failure data, it is challenging to reliably screen out those rare defects within a high-dimensional input feature space formed by the relevant parametric test measurements. In this paper, we study several unsupervised learning techniques based on six industrial test datasets, and propose to train a more robust unsupervised learning model by self-labeling the training data via a set of transformations. Using the labeled data we train a multi-class classifier through supervised training. The goodness of the multiclass classification decisions with respect to an unseen input data is used as a normality score to defect anomalies. Furthermore, we propose to use reversible information lossless transformations to retain the data information and boost the performance and robustness of the proposed self-labeling approach.
Hanbin Hu, Peng Li 0001
ITC4
2020 Temporal Spike Sequence Learning via Backpropagation for Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are well suited for spatio-temporal learning and implementations on energy-efficient event-driven neuromorphic processors. However, existing SNN error backpropagation (BP) methods lack proper handling of spiking discontinuities and suffer from low performance compared with the BP methods for traditional artificial neural networks. In addition, a large number of time steps are typically required to achieve decent performance, leading to high latency and rendering spike-based computation unscalable to deep architectures. We present a novel Temporal Spike Sequence Learning Backpropagation (TSSL-BP) method for training deep SNNs, which breaks down error backpropagation across two types of inter-neuron and intra-neuron dependencies and leads to improved temporal learning precision. It captures inter-neuron dependencies through presynaptic firing times by considering the all-or-none characteristics of firing activities and captures intra-neuron dependencies by handling the internal evolution of each neuronal state in time. TSSL-BP efficiently trains deep SNNs within a much shortened temporal window of a few steps while improving the accuracy for various image classification datasets including CIFAR10.
Peng Li 0001
NeurIPS2
2020 Quantized synchronization of memristive neural networks with time-varying delays via super-twisting algorithm
Bo Sun 0002, Shiping Wen 0001, Tingwen Huang, Yiran Chen 0001, Peng Li 0001
Neurocomputing6
2020 Rethinking the performance comparison between SNNS and ANNS
Lei Deng 0003, Yujie Wu 0002, Xing Hu 0001, Ling Liang 0003, Yufei Ding 0001, Guoqi Li 0002, Guang-She Zhao, Peng Li 0001, Yuan Xie 0001
Neural Networks8
2020 Sliding mode control of neural networks via continuous or periodic sampling event-triggering algorithm
Shiqin Wang, Yuting Cao, Tingwen Huang, Yiran Chen 0001, Peng Li 0001, Shiping Wen 0001
Neural Networks5
2020 A smoothing neural network for minimization l1-lp in sparse signal reconstruction with measurement noises
Xing He 0001, Tingwen Huang, Junjian Huang, Peng Li 0001
Neural Networks5
2019 Enabling High-Dimensional Bayesian Optimization for Efficient Failure Detection of Analog and Mixed-Signal Circuits
abstract
With increasing design complexity and stringent robustness requirements in application such as automotive electronics, analog and mixed-signal (AMS) verification becomes akey bottleneck. Rare failure detection in a high-dimensional parameter space using minimal expensive simulation data is a major challenge. We address this challenge under a Bayesian learning framework using Bayesian optimization (BO). We formulate the failure detection as a BO problem where a chosen acquisition function is optimized to select the next (set of) optimal simulation sampling point(s) such that rare failures may be detected using a small amount of data. While providing an attractive black-box solution to design verification, in practice BO is limited in its ability in dealing with high-dimensional problems. We propose to use random embedding to effectively reduce the dimensionality of a given verification problem to improve both the quality of BO-based optimal sampling and computational efficiency. We demonstrate the success of the proposed approach on detecting rare design failures under high-dimensional process variations which are completely missed by competitive smart sampling and BO techniques without dimension reduction.
Hanbin Hu, Peng Li 0001, Jianhua Z. Huang
DAC2
2019 Spike-Train Level Backpropagation for Training Deep Recurrent Spiking Neural Networks
abstract
Spiking neural networks (SNNs) well support spatiotemporal learning and energy-efficient event-driven hardware neuromorphic processors. As an important class of SNNs, recurrent spiking neural networks (RSNNs) possess great computational power. However, the practical application of RSNNs is severely limited by challenges in training. Biologically-inspired unsupervised learning has limited capability in boosting the performance of RSNNs. On the other hand, existing backpropagation (BP) methods suffer from high complexity of unrolling in time, vanishing and exploding gradients, and approximate differentiation of discontinuous spiking activities when applied to RSNNs. To enable supervised training of RSNNs under a well-defined loss function, we present a novel Spike-Train level RSNNs Backpropagation (ST-RSBP) algorithm for training deep RSNNs. The proposed ST-RSBP directly computes the gradient of a rated-coded loss function defined at the output layer of the network w.r.t tunable parameters. The scalability of ST-RSBP is achieved by the proposed spike-train level computation during which temporal effects of the SNN is captured in both the forward and backward pass of BP. Our ST-RSBP algorithm can be broadly applied to RSNNs with a single recurrent layer or deep RSNNs with multiple feed-forward and recurrent layers. Based upon challenging speech and image datasets including TI46, N-TIDIGITS, Fashion-MNIST and MNIST, ST-RSBP is able to train RSNNs with an accuracy surpassing that of the current state-of-art SNN BP algorithms and conventional non-spiking deep learning models.
Peng Li 0001
NeurIPS2
2019 Energy-efficient FPGA Spiking Neural Accelerators with Supervised and Unsupervised Spike-timing-dependent-Plasticity
abstract
The liquid state machine (LSM) is a model of recurrent spiking neural networks (SNNs) and provides an appealing brain-inspired computing paradigm for machine-learning applications. Moreover, operated by processing information directly on spiking events, the LSM is amenable to efficient event-driven hardware implementation. However, training SNNs is, in general, a difficult task as synaptic weights shall be updated based on neural firing activities while achieving a learning objective. In this article, we explore bio-plausible spike-timing-dependent-plasticity (STDP) mechanisms to train liquid state machine models with and without supervision. First, we employ a supervised STDP rule to train the output layer of the LSM while delivering good classification performance. Furthermore, a hardware-friendly unsupervised STDP rule is leveraged to train the recurrent reservoir to further boost the performance. We pursue efficient hardware implementation of FPGA LSM accelerators by performing algorithm-level optimization of the two proposed training rules and exploiting the self-organizing behaviors naturally induced by STDP. Several recurrent spiking neural accelerators are built on a Xilinx Zync ZC-706 platform and trained for speech recognition with the TI46 speech corpus as the benchmark. Adopting the two proposed unsupervised and supervised STDP rules outperforms the recognition accuracy of a competitive non-STDP baseline training algorithm by up to 3.47%.
Yu Liu 0028, Sai Sourabh Yenamachintala, Peng Li 0001
ACM J. Emerg. Technol. Comput. Syst.3
2019 Taming the Stability-Constrained Performance Optimization Challenge of Distributed On-Chip Voltage Regulation
abstract
Distributed on-chip voltage regulation is promising for addressing many IC power delivery challenges. However, complex interactions between active regulators and the surrounding parasitic passive RLC network cause stability concern. The recently developed hybrid stability technique provides a unique opportunity for coping with stability of distributed on-chip regulation and enabling efficient localized system design. However, the inherent conservativeness of the hybrid stability theorem (HST) leads to large pessimism in stability evaluation and hence causes overdesign. In this paper, the above challenge is addressed by extending the HST with an optimal frequency-dependent system partitioning technique which can significantly reduce the amount of pessimism in stability analysis. To put the proposed approach on a firm theoretical footing, we prove that the partitioning technique removes the conservativeness without altering the physical system and key theoretical properties of the partitioned blocks are maintained under certain constraints. Upon this, an efficient stability-ensuring power delivery design methodology using an automated design flow is developed to significantly improve power delivery performance. Within a large design space, the proposed approach ensures stability and improves system performance by up to 53%, measured by a figure of merit (FOM), when compared to the classical phase margin design approach, which provides no guarantee of stability. Furthermore, on average our approach boosts the FOM by 113% while consuming 11% less power compared to a reference hybrid stability approach.
Xin Zhan, Peng Li 0001, Edgar Sánchez-Sinencio
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Power Management for Multicore Processors via Heterogeneous Voltage Regulation and Machine Learning Enabled Adaptation
abstract
This work is based on the vision that the ultimate power integrity and efficiency may be best achieved via a heterogeneous chain of voltage processing starting from onboard switching voltage regulators (VRs), to on-chip switching VRs, and finally to networks of distributed on-chip linear VRs. As such, we propose a heterogeneous voltage regulation (HVR) architecture encompassing regulators with complimentary characteristics in response time, size, and efficiency. By exploring the rich heterogeneity and tunability in HVR, we develop systematic workload-aware power management policies to adapt heterogeneous VRs with respect to workload change at multiple temporal scales to significantly improve system power efficiency while providing a guarantee for power integrity. The proposed techniques are further supported by hardware-accelerated machine learning (ML) prediction of nonuniform spatial workload distributions for more accurate HVR adaptation at fine time granularity. Our evaluations based on the PARSEC benchmark suite show that the proposed adaptive three-stage HVR reduces the total system energy dissipation by up to 23.9% and 15.7% on average compared with the conventional static two-stage voltage regulation using off-chip and on-chip switching VRs. Compared with the three-stage static HVR, our runtime control reduces system energy by up to 17.9% and 12.2% on average. Furthermore, the proposed ML prediction offers up to 4.1% reduction of system energy.
Xin Zhan, Edgar Sánchez-Sinencio, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2018 HFMV: hybridizing formal methods and machine learning for verification of analog and mixed-signal circuits
abstract
With increasing design complexity and robustness requirement, analog and mixed-signal (AMS) verification manifests itself as a key bottleneck. While formal methods and machine learning have been proposed for AMS verification, these two techniques suffer from their own limitations, with the former being specifically limited by scalability and the latter by the inherent uncertainty in learning-based models. We present a new direction in AMS verification by proposing a hybrid formal/machine-learning verification technique (HFMV) to combine the best of the two worlds. HFMV adds formalism on the top of a probabilistic learning model while providing a sense of coverage for extremely rare failure detection. HFMV intelligently and iteratively reduces uncertainty of the learning model by a proposed formally-guided active learning strategy and discovers potential rare failure regions in complex high-dimensional parameter spaces. It leads to reliable failure prediction in the case of a failing circuit, or a high-confidence pass decision in the case of a good circuit. We demonstrate that HFMV is able to employ a modest amount of data to identify hard-to-find rare failures which are completely missed by state-of-the-art sampling methods even with high volume sampling data.
Hanbin Hu, Qingran Zheng, Peng Li 0001
DAC4
2018 Design and architectural co-optimization of monolithic 3D liquid state machine-based neuromorphic processor
abstract
A liquid state machine (LSM) is a powerful recurrent spiking neural network shown to be effective in various learning tasks including speech recognition. In this work, we investigate design and architectural co-optimization to further improve the area-energy efficiency of LSM-based speech recognition processors with monolithic 3D IC (M3D) technology. We conduct fine-grained tier partitioning, where individual neurons are folded, and explore the impact of shared memory architecture and synaptic model complexity on the power-performance-area-accuracy (PPAA) benefit of M3D LSM-based speech recognition. In training and classification tasks using spoken English letters, we obtain up to 70.0% PPAA savings over 2D ICs.
Bon Woong Ku, Yu Liu 0028, Yingyezhe Jin, Sandeep Kumar Samal, Peng Li 0001, Sung Kyu Lim
DAC5
2018 Parallelizable Bayesian optimization for analog and mixed-signal rare failure detection with high coverage
abstract
Due to inherent complex behaviors and stringent requirements in analog and mixed-signal (AMS) systems, verification becomes a key bottleneck in the product development cycle. For the first time, we present a Bayesian optimization (BO) based approach to the challenging problem of verifying AMS circuits with stringent low failure requirements. At the heart of the proposed BO process is a delicate balancing between two competing needs: exploitation of the current statistical model for quick identification of highly-likely failures and exploration of undiscovered design space so as to detect hard-to-find failures within a large parametric space. To do so, we simultaneously leverage multiple optimized acquisition functions to explore varying degrees of balancing between exploitation and exploration. This makes it possible to not only detect rare failures which other techniques fail to identify, but also do so with significantly improved efficiency. We further build in a mechanism into the BO process to enable detection of multiple failure regions, hence providing a higher degree of coverage. Moreover, the proposed approach is readily parallelizable, further speeding up failure detection, particularly for large circuits for which acquisition of simulation/measurement data is very time-consuming. Our experimental study demonstrates that the proposed approach is very effective in finding very rare failures and multiple failure regions which existing statistical sampling techniques and other BO techniques can miss, thereby providing a more robust and cost-effective methodology for rare failure detection.
Hanbin Hu, Peng Li 0001, Jianhua Z. Huang
ICCAD2
2018 Area-efficient and low-power face-to-face-bonded 3D liquid state machine design
abstract
As small-form-factor and low-power end devices matter in the cloud networking and Internet-of-Things Era, the bio-inspired neuromorphic architectures attract great attention recently in the hope of reaching the energy-efficiency of brain functions. Out of promising solutions, a liquid state machine (LSM), that consists of randomly and recurrently connected reservoir neurons and trainable readout neurons, has shown a great promise in delivering brain-inspired computing power. In this work, we adopt the state-of-the-art face-to-face (F2F)-bonded 3D IC flow named Compact-2D [4] to the LSM processor design, and study the power-area-accuracy benefits of 3D LSM ICs targeting the next generation commercial-grade neuromorphic computing platforms. First, we analyze how the different size and connection density of a reservoir in the LSM architecture affects the learning performance using the real-world speech recognition benchmark. Also, we explore how much the power-area design overhead should be paid off to enable better classification accuracy. Based on the power-area-accuracy trade-off, we implement a F2F-bonded 3D LSM IC using the optimal LSM architecture, and finally justify that 3D integration practically benefits the LSM processor design in huge form factor and power savings while preserving the best learning performance.
Bon Woong Ku, Yu Liu 0028, Yingyezhe Jin, Peng Li 0001, Sung Kyu Lim
ICCAD4
2018 Hybrid Macro/Micro Level Backpropagation for Training Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are positioned to enable spatio-temporal information processing and ultra-low power event-driven neuromorphic hardware. However, SNNs are yet to reach the same performances of conventional deep artificial neural networks (ANNs), a long-standing challenge due to complex dynamics and non-differentiable spike events encountered in training. The existing SNN error backpropagation (BP) methods are limited in terms of scalability, lack of proper handling of spiking discontinuities, and/or mismatch between the rate-coded loss function and computed gradient. We present a hybrid macro/micro level backpropagation (HM2-BP) algorithm for training multi-layer SNNs. The temporal effects are precisely captured by the proposed spike-train level post-synaptic potential (S-PSP) at the microscopic level. The rate-coded errors are defined at the macroscopic level, computed and back-propagated across both macroscopic and microscopic levels. Different from existing BP methods, HM2-BP directly computes the gradient of the rate-coded loss function w.r.t tunable parameters. We evaluate the proposed HM2-BP algorithm by training deep fully connected and convolutional SNNs based on the static MNIST [14] and dynamic neuromorphic N-MNIST [26]. HM2-BP achieves an accuracy level of 99.49% and 98.88% for MNIST and N-MNIST, respectively, outperforming the best reported performances obtained from the existing SNN BP algorithms. Furthermore, the HM2-BP produces the highest accuracies based on SNNs for the EMNIST [3] dataset, and leads to high recognition accuracy for the 16-speaker spoken English letters of TI46 Corpus [16], a challenging patio-temporal speech recognition benchmark for which no prior success based on SNNs was reported. It also achieves competitive performances surpassing those of conventional deep learning models when dealing with asynchronous spiking streams.
Yingyezhe Jin, Peng Li 0001
NeurIPS3
2018 Online Adaptation and Energy Minimization for Hardware Recurrent Spiking Neural Networks
abstract
The Liquid State Machine (LSM) is a promising model of recurrent spiking neural networks that provides an appealing brain-inspired computing paradigm for machine-learning applications such as pattern recognition. Moreover, processing information directly on spiking events makes the LSM well suited for cost- and energy-efficient hardware implementation. In this article, we systematically present three techniques for optimizing energy efficiency while maintaining good performance of the proposed LSM neural processors from both an algorithmic and hardware implementation point of view. First, to realize adaptive LSM neural processors, thus boost learning performance, we propose a hardware-friendly Spike-Timing Dependent Plastic (STDP) mechanism for on-chip tuning. Then, the LSM processor incorporates a novel runtime correlation-based neuron gating scheme to minimize the power dissipated by reservoir neurons. Furthermore, an activity-dependent clock gating approach is presented to address the energy inefficiency due to the memory-intensive nature of the proposed neural processors. Using two different real-world tasks of speech and image recognition to benchmark, we demonstrate that the proposed architecture boosts the average learning performance by up to 2.0% while reducing energy dissipation by up to 29% compared to a baseline LSM with little extra hardware overhead on a Xilinx Virtex-6 FPGA.
Yu Liu 0028, Yingyezhe Jin, Peng Li 0001
ACM J. Emerg. Technol. Comput. Syst.3
2018 Design Space Exploration of Distributed On-Chip Voltage Regulation Under Stability Constraint
Xin Zhan, Joseph Riad, Peng Li 0001, Edgar Sánchez-Sinencio
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Convergence-Boosted Graph Partitioning using Maximum Spanning Trees for Iterative Solution of Large Linear Circuits
abstract
The ability to solve large linear systems efficiently has been a key part of building a successful circuit simulator such as ones for power grid analysis. Partition-based iterative methods have been widely adopted due to their divide-and-conquer nature and amenability to parallel implementation. However such methods rely heavily on the quality of circuit graph partitioning, and it is extremely challenging to develop a robust partitioning scheme that always generates near-optimal results. In this paper, we present a new line of thinking by integrating two rather distinct schools of iterative methods: partitioning based and support graph based. The former enjoys ease of parallelization, however, lacks a direct control of the numerical properties of the produced partitions. In contrast, the latter operates on the maximum spanning tree (MST) of the circuit graph, which is optimized for fast numerical convergence, but is bottlenecked by its difficulty of parallelization. By combining the two, we propose a partitioning-based preconditioner based on the theory of support graph. The circuit partitioning is guided by the MST of the underlying circuit graph, offering essential guidance for achieving fast convergence. The resulting block-Jacobi-like preconditioner maximizes the numerical benefit inherited from support graph theory while lending itself to straightforward parallelization as a partition-based method. The experimental results on IBM power grid suite and synthetic power grid benchmarks show that our proposed method speeds up the DC simulation by up to 11.5X over an state-of-the-art direct solver.
Peng Li 0001
DAC3
2017 Noise-sensitive feedback loop identification in linear time-varying analog circuits
abstract
The continuing scaling of VLSI technology and design complexity has rendered robustness of analog circuits a significant concern. Parasitic effects may introduce unexpected marginal instability within multiple noise-sensitive loops and hence jeopardize circuit operation and processing precision. The Loop Finder algorithm has been recently proposed to allow detection of noise-sensitive return loops for circuits that are described using a linear time-invariant (LTI) system model. However, many practical circuits such as switched-capacitor filters and mixers present time-varying behaviors which are intrinsically coupled with noise propagation and introduce new noise generation mechanisms. For the first time, we take an in-depth look into the marginal instability of linear periodically time-varying (LPTV) analog circuits and further develop an algorithm for efficient identification of noise-sensitive loops, unifying the solution to noise sensitivity analysis for both LTI and LPTV circuits.
Ang Li 0005, Peng Li 0001, Tingwen Huang, Edgar Sánchez-Sinencio
DATE2
2017 Calcium-modulated supervised spike-timing-dependent plasticity for readout training and sparsification of the liquid state machine
abstract
The Liquid State Machine (LSM) is a promising model of recurrent spiking neural networks. It consists of a fixed recurrent network, or the reservoir, which projects to a readout layer through plastic readout synapses. The classification performance is highly dependent on the training of readout synapses which tend to be very dense and contribute significantly to the overall network complexity. We present a unifying biologically inspired calcium-modulated supervised spike-timing-dependent plasticity (STDP) approach to training and sparsification of readout synapses, where supervised temporal learning is modulated by the post-synaptic firing level characterized by the post-synaptic calcium concentration. The proposed approach prevents synaptic weight saturation, boosts learning performance, and sparsifies the connectivity between the reservoir and readout layer. Using the recognition rate of spoken English letters adopted from the TI46 speech corpus as a measure of performance, we demonstrate that the proposed approach outperforms a baseline supervised STDP mechanism by up to 25%, and a competitive non-STDP spike-dependent training algorithm by up to 2.7%. Furthermore, it can prune out up to 30% of readout synapses without causing significant performance degradation.
Yingyezhe Jin, Peng Li 0001
IJCNN2
2017 Navigating mobile robots to target in near shortest time using reinforcement learning with spiking neural networks
abstract
The autonomous navigation of mobile robots in unknown environments is of great interest in mobile robotics. This article discusses a new strategy to navigate to a known target location in an unknown environment using a combination of the “go-to-goal” approach and reinforcement learning with biologically realistic spiking neural networks. While the “go-to-goal” approach itself might lead to a solution for most environments, the added neural reinforcement learning in this work results in a strategy that takes the robot from a starting position to a target location in a near shortest possible time. To achieve the goal, we propose a reinforcement learning approach based on spiking neural networks. The presented biologically motivated delayed reward mechanism using eligibility traces results in a greedy approach that leads the robot to the target in a close to shortest possible time.
Amarnath Mahadevuni, Peng Li 0001
IJCNN2
2017 Biologically inspired reinforcement learning for mobile robot collision avoidance
abstract
Collision avoidance is a key technology enabling applications such as autonomous vehicles and robots. Various reinforcement learning techniques such as the popular Q-learning algorithms have emerged as a promising solution for collision avoidance in robotics. While spiking neural networks (SNNs), the third generation model of neural networks, have gained increased interest due to their closer resemblance to biological neural circuits in the brain, the application of SNNs to mobile robot navigation has not been well studied. Under the context of reinforcement learning, this paper aims to investigate the potential of biologically-motivated spiking neural networks for goal-directed collision avoidance in reasonably complex environments. Unlike the existing additive reward-modulated spike-timing dependent plasticity learning rule (A-RM-STDP), for the first time, we explore a new multiplicative RM-STDP scheme (M-RM-STDP) for the targeted application. Furthermore, we propose a more biologically plausible feed-forward spiking neural network architecture with fine-grained global rewards. Finally, by combining the above two techniques we demonstrate a further improved solution to collision avoidance. Our proposed approaches not only completely outperform Q-learning for cases where Q-learning can hardly reach the target without collision, but also significantly outperform a baseline SNN with A-RM-STDP in terms of both success rate and the quality of navigation trajectories.
Myung Seok Shim, Peng Li 0001
IJCNN2
2017 Exploring sparsity of firing activities and clock gating for energy-efficient recurrent spiking neural processors
abstract
As a model of recurrent spiking neural networks, the Liquid State Machine (LSM) offers a powerful brain-inspired computing platform for pattern recognition and machine learning applications. While operated by processing neural spiking activities, the LSM naturally lends itself to an efficient hardware implementation via exploration of typical sparse firing patterns emerged from the recurrent neural network and smart processing of computational tasks that are orchestrated by different firing events at runtime. We explore these opportunities by presenting a LSM processor architecture with integrated on-chip learning and its FPGA implementation. Our LSM processor leverage the sparsity of firing activities to allow for efficient event-driven processing and activity-dependent clock gating. Using the spoken English letters adopted from the TI46 [1] speech recognition corpus as a benchmark, we show that the proposed FPGA-based neural processor system is up to 29% more energy efficient than a baseline LSM processor with little extra hardware overhead.
Yu Liu 0028, Yingyezhe Jin, Peng Li 0001
ISLPED3
2017 Performance and robustness of bio-inspired digital liquid state machines: A case study of speech recognition
Yingyezhe Jin, Peng Li 0001
Neurocomputing2
2017 Energy efficient parallel neuromorphic architectures with approximate arithmetic on FPGA
Qian Wang 0003, Youjie Li, Botang Shao, Siddhartha Dey, Peng Li 0001
Neurocomputing5
2017 Multiharmonic Small-Signal Modeling of Low-Power PWM DC-DC Converters
abstract
Small-signal models of pulse-width modulation (PWM) converters are widely used for analyzing stability and play an important role in converter design and control. However, existing small-signal models either are based on averaged DC behaviors, and hence are unable to capture frequency responses that are faster than the switching frequency, or greatly approximate these high-frequency responses. We address the severe limitations of the existing models by proposing a multiharmonic model that provides a complete small-signal characterization of both DC averages and high-order harmonic responses. The proposed model captures important high-frequency overshoots and undershoots of the converter response, which are otherwise unaccounted for by the existing techniques. In two converter examples, the proposed model corrects the misleading results of the existing models by providing truthful characterization of the overall converter AC response and offers important guidance for converter design and closed-loop control.
Dani A. Tannir, G. Peter Fang, Wei Dong 0002, Peng Li 0001
ACM Trans. Design Autom. Electr. Syst.7
2016 Relevance vector and feature machine for statistical analog circuit characterization and built-in self-test optimization
abstract
Aiding design and test optimization of analog circuits requires accurate models that can reliably capture complex dependencies of circuit performances on essential circuit and device parameters, and test signatures. We present a novel Bayesian learning technique, namely relevance vector and feature machine (RVFM), for characterizing analog circuits with sparse statistical regression models. RVFM not only produces accurate models learned from a moderate amount of simulation or measurement data, but also computes a probabilistically inferred weighting factor quantifying the criticality of each parameter as part of the overall learning framework, hence offering a powerful enabler for variability modeling, failure diagnosis, and test development. Compared to other popular learning-based techniques, the proposed RVFM produces more accurate models, requires less amount of training data, and extracts more reliable parametric ranking. The effectiveness of RVFM is demonstrated in terms of the statistical variability modeling of a low-dropout regulator (LDO) and the built-in self-test (BIST) development of a charge-pump phase-locked loop (PLL).
Honghuang Lin, Peng Li 0001
DAC2
2016 Distributed on-chip regulation: theoretical stability foundation, over-design reduction and performance optimization
abstract
While distributed on-chip voltage regulation offers an appealing solution to power delivery, designing power delivery networks (PDNs) with distributed on-chip voltage regulators with guaranteed stability is challenging because of the complex interactions between active regulators and the bulky passive network. The recently developed hybrid stability theory provides an efficient stability checking and design approach, giving rise to highly desirable localized design of PDNs. However, the inherent conservativeness of the hybrid stability criteria can lead to pessimism in stability evaluation and hence large over-design. We address this challenge by proposing an optimal frequency-dependent system partitioning technique to significantly reduce the amount of pessimism in stability analysis. With theoretical rigor, we show how to partition a PDN system by employing optimal frequency-dependent impedance splitting between the passive network and voltage regulators while maintaining the desired theoretical properties of the partitioned system blocks upon which the hybrid stability principle is anchored. We demonstrate a new stability-ensuring PDN design approach with the proposed over-design reduction technique using an automated optimization flow which significantly boosts regulation performance and power efficiency.
Xin Zhan, Peng Li 0001, Edgar Sánchez-Sinencio
DAC2
2016 Multi-harmonic nonlinear modeling of low-power PWM DC-DC converters operating in CCM and DCM
Dani A. Tannir, Peng Li 0001
DATE4
2016 D-LSM: Deep Liquid State Machine with unsupervised recurrent reservoir tuning
abstract
The Liquid State Machine (LSM) is a biologically plausible model of computation for recurrent spiking neural networks, which offers promising solutions to real-world applications in both software and hardware based systems. At the same time, deep feedforward rate-based neural networks such as convolutional neural networks (CNNs) have achieved great success in many computer vision related applications. However, a systematic exploration of deep recurrent spiking neural networks is lacking. We propose a new model of Deep Liquid State Machine (D-LSM), which simultaneously explores the powers of recurrent spiking networks and deep architectures. D-LSM consists of multiple basic LSM processing and pooling stages. Recurrent reservoir networks across different LSM stages act as nonlinear filters capable of extracting spatio-temporal features of increasingly higher levels from the input. We propose to train the D-LSM practically by adopting unsupervised training (e.g. through STDP) for recurrent reservoirs and spike-based supervised rules for the final readout stage. The perspective of realizing D-LSM based hardware processors is also presented.
Qian Wang 0003, Peng Li 0001
ICPR2
2016 AP-STDP: A novel self-organizing mechanism for efficient reservoir computing
abstract
The Liquid State Machine (LSM) exploits the computation capability of recurrent spiking neural networks by incorporating a randomly generated reservoir, which is often fixed. This standard choice relaxes the challenging need for training the complex recurrent reservoir. The fixed reservoir is used as a generic kernel to map the temporal input signals to the internal network dynamics, and a readout layer is trained to extract the information embedded in the network dynamics to facilitate pattern classification. However, the question of how to effectively tune the reservoir for given computational tasks remains to be answered. In this paper, we propose a novel Activity-based Probabilistic Spiking-Timing Dependent Plastic (AP-STDP) mechanism for self-organizing reservoirs. Compared to conventional STDP mechanisms, the proposed rule improves tuning efficiency, prevents the saturation of synaptic memory, and boosts performance. We assess the internal representation ability of the proposed self-organizing mechanism via principal component analysis (PCA) and show that the proposed method is advantageous over other STDP algorithms. Using the spoken English letters adopted from the TI46 speech corpus for performance benchmarking, we demonstrate that AP-STDP consistently outperforms other STDP mechanisms regardless of reservoir size, and is able to boost the performance of the isolated spoken English letter recognition by 2.7% with a small reservoir size.
Yingyezhe Jin, Peng Li 0001
IJCNN2
2016 Liquid state machine based pattern recognition on FPGA with firing-activity dependent power gating and approximate computing
abstract
This paper presents an FPGA architecture and implementation of the Liquid State Machine, a spiking neural network model, for real world pattern recognition problems. The proposed architecture consists of a parallel digital reservoir with fixed synapses, and a readout stage that is tuned by a biologically plausible supervised learning rule. When evaluated using the TI46 speech corpus, a widely adopted speech recognition benchmark, the presented FPGA neuromorphic processors demonstrate highly competitive recognition performance and provide a runtime speedup of 88X over the 2.3 GHz AMD OpteronTM Processor. A number of critical design issues such as interconnection of liquid neurons, storage of synaptic weights and design of arithmetic blocks are addressed in this work. More importantly, it is shown that the unique computational structure and inherent resilience of the liquid state machine can be leveraged for highly efficient FPGA implementation. For t Iiis, it is demonstrated that the proposed firing-activity based power gating and approximate arithmetic computing with runtime adjustable precision can lead to up to 30.2% reduction in power and energy dissipation without greatly impacting speech recognition performance.
Qian Wang 0003, Youjie Li, Peng Li 0001
ISCAS3
2016 Neuromorphic Processors with Memristive Synapses: Synaptic Interface and Architectural Exploration
abstract
Due to their nonvolatile nature, excellent scalability, and high density, memristive nanodevices provide a promising solution for low-cost on-chip storage. Integrating memristor-based synaptic crossbars into digital neuromorphic processors (DNPs) may facilitate efficient realization of brain-inspired computing. This article investigates architectural design exploration of DNPs with memristive synapses by proposing two synapse readout schemes. The key design tradeoffs involving different analog-to-digital conversions and memory accessing styles are thoroughly investigated. A novel storage strategy optimized for feedforward neural networks is proposed in this work, which greatly reduces the energy and area cost of the memristor array and its peripherals.
Qian Wang 0003, Yongtae Kim 0001, Peng Li 0001
ACM J. Emerg. Technol. Comput. Syst.3
2016 Robust and Efficient Transistor-Level Envelope-Following Analysis of PWM/PFM/PSM DC-DC Converters
abstract
The envelope-following (EF) simulation of practical dc-dc converters is challenging due to the presence of digital behavior, strong nonlinearity, complex frequency module schemes, and feedback loops. This paper presents a novel EF method for time-domain analysis of dc-dc converters-based upon a numerically robust time-delayed phase condition to track the envelopes of circuit states under a varying switching frequency. We further develop an EF technique that is applicable to both fixed and varying switching frequency operations, thereby providing a unifying solution to converters with pulse width modulation, pulse frequency modulation, and pulse skipping modulation. By adopting three fast simulation techniques, our proposed EF method achieves higher speedup without composing the accuracy of the results. The robustness and efficiency of the proposed method are demonstrated using several dc-dc converter and oscillator circuits modeled using the industrial standard BSIM4 transistor models. A significant runtime speedup of up to 30X with respect to the conventional transient analysis is achieved for several dc-dc converters with strong nonlinear switching characteristics.
Peng Li 0001, Suming Lai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Accurate Modeling of Nonideal Low-Power PWM DC-DC Converters Operating in CCM and DCM using Enhanced Circuit-Averaging Techniques
abstract
The development of enhanced modeling techniques for the simulation of switched-mode Pulse Width Modulated (PWM) DC-DC power converters using circuit averaging is the main focus of this article. The circuit-averaging technique has traditionally been used to model the behavior of PWM DC-DC converters without considering important nonideal characteristics of the switching devices. As a result, most of these existing approaches present simplified models that are ideal or linearized, and do not accurately account for the performance characteristics of the converter. This is especially problematic for low-power applications. In this article, we present an enhanced nonideal behavioral circuit-averaged model that makes the simulation of DC-DC converters both computationally efficient and accurate, thereby presenting an important tool for circuit designers. Experimentally, we show that our Verilog-A-based new model allows for accurate simulation of both Buck- and Boost-type PWM converters operating in either CCM or DCM modes while providing more than one order of magnitude speedup over the transistor-level simulation.
Dani A. Tannir, Peng Li 0001
ACM Trans. Design Autom. Electr. Syst.3
2015 Leveraging emerging nonvolatile memory in high-level synthesis with loop transformations
abstract
To mitigate the “Power Wall” challenges for both mobile devices and data centers, accelerator-rich architecture with normally-off mode has been intensively studied recently. Power/energy optimization in high-level synthesis for accelerator design is critical for such accelerator-rich architecture. The emerging nonvolatile memory (NVM), offers many benefits such as ultra-low leakage power, high density, and instant power-on/off, and therefore is a promising alternative for the hardware accelerator design to achieve further power reduction. However, such NVM suffers from large write energy and latency, which brings new challenges for the buffer allocation in the custom accelerator design. This paper presents the first framework that optimizes NVM allocation in high-level synthesis for custom accelerator design, considering loop transformations. It solves the loop transformation, buffer allocation, and buffer type selection to minimize the memory power consumption, while under area, bandwidth, and performance constraints. This paper formulates the optimization problem, and solves it with a problem-specific designed stimulated annealing solution. Experiments demonstrate 32% extra power reduction compared with the previous method without optimizing loop transformations.
Shuangchen Li, Ang Li 0005, Yuan Zhe, Yongpan Liu, Peng Li 0001, Guangyu Sun 0003, Yu Wang 0002, Huazhong Yang, Yuan Xie 0001
ISLPED5
2015 A Reconfigurable Digital Neuromorphic Processor with Memristive Synaptic Crossbar for Cognitive Computing
abstract
This article presents a brain-inspired reconfigurable digital neuromorphic processor (DNP) architecture for large-scale spiking neural networks. The proposed architecture integrates an arbitrary number of N digital leaky integrate-and-fire (LIF) silicon neurons to mimic their biological counterparts and on-chip learning circuits to realize spike-timing-dependent plasticity (STDP) learning rules. We leverage memristor nanodevices to build an N × N crossbar array to store not only multibit synaptic weight values but also network configuration data with significantly reduced area overhead. Additionally, the crossbar array is designed to be accessible both column- and row-wise to expedite the synaptic weight update process for learning. The proposed digital pulse width modulator (PWM) produces binary pulses with various durations for reading and writing the multilevel memristive crossbar. The proposed column based analog-to-digital conversion (ADC) scheme efficiently accumulates the presynaptic weights of each neuron and reduces silicon area overhead by using a shared arithmetic unit to process the LIF operations of all N neurons. With 256 silicon neurons, learning circuits and 64K synapses, the power dissipation and area of our DNP are 6.45 mW and 1.86 mm 2 , respectively, when implemented in a 90-nm CMOS technology. The functionality of the proposed DNP architecture is demonstrated by realizing an unsupervised-learning based character recognition system.
Yongtae Kim 0001, Yong Zhang 0049, Peng Li 0001
ACM J. Emerg. Technol. Comput. Syst.3
2015 Circuit design and exponential stabilization of memristive neural networks
Shiping Wen 0001, Tingwen Huang, Zhigang Zeng, Yiran Chen 0001, Peng Li 0001
Neural Networks5
2015 Circuit Performance Classification With Active Learning Guided Sampling for Support Vector Machines
abstract
Leveraging machine learning has been proven as a promising avenue for addressing many practical circuit design and verification challenges. We demonstrate a novel active learning guided machine learning approach for characterizing circuit performance. When employed under the context of support vector machines (SVMs), the proposed probabilistically weighted active learning approach is able to dramatically reduce the size of the training data, leading to significant reduction of the overall training cost. The proposed active learning approach is extended to the training of asymmetric SVM classifiers, which is further sped up by a global acceleration scheme. We demonstrate the excellent performance of the proposed techniques using four case studies: 1) dc/dc converter ripple noise analysis; 2) phase-locked loop lock-time verification; 3) reliability analysis of a ring oscillator with respect to process variations and initial conditions; and 4) prediction of chip peak temperature using a limited number of on-chip temperature sensors.
Honghuang Lin, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 A Digital Liquid State Machine With Biologically Inspired Learning and Its Application to Speech Recognition
abstract
This paper presents a bioinspired digital liquid-state machine (LSM) for low-power very-large-scale-integration (VLSI)-based machine learning applications. To the best of the authors' knowledge, this is the first work that employs a bioinspired spike-based learning algorithm for the LSM. With the proposed online learning, the LSM extracts information from input patterns on the fly without needing intermediate data storage as required in offline learning methods such as ridge regression. The proposed learning rule is local such that each synaptic weight update is based only upon the firing activities of the corresponding presynaptic and postsynaptic neurons without incurring global communications across the neural network. Compared with the backpropagation-based learning, the locality of computation in the proposed approach lends itself to efficient parallel VLSI implementation. We use subsets of the TI46 speech corpus to benchmark the bioinspired digital LSM. To reduce the complexity of the spiking neural network model without performance degradation for speech recognition, we study the impacts of synaptic models on the fading memory of the reservoir and hence the network performance. Moreover, we examine the tradeoffs between synaptic weight resolution, reservoir size, and recognition performance and present techniques to further reduce the overhead of hardware implementation. Our simulation results show that in terms of isolated word recognition evaluated using the TI46 speech corpus, the proposed digital LSM rivals the state-of-the-art hidden Markov-model-based recognizer Sphinx-4 and outperforms all other reported recognizers including the ones that are based upon the LSM or neural networks.
Yong Zhang 0049, Peng Li 0001, Yingyezhe Jin, Yoonsuck Choe
IEEE Trans. Neural Networks Learn. Syst.2
2015 Decoupling Capacitance Design Strategies for Power Delivery Networks with Power Gating
abstract
Power gating is a widely used leakage power saving strategy in modern chip designs. However, power gating introduces unique power integrity issues and trade-offs between switching and rush current (wake-up) supply noises. At the same time, the amount of power saving intrinsically trades off with power integrity. In addition, these trade-offs significantly vary with supply voltage. In this article, we propose systemic decoupling capacitors (decaps) optimization strategies that optimally trade-off between power integrity and leakage saving. Specially, new global decap and reroutable decap design concepts are proposed to relax the tight interaction between power integrity and leakage saving of power gated PDNs with a single supply voltage level. Furthermore, we propose a flexible decap allocation technique to deal with the design trade-offs under multiple supply voltage levels. The proposed strategies are implemented in an automatic design flow for choosing the optimal amount of local decaps, global decaps and reroutable decaps. The conducted experiments demonstrate that leakage saving can be increased significantly compared with the conventional PDN design approach with a single supply voltage level using the proposed techniques without jeopardizing power integrity. For PDN designs operating at two supply voltage levels, the optimal performance is achieved at each voltage level.
Tong Xu 0004, Peng Li 0001, Savithri Sundareswaran
ACM Trans. Design Autom. Electr. Syst.2
2015 Energy Efficient Approximate Arithmetic for Error Resilient Neuromorphic Computing
abstract
This brief proposes a novel design scheme for approximate adders and comparators to significantly reduce energy consumption while maintaining a very low error rate. The considerably improved error rate and critical path delay stem from the employed carry prediction technique that leverages the information from less significant input bits in a parallel manner. The proposed designs have been adopted in a VLSI-based neuromorphic character recognition chip with unsupervised learning implemented on chip. The approximation errors of the proposed arithmetic units have been shown to have negligible impact on the training process while archiving good energy efficiency.
Yongtae Kim 0001, Yong Zhang 0049, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2015 A Parallel Digital VLSI Architecture for Integrated Support Vector Machine Training and Classification
abstract
This paper presents a parallel digital VLSI architecture for combined support vector machine (SVM) training and classification. For the first time, cascade SVM, a powerful training algorithm, is leveraged to significantly improve the scalability of hardware-based SVM training and develop an efficient parallel VLSI architecture. The presented architecture achieves excellent scalability by spreading the training workload of a given data set over multiple SVM processing units with minimal communication overhead. Hardware-friendly implementation of the cascade algorithm is employed to achieve low hardware overhead and allow for training over data sets of variable size. In the proposed parallel cascade architecture, a multilayer system bus and multiple distributed memories are used to fully exploit parallelism. In addition, the proposed architecture is rather flexible and can be tailored to realize hybrid use of hardware parallel processing and temporal reuse of processing resources, leading to good tradeoffs between throughput, silicon overhead and power dissipation. Several parallel cascade SVM processors have been designed with a commercial 90-nm CMOS technology, which provide up to a 561× training time speedup and a significant estimated 21 859× energy reduction compared with the software SVM algorithm running on a 45-nm commercial general-purpose CPU.
Qian Wang 0003, Peng Li 0001, Yongtae Kim 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Parallel Hierarchical Reachability Analysis for Analog Verification
abstract
Formal methods such as reachability analysis often suffer from state space explosion in the verification of complex analog circuits. This paper proposes a parallel hierarchical SMT-based reachability analysis technique based on circuit decomposition. Circuits are systematically decomposed into subsystems with less complex transient behaviors which can be solved in parallel. Then a simulation-assisted SMT-based reachability analysis approach is adopted to conservatively approximate the reachable spaces in each subsystem with support function representations. We formally develop a general decomposition algorithm without overapproxmiation in system reconstruction. The efficiency of this general methodology is further optimized with an efficient parallel implementation strategy.
Honghuang Lin, Peng Li 0001
DAC2
2014 Approximate property checking of mixed-signal circuits
abstract
Growing circuit complexity and design uncertainty has made it difficult to predict whether large circuits meet target property specifications. To address this, we conservatively approximate the failure probability estimate by defining an interval that bounds this probability. Doing so using an arbitrary sampling distribution requires a learner. Given that the learner's knowledge is imperfect, the interval must first capture its uncertainty. An ensemble of such learners can then be used to compensate for the bias. Lastly, we develop an adaptive sampling scheme to tighten the obtained interval with increased simulation resources, thus controlling the accuracy vs. turn-around-time trade-off.
Parijat Mukherjee, Chirayu S. Amin, Peng Li 0001
DAC3
2014 Leveraging pre-silicon data to diagnose out-of-specification failures in mixed-signal circuits
abstract
Diagnosing out-of-specification failures in mixed-signal circuits has become increasingly challenging due to: (1) failures caused by interactions between input-signal conditions and design uncertainties, and (2) the need to identify critical input and uncertainty conditions that cause these regions. We propose a simulation-driven approach that first uses ensemble learning to extract if -- then rules that naturally solve both problems. By ranking, pruning and clustering these rules, we then construct non-linear failure regions which can be directly employed for pre-silicon debug, as demonstrated on a phase-locked loop circuit. Furthermore, these regions can be used to guide test pattern generation and/or assist with post-silicon debug.
Parijat Mukherjee, Peng Li 0001
DAC2
2014 A unifying and robust method for efficient envelope-following simulation of PWM/PFM DC-DC converters
abstract
The envelope-following (EF) simulation of practical DC-DC converters is challenging due to the presence of digital behavior, strong nonlinearity, complex frequency module schemes and feedback loops. This paper presents a novel EF method for time-domain analysis of DC-DC converters based upon a numerically robust time-delayed phase condition to track the envelopes oaf circuit states under a varying switching frequency. We further develop an EF technique that is applicable to both fixed and varying switching frequency operations, thereby providing a unifying solution to converters with pulse-width modulation (PWM) and/or pulse-frequency modulation (PFM). The robustness and efficiency of the proposed method are demonstrated using several DC-DC converter and oscillator circuits modeled using the industrial standard BSIM4 transistor models. A significant runtime speedup of 30× with respect to the conventional transient analysis is achieved for PFM DC-DC converters with strong nonlinear switching characteristics.
Peng Li 0001, Suming Lai
ICCAD2
2014 A model for array-based approximate arithmetic computing with application to multiplier and squarer design
abstract
We propose a general model for array-based approximate arithmetic computing to trade off accuracy for significant reduction in energy consumption, which is realized by identifying input signatures for efficient compensation of approximation errors. Under this model, our approximate 16x16 bits fixed-width Booth multiplier consumes 44.96% and 28.33% less energy and area compared with the most accurate fixed-width Booth multiplier. Furthermore, it reduces average error, max error and mean square by 10.46%, 30.77% and 21.26%, respectively, when compared with the best reported approximate design. Using the same approach, significant energy consumption, area and error reduction is achieved for a squarer unit.
Botang Shao, Peng Li 0001
ISLPED2
2014 Understanding SRAM Stability via Bifurcation Analysis: Analytical Models and Scaling Trends
abstract
In the past decades, aggressive scaling of transistor feature size has been a primary force driving higher Static Random Access Memory (SRAM) integration density. Due to technology scaling, nanometer SRAM designs become increasingly vulnerable to stability challenges. The traditional way of analyzing stability is through the use of Static Noise Margins (SNMs). SNMs are not capable of capturing the key nonlinear dynamics associated with memory operations, leading to imprecise characterization of stability. This work rigorously develops dynamic stability concepts and, more importantly, captures them in physically based analytical models. By leveraging nonlinear stability theory, we develop analytical models that characterize the minimum required amplitude and duration of injected current noises that can flip the SRAM state. These models, which are parameterized in key design, technology, and operating condition parameters, provide important design insights and offer a basis for predicting scaling trends of SRAM dynamic stability.
Yenpo Ho, Garng M. Huang, Peng Li 0001
ACM Trans. Design Autom. Electr. Syst.3
2013 Verification of digitally-intensive analog circuits via kernel ridge regression and hybrid reachability analysis
abstract
The emergence of digitally-intensive analog circuits introduces new challenges to formal verification due to increased digital design content, and non-ideal digital effects such as finite resolution, round-off error and overflow. We propose a machine learning approach to convert digital blocks to conservative analog approximations via the use of kernel ridge regression. These learned models are then adopted in a hybrid formal reachability analysis framework where the support function based manipulations are developed to efficiently handle the large linear portion of the design and the more general satisfiability modulo theories technique is applied to the remaining nonlinear portion. The efficiency of the proposed method is demonstrated for the locked time verification of a digitally intensive phase locked loop.
Honghuang Lin, Peng Li 0001, Chris J. Myers
DAC2
2013 An energy efficient approximate adder with carry skip for error resilient neuromorphic VLSI systems
abstract
We propose a novel approximate adder design to significantly reduce energy consumption with a very moderate error rate. The significantly improved error rate and critical path delay stem from the employed carry prediction technique that leverages the information from less significant input bits in a parallel manner. An error magnitude reduction scheme is proposed to further reduce amount of error once detected with low cost. Implemented in a commercial 90 nm CMOS process, it is shown that the proposed adder is up to 2.4× faster and 43% more energy efficient over traditional adders while having an error rate of only 0.18%. The proposed adder has been adopted in a VLSI-based neuromorphic character recognition chip using unsupervised learning. The approximation errors of the proposed adder have been shown to have negligible impact on the training process. Moreover, the energy savings of up to 48.5% over traditional adders is achieved for the neuromorphic circuit with scaled supply level. Finally, we achieve error-free operations by including a low-overhead error correction logic.
Yongtae Kim 0001, Yong Zhang 0049, Peng Li 0001
ICCAD3
2013 Localized Stability Checking and Design of IC Power Delivery With Distributed Voltage Regulators
abstract
Placing multiple voltage regulators onto the die is an effective way of enabling distributed on-chip voltage regulation and provides significant benefits in suppressing various types of power supply noise. However, the complex interactions between the active voltage regulators and the large passive subnetwork may render the complete power delivery network (PDN) unstable, leading to design failures. While traditional stability measures such as phase margin are not applicable to regulated PDNs that have a large number of loops, a brute-force analysis of network stability can be impractical due to the high complexity of a given PDN. We present a hybrid stability margin concept and the associated stability-checking method for PDNs with integrated linear low-dropout voltage regulators (LDOs). With theoretical rigor, the proposed approach is local in the sense that the stability of the entire network can be efficiently examined through a hybrid stability constraint that is defined locally for individual LDOs. In the same spirit, we propose a localized LDO design methodology that optimizes individual LDOs in a stand-alone manner while ensuring the network-level stability. Key circuit-level design considerations and tradeoffs involved in stability-ensuring LDO design are also discussed.
Suming Lai, Boyuan Yan, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Simulation-Assisted Formal Verification of Nonlinear Mixed-Signal Circuits With Bayesian Inference Guidance
abstract
The pressing need for the verification of analog and mixed-signal (AMS) designs is driven by increased design complexity and the integration of such circuits into SoCs. However, verification of AMS circuits remains a significant challenge. This paper proposes a simulation-assisted formal verification methodology that leverages SMT-based satisfiability techniques to tackle the challenges arising from the inherent analog and/or hybrid nature of AMS systems. Although state-of-the-art SMT solvers, in the worst-case scenario, still have exponential complexity in the number of constraints, the main focus of this paper is to first formally formulate the verification task into an SMT problem, then accelerate the verification by using simulation assistance. To verify the nonlinear dynamics, randomly sampled simulations are first applied to quickly explore the reachable state space, and then a nonlinear SMT solver is invoked to ensure the conservativeness. To achieve optimal efficiency, the tradeoff between the runtime costs of simulation and SMT solving is analyzed by means of a Bayesian inference-based technique that dynamically learns from the simulation history. This paper demonstrates the feasibility and efficacy of the proposed methodology on conservative verification of dynamic properties of nonlinear AMS circuits.
Leyi Yin, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 IC power delivery: Voltage regulation and conversion, system-level cooptimization and technology implications
abstract
Modern IC power delivery systems encompass large on-chip passive power grids and active on-chip or off-chip voltage converters and regulators. While there exists little work targeting on holistic design of such complex IC subsystems, the optimal system-level design of power delivery is critical for achieving power integrity and power efficiency. In this article, we conduct a systematic design analysis on power delivery networks that incorporate Buck Converters (BCs) and on-chip Low-Dropout voltage regulators (LDOs) for the entire chip power supply. The electrical interactions between active voltage converters, regulators as well as passive power grids and their influence on key system design specifications are analyzed comprehensively. With the derived design insights, the system-level codesign of a complete power delivery network is facilitated by a proposed automatic optimization flow in which key design parameters of buck converters and on-chip LDOs as well as on-chip decoupling capacitance are jointly optimized. The experimental results demonstrate significant performance improvements resulted from the proposed system cooptimization in terms of achievable area overhead, supply noise and power efficiency. Impacts of different decoupling capacitance technologies are also investigated.
Zhiyu Zeng, Suming Lai, Peng Li 0001
ACM Trans. Design Autom. Electr. Syst.3
2013 Fast Thermal Analysis on GPU for 3D ICs With Integrated Microchannel Cooling
abstract
Effective thermal management for 3D integrated circuits (3D ICs) is becoming increasingly challenging due to the ever-increasing power density and chip design complexity; traditional heat sinks are expected to quickly reach their limits for meeting the cooling needs of 3D ICs. Alternatively, the integrated liquid-cooled microchannel heat sink has become one of the most effective solutions. In this paper, we present fast multigrid and block tridiagonally preconditioned graphics processing unit (GPU) based thermal simulation algorithms for 3D ICs. Unlike the CPU-based solver development in which existing sophisticated numerical simulation tools (matrix solvers) can be readily adopted and implemented, GPU-based thermal simulation demands more effort in the algorithm and data structure design phase, and requires careful consideration of GPU's thread/memory organization, data access/communication patterns, arithmetic intensity, as well as its hardware occupancies. As shown by various experimental results, our GPU-based 3D thermal simulation solvers can achieve more than 360× speedups over the best available direct solvers and more than 35× speedups over the CPU-based iterative solvers, without loss of accuracy.
Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Stability assurance and design optimization of large power delivery networks with multiple on-chip voltage regulators
abstract
Distributive on-chip voltage regulation is appealing to solving the power integrity problems in nowadays high-end SoCs. Nevertheless, ensuring the stability of large-scale power delivery networks regulated by a multiplicity of voltage regulators is challenging due to the size of the system and complex interactions between the regulators and the large loading network. We present a theoretically elegant framework that provides a rigorous guarantee for network stability. We further develop a practical design approach that largely decouples the design of linear voltage regulators from that of the complex load, making it feasible to ensure the stability of the complete network. The presented design approach has been successfully applied to several design examples with guaranteed stability and competitive performances.
Suming Lai, Boyuan Yan, Peng Li 0001
ICCAD3
2012 Design analysis of IC power delivery
abstract
Power delivery design plays a critical role in ensuring power delivery integrity and achieving overall design power efficiency. A power delivery network (PDN) consists of a multiplicity of passive and active components, which interact with each other in a complex manner. Under this context, power delivery design is a multifaceted problem and a holistic system optimization involving joint design of the passive distribution sub-network and active voltage converters and regulators is indispensable. This paper provides a succinct review of the related PDN analysis and design problems and motivates the need for system-level co-optimization of key design parameters in order to optimally trade off between power supply noise, power efficiency, area overhead and stability.
Peng Li 0001
ICCAD1
2012 Classifying circuit performance using active-learning guided support vector machines
abstract
Leveraging machine learning has been proven as a promising avenue for addressing many practical circuit design and verification challenges. We demonstrate a novel active learning guided machine learning approach for characterizing circuit performance. When employed under the context of support vector machines, the proposed probabilistically weighted active learning approach is able to dramatically reduce the size of the training data, leading to significant reduction of the overall training cost. The proposed active learning approach is extended to the training of asymmetric support vector machine classifiers, which is further sped up by a global acceleration scheme. We demonstrate the excellent performance of the proposed techniques using three case studies: PLL lock-time verification, SRAM yield analysis and prediction of chip peak temperature using a limited number of on-chip temperature sensors.
Honghuang Lin, Peng Li 0001
ICCAD2
2012 Verifying dynamic properties of nonlinear mixed-signal circuits via efficient SMT-based techniques
abstract
The pressing need for the verification of analog and mixed-signal (AMS) designs is driven by increased design complexity and the integration of such circuits into SoCs. However, verification of AMS circuits remains as a significant challenge. We propose a methodology that leverages SMT-based Satisfiability techniques to tackle the challenges arising from the inherent analog and/or hybrid natures of AMS systems. We demonstrate the feasibility and efficacy of the proposed methodology on conservative verification of dynamic properties of nonlinear AMS circuits.
Leyi Yin, Peng Li 0001
ICCAD3
2012 Load-aware stochastic feedback control for DVFS with tight performance guarantee
Peng Li 0001
VLSI-SoC2
2012 Linking brain behavior to underlying cellular mechanisms via large-scale brain modeling and simulation
Yong Zhang 0049, Boyuan Yan, Mingchao Wang, Jingzhen Hu, Haokai Lu, Peng Li 0001
Neurocomputing6
2012 Efficient Identification of Unstable Loops in Large Linear Analog Integrated Circuits
abstract
Stability analysis is one of the key challenges in analog circuit design. As feature sizes continue to shrink and the effect of parasitics becomes more dominant, we are forced to deal with stability analysis of increasingly complex multiloop structures with potentially hundreds of loops-a task that can no longer be dealt with using traditional methods. An automated stability checker tool that detects sources of potential ringing behavior within a reasonable turnaround time has thus been made necessary. Such a tool would not just help in debug but could also serve as a postlayout validation tool. We thus present an efficient loop finder algorithm to identify sources of ringing in large linear analog circuits. At the heart of our automated stability checker are two newly developed computationally efficient algorithms-the first to detect all poles within a given region of interest with a high degree of confidence and the second to extract second-order approximations of node impedance transfer functions given these pole locations. In this paper, we discuss these algorithms in detail, propose various optimization heuristics to further speed up the pole discovery algorithm, and then go on to develop a parallel implementation of both these underlying algorithms. It is demonstrated that these approaches together allow us to outperform the original loop finder algorithm based on direct eigen methods by two to four orders of magnitude and thus enable stability analysis of even larger extracted industrial designs than was previously possible while providing reasonable turnaround time.
Parijat Mukherjee, G. Peter Fang, Rod Burt, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2012 Quantifying Dynamic Stability of Genetic Memory Circuits
abstract
Bistability/Multistability has been found in many biological systems including genetic memory circuits. Proper characterization of system stability helps to understand biological functions and has potential applications in fields such as synthetic biology. Existing methods of analyzing bistability are either qualitative or in a static way. Assuming the circuit is in a steady state, the latter can only reveal the susceptibility of the stability to injected DC noises. However, this can be inappropriate and inadequate as dynamics are crucial for many biological networks. In this paper, we quantitatively characterize the dynamic stability of a genetic conditional memory circuit by developing new dynamic noise margin (DNM) concepts and associated algorithms based on system theory. Taking into account the duration of the noisy perturbation, the DNMs are more general cases of their static counterparts. Using our techniques, we analyze the noise immunity of the memory circuit and derive insights on dynamic hold and write operations. Considering cell-to-cell variations, our parametric analysis reveals that the dynamic stability of the memory circuit has significantly varying sensitivities to underlying biochemical reactions attributable to differences in structure, time scales, and nonlinear interactions between reactions. With proper extensions, our techniques are broadly applicable to other multistable biological systems.
Yong Zhang 0049, Peng Li 0001, Garng M. Huang
IEEE ACM Trans. Comput. Biol. Bioinform.2
2011 Automatic stability checking for large linear analog integrated circuits
abstract
Stability analysis is one of the key challenges in the design of large linear analog circuits with complex multi-loop structures. In this paper, we present an efficient loop finder algorithm to identify potentially unstable loops in such circuits. At the heart of our automated stability checker lie two newly developed computationally efficient algorithms --- the first to detect all poles within a given region of interest and the second to extract second order approximations of node impedance transfer functions given these pole locations. It is shown that the proposed technique outperforms existing stability methods by more than one order of magnitude for medium sized circuits and enables stability analysis of large extracted industrial designs which was previously infeasible.
Parijat Mukherjee, G. Peter Fang, Rod Burt, Peng Li 0001
DAC4
2011 Decoupling for power gating: sources of power noise and design strategies
abstract
Power gating is essential for controlling leakage power dissipation of modern chip designs. However, power gating introduces unique power delivery integrity issues and tradeoffs between switching and rush current (wake-up) supply noises. In addition, in power-gated power delivery networks (PDNs), the amount of power saving intrinsically trades off with power integrity. In this paper, we propose systemic decoupling capacitance optimization strategies that optimally balance between switching and rush current noises, and tradeoff between power integrity and wake-up time, hence power saving. Furthermore, we propose a novel re-routable decoupling capacitance concept to break the tight interaction between power integrity and power saving, providing further improved tradeoffs between the two. Our design strategies have been implemented in a simulation-based optimization flow and the conducted experimental results have demonstrated significant improvement on leakage power saving through the presented techniques.
Tong Xu 0004, Peng Li 0001, Boyuan Yan
DAC2
2011 High effective-resolution built-in jitter characterization with quantization noise shaping
abstract
A novel built-in jitter characterization architecture combining quantization noise shaping and a partial Vernier delay structure is proposed for high resolution jitter measurement. The effective resolution is optimized at the system level as well as the circuit level. Using 90nm CMOS technology, an area of 0.008mm2 is occupied. The power consumption is 1.85mW. An effective resolution of 1.5ps is achieved.
Leyi Yin, Yongtae Kim 0001, Peng Li 0001
DAC3
2011 Fast static analysis of power grids: Algorithms and implementations
abstract
Large VLSI on-chip power delivery networks (PDN) are challenging to analyze due to sheer network complexity. In this paper, three power grid solvers developed in our group: a direct solver using Cholesky decomposition, a GPU-based multigrid preconditioning solver, and a partitioning-based solver using spatial locality, are reviewed. Following the requirements of TAU 2011 Power Grid Simulation Contest, single-threaded versions of these solvers are implemented and their performances are evaluated in terms of runtime, memory, maximum error and average error. The experimental results show that for the published IBM power grid benchmarks, the direct solver has the best overall performance among the three.
Zhiyu Zeng, Tong Xu 0004, Peng Li 0001
ICCAD4
2011 Simulation of large neuronal networks with biophysically accurate models on graphics processors
abstract
Efficient simulation of large-scale mammalian brain models provides a crucial computational means for understanding complex brain functions and neuronal dynamics. However, such tasks are hindered by significant computational complexities. In this work, we attempt to address the significant computational challenge in simulating large-scale neural networks based on biophysically plausible Hodgkin-Huxley (HH) neuron models. Unlike simpler phenomenological spiking models, the use of HH models allows one to directly associate the observed network dynamics with the underlying biological and physiological causes, but at a significantly higher computational cost. We exploit recent commodity massively parallel graphics processors (GPUs) to alleviate the significant computational cost in HH model based neural network simulation. We develop look-up table based HH model evaluation and efficient parallel implementation strategies geared towards higher arithmetic intensity and minimum thread divergence. Furthermore, we adopt and develop advanced multi-level numerical integration techniques well suited for intricate dynamical and stability characteristics of HH models. On a commodity GPU card with 240 streaming processors, for a neural network with one million neurons and 200 million synaptic connections, the presented GPU neural network simulator is about 600X faster than a basic serial CPU based simulator, 28X faster than the CPU implementation of the proposed techniques, and only two to three times slower than the GPU based simulation using simpler phenomenological spiking models.
Mingchao Wang, Boyuan Yan, Jingzhen Hu, Peng Li 0001
IJCNN4
2011 Hierarchical Multialgorithm Parallel Circuit Simulation
abstract
The emergence of multicore and many-core processors has introduced new opportunities and challenges to electronic design automation research and development. While the availability of increasing parallel computing power holds new promise to address many challenges in computer-aided design (CAD), the leverage of hardware parallelism can only be possible with a new generation of parallel CAD applications. In this paper, we propose a novel hierarchical multialgorithm (MA) parallel circuit simulation approach and its multicore implementation to expedite one of the most fundamental CAD applications: transistor-level transient circuit simulation. In our parallel circuit simulation approach, we create two levels of parallelism. At the higher level of parallelism, we start multiple simulation algorithms in parallel for a given simulation task. Interalgorithm communication is established to enable simulation algorithms to exchange useful information so that they could advance faster than without doing so. At the lower level of parallelism, each algorithm within the MA framework utilizes fine-grained parallel techniques such as parallel device evaluation and parallel matrix solve to fully harness the available hardware resources. By combining the two levels of parallelism, the computing power of the multicore or many-core processor platforms can be fully utilized to achieve superlinear speedup in circuit simulation.
Xiaoji Ye, Wei Dong 0002, Peng Li 0001, Sani R. Nassif
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 Hierarchical Analog/Mixed-Signal Circuit Optimization Under Process Variations and Tuning
abstract
A hierarchical optimization methodology is presented to achieve robust analog/mixed-signal circuit design with consideration of process variations. Hierarchical optimization using building circuit block Pareto models is an efficient approach for optimizing nominal performances of large analog circuits. However, yield-aware system optimization, as dictated by the need for safeguarding chip manufacturability in scaled technologies, is completely nontrivial. Two fundamental difficulties are addressed for achieving such a methodology: yield-aware Pareto performance characterization at the building block level and yield-aware optimization problem formulation at the system level. In addition, postsilicon tuning in complex mixed-signal system designs is investigated and the proposed optimization framework is extended for such systems. The presented methodology is demonstrated by hierarchical optimization of a phased-locked loop consisting of multiple building blocks and self-tuning function blocks.
Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2011 Parallel circuit simulation with adaptively controlled projective integration
abstract
In this article, a parallel transient circuit simulation approach based on an adaptively-controlled time-stepping scheme is proposed. Different from the widely-used implicit numerical integration techniques in most transient simulators, this work exploits the recently-developed explicit telescopic projective numerical integration method for efficient parallel circuit simulation. Because telescopic projective integration addresses the well-known stability issue of explicit numerical integrations by adopting combinations of inner integrators and outer integrators in a multilevel fashion, the simulation time-step is no longer limited by the smallest time constant in the circuit. With dynamic control of telescopic projective integration, the proposed projective integration framework not only leads to noticeable efficiency improvement in circuit simulation, it also lends itself to straightforward parallelization due to its explicit nature. The latter has led to encouraging runtime efficiencies, observed on shared-memory parallel platforms. In addition to solving standard initial-value problems (IVPs) of differential equations, the same telescopic integration framework is adopted for solving final-value problems (FVPs), where the system is integrated backwards in time. Through a new elegant formulation, we show how an IVP and FVP can be simultaneously solved to allow for a coarse-grained bidirectional parallel circuit simulation scheme. Such a bidirectional approach is demonstrated in the context of parallel shooting-Newton-based steady-state circuit analysis. The proposed bidirectional approach has unique and favorable properties: the solutions of the two ODE problems are completely data-independent with built-in automatic load balancing.
Wei Dong 0002, Peng Li 0001
ACM Trans. Design Autom. Electr. Syst.2
2011 Locality-Driven Parallel Static Analysis for Power Delivery Networks
abstract
Large VLSI on-chip Power Delivery Networks (PDNs) are challenging to analyze due to the sheer network complexity. In this article, a novel parallel partitioning-based PDN analysis approach is presented. We use the boundary circuit responses of each partition to divide the full grid simulation problem into a set of independent subgrid simulation problems. Instead of solving exact boundary circuit responses, a more efficient scheme is proposed to provide near-exact approximation to the boundary circuit responses by exploiting the spatial locality of the flip-chip-type power grids. This scheme is also used in a block-based iterative error reduction process to achieve fast convergence. Detailed computational cost analysis and performance modeling is carried out to determine the optimal (or near-optimal) number of partitions for parallel implementation. Through the analysis of several large power grids, the proposed approach is shown to have excellent parallel efficiency, fast convergence, and favorable scalability. Our approach can solve a 16-million-node power grid in 18 seconds on an IBM p5-575 processing node with 16 Power5+ processors, which is 18.8X faster than a state-of-the-art direct solver.
Zhiyu Zeng, Peng Li 0001, Vivek Sarin
ACM Trans. Design Autom. Electr. Syst.3
2011 Parallel On-Chip Power Distribution Network Analysis on Multi-Core-Multi-GPU Platforms
abstract
The challenging task of analyzing on-chip power (ground) distribution networks with multimillion node complexity and beyond is key to today's large chip designs. For the first time, we show how to exploit recent massively parallel single-instruction multiple-thread (SIMT)-based graphics processing unit (GPU) platforms to tackle large-scale power grid analysis with promising performance. Several key enablers including GPU-speciflc algorithm design, circuit topology transformation, workload partitioning, performance tuning are embodied in our GPU-accelerated hybrid multigrid (HMD) algorithm (GpuHMD) and its implementation. We also demonstrate that using the HMD solver as a preconditioner, the conjugate gradient solver can converge much faster to the true solution with good robustness. Extensive experiments on industrial and synthetic benchmarks have shown that for DC power grid analysis using one GPU, the proposed simulation engine achieves up to 100× runtime speedup over a state-of-the-art direct solver and more than 50× speedup over the CPU based multigrid implementation, while utilizing a four-core-four-GPU system, a grid with eight million nodes can be solved within about 1 s. It is observed that the proposed approach scales favorably with the circuit complexity, at a rate about 1 s per two million nodes on a single GPU card.
Zhiyu Zeng, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2010 Parallel program performance modeling for runtime optimization of multi-algorithm circuit simulation
abstract
With the increasing popularity of multi-core processors and the promise of future many-core systems, parallel CAD algorithm development has attracted a significant amount of research effort. However, a highly relevant issue, parallel program performance modeling has received little attention in the EDA community. Performance modeling serves the critical role of guiding parallel algorithm design and provides a basis for runtime performance optimization. We propose a systematic composable approach for the performance modeling of a recently developed hierarchical multi-algorithm parallel circuit simulation (HMAPS) approach. The unique integration of inter- and intra-algorithm parallelisms allows a multiplicity of parallelisms to be exploited in HMAPS and also creates interesting modeling challenges in forms of complex performance tradeoffs and large runtime configuration space. We model the performances of key subtask entities as functions of workload and parallelism. We address significant complications introduced by inter-algorithm interactions in terms of memory contention and collaborative simulation behavior via novel penalty and statistical based modeling. The proposed approach is able to accurately predict the parallel performance of a given HMAPS configuration and hence enables the runtime optimization of the parallel simulation code.
Xiaoji Ye, Peng Li 0001
DAC2
2010 Exploiting reconfigurability for low-cost in-situ test and monitoring of digital PLLs
abstract
We exploit the reconfigurability of recent all-digital PLL designs to provide novel in-situ output jitter test and diagnosis abilities under multiple parametric variations of key analog building blocks. Digital signatures are collected and processed under specifically designed loop filter configurations to facilitate low-cost high-accuracy performance prediction and diagnosis.
Leyi Yin, Peng Li 0001
DAC2
2010 Tradeoff analysis and optimization of power delivery networks with on-chip voltage regulation
abstract
Integrating a large number of on-chip voltage regulators holds the promise of solving many power delivery challenges through strong local load regulation and facilitates system-level power management. The quantitative understanding of such complex power delivery networks (PDNs) is hampered by the large network complexity and interactions between passive on-die/package-level circuits and a multitude of nonlinear active regulators. We develop a fast combined GPU-CPU analysis engine encompassing several simulation strategies, optimized for various subcomponents of the network. Using accurate quantitative analysis, we demonstrate the significant performance improvement brought by onchip low-dropout regulators (LDOs) in terms of suppressing high-frequency local voltage droops and avoiding the mid-frequency resonance caused by off-chip inductive parasitics. We perform comprehensive analysis on the tradeoffs among overhead of on-chip LDOs, maximum voltage droop and overall power efficiency. We conduct systematic design optimization by developing a simulation-based nonlinear optimization strategy that determines the optimal number of on-chip LDOs required and on-board input voltage, and the corresponding voltage droop and power efficiency for PDNs with multiple power domains.
Zhiyu Zeng, Xiaoji Ye, Peng Li 0001
DAC4
2010 Separatrices in high-dimensional state space: system-theoretical tangent computation and application to SRAM dynamic stability analysis
abstract
Shrinking access cycle times and the employment of dynamic read/write assist circuits have made the use of standard static noise margins increasingly problematic for scaled SRAM designs. Recently proposed dynamic noise margins precisely characterize dynamic stability using the concept of stability boundaries, or separatrices, and provide elegant separatrix tracing algorithm. However, the present separatrix characterization method is only efficient in the two-dimensional state space and hence not practically applicable to fully extracted SRAM designs with additional parasitics. We present a rigorous system-theoretical approach for computing the tangent approximation to the separatrix in the high-dimensional space. Using this as a basis, we develop fast method based on tangent approximation and exact iterative-refinement method for analyzing SRAM dynamic stability. The proposed algorithms have been implemented as a SPICE-like CAD tool and are broadly applicable to efficient computation of dynamic noise margins.
Yong Zhang 0049, Peng Li 0001, Garng M. Huang
DAC2
2010 Fast thermal analysis on GPU for 3D-ICs with integrated microchannel cooling
abstract
While effective thermal management for 3D-ICs is becoming increasingly challenging due to the ever increasing power density and chip design complexity, traditional heat sinks are expected to quickly reach their limits for meeting the cooling needs of 3D-ICs. Alternatively, integrated liquid-cooled microchannel heat sink becomes one of the most effective solutions. For the first time, we present fast GPU-based thermal simulation methods for 3D-ICs with integrated microchannel cooling. Based on the physical heat dissipation paths of 3D-ICs with integrated microchannels, we propose novel preconditioned iterative methods that can be efficiently accelerated on GPU's massively parallel computing platforms. Unlike the CPU-based solver development environment in which many existing sophisticated numerical simulation methods (matrix solvers) can be readily adopted and implemented, GPU-based thermal simulation demands more efforts in the algorithm and data structure design phase, and requires careful consideration of GPU's thread/memory organizations, data access/communication patterns, arithmetic intensity, as well as the hardware occupancies. As shown in various experimental results, our GPU-based 3D thermal simulation solvers can achieve up to 360X speedups over the best available direct solvers and more than 35X speedups compared with the CPU-based iterative solvers, without loss of accuracy.
Peng Li 0001
ICCAD2
2010 On behavioral model equivalence checking for large analog/mixed signal systems
abstract
This paper presents a systematic, hierarchical, optimization based semi-formal equivalence checking methodology for large analog/mixed signal systems such as PLLs, ADCs and I/O's. We verify the equivalence between a behavioral model and its electrical implementation over a limited, but highly likely, input space defined as the Constrained Behavioral Input Space. Further, we clearly distinguish between the behavioral and electrical domains and define mappings between the two domains to allow for calculation of deviation between the behavioral and electrical implementation. The verification problem is then formulated as an optimization problem which is solved by interfacing a SQP based optimizer with commercial circuit simulation tools. The proposed methodology is then applied for equivalence checking of a PLL as a test case.
Peng Li 0001
ICCAD2
2010 On-the-fly runtime adaptation for efficient execution of parallel multi-algorithm circuit simulation
abstract
The past several years have witnessed a significant interest in developing parallel CAD algorithms and implementations that exploit various multi-core and distributed computing hardware. In addition to fundamental parallel algorithm design, the ability in modeling parallel performance and facilitating runtime optimization is indispensable for achieving good efficiency for complex parallel CAD applications. Under the context of a recently developed hierarchical multi-algorithm parallel circuit simulation (HMAPS) framework, we demonstrate a runtime optimization approach that allows for automatic on-the-fly reconfiguration of the parallel simulation code. We show how the runtime information, collected as parallel simulation proceeds, can be combined with static parallel performance models to enable dynamic adaptation of parallel simulation execution for improved performance and robustness. Our results have shown that the proposed approach not only finds the near-optimal code configuration over a large configuration space, it also outperforms multi-algorithm circuit simulation assisted only with static pre-runtime parallel performance modeling.
Xiaoji Ye, Peng Li 0001
ICCAD2
2010 Accurate clock mesh sizing via sequential quadraticprogramming
abstract
Clock mesh is widely used in microprocessor designs for achieving low clock skew and high variation tolerance. Clock mesh optimization is a very difficult problem because it has highly-connected structure and requires accurate delay models which are computationally expensive. Existing methods on clock network optimization are either restricted to clock trees, which are easy to be separated into smaller problems, or naive heuristics based on crude delay models. In this paper, we propose a clock mesh sizing algorithm which is aimed to minimize mesh wire area with consideration of clock skew constraints. This algorithm is a systematic solution search through rigorous Sequential Quadratic Programming (SQP). The SQP is guided by an efficient adjoint sensitivity analysis which has near-SPICE-level accuracy and faster-than-SPICE speed. Experimental results on various benchmark circuits indicate that our algorithm leads to significant wire area reduction while maintaining low clock skew.
Venkata Rajesh Mekala, Yifang Liu, Xiaoji Ye, Jiang Hu 0001, Peng Li 0001
ISPD5
2010 Scalable Analysis of Mesh-Based Clock Distribution Networks Using Application-Specific Reduced Order Modeling
abstract
Clock meshes possess inherent low clock skews and excellent immunity to process-voltage-temperature variations, and have increasingly found their way to high-performance integrated circuit designs. However, analysis of such massively coupled networks is significantly hindered by the sheer size of the network and tight coupling between non-tree interconnects and large numbers of clock drivers. While the SPICE simulation of large clock meshes is often intractable, standard interconnect model order reduction algorithms also fail due to the large number of input/output ports introduced by clock drivers. The presented approach is motivated by the key observation of the steady-state operation of the clock networks while its efficiency is facilitated by exploringnewclock-mesh specific harmonic-weighted model order reduction algorithm and locality analysis via port sliding. The scalability of the analysis is significantly improved by eliminating the need for computing infeasible multi-port passive reduced order interconnect models with large port count and decomposing the overall task into very tractable and naturally parallelizable model generation and fast Fourier transform/inverse-fast Fourier transform operations, all on a per driver or per sink basis. We demonstrate the application of our approach by feasibly analyzing large clock meshes with excellent accuracy.
Xiaoji Ye, Peng Li 0001, Min Zhao 0001, Rajendran Panda, Jiang Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Discrete Buffer and Wire Sizing for Link-Based Non-Tree Clock Networks
abstract
Clock network is a vulnerable victim of variations as well as a main power consumer in many integrated circuits. Recently, link-based non-tree clock network attracts people's attention due to its appealing tradeoff between variation tolerance and power overhead. In this work, we investigate how to optimize such clock networks through buffer and wire sizing. A two-stage hybrid optimization approach is proposed. It considers the realistic constraint of discrete buffer/wire sizes and is based on accurate delay models. In order to provide reliable and efficient guidance for the optimization, we suggest to apply support vector machine (SVM)-based machine learning as a surrogate for expensive circuit-level simulation. Experimental results on benchmark circuits show that our sizing method can reduce clock skew by 45% on average with very small increase on power dissipation.
Rupak Samanta, Jiang Hu 0001, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2010 Combinatorial Algorithms for Fast Clock Mesh Optimization
abstract
Clock mesh has been widely used to distribute the clock signal across the chip. Clock mesh is driven by a top-level tree and a set of mesh buffers. We present fast and efficient combinatorial algorithms to simultaneously identify the candidate locations as well as sizes of the buffers driving the clock mesh. We show that such a sizing offers a better solution than inserting buffers of uniform size across the mesh. Due to the high redundancy, a mesh architecture offers high tolerance toward variations in clock skew. However, such a redundancy comes at the expense of mesh wire length and power dissipation. Based on survivable network theory, we formulate the problem to reduce the clock mesh by retaining only those edges that are critical to maintain redundancy. Such a formulation offers designer the option to tradeoff between power and tolerance to process variations. We present efficient postprocessing techniques to reduce the size of the mesh buffers after mesh reduction. Experimental results indicate that our techniques can result inpower savings up to 28% with less than 3.3% delay penalty. We also present driver models that can help in simulating the clock mesh. Such models achieve near-HSPICE accuracy with significant speedup in run time.
Ganesh Venkataraman, Jiang Hu 0001, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2009 Parallelizable stable explicit numerical integration for efficient circuit simulation
abstract
This work exploits the recently developed telescopic projective numerical integration method for efficient parallel circuit simulation. Stable explicit numerical integration is achieved by adopting an explicit inner integrator (e.g. forward Euler) in a multi-level telescopic projective framework, thereby addressing the well-known stability limitation of many explicit numerical integration methods. In the presented approach, the effective time step of the entire multilevel integration is no longer limited by the smallest time constant of the circuit so as to safeguard stability. Rather, it is controlled solely by the accuracy requirement. This makes it possible to explore the natural parallelizability of such explicit integration method for parallel circuit simulation. We demonstrate the potential of the presented approach and its parallel implementation on multi-core machines with encouraging initial results.
Wei Dong 0002, Peng Li 0001
DAC2
2009 Closed-loop modeling of power and temperature profiles of FPGAs
abstract
In recent times, the contribution of leakage power to the total power consumption of a chip has been increasing at an alarming rate. Leakage power is expected to exceed dynamic power in newer process technologies. Since leakage exhibits an exponential increase with temperature, it is possible that the high leakage of an IC causes a temperature increase, which in turn causes an increase in leakage, and so on, until the IC fails due to overheating. At the very least, this may cause the temperature and power consumption of the IC to be poorly estimated by traditional thermal or power modeling techniques. We developed a framework to model this situation in an FPGA context. Our CAD framework accurately models the total power consumption of the design at a given temperature, finds the thermal profile of the IC under this power consumption, and then uses this new thermal information to update the power consumption. This is iterated until the temperature of the IC converges, or until the temperatures on the die exceed a safe value. The iterations are very fast, due to the use of accurate and compact mathematical macromodels for leakage and temperature computation in the inner loop. We have exhaustively verified the fidelity of all our leakage macromodels. They estimate the leakage, at any temperature, to within 3% of the values generated by SPICE, while providing greater than four orders of magnitude speedup over explicit SPICE runs. Our experiments show that this model helps avoid an incorrect estimation of chip temperature and total power consumption, and also helps detect the increase in device temperature beyond a safe value. The average (maximum) error of our temperature estimates has been found to be within 1% (2.5%) compared to a full-chip 3D temperature modeling tool.
Kanupriya Gulati, Sunil P. Khatri, Peng Li 0001
FPGA3
2009 Final-value ODEs: Stable numerical integration and its application to parallel circuit analysis
abstract
While solving initial-value ODEs is the de facto approach to time-domain circuit simulation, the opposite act, solving final-value ODEs, has been neglected for a long time. Stable numerical integration of initial-value ODEs involves significant complications; the application of standard integration methods simply leads to instability. We show that not only practically meaningful applications of final-value ODE problems exist, but also the inherent stability challenges may be addressed by recently proposed numerical methods. Furthermore, we demonstrate an elegant bi-directional parallel circuit simulation scheme, where one time-domain simulation task is sped up by simultaneously solving initial and final-value ODEs, one from each end of the time axis. The proposed approach has unique and favorable properties: the solutions of the two ODE problems are completely data independent with built-in automatic load balancing. As a specific application study, we demonstrate the proposed technique under the contexts of parallel digital timing simulation and the shooting-Newton based steady-state analysis.
Wei Dong 0002, Peng Li 0001
ICCAD2
2009 Nonvolatile memristor memory: Device characteristics and design implications
abstract
The search for new nonvolatile universal memories is propelled by the need for pushing power-efficient nanocomputing to the next higher level. As a potential contender for the next-generation memory technology of choice, the recently found "the missing fourth circuit element", memristor, has drawn a great deal of research interests. In this paper, we characterize the fundamental electrical properties of memristor devices by encapsulating them into a set of compact closed-form expressions. Our derivations provide valuable design insights and allow a deeper understanding of key design implications of memristor-based memories. In particular, we investigate the design of read and write circuits and analyze data integrity and noise-tolerance issues.
Yenpo Ho, Garng M. Huang, Peng Li 0001
ICCAD3
2009 Leveraging efficient parallel pattern search for clock mesh optimization
abstract
Mesh-based clock distribution network has been employed in many high-performance microprocessor designs due to its favorable properties such as low clock skew and robustness. Such clock distributions are usually highly complex. While the simulation of clock meshes is already time consuming, tuning such networks under tight performance constraints is a more daunting task. In this paper, we address the challenging task of driver size optimization with a goal of skew minimization. The expensive objective function evaluations and difficulty in getting explicit sensitivity information make this problem intractable to standard optimization methods. We propose to explore the recently developed asynchronous parallel pattern search (APPS) method for efficient driver size tuning. While being a search-based method, APPS not only provides the desirable derivative-free optimization capability, but is also amenable to parallelization and possesses appealing theoretically rigorous convergence properties. We show how such a method can lead to powerful parallel sizing optimization of large clock meshes with significant runtime and quality advantages over the traditional sequential quadratic programming (SQP) method. We also show how design-specific properties and speeding-up techniques can be exploited to make the optimization even more efficient while maintaining the convergence of APPS in a practical sense.
Xiaoji Ye, Srinath Narasimhan, Peng Li 0001
ICCAD3
2009 Gene-regulatory memories: Electrical-equivalent modeling, simulation and parameter identification
abstract
The development of gene-regulatory memory circuits provides key understandings of biological information storage and enables new biological applications. Computer models and simulations can provide quantitative analysis and prediction of the behaviors and functions of genetic networks, thereby providing valuable verification and design guidance. In this paper, we model the nonlinear dynamics associated with various chemical reactions in gene-regulatory memory networks using chemical reaction equations. These reaction equations are mapped into a set of electrical-equivalent models and the network is simulated by an extended SPICE-like circuit simulation environment. Furthermore, we address the practical difficulty in direct characterization of network model parameters by developing a simulation-driven Bayesian framework for parameter identification. To ensure the reliable identification of key system properties, we propose a two-step structure-preserving parameter identification approach. The first step infers bistability, the most critical characteristics of a memory device; and the second step is geared towards identifying dynamical properties of the network while maintaining the identified bistability. We demonstrate the proposed approaches through extensive simulations that well agree with established biological understandings and identified networks that recreate measured circuit responses in a statistical sense.
Yong Zhang 0049, Peng Li 0001
ICCAD2
2009 A Parallel Harmonic-Balance Approach to Steady-State and Envelope-Following Simulation of Driven and Autonomous Circuits
abstract
In this paper, we present a parallel harmonic-balance approach, applicable to the steady-state and envelope-following analyses of both driven and autonomous circuits. Our approach is centered on a naturally parallelizable preconditioning technique that speeds up the core computation in harmonic-balance-based analysis. As a coarse-grained parallel approach by algorithm construction, the proposed method facilitates parallel computing via the use of domain knowledge and simplifies parallel programming compared with fine-grained strategies. The proposed parallel preconditioning technique can be combined with more conventional parallel approaches such as parallel device model evaluation, parallel fast Fourier transform operation, and parallel matrix-vector product to further improve runtime efficiency. In our message-passing-interface-based implementation over a cluster of workstations and multithreading-based implementation on a shared-memory machine, favorable runtime speedups with respect to the conventional serial approaches and the serial implementations of the same parallel algorithms are achieved.
Wei Dong 0002, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 An On-the-Fly Parameter Dimension Reduction Approach to Fast Second-Order Statistical Static Timing Analysis
abstract
While first-order statistical static timing analysis (SSTA) techniques enjoy good runtime efficiency desired for tackling large industrial designs, more accurate second-order SSTA techniques have been proposed to improve the analysis accuracy, but at the cost of high computational complexity. Although many sources of variations may impact the circuit performance, considering a large number of inter- and intra-die variations in the traditional SSTA is very challenging. In this paper, we address the analysis complexity brought by high parameter dimensionality in SSTA and propose an accurate yet fast second-order SSTA algorithm based on novel on-the-fly parameter dimension reduction techniques. By developing a reduced rank regression (RRR)-based approach and a method of moments (MOM)-based parameter reduction algorithm within the block-based SSTA flow, we demonstrate that accurate second-order SSTA can be extended to a much higher parameter dimensionality than what is possible before. Our experimental results have shown that the proposed parameter reductions can achieve up to 10times parameter dimension reduction and lead to significantly improved second-order SSTA under a large set of process variations.
Peng Li 0001, Yaping Zhan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 Performance-Oriented Parameter Dimension Reduction of VLSI Circuits
abstract
To account for the growing process variability in modern VLSI technologies, circuit models parameterized in a multitude of parametric variations are becoming increasingly indispensable in robust circuit design. However, the high parameter dimensionality can introduce significant complexity and may even render variation-aware performance analysis and optimization completely intractable. We present a performance-oriented parameter dimension reduction framework to reduce the modeling complexity associated with high parameter dimensionality. Our framework has a theoretically sound statistical basis, namely, reduced rank regression (RRR) and its various extensions that we have introduced for more practical VLSI circuit modeling. For a variety of VLSI circuits including interconnects and CMOS digital circuits, it is shown that this parameter reduction framework can provide more than one order of magnitude reduction in parameter dimensionality. Such parameter reduction immediately leads to reduced simulation cost in sampling-based performance analysis, and more importantly, highly efficient parameterized sub-circuit models that are instrumental in tackling the complexity of variation-tolerance VLSI system design.
Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2008 WavePipe: parallel transient simulation of analog and digital circuits on multi-core shared-memory machines
abstract
While the emergence of multi-core shared-memory machines offers a promising computing solution to ever complex chip design problems, new parallel CAD methodologies must be developed to gain the full benefit of these increasingly parallel computing systems. We present a parallel transient simulation methodology and its multi-threaded implementation for general analog and digital ICs. Our new approach, Waveform Pipelining (abbreviated as WavePipe), exploits coarsegrained application-level parallelism by simultaneously computing circuit solutions at multiple adjacent time points in a way resembling hardware pipelining. There are two embodiments in WavePipe: backward and forward pipelining schemes. While the former creates independent computing tasks that contribute to a larger future time step by moving backwards in time, the latter performs predictive computing along the forward direction of the time axis. Unlike existing relaxation methods, WavePipe facilitates parallel circuit simulation without jeopardying convergence and accuracy. As a coarse-grained parallel approach, WavePipe not only requires low parallel programming effort, more importantly, it creates new avenues to fully utilize increasingly parallel hardware by going beyond conventional finer grained parallel device model evaluation and matrix solutions.
Wei Dong 0002, Peng Li 0001, Xiaoji Ye
DAC2
2008 SRAM dynamic stability: theory, variability and analysis
abstract
Technology scaling in sub-100 nm regime has significantly shrunk the SRAM stability margins in data retention, read and write operations. Conventional static noise margins (SNMs) are unable to capture nonlinear cell dynamics and become inappropriate for state-of-the-art SRAMs with shrinking access time and/or advanced dynamic read-write-assist circuits. Using the insights gained from rigorous nonlinear system theory, we define the much needed SRAM dynamic noise margins (DNMs). The newly defined DNMs not only capture key SRAM nonlinear dynamical characteristics but also provide valuable design insights. Furthermore, we show how system theory can be exploited to develop CAD algorithms that can analyze SRAM dynamic stability characteristics three orders of magnitude faster than a brute-force approach while maintaining SPICE-level accuracy. We also demonstrate a parametric dynamic stability analysis approach suitable for low-probability cell failures, leading to three orders of magnitude runtime speedup for yield analysis under high-sigma parameter variations.
Wei Dong 0002, Peng Li 0001, Garng M. Huang
ICCAD2
2008 Multigrid on GPU: tackling power grid analysis on parallel SIMT platforms
abstract
The challenging task of analyzing on-chip power (ground) distribution networks with multi-million node complexity and beyond is key to todaypsilas large chip designs. For the first time, we show how to exploit recent massively parallel single-instruction multiple-thread (SIMT) based graphics processing unit (GPU) platforms to tackle power grid analysis with promising performance. Several key enablers including GPU-specific algorithm design, circuit topology transformation, workload partitioning, performance tuning are embodied in our GPU-accelerated hybrid multigrid algorithm, GpuHMD, and its implementation. In particular, a proper interplay between algorithm design and SIMT architecture consideration is shown to be essential to achieve good runtime performance. Different from the standard CPU based CAD development, care must be taken to balance between computing and memory access, reduce random memory access patterns and simplify flow control to achieve efficiency on the GPU platform. Extensive experiments on industrial and synthetic benchmarks have shown that the proposed GpuHMD engine can achieve 100times runtime speedup over a state-of-the-art direct solver and be more than 15times faster than the CPU based multigrid implementation. The DC analysis of a 1.6 million-node industrial power grid benchmark can be accurately solved in three seconds with less than 50 MB memory on a commodity GPU. It is observed that the proposed approach scales favorably with the circuit complexity, at a rate about one second per million nodes.
Peng Li 0001
ICCAD2
2008 MAPS: multi-algorithm parallel circuit simulation
abstract
The emergence of multi-core and many-core processors has introduced new opportunities and challenges to EDA research and development. While the availability of increasing parallel computing power holds new promise to address many computing challenges in CAD, the leverage of hardware parallelism can only be possible with a new generation of parallel CAD applications. In this paper, we propose a novel multi-algorithm parallel circuit simulation approach (MAPS) and its multi-core implementation to expedite one of the most fundamental CAD applications: transistor-level transient circuit simulation. MAPS starts multiple simulation algorithms in parallel for a given simulation task. By properly synchronizing these algorithms on-the-fly, we exploit the diversity in simulation algorithms to achieve possibly superlinear overall speedup in transient simulation. In addition, our unique multi-algorithm framework allows unique safe exploration of simulation methods that are conventionally discarded due to convergence concerns. As a coarse grained parallel simulation approach, the implementation of MAPS demands a minimum of parallel programming effort and allows for reuse of existing serial simulation codes.
Xiaoji Ye, Wei Dong 0002, Peng Li 0001, Sani R. Nassif
ICCAD3
2008 Yield-aware hierarchical optimization of large analog integrated circuits
abstract
Hierarchical optimization using building circuit block pareto performance models is an efficient and well established approach for optimizing the nominal performances of large analog circuits. However, the extension to yield-aware hierarchical methodology, as dictated by the need for safeguarding chip manufacturability in scaled technologies, is completely nontrivial. We address two fundamental difficulties in achieving such a methodology: yield-aware pareto performance characterization at the building block level and yield-aware system-level optimization problem formulation. It is shown that our approach is not only able to effectively capture the block performance trade-offs at different yield levels, but also correctly formulate the whole system yield and efficiently perform system-level optimization in presence of process variations. Our approach extends the efficiency of hierarchical analog optimization, enjoyed for improving nominal circuit performances, to yield-aware optimization. Our methodology is demonstrated by the hierarchical optimization of a phased locked loop (PLL) consisting of multiple circuit blocks.
Peng Li 0001
ICCAD2
2008 Modeling dynamic stability of SRAMS in the presence of single event upsets (SEUs)
abstract
SRAM yield is very important from an economics viewpoint, because of the extensive use of memory in modern processors and SOCs. Therefore, SRAM stability analysis tools have become essential. SRAM stability analysis based on static noise margin (SNM) often results in pessimistic designs because SNM cannot capture the transient behavior of the noise. Therefore, to improve accuracy, dynamic stability analysis is required. The model presented in this paper performs dynamic stability analysis of an SRAM cell in the presence of an SEU event. The experimental results demonstrate that our model is very accurate, with a critical charge estimation error of 2.5% compared to HSPICE. The run-time of our model is also significantly lower (1200× lower) than the HSPICE run-time. Thus, our model enables the SRAM designer to quickly and accurately analyze stability during the design phase.
Rajesh Garg, Peng Li 0001, Sunil P. Khatri
ISCAS2
2008 Discrete buffer and wire sizing for link-based non-tree clock networks
abstract
Clock network is a vulnerable victim of variations as well as a main power consumer in many integrated circuits. Recently, link-based non-tree clock network attracts people’s attention due to its appealing tradeoff between variation tolerance and power overhead. In this work, we investigate how to optimize such clock networks through buffer and wire sizing. A two-stage hybrid optimization approach is proposed. It considers the realistic constraint of discrete buffer/wire sizes and is based on accurate delay models. In order to provide reliable and efficient guidance for the optimization, we suggest to apply SVM (Support Vector Machine) based machine learning as a surrogate for expensive circuit-level simulation. Experimental results on benchmark circuits show that our sizing method can reduce clock skew by 43 % on average with very small increase on power dissipation.
Rupak Samanta, Jiang Hu 0001, Peng Li 0001
ISPD3
2008 A Preconditioned Hierarchical Algorithm for Impedance Extraction of Three-Dimensional Structures With Multiple Dielectrics
abstract
This paper presents the first boundary element method (BEM) impedance extraction algorithm for interconnects with multiple dielectrics. Multiple dielectrics are common in integrated circuits and packages. However, previous BEM algorithms, including FastImp and FastPep, assume uniform dielectric due to their limitation, thus causing considerable errors. Our algorithm introduces a circuit formulation which makes it possible to utilize either multilayer Green's function or equivalent charge method to extract impedance in multiple dielectrics. The novelty of the formulation is the reduction of the unknowns and the application of hierarchical data structure. The hierarchical data structure permits efficient sparsification transformation and preconditioners to accelerate the linear equation solver. Experimental results demonstrate that the new algorithm is accurate and efficient. For uniform dielectric problems, our algorithm is more accurate than FastImp while its number of unknowns is ten times less than that of FastImp. For multiple dielectric problems, its relative error with respect to HFSS is below 3%.
Peng Li 0001, Vivek Sarin, Weiping Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2008 Statistical Static Timing Analysis Considering Process Variation Model Uncertainty
abstract
Increasing variability in modern manufacturing processes makes it important to predict the yields of chip designs at early design stage. In recent years, a number of statistical static timing analysis (SSTA) and statistical circuit optimization techniques have emerged to quickly estimate the design yield and perform robust optimization. These statistical methods often rely on the availability of statistical process variation models whose accuracy, however, is severely hampered by the limitations in test structure design, test time, and various sources of inaccuracy inevitably incurred in process characterization. To consider model characterization inaccuracy, we present an efficient importance sampling based optimization framework that can translate the uncertainty in process models to the uncertainty in circuit performance, thus offering the desired statistical best/worst case circuit analysis capability accounting for the unavoidable complexity in process characterization. Furthermore, our new technique provides valuable guidance to process characterization. Examples are included to demonstrate the application of our general analysis framework under the context of SSTA.
Wei Dong 0002, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 Accelerating Harmonic Balance Simulation Using Efficient Parallelizable Hierarchical Preconditioning
abstract
An efficient parallelizable hierarchical preconditioning framework is proposed to accelerate the simulation of analog and RF circuits using Harmonic Balance (HB). We show that the proposed hierarchical preconditioning scheme not only brings significant runtime speedup over the conventional block diagonal (BD) preconditioner for driven circuits, but also enhances the simulation of autonomous circuits. Due to the specific algorithm construction, this hierarchical preconditioning can be achieved by solving a set of linear matrix problems with tree-like data dependency. Hence, it can be naturally parallelized to further improve the simulation efficiency. An efficient parallel HB simulation engine has been developed by distributing the simulation workload over a cluster of machines via message passing interface (MPI). Our experiments have shown that this new parallel HB algorithm can achieve the significant runtime speedup over the conventional serial preconditioning technique.
Wei Dong 0002, Peng Li 0001
DAC2
2007 Fast Second-Order Statistical Static Timing Analysis Using Parameter Dimension Reduction
abstract
The ability to account for the growing impacts of multiple process variations in modern technologies is becoming an integral part of nanometer VLSI design. Under the context of timing analysis, the need for combating process variations has sparkled a growing body of statistical static timing analysis (SSTA) techniques. While first-order SSTA techniques enjoy good runtime efficiency desired for tackling large industrial designs, more accurate second-order SSTA techniques have been proposed to improve the analysis accuracy, but at the cost of high computational complexity. Although many sources of variations may impact the circuit performance, considering a large number of inter-die and intra-die variations in the traditional SSTA analysis is very challenging. In this paper, we address the analysis complexity brought by high parameter dimensionality in static timing analysis and propose an accurate yet fast second-order SSTA algorithm based upon novel parameter dimension reduction. By developing reduced-rank regression based parameter reduction algorithms within block-based SSTA flow, we demonstrate that accurate second order SSTA analysis can be extended to a much higher parameter dimensionality than what is possible before. Our experimental results have shown that the proposed parameter reduction can achieve up to 10X parameter dimension reduction and lead to significantly improved second-order SSTA analysis under a large set of process variations.
Peng Li 0001, Yaping Zhan
DAC2
2007 Statistical Leakage Power Minimization Using Fast Equi-Slack Shell Based Optimization
abstract
Leakage power is becoming an increasingly important component of total chip power consumption for nanometer IC designs. Minimization of leakage power unavoidably enforces the consideration of the key sources of process variations, namely transistor channel length and threshold variations, since both have a significant impact on timing and leakage power. However, the statistical nature of chip performances often requires the use of expensive statistical analysis and optimization techniques in a leakage minimization task, contributing to high computational complexity. Further, the commonly used discrete cell libraries bring specific difficulty for design optimization and render pure continuous sizing and VT optimization algorithm suboptimal. In this paper, we present a fast yet effective approach to statistical leakage power reduction via gate sizing and multiple VT assignment. The proposed technique achieves the runtime efficiency via the use of the novel concept of equi-slack shells and performs fast leakage power reduction on the basis of shells while maintaining the timing yield. When combined with a finer grained gate-based post tuning step, the presented technique achieves Superior runtime efficiency while offering significant leakage power reduction.
Xiaoji Ye, Yaping Zhan, Peng Li 0001
DAC3
2007 A Framework for Accounting for Process Model Uncertainty in Statistical Static Timing Analysis
abstract
In recent years, a large body of statistical static timing analysis and statistical circuit optimization techniques have emerged, providing important avenues to account for the increasing process variations in design. The realization of these statistical methods often demands the availability of statistical process variation models whose accuracy, however, is severely hampered by limitations in test structure design, test time and various sources of inaccuracy inevitably incurred in process characterization. Consequently, it is desired that statistical circuit analysis and optimization can be conducted based upon imprecise statistical variation models. In this paper, we present an efficient importance sampling based optimization framework that can translate the uncertainty in the process models to the uncertainty in parametric yield, thus offering the very much desired statistical best/worst-case circuit analysis capability accounting for unavoidable complexity in process characterization. Unlike the previously proposed statistical learning and probabilistic interval based techniques, our new technique efficiently computes tight bounds of the parametric circuit yields based upon bounds of statistical process model parameters while fully capturing correlation between various process variations. Furthermore, our new technique provides valuable guidance to process characterization. Examples are included to demonstrate the application of our general analysis framework under the context of statistical static timing analysis.
Wei Dong 0002, Peng Li 0001
DAC4
2007 Efficient VCO phase macromodel generation considering statistical parametric variations
abstract
With the growing concern of process variability, parame- terized circuit models are becoming increasingly important for circuit design and verification. Although techniques exist to extract compact VCO phase macromodels, a direct parametrization of VCO macromodels over a large set of parametric variations not only results in highly complex models, but also leads to significantly high computational cost. In this paper, an efficient parameterized VCO phase model generation technique is presented to capture the impacts of statistical parametric variations. The model extraction cost of our approach is significantly reduced by exploiting circuit-specific parameter dimension reduction, which effectively reduces the parameter space dimension over which the phase model needs to be extracted. The application of parameter reduction is facilitated by a novel and fast time-domain sampling technique that provides the essential statistical correlation data. Our numerical experiments have shown that the proposed model generation approach is more efficient than brute-force parametric modeling while producing accurate parameterized phase models that can capture large range parametric variations.
Wei Dong 0002, Peng Li 0001
ICCAD3
2007 A methodology for timing model characterization for statistical static timing analysis
abstract
While the increasing need for addressing process variability in sub-90nm VLSI technologies has sparkled a large body of statistical timing and optimization research, the realization of these techniques heavily depends on the availability of timing models that feed the statistical timing analysis engine. To target at this critical but less explored territory, in this paper, we present numerical and statistical modeling techniques that are suitable for the underlying timing model characterization infrastructure of statistical timing analysis. Our techniques are centered around the understanding that while the widening process variability calls for accurate non-first-order timing models, their deployment requires well-controlled characterization techniques to cope with the complexity and scalability. We present a methodology by which timing variabilities in interconnects and nonlinear gates are translated efficiently into quadratic timing models suitable for accurate statistical timing analysis. Specific parameter reduction techniques are developed to control the characterization cost that is a function of number of variation sources. The proposed techniques are extensively demonstrated under the context of logic stage timing characterization involving interactions between logic gates and interconnects.
Peng Li 0001
ICCAD2
2007 Analysis of large clock meshes via harmonic-weighted model order reduction and port sliding
abstract
Clock meshes posses inherent low clock skews and excellent immunity to PVT variations, and have increasingly found their way to high-performance IC designs. However, analysis of such massively coupled networks is significantly hindered by the sheer size of the network and tight coupling between non-tree interconnects and large numbers of clock drivers. The presented Harmonic-weighted model order reduction algorithm is motivated by the key observation of the steady-state operation of the clock networks, and its efficiency is facilitated by the locality analysis via port sliding. The scalability of the analysis is significantly improved by eliminating the need of computing infeasible multi-port passive reduced order interconnect models with large port count. And the overall task is decomposed into tractable and naturally parallelizable model generation and FFT/Inverse-FFT operations, all on a per driver or per sink basis.
Xiaoji Ye, Peng Li 0001, Min Zhao 0001, Rajendran Panda, Jiang Hu 0001
ICCAD2
2007 Impedance extraction for 3-D structures with multiple dielectrics using preconditioned boundary element method
abstract
In this paper, we present the first BEM impedance extraction algorithm for multiple dielectrics. The effect of multiple dielectrics is significant and efficient modeling is challenging. However, previous BEM algorithms, including Fastlmp and EastPep, assume uniform dielectric, thus causing considerable errors. The new algorithm introduces a circuit formulation which makes it possible to utilizes either multilayer Green's function or equivalent charge method to extract impedance in multiple dielectrics. The novelty of the formulation is the reduction of the number of unknowns and the application of the hierarchical data structure. The hierarchical data structure permits efficient sparsification transformation and preconditioners to accelerate the linear equation solver. Experimental results demonstrate that the new algorithm is accurate and efficient. For uniform dielectric problems, the new algorithm is one magnitude faster than Fastlmp, while its results differ from Fastlmp within 2%. For multiple dielectrics problems, its relative error with respect to HFSS is below 3%.
Peng Li 0001, Vivek Sarin, Weiping Shi
ICCAD2
2007 Yield-aware analog integrated circuit optimization using geostatistics motivated performance modeling
abstract
Automated circuit optimization is an important component of complex analog integrated circuit design. Today's analog designs must be optimized not only for nominal performance but also for robustness in order to maintain a reasonable yield with highly scaled VLSI technologies. The complex nature of analog/mixed-signal systems, however, makes this yield-aware analog circuit optimization extremely difficult and costly. In this paper, we adopt a geostatistics motivated approach (i.e. Kriging model) for efficient extraction of yield-aware Pareto front performance models for analog circuits. An iterative search based optimization approach is proposed to efficiently seek optimal performance tradeoffs under yield constraints in high-dimensional design parameter and process variation spaces. Our experiments confirm that the generated yield-aware Pareto fronts are accurate and the optimization procedure is very efficient. The latter is achieved by the well controlled iterative update scheme in the presented techniques which avoids an excessive number of time consuming transistor-level simulations.
Peng Li 0001
ICCAD2
2007 A methodology for systematic built-in self-test of phase-locked loops targeting at parametric failures
abstract
Test of phase-locked loops (PLLs) has been hampered by the complex mixed-signal nature of the system operation. While several built-in self-test (BIST) schemes have been proposed to reduce the cost of PLL test, a systematic BIST development methodology, specially targeting at the growing parametric failures in nanometer VLSI technologies, is yet to be developed. In this paper, we first present a detailed bottom- up parametric PLL macromodeling approach that is developed to realistically map the device-level process variations to the variations in system-level performances. Our parametric modeling techniques allow us to examine the correlations between the system performances and specific BIST measurements feasibly through behavioral-levels simulations. By exploiting our modeling infrastructure, an efficient methodology is then developed to facilitate evaluation and optimization of PLL BIST schemes. The proposed methodology is enabled by novel circuit-level macromodeling and powerful statistical dimension reduction techniques, the latter of which are employed to cope with the challenges imposed by the large number of process variations that must be considered. The application of our BIST development methodology is demonstrated by generating optimized BIST schemes that produce low mis-prediction levels for detection of parametric failures of charge-pump PLLs.
Peng Li 0001
ITC2
2007 Hierarchical Harmonic-Balance Methods for Frequency-Domain Analog-Circuit Analysis
abstract
As a widely adopted frequency-domain method, harmonic balance (HB) provides efficient steady-state circuit analysis for analog and RF circuits. The conventional matrix-implicit Krylov subspace technique with the block-diagonal (BD) preconditioner has made it possible to compute the steady-state responses of large-scale circuits. However, not all HB problems, particularly strongly nonlinear circuit problems, can be solved reliably or efficiently using the standard BD-preconditioning technique. In this paper, hierarchical HB methods are proposed wherein robust preconditioning is provided via solution of a set of approximate linearized HB problems of progressively smaller size across multiple levels of the problem hierarchy. These subproblems are constructed using the same matrix-implicit formulation to retain the memory efficiency of Krylov subspace methods. Moreover, the number of allocated Krylov subspace matrix solvers, hence the memory usage, is significantly reduced via a recently introduced solver-sharing technique. The efficiency of our hierarchical preconditioning technique is further improved by adopting a one-step correction to the standard BD preconditioner and a multigrid-motivated iterative scheme. It has been shown that the proposed approaches can achieve up to 10 runtime speedup over the popular BD preconditioner and robust convergence even for strongly nonlinear circuits for which the BD preconditioner fails to converge.
Wei Dong 0002, Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Utilizing Redundancy for Timing Critical Interconnect
abstract
Conventionally, the topology of signal net routing is almost always restricted to Steiner trees, either unbuffered or buffered. However, introducing redundant paths into the topology (which leads to non-tree) may significantly improve timing performance as well as tolerance to open faults and variations. These advantages are particularly appealing for timing critical net routings in nanoscale VLSI designs where interconnect delay is a performance bottleneck and variation effects are increasingly remarkable. We propose Steiner network construction heuristics which can generate either tree or non-tree with different slack-wirelength tradeoff, and handle both long path and short path constraints. We also propose heuristics for simultaneous Steiner network construction and buffering, which may provide further improvement in slack and resistance to variations. Furthermore, incremental non-tree delay update techniques are developed to facilitate fast Steiner network evaluations. Extensive experiments in different scenarios show that our heuristics usually improve timing slack by hundreds of pico seconds compared to traditional approaches. When process variations are considered, our heuristics can significantly improve timing yield because of nominal slack improvement and delay variability reduction.
Shiyan Hu 0001, Qiuyang Li, Jiang Hu 0001, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Characterizing Multistage Nonlinear Drivers and Variability for Accurate Timing and Noise Analysis
abstract
Nanoscale device characteristics and noise coupling have rendered traditional waveform-based gate delay models increasingly difficult to adopt. While the widely adopted delay models are built upon the assumption of simple ramp-like signal waveforms, realistic signal shapes in nanoscale designs can be far more complex. The need for considering process-voltage-temperature (PVT) variations imposes further accuracy requirement on gate models. We present a parameterizable waveform independent gate model (PWiM) where no assumption is made upon the input waveforms. The PWiM model is constructed by encapsulating the driver's intrinsic nonlinear dc and dynamic characteristics, which are important to model for complex signal waveforms, via novel and yet easy-to-implement characterization steps. As such, PWiM can provide near-SPICE accuracy for input signals that significantly deviate from simple ramps. While recently developed current-based models can only be applied to single channel-connected component, PWiM can work for multistage cells leading to improved library compactness and analysis efficiency. Our experiments have indicated that the proposed driver model not only provides up to two orders of magnitude speedups over SPICE for delay and noise analysis, it also offers accurate assessment of performance variability introduced by process and environmental variations.
Peng Li 0001, Emrah Acar
IEEE Trans. Very Large Scale Integr. Syst.1
2007 Fast Variational Interconnect Delay and Slew Computation Using Quadratic Models
abstract
Interconnects constitute a dominant source of circuit delay for modern chip designs. The variations of critical dimensions in modern VLSI technologies lead to variability in interconnect performance that must be fully accounted for in timing verification. However, handling a multitude of inter-die/intra-die variations and assessing their impacts on circuit performance can dramatically complicate the timing analysis. In this paper, a practical interconnect delay and slew analysis technique is presented to facilitate efficient evaluation of wire performance variability. By harnessing a collection of computationally efficient procedures and closed-form formulas, process variations are directly mapped into the variability of the output delay and slew. An efficient method based on sensitivity analysis is implemented to calculate driving point models under variations for gate-level timing analysis. The proposed adjoint technique not only provides statistical performance variations of the interconnect network under analysis, but also produces delay and slew expressions parameterized in the underlying process variations in a quadratic parametric form. As such, it can be harnessed to enable statistical timing analysis while considering important statistical correlations. Our experimental results have indicated that the presented analysis is accurate regardless of location of sink nodes and it is also robust over a wide range of process variations.
Xiaoji Ye, Frank Liu 0001, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2006 Steiner network construction for timing critical nets
abstract
Conventionally, signal net routing is almost always implemented asSteiner trees. However, non-tree topology is often superior on timing performance as well as tolerance to open faults and variations. These advantages are particularly appealing for timing critical net routings in nano-scale VLSI designs where interconnect delay is a performance bottleneck and variation effects are increasingly remarkable. We propose Steiner network construction heuristics which can generate either tree or non-tree with different slack-wirelength tradeoff, and handle both long path and short path constraints. Incremental non-tree delay update techniques are developed to facilitate fast Steiner network evaluations. Extensive experiments in different scenarios show that our heuristics usually improve timing slack by hundreds of pico seconds compared to traditional tree approaches.
Shiyan Hu 0001, Qiuyang Li, Jiang Hu 0001, Peng Li 0001
DAC4
2006 Model order reduction of linear networks with massive ports via frequency-dependent port packing
abstract
Model order reduction has been a driving force for reducing analysis complexity of VLSI systems containing large linear networks. However, most existing reduction techniques are only applicable to networks with a small number of ports, failing to fulfill an even stronger need of reducing massively interconnected subsystems such as power grids and wide buses. In this paper, a port packing scheme is presented wherein the correlation between circuit ports is explored in a frequency-dependent manner. In the proposed McPack (multiport circuit packing) algorithm, port packing is combined with a practical realization of the recently developed tangential interpolation scheme for model reduction. McPack performs feasible moment matching for networks with many ports in the sense of tangential interpolation. With guaranteed passivity, extensibility to multi-point expansion as well as comparable complexity, McPack systematically introduces frequency-domain port packing into the existing projection-based model order reduction framework. For several large networks with high port count, the presented algorithm is shown to be significantly more accurate than the standard block-moment matching algorithm as well as other recently developed alternative.
Peng Li 0001, Weiping Shi
DAC1
2006 Lookup table based simulation and statistical modeling of Sigma-Delta ADCs
abstract
Sigma-Delta (ΕΔ) ADCs have been widely adopted in data conversion applications due to the good performance. However, oversampling and complex circuit behavior render the simulation of these designs prohibitively time consuming. In this paper, a lookup table (LUT) based modeling technique is presented for efficient analysis of ΕΔ ADCs. In the proposed approach, various transistor-level circuit non-idealities are systematically characterized at the building-block level and the complete ADC is simulated much more efficiently using these table models. As such, our approach can provide up to four orders of magnitude runtime speedup over SPICE-like simulators, hence significantly shortening the CPU time required for evaluating system performances such as SNDR (signal-noise-distortion-ratio). The proposed LUT modeling technique is further extended to assess performance variations due to parameter fluctuations. The resulting parameterized LUT modeling technique not only facilitates scalable performance variation analysis of complex ΕΔ ADC designs, but also allows us to feasibly extract statistical performance correlation models for low-cost test solutions.
Peng Li 0001
DAC2
2006 Performance-oriented statistical parameter reduction of parameterized systems via reduced rank regression
abstract
Process variations in modern VLSI technologies are growing in both magnitude and dimensionality. To assess performance variability, complex simulation and performance models parameterized in a high-dimensional process variation space are desired. However, the high parameter dimensionality, imposed by a large number of variation sources encountered in modern technologies, can introduce significant complexion in circuit analysis and may even render performance variability analysis completely intractable. We address the challenge brought by high-dimensional process variations via a new performance-oriented parameter dimension reduction technique. The basic premise behind our approach is that the dimensionality of performance variability is determined not only by the statistical characteristics of the underlying process variables, but also by the structural information imposed by a given design. Using the powerful reduced rank regression (RRR) and its extension as a vehicle for variability modeling, we are able to systematically identify statistically significant reduced parameter sets and compute not only reduced-parameter but also reduced-parameter-order models that are far more efficient than what was possible before. For a variety of interconnect modeling problems, it is shown that the proposed parameter reduction technique can provide more than one order of magnitude reduction in parameter dimensionality. Such parameter reduction immediately leads to reduced simulation cost in sampling-based performance analysis, and more importantly, highly efficient parameterized interconnect reduced order models. As a general parameter dimension reduction methodology, it is anticipated that the proposed technique is broadly applicable to a variety of statistical circuit modeling problems, thereby offering a useful framework for controlling the complexity of statistical circuit analysis.
Peng Li 0001
ICCAD2
2006 Combinatorial algorithms for fast clock mesh optimization
abstract
We present a fast and efficient combinatorial algorithm to simultaneously identify the candidate locations as well as the sizes of the buffers driving a clock mesh. Due to the high redundancy, a mesh architecture offers high tolerance towards variation in the clock skew. However, such a redundancy comes at the expense of mesh wire length and power dissipation. Based on survivable network theory, we formulate the problem to reduce the clock mesh by retaining only those edges that are critical to maintain redundancy. Such a formulation offers designer the option to trade-off between power and tolerance to process variations. Experimental results indicate that our techniques can result in power savings up to 28% with less than 4% delay penalty.
Ganesh Venkataraman, Jiang Hu 0001, Peng Li 0001
ICCAD4
2006 Practical variation-aware interconnect delay and slew analysis for statistical timing verification
abstract
Interconnects constitute a dominant source of circuit delay for modern chip designs. The variations of critical dimensions in modern VLSI technologies lead to variability in interconnect performance that must be fully accounted for in timing verification. However, handling a multitude of inter-die/intra-die variations and assessing their impacts on circuit performance can dramatically complicate the timing analysis. In this paper, a practical interconnect delay and slew analysis technique is presented to facilitate efficient evaluation of wire performance variability. By harnessing a collection of computationally efficient procedures and closed-form formulas, process and input signal variations are directly mapped into the variability of the output delay and slew. Since our approach produces delay and slew expressions parameterized in the underlying process variations, it can be harnessed to enable statistical timing analysis while considering important statistical correlations. Our experimental results have indicated that the presented analysis is accurate regardless of location of sink nodes and it is also robust over a wide range of process variations.
Xiaoji Ye, Peng Li 0001, Frank Liu 0001
ICCAD2
2006 Statistical Sampling-Based Parametric Analysis of Power Grids
abstract
A statistical sampling-based parametric analysis is presented for analyzing large power grids in a "localized" fashion. By combining random walks with the notion of "importance sampling," the proposed technique is capable of efficiently computing the impacts of multiple circuit parameters on selected network nodes. A "new localized" sensitivity analysis is first proposed to solve not only the nominal node response but also its sensitivities with respect to multiple parameters using a single run of the random walks algorithm. This sampling-based technique is further extended from the first-order sensitivity analysis to a more general second-order analysis. By exploiting the natural spatial locality inherent in the proposed algorithm formulation, the second-order analysis can be performed efficiently even for a large number of global and local variation sources. The theoretical convergence properties of three importance sampling estimators for power grid analysis are presented, and their effectiveness is compared experimentally on several examples. The superior performance of the proposed technique is demonstrated by analyzing several large power grids under process and current loading variations to which the application of the existing brute-force simulation techniques becomes completely infeasible
Peng Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2006 IC thermal simulation and modeling via efficient multigrid-based approaches
abstract
The ever-increasing power consumption and packaging density of integrated systems creates on-chip temperatures and gradients that can have a substantial impact on performance and reliability. While it is conceptually understood that a thermal equivalent circuit can be constructed to characterize the temperature gradients across the chip, direct and iterative solutions of the corresponding three-dimensional (3-D) equations are often intractable for a full-chip analysis. Integrated circuit (IC)-specific multigrid (MG) techniques for fast chip level thermal steady-state and transient simulation are proposed. This approach avoids an explicit construction of the matrix problem that is intractable for most full-chip problems. Specific MG treatments are proposed to cope with the strong anisotropy of the full-chip thermal problem that is created by the vast difference in material thermal properties and chip geometries. Importantly, this paper demonstrates that only with careful thermal modeling assumptions and appropriate choices for grid hierarchy, MG operators, and smoothing steps across grid points can a full-chip thermal problem be accurately and efficiently analyzed. This paper further speeds up the large thermal transient simulations by incorporating reduced-order thermal models that can be efficiently extracted under the same MG framework. The experiments carried out in this work have shown that the proposed methodology provides sufficient efficiency in both runtime and memory usage.
Peng Li 0001, Lawrence T. Pileggi, Mehdi Asheghi, Rajit Chandra
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2005 Power grid simulation via efficient sampling-based sensitivity analysis and hierarchical symbolic relaxation
abstract
On-chip supply networks are playing an increasingly important role for modern nanometer-scale designs. However, the ever growing sizes of power grids make the analysis problem extremely difficult thereby introducing severe challenges in design and optimization. The inherent analysis complexity calls for innovations in simulation techniques that must provide appropriate accuracy, efficiency as well as the tradeoff thereof to aid design verification and optimization. In this paper, we first present a sampling-based sensitivity analysis by employing the notation of importance sampling in a Monte Carlo based circuit simulation framework. This technique allows the extraction of multi-parameter sensitivities for the node voltages of interest in the same Monte Carlo runs that are used for computing the nominal voltage values. For more efficient nonstructured whole-grid solution approaches, we further introduce a new direct solution method by embedding symbolic relaxation steps in a hierarchical fashion. As a direct method, the proposed hierarchical symbolic relaxation is suitable to both dc and transient analyses. Circuit examples are included to demonstrate the efficacy of the proposed techniques.
Peng Li 0001
DAC1
2005 Specification Test Compaction for Analog Circuits and MEMS
abstract
Testing a non-digital integrated system against all of its specifications can be quite expensive due to the elaborate test application and measurement setup required. We propose to eliminate redundant tests by employing /spl epsi/-SVM based statistical learning. The application of the proposed methodology to an operational amplifier and a MEMS accelerometer reveal that redundant tests can be statistically identified from a complete set of specification-based tests, with negligible error. Specifically, after eliminating five of eleven specification-based tests for an operational amplifier, the defect escape and yield loss is small at 0.6% and 0.9%, respectively. For the accelerometer, defect escape of 0.2% and yield loss of 0.1% occurs when the hot and cold tests are eliminated. For the accelerometer, this level of compaction would reduce test cost by more than half.
Sounil Biswas, Peng Li 0001, R. D. (Shawn) Blanton, Lawrence T. Pileggi
DATE2
2005 Modeling Interconnect Variability Using Efficient Parametric Model Order Reduction
abstract
Assessing IC manufacturing process fluctuations and their impacts on IC interconnect performance has become unavoidable for modern DSM designs. However, the construction of parametric interconnect models is often hampered by the rapid increase in computational cost and model complexity. In this paper we present an efficient yet accurate parametric model order reduction algorithm for addressing the variability of IC interconnect performance. The efficiency of the approach lies in a novel combination of low-rank matrix approximation and multi-parameter moment matching. The complexity of the proposed parametric model order reduction is as low as that of a standard Krylov subspace method when applied to a nominal system. Under the projection-based framework, our algorithm also preserves the passivity of the resulting parametric models.
Peng Li 0001, Frank Liu 0001, Xin Li 0001, Lawrence T. Pileggi, Sani R. Nassif
DATE1
2005 Variational analysis of large power grids by exploring statistical sampling sharing and spatial locality
abstract
We propose a parametric random walk algorithm to facilitate a feasible evaluation of a few critical network nodes under the influence of a large number of variation sources in a power grid. By combining statistical sampling sharing with random walks, we devise an efficient localized sensitivity analysis for large power distribution networks such that the analysis can be conducted without solving the complete network. We further show that this sampling-based parametric analysis can be extended from the first order sensitivity analysis to a more accurate second order analysis. By exploiting the natural spatial locality inherent in our algorithm formulation, the second order parametric analysis can be conducted very efficiently even for a large number of global and local variation sources. The proposed approach is demonstrated by analyzing large power grids under the influence of process and current loading variations to which the application of the standard brutal-force circuit simulation becomes completely infeasible. Our results have demonstrated the superior performance of the proposed algorithm both in terms of accuracy and runtime.
Peng Li 0001
ICCAD1
2005 Parameterized interconnect order reduction with explicit-and-implicit multi-parameter moment matching for inter/intra-die variations
abstract
In this paper we propose a novel parameterized interconnect order reduction algorithm, CORE, to efficiently capture both inter-die and intra-die variations. CORE applies a two-step explicit-and-implicit scheme for multiparameter moment matching. As such, CORE can match significantly more moments than other traditional techniques using the same model size. In addition, a recursive Arnoldi algorithm is proposed to quickly construct the Krylov subspace that is required for parameterized order reduction. Applying the recursive Arnoldi algorithm significantly reduces the computation cost for model generation. Several RC and RLC interconnect examples demonstrate that CORE can provide up to 10/spl times/ better modeling accuracy than other traditional techniques, while achieving smaller model complexity (i.e. size). It follows that these interconnect models generated by CORE can provide more accurate simulation result with cheaper simulation cost, when they are utilized for gate-interconnect co-simulation.
Xin Li 0001, Peng Li 0001, Lawrence T. Pileggi
ICCAD2
2005 Practical techniques to reduce skew and its variations in buffered clock networks
abstract
Clock skew is becoming increasingly difficult to control due to variations. Link based non-tree clock distribution is a cost-effective technique for reducing clock skew variations. However, previous works based on this technique were limited to unbuffered clock networks and neglected spatial correlations in the experimental validation. In this work, we overcome these shortcomings and make the link based non-tree approach feasible for realistic designs. The short circuit risk and multi-driver delay issues in buffered non-tree clock networks are investigated. Our approach is validated with SPICE based Monte Carlo simulations, considering spatial correlations among variations. The experimental results show that our approach can reduce the maximal skew by 47%, improve the skew yield from 15% to 73% on average with a decrease on the total wire and buffer capacitance.
Ganesh Venkataraman, Nikhil Jayakumar, Jiang Hu 0001, Peng Li 0001, Sunil P. Khatri, Anand Rajaram, Patrick McGuinness, Charles J. Alpert
ICCAD4
2005 A Waveform Independent Gate Model for Accurate Timing Analysis
abstract
In nanoscale regime, it is becoming increasingly difficult to model signal shapes using simple ramp-like waveforms due to various noise coupling effects. We present an accurate waveform independent gate (WiM) model without any assumption of signal waveforms. Our model can be applied to arbitrary gate inputs while maintaining excellent near-SPICE accuracy. The application of the proposed gate modeling technique is demonstrated under the context of the gate-level timing simulation.
Peng Li 0001, Emrah Acar
ICCD1
2005 Temperature-Dependent Optimization of Cache Leakage Power Dissipation
abstract
Leakage power consists of an increasing portion of the total power consumption for modern IC designs. Due to the strong inter-dependency between leakage and temperature, it becomes imperative to consider the thermal effects while optimizing the leakage power. In this paper, we present a temperature-dependent optimization methodology for on-chip caches. By integrating fast yet accurate coupled thermal-leakage simulations into an optimization flow, we are able to optimally tradeoff between the cache performance and leakage power while considering realistic on-chip temperature distribution. Our analysis indicates that for future memory intensive designs, the lack of chip temperature information can cause a significant error in the leakage power estimation, thus leading to non-optimal cache designs. Our results further imply that the optimization of cache performance and leakage power shall be attacked as part of the whole system design task in which chip-level floor planning and its thermal impacts are fully addressed.
Peng Li 0001, Yangdong Deng, Lawrence T. Pileggi
ICCD1
2005 Compact reduced-order modeling of weakly nonlinear analog and RF circuits
abstract
A compact nonlinear model order-reduction method (NORM) is presented that is applicable for time-invariant and periodically time-varying weakly nonlinear systems. NORM is suitable for model order reduction of a class of weakly nonlinear systems that can be well characterized by low-order Volterra functional series. The automatically extracted macromodels capture not only the first-order (linear) system properties, but also the important second-order effects of interest that cannot be neglected for a broad range of applications. Unlike the existing projection-based reduction methods for weakly nonlinear systems, NORM begins with the general matrix-form Volterra nonlinear transfer functions to derive a set of minimum Krylov subspaces for order reduction. Moment matching of the nonlinear transfer functions by projection of the original system onto this set of minimum Krylov subspaces leads to a significant reduction of model size. As we will demonstrate as part of comparison with existing methods, the efficacy of model reduction for weakly nonlinear systems is determined by the achievable model compactness. Our results further indicate that a multipoint version of NORM can substantially improve the model compactness for nonlinear system reduction. Furthermore, we show that the structure of the nonlinear system can be exploited to simplify the reduced model in practice, which is particularly effective for circuits with sharp frequency selectivity. We demonstrate the practical utility of NORM and its extension for macromodeling weakly nonlinear RF communication circuits with periodically time-varying behavior.
Peng Li 0001, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2004 A frequency relaxation approach for analog/RF system-level simulation
abstract
The increasing complexity of today's mixed-signal integrated circuits necessitates both top-down and bottom-up system-level verification. Time-domain state-space modeling and simulation approaches have been successfully applied for such purposes (e.g. Simulink); however, analog circuits are often best analyzed in the frequency domain. Circuit-level analyses, such as harmonic balance, have been successfully extended to the frequency domain [2], but these algorithms are impractical for simulating large systems with wide-band input and noise signals. In this paper we proposed a frequency-domain approach for analog/RF system-level simulation that is capable of capturing various second order effects (e.g. nonlinearity, noise, etc.) for both time-invariant and time-varying systems with wide-band inputs. The simulator directly evaluates the frequency domain response at each node via a relaxation scheme that is proven to be convergent under typical circuit conditions. Our experimental results demonstrate the accuracy and efficiency of the proposed simulator under various wide-band input and noise excitations.
Xin Li 0001, Yang Xu 0017, Peng Li 0001, Padmini Gopalakrishnan, Lawrence T. Pileggi
DAC3
2004 Efficient harmonic balance simulation using multi-level frequency decomposition
abstract
Efficient harmonic balance (HB) simulation provides a useful tool for the design of RF and microwave integrated circuits. For practical circuits that can contain strong nonlinearities, however, HB problems cannot be solved reliably or efficiently using conventional techniques. Various preconditioning techniques have been proposed to facilitate a robust and efficient analysis based on Krylov subspace linear solvers. In This work we introduce a multi-level frequency domain preconditioner based on a hierarchical frequency decomposition approach. At each Newton iteration, we recursively solve a set of smaller problems to provide an effective preconditioner for the large linearized HB problem. Compared to the standard single-level block diagonal preconditioner, our experiments indicate that our approach provides a more robust, memory efficient solution while offering a 2-9/spl times/ speedup for several strongly nonlinear HB problems in our experiments.
Peng Li 0001, Lawrence T. Pileggi
ICCAD1
2004 Efficient full-chip thermal modeling and analysis
abstract
The ever-increasing power consumption and packaging density of integrated systems creates on-chip temperatures and gradients that can have a substantial impact on performance and reliability. While it is conceptually understood that a thermal equivalent circuit can be constructed to characterize the temperature gradients across the chip, direct and iterative solutions of the corresponding 3D equations are often intractable for a full-chip analysis. Multigrid accelerated iterative methods can be applied to solve the equivalent circuit problem that is provably symmetric positive definite; however, explicitly building the matrix problem is intractable for most full-chip problems. In This work we present a multigrid iterative approach for the full-chip thermal analysis which does not require explicit construction of the equivalent circuit matrix. We propose specific multigrid treatments to cope with the strong anisotropy of the full-chip thermal problem that is created by the vast difference in material thermal properties and chip geometries. Importantly, we demonstrate that only with careful thermal modeling assumptions and appropriate choices for grid hierarchy, multigrid operators and smoothing steps across grid points, can we accurately and efficiently analyze a full-chip thermal problem. Experimental results demonstrate the efficacy of the proposed multigrid methodology. Our prototyped thermal simulator is able to solve a steady-state problem with more than 10 million unknowns in 125 CPU seconds with a peak memory usage of 231 mega bytes.
Peng Li 0001, Lawrence T. Pileggi, Mehdi Asheghi, Rajit Chandra
ICCAD1
2003 A frequency separation macromodel for system-level simulation of RF circuits
abstract
In this paper we propose a frequency-separation methodology to generate system-level macromodels for analog and RF circuits. The proposed macromodels are similar in form to those based on Volterra kernel calculations, but are much simpler in terms of characterization and overall model complexity, and can be derived from existing device models. This simplicity is realized by applying some basic assumptions on the form of the input excitations, and via separation of the nonlinearities from the dynamic behavior. In addition, by further separating the ideal model functionality, this macromodel is applicable to strongly nonlinear components such as mixers. While time-varying Volterra series models have been proposed for mixers with a fixed local oscillation (LO) signal, the proposed frequency separation model is completely general and can capture the variations of the LO input during a system-level simulation. The proposed macromodels are demonstrated in a system-level simulation tool based on Simulink for efficient evaluation of the entire RF system and associated components. A GSM receiver system in 0.25μm CMOS process is used to demonstrate the efficacy of these macromodels in our system-level simulation environment.
Xin Li 0001, Peng Li 0001, Yang Xu 0017, Robert Dimaggio, Lawrence T. Pileggi
ASP-DAC2
2003 Nonlinear distortion analysis via linear-centric models
abstract
An efficient distortion analysis methodology is presented for analog and RF circuits that utilizes linear-centric circuit models to generate individual distortion contributions due to the various circuit nonlinearities. The per-nonlinearity distortion results are obtained via a straightforward post-simulation step that is simpler and more efficient than the Volterra series based approaches and do not require the high order device model derivatives. For this reason the order of analysis can be significantly higher than that for a Volterra series implementation while fully accounting for all nonlinearity effects. The proposed methodology is not restricted to weakly nonlinear circuits, but can also analyze per-nonlinearity distortion for active switching mixers and switch capacitor circuits when they are modeled as periodically time-varying weakly nonlinear systems. While Volterra series have also been attempted for this same class of circuits, the requirement of device models for all of the high order model derivatives makes such analysis somewhat impractical. The proposed methodology provides important design insights regarding the relationships between design parameters and circuit linearity, hence the overall system performance. Circuit examples are used to demonstrate the efficacy of the proposed approach, and interesting insights are observed for RF switching mixers in particular.
Peng Li 0001, Lawrence T. Pileggi
ASP-DAC1
2003 Analog and RF circuit macromodels for system-level analysis
abstract
Design and validation of mixed-signal integrated systems require system-level model abstractions. Generalized Volterra series based models have been successfully applied for analog and RF component macromodels, but their complexity can sometimes limit their utility for time-varying systems and large circuits with complex device models or numerous parasitics. In this paper we propose simple and efficient analog and RF circuit macromodels that provide accurate model abstractions for large, complex time-varying circuits over frequency bands of interest. By starting with the system-level block diagram model structures and focusing on the narrow RF bands, the proposed macromodels can efficiently capture the nonlinear behavior as well as the impact of RLC coupling parasitics via compact reduced-order model forms. While the macromodel can trade accuracy for simplicity in terms of the number of frequency expansion points, we find that expansion about one frequency point provides the accuracy required for system-level analysis of most RF and narrow-band analog components. The macromodel form corresponds to block diagram structures that are easily incorporated into our system-level simulation tool based on Simulink.
Xin Li 0001, Peng Li 0001, Yang Xu 0017, Lawrence T. Pileggi
DAC2
2003 NORM: compact model order reduction of weakly nonlinear systems
abstract
This paper presents a compact Nonlinear model Order Reduction Method (NORM) that is applicable for time-invariant and time-varying weakly nonlinear systems. NORM is suitable for reducing a class of weakly nonlinear systems that can be well characterized by low order Volterra functional series. Unlike existing projection based reduction methods [6]-[8], NORM begins with the general matrix-form Volterra nonlinear transfer functions to derive a set of minimum Krylov subspaces for order reduction. Direct moment matching of the nonlinear transfer functions by projection of the original system onto this set of minimum Krylov subspaces leads to a significant reduction of model size. As we will demonstrate as part of our comparison with existing methods, the efficacy of model order for weakly nonlinear systems is determined by the extend to which models can be reduced. Our results further indicate that a multiple-point version of NORM can substantially reduce the model size and approach the ultimate model compactness that is achievable for nonlinear system reduction. We demonstrate the practical utility of NORM for macro-modeling weakly nonlinear RF circuits with time-varying behavior.
Peng Li 0001, Lawrence T. Pileggi
DAC1
2003 Noise Macromodel for Radio Frequency Integrated Circuits
abstract
Noise performance is a critical analog and RF circuit design constraint, and can impact the selection of the IC system-level architecture. It is therefore imperative that some model of the noise is represented at the highest levels of abstraction during the design process. In this paper we propose a noise macromodel for analog circuits and demonstrate it by way of implementation in a system level simulator based on MATLAB. We also explain our process of macromodel extraction via reformulation of frequency-domain noise analysis results, and the corresponding steps of model order reduction. The results demonstrate the efficacy of this macromodel for frequency domain system level simulation.
Yang Xu 0017, Xin Li 0001, Peng Li 0001, Lawrence T. Pileggi
DATE3
2003 A Hybrid Approach to Nonlinear Macromodel Generation for Time-Varying Analog Circuits
Peng Li 0001, Xin Li 0001, Yang Xu 0017, Lawrence T. Pileggi
ICCAD1
2003 Efficient per-nonlinearity distortion analysis for analog and RF circuits
abstract
An efficient distortion analysis methodology is presented for analog and RF circuits that utilizes linear-centric circuit models to generate individual distortion contributions due to each nonlinear component in a circuit. The per-nonlinearity distortion results are obtained via a straightforward post-simulation step that is simpler and more efficient than the Volterra series-based approaches and does not require high-order device-model derivatives. For this reason, the order of analysis can be significantly higher than that for a Volterra series-based implementation while fully accounting for all distortion effects using most existing device models. Moreover, the proposed methodology can also analyze per-nonlinearity distortion for active switching mixers and switch capacitor circuits when they are modeled as periodically time-varying weakly nonlinear systems. The proposed methodology provides important design insights regarding the relationships between design parameters and circuit linearity, hence, the overall system performance. Circuit examples are used to demonstrate the efficacy of the proposed approach, and interesting insights are observed for RF switching mixers in particular.
Peng Li 0001, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2002 A Linear-Centric Modeling Approach to Harmonic Balance Analysis
abstract
In this paper, we propose a new harmonic balance simulation methodology based on a linear-centric modeling approach. A linear circuit representation of the nonlinear devices and associated parasitics is used along with corresponding time and frequency domain inputs to solve for the nonlinear steady-state response via successive chord (SC) iterations. For our circuit examples, this approach is shown to be up to 60/spl times/ more run-time efficient than traditional Newton-Raphson (N-R) based iterative methods, while providing the same level of accuracy. This SC-based approach converges as reliably as the N-R approaches, including for circuit problems which cause alternative relaxation-based harmonic balance approaches to fail. The efficacy of this linear-centric methodology further improves with increasing model complexity, the inclusion of interconnect parasitics and other analyses that are otherwise difficult with traditional nonlinear models.
Peng Li 0001, Lawrence T. Pileggi
DATE1