Sheldon X.-D. Tan

dblp:t/SXDTan · also Xiang-Dong Tan · DBLP profile ↗
← Back
221ranked-venue papers
17as first author
35since 2021 · last 2026
0000-0003-2119-6869ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 217 · 17 first-author · 34 since 2021Software engineering, systems software and programming languages · 17 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 WarPGNN: A Parametric Thermal Warpage Analysis Framework with Physics-aware Graph Neural Network
Haotian Lu 0002, Jincong Lu, Sachin Sachdeva, Sheldon X.-D. Tan
ISLPED4
2025 Hybrid Temporal Computing for Lower Power Hardware Accelerators
abstract
In this paper, we propose a new hybrid temporal computing (HTC) framework that leverages both pulse rate and temporal data encoding to design ultra-low energy hardware accelerators. Our approach is inspired by the recently proposed temporal computing, or race logic, which encodes data values as single delays, leading to significantly lower energy consumption due to minimized signal switching. The new HTC framework overcomes the inherent limitations of race logic by encoding signals in both temporal and pulse rate formats for multiplication and in temporal format for propagation. We demonstrate how HTC multiplication is performed for both unipolar and bipolar data encoding while consuming reduced switching energy. Additionally, we implement two widely used hardware accelerators: a Finite Impulse Response (FIR) filter and a Discrete Cosine Transform (DCT)/iDCT. Experimental results show that compared to the CBSC MAC, the HTC MAC reduces power consumption by 45.2% and area footprint by 50.13%. Compared to the CBSC design, the HTC-based FIR filter reduces power consumption by 36.61% and area cost by 45.85%. The HTC-based DCT filter retains the quality of the original image with a decent PSNR, while consuming 23.34% less power and occupying 18.20% less area than the CBSC MAC-based DCT filter.
Maliha Tasnim, Sachin Sachdeva, Sheldon X.-D. Tan
ASP-DAC4
2025 Power Map Characterization and Modeling for Commercial CPU/GPUs Considering Temperature Dependence
abstract
In this paper, we address the challenge of accurate full-chip power mapping for commercial off-the-shelf CPU and GPU processors, explicitly considering temperature dependence. It is well known that both dynamic and leakage power are strongly temperature-dependent; however, existing power estimation methods for real chips often neglect this critical factor. To mitigate this, we characterize temperature-dependent spatial power maps for the first time on commercial processors, including the AMD Radeon RX 6400 GPU and Qualcomm Snapdragon 680 (SM6225) CPU. Using a back-side cooling infrared (IR) thermal imaging system, we capture full-chip thermal maps under different cooling conditions while running identical workloads. These thermal maps are converted into power maps using first-principles-based methods. By repeating this process across varying cooling environments, we collect power maps corresponding to different average chip temperatures. Our experimental results confirm that both total power and spatial power distributions vary significantly with cooling conditions, even under the same workload. We then train machine learning models using real-time performance and utilization metrics—collected via AMD Adrenalin Edition and Qualcomm Snapdragon Profiler—to capture these thermal effects. Two deep neural network architectures are explored: a transformer-based model, ChipPowerMap, and a CNN-based decoder model. We compare their performance in accurately predicting temperature-aware full-chip power maps. Numerical results highlight the effectiveness of ChipPowerMap in achieving highly accurate thermal map predictions, boasting an RMSE of only 67.88mW/mm2or 0.97% of the full-scale error. It also outperforms the CNN-based method by 1.62x in terms of accuracy on average. Besides, the proposed model offers real-time estimation with a rapid speed of 25ms on the target chip.
Jincong Lu, Sachin Sachdeva, Haotian Lu 0002, Sheldon X.-D. Tan
ISLPED4
2025 PISOV: Physics-Informed Separation of Variables Solvers for Full-Chip Thermal Analysis
abstract
Thermal issues are becoming increasingly critical due to rising power densities in high-performance chip design. The need for fast and precise full-chip thermal analysis is evident. Although machine learning (ML)-based methods have been widely used in thermal simulation, their training time remains a challenge. In this article, we proposed a novel physics-informed separation of variables solver (PISOV) to significantly reduce training time for fast full-chip thermal analysis. Inspired by the recently proposed ThermPINN, we employ a least-square regression method to calculate the unknown coefficients of the cosine series. The proposed PISOV method combines physics-informed neural network (PINN) and separation of variables (SOVs) methods. Due to the matrix-solving method of PISOV, its speed is much faster than that of ThermPINN. On top of PISOV, we parameterize effective convection coefficients and power values for surrogate model-based uncertainty quantification (UQ) analysis by using neural networks, a task that cannot be accomplished by the SOV method. In the parameterized PISOV, we only need to calculate once to obtain all parameterized results of the hyperdimensional partial differential equations. Additionally, we study the impact of sampling methods (such as grid, uniform, Sobol, Latin hypercube sampling (LHS), Halton, and Hammersly) and hybrid sampling methods on the accuracy of PISOV and parameterized PISOV. Numerical results show that PISOV can achieve a speedup of$245\times $, and$10^{4}\times $over ThermPINN, and PINN, respectively. Among different sampling methods, the Hammersley sampling method yields the best accuracy.
Liang Chen 0025, Wenxing Zhu, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 BPINN-EM: Fast Stochastic Analysis of Electromigration Damage using Bayesian Physics-Informed Neural Networks
abstract
Electromigration (EM) induced aging and degradation in interconnect wires is inherently a stochastic process, with lifetime typically measured in terms of mean time to failure at both wire and circuit levels. However, existing approaches still incur high computing costs, as computing both means and variances is generally expensive. In this work, we propose a novel fast variational analysis framework to tackle the challenges of stochastic estimation of EM stress evolution in multi-segment interconnect wires. We utilize Bayesian networks in conjunction with the recently introduced hierarchical (two-step) physics-informed neural networks (PINN). The resulting method, termed BPINN-EM, enables rapid variational stress analysis of metal wires by leveraging the robust uncertainty quantification capability of Bayesian networks with expedited training over small dataset. Moreover, we devise BPINN-EM to incorporate Bayesian networks only in the first stage of the hierarchical PINN, thereby circumventing the need for sampling across the entire PINN level during training and significantly reducing training costs. Our results on several general multi-segment interconnect structure demonstrate that the proposed BPINN-EM approach is much more efficient than conventional baselines and state-of-the-art algorithms. Compared to a Monte Carlo simulator implemented in COMSOL, BPINN-EM offers a 240× speedup. Moreover, compared to the recently proposed EMSpice simulated by the Monte Carlo method, the new method provides more than an 85× speedup with almost no loss of accuracy.
Subed Lamichhane, Mohammadamir Kavousi, Sheldon X.-D. Tan
ICCAD3
2024 Exploring BTI aging effects on spatial power density and temperature profiles of VLSI chips
Sachin Sachdeva, Jincong Lu, Hussam Amrouch, Sheldon X.-D. Tan
Integr.4
2024 Fast and Scaled Counting-Based Stochastic Computing Divider Design
abstract
This article presents novel designs for stochastic computing (SC)-based dividers, which promise low latency, high energy efficiency as well as high accuracy for error-tolerant arithmetic operations. We first introduce CBDIV, which is based on the recently proposed counter-based SC concept and correlation based SC to perform division. Then we introduce FSCDIV, which further improves the accuracy of CBDIV by applying a scaling strategy and mitigating the latency by optimizing the counting scheme. The FSCDIV will equally scale up the divider and dividend before the division process, and thereby avoid large relative error when both input values of the divider and dividend are small. The proposed fast counting method, accelerates FSCDIV by counting new bit pair (0-1 pair) among only half of the stochastic number bitstream instead of the entire bitstream, resulting in almost half of the counting latency and one-fourth of the overall division operation latency. The experimental results demonstrate that the proposed CBDIV, implemented in a 32nm technology node, outperforms state-of-the-art works by 77.8% in accuracy, 37.1% in delay, 21.5% in area, 50.6% in area delay product (ADP), and 25.9% in power consumption. Compared to the fixed-point division baseline, CBDIV also achieves a 31.9% reduction in energy consumption and is more energy-efficient than existing SC-based dividers for binary inputs and outputs required in efficient image processing implementations. Moreover, we demonstrate that FSCDIV improves delay by 56.4%, ADP by 16.0%, energy consumption by 45.0%, and accuracy by 61.2%. We also evaluate CBDIV and FSCDIV designs in a contrast stretch image processing workload, and the results show that the proposed designs can improve the image quality by up to 18.3 dB on average when compared to state-of-the-art works.
Shuyuan Yu, Maliha Tasnim, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Learning Based Spatial Power Characterization and Full-Chip Power Estimation for Commercial TPUs
abstract
In this paper, we propose a novel approach for the real-time estimation of chip-level spatial power maps for commercial Google Coral M.2 TPU chips based on a machine-learning technique for the first time. The new method can enable the development of more robust runtime power and thermal control schemes to take advantage of spatial power information such as hot spots that are otherwise not available. Different from the existing commercial multi-core processors in which real-time performance-related utilization information is available, the TPU from Google does not have such information. To mitigate this problem, we propose to use features that are related to the workloads of running different deep neural networks (DNN) such as the hyperparameters of DNN and TPU resource information generated by the TPU compiler. The new approach involves the offline acquisition of accurate spatial and temporal temperature maps captured from an external infrared thermal imaging camera under nominal working conditions of a chip. To build the dynamic power density map model, we apply generative adversarial networks (GAN) based on the workload-related features. Our study shows that the estimated total powers match the manufacturer's total power measurements extremely well. Experimental results further show that the predictions of power maps are quite accurate, with the RMSE of only 4.98mW/mm2, or 2.6% of the full-scale error. The speed of deploying the proposed approach on an Intel Core i7-10710U is as fast as 6.9ms, which is suitable for real-time estimation.
Jincong Lu, Wentian Jin, Sachin Sachdeva, Sheldon X.-D. Tan
ASP-DAC5
2023 PAALM: Power Density Aware Approximate Logarithmic Multiplier Design
abstract
Approximate hardware designs can lead to significant power or energy reduction. However, a recent study showed that approximated designs might lead to unwanted higher temperature and related reliability issues due to the increased power density. In this work, we try to mitigate this important problem by proposing a novel power density aware approximate logarithmic multiplier (called PAALM) design for the first time. The new multiplier design is based on the approximate logarithmic multiplier (ALM) framework due to its rigorous mathematics based foundation. The idea is to re-design the high computing switch activities of existing ALM designs based on equivalent mathematical formula so that the power density can be reduced at no accuracy loss while at costs of some area overheads. Our results show that the proposed PAALM design can improve 11.5%/5.7% of power density and 31.6%/70.8% of area with 8/16-bit precision when compared with the fixed-point multiplier baseline, respectively. And also achieves extremely low error bias: -0.17/0.08 for 8/16-bit precision, respectively. On top of this, we further implement the PAALM design in a Convolutional Neural Network (CNN) and test it on CIFAR10 dataset. The results show that with error compensation, PAALM can achieve the same inference accuracy as the fixed-point multiplier baseline. We also evaluate the PAALM in a discrete cosine transformation (DCT) application. The results show that with error compensation, PAALM can improve the image quality of 8.6dB in average when compared to the ALM design.
Shuyuan Yu, Sheldon X.-D. Tan
ASP-DAC2
2023 Fast Full-Chip Parametric Thermal Analysis Based on Enhanced Physics Enforced Neural Networks
abstract
In this work, we propose a fast full-chip thermal numerical analysis approach based on an enhanced physics-informed neural networks (PINN) framework. The new method, called ThermPINN, leverages both PINN-based DNN optimization framework and analytic solutions of simplified thermal problems for solving thermal partial differential equations (PDE). The resulting ThermPINN leads to more efficient training speed of DNN networks and more scalability for solving large PDE problems. Specifically, we propose to partially enforce physics laws based on closely related analytic solutions to simpler problems. As a result, we are able to significantly reduce the number of variables in the loss function and easily meet boundary conditions. To consider the impact of various ambient temperatures and effective convection coefficients, which are influenced by different design parameters and run-time conditions, we develop a parameterized thermal analysis technique. This technique enables design space exploration and uncertainty quantification (UQ), which are critical for ensuring the reliability of integrated circuits under various operating conditions. The numerical results on alpha21264 processor show that the proposed ThermPINN has 2× speedup and 3× better accuracy over the state-of-the-art thermal simulator, VarSim. The experimental results for 2-D full-chip thermal analysis of 3171 cases show that the proposed parameterized ThermPINN considering both training and inference time can achieve a 6× speedup over commercial COMSOL with an average mean absolute error (AE) of 0.47 K. In terms of training time, the proposed parameterized ThermPINN is 11× faster than the parameterized plain PINN with similar accuracy. The UQ analysis with 5000 samples for maximum temperature propagated from ambient temperature shows that the parameterized ThermPINN and parameterized plain PINN are 113× and 22× faster than COMSOL, respectively.
Liang Chen 0025, Jincong Lu, Wentian Jin, Sheldon X.-D. Tan
ICCAD4
2023 PostPINN-EM: Fast Post-Voiding Electromigration Analysis Using Two-Stage Physics-Informed Neural Networks
abstract
In this paper, we propose a novel machine learning-based approach, called PostPInn- Em, for solving the partial differential equations for stress evolution in a confined metal interconnect multi-segment trees during the post-voiding stage for fast electromigration (EM) check for interconnects. The new approach is based on an enhanced two-stage Physics-Informed Neural Networks (PINN) framework in which the physics law for a single wire is enforced first and then atomic flux conservation and stress continuity at the inter-segment junctions of wire segments are then fulfilled to reduce the number of variables of loss functions for the fast training process. Existing two-stage PINN method uses supervised learning method for modeling a single wire under various atomic flux conditions for the first stage, which turns out to be much more difficult for post-voiding phase due to arbitrary non-zero initial conditions. To mitigate this problem, we propose a new closed-form parameterized formula for stress solution of single wires with variable bound-ary conditions based on the Laplace transformation methods. Furthermore, we derive the analytic solutions for wire segment with and without voiding as not all the wire segments will have voids during the post-voiding phase. Numerical results on some synthesized multi-segment interconnects show that the proposed PostPINN-EM can achieve more than 100X speedup compared to FEM based tool COMSOL with the expense of less than 1% accuracy. Compared to the state of the art tool EMspice v1.0 [1], this method can achieve more than 25X speedup with similar accuracy compared to golden results from COMSOL.
Subed Lamichhane, Wentian Jin, Liang Chen 0025, Mohammadamir Kavousi, Sheldon X.-D. Tan
ICCAD5
2023 Real-time Thermal Map Estimation for AMD Multi-Core CPUs Using Transformer
abstract
This paper presents a novel approach for real-time estimation of spatial thermal maps for the commercial AMD Ryzen 7 4800U 8-core microprocessor using a transformer-based machine learning method. The proposed method, called ThermTransformer, leverages real-time performance metrics of the AMD chip, provided by uProf 4.0, to accurately estimate transient thermal maps. These maps can be valuable for dynamic thermal, power, and reliability controls requiring higher accuracy. Unlike traditional Convolutional Neural Networks (CNN) designed for image data or Recurrent Neural Networks (RNN) suitable for transient data, ThermTransformer is based on a modified self-attention architecture. It takes time-series performance metrics information as input and directly generates transient thermal images. Our results demonstrate that this transformer-based method achieves the best of both worlds - surpassing CNN in prediction quality and performing well for transient data. Exper-imental results reveal that ThermTransformer achieves highly accurate predictions of power maps, with an RMSE of only 0.36°C or 0.8% of the full-scale error. Additionally, it outperforms the recently proposed GAN-based thermal map estimation method, ThermGAN, by 1.66x and the LSTM-based thermal prediction method, RealMaps, by 6.09x in terms of accuracy on average. Furthermore, the proposed approach can be efficiently deployed on the target chip, providing real-time estimation with a speed as fast as 14ms.
Jincong Lu, Sheldon X.-D. Tan
ICCAD3
2023 MAGIC-DHT: Fast in-memory computing for Discrete Hadamard Transform
abstract
Discrete Hadamard transform (DHT) is a signal processing tool that decomposes an arbitrary input vector into a superposition of Walsh functions. Due to its wide range of applications in processing big data, a fast and energy-efficient hardware design for DHT with high throughput capability is essential. Processing in memory (PIM) allows the in-place computation to reduce the data traffic, which is a major speed bottleneck in the existing computing. In this work, we propose an efficient hybrid parallel PIM-based computation for DHT. Our proposed method explores the recursive computation of DHT and is based on the memristor-aided logic (MAGIC) gates in which the arithmetic operations are carried out via simple logic NOR operation. We propose two in-memory computing methods for the DHT encoding process. At the arithmetic level, to improve efficiency, we propose to share the intermediate results between addition and subtraction in DHT in the first method called MAGIC-DHT-1D which provides an average speedup of 1.12× over the recently proposed DigitalPIM for 1D DHT. Furthermore,MAGIC-DHT-1D also outperforms SIMPLER in terms of energy and energy density in average. We also propose a second method, called MAGIC-DHT-2D, to share the carrier independent computation cycles among multi-bit parallel addition and subtraction. At the algorithm level, we also explore both row and column-based PIM NOR computing in the same crossbar to avoid the transposition operation required in the 2D DHT process. MAGIC-DHT-2D provides an average speedup of 4.84× and 7.25× over two state-of-the-art methods DigitalPIM and SIMPLER, respectively for each each complete set of 2D DHT computing cycles. Our numerical results further show that our proposed optimized methods can lead up to 56.19× and 6.90× speed-up, as well as 57.84× and 5.96× higher throughput over NVIDIA RTX Titan GPU to compute 1D DHT and 2D DHT, respectively.
Maliha Tasnim, Chinmay Raje, Shuyuan Yu, Elaheh Sadredini, Sheldon X.-D. Tan
Integr.5
2023 Hot-spot aware thermoelectric array based cooling for multicore processors
Sheriff Sadiqbatcha, Liang Chen 0025, Cuong Thi, Sachin Sachdeva, Hussam Amrouch, Sheldon X.-D. Tan
Integr.7
2023 Linear Time Electromigration Analysis Based on Physics-Informed Sparse Regression
abstract
In this work, we propose a novel physics-informed sparse regression (PISR) framework to solve stress evolution (described by Korhonen’s equations) in general multisegment wires using an unsupervised learning scheme. Unlike the existing physics-informed neural network (PINN) framework, the PISR method trains the trainable weights through the Moore–Penrose generalized inverse algorithm used in extreme learning machine (ELM), which is extremely faster than the backpropagation algorithm. To improve the accuracy of PISR for complex multisegment interconnects, we employ domain decomposition schemes in both space and time. For each subdomain, we use different trainable weights but the same shared neural network to represent each subsolution, which leads to more efficient memory usage. Furthermore, we propose to use sparse matrix techniques to accelerate the training speed of the PISR method and prove that the resulting PISR has linear time complexity for analyzing tree-structured interconnects. Finally, we divide the time into many time intervals and apply an autoregressive model to simulate each time interval in sequence to further improve scalability and reduce memory cost so that the PISR method can perform EM analysis for large-scale multisegment interconnects. Experimental results on different kinds of interconnect structures show that the proposed PISR method has the same accuracy level as the numerical methods. The results on$N_{T}$T-junctions interconnect trees show that the proposed PISR method indeed demonstrates true linear time complexity. Furthermore, PISR can deliver$8.9\times $,$20.6\times $, and$1284\times $speedups over the recently proposed semi-analytic method (ASOV), finite difference method accelerated with model order reduction (FDM-MOR), FDM for the interconnect with$N_{T}= 5000$, respectively. Furthermore, we show that PISR also achieves an$818\times $speedup in training over the plain PINN method based on the traditional backpropagation algorithm.
Liang Chen 0025, Wentian Jin, Mohammadamir Kavousi, Subed Lamichhane, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Thermoelectric Cooler Modeling and Optimization via Surrogate Modeling Using Implicit Physics-Constrained Neural Networks
abstract
Thermoelectric cooler (TEC) is a promising active cooling device to remove the localized hot spots precisely in VLSI chips. In this article, we use a novel implicit physics-constrained neural networks (called IPCNNs) to build a surrogate model for the single TEC device with the reduction from 3-D to 1-D. First, the surrogate model represented by the deep neural networks (DNNs) allows parameterization of key design and running parameters, such as current density, length, and thermal boundary conditions of the TEC. Second, the proposed method tries to partition the physics laws into two different groups, which then are enforced by supervised learning and physics-informed neural networks (PINNs) framework sequentially. Such implicit PCNN scheme can lead to much faster training speed and better convergent accuracy for the unsupervised training. The existing plain PINN enforces all the physics laws via the loss functions and the network tends to have very slow training speed and a large convergent error for large problems. An extreme learning machine (ELM) is used for the networks in the first stage. Compared with fully connected network (FCN) trained by the traditional back-propagation algorithm, ELM can be easily trained and converges much faster. Furthermore, by leveraging the differential nature of the DNN model, we can directly estimate the derivative of the cooling heat flux with respect to current density instead of using a finite difference approximation. The calculated derivatives are used to find the optimal current density to achieve maximum cooling heat flux via Newton’s method. Last but not least, we propose a novel hybrid finite element neural network (FENN) method to perform thermal analysis of the VLSI chip system with the TEC device. The DNN model is embedded into COMSOL through the heat flux boundary conditions. Experimental results show that the machine learning-based method can achieve about$8.5\times $speedup with good accuracy than the COMSOL-based finite element method. Furthermore, the proposed IPCNN is more stable and accurate than the existing PINN. The proposed FENN can have a$5.1\times $speedup and$5.4\times $memory reduction over the traditional numerical method.
Liang Chen 0025, Wentian Jin, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Hot-Trim: Thermal and Reliability Management for Commercial Multicore Processors Considering Workload Dependent Hot Spots
abstract
This work proposes a new dynamic thermal and reliability management framework via task mapping and migration to improve thermal performance and reliability of commercial multicore processors considering workload-dependent thermal hot spot stress. The new method is motivated by the observation that different workloads activate different spatial power and thermal hot spots within each core of processors. Existing run-time thermal management, which is based on on-chip location-fixed thermal sensor information, can lead to suboptimal management solutions as the temperatures provided by those sensors may not be the true hot spots. The new method, called Hot-Trim, utilizes a machine learning-based approach to characterize the power density hot spots across each core, then a new task mapping/migration scheme is developed based on the hot spot stresses. Compared to existing works, the new approach is the first to optimize VLSI reliabilities by exploring workload-dependent power hot spots. The advantages of the proposed method over the Linux baseline task mapping and the temperature-based mapping method are demonstrated and validated on real commercial chips. Experiments on a real Intel Core i7 quad-core processor executing PARSEC-3.0 and SPLASH-2 benchmarks show that, compared to the existing Linux scheduler, core and hot spot temperature can be lowered by 1.15 °C–1.31 °C. In addition, Hot-Trim can improve the chip’s electro-migration (EM), negative biased temperature instability, and hot-carrier-injection (HCI) related reliability by 30.2%, 7.0%, and 31.1%, respectively, compared to Linux baseline without any performance degradation. Furthermore, it improves EM and HCI-related reliability by 29.6% and 19.6%, respectively, and at the same time even further reduces the temperature by half a degree compared to the conventional temperature-based mapping technique.
Sheriff Sadiqbatcha, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 GridNetOpt: Fast Full-Chip EM-Aware Power Grid Optimization Accelerated by Deep Neural Networks
abstract
This article presents a fast full-chip electromigration (EM) aware IR drop constrained optimization framework, namedGridNetOpt, for on-chip power grid networks accelerated by deep neural networks (DNNs). Compared to the existing linear programming-based methods, the new method employs more flexible conjugate gradient-based optimization to size the wire width of the power grids. To mitigate the high cost of sensitivity calculation of the adjoint network using full-chip IR drop analysis at every iteration step, the sensitivity is computed via a trained conditional generative adversarial network (CGAN). The new method exploits the differentiable characteristics of DNNs for fast sensitivity computation. The sensitivity, which is the node voltage with respect to wire resistance, will guide the search direction during the optimization process. In order to consider more accurate EM failure effects, the training data is obtained from the power grids under different wire widths and current loads analyzed by a state-of-the-art full-chip multiphysics-based coupled EM-IR drop analysis tool. This is in contrast with the existing linear programming-based methods, in which only immortal wires or wires with nonzero resistance can be dealt with. Numerical results on a number of synthesized power grid benchmarks from ARM Cortex-M0 processor designs show that the proposedGridNetOptcan lead to at least an order of magnitude speedup over the conjugate gradient-based method using the traditional adjoint network method. Compared to the previous localized power grid fixing work withGridNet,GridNetOptleads to smaller area overhead for all the benchmarks we tested. It can also reduce IR drops for power grid circuits with immortal wires, which is not possible with the localizedGridNetmethod.
Han Zhou 0002, Wentian Jin, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Fast Thermal Analysis for Chiplet Design based on Graph Convolution Networks
abstract
2.5D chiplet-based technology promises an efficient integration technique for advanced designs with more functionality and higher performance. Temperature and related thermal optimization, heat removal are of critical importance for temperature-aware physical synthesis for chiplets. This paper presents a novel graph convolutional networks (GCN) architecture to estimate the thermal map of the 2.5D chiplet-based systems with the thermal resistance networks built by the compact thermal model (CTM). First, we take the total power of all chiplets as an input feature, which is a global feature. This additional global information can overcome the limitation that the GCN can only extract local information via neighborhood aggregation. Second, inspired by convolutional neural networks (CNN), we add skip connection into the GCN to pass the global feature directly across the hidden layers with the concatenation operation. Third, to consider the edge embedding feature, we propose an edge-based attention mechanism based on the graph attention networks (GAT). Last, with the multiple aggregators and scalers of principle neighborhood aggregation (PNA) networks, we can further improve the modeling capacity of the novel GCN. The experimental results show that the proposed GCN model can achieve an average RMSE of 0.31 K and deliver a 2.6× speedup over the fast steady-state solver of open-source HotSpot based on SuperLU. More importantly, the GCN model demonstrates more useful generalization or transferable capability. Our results show that the trained GCN can be directly applied to predict thermal maps of six unseen datasets with acceptable mean RMSEs of less than 0.67 K without retraining via inductive learning.
Liang Chen 0025, Wentian Jin, Sheldon X.-D. Tan
ASP-DAC3
2022 Fast Electromigration Stress Analysis Considering Spatial Joule Heating Effects
abstract
Temperature gradient due to Joule heating has huge impacts on the electromigration (EM) induced failure effects. However, Joule heating and related thermomigration (TM) effects were less investigated in the past for physics-based EM analysis for VLSI chip design. In this work, we propose a new spatial temperature aware transient EM induced stress analysis method. The new method consists of two new contributions: First, we propose a new TM-aware void saturation volume estimation method for fast immortality check in the post-voiding phase for the first time. We derive the analytic formula to estimate the void saturation in the presence of spatial temperature gradients due to Joule heating. Second, we develop a fast numerical solution for EM-induced stress analysis for multi-segment interconnect trees considering TM effect. The new method first transforms the coupled EM-TM partial differential equations into linear time-invariant ordinary differential equations (ODEs). Then extended Krylov subspace-based reduction technique is employed to reduce the size of the original system matrices so that they can be efficiently simulated in the time domain. The proposed method can perform the simulation process for both void nucleation and void growth phases under time-varying input currents and position-dependent temperatures. The numerical results show that, compared to the recently proposed semi-analytic EM-TM method, the proposed method can lead to about 28x speedup on average for the interconnect with up to 1000 branches for both void nucleation and growth phases with negligible errors.
Mohammadamir Kavousi, Liang Chen 0025, Sheldon X.-D. Tan
ASP-DAC3
2022 HEALM: Hardware-Efficient Approximate Logarithmic Multiplier with Reduced Error
abstract
In this work, we propose a new approximate logarithm multipliers (ALM) based on a novel error compensation scheme. The proposed hardware-efficient ALM, named HEALM, first determines the truncation width for mantissa summation in ALM. Then the error compensation or reduction is performed via a lookup table, which stores reduction factors for different regions of input operands. This is in contrast to an existing approach, in which error reduction is performed independently of the width truncation of mantissa summation. As a result, the new design will lead to more accurate result with both reduced area and power. Furthermore, different from existing approaches which will either introduce resource overheads when doing error improvement or lose accuracy when saving area and power, HEALM can improve accuracy and resource consumption at the same time. Our study shows that 8-bit HEALM can achieve up to 2.92%, 9.30%, 16.08%, 17.61% improvement in mean error, peak error, area, power consumption respectively over REALM, which is the state of art work with the same number of bits truncated. We also propose a single error coefficient mode named HEALM-TA-S, which improves the ALM design with a truncation adder (TA) for mantissa summation. Furthermore, we evaluate the proposed HEALM design in a discrete cosine transformation (DCT) application. The result shows that with different values of k, HEALM-TA can improve the image quality upon the ALM baseline by 7.8~17.2dB in average and HEALM-SOA can improve 2.9~15.8dB in average, respectively. Besides, HEALM-TA and HEALM-SOA outperform all the state of art works with k = 2, 3, 4 on the image quality. And the single coefficient mode, HEALM-TA-S, can improve the image quality upon the baseline up to 4.1dB in average with extremely low resource consumption.
Shuyuan Yu, Maliha Tasnim, Sheldon X.-D. Tan
ASP-DAC3
2022 Scaled-CBSC: scaled counting-based stochastic computing multiplication for improved accuracy
abstract
Stochastic computing (SC) can lead area-efficient implementation of logic designs. Existing SC multiplication, however, suffers a long-standing problem: large multiplication error with small inputs due to its intrinsic nature of bit-stream based computing. In this article, we propose a new scaled counting-based SC multiplication approach, called Scaled-CBSC, to mitigate this issue by introducing scaling bits to ensure the bit '1' density of the stochastic number is sufficiently large. The idea is to convert the "small" inputs to "large" inputs, thus improve the accuracy of SC multiplication. But different from an existing stream-bit based approach, the new method uses the binary format and does not require stochastic addition as the SC multiplication always starts with binary numbers. Furthermore, Scaled-CBSC only requires all the numbers to be larger than 0.5 instead of arbitrary defined threshold, which leads to integer numbers only for the scaling term. The experimental results show that the 8-bit Scaled-CBSC multiplication with 3 scaling bits can achieve up to 46.6% and 30.4% improvements in mean error and standard deviation, respectively; reduce the peak relative error from 100% to 1.8%; and improve 12.6%, 51.5%, 57.6%, 58.4% in delay, area, area-delay product, energy consumption, respectively, over the state of art work. Furthermore, we evaluate the proposed multiplication approach in a discrete cosine transformation (DCT) application. The results show that with 3 scaling bits, 8-bit scaled counting-based SC multiplication can improve the image quality with 5.9dB upon the state of art work in average.
Shuyuan Yu, Sheldon X.-D. Tan
DAC2
2022 HierPINN-EM: Fast Learning-Based Electromigration Analysis for Multi-Segment Interconnects Using Hierarchical Physics-Informed Neural Network
abstract
Electromigration (EM) becomes a major concern for VLSI circuits as the technology advances in the nanometer regime. The crux of problem is to solve the partial differential Korhonen equations, which remains challenging due to the increasing integrated density. Recently, scientific machine learning has been explored to solve partial differential equations (PDE) due to breakthrough success in deep neural networks and existing approach such as physics-informed neural networks (PINN) shows promising results for some small PDE problems. However, for large engineering problems like EM analysis for large interconnect trees, it was shown that the plain PINN does not work well due the to large number of variables. In this work, we propose a novel hierarchical PINN approach, HierPINN-EM for fast EM induced stress analysis for multi-segment interconnects. Instead of solving the interconnect tree as a whole, we first solve EM problem for one wire segment under different boundary and geometrical parameters using supervised learning. Then we apply unsupervised PINN concept to solve the whole interconnects by enforcing the physics laws in the boundaries for all wire segments. In this way, HierPINN-EM can significantly reduce the number of variables at plain PINN solver. Numerical results on a number of synthetic interconnect trees show that HierPINN-EM can lead to orders of magnitude speedup in training and more than 79× better accuracy over the plain PINN method. Furthermore, HierPINN-EM yields 19% better accuracy with 99% reduction in training cost over recently proposed Graph Neural Network-based EM solver, EMGraph.
Wentian Jin, Liang Chen 0025, Subed Lamichhane, Mohammadamir Kavousi, Sheldon X.-D. Tan
ICCAD5
2022 GPUCalorie: Floorplan Estimation for GPU Thermal Evaluation
abstract
GPUs are massively parallel architecture that consume significant power, which lead to high thermal output. Thermal constraints of GPUs are one of the major limitations in high performance, mobile and embedded applications. However, accurate thermal modeling tools for GPUs are lacking for researchers. We identify that limiting factors to further research are the absence of GPU floorplans necessary for thermal modeling, validated thermal trends, and outdated component-level power models. To this end, we present GPUCalorie, a thermal modeling methodology using specialized infrared thermography setup for measuring and validating thermal behaviors of real GPUs. We validate a floorplan of Nvidia’s GTX1050 identified through our infrared thermography setup. We validate the GPUCalorie identified floorplan against a real GTX1050 GPU, showing 10% error for the thermal map.
Marcus Chow, Ali Jahanshahi, Ana Cardenas Beltran, Sheldon X.-D. Tan, Daniel Wong 0001
ISPASS4
2022 Real-Time Full-Chip Thermal Tracking: A Post-Silicon, Machine Learning Perspective
abstract
This article presents a novel approach to real-time tracking of full-chip heatmaps for off-the-shelf microprocessors based on machine-learning. The proposed post-silicon approach, named RealMaps, only uses the existing temperature sensors and workload-independent utilization information. RealMaps does not require any knowledge of the proprietary design or manufacturing process-specific details of the chip. Consequently, the methods presented in this work can be implemented by either the original chip manufacturer or a third party alike. The approach involves offline acquisition of spatial heatmaps using a thermal imaging setup. To build the dynamic thermal model, a temporal-aware long-short-term-memory neutral network is trained with system-level features as inputs. 2D discrete cosine transformation (DCT) is performed on the heatmaps so that they can be expressed with just a few dominant DCT coefficients. This allows the model to be built to estimate just the dominant spatial features of the heatmaps, rather than the entire heatmap images, making it significantly more efficient. Experimental results from two commercial chips show that RealMaps can estimate the full-chip heatmaps with 0.9C and 1.2C root-mean-square-error respectively and take only 0.4ms for each inference. Compared to the state of the art pre-silicon approach, RealMaps shows similar accuracy, but with much less computational cost.
Sheriff Sadiqbatcha, Hussam Amrouch, Sheldon X.-D. Tan
IEEE Trans. Computers4
2022 Electrothermal Simulation and Optimal Design of Thermoelectric Cooler Using Analytical Approach
abstract
In this article, electrothermal modeling and simulation of thermoelectric cooling (TEC) in the package design of VLSI systems are performed by solving coupled heat conduction and current continuity equations. We propose a new analytical solution to the coupled partial differential equations (PDEs) which describe temperature and voltage with the reduction from 3-D to 1-D. In addition to this, we derive new analytic expressions for two key performance metrics for TEC devices: 1) the maximum temperature difference and 2) the maximum heat-flux pumping capability, which can be guided for the optimal design of thermoelectric cooler to achieve the maximum cooling performance. Furthermore, for the first time, we observe that when the dimensionless figure of merit$ZT_{0}$value is larger than 1, there is no maximum heat-flux value, which means the heat dissipation due to the Peltier and Fourier transfer effects is larger than the heat generation caused by the Joule heating effect, which can lead to more efficient TEC cooling design. The accuracy of the proposed 1-D formulas is verified by a 3-D finite element method using COMSOL software. The compact model delivers many orders of magnitude speedup and memory saving compared to COMSOL with marginal accuracy loss. Compared with the conventional simplified 1-D energy equilibrium model, the proposed analytical coupled multiphysics model is more robust and accurate.
Liang Chen 0025, Sheriff Sadiqbatcha, Hussam Amrouch, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Full-Chip Power Density and Thermal Map Characterization for Commercial Microprocessors Under Heat Sink Cooling
abstract
In this article, we address the problem of accurate full-chip power and thermal map estimation for commercial off-the-shelf multicore processors. Processors operating with heat sink cooling remains a challenging problem due to the difficulty in direct measurement. We first propose an accurate full-chip steady-state power density map estimation method for commercial multicore microprocessors. The new method consists of a few steps. First, 2-D spatial Laplace operation is performed on the measured thermal maps (images) without heat sink to obtain the so-calledraw power maps. Then, a novel scheme is developed to generate the true power density maps from the raw power density maps. The new approach is based on thermal measurements of the processor with back-side cooling using an advanced infrared (IR) thermal imaging system. FEM thermal model constructed in COMSOL Multiphysics is used to validate the estimated power density maps and thermal conductivity. Later, this work creates a high-fidelity FEM thermal model with heat sink and reconstructs the full-chip thermal maps while the heat sink is on. Ensuring that power maps are similar under back cooling and heat sink cooling settings, the reconstructed thermal maps are verified by the matching between the on-chip thermal sensor readings and the corresponding elements of thermal maps. Experiments on an Intel i7-8650U 4-core processor with back cooling shows 96% similarity (2-D correlation) between the measured thermal maps and the thermal maps reconstructed from the estimated power maps, with 1.3 °C average absolute error. Under heat sink cooling, the average absolute error is 2.2 °C over a 56 °C temperature range and about 3.9% error between the computed and the real thermal maps at the sensor locations. Furthermore, the proposed power map estimation method achieves higher resolution and at least$100\times $speedup than a recently proposed state-of-art Blind Power Identification method.
Sheriff Sadiqbatcha, Michael O'Dea, Hussam Amrouch, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 EMGraph: Fast Learning-Based Electromigration Analysis for Multi-Segment Interconnect Using Graph Convolution Networks
abstract
Electromigration (EM) becomes a major concern for VLSI circuits as the technology advances in the nanometer regime. With Korhonen equations, EM assessment for VLSI circuits remains challenged due to the increasing integrated density. VLSI multisegment interconnect trees can be naturally viewed as graphs. Based on this observation, we propose a new graph convolution network (GCN) model, which is called EMGraph considering both node and edge embedding features, to estimate the transient EM stress of interconnect trees. Compared with recently proposed generative adversarial network (GAN) based stress image-generation method, EMGraph model can learn more transferable knowledge to predict stress distributions on new graphs without retraining via inductive learning. Trained on the large dataset, the model shows less than 1.5% averaged error compared to the ground truth results and is orders of magnitude faster than both COMSOL and state-of-the-art method. It also achieves smaller model size, $4\times$ accuracy and $14\times$ speedup over the GAN-based method.
Wentian Jin, Liang Chen 0025, Sheriff Sadiqbatcha, Shaoyi Peng, Sheldon X.-D. Tan
DAC5
2021 COSAIM: Counter-based Stochastic-behaving Approximate Integer Multiplier for Deep Neural Networks
abstract
In this work, we propose a new counter-based stochastic-behaving approximate integer unsigned multiplier, called COSAIM, for many emerging error tolerant application workloads such as deep neural networks. Unlike existing approximate multipliers, which are based on some deterministic ad-hoc methods or mathematical formula, the new design is an improved stochastic multiplier, which performs improved sequential counting for multiplication operation in a deterministic way. In this work, we further improve the counting efficiency by introducing approximate schemes to significantly speed up the counting process, which leads to significant clock cycle reduction with no accuracy loss. COSAIM bears all the advantages of stochastic computing such as built-in configurability for progressive performance-accuracy trade-off. At the same time, it shows very small latency and high energy efficiency. Our evaluation shows that the COSAIM with error improvement operation can achieve very low error bias (0.06%), along with lower mean error (0.30% to 3.49%), and low peak errors (around 1.81%) with variance of 1. $47\times 10^{-4}$ %. Experimental results obtained from Xilinx ISE show that compared with the 8-bit exact multiplier baseline, COSAIM can save up to 53.95%, 32.84%, 52.24%, 21.05% in area, power, energy and the product Area. 1/Throughput, respectively. Furthermore, by doing shared parallel design, COSAIM can further lead to improvements in area, power and energy reduction by 60.44%, 53.33% and 68.54%, respectively compared to the baseline. We also implement COSAIM in a Convolution Neural Network (CNN) and test it on CIFAR10 dataset and find that CNN with COSAIM delivers similar inference accuracy compared to some state of art approximate multipliers.
Shuyuan Yu, Sheldon X.-D. Tan
DAC3
2021 Data-Driven Electrostatics Analysis based on Physics-Constrained Deep learning
abstract
Computing the electric potential and electric field is important for modeling and analysis of VLSI chip and high speed circuits. For instance, it is an important step for DC analysis for high speed circuits as well as dielectric reliability and capacitance extraction for VLSI interconnects. In this paper, we propose a new data-driven meshless 2D analysis method, called PCEsolve, of electric potential and electric fields based on the physics-constrained deep learning scheme. We show how to formulate the differential loss functions to consider the Laplace differential equations with voltage boundary conditions for typical electrostatic analysis problem so that the supervised learning process can be carried out. We apply the resulting PCEsolve solver to calculate electric potential and electric field for VLSI interconnects with complicated boundaries. We show the potential and limitations of physics-constrained deep learning for practical electrostatics analysis. Our study for purely label-free training (in which no information from FEM solver is provided) shows that PCEsolve can get accurate results around the boundaries, but the accuracy degenerates in regions far away from the boundaries. To mitigate this problem, we explore to add some simulation data or labels at collocation points derived from FEM analysis and resulting PCEsolve can be much more accurate across all the solution domain. Numerical results demonstrate that the PCEsolve achieves an average error rate of 3.6% on 64 cases with random boundary conditions and it is 27.5× faster than COMSOL on test cases. The speedup can be further boosted to ~ 38000× in single-point estimations. We also study the impacts of weights on different components of loss functions to improve the model accuracy for both voltage and electric field.
Wentian Jin, Shaoyi Peng, Sheldon X.-D. Tan
DATE3
2021 Special Session: Machine Learning for Semiconductor Test and Reliability
abstract
With technology scaling approaching atomic levels, IC test and diagnosis of complex System-on-Chips (SoCs) become overwhelming challenging. In addition, sustaining the reliability of transistors as well as circuits at such extreme feature sizes, for the entire projected lifetime, also become profoundly difficult. This holds even more when it comes to emerging technologies that go beyond convectional CMOS in which the underlying physics are not yet fully understood. In this special session paper, we describe the usage of machine learning in several test and reliability related areas. First, we demonstrate the vital role that machine learning can play in IC test showing the importance of explainability as a frontier for machine learning in IC test. Afterwards, we discuss how novel physics-informed neural networks can be employed to model electrostatic problems in VLSI designs. This is essential to mitigate the deleterious effects of of time dependent dielectric breakdown, which is the key source of reliability degradations. Finally, we discuss the major sources of reliability degradations at the transistor level in advanced technology nodes such as transistor aging phenomena and self-heating effects as well as we demonstrate how machine learning approaches can further help in developing reliable emerging technologies.
Hussam Amrouch, Animesh Basak Chowdhury, Wentian Jin, Ramesh Karri, Farshad Khorrami, Prashanth Krishnamurthy, Ilia Polian, Victor M. van Santen, Benjamin Tan 0001, Sheldon X.-D. Tan
VTS10
2021 Robust power grid network design considering EM aging effects for multi-segment wires
Han Zhou 0002, Liang Chen 0025, Sheldon X.-D. Tan
Integr.3
2021 A Fast Semi-Analytic Approach for Combined Electromigration and Thermomigration Analysis for General Multisegment Interconnects
abstract
Considering temperature gradient or thermomigration (TM) impacts on electromigration (EM) due to Joule heating was less studied in the past. In this article, we propose a new semi-analytical stress transient analysis method to consider both EM and TM effects for general multisegment interconnects. The new method is based on the separation of variables (SOVs) approach to find the analytic solution of coupled EM-TM partial differential equation (PDE). The algorithm consists of several steps. We first develop analytic solutions to compute the steady-state temperature distribution of multisegment wires. Based on this, we derive closed-form solutions for steady-state hydrostatic stress distribution in the context of thermal gradients due to Joule heating for multisegment interconnect wires. With the steady-state stress distribution, the coupled EM-TM PDE can be homogenized and solved by the SOV method. To deal with temperature/position-dependent diffusivity of metal migration process due to nonuniform temperature distribution, we utilize a piecewise linear technique to approximate the position-dependent diffusivity. The numerical results on multisegment interconnects show that the proposed method has negligible error loss compared to commercial finite element analysis software COMSOL but is about an order of magnitude faster than COMSOL with 10× less memory footprint. The numerical results further show that temperature gradient due to Joule heating indeed has significant impacts on the EM failure process.
Liang Chen 0025, Sheldon X.-D. Tan, Zeyu Sun 0001, Shaoyi Peng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Post-Silicon Heat-Source Identification and Machine-Learning-Based Thermal Modeling Using Infrared Thermal Imaging
abstract
In this article, we present a novel post-silicon approach to locating the dominant heat sources on commercial multicore processors using heatmaps measured via an infrared (IR) thermal imaging setup. To locate the heat sources, 2-D spatial Laplacian transformation is performed on the heatmaps followed by K-means clustering to find the dominant power/heat-source clusters. This is an exclusively post-silicon approach that does not require any knowledge of the underlying design of the commercial chips other than the information that is publicly available. Since the identified clusters are the thermally vulnerable areas on the die, we then propose a machine-learning-based framework to deriving a thermal model capable of estimating their temperatures during online use. Our approach involves collecting transient temperature data of the aforementioned heat sources and synchronized high-level performance metrics from the chip, and training a long-short-term-memory (LSTM) neural network (NN) that uses the performance metrics as inputs to estimate the temperatures of the identified heat sources in real time. Since the model is meant for real-time use, we explore methods of reducing the performance overhead and inference time of the model. This includes a novel power correlation-based approach to identifying the thermally irrelevant performance metrics and eliminating them in order to reduce the input dimensionality of the model, and an analysis on network sizing to determine the ideal NN configuration for the problem at hand. The model is trained and tested exclusively using measured thermal data from commercial multicore processors. The experimental results from two Intel multicore processors (i5-3337U and i7-8650U) show that the proposed approach achieves very high accuracy (root-mean-square error: 0.55 °C-0.93 °C) in estimating the temperatures of all the identified heat sources on the chip.
Sheriff Sadiqbatcha, Hengyang Zhao, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 Fast Physics-Based Electromigration Analysis for Full-Chip Networks by Efficient Eigenfunction-Based Solution
abstract
Electromigration (EM) becomes one of the most challenging reliability issues for current and future ICs in 10-nm technology and below. In this article, a novel method is proposed for the EM hydrostatic stress analysis on 2-D multibranch interconnect trees, which is the foundation of the EM reliability assessment for large-scale on-chip interconnect networks, such as on-chip power grid networks. The proposed method, which is based on an eigenfunction technique, could efficiently calculate the hydrostatic stress evolution for multibranch interconnect trees stressed with different current densities and nonuniformly distributed thermal effects. The proposed method solves the partial differential equations of transient EM stress more efficiently since it does not require any discretization either spatially or temporally, which is in contrast to numerical methods, such as the finite difference method and finite element method. The accuracy of the proposed transient analysis approach is validated against the analytical solution and commercial tools. The convergence of the proposed method is demonstrated by numerical experiments on practical power/ground networks, showing that only a small number of eigenfunction terms are necessary for the accurate solution. Thanks to its analytical nature, the proposed method is also utilized in efficient EM analysis techniques, such as searching for the void nucleation time by a modified bisection algorithm. The numerical results show that the proposed method is 10X-100X faster than the finite difference method and scales better for larger interconnect trees.
Shaobin Ma, Sheldon X.-D. Tan, Chase Cook, Liang Chen 0025, Jianlei Yang 0001, Wenjian Yu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 An Adaptive Electromigration Assessment Algorithm for Full-chip Power/Ground Networks
abstract
In this paper, an adaptive algorithm is proposed to perform electromigration (EM) assessment for full-chip power/ground networks. Based on the eigenfunction solutions, the proposed method improves the efficiency by properly selecting the eigenfunction terms and utilizing the closed-form eigenfunctions for commonly seen interconnect wires such as T-shaped or cross-shaped wires. It is demonstrated that the proposed method can trad-off well among the accuracy, efficiency and applicability of the eigenfunction based methods. The experimental results show that the proposed method is about three times faster than the finite difference method and other eigenfunction based methods.
Shaobin Ma, Sheldon X.-D. Tan, Liang Chen 0025
ASP-DAC3
2020 Machine Learning Based Online Full-Chip Heatmap Estimation
abstract
Runtime power and thermal control is crucial in any modern processor. However, these control schemes require accurate real-time temperature information, ideally of the entire die area, in order to be effective. On-chip temperature sensors alone cannot provide the full-chip temperature information since the number of sensors that are typically available is very limited due to their high area and power overheads. Furthermore, as we will demonstrate, the peak locations within hot-spots are not stationary and are very workload dependent, making it difficult to rely on fixed temperature sensors alone. Therefore, we propose a novel approach to real-time estimation of fullchip transient heatmaps for commercial processors based on machine learning. The model derived in this work supplements the temperature data sensed from the existing on-chip sensors, allowing for the development of more robust runtime power and thermal control schemes that can take advantage of the additional thermal information that is otherwise not available. The new approach involves offline acquisition of accurate spatial and temporal heatmaps using an infrared thermal imaging setup while nominal working conditions are maintained on the chip. To build the dynamic thermal model, we apply LongShort-Term-Memory (LSTM) neutral networks with system-level variables such as chip frequency, instruction counts, and other performance metrics as inputs. To reduce the dimensionality of the model, 2D spatial discrete cosine transformation (DCT) is first performed on the heatmaps so that they can be expressed with just their dominant DCT frequencies. Our study shows that only 6×6 DCT coefficients are required to maintain sufficient accuracy across a variety of workloads. Experimental results show that the proposed approach can estimate the full-chip heatmaps with less than 1.4°C root-mean-square-error and take only ~19ms for each inference which suits well for real-time use.
Sheriff Sadiqbatcha, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
ASP-DAC6
2020 Reliable Power Grid Network Design Framework Considering EM Immortalities for Multi-Segment Wires
abstract
This paper presents a new power grid network design and optimization technique that considers the new EM immortality constraint due to EM void saturation volume for multi-segment interconnects. Void may grow to its saturation volume without changing the wire resistance significantly. However, this phenomenon was ignored in existing EM-aware optimization methods. By considering this new effect, we can remove more conservativeness in the EM-aware on-chip power grid design. Along with recently proposed nucleation phase immortality constraint for multi-segment wires, we show that both EM immortality constraints can be naturally integrated into the existing programming based power grid optimization framework. To further mitigate the overly conservative problem of existing immortality-constrained optimization methods, we further explore two strategies: first we size up failed wires to meet one of the immorality conditions subject to design rules; second, we consider the EM-induced aging effects on power supply networks for a target lifetime, which allows some short-lifetime wires to fail and optimizes the rest of the wires. Numerical results on a number of IBM-format power grid networks demonstrate that the new method can reduce more power grid area compared to the existing EM-immortality constrained optimizations. Furthermore, the new method can optimize power grids with nucleated wires, which would not be possible with the existing methods.
Han Zhou 0002, Shuyuan Yu, Zeyu Sun 0001, Sheldon X.-D. Tan
ASP-DAC4
2020 Run-Time Accuracy Reconfigurable Stochastic Computing for Dynamic Reliability and Power Management: Work-in-Progress
abstract
In this paper, we propose a novel accuracy-reconfigurable stochastic computing (ARSC) framework for dynamic reliability and power management. Different than the existing stochastic computing works, where the accuracy versus power/energy trade-off is carried out in the design time, the new ARSC design can change accuracy or bit-width of the data in the run-time so that it can accommodate the long-term aging effects by slowing the system clock frequency at the cost of accuracy while maintaining the throughput of the computing. We validate the ARSC concept on a discrete cosine transformation (DCT) and inverse DCT designs for image compressing/decompressing applications, which are implemented on Xilinx Spartan-6 family XC6SLX45 platform. Experimental results show that the new design can easily mitigate the long-term aging induced effects by accuracy trade-off while maintaining the throughput of the whole computing process using simple frequency scaling. We further show that one-bit precision loss for the input data, which translated to 3.44dB of the accuracy loss in term of Peak Signal to Noise Ratio (PSNR) for images, we can sufficiently compensate the NBTI induced aging effects in 10 years while maintaining the pre-aging computing throughput of 7.19 frames per second. At the same time, we can save 74% power consumption by 10.67dB of accuracy loss. The proposed ARSC computing framework also allows much aggressive frequency scaling, which can lead to order of magnitude power savings compared to the traditional dynamic voltage and frequency scaling (DVFS) techniques.
Shuyuan Yu, Han Zhou 0002, Shaoyi Peng, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
CASES6
2020 Accurate Power Density Map Estimation for Commercial Multi-Core Microprocessors
abstract
In this work, we propose an accurate full chip steady-state power density map estimation method for the commercial multi-core microprocessors. The new approach is based on the measured steady-state thermal maps (images) from an advanced infrared (IR) thermal imaging system to ensure its accuracy. The new method consists of a few steps. First, based on the first principle of heat transfer, 2D spatial Laplace operation is performed on the given thermal map to obtain the so-called raw power density map, which consists of both positive and negative values due to the steady-state nature and boundary conditions of the microprocessors. Then based on the total power of the microprocessor from an online CPU monitoring tool, we develop a novel scheme to generate the actual real positive- only power density map from the raw power density map. At the same time, we develop a novel approach to estimating the effective thermal conductivity of the microprocessors. To further validate the power density map and the estimated actual thermal conductivity of the microprocessors, we construct a thermal model with COMSOL, which mimics the real experimental set up of measurement used in the IR imaging system. Then we compute the thermal maps from the estimated power density maps to ensure the computed thermal maps match the measured thermal maps using FEM method. Experimental results on intel i7-8650U 4-core processor show 1.8°C root-mean-square- error (RMSE) and 96% similarity (2D correlation) between the computed thermal maps and the measured thermal maps.
Sheriff Sadiqbatcha, Wentian Jin, Sheldon X.-D. Tan
DATE4
2020 Full-Chip Thermal Map Estimation for Commercial Multi-Core CPUs with Generative Adversarial Learning
abstract
In this paper, we propose a novel transient full-chip thermal map estimation method for multi-core commercial CPU based on the data-driven generative adversarial learning method. We treat the thermal modeling problem as an image-generation problem using the generative neural networks. In stead of using traditional functional unit powers as input, the new models are directly based on the measurable real-time high level chip utilizations and thermal sensor information of commercial chips without any assumption of additional physical sensors requirement. The resulting thermal map estimation method, called ThermGAN can provide tool-accurate full-chip transient thermal maps from the given performance monitor traces of commercial off-the-shelf multi-core processors. In our work, both generator and discriminator are composed of simple convolutional layers with Wasserstein distance as loss function. ThermGAN can provide the transient and real-time thermal map without using any historical data for training and inferences, which is contrast with a recent RNN-based thermal map estimation method in which historical data is needed. Experimental results show the trained model is very accurate in thermal estimation with an average RMSE of 0.47°C, namely, 0.63% of the full-scale error. Our data further show that the speed of the model is faster than 7.5ms per inference, which is two orders of magnitude faster than the traditional finite element based thermal analysis. Furthermore, the new method is ~4x more accurate than recently proposed LSTM-based thermal map estimation method and has faster inference speed. It also achieves ~2x accuracy with much less computational cost than a state-of-the-art pre-silicon based estimation method.
Wentian Jin, Sheriff Sadiqbatcha, Sheldon X.-D. Tan
ICCAD4
2020 Electromigration Immortality Check considering Joule Heating Effect for Multisegment Wires
abstract
Electromigration (EM) is still the most important reliability concern for VLSI systems, especially at the nanometer regime. EM immortality check is an important step for full-chip EM signoff analysis. In this paper, we propose a new electromigration (EM) immortality check method for multi-segment interconnect considering the impacts of Joule heating induced temperature gradient. Temperature gradients from metal Joule heating, called thermal migration, can be a significant force for the metal atomic migrations, and these impacts get more significant as technology scales down. Compared to existing methods, the new method can consider the spatial temperature gradient due to Joule heating for multi-segment wires for the first time. We derive the analytic solution for the resulting steady-state EM-thermal migration stress distribution problem. Then we develop the new temperature-aware voltage-based EM immortality check method considering the multi-segment temperature migration effects, which carries all the benefits of the recently proposed voltage-based EM immortality method for multi-segment interconnects. Numerical results on an IBM power grid and self synthesized power delivery networks show that the proposed temperature-aware EM immortality check method is much more accurate than recently proposed state of the art EM immortality method.
Mohammadamir Kavousi, Liang Chen 0025, Sheldon X.-D. Tan
ICCAD3
2020 GridNet: Fast Data-Driven EM-Induced IR Drop Prediction and Localized Fixing for On-Chip Power Grid Networks
abstract
Electromigration (EM) is a major failure effect for on-chip power grid networks of deep submicron VLSI circuits. EM degradation of metal grid lines can lead to excessive voltage drops (IR drops) before the target lifetime. In this paper, we propose a fast data-driven EM-induced IR drop analysis framework for power grid networks, named GridNet, based on the conditional generative adversarial networks (CGAN). It aims to accelerate the incremental full-chip EM-induced IR drop analysis, as well as IR drop violation fixing during the power grid design and optimization. More importantly, GridNet can naturally leverage the differentiable feature of deep neural networks (DNN) to obtain the sensitivity information of node voltage with respect to the wire resistance (or width) with marginal cost. Grid-Net treats continuous time and the given electrical features as input conditions, and the EM-induced time-varying voltage of power grid networks as the conditional outputs, which are represented as data series images. We show that GridNet is able to learn the temporal dynamics of the aging process in continuous time domain. Besides, we can take advantage of the sensitivity information provided by GridNet to perform efficient localized IR drop violation fixing in the late stage design and optimization. Numerical results on 36000 synthesized power grid network samples demonstrate that the new method can lead to 105× speedup over the recently proposed full-chip coupled EM and IR drop analysis tool. We further show that localized IR drop violation fix for the same set of power grid networks can be performed remarkably efficiently using the cheap sensitivity computation from GridNet.
Han Zhou 0002, Wentian Jin, Sheldon X.-D. Tan
ICCAD3
2020 EM-GAN: Data-Driven Fast Stress Analysis for Multi-Segment Interconnects
abstract
Electromigration (EM) analysis for complicated interconnects requires the solving of partial differential equations, which is expensive. In this paper, we propose a fast transient hydrostatic stress analysis for EM failure assessment for multisegment interconnects using generative adversarial networks (GANs). Our work is inspired by the image synthesis and feature of generative deep neural networks. The stress evaluation of multi-segment interconnects, modeled by partial differential equations, can be viewed as time-varying 2D-images-to-image problem where the input is the multi-segment interconnects topology with current densities and the output is the EM stress distribution in those wire segments at the given aging time. We show that the conditional GAN can be exploited to attend the temporal dynamics for modeling the time-varying dynamic systems like stress evolution over time. The resulting algorithm, called EM-GAN, can quickly give accurate stress distribution of a general multi-segment wire tree for a given aging time, which is important for full-chip fast EM failure assessment. Our experimental results show that the EM-GAN shows 6.6% averaged error compared to COMSOL simulation results with orders of magnitude speedup. It also delivers 8.3× speedup over state-of-the-art analytic based EM analysis solver.
Wentian Jin, Sheriff Sadiqbatcha, Zeyu Sun 0001, Han Zhou 0002, Sheldon X.-D. Tan
ICCD5
2020 Full-chip wire-oriented back-end-of-line TDDB hotspot detection and lifetime analysis
Shaoyi Peng, Ertugrul Demircan, Mehul D. Shroff, Sheldon X.-D. Tan
Integr.4
2020 Accelerating Electromigration Aging: Fast Failure Detection for Nanometer ICs
abstract
For practical testing and detection of electromigration (EM) induced failures in dual damascene copper interconnects, one critical issue is creating stressing conditions to induce the chip to fail exclusively under EM in a very short period of time so that EM sign-off and validation can be carried out efficiently. Existing acceleration techniques, which rely on increasing temperature and current densities beyond the known limits, also accelerate other reliability effects making it very difficult, if not impossible, to test EM in isolation. In this paper, we propose novel EM wear-out acceleration techniques to address the aforementioned issue. First, we show that multisegment interconnects with reservoir and sink structures can be exploited to significantly speedup the EM wear-out process. Based on this observation, we propose three strategies to accelerate EM induced failure: 1) reservoir-enhanced acceleration; 2) sink-enhanced acceleration; and 3) a hybrid method that combines both reservoir and sink structures. We then propose several configurable interconnect structures that exploit atomic reservoirs and sinks for accelerated EM testing. Such configurable interconnect structures are very flexible and can be used to achieve significant lifetime reductions at the cost of some routing resources. Using the proposed technique, EM testing can be carried out at nominal current densities, and at a much lower temperature compared to traditional testing methods. This is the most significant contribution of this paper since, to our knowledge, this is the only method that allows EM testing to be performed in a controlled environment without the risk of invoking other reliability effects that are also accelerated by elevated temperature and current density. The simulation results show that using the proposed method, we can reduce the EM lifetime of a chip from ten years down to a few hours (about 105× acceleration) under the 150 °C temperature limit, which is sufficient for practical EM testing of typical nanometer CMOS ICs.
Sheriff Sadiqbatcha, Zeyu Sun 0001, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Leakage-Aware Predictive Thermal Management for Multicore Systems Using Echo State Network
abstract
Leakage power is becoming significant in new generation IC chips. As leakage power is nonlinearly related to temperature, it is challenging to manage the thermal behavior of today's multicore systems, since thermal management becomes a nonlinear control problem. In this paper, a new predictive dynamic thermal management (DTM) method with neural network thermal model is proposed to naturally consider the inherent nonlinearity between leakage and temperature. We start with analyzing the problems of using recurrent neural network (RNN) to build the nonlinear thermal model, and point out that there is exploding gradient induced long-term dependencies problem, leading to large model prediction errors. Based on this analysis, we further propose to use echo state network (ESN), which is a special type of RNN, as the leakage-aware nonlinear thermal model. We theoretically and experimentally show that ESN achieves much higher accuracy by completely avoiding the long-term dependencies problem. On top of this nonlinear ESN thermal model, we propose a novel model predictive control (MPC) scheme called ESN MPC, which uses iterative steps to find the optimal future power recommendations for thermal management. Being able to consider the leakage-temperature nonlinear effects and equipped with advanced control technique, the new method achieves an overall high quality temperature management with smooth and accurate temperature tracking. The experimental results show the new method outperforms the state-of-the-art leakage-aware multicore DTM method in both temperature management quality and computing overhead.
Hai Wang 0002, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Fast Analytic Electromigration Analysis for General Multisegment Interconnect Wires
abstract
Electromigration (EM) is considered to be one of the most important reliability issues for current and future ICs in 10-nm technology and below. In this article, we propose a fast analytic solution to compute the stress evolution in the confined multisegment interconnect wires. The new method, called the accelerated separation of variables (ASOV) method, aims to find the analytic solutions of the partial differential equations of stress in confined interconnect metals based on the SOV method. It offers several improvements over the existing plain SOV-based method. First, we show that the accuracy of the solution depends on the structure of the interconnects. As a result, the number of required eigenvalues is structure and problem dependent, instead of fixed numbers used by the existing SOV method. Second, for the straight line multisegment and star-structured multiterminal interconnects, analytical expressions are formulated to calculate the eigenvalues directly instead of using numerical methods as in the existing SOV method. Third, we propose a linear Gaussian elimination (GE) algorithm by exploiting the banded structure with the serrated-edge form of the transcendental matrix, which can significantly speed up GE process, and is the key computing step in the SOV-based solution framework. Fourth, instead of using the simple bisection search, we propose to use an enhanced determinant-based secant iterative method to find the eigenvalues of the transcendental matrix. Numerical results show that a good agreement is achieved between analytical and numerical results on two special cases, and the resulting algorithm can lead to 3-5X speedup over the existing plain SOV-based solution on a number of multisegment interconnects benchmarks.
Liang Chen 0025, Sheldon X.-D. Tan, Zeyu Sun 0001, Shaoyi Peng
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Hot Spot Identification and System Parameterized Thermal Modeling for Multi-Core Processors Through Infrared Thermal Imaging
abstract
Accurate thermal models suitable for system level dynamic thermal, power and reliability regulation and management are vital for many commercial multi-core processors. However, developing such accurate thermal models and identifying the related thermal-power relevant spatial locations for commercial processors is a challenging task due to the lack of information and available tools. Existing tools such as HotSpot-like thermal models may suffer from inaccuracy or inefficiency for online applications, primarily because most rely on parameters that cannot be precisely quantified, such as power-traces, while others are numerical methods not suitable for runtime use. In this work, we propose a novel approach to automatically detecting the major heat-sources on a commercial multi-core microprocessor using an infrared thermal imaging setup. Our approach involves a number of steps including 2D discrete cosine transformation filter for noise reduction on the measured thermal maps, and Laplacian transformation followed by K-mean clustering for heat-source identification. Since the identified heat-sources are the thermally vulnerable areas of the die, we propose a novel approach to deriving a thermal model capable of predicting their temperatures during runtime. We apply Long-Short-Term-Memory (LSTM) networks to build a dynamic thermal model which uses system-level variables such as chip frequency, voltage and instruction count as inputs. The model is trained and tested exclusively using measured thermal data from a commercial multi-core processor. Experimental results show that the proposed thermal model achieves very high accuracy (root-mean-square-error: 2.04°C to 2.57° C) in predicting the temperature of all the identified heat-sources on the chip.
Sheriff Sadiqbatcha, Hengyang Zhao, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
DATE5
2019 Reliability based hardware Trojan design using physics-based electromigration models
Chase Cook, Sheriff Sadiqbatcha, Zeyu Sun 0001, Sheldon X.-D. Tan
Integr.4
2019 GPU-based Ising computing for solving max-cut combinatorial optimization problems
Chase Cook, Hengyang Zhao, Takashi Sato 0001, Masayuki Hiromoto, Sheldon X.-D. Tan
Integr.5
2019 GDP: A Greedy Based Dynamic Power Budgeting Method for Multi/Many-Core Systems in Dark Silicon
abstract
Dark silicon phenomenon is significant in today's multi/many-core systems manufactured using new generation technology. In order to enhance performance of dark silicon systems, power budget constrained dynamic optimizations are performed in various ways including dynamic voltage and frequency scaling (DVFS) and task scheduling. However, power budgets given by existing methods are generally over pessimistic, which greatly limit the capability of dynamic performance optimization methods. In order to resolve this problem, we propose a dynamic power budgeting method, called Greedy based Dynamic Power (GDP). Different from existing methods, which are steady state based and ignore active core distributions, GDP formulates the power budgeting problem as a thermal-constrained combinational power optimization problem. To efficiently solve this problem, we propose two new ideas: first, we transform the original power-optimization problem to an easier solving temperature-optimization problem; second, we employ a more efficient greedy based algorithm that finds a sub-optimal active core distribution which maximizes power budget. The new method can consider current temperature states and transient thermal effects, which were ignored by existing methods. Both theoretical studies and experimental results show that GDP outperforms existing methods by providing a higher and less pessimistic power budget with low computing cost and guaranteed thermal safety.
Hai Wang 0002, Diya Tang, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030
IEEE Trans. Computers4
2019 Saturation-Volume Estimation for Multisegment Copper Interconnect Wires
abstract
Recently, many physics-based electromigration (EM) models have been proposed to mitigate the over-conservativeness of existing Black-Blech-based EM models. To assess EM failures, one needs to estimate how the void grows after the nucleation phase. As a result, it is important to estimate the saturation volume of a void, which is important for EM mortality check. The existing saturation void volume model only works for a single wire segment. In this paper, we propose a new model for fast estimation of the void's saturation volume for general multisegment interconnect wires. The new model is based on the fundamental atom conservation at the steady-state condition of void growth phases. The new formula agrees with the existing saturation void volume formula for the single-segment wire case and is a natural extension of the single-segment case to general multisegment wires. In addition, we consider the impacts of the void volume on final stress distributions of the wire to further improve the accuracy of the proposed formula. Based on the new formula, we propose a new EM immortality check flow, which considers both the recently proposed EM immortality in the void nucleation phase and the void saturation volume in the growth phase. The new flow can further reduce the conservativeness of the existing EM failure effect analysis. The numerical results show that the proposed formula agrees well with a published work for two-segment cases, which are supported by experimental data. The formula is also validated by the recently proposed physics-based 3-D finite-element (FEM) analysis tool for general multisegment interconnect wires. We also demonstrate new EM immortality check flow to quickly identify the new type of immortal wires, which are nucleated but with smaller-than-critical voids.
Zeyu Sun 0001, Sheriff Sadiqbatcha, Hengyang Zhao, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.4
2019 EM-Aware and Lifetime-Constrained Optimization for Multisegment Power Grid Networks
abstract
This paper proposes a new power-ground (P/G) network sizing technique based on the recently proposed fast electromigration (EM) immortality check method for general multisegment interconnect wires and a new physics-based EM assessment technique for more accurate time to failure analysis. This paper first shows that the new P/G optimization problem, subject to the voltage IR drop and new EM constraints, can still be formulated as an efficient sequence of linear programing problem, where the optimization is carried out in two linear programing phases in each iteration. The new optimization will ensure that none of the wires fail if all the constraints are satisfied. However, requiring all the wires to be EM immortal can be overconstrained. To mitigate this problem, the first improvement is by means of adding reservoir branches to the mortal wires whose lifetime cannot be made immortal by wire sizing. This is a very effective approach as long as there is a sufficient reservoir area. The second improvement is to consider the aging effects of interconnect wires in the P/G networks. The idea is to allow some short-lifetime wires to fail and optimize the rest of the wires while considering the additional resistance caused by the failed wire segments. In this way, the resulting P/G networks can be optimized, such that the target lifetime of the whole P/G networks can be ensured and will become more robust and aging-aware over the expected lifetime of the chip. Numerical results on a number of IBM and self-generated power supply networks demonstrate that the new method can effectively reduce the area of the networks while ensuring immortality or enforcing target lifetime for all the wires, which is not the case for the existing current-density-constrained optimization methods.
Han Zhou 0002, Zeyu Sun 0001, Sheriff Sadiqbatcha, Naehyuck Chang, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.5
2018 Accelerating electromigration aging for fast failure detection for nanometer ICs
abstract
For practical testing and detection of electromigration (EM) induced failures in dual damascene copper interconnects in today's and future sub-10nm ICs, one critical issue is how to create stressing conditions so that the chip will fail exclusively under EM in a very short period of time so that EM signoff and validation can be carried out efficiently. In this work, we propose novel EM wearout-acceleration techniques for practical VLSI chips. We will first review the recently proposed three-phase physics-based EM models and discuss the important factors contributing to the EM aging process. Then we propose a new formula for fast estimation of the void's saturation volume for general multi-segment interconnect wires, which is important for EM mortality check. We then investigate two strategies to accelerate the EM failure process: reservoir-enhanced acceleration and temperature-based acceleration. First we show that multi-segment interconnects with reservoir structures and their stressing currents can be exploited to significantly speedup the EM wearout process. Such configurable reservoir-based wires are very flexible and can achieve various EM accelerations at the costs of some routing resources. Additionally, we show that further acceleration can be achieved by increasing temperature. On average, 10% increase in temperature yields about 10X wearout acceleration. However, purely temperature based acceleration is not possible since practical VLSI chips have temperature limitations which must be strictly enforced to ensure the chip only fails under EM, and not due to other reliability effects. In this study, we show that it is possible to achieve significantly high acceleration while staying within the feasible operating zones by combining the two acceleration techniques. Experimental results show that by combining temperature and reservoir accelerations, we can reduce the EM lifetime of a chip from 10 years down to a few hours (about 105acceleration) under the 150°C temperature limit, which is sufficient for practical EM testing of typical nanometer CMOS ICs.
Zeyu Sun 0001, Sheriff Sadiqbatcha, Hengyang Zhao, Sheldon X.-D. Tan
ASP-DAC4
2018 Electromigration-lifetime constrained power grid optimization considering multi-segment interconnect wires
abstract
Electromigration (EM) remains the top killer of copperbased damascene interconnects in 10nm and beyond technologies. On-chip power/ground (P/G) networks are most vulnerable to EM failures due to large and unidirectional current flows, thus proper sizing and even routing of power grid networks are critical for the full-chip EM sign-off. In this paper, we propose a new P/G network sizing technique based on recently proposed fast EM immortality check method for general multi-segment interconnect wires and physics-based EM assessment technique for fast time to failure analysis. We first show that the new P/G optimization problem subjected to voltage IR drop and new EM constraints can still be formulated as a sequence of linear programming (SLP) problem. The new optimization will ensure that all the wires will not fail if all the constraints are satisfied. To consider EM-induced aging effects on power supply networks for the target lifetime and mitigate the over-conservation of the first optimization formulation, we further propose an aging-aware P/G optimization method, which allows some short-lived wires to fail or to age and optimizes the rest of the wires considering resistance increase of those failed wire segments. In this way, the P/G networks can be optimized more effectively and become more robust and aging-aware. Numerical results on a number of IBM and self-generated power supply networks show that the new approach can effectively reduce the area of the network while ensuring immortality or improving target lifetime of all the wires, which is not the case for the existing current density constrained optimization method.
Han Zhou 0002, Yijing Sun, Zeyu Sun 0001, Hengyang Zhao, Sheldon X.-D. Tan
ASP-DAC5
2018 Multi-physics-based FEM analysis for post-voiding analysis of electromigration failure effects
abstract
In this paper, we propose a new multi-physics finite element method (FEM) based analysis method for void growth simulation of confined copper interconnects. This new method for the first time considers three important physics simultaneously in the EM failure process and their time-varying interactions: the hydrostatic stress in the confined interconnect wire, the current density and Joule heating induced temperature. As a result, we end up with solving a set of coupled partial differential equations which consist of the stress diffusion equation (Korhonen's equation), the phase field equation (for modeling void boundary move), the Laplace equation for current density and the heat diffusion equation for Joule heating and wire temperature. In the new method, we show that each of the physics will have different physical domains and differential boundary conditions, and how such coupled multi-physics transient analysis was carried out based on FEM and different time scales are properly handled. Experiment results show that by considering all three coupled physics - the stress, current density, and temperature - and their transient behaviors, the proposed FEM EM solver can predict the unique transient wire resistance change pattern for copper interconnect wires, which were well observed by the published experiment data. We also show that the simulated void growth speed is less conservative than recently proposed compact EM model.
Hengyang Zhao, Sheldon X.-D. Tan
ICCAD2
2018 Detection of counterfeited ICs via on-chip sensor and post-fabrication authentication policy
Taeyoung Kim 0001, Sheldon X.-D. Tan, Chase Cook, Zeyu Sun 0001
Integr.2
2018 Recent advances in EM and BTI induced reliability modeling, analysis and optimization (invited)
Sheldon X.-D. Tan, Hussam Amrouch, Taeyoung Kim 0001, Zeyu Sun 0001, Chase Cook, Jörg Henkel
Integr.1
2018 A Fast Leakage-Aware Full-Chip Transient Thermal Estimation Method
abstract
Accurate and fast thermal estimation is important for the runtime thermal regulation of modern microprocessors due to excessive on-chip temperatures. However, due to the nonlinear relationship between the leakage power and temperature, full-chip thermal estimation methods suffer slow speed and scalability issue when the increasing static leakage power is considered. In this work, we propose a new fast leakage-aware full-chip thermal estimation method. Unlike traditional methods, which use iteration to handle the leakage-temperature nonlinearity dependency issue, the new method applies a dynamic linearization algorithm, which adaptively transforms the original nonlinear thermal model into a number of local linear thermal models. In order to further improve the thermal estimation efficiency, a specially-designed adaptive model order reduction method is integrated into the thermal estimation framework to generate local compact thermal models. Our numerical results show that the new method is able to accurately estimate full-chip transient temperature distribution by fully considering the nonlinear leakage-temperature dependency with fast speed. On different chips with core number ranging from 9 to 36, it achieved 85x to 589x speedup in average against traditional iteration based method, with average thermal estimation error to be around 0.2°C.
Hai Wang 0002, Jiachun Wan, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Yuan Yuan 0030, Keheng Huang, Zhenghong Zhang
IEEE Trans. Computers3
2018 Fast Electromigration Immortality Analysis for Multisegment Copper Interconnect Wires
abstract
In this paper, we present a novel and fast electromigration (EM) immortality check for general multisegment interconnect wires. Instead of using current density as the key parameter, as in traditional EM analysis methods based on Black's equation and the Blech limit, the new method estimates the EM-induced steady-state stress in general multisegment copper interconnect wires based on a novel parameter, Critical EM Voltage, VCrit,EM. We show that the VCrit,EM is essentially the natural, but important, extension of the Blech limit concept, which describes the EM immortality condition for a single segment wire, to more general multisegment interconnect wires. The proposed method, called voltage-based EM (VBEM) method, mitigates the problem of current-density-based EM criteria, which can only be applied to a single wire. The new VBEM method can naturally comprehend the impact of the topology of the wire structure on EM-induced stress. As a result, this new VBEM analysis method is very amenable to addressing EM violations, as it brings new optimization capabilities to the physical design flow. The VBEM stress estimation method is based on the fundamental steady-state stress equations. This approach avoids computationally intensive numerical methods and can be implemented in CAD tools very easily, as we demonstrate on real design examples. We also show that the proposed VBEM analysis method agrees with results from the finite difference method in the steady state through one example and also agrees with one published closed-form expression of steady-state stress for a special 3-terminal wire case. Furthermore, we compare VBEM against the COMSOL finite element analysis tool and another published EM numerical simulator XSim, validated by measured results, which shows that VBEM agrees with both of them very well in terms of accuracy and thus further validates the proposed method. We also study the impact of current crowding in practical interconnect wires on the estimated steady-state stress, which are shown to be not significant if the length of the wire is much greater than its width. An extension of the VBEM method to consider the significant current crowding effects is also shown and additionally, we analyze mesh-structured interconnect wires and demonstrate that the proposed VBEM method is correct and accurate on such structures.
Zeyu Sun 0001, Ertugrul Demircan, Mehul D. Shroff, Chase Cook, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2018 Thermal-Sensor-Based Occupancy Detection for Smart Buildings Using Machine-Learning Methods
abstract
In this article, we propose a novel approach to detect the occupancy behavior of a building through the temperature and/or possible heat source information. The new method can be used for energy reduction and security monitoring for emerging smart buildings. Our work is based on a building simulation program, EnergyPlus, from the Department of Energy. EnergyPlus can model various time-series inputs to a building such as ambient temperature; heating, ventilation, and air-conditioning (HVAC) inputs; power consumption of electronic equipment; lighting; and number of occupants in a room, sampled each hour, and produce resulting temperature traces of zones (rooms). Two machine-learning-based approaches for detecting human occupancy of a smart building are applied herein, namely support vector regression (SVR) and recurrent neural network (RNN). Experimental results with SVR show that the four-feature model provides accurate detection rates, giving a 0.638 average error and 5.32% error rate, and the five-feature model delivers a 0.317 average error and 2.64% error rate. This indicates that SVR is a viable option for occupancy detection. In the RNN method, Elman’s RNN can estimate occupancy information of each room of a building with high accuracy. It has local feedback in each layer and, for a five-zone building, it is very accurate for occupancy behavior estimation. The error level, in terms of number of people, can be as low as 0.0056 on average and 0.288 at maximum, considering ambient, room temperatures, and HVAC powers as detectable information. Without knowing HVAC powers, the estimation error can still be 0.044 on average, and only 0.71% estimated points have errors greater than 0.5. Our article further shows that both methods deliver similar accuracy in the occupancy detection. But the SVR model is more stable for adding or removing features of the system, while the RNN method can deliver more accuracy when the features used in the model do not change a lot.
Hengyang Zhao, Qi Hua, Haibao Chen, Yaoyao Ye, Hai Wang 0002, Sheldon X.-D. Tan, Esteban Tlelo-Cuautle
ACM Trans. Design Autom. Electr. Syst.6
2018 Fast Electromigration Stress Evolution Analysis for Interconnect Trees Using Krylov Subspace Method
Chase Cook, Zeyu Sun 0001, Ertugrul Demircan, Mehul D. Shroff, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.5
2018 Physics-Based Compact TDDB Models for Low-k BEOL Copper Interconnects With Time-Varying Voltage Stressing
abstract
Time-dependent dielectric breakdown (TDDB) is one of the important failure mechanisms for copper (Cu) interconnects. This problem becomes more severe as the pitch between wires is shrinking and low-k dielectric materials (low electrical and mechanical strength) are used. Many TDDB models have been proposed based on different physics kinetics in the past. Recently, a physics-based TDDB model, which is based on the breakdown concept of electric path generation, has been proposed and has shown advantage over widely accepted existing electrostatic field-based TDDB assessment. However, determination of the time-to-failure from this model includes time-consuming finite-element method (FEM). In this paper, we try to mitigate this problem by developing fast time to failure evaluation method based on the closed form solution of the ion diffusion partial differential equations. We show that the location of the minimum concentration can be determined by the dominant terms with sufficient accuracy and the time to failure can also be computed with a few dominant terms. On top of this, we also consider the time-varying stressing voltages, which is commonly seen in practical VLSI chips. We propose to develop the equivalent dc stressing voltage, which is parameterized in terms of amplitude, duty cycle, and period for periodic stressing voltage waveforms using regression-based method. We further validate the proposed analytic TDDB concentration and time to failure formula, and the equivalent dc stressing voltage compact model against the results of an FEM analysis using COMSOL. Numerical results further show that the new compact TDDB model can lead to three orders of magnitude speedup with less than 1% error against the existing FEM results.
Shaoyi Peng, Han Zhou 0002, Taeyoung Kim 0001, Haibao Chen, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.5
2018 Recovery-Aware Proactive TSV Repair for Electromigration Lifetime Enhancement in 3-D ICs
abstract
Electromigration (EM) becomes a major reliability concern in 3-D integrated circuits (3-D ICs). To mitigate this problem, a typical solution is to use through-silicon via (TSV) redundancy in a reactive manner, maintaining the operability of a 3-D chip in the presence of EM failures by detecting and replacing faulty TSVs with spares. In this paper, we explore an alternative, more preferred approach to enhance the EM-related lifetime reliability of TSV grid, in which redundancy is used proactively to allow nonfaulty TSVs to be temporarily deactivated. In this way, EM wear-out can be extended by exploiting its recovery property. The proposed solution is based on two consecutive stages, in which TSV redundancy allocation and TSV repair are finalized at both design-time and runtime, respectively. Applied to 3-D benchmark designs, the recovery-aware proactive repair approach increases EM-related lifetime reliability (measured in mean-time-to-failure) of the entire TSV grid by up to 12× relative to the conventional reactive method, with similar area overhead. In addition, a runtime dynamic recovery approach is proposed to further improve EM-related lifetime reliability to account for stress variation across different chips and over the operational lifetime.
Shengcheng Wang, Taeyoung Kim 0001, Zeyu Sun 0001, Sheldon X.-D. Tan, Mehdi Baradaran Tahoori
IEEE Trans. Very Large Scale Integr. Syst.4
2018 Postvoiding FEM Analysis for Electromigration Failure Characterization
Hengyang Zhao, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Physics-based electromigration modeling and assessment for multi-segment interconnects in power grid networks
abstract
Electromigration (EM) is considered to be one of the most important reliability issues for current and future ICs in 10nm technology and below. In this paper we focus on the EM stress evaluation for one-dimensional multi-segment interconnect wires in which all the segments have the same direction, which is a common routing structure for power grid networks. The proposed method, which is based on integral transform technique, could efficiently calculate the hydrostatic stress evolution for multi-segment metal wires stressed with different current densities. The new method can also naturally consider the pre-existing residual stresses coming from thermal or other stress sources. Based on this new transient EM assessment method, a full-chip assessment algorithm for power grid networks is then proposed. The new algorithm is also based on the IR-drop metrics for failure assessment of the power grid networks. However, it finds the precise location and time of EM-induced void nucleation by directly checking the time-changing hydrostatic stresses of all the wires. The resulting EM assessment method can ensure sufficient accuracy of the EM verification for large scale power grid networks without sacrificing the efficiency. The accuracy of the proposed transient analysis approach is validated against the numerical analysis. Also the resulting EM-aware full-chip power grid reliability analysis has been demonstrated and compared with existing methods.
Sheldon X.-D. Tan, Yici Cai, Shengqi Yang
DATE4
2017 Recovery-aware proactive TSV repair for electromigration in 3D ICs
abstract
Electromigration (EM) becomes a major reliability concern in three-dimensional integrated-circuits (3D ICs). To mitigate this problem, a typical solution is to use TSV redundancy in a reactive manner, maintaining the operability of a 3D chip in the presence of EM failures by detecting and replacing faulty TSVs with spares. In this work, we explore an alternative, more preferred approach to enhance the EM-related lifetime reliability of TSV grid, in which redundancy is used proactively to allow non-faulty TSVs to be temporarily deactivated. In this way, EM wear-out can be reversed by exploiting its recovery property. Applied to 3D benchmark designs, the recovery-aware proactive repair approach increases EM-related lifetime reliability (measured in mean-time-to-failure) of the entire TSV grid by up to 12X relative to the conventional reactive method, with less area overhead.
Shengcheng Wang, Hengyang Zhao, Sheldon X.-D. Tan, Mehdi Baradaran Tahoori
DATE3
2017 Leveraging recovery effect to reduce electromigration degradation in power/ground TSV
abstract
With increasing temperature and current density, electromigration (EM) becomes a major interconnect reliability challenge in power distribution networks (PDNs) of three-dimensional integrated-circuits (3D ICs). In order to improve the EM reliability of power/ground (P/G) through-silicon-vias (TSVs), the conventional solution is to use larger TSVs in order to decrease the current densities. In this work we exploit the recovery effects for EM reliability improvement by periodically deactivating P/G TSVs. In order to predict EM-related lifetime for TSV accurately, a novel three-phase EM model is proposed with a focus on single damascene via-last process. Different from existing TSV EM models, the new TSV EM model considers the nucleation phase and the impacts of initial thermo-mechanical stress, which is significant for the TSVs in addition to this recovery effect modeling. Furthermore, a recovery-aware repair architecture is developed for EM reliability improvement. Applied to 3D benchmark designs, the proposed repair approach increases EM-related lifetime of the P/G TSV grid by 4.4X in average relative to the conventional TSV sizing method, with negligible area overhead.
Shengcheng Wang, Zeyu Sun 0001, Sheldon X.-D. Tan, Mehdi Baradaran Tahoori
ICCAD4
2017 Fast physics-based electromigration analysis for multi-branch interconnect trees
abstract
Electromigration (EM) becomes one of the most challenging reliability issues for current and future ICs in 10nm technology and below. In this paper, we propose a new analyses method for the EM hydrostatic stress evolution for multi-branch interconnect trees, which is the foundation of the EM reliability assessment for large scale on-chip interconnect networks, such as power grid networks. The proposed method, which is based on eigenfunctions technique, could efficiently calculate the hydrostatic stress evolution for multi-branch interconnect trees stressed with different current densities and non-uniformly distributed thermal effects. The new method can also accommodate the pre-existing residual stresses coming from thermal or other stress sources. The proposed method solves the partial differential equations of EM stress more efficiently since it does not require any discretization either spatially or temporally, which is in contrast to numerical methods such as finite difference method and finite element method. The accuracy of the proposed transient analysis approach is validated against the analytical solution and commercial tools. The efficiency of the proposed method is demonstrated and compared to finite difference method. The proposed method is 10X~100X times faster than finite difference method and scales better for larger interconnect trees.
Sheldon X.-D. Tan, Chase Cook, Shengqi Yang
ICCAD4
2017 Dynamic electromigration modeling for transient stress evolution and recovery under time-dependent current and temperature stressing
Xin Huang 0003, Valeriy Sukharev, Taeyoung Kim 0001, Sheldon X.-D. Tan
Integr.4
2017 Editorial
abstract
As I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design.
Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.37
2017 Energy and Lifetime Optimizations for Dark Silicon Manycore Microprocessor Considering Both Hard and Soft Errors
abstract
In this paper, we propose a new energy and lifetime optimization techniques for emerging dark silicon manycore microprocessors considering both hard long-term reliability effects (hard errors) and transient soft errors, which have been studied less in the past. We consider a recently proposed physics-based electromigration (EM) reliability model to predict the EM-induced reliability. We employ both dynamic voltage and frequency scaling (DVFS) and dark silicon core state using ON/OFF switching action as the two control knobs. We show that on-chip power consumption has different (even contradicting) impacts on soft and hard reliability effects. This paper also shows that soft error should be mitigated by other techniques if aggressive low power and high long-term reliability are pursued. We focus on two optimization techniques for improving lifetime and reducing energy. To optimize EM-induced lifetime, we first apply the adaptive Q-learning-based method, which is suitable for dynamic runtime operation as it can provide cost-effective yet good solutions. The second lifetime optimization approach is the mixed-integer linear programming (MILP) method, which typically yields better solutions but at higher computational costs. To optimize the energy of a dark silicon chip subject to the both hard and soft reliability effects, power budgets, and performance limits, the Q-learning method has been applied as well. A large class of multithreaded applications is used as our benchmarks to validate and compare the proposed dynamic reliability management methods. Experimental results on a 64-core dark silicon chip show that the proposed DRM algorithm can effectively manage and optimize the lifetime of a dark silicon microprocessor under the given power budget and performance limit. Also, the proposed energy optimization can effectively manage and optimize energy consumption subject to both hard and soft-error rates, power budget, and performance limits as constraints. We also show that the under tightened power and performance constraints, we cannot satisfy both hard and soft errors at the same time as there is no simple tradeoff between performance/power and reliability in this case. Some other soft-error mitigation techniques are required in this case.
Taeyoung Kim 0001, Zeyu Sun 0001, Haibao Chen, Hai Wang 0002, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Electromigration recovery modeling and analysis under time-dependent current and temperature stressing
abstract
Electromigration (EM) has been considered to be the major reliability issue for current and future VLSI technologies. Current EM reliability analysis is overloaded by over-conservative and simplified EM models. Particularly the transient recovery effect in the EM-induced stress evolution kinetics has never been treated properly in all the existing analytical EM models. In this article, we propose a new physics-based dynamic compact EM model, which for the first time, can accurately predict the transient hydrostatic stress recovery effect in a confined metal wire. The new dynamic EM model is based on the direct analytical solution of one-dimensional Korhonen's equation with load driven by any unipolar or bipolar current waveforms under varying temperature. We show that the EM recovery effect can be quite significant even under unidirectional current loads. This healing process is sensitive to temperature, and higher temperatures lead to faster and more complete recovery. Such effect can be further exploited to significantly extend the lifetime of the interconnect wires if the chip current or power can be properly regulated and managed. As a result, the new dynamic EM model can be incorporated with existing dynamic thermal/power/reliability management and optimization approaches, devoted to reliability-aware optimization at multiple system levels (chip/server/rack/data centers). Presented results show that the proposed EM model agrees very well with the numerical analysis results under any time-varying current density and temperature profiles.
Xin Huang 0003, Valeriy Sukharev, Taeyoung Kim 0001, Haibao Chen, Sheldon X.-D. Tan
ASP-DAC5
2016 Thermal modeling for energy-efficient smart building with advanced overfitting mitigation technique
abstract
Building energy accounts large amount of the total energy consumption, and smart building energy control leads to high energy efficiency and significant energy savings. A compact and accurate building thermal model is important for designing the efficient energy control system. In this paper, we propose an accurate thermal behavior modeling technique for general and complicated buildings. This new modeling technique builds compact thermal model by system identification using temperature and power data obtained from EnergyPlus software, which can provide realistic temperature, weather and power data for buildings. In order to make the best use of data from EnergyPlus and avoid the overfitting problem associated with the system identification method, a cross-validation technique is employed to generate multiple thermal models to find the optimal model order. The final model is then generated by performing a regular system identification using the previously selected order. Experimental results from a case study of a 5-zone building have shown that the proposed method is able to find the optimal model order, and the building models built by the proposed method can achieve 1-3% average errors and less than 10-18% maximum errors for the estimation of zone temperatures for about a one year period.
Wandi Liu, Hai Wang 0002, Hengyang Zhao, Shujuan Wang, Haibao Chen, Yuzhuo Fu, Jian Ma 0002, Xin Li 0001, Sheldon X.-D. Tan
ASP-DAC9
2016 Physics-based full-chip TDDB assessment for BEOL interconnects
abstract
As technology advances, Time-Dependent Dielectric Breakdown (TDDB) has become one of the major reliability threats for Copper/low-k interconnects. This article presents a novel approach, techniques, and flow for the physics-based chip-scale assessment of backend low-k TDDB. In our work, the breakdown development is considered as the complementary combination of electric current path generation by means of diffusing metal ions and field-based hoping conductivity of the current carriers. It replaces the widely accepted across-layout electrostatic field based TDDB assessment. As a result, the model generated time-to-failure (TTF) is governed by kinetics of the electric current path generation, which is controlled by a time-dependent minimum metal ion concentration in the inter-metal dielectrics (IMD) gap-fill. Finite element analysis (FEA)-based simulations are used for populating the set of lookup tables, which provide a time to breakdown for any interconnect pattern with given geometries and voltages. A pattern-matching technique is used for extracting from the layout all patterns belonging to different classes of pattern shapes with different geometries, locations and electric loads. Experimental results obtained on a test chip show that upon the calibration the proposed flow provides a capability to evaluate chip-scale low-k TDDB reliability based on the calculated TTF and detect most leaking shapes in the layout.
Xin Huang 0003, Valeriy Sukharev, Zhongdong Qi, Taeyoung Kim 0001, Sheldon X.-D. Tan
DAC5
2016 Invited - Cross-layer modeling and optimization for electromigration induced reliability
abstract
In this paper, we propose a new approach for cross-layer electromigration (EM) induced reliability modeling and optimization at physics, system and datacenter levels. We consider a recently proposed physics-based electromigration (EM) reliability model to predict the EM reliability of full-chip power grid networks for long-term failures. We show how the new physics-based dynamic EM model at the physics level can be abstracted at the system level and even at the datacenter level. Our datacenter system-level power model is based on the BigHouse simulator. To speed up the online optimization for energy in a datacenter, we propose a new combined datacenter power and reliability compact model using a learning based approach in which a feed-forward neural network (FNN) is trained to predict energy and long term reliability for each processor under datacenter scheduling and workloads. To optimize the energy and reliability of a datacenter, we apply the efficient adaptive Q-learning based reinforcement learning method. Experimental results show that the proposed compact models for the datacenter system trained with different workloads under different cluster power modes and scheduling policies are able to build accurate energy and lifetime. Moreover, the proposed optimization method effectively manages and optimizes data-center energy subject to reliability, given power budget and performance.
Taeyoung Kim 0001, Zeyu Sun 0001, Chase Cook, Hengyang Zhao, Ruiwen Li, Daniel Wong 0001, Sheldon X.-D. Tan
DAC7
2016 Learning-based dynamic reliability management for dark silicon processor considering EM effects
Taeyoung Kim 0001, Xin Huang 0003, Haibao Chen, Valeriy Sukharev, Sheldon X.-D. Tan
DATE5
2016 Dynamic reliability management for near-threshold dark silicon processors
abstract
In this article, we propose a new dynamic reliability management (DRM) techniques at the system level for emerging low power dark silicon manycore microprocessors operating in near-threshold region. We mainly consider the electromigration (EM) failures. To leverage the EM recovery effects, which was ignored in the past, at the system-level, we propose a new equivalent DC current model to consider recovery effects for general time-varying current waveforms so that existing compact EM model can be applied. The new equivalent DC current is calculated in two steps: firstly, the equivalent square waveform is calculated so that peak and terminal stresses are matched, secondly, the parameterized equivalent DC current is derived in terms of the parameters of the periodic fitted square waveforms from the first step. The new recovery EM model can allow EM-induced lifetime to be better managed at the system level. The system level energy optimization problem considering EM lifetime subject to power and performance constraints is framed by seeking the best dark silicon cores' voltage and on/off status. The resulting problem is solved by the State-Action-Reward-State-Action (SARSA) reinforcement learning algorithm. Experimental results on a 64-core near-threshold dark silicon processor show that the new equivalent EM DC currents can fully exhibit the recovery effects at the system-level so that trade-off between EM lifetime and energy/performance can be easily made. We further show that the proposed learning-based energy optimization can effectively manage and optimize energy subject to reliability, given power budget and performance limits. When the recovery effects are considered, the new optimization method can achieve 8.6× longer lifetime at the costs of 2.0× more energy and 3.3× more performance degradation.
Taeyoung Kim 0001, Zeyu Sun 0001, Chase Cook, Jagadeesh Gaddipati, Hai Wang 0002, Haibao Chen, Sheldon X.-D. Tan
ICCAD7
2016 Voltage-based electromigration immortality check for general multi-branch interconnects
abstract
As VLSI technology features are pushed to the limit with every generation and with the introduction of new materials and increased current densities to satisfy the performance demands, Electromigration (EM) is projected to be a key reliability issue for current and future VLSI technologies. Existing EM signoff mainly relies on current density-based assessment using Black's equation and Blech product. This model does not work well for multi-branch interconnect wires as the stresses developed in each wire segment is not independent of one another. In this paper, we present a novel and fast EM Immortality check for general multi-branch interconnect trees. Instead of using current density as the key parameter as in traditional methods, the new method estimates the EM-induced stress in general multi-branch interconnects based on the terminal voltages or potentials. It can be viewed as the Blech product for multi-branch interconnects for fast check of EM immortality of wires. Besides, this voltage-based EM (VBEM) assessment technique can naturally comprehend the impact of the topology of the wire structure on EM-induced stress. As a result, this new VBEM analysis method is very amenable to EM violation fixing as it brings new capabilities to the physical design stage. The VBEM stress estimation method is based on the fundamental steady-state stress equations. This approach eliminates the need for complex look-up tables for different geometries and can be implemented in CAD tools very easily as we demonstrate on real design examples. We show that its solution is consistent with the physics-based dynamic EM stress evaluations from the numerical analysis by COMSOL.
Zeyu Sun 0001, Ertugrul Demircan, Mehul D. Shroff, Taeyoung Kim 0001, Xin Huang 0003, Sheldon X.-D. Tan
ICCAD6
2016 Learning-based occupancy behavior detection for smart buildings
abstract
In this article, we propose a novel method to detect the occupancy behavior of a building through the temperature and/or possible heat source information, which can be used for energy reduction, security monitoring for emerging smart buildings. Our work is based on a realistic building simulation program, EnergyPlus, from Department of Energy. EnergyPlus can model the various time-series inputs to a building such as ambient temperature, heating, ventilation, and air-conditioning (HVAC) inputs, power consumption of electronic equipment, lighting and number of occupants in a room sampled in each hour and produce resulting temperature traces of zones (rooms). The new approach is based on a learning based approach in which a recurrent neutral network (RNN) is trained to detect the number of people in a room based on the room temperature and other information such as ambient temperature, and other related heat sources. We applied the Elman's recurrent neural network (ELNN), which has local feedbacks in each layer. We use an empirical formula to calculate the RNN layer number and layer size to configure RNN architecture to avoid overfitting and under-fitting problems. Experimental results from a case study of a 5-zone building show that ELNN can lead to very accurate occupancy behavior estimation. The error level, in terms of number of people, can be as low as 0.0056 on average and 0.288 at maximum when we consider ambient, room temperatures and HVAC powers as detectable information. Without knowing HVAC powers, estimation error can still be 0.044 on average, and only 0.71% estimated points have errors greater than 0.5.
Hengyang Zhao, Zhongdong Qi, Shujuan Wang, Kambiz Vafai, Hai Wang 0002, Haibao Chen, Sheldon X.-D. Tan
ISCAS7
2016 Parallel GMRES solver for fast analysis of large linear dynamic systems on GPU platforms
Sheldon X.-D. Tan, Hengyang Zhao, Xuexin Liu, Hai Wang 0002, Guoyong Shi
Integr.2
2016 Electromigration assessment for power grid networks considering temperature and thermal stress effects
Xin Huang 0003, Valeriy Sukharev, Jun-Ho Choy, Marko Chew, Taeyoung Kim 0001, Sheldon X.-D. Tan
Integr.6
2016 Editorial: Special Issue on The 14th International Conference on Computer-Aided Design and Computer Graphics (CAD/Graphics 2015)
Xin Li 0001, Sheldon X.-D. Tan, Yu Wang 0002
Integr.2
2016 Analytical Modeling and Characterization of Electromigration Effects for Multibranch Interconnect Trees
abstract
Electromigration (EM) in very large scale integration (VLSI) interconnects has become one of the major reliability issues for current and future VLSI technologies. However, existing EM modeling and analysis techniques are mainly developed for a single wire. For practical VLSI chips, the elemental EM reliability unit called interconnect tree is a multibranch interconnect segment consisting of a continuously connected, highly conductive metal (Cu) lines terminated by diffusion barriers and located within the single level of metallization. The EM effects in those branches are not independent and have to be considered simultaneously. In this paper, we demonstrate, for the first time, a first principle-based analytical solution of this problem. We have derived the analytical expressions describing the hydrostatic stress evolution in several typical interconnect trees: 1) the straight-line three-terminal wires; 2) the T-shaped four-terminal wires; and 3) the cross-shaped five-terminal wires. The new approach solves the stress evolution in a multibranch tree by de-coupling the individual segments through the proper boundary conditions (BCs) accounting the interactions between different branches. By using Laplace transformation technique, analytical solutions are obtained for each type of the interconnect trees. The analytical solutions in terms of a set of auxiliary basis functions using the complementary error function agree well with the numerical analysis results. Our analysis further demonstrates that using the first two dominant basis functions can lead to 0.5% error, which is sufficient for practical EM analysis.
Haibao Chen, Sheldon X.-D. Tan, Xin Huang 0003, Taeyoung Kim 0001, Valeriy Sukharev
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Physics-Based Electromigration Models and Full-Chip Assessment for Power Grid Networks
abstract
This paper presents a novel approach and techniques for physics-based electromigration (EM) assessment in power delivery networks of very large scale integration systems. An increase in the voltage drop above the threshold level, caused by EM-induced increase in resistances of the individual interconnect branches, is considered as a failure criterion. It replaces a currently employed conservative weakest branch criterion, which does not account an essential redundancy for current propagation existing in the power-ground (P/G) networks. EM-induced increase in the resistance of the individual grid branches is described in the approximation of the recently developed physics-based formalism for void nucleation and growth. An approach to calculation of the void nucleation times in the group of branches comprising the interconnect tree is implemented. As a result, P/G networks become time-varying linear networks. A developed technique for calculating the hydrostatic stress evolution inside a multibranch interconnect tree allows to avoid over optimistic prediction of the time-to-failure made with the Blech-Black analysis of individual branches of interconnect tree. Experimental results obtained on a number of International Business Machines Corporation benchmark circuits show that the proposed method will lead to less conservative estimation of the lifetime than the existing Black-Blech-based methods. It also reveals that the EM-induced failure is more likely to happen at the place where the hydrostatic stress predicted by the initial current density is large and is more likely to happen at longer times when the saturated void volume effect is taken into account.
Xin Huang 0003, Armen Kteyan, Sheldon X.-D. Tan, Valeriy Sukharev
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Hierarchical Dynamic Thermal Management Method for High-Performance Many-Core Microprocessors
abstract
It is challenging to manage the thermal behavior of many-core microprocessors while still keeping them running at high performance since the control complexity increases as the core number increases. In this article, a novel hierarchical dynamic thermal management method is proposed to overcome this challenge. The new method employs model predictive control (MPC) with task migration and a DVFS scheme to ensure smooth control behavior and negligible computing performance sacrifice. In order to be scalable to many-core systems, the hierarchical control scheme is designed with two levels. At the lower level, the cores are spatially clustered into blocks, and local task migration is used to match current power distribution with the optimal distribution calculated by MPC. At the upper level, global task migration is used with the unmatched powers from the lower level. A modified iterative minimum cut algorithm is used to assist the task migration decision making if the power number is large at the upper level. Finally, DVFS is applied to regulate the remaining unmatched powers. Experiments show that the new method outperforms existing methods and is very scalable to manage many-core microprocessors with small performance degradation.
Hai Wang 0002, Jian Ma 0002, Sheldon X.-D. Tan, Chi Zhang 0029, He Tang 0003, Keheng Huang, Zhenghong Zhang
ACM Trans. Design Autom. Electr. Syst.3
2016 Statistical Rare-Event Analysis and Parameter Guidance by Elite Learning Sample Selection
abstract
Accurately estimating the failure region of rare events for memory-cell and analog circuit blocks under process variations is a challenging task. In this article, we propose a new statistical method, called EliteScope , to estimate the circuit failure rates in rare-event regions and to provide conditions of parameters to achieve targeted performance. The new method is based on the iterative blockade framework to reduce the number of samples, but consists of two new techniques to improve existing methods. First, the new approach employs an elite-learning sample-selection scheme, which can consider the effectiveness of samples and well coverage for the parameter space. As a result, it can reduce additional simulation costs by pruning less effective samples while keeping the accuracy of failure estimation. Second, the EliteScope identifies the failure regions in terms of parameter spaces to provide a good design guidance to accomplish the performance target. It applies variance-based feature selection to find the dominant parameters and then determine the in-spec boundaries of those parameters. We demonstrate the advantage of our proposed method using several memory and analog circuits with different numbers of process parameters. Experiments on four circuit examples show that EliteScope achieves a significant improvement on failure-region estimation in terms of accuracy and simulation cost over traditional approaches. The 16b 6T-SRAM column example also demonstrates that the new method is scalable for handling large problems with large numbers of process variables.
Taeyoung Kim 0001, Hosoon Shin, Sheldon X.-D. Tan, Xin Li 0001, Haibao Chen, Hai Wang 0002
ACM Trans. Design Autom. Electr. Syst.4
2016 Corrections to "GPU-Accelerated Parallel Sparse LU Factorization Method for Fast Circuit Analysis"
abstract
In the above-named work, Fig. 5 has an error (see the old Fig. 5). The new Fig. 5 is shown here. The following is the detailed explanation for the correction.
Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.2
2016 GPU-Accelerated Parallel Sparse LU Factorization Method for Fast Circuit Analysis
abstract
Lower upper (LU) factorization for sparse matrices is the most important computing step for circuit simulation problems. However, parallelizing LU factorization on the graphic processing units (GPUs) turns out to be a difficult problem due to intrinsic data dependence and irregular memory access, which diminish GPU computing power. In this paper, we propose a new sparse LU solver on GPUs for circuit simulation and more general scientific computing. The new method, which is called GPU accelerated LU factorization (GLU) solver (for GPU LU), is based on a hybrid right-looking LU factorization algorithm for sparse matrices. We show that more concurrency can be exploited in the right-looking method than the left-looking method, which is more popular for circuit analysis, on GPU platforms. At the same time, the GLU also preserves the benefit of column-based left-looking LU method, such as symbolic analysis and columnlevel concurrency. We show that the resulting new parallel GPU LU solver allows the parallelization of all three loops in the LU factorization on GPUs. While in contrast, the existing GPU-based left-looking LU factorization approach can only allow parallelization of two loops. Experimental results show that the proposed GLU solver can deliver 5.71χ and 1.46x speedup over the single-threaded and the 16-threaded PARDISO solvers, respectively, 19.56x speedup over the KLU solver, 47.13x over the UMFPACK solver, and 1.47x speedup over a recently proposed GPU-based left-looking LU solver on the set of typical circuit matrices from the University of Florida (UFL) sparse matrix collection. Furthermore, we also compare the proposed GLU solver on a set of general matrices from the UFL, GLU achieves 6.38x and 1.12x speedup over the singlethreaded and the 16-threaded PARDISO solvers, respectively, 39.39x speedup over the KLU solver, 24.04x over the UMFPACK solver, and 2.35x speedup over the same GPU-based left-looking LU solver. In addition, comparison on self-generated RLC mesh networks shows a similar trend, which further validates the advantage of the proposed method over the existing sparse LU solvers.
Sheldon X.-D. Tan, Hai Wang 0002, Guoyong Shi
IEEE Trans. Very Large Scale Integr. Syst.2
2015 New electromigration modeling and analysis considering time-varying temperature and current densities
abstract
Electromigration (EM) is projected to be the major reliability issue for current and future VLSI technologies. However, existing EM models and assessment techniques are mainly based on the constant current density and temperature. Such models will not work well at the system level as the current density (power) and temperature are changing with time due to different tasks (their loans) applied at run time. Existing EM approaches using average current density or temperature, however, will lead to significant errors as shown in this work. In this paper, we propose a new physics-based EM model considering time-varying temperature and current density, which reflects a more practical chip working conditions especially for multi-core and emerging 3D ICs. We study the impacts of the time-varying current densities and temperature profiles on EM-induced lifetime of a wire for both nucleation phase and growth phase. We propose a fast stress calculation method for given time-varying temperature and current densities for the nucleation phase. We further develop new formulae to compute the resistance changes in growth phase due to changing temperature and current densities. Experimental results show that the proposed method shows an excellent agreement with the detailed numerical analysis but with much improved efficiency.
Haibao Chen, Sheldon X.-D. Tan, Xin Huang 0003, Valeriy Sukharev
ASP-DAC2
2015 GPU-accelerated parallel Monte Carlo analysis of analog circuits by hierarchical graph-based solver
abstract
In this article, we propose a new parallel matrix solver, which is very amenable for Graphic Process Unit (GPU) based fine-grain massively-threaded parallel computing. The new method is based on the graph-based symbolic analysis technique to generate the computing sequence of determinants in terms of determinant decision diagrams (DDDs). DDD represents very simple data dependence and data parallelism, which can be explored much easier by GPU massively-threaded parallel computing than existing LU-based methods. The new method is based on the hierarchical determinant decision diagrams (HDDDs). Inspired by the inherent data parallelism and simple data dependence in the evaluation process of HDDD, we design GPU-amenable continuous data structures to enable fast memory access and evaluation of massive parallel threads. In addition to parallelism in DDD graph, the new algorithm can naturally explore data independence existing in Monte Carlo and frequency domain analysis. The resulting algorithm is a general-purpose matrix solver suitable for fine-grain massive GPU-based computing for any circuit matrices. Experimental results show that the new evaluation algorithm can achieve about two orders of magnitude speedup over the serial CPU based evaluation and more than 4× speedup over numerical SPICE-based simulation method on some large analog circuits.
Sheldon X.-D. Tan
ASP-DAC2
2015 Interconnect reliability modeling and analysis for multi-branch interconnect trees
abstract
Electromigration (EM) in VLSI interconnects has become one of the major reliability issues for current and future VLSI technologies. However, existing EM modeling and analysis techniques are mainly developed for a single wire. For practical VLSI chips, the interconnects such as clock and power grid networks typically consist of multi-branch metal segments representing a continuously connected, highly conductive metal (Cu) lines within one layer of metallization, terminating at diffusion barriers. The EM effects in those branches are not independent and they have to be considered simultaneously. In this paper, we demonstrate, for the first time, a first principle based analytic solution of this problem. We investigate the analytic expressions describing the hydrostatic stress evolution in several typical interconnect trees: the straight-line 3-terminal wires, the T-shaped 4-terminal wires and the cross-shaped 5-terminal wires. The new approach solves the stress evolution in a multi-branch tree by de-coupling the individual segments through the proper boundary conditions accounting the interactions between different branches. By using Laplace transformation technique, analytical solutions are obtained for each type of the interconnect trees. The analytical solutions in terms of a set of auxiliary basis functions using the complementary error function agree well with the numerical analysis results. Our analysis further demonstrates that using the first two dominant basis functions can lead to 0.5% error, which is sufficient for practical EM analysis.
Haibao Chen, Sheldon X.-D. Tan, Valeriy Sukharev, Xin Huang 0003, Taeyoung Kim 0001
DAC2
2015 From Robust Chip to Smart Building: CAD Algorithms and Methodologies for Uncertainty Analysis of Building Performance
abstract
Buildings consume about 40% of the total energy use in the U.S. and, hence, accurately modeling, analyzing and optimizing building energy is considered as an extremely important task today. Towards this goal, uncertainty/sensitivity analysis has been proposed to identify the critical physical and environmental parameters contributing to building energy consumption. In this paper, we propose to apply sparse regression techniques to uncertainty/sensitivity analysis of smart buildings. We consider the orthogonal matching pursuit (OMP) algorithm as a case study to demonstrate its superior efficacy over other conventional approaches. Experimental results reveal that OMP achieves up to 18.6× runtime speedups over the conventional least-squares fitting method without surrendering any accuracy.
Xiaoming Chen 0003, Xin Li 0001, Sheldon X.-D. Tan
ICCAD3
2015 EM-Based on-Chip Aging Sensor for Detection and Prevention of Counterfeit and Recycled ICs
abstract
The counterfeiting and recycled integrated circuits (ICs) has become a major security threat for commercial and military systems. In addition to the huge economic impacts, they post significant security and safety threats on those systems. In this paper, we propose a new lightweight on-chip aging sensor, which is based on the electromigration (EM)-induced aging effects for fast detection and prevention of recycled ICs. Our new EM-based aging sensor exploits the natural aging/failure mechanism of interconnect wires to time the aging of the chip. Compared with existing aging sensor, the new aging sensor can provide more accurate prediction of the chip usage time at smaller area footprints due to its simple structure. The new sensor is based on a newly proposed physics-based stress evolution model of EM effects for accurate prediction of the EM failure. As a result, we can design the interconnect wire structures based on copper interconnect technology so that the resulting wires will have detectable EM failure at a specific time with sufficient accuracy. In order to mitigate the problem of the inherent variations in the metal grain sizes and assess its impacts on the nucleation time of metal wires, a number of parallel properly structured wires are used in the sensor. The parameters of the wires are optimized using the new EM model. Our statistical and variational analysis shows that the proposed aging sensor can accurately predict the targeted failure times in the presence of both inherent uncertainties. Our study also shows that more parallel wires will lead to more accurate statistical predictions at costs of more areas.
Xin Huang 0003, Sheldon X.-D. Tan
ICCAD3
2015 Learning Based Compact Thermal Modeling for Energy-Efficient Smart Building Management: (invited)
abstract
In this article, we propose a new behavioral thermal modeling method for fast building performance analysis, which is critical for energy-efficient smart building control and management. The new approach is based on two recurrent neutral network architecture to obtain the compact nonlinear thermal models for complicated building. We start with a more realistic building simulation program, EnergyPlus, from Department of Energy, to model some practical buildings such as office buildings and data centers. EnergyPlus can model the various time-series inputs to a building such as ambient temperature, heating, ventilation, and air-conditioning (HVAC) inputs, power consumption of electronic equipment, lighting and number of occupants in a room sampled in each hour and produce resulting temperature traces of zones (rooms). In this work, we apply two recurrent neural network (RNN) architectures to build the non-linear compact thermal model of the building: one is non-linear state-space RNN architecture (NLSS), which has global feedbacks, and the other one is Elman's RNN architecture (ELNN), which has local feedbacks in each layer. We give a simple formula to calculate the RNN layer number, layer size to configure RNN architecture to avoid overfitting and underfitting problems. A cross-validation based training technique is further applied to improve predictable accuracy of models. Experimental results from a case study of three buildings show that ELNN and NLSS can both build very accurate building thermal models for the 2-zone and 5-zone building cases: both of them have average errors from around 1% to 1.5% for the two buildings. For the more complex 6-zone building case, ELNN outperforms NLSS with maximum errors 16% against 23%. But both methods have 2.2% average errors.
Hengyang Zhao, Daniel Quach, Shujuan Wang, Hai Wang 0002, Haibao Chen, Xin Li 0001, Sheldon X.-D. Tan
ICCAD7
2015 H-Matrix-Based Finite-Element-Based Thermal Analysis for 3D ICs
abstract
In this article, we propose an efficient finite-element-based (FE-based) method for both steady and transient thermal analyses of high-performance integrated circuits based on the hierarchical matrix ( H -matrix) representation. H -matrix has been shown to provide a data-sparse way to approximate the matrices and their inverses with almost linear-space and time complexities. In this work, we apply the H -matrix concept for solving heating diffusion problems modeled by parabolic partial differential equations (PDEs) based on the finite element method. We show that the matrix from a FE-based steady and transient thermal analysis can be represented by H -matrix without any approximation, and its inverse and Cholesky factors can be evaluated by H -matrix with controlled accuracy. We then show and prove that the memory and time complexities of the solver are bounded by O ( k 1 N log N ) and O ( k 1 2 N log 2 N ), respectively, where k 1 is a small quantity determined by accuracy requirements and N is the number of unknowns in the system. The comparison with existing product-quality LU solvers, CSPARSE and UMFPACK, on a number of 3D IC thermal matrices, shows that the new method is much more memory efficient than these methods, which however prevents CPU time comparison with those methods on large examples. But the proposed method can solve all the given thermal circuits with decent scalabilities, which shows good agreement with the predicted theoretical results.
Haibao Chen, Ying-Chi Li, Sheldon X.-D. Tan, Xin Huang 0003, Hai Wang 0002, Ngai Wong 0001
ACM Trans. Design Autom. Electr. Syst.3
2015 Task Migrations for Distributed Thermal Management Considering Transient Effects
abstract
In this brief, a new distributed thermal management scheme using task migrations based on a new temperature metric called effective initial temperature is proposed to reduce the on-chip temperature variance and the occurrence of hot spots for many-core microprocessors. The new temperature metric derived from frequency domain moment matching technique incorporates both initial temperature and other transient effects to make optimized task migration decisions, which leads to more effective reduction of hot spots in the experiments on a 100-core microprocessor than the existing distributed thermal management methods.
Zao Liu, Sheldon X.-D. Tan, Xin Huang 0003, Hai Wang 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2015 A GPU-Accelerated Parallel Shooting Algorithm for Analysis of Radio Frequency and Microwave Integrated Circuits
abstract
This paper presents a new parallel shooting-Newton method based on a graphic processing unit (GPU)-accelerated periodic Arnoldi shooting solver (GAPAS) for fast periodic steady-state analysis of radio frequency/millimeter-wave integrated circuits. The new algorithm first explores a periodic structure of the state matrix by using a periodic Arnoldi algorithm for computing the resulting structured Krylov subspace in the generalized minimal residual (GMRES) solver. The resulting periodic Arnoldi shooting method is very amenable for massive parallel computing, such as GPUs. Second, the periodic Arnoldi-based GMRES solver in the shooting-Newton method is parallelized on the recent NVIDIA Tesla GPU platforms. We further explore CUDA GPUs features, such as coalesced memory access and overlapping transfers with computation to boost the efficiency of the resulting parallel GAPAS method. Experimental results from several industrial examples show that when compared with the state-of-the-art implicit GMRES method under the same accuracy, the new parallel shooting-Newton method can lead up to $8\times$ speedup.
Xuexin Liu, Hao Yu 0001, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Parallel Thermal Analysis of 3-D Integrated Circuits With Liquid Cooling on CPU-GPU Platforms
abstract
In this brief, we propose an efficient parallel finite difference-based thermal simulation algorithm for 3-D-integrated circuits (ICs) using generalized minimum residual method (GMRES) solver on CPU-graphic processing unit (GPU) platforms. First, the new method starts from basic physics-based heat equations to model 3-D-ICs with intertier liquid cooling microchannels and directly solves the resulting partial differential equations. Second, we develop a new parallel GPU-GMRES solver to compute the resulting thermal systems on a CPU-GPU platform. We also explore different preconditioners (implicit and explicit) and study their performances on thermal circuits and other types of matrices. Experimental results show the proposed GPU-GMRES solver can deliver orders of magnitudes speedup over the parallel LU-based solver and up to 4× speedup over CPU-GMRES for both dc and transient thermal analyzes on a number of thermal circuits and other published problems.
Xuexin Liu, Kuangya Zhai, Zao Liu, Sheldon X.-D. Tan, Wenjian Yu
IEEE Trans. Very Large Scale Integr. Syst.5
2014 Time-domain performance bound analysis for analog and interconnect circuits considering process variations
abstract
Time-Domain worst case or performance bound estimation for analog integrated circuits and interconnect circuits are crucial for both analog and digital circuit design and optimization in the presence of process variations. In this paper, we present a novel non-Monte-Carlo (MC) performance bound analysis technique in time domain. The new method consists of several steps. First the symbolic transient modified nodal analysis (MNA) formulation of the circuit matrices of (linearized) analog and interconnect circuits at a time step is formed. Then the closed-form expressions of the interested performance in terms of variational parameters of the circuit matrices of (linearized) analog and interconnect circuits are derived via a graph-based symbolic analysis method. Then time-domain performance response bound of current time step are obtained by a nonlinear constrained optimization process subject to the parameter variations and variational circuit state bounds computed from the previous time step. We study the bounds computed by the proposed against the different sigma bounds by the standard MC method, which shows that the proposed method is more efficient for computing high sigma bounds than the MC method. Experimental results show that the new method can deliver order of magnitudes speedup over the standard Monte Carlo simulation on some typical analog circuits and interconnect circuits with high accuracy.
Sheldon X.-D. Tan, Yici Cai, Puying Tang
ASP-DAC2
2014 Physics-based Electromigration Assessment for Power Grid Networks
abstract
This paper presents a novel approach and techniques for physics-based electromigration (EM) assessment in power delivery networks of VLSI systems. An increase in the voltage drop above the threshold level, caused by EM-induced increase in resistances of the individual interconnect segments, is considered as a failure criterion. It replaces a currently employed conservative weakest segment criterion, which does not account an essential redundancy for current propagation existing in the power-ground (p/g) networks. EM-induced increase in the resistance of the individual grid segments is described in the approximation of the recently developed physics-based formalism for void nucleation and growth. A statistical approach to calculation of the void nucleation times in the group of branches comprising the interconnect tree is implemented. As a result, p/g networks become time-varying linear networks. A developed technique for calculating the hydrostatic stress evolution inside a multi-branch interconnect tree allows to avoid over optimistic prediction of the time-to-failure (TTF) made with the Blech-Black analysis of individual branches of interconnect tree. Experimental results obtained on a number of IBM benchmark circuits validate the proposed methodology.
Xin Huang 0003, Valeriy Sukharev, Sheldon X.-D. Tan
DAC4
2014 Battery Management and Application for Energy-Efficient Buildings
abstract
As the building stock consumes 40% of the U.S. primary energy consumption, it is critically important to improve building energy efficiency. This involves reducing the total energy consumption of buildings, reducing the peak energy demand, and leveraging renewable energy sources, etc. To achieve such goals, hybrid energy supply has becoming popular, where multiple energy sources such as grid electricity, on-site fuel cell generators, solar, wind, and battery storage are scheduled together to improve energy efficiency.
Tianshu Wei, Taeyoung Kim 0001, Sangyoung Park, Qi Zhu 0002, Sheldon X.-D. Tan, Naehyuck Chang, Sadrul Ula, Mehdi Maasoumy
DAC5
2014 Lifetime optimization for real-time embedded systems considering electromigration effects
abstract
In this article, we propose a new lifetime task optimization technique for real-time embedded processors considering the electromigration-induced reliability. The new approach is based on a recently proposed physics-based electromigration (EM) model for more accurate EM assessment of a power grid network at the chip level. We apply the dynamic voltage and frequency scaling (DVFS) (by selecting the performance states or p-states of the tasks to manage the power) and thus the lifetime of the processor running different tasks over their periods. We consider both single-rate and multi-rate embedded systems with preemption. To model the mean-time-to-failure (MTTF) of a task for a given p-state, response surface modeling is applied. We then frame the reliability optimization problem as the continuous constrained nonlinear optimization problem in which the system EM-induced reliability is maximized subject to the timing constraints, which is further solved by simulated annealing method. Experimental results show that for low utilization systems, significant reliability improvement can be achieved with even smaller power consumption than existing reliability-ignore scheduling method. The proposed method can lead to near Pareto's front trade-off between the power/energy and the lifetime compared to the existing task scheduling method.
Taeyoung Kim 0001, Bowen Zheng 0001, Haibao Chen, Qi Zhu 0002, Valeriy Sukharev, Sheldon X.-D. Tan
ICCAD6
2014 IR-drop based electromigration assessment: parametric failure chip-scale analysis
abstract
This paper presents a novel approach and techniques for electromigration (EM) assessment in power delivery networks. An increase in the voltage drop above the threshold level, caused by EM-induced increase in resistances of the individual interconnect segments, is considered as a failure criterion. This criterion replaces a currently employed conservative weakest segment criterion, which does not account an essential redundancy for current propagation existing in the power-ground (p/g) networks. EM-induced increase in the resistance of the individual grid segments is described in the approximation of the physics-based formalism for void nucleation and growth. A developed technique for calculating the hydrostatic stress distribution inside a multi branch interconnect tree allows to avoid over optimistic prediction of the time to failure made with the Blech-Black analysis of individual branches of interconnect segment. Experimental results obtained on the IBM benchmark circuit validate the proposed methods.
Valeriy Sukharev, Xin Huang 0003, Haibao Chen, Sheldon X.-D. Tan
ICCAD4
2014 Compact thermal modeling for packaged microprocessor design with practical power maps
Zao Liu, Sheldon X.-D. Tan, Hai Wang 0002, Yingbo Hua, Ashish Gupta 0007
Integr.2
2014 Compact Lateral Thermal Resistance Model of TSVs for Fast Finite-Difference Based Thermal Analysis of 3-D Stacked ICs
abstract
Thermal issue is the leading design constraint for 3-D stacked integrated circuits (ICs) and through silicon vias (TSVs) are used to effectively reduce the temperature of 3-D ICs. Normally, TSV is considered as a good thermal conductor in its vertical direction, and its vertical thermal resistance has been well modeled. However, lateral heat transfer of TSVs, which is also important, was largely ignored in the past. In this paper, we propose an accurate physics-based model for lateral thermal resistance of TSVs in terms of physical and material parameters, and study the conditions for model accuracy. For TSV arrays or farm, we show that the space or pitch between TSVs has a significant impact on TSV thermal behavior and should be properly considered in the TSV models. The proposed lateral thermal resistance model is fully compatible with the existing modeling approaches, and thus we could build a more accurate complete TSV thermal model. The new TSV thermal model can be easily integrated into a finite difference (FD) based thermal analysis framework to improve analysis efficiency. The accuracy of the model is validated against a commercial finite element tool-COMSOL. Experimental results show that the improved TSV thermal model (with proposed lateral thermal model) could greatly improve the accuracy of FD method in thermal simulation comparing with the existing method.
Zao Liu, Sahana Swarup, Sheldon X.-D. Tan, Haibao Chen, Hai Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Performance bound and yield analysis for analog circuits under process variations
abstract
Yield estimation for analog integrated circuits are crucial for analog circuit design and optimization in the presence of process variations. In this paper, we present a novel analog yield estimation method based on performance bound analysis technique in frequency domain. The new method first derives the transfer functions of linear (or linearized) analog circuits via a graph-based symbolic analysis method. Then frequency response bounds of the transfer functions in terms of magnitude and phase are obtained by a nonlinear constrained optimization technique. To predict yield rate, bound information are employed to calculate Gaussian distribution functions. Experimental results show that the new method can achieve similar accuracy while delivers 20 times speedup over Monte Carlo simulation of HSPICE on some typical analog circuits.
Xuexin Liu, Adolfo Adair Palma-Rodriguez, Santiago Rodriguez-Chavez, Sheldon X.-D. Tan, Esteban Tlelo-Cuautle, Yici Cai
ASP-DAC4
2013 Compact nonlinear thermal modeling of packaged integrated systems
abstract
This paper proposes a new thermal nonlinear modeling technique for packaged integrated systems. Thermal behavior of complicated systems like packaged electronic systems may exhibit nonlinear and temperature dependent properties. As a result, it is difficult to use a low order linear model to approximate the thermal behavior of the packaged integrated systems without accuracy loss. In this paper, we try to mitigate this problem by using piecewise linear (PWL) approach to characterizing the thermal behavior of those systems. The new method (called ThermSubPWL), which is the first proposed approach to nonlinear thermal modeling problem, identifies the linear local models for different temperature ranges using the subspace identification method. A linear transformation method is proposed to transform all the identified linear local models to the common state basis to build the continuous piecewise linear model. Experimental results validate the proposed method on a realistic packaged integrated system modeled via the multi-domain/physics commercial tool, COMSOL, under practical power signal inputs. The new piecewise models can lead to much smaller model order without accuracy loss, which translates to significant savings in both the simulation time and the time required to identify the reduced models compared to applying the high order models.
Zao Liu, Sheldon X.-D. Tan, Hai Wang 0002, Sahana Swarup, Ashish Gupta 0007
ASP-DAC2
2013 Dynamic thermal management for multi-core microprocessors considering transient thermal effects
abstract
Dynamic thermal management method is a viable way to effectively mitigate the thermal emergences. In this paper, a new thermal management scheme is proposed to reduce the on-chip temperature variance and the occurrence of hot spots by considering more transient thermal effects. The new method performs the task migrations to reduce the temperature variations across the chip. Instead of intuitively assigning the heavy tasks to the low temperature cores to balance the thermal profile based on steady state thermal analysis, the proposed method applies moment matching based transient thermal analysis techniques for fast thermal estimation and prediction to guide the migration process. We show that by considering the dominant temperature moment component, the resulting algorithm can lead to significant reduction of hot spots without full transient thermal simulation. Our experimental results on a 16 core microprocessor demonstrate that the proposed method can reduce the number of the hot spots by 50% compared to the simple lowest temperature based task scheduling method, leading to more uniform on-chip temperature distribution across the microprocessor cores.
Zao Liu, Tailong Xu, Sheldon X.-D. Tan, Hai Wang 0002
ASP-DAC3
2013 A power-driven thermal sensor placement algorithm for dynamic thermal management
abstract
On-chip physical thermal sensors play a vital role for accurately estimating the full-chip thermal profile. How to place physical sensors such that both the number of thermal sensors and the temperature estimation errors are minimized becomes important for on-chip dynamic thermal management of today's high-performance microprocessors. In this paper, we present a new systematic thermal sensor placement algorithm. Different from the traditional thermal sensor placement algorithms where only the temperature information is explored, the new placement method takes advantage of functional unit power information by exploiting the correlation of power estimation errors among functional blocks. The new power-driven placement algorithm applies the correlation clustering algorithm to determine both the locations of sensors and the number of sensors automatically such that the temperature estimation errors can be minimized. Experimental results on a dual-core architecture show that the new thermal sensor placements yield more accurate full-chip temperature estimation compared to the uniform and the k-means based placement approaches.
Hai Wang 0002, Sheldon X.-D. Tan, Sahana Swarup, Xuexin Liu
DATE2
2013 Compact lateral thermal resistance modeling and characterization for TSV and TSV array
abstract
Thermal issues are among the major concerns for 3D stacked ICs, and Through silicon vias (TSVs) are used to effectively reduce the temperature of 3D ICs. Normally, TSV is considered as a good thermal conductor in its vertical direction, and its vertical thermal resistance has been studied extensively. However, lateral heat transfer of TSVs, which is also important, was largely ignored in the past. In this paper, we propose an accurate physics-based model for lateral resistance of TSVs in terms of physical and material parameters, and discuss the conditions valid for model accuracy. In addition to modeling the lateral thermal resistance of a single TSV, the proposed thermal model is also applicable to TSV arrays or TSV farms. We show that the TSV insulation linear and space between TSVs could impose a significant impact on TSV thermal behavior. The new TSV thermal model can be easily integrated into a finite difference based thermal analysis framework to improve analysis efficiency. The accuracy of the model is validated against a commercial finite element tool - COMSOL. Experimental results show that the proposed TSV lateral thermal resistance model is very accurate for both a single TSV and TSV arrays.
Zao Liu, Sahana Swarup, Sheldon X.-D. Tan
ICCAD3
2013 Parallel power grid analysis using preconditioned GMRES solver on CPU-GPU platforms
abstract
In this paper, we propose an efficient parallel dynamic linear solver, called GPU-GMRES, for transient analysis of large power grid networks. The new method is based on the preconditioned generalized minimum residual (GMRES) iterative method implemented on heterogeneous CPU-GPU platforms. The new solver is very robust and can be applied to power grids with different structures and other applications like thermal analysis. The proposed GPU-GMRES solver adopts the very general and robust incomplete LU (ILU) based preconditioner. We show that by properly selecting the right amount of fill-ins in the incomplete LU factors, a good trade-off between GPU efficiency and GMRES convergence rate can be achieved for the best overall performance. Such a tunable feature makes this algorithm very adaptive to different problems. Furthermore, we properly partition the major computing tasks in GMRES solver to minimize the data traffic between CPU and GPU, which further boosts performance of the proposed method. Experimental results on the set of published IBM benchmark circuits and mesh-structured power grid networks show that the GPU-GMRES solver can deliver order of magnitudes speedup over the direct LU solver UMFPACK. GPU-GMRES can also deliver 3-10× speedup over the CPU implementation of the same GMRES method on transient analysis.
Xuexin Liu, Hai Wang 0002, Sheldon X.-D. Tan
ICCAD3
2013 Statistical full-chip total power estimation considering spatially correlated process variations
Zhigang Hao, Sheldon X.-D. Tan, Guoyong Shi
Integr.2
2013 Performance bound analysis of analog circuits in frequency- and time-domain considering process variations
abstract
In this article, we propose a new performance bound analysis of analog circuits considering process variations. We model the variations of component values as intervals measured from tested chips and manufacture processes. The new method first applies a graph-based analysis approach to generate the symbolic transfer function of a linear(ized) analog circuit. Then the frequency response bounds (maximum and minimum) are obtained by performing nonlinear constrained optimization in which magnitude or phase of the transfer function is the objective function to be optimized subject to the ranges of process variational parameters. The response bounds given by the optimization-based method are very accurate and do not have the over-conservativeness issues of existing methods. Based on the frequency-domain bounds, we further develop a method to calculate the time-domain response bounds for any arbitrary input stimulus. Experimental results from several analog benchmark circuits show that the proposed method gives the correct bounds verified by Monte Carlo analysis while it delivers one order of magnitude speedup over Monte Carlo for both frequency-domain and time-domain bound analyses. We also show analog circuit yield analysis as an application of the frequency-domain variational bound analysis.
Xuexin Liu, Sheldon X.-D. Tan, Adolfo Adair Palma-Rodriguez, Esteban Tlelo-Cuautle, Guoyong Shi
ACM Trans. Design Autom. Electr. Syst.2
2013 Composable thermal modeling and simulation for architecture-level thermal designs of multicore microprocessors
abstract
Efficient temperature estimation is vital for designing thermally efficient, lower power and robust integrated circuits in nanometer regime. Thermal simulation based on the detailed thermal structures no longer meets the demanding tasks for efficient design space exploration. The compact and composable model-based simulation provides a viable solution to this difficult problem. However, building such thermal models from detailed thermal structures was not well addressed in the past. In this article, we propose a new compact thermal modeling technique, called ThermComp , standing for thermal modeling with composable modules. ThermComp can be used for fast thermal design space exploration for multicore microprocessors. The new approach builds the composable model from detailed structures for each basic module using the finite difference method and reduces the model complexity by the sampling-based model order reduction technique. These composable models are then used to assemble different multicore architecture thermal models and realized into SPICE-like netlists. The resulting thermal models can be simulated by the general circuit simulator SPICE. ThermComp tries to preserve the accuracy of fine-grained models with the speed of coarse-grained models. Experimental results on a number of multicore microprocessor architectures show the new approach can easily build accurate thermal systems from compact composable models for fast architecture thermal analysis and optimization and is much faster than the existing HotSpot method with similar accuracy.
Hai Wang 0002, Sheldon X.-D. Tan, Ashish Gupta 0007, Yuan Yuan 0030
ACM Trans. Design Autom. Electr. Syst.2
2013 Symbolic Moment Computation for Statistical Analysis of Large Interconnect Networks
abstract
The shrinking technology feature size and dense large-scale integration make process variation a challenging issue directly confronting the latest design automation tools. Process variation causes severe variation in interconnect networks, including very large-scale integrated interconnect structures, such as clock trees, clock mesh, power-ground networks, and other wiring structures in 3-D integrated circuits. The traditional moment computation techniques are only partly useful for analyzing such variational problems, however, their computational efficiency cannot meet the quickly rising needs, such as statistical analysis. This paper presents a novel symbolic moment calculator (SMC) for variational interconnect analysis. The moment calculator is constructed in a regular data structure that incorporates binary decision diagrams for data storage and computation. Given an interconnect circuit, such a computation diagram has to be constructed only once and can be repeatedly invoked for computation of moments with varying parameter values. Also, the SMC is friendly to interconnect synthesis in that it can be incrementally modified according to the modifications made to the circuit structure. Applications of the SMC for fast moment computation, sensitivity analysis, and statistical timing analysis are addressed. Significant efficiency is demonstrated comparing to other existing methods.
Zhigang Hao, Guoyong Shi, Sheldon X.-D. Tan, Esteban Tlelo-Cuautle
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Time-domain performance bound analysis of analog circuits considering process variations
abstract
In this paper, we propose a new time-domain performance bound analysis method for analog circuits considering process variations. The proposed method, called TIDBA, consists of several steps to compute the bound performances in time domain. First the performance bound in frequency domain is computed for a linearized analog circuits by an variational symbolic analysis method and the Kharitonov's functions. Then the time domain performance bound is computed via a new general-signal transient bound analysis method. The new algorithm can give transient lower bound and upper bound of the performance variations affected analog circuits accurately and reliably. Experimental results from two industry benchmark circuits show that TIDBA gives the correct bounds for the Monte Carlo analysis while it delivers one order of magnitude speedup over the Monte Carlo method.
Xuexin Liu, Sheldon X.-D. Tan, Zhigang Hao, Guoyong Shi
ASP-DAC2
2012 Parallel statistical analysis of analog circuits by GPU-accelerated graph-based approach
abstract
In this paper, we propose a new parallel statistical analysis method for large analog circuits using determinant decision diagram (DDD) based graph technique based on GPU platforms. DDD-based symbolic analysis technique enables exact symbolic analysis of vary large analog circuits. But we show that DDD-based graph analysis is very amenable for massively threaded based parallel computing based on GPU platforms. We design novel data structures to represent the DDD graphs in the GPUs to enable fast memory access of massive parallel threads for computing the numerical values of DDD graphs. The new method is inspired by inherent data parallelism and simple data independence in the DDD-based numerical evaluation process. Experimental results show that the new evaluation algorithm can achieve about one to two order of magnitudes speedup over the serial CPU based evaluations and 2-3 times speedup over numerical SPICE-based simulation method on some large analog circuits.
Xuexin Liu, Sheldon X.-D. Tan, Hai Wang 0002
DATE2
2012 A GPU-accelerated envelope-following method for switching power converter simulation
abstract
In this paper, we propose a new envelope-following parallel transient analysis method for the general switching power converters. The new method first exploits the parallelisim in the envelope-following method and parallelize the Newton update solving part, which is the most computational expensive, in GPU platforms to boost the simulation performance. To further speed up the iterative GMRES solving for Newton update equation in the envelope-following method, we apply the matrix-free Krylov basis generation technique, which was previously used for RF simulation. Last, the new method also applies more robust Gear-2 integration to compute the sensitivity matrix instead of traditional integration methods. Experimental results from several integrated on-chip power converters show that the proposed GPU envelope-following algorithm leads to about 10× speedup compared to its CPU counterpart, and 100× faster than the traditional envelop-following methods while still keeps the similar accuracy.
Xuexin Liu, Sheldon X.-D. Tan, Hai Wang 0002, Hao Yu 0001
DATE2
2012 Runtime power estimator calibration for high-performance microprocessors
abstract
Accurate runtime power estimation is important for on-line thermal/power regulation on today's high performance processors. In this paper, we introduce a power calibration approach with the assistance of on-chip physical thermal sensors. It is based on a new error compensation method which corrects the errors of power estimations using the feedback from physical thermal sensors. To deal with the problem of limited number of physical thermal sensors, we propose a statistical power correlation extraction method to estimate powers for places without thermal sensors. Experimental results on standard SPEC benchmarks show the new method successfully calibrates the power estimator with very low overhead introduced.
Hai Wang 0002, Sheldon X.-D. Tan, Xuexin Liu, Ashish Gupta 0007
DATE2
2012 Localized relaxation theory of circuits and its applications in electro-thermal analyses
Zuying Luo, Guoxing Zhao, Joseph A. Gordon, Sheldon X.-D. Tan
Sci. China Inf. Sci.4
2012 Fast timing analysis of clock networks considering environmental uncertainty
Hai Wang 0002, Hao Yu 0001, Sheldon X.-D. Tan
Integr.3
2012 A Fast Non-Monte-Carlo Yield Analysis and Optimization by Stochastic Orthogonal Polynomials
abstract
Performance failure has become a significant threat to the reliability and robustness of analog circuits. In this article, we first develop an efficient non-Monte-Carlo (NMC) transient mismatch analysis, where transient response is represented by stochastic orthogonal polynomial (SOP) expansion under PVT variations and probabilistic distribution of transient response is solved. We further define performance yield and derive stochastic sensitivity for yield within the framework of SOP, and finally develop a gradient-based multiobjective optimization to improve yield while satisfying other performance constraints. Extensive experiments show that compared to Monte Carlo-based yield estimation, our NMC method achieves up to 700 X speedup and maintains 98% accuracy. Furthermore, multiobjective optimization not only improves yield by up to 95.3% with performance constraints, it also provides better efficiency than other existing methods.
Fang Gong, Xuexin Liu, Hao Yu 0001, Sheldon X.-D. Tan, Junyan Ren, Lei He 0001
ACM Trans. Design Autom. Electr. Syst.4
2012 Fast Statistical Full-Chip Leakage Analysis for Nanometer VLSI Systems
abstract
In this article, we present a new full-chip statistical leakage estimation considering the spatial correlation condition (strong or weak). The new algorithm can deliver linear time, O ( N ), time complexity, where N is the number of grids on chip. The proposed algorithm adopts a set of uncorrelated virtual variables over grid cells to represent the original physical random variables and the cell size is determined by the spatial correlation length. In this way, each physical variable is always represented by virtual variables locally. We prove the number of neighbor cells for each grid cell is not related to the condition of spatial correlation (from no correlation to 100% correlated), which leads to linear time complexity in terms of number of gates. We compute the gate leakage by the orthogonal polynomials-based collocation method. The total leakage of a whole chip can be computed by simply summing up the coefficients of corresponding orthogonal polynomials in each grid cell. Furthermore, we develop a look-up table to cache statistical information for each type of gate instead of calculating leakage for every single instance of gate on a chip. As a result, a new statistical leakage characterization in Standard Cell Library (SCL) is put forward. Furthermore, an incremental analysis algorithm is proposed to update the chip-level statistical leakage information efficiently after a few changes are made. The proposed method has no restrictions on static leakage models, or types of leakage distributions. The large circuit examples in 45nm CMOS process demonstrate the proposed algorithm is 1000X faster than a recently proposed grid-based method with similar accuracy and many orders of magnitude times speedup over the Monte Carlo method. Experimental results also show the incremental analysis provides about 10X further speedup. We expect the incremental analysis could achieve more speedup over the full leakage analysis for larger problem sizes.
Ruijing Shen, Sheldon X.-D. Tan, Hai Wang 0002, Jinjun Xiong
ACM Trans. Design Autom. Electr. Syst.2
2012 Compact Modeling of Interconnect Circuits over Wide Frequency Band by Adaptive Complex-Valued Sampling Method
abstract
In this article, we propose a new model order-reduction method for compact modeling of interconnect circuits over wide frequency band using a novel complex-valued adaptive sampling and error estimation scheme. We address the outstanding error control problems in the existing sampling-based reduction framework over a frequency band. Our new method, WBMOR , explicitly and efficiently computes the exact residual errors to guide the sampling process. We show by sampling along the imaginary axis and performing a new complex-valued reduction that the reduced model will match exactly with the original model at the sample points. Additionally, we show in theory that the proposed method can achieve the error bound over a given frequency range. In practice, the new algorithm can help designers choose the best order of the reduced model for the given frequency range and error bound via the adaptive sampling scheme. In addition, WBMOR can perform wideband accurate reductions of interconnect circuits for analog and RF applications where model accuracy needs to be maintained over a wide frequency range. We compare several sampling schemes such as Monte Carlo, logarithmic, recently proposed resampling, and ARMS methods. Experimental results on a number of RLC circuits show that WBMOR is much more efficient than all the other sampling methods, including the recently proposed resampling and ARMS schemes with the same reduction orders. Compared with the traditional real-valued sampling methods, the complex-valued sampling method is more accurate for the same computational cost.
Hai Wang 0002, Sheldon X.-D. Tan, Ryan Rakib
ACM Trans. Design Autom. Electr. Syst.2
2012 General Parameterized Thermal Modeling for High-Performance Microprocessor Design
abstract
This paper proposes a new parameterized dynamic thermal modeling algorithm for emerging thermal-aware design and optimization for high-performance microprocessor design at architecture and package levels. Compared with existing behavioral thermal modeling algorithms, the proposed method can build the compact models from more general transient power and temperature waveforms used as training data. Such an approach can make the modeling process much easier and less restrictive than before and, thus, more amenable for practical measured data. The new method, called ParThermSID, consists of two steps. First, the response surface method based on second-order polynomials is applied to build the parameterized models at each time point for all of the given sampling nodes in the parameter space. Second, an improved subspace system identification method, called ThermSID, is employed to build the discrete state space models, by construction of the Hankel matrix and state space realization, for each time-varying coefficient of the polynomials generated in the first step. To overcome the overfitting problems of the subspace method, the new method employs an overfitting mitigation technique to improve model accuracy and predictive ability. Experimental results on a practical quad-core microprocessor show that the generated parameterized thermal model matches the given data very well. The compact models generated by ParThermSID also offer two orders of magnitude speedup over the commercial thermal analysis tool FloTHERM on the given example. The results also show that ThermSID is more accurate than the existing ThermPOF method.
Thom Jefferson A. Eguia, Sheldon X.-D. Tan, Ruijing Shen, Eduardo H. Pacheco, Murli Tirumala, Lingli Wang
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Decentralized and Passive Model Order Reduction of Linear Networks With Massive Ports
abstract
It is well known that model order reduction for circuits with many terminals remains a challenging problem. One reason is that existing approaches are based on a centralized framework, in which each input-output pair is implicitly assumed to be equally interacted and the matrix-valued transfer function is assumed to be fully populated. In this paper, we attempt to address this long-standing problem using a decentralized model order reduction scheme, in which a multi-input multi-output system is decoupled into a number of subsystems and each subsystem corresponds to one output and several dominant inputs. The decoupling process is based on the relative gain array, which measures the degree of interaction of each input-output pair. For each decoupled subsystem, passive reduction can be easily achieved using existing reduction techniques. The proposed method is suitable for resistance-dominant interconnects such as on-chip power grids, substrate planes where extremely compact models can be obtained. Simulation results demonstrate the advantage of the proposed method compared to the existing approaches.
Boyuan Yan, Sheldon X.-D. Tan, Lingfei Zhou, Jie Chen 0005, Ruijing Shen
IEEE Trans. Very Large Scale Integr. Syst.2
2011 A structured parallel periodic Arnoldi shooting algorithm for RF-PSS analysis based on GPU platforms
abstract
The recent multi/many-core CPUs or GPUs have provided an ideal parallel computing platform to accelerate the time-consuming analysis of radio-frequency/millimeter-wave (RF/MM) integrated circuit (IC). This paper develops a structured shooting algorithm that can fully take advantage of parallelism in periodic steady state (PSS) analysis. Utilizing periodic structure of the state matrix of RF/MM-IC simulation, a cyclic-block-structured shooting-Newton method has been parallelized and mapped onto recent GPU platforms. We first present the formulation of the parallel cyclic-block-structured shooting-Newton algorithm, called periodic Arnoldi shooting method. Then we will present its parallel implementation details on GPU. Results from several industrial examples show that the structured parallel shooting-Newton method on Tesla's GPU can lead to speedups of more than 20× compared to the state-of-the-art implicit GMRES methods under the same accuracy on the CPU.
Xuexin Liu, Hao Yu 0001, Jacob Relles, Sheldon X.-D. Tan
ASP-DAC4
2011 Performance bound analysis of analog circuits considering process variations
abstract
In this paper, we propose a new performance bound analysis of analog circuits considering process variations. We model the variations of component values as intervals measured from tested chip and manufacture processes. The new method applies a graph-based symbolic analysis and affine interval arithmetic to derive the variational transfer functions of analog circuits (linearized) with variational coefficients in forms of intervals. Then the frequency response bounds (maximum and minimum) are obtained by performing analysis of a finite number of transfer functions given by the Kharitonov's polynomial functions. We show that symbolic de-cancellation is critical for the affine interval analysis. The response bound given by the Kharitonov's functions are conservative given the correlations among coefficient intervals in transfer functions. Experimental results demonstrate the effectiveness of the proposed compared to the Monte Carlo method.
Zhigang Hao, Sheldon X.-D. Tan, Ruijing Shen, Guoyong Shi
DAC2
2011 Full-chip runtime error-tolerant thermal estimation and prediction for practical thermal management
abstract
Temperature estimation and prediction are critical for online regulation of temperature and hot spots on today's high performance processors. In this paper, we present a new method, called FRETEP, to accurately estimate and predict the full-chip temperature at runtime under more practical conditions where we have inaccurate thermal model, less accurate power estimations and limited number of on-chip physical thermal sensors. FRETEP employs a number of new techniques to address this problem. First, we propose a new thermal sensor based error compensation method to correct the errors due to the inaccuracies in thermal model and power estimations. Second, we raise a new correlation based method for error compensation estimation with limited number of thermal sensors. Third, we optimize the compact modeling technique and integrate it into the error compensation process in order to perform the thermal estimation with error compensation at runtime. Last but not least, to enable accurate temperature prediction for the emerging predictive thermal management, we design a full-chip thermal prediction framework employing time series prediction method. Experimental results show FRETEP accurately estimates and predicts the full-chip thermal behavior with very low overhead introduced and compares very favorably with the Kalman filter based approach on standard SPEC benchmarks.
Hai Wang 0002, Sheldon X.-D. Tan, Guangdeng Liao, Rafael Quintanilla, Ashish Gupta 0007
ICCAD2
2010 Efficient power grid integrity analysis using on-the-fly error check and reduction
abstract
In this paper, we present a new voltage IR drop analysis approach for large on-chip power delivery networks. The new approach is based on recently proposed sampling based reduction technique to reduce the circuit matrices before the simulation. Due to the disruptive nature of tap current waveforms in typical industry power grid networks, input current sources typically has wide frequency power spectrum. To avoid the excessively sampling, the new approach introduces an error check mechanism and on-the-fly error reduction scheme during the simulation of the reduced circuits to improve the accuracy of estimating the the large IR drops. The proposed method presents a new way to combine model order reduction and simulation to achieve the overall efficiency of simulation. The new method can also easily trade errors for speed for different applications. Experimental results show the proposed IR drop analysis method can significantly reduce the errors of the existing ETBR method at the similar computing cost, while it can have 10X and more speedup over the the commercial power grid simulator in UltraSim with about 1-2% errors on a number of real industry benchmark circuits.
Sheldon X.-D. Tan, Ning Mi, Yici Cai
ASP-DAC2
2010 Wideband reduced modeling of interconnect circuits by adaptive complex-valued sampling method
abstract
In this paper, we propose a new wideband model order reduction method for interconnect circuits by using a novel adaptive sampling and error estimation scheme. We try to address the outstanding error control problems in the existing sampling-based reduction framework. In the new method, called WBMOR, we explicitly compute the exact residual errors to guide the sampling process. We show that by sampling along the imaginary axis and performing a new complex-valued reduction, the reduced model will match exactly with the original model at the sample points. We show theoretically that the proposed method can achieve the error bound over a given frequency range. Practically the new algorithm can help designers choose the best order of the reduced model for the given frequency range and error bound via adaptive sampling scheme. As a result, it can perform wideband accurate reductions of interconnect circuits for analog and RF applications. We compare several sampling schemes such as linear, logarithmic, and recently proposed re-sampling methods. Experimental results on a number of RLC circuits show that WBMOR is much more accurate than all the other simple sampling methods and the recently proposed re-sampling scheme with the same reduction orders. Compared with the real-valued sampling methods, the complex-valued sampling method is more accurate for the same computational costs.
Hai Wang 0002, Sheldon X.-D. Tan, Gengsheng Chen
ASP-DAC2
2010 Efficient model reduction of interconnects via double gramians approximation
abstract
The gramian approximation methods have been proposed recently to overcome the high computing costs of classical balanced truncation based reduction methods. But those methods typically gain efficiency by projecting the original system only onto one dominant subspace of the approximate system gramian (for instance using only controllability gramian). This single gramian reduction method can lead to large errors as the subspaces of controllability and observability can be quite different for general interconnects with unsymmetric system matrices. In this paper, we propose a fast balanced truncation method where the system is balanced in terms of two approximate gramians as achieved in the classical balanced truncation method. The novelty of the new method is that we can keep the similar computing costs of the single gramian method. The proposed algorithm is based on a generalized SVD-based balancing scheme such that the dominant subspace of the approximate gramian product can be obtained in a very efficient way without explicitly forming the gramians. Experimental results on a number of published benchmarks show that the proposed method is much more accurate than the single gramian method with similar computing costs.
Boyuan Yan, Sheldon X.-D. Tan, Gengsheng Chen, Yici Cai
ASP-DAC2
2010 A fast analog mismatch analysis by an incremental and stochastic trajectory piecewise linear macromodel
abstract
To cope with an increasing complexity when analyzing analog mismatch in sub-90nm designs, this paper presents a fast non-Monte-Carlo method to calculate mismatch in time domain. The local random mismatch is described by a noise source with an explicit dependence on geometric parameters, and is further expanded by stochastic orthogonal polynomials (SOPs). This forms a stochastic differential-algebra-equation (SDAE). To deal with large-scale problems, the SDAE is linearized at a number of snapshots along the nominal transient trajectory, and hence is naturally embedded into a trajectory-piecewise-linear (TPWL) macromodeling. The TPWL is improved with a novel incremental aggregation of subspaces identified at those snapshots. Experiments show that the proposed method, isTPWL, is hundreds of times faster than Monte-Carlo method with a similar accuracy. In addition, our macromodel further reduces runtime by up to 25X, and is faster to build and more accurate to simulate compared to existing approaches.
Hao Yu 0001, Xuexin Liu, Hai Wang 0002, Sheldon X.-D. Tan
ASP-DAC4
2010 A robust periodic arnoldi shooting algorithm for efficient analysis of large-scale RF/MM ICs
abstract
The verification of large radio-frequency/millimeter-wave (RF/MM) integrated circuits (ICs) has regained attention for high-performance designs beyond 90nm and 60GHz. The traditional time-domain verification by standard Krylov-subspace based shooting method might not be able to deal with newly increased verification complexity. The numerical algorithms with small computational cost yet superior convergence are highly desired to extend designers' creativity to probe those extremely challenging designs of RF/MM ICs. This paper presents a new shooting algorithm for periodic RF/MM-IC systems. Utilizing a periodic structure of the state matrix, a periodic Arnoldi shooting algorithm is developed to exploit the structured Krylov-subspace. This leads to an improved efficiency and convergence. Results from several industrial examples show that the proposed periodic Arnoldi shooting method, called PAS, is 1000 times faster than the direct-LU and the explicit GMRES methods. Moreover, when compared to the existing industrial standard, a matrix-free GMRES with non-structured Krylov-subspace, the new PAS method reduces iteration number and runtime by 3 times with the same accuracy.
Xuexin Liu, Hao Yu 0001, Sheldon X.-D. Tan
DAC3
2010 A linear algorithm for full-chip statistical leakage power analysis considering weak spatial correlation
abstract
Full-chip statistical leakage power analysis typically requires quadratic time complexity in the presence of spatial correlation. When spatial correlation are strong (with large spatial correlation length), efficient linear time complexity analysis can be attained as the number of variational variables can be significantly reduced. However this is not the case for circuits where gate leakage currents are weakly correlated. In this paper, we present a linear time algorithm for statistical leakage power analysis in the presence of weak spatial correlation. The new algorithm exploits the fact that gate leakage current can be efficiently computed locally when correlation is weak. We adopt a newly proposed spatial correlation model where a new set of location-dependent uncorrelated variables are defined over virtual grids to represent the original physical random variables via fitting. To compute the leakage current of a gate on the new set of variables, the new method uses the orthogonal polynomials based collocation method, which can be applied to any gate leakage models. The total leakage currents are then computed by simply summing up the resulting orthogonal polynomials (their coefficients) on the new set of variables for all gates. Experimental results show that the proposed method is about two orders of magnitude faster than the recently proposed grid-based method [3] with similar accuracy and many orders of magnitude times over the Monte Carlo method.
Ruijing Shen, Sheldon X.-D. Tan, Jinjun Xiong
DAC2
2010 General behavioral thermal modeling and characterization for multi-core microprocessor design
abstract
This paper proposes a new architecture-level thermal modeling method to address the emerging thermal related analysis and optimization problem for high-performance multi-core microprocessor design. The new approach builds the thermal behavioral models from the measured or simulated thermal and power information at the architecture level for multi-core processors. Compared with existing behavioral thermal modeling algorithms, the proposed method can build the behavioral models from given arbitrary transient power and temperature waveforms used as the training data. Such an approach can make the modeling process much easier and less restrictive than before, and more amenable for practical measured data. The new method is based on a subspace identification method to build the thermal models, which first generates a Hankel matrix of Markov parameters, from which state matrices are obtained through minimum square optimization. To overcome the overfitting problems of the subspace method, the new method employs an overfitting mitigation technique to improve model accuracy and predictive ability. Experimental results on a real quad-core microprocessor show that ThermSID is more accurate than the existing ThermPOF method. Furthermore, the proposed overfitting mitigation technique is shown to significantly improve modeling accuracy and predictability.
Thom Jefferson A. Eguia, Sheldon X.-D. Tan, Ruijing Shen, Eduardo H. Pacheco, Murli Tirumala
DATE2
2010 General switch box modeling and optimization for FPGA routing architectures
abstract
This paper explores the FPGA routing architecture based on a new concept of “general switch box (GSB)” to improve the performance of FPGA. Compared with the existing CB/SB routing architecture and CS-box architecture, the proposed GSB architecture has much larger exploration space. Experimental results with MCNC benchmark circuits show that the performance of FPGAs with GSB is about 24.3% better than the CB/SB architecture with the same segment distribution in terms of product of channel width and delay using 0.17% less routing switches for the single wire length. For the two types of wire segments, we propose an architecture with 13.3% performance improvement at the cost of about 0.8% increase in switch number compared to the single wire length GSB architecture.
Kejie Ma, Lingli Wang, Xuegong Zhou, Sheldon X.-D. Tan, Jiarong Tong
FPT4
2010 A linear statistical analysis for full-chip leakage power with spatial correlation
abstract
In this paper, we present an approved linear-time algorithm for statistical leakage analysis in the present of any spatial correlation condition (strong or weak). The new algorithm adopts a new set of uncorrelated variables over virtual grids to represent the original physical random variables and the grid size (thus of number of new random variables) is determined by the spatial correlation length. In this way, each physical variable is always represented by virtual variables locally. We prove that the number of neighboring virtual grids for each grid is not related to condition of spatial correlation, which leads to linear time complexity in terms of number of gates. We compute the gate leakage by the orthogonal polynomials based collocation method. The total leakage of a whole chip can be computed by simply summing up the coefficients of corresponding orthogonal polynomials for each grid. Furthermore, look-up table can be created to cache statistical information for each type of gates in library instead of calculating leakage for every single gate on chip. As a result, we end up with O(N) time complexity, where N is the number of grids on chip. The proposed method has no restrictions on static leakage models, types of statistical distributions for leakage currents. Experimental results show that the proposed method is about 1000X faster than the recently proposed grid-based method [2] with similar accuracy and many orders of magnitude times over the Monte Carlo method.
Ruijing Shen, Sheldon X.-D. Tan, Jinjun Xiong
ACM Great Lakes Symposium on VLSI2
2010 Statistical analysis of large on-chip power grid networks by variational reduction scheme
Sheldon X.-D. Tan
Integr.2
2010 Statistical modeling and analysis of chip-level leakage power by spectral stochastic method
Ruijing Shen, Sheldon X.-D. Tan, Ning Mi, Yici Cai
Integr.2
2010 Parameterized architecture-level dynamic thermal models for multicore microprocessors
abstract
In this article, we propose a new architecture-level parameterized dynamic thermal behavioral modeling algorithm for emerging thermal-related design and optimization problems for high-performance multicore microprocessor design. We propose a new approach, called ParThermPOF , to build the parameterized thermal performance models from the given accurate architecture thermal and power information. The new method can include a number of variable parameters such as the locations of thermal sensors in a heat sink, different components (heat sink, heat spreader, core, cache, etc.), thermal conductivity of heat sink materials, etc. The method consists of two steps: first, a response surface method based on low-order polynomials is applied to build the parameterized models at each time point for all the given sampling nodes in the parameter space. Second, an improved Generalized Pencil-Of-Function (GPOF) method is employed to build the transfer-function-based behavioral models for each time-varying coefficient of the polynomials generated in the first step. Experimental results on a practical quad-core microprocessor show that the generated parameterized thermal model matches the given data very well. The compact models by ParThermPOF offer two order of magnitudes speedup over the commercial thermal analysis tool FloTHERM on the given examples. ParThermPOF is very suitable for design space exploration and optimization where both time and system parameters need to be considered.
Sheldon X.-D. Tan, Eduardo H. Pacheco, Murli Tirumala
ACM Trans. Design Autom. Electr. Syst.2
2010 Variational Capacitance Extraction and Modeling Based on Orthogonal Polynomial Method
abstract
In this paper, we propose a novel statistical capacitance extraction method for interconnect conductors considering process variations. The new method is called statCap, where orthogonal polynomials are used to represent the statistical processes in a deterministic way. We first show how the variational potential coefficient matrix is represented in a first-order form using Taylor expansion and orthogonal decomposition. Then, an augmented potential coefficient matrix, which consists of the coefficients of the polynomials, is derived. After this, corresponding augmented system is solved to obtain the variational capacitance values in the orthogonal polynomial form. Finally, we present a method to extend statCap to the second-order form to give more accurate results without loss of efficiency compared to the linear models. We show the derivation of the analytic second-order orthogonal polynomials for the variational capacitance integral equations. Experimental results show that statCap is two orders of magnitude faster than the recently proposed statistical capacitance extraction method based on the spectral stochastic collocation approach and many orders of magnitude faster than the Monte Carlo method for several practical conductor structures.
Ruijing Shen, Sheldon X.-D. Tan, Wenjian Yu, Yici Cai, Gengsheng Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Fast Analysis of a Large-Scale Inductive Interconnect by Block-Structure-Preserved Macromodeling
abstract
Abstract—To efficiently analyze the large-scale interconnect dominant circuits with inductive couplings (mutual inductances), this paper introduces a new state matrix, called VNA, to stamp inverse-inductance elements by replacing inductive-branch current with flux. The state matrix under VNA is diagonal-dominant, sparse, and passive. To further explore the sparsity and hierarchy at the block level, a new matrix-stretching method is introduced to reorder coupled fluxes into a decoupled state matrix with a bordered block diagonal (BBD) structure. A corresponding block-structure-preserved model-order reduction, called BVOR, is developed to preserve the sparsity and hierarchy of the BBD matrix at the block level. This enables us to efficiently build and simulate the macromodel within a SPICE-like circuit simulator. Experiments show that our method achieves up to 7 faster modeling building time, up to 33 faster simulation time, and as much as 67 smaller waveform error compared to SAPOR [a second-order reduction based on nodal analysis (NA)] and PACT (a first-order 2 2 structured reduction based on modified NA). Index Terms—Circuit simulation, high-speed interconnect model, model-order reduction. I.
Hao Yu 0001, Chunta Chu, Yiyu Shi 0001, David Smart, Lei He 0001, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.6
2009 Statistical analysis of on-chip power grid networks by variational extended truncated balanced realization method
abstract
In this paper, we present a novel statistical analysis approach for large power grid network analysis under process variations. The new algorithm is very efficient and scalable for huge networks with a large number of variational variables. This approach, called varETBR for variational extended truncated balanced realization, is based on model order reduction techniques to reduce the circuit matrices before the variational simulation. It performs the parameterized reduction on the original system using variation-bearing subspaces. varETBR calculates variational response Gramians by Monte-Carlo based numerical integration considering both system and input source variations for generating the projection subspace. varETBR is very scalable for the number of variables and is flexible for different variational distributions and ranges as demonstrated in experimental results. After the reduction, Monte-Carlo based statistical simulation is performed on the reduced system and the statistical responses of the original system are obtained thereafter. Experimental results, on a number of IBM benchmark circuits [15] up to 1.6 million nodes, show that the varETBR can be 4500X faster than the Monte-Carlo method and is much more scalable than one of the recently proposed approaches.
Sheldon X.-D. Tan, Gengsheng Chen, Xuan Zeng 0001
ASP-DAC2
2009 Statistical modeling and analysis of chip-level leakage power by spectral stochastic method
abstract
In this paper, we present a novel statistical full-chip leakage power analysis method. The new method can provide a general framework to derive the full-chip leakage current or power in a closed form in terms of the variational parameters, such as the channel length, the gate oxide thickness, etc. It can accommodate various spatial correlations. The new method employs the orthogonal polynomials to represent the variational gate leakages in a closed form first, which is generated by a fast multi-dimensional Gaussian quadrature method. The total leakage currents then are computed by simply summing up the resulting orthogonal polynomials (their coefficients). Unlike many existing approaches, no grid-based partitioning and approximation are required. Instead, the spatial correlations are naturally handled by orthogonal decompositions. The proposed method is very efficient and it becomes linear in the presence of strong spatial correlations. Experimental results show that the proposed method is about 10× faster than the recently proposed method [4] with constant better accuracy.
Ruijing Shen, Ning Mi, Sheldon X.-D. Tan, Yici Cai, Xianlong Hong
ASP-DAC3
2009 Fast analysis of nontree-clock network considering environmental uncertainty by parameterized and incremental macromodeling
abstract
It is challenging to verify clock-skew for large-scale nontree clock network with environmental uncertainties such as supply voltage fluctuation and thermal temperature gradient. This paper presents a fast clock-skew analysis via parameterized incremental truncated-balanced-realization, called piTBR method. Environmental uncertainties are parametrically and structurally added into the state equation of clock network. A compact macromodel is obtained by the subspace projection constructed from the singular value decomposition (SVD) of circuit output waveforms. To reduce the computational cost, we propose an incremental SVD method that only needs to partially update the projection matrix by analyzing the perturbed output waveform owning to environmental uncertainties. Experiments on a number of clock networks show that compared with the macromodeling by the fast TBR method, our method reduces the computational cost in the order of 100× with a similar accuracy. In addition, compared with the macromodeling by the Krylov-subspace-based method, our method reduces the waveform error by 2× with a similar runtime.
Hai Wang 0002, Hao Yu 0001, Sheldon X.-D. Tan
ASP-DAC3
2009 GPU friendly fast Poisson solver for structured power grid network analysis
abstract
In this paper, we propose a novel simulation algorithm for large scale structured power grid networks. The new method formulates the traditional linear system as a special two-dimension Poisson equation and solves it using an analytical expressions based on FFT technique. The computation complexity of the new algorithm is O(NlgN), which is much smaller than the traditional solver's complexity O(N1.5) for sparse matrices, such as the SuperLU solver and the PCG solver. Also, due to the special formulation, graphic process unit (GPU) can be explored to further speed up the algorithm. Experimental results show that the new algorithm is stable and can achieve 100X speed up on GPU over the widely used SuperLU solver with very little memory footprint.
Yici Cai, Wenting Hou, Liwei Ma, Sheldon X.-D. Tan, Pei-Hsin Ho
DAC5
2009 An efficient decoupling capacitance optimization using piecewise polynomial models
abstract
This paper proposes an efficient decoupling (decaps) capacitance optimization algorithm to reduce the voltage noise of on-chip power grid networks. The new method is based on the efficient charge formulation of the decap allocation problem. But different from the existing work [12], the new method applies the more accurate piecewise polynomial micromodels to estimate the voltage noises during the linear programming process. The resulting method overcomes the over-estimation problem, which plagues the existing method. The proposed method has the best of two worlds: it has the efficiency of the charge-based methods and the accuracy of the sensitivity-based methods. Experimental results demonstrate that the proposed method leads to the decap values similar to that of the sensitivity-based methods, which give the best reported results and are much better than the existing charge-based method, and at the same time, it enjoys the similar efficiency of the charge-based method.
Yici Cai, Sheldon X.-D. Tan, Xianlong Hong, Jacob Relles
DATE3
2009 Decoupling capacitance efficient placement for reducing transient power supply noise
abstract
Decoupling capacitance (decap) is an efficient way to reduce transient noise in on-chip power supply networks. However, excessive decap may cause more leakage power, chip resource waste, and even lead to more design iterations. In this paper, we present a novel decap-efficient placement algorithm for transient power supply noise reduction. In contrast to traditional design flow, our approach considers decap impacts at the placement stage to seek the placement minimizing decap requirements while still satisfying the traditional placement objectives. In the new method, we first devise a fast procedure to assess the decap requirement for the force-based placement framework, in which the required decap is modeled as a density function over the chip. Then, we build a corresponding supply and demand system to adjust the placement in favor of minimizing decap. Finally, we develop a decap efficient placement algorithm with a new force induced by imbalance between power supply and power demands. Experimental results show that the new combined placement and decap optimization flow could reduce the minimum decap area by 35% with a wire length increase of only 0.5% at nearly the same computational cost, which is efficient for practical problems.
Yici Cai, Qiang Zhou 0001, Sheldon X.-D. Tan, Thom Jefferson A. Eguia
ICCAD4
2009 Localized Statistical 3D Thermal Analysis Considering Electro-Thermal Coupling
abstract
In this paper, we propose a novel method for analyzing fewer hot spots in a chip. The method, called SNSOR (Single-Node Successive Over Relaxation), is based on a novel localized relaxed iterative approach to perform statistical analysis on one hot spot at a time. Based on SNSOR, we propose an approximation method, called ET-SNSOR (Electro-Thermal SNSOR), to deal with the electro-thermal coupling (ETC) effects. ET-SNSOR first uses the iterative method to update correlations from ETC, and then computes standard temperature deviations for hot spots, according to ETC and updated correlations. Experiments show that ET-SNSOR is three orders of magnitude faster than the Monte-Carlo method with small errors (less than 4.76% on maximum). It only takes an average of 0.18 second to analyze one hot spot statistically for a large test case of 1.3M nodes with ETC effects.
Zuying Luo, Jeffrey Fan, Sheldon X.-D. Tan
ISCAS3
2009 Hierarchical Krylov subspace based reduction of large interconnects
Sheldon X.-D. Tan, Lifeng Wu 0002
Integr.2
2009 Multiple block structure-preserving reduced order modeling of interconnect circuits
Ning Mi, Sheldon X.-D. Tan, Boyuan Yan
Integr.2
2009 Architecture-Level Thermal Characterization for Multicore Microprocessors
abstract
This paper investigates a new architecture-level thermal characterization problem from a behavioral modeling perspective to address the emerging thermal related analysis and optimization problems for high-performance multicore microprocessor design. We propose a new approach, calledThermPOF, to build the thermal behavioral models from the measured or simulated thermal and power information at the architecture level.ThermPOFfirst builds the behavioral thermal model using the generalized pencil-of-function (GPOF) method. Owing to the unique characteristics of transient temperature changes at the chip level, we propose two new schemes to improve the GPOF. First, we apply a logarithmic-scale sampling scheme instead of the traditional linear sampling to better capture the temperature changing behaviors. Second, we modify the extracted thermal impulse response such that the extracted poles from GPOF are guaranteed to be stable without accuracy loss. To further reduce the model size, a Krylov subspace-based reduction method is performed to reduce the order of the models in the state-space form. Experimental results on a real quad-core microprocessor show that generated thermal behavioral models match the given temperature very well.
Sheldon X.-D. Tan, Eduardo H. Pacheco, Murli Tirumala
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Hierarchical Krylov subspace reduced order modeling of large RLC circuits
abstract
In this paper, we propose a new model order reduction approach for large interconnect circuits using hierarchical decomposition and Krylov subspace projection-based model order reduction. The new approach, called hiePrimor, first partitions a large interconnect circuit into a number of smaller subcircuits and then performs the projection-based model order reduction on each of subcircuits in isolation and on the top level circuit thereafter. The new approach can exploit the parallel computing to speed up the reduction process. Theoretically we show hiePrimor can have the same accuracy as the flat reduction method given the same reduction order and it can also preserves the passivity of the reduced models as well. We also show that partitioning is important for hierarchical projection-based reduction and the minimum-span objective should be required to archive best performance for hierarchical reduction. The proposed method is suitable for reducing large global interconnects like coupled bus, transmission lines, large clock nets in the post layout stage. Experimental results demonstrate that hiePrimor can be significantly faster than flat projection method like PRIMA and be order of magnitude faster than PRIMA with parallel computing without loss of accuracy.
Sheldon X.-D. Tan
ASP-DAC2
2008 Architecture-level thermal behavioral characterization for multi-core microprocessors
abstract
In this paper, we investigate a new architecture-level thermal characterization problem from behavioral modeling perspective to address the emerging thermal related analysis and optimization problems for high-performance multi-core microprocessor design. We propose a new approach, called ThermPOF, to build the thermal behavioral models from the measured architecture thermal and power information. ThermPOF first builds the behavioral thermal model using generalized pencil-of-function (GPOF) method. And then to effectively model transient temperature changes, we proposed two new schemes to improve the GPOF. First we apply logarithmic-scale sampling instead of traditional linear sampling to better capture the temperature changing characteristics. Second, we modify the extracted thermal impulse response such that the extracted poles from GPOF are guaranteed to be stable without accuracy loss. To further reduce the model size, Krylov subspace based model order reduction is performed to reduce the order of the models in the state-space form. Experimental results on a practical quad-core microprocessor show that generated thermal behavioral models match the measured data very well.
Sheldon X.-D. Tan, Murli Tirumala
ASP-DAC2
2008 DeMOR: decentralized model order reduction of linear networks with massive ports
abstract
Model order reduction is an efficient technique to reduce the system complexity while producing a good approximation of the input-output behavior. However, the efficiency of reduction degrades as the number of ports increases, which remains a long-standing problem. The reason for the degradation is that existing approaches are based on a centralized framework, where each input-output pair is implicitly assumed to be equally interacted and the matrix-valued transfer function has to be assumed to be fully populated. In this paper, a decentralized model order reduction scheme is proposed, where a multi-input multi-output (MIMO) system is decoupled into a number of subsystems and each subsystem corresponds to one output and several dominant inputs. The decoupling process is based on the relative gain array (RGA), which measures the degree of interaction of each input-output pair. Our experimental results on a number of interconnect circuits show that most of the input-output interactions are usually insignificant, which can lead to extremely compact models even for systems with massive ports. The reduction scheme is very amenable for parallel computing as each decoupled subsystem can be reduced independently.
Boyuan Yan, Lingfei Zhou, Sheldon X.-D. Tan, Jie Chen 0005, Bruce McGaughy
DAC3
2008 ETBR: Extended Truncated Balanced Realization Method for On-Chip Power Grid Network Analysis
abstract
In this paper, we present a novel simulation approach for power grid network analysis. The new approach, called ETBR for extended truncated balanced realization, is based on model order reduction techniques to reduce the before the simulation. Different from the (improved) extended Krylov subspace methods EKS/IEKS [15, 2], ETBR performs fast truncated balanced realization on response Grammian to reduce the original system with the similar computation costs of EKS. ETBR also avoids the adverse explicit moment representation of the input signals. Instead, it uses spectrum representation of input signals by fast Fourier transformation. As a result, ETBR is more flexible for different types of input sources and can better capture the high frequency contents than EKS, and this leads to more accurate results especially for fast changing input signals. Experimental results on a number of large networks (up to one million nodes) show that, given the same order of the reduced model, ETBR is indeed more accurate than the EKS method especially for input sources rich in high-frequency components. ETBR also shows similar computation costs of EKS and less memory consumption than EKS.
Sheldon X.-D. Tan, Bruce McGaughy
DATE2
2008 Variational capacitance modeling using orthogonal polynomial method
abstract
In this paper, we propose a novel statistical capacitance extraction method for interconnects considering process variations. The new method, called statCap, is based on the spectral stochastic method where orthogonal polynomials are used to represent the statistical processes in a deterministic way. We first show how the variational potential coefficient matrix is represented in a first-order form using Taylor expansion and orthogonal decomposition. Then an augmented potential coefficient matrix, which consists of the coefficients of the polynomials, is derived. After that, corresponding augmented system is solved to obtain the variational capacitance values in the orthogonal polynomial form. Experimental results show that our method is two orders of magnitude faster than the recently proposed statistical capacitance extraction method based on the spectral stochastic collocation approach and many orders of magnitude faster than the Monte Carlo method for several practical interconnect structures.
Gengsheng Chen, Ruijing Shen, Sheldon X.-D. Tan, Wenjian Yu, Jiarong Tong
ACM Great Lakes Symposium on VLSI4
2008 FEKIS: a fast architecture-level thermal analyzer for online thermal regulation
abstract
Owning to increasing power consumption and the corresponding heat dissipated on die, efficient on-chip temperature regulation becomes imperative for today's high performance microprocessors. Temperature tracking based on the on-chip thermal sensors is not sufficient as the temperature hot spots keep changing with the load. One way to mitigate this problem is by means of software sensors, where temperature of any location is computed based on realtime power information and calibrated with the physical sensors. In this paper, we present a very efficient numerical thermal analyzer, which is suitable for fast temperature tracking and online thermal regulation. The proposed method, called FEKIS, combines two existing numerical techniques: extended Krylov subspace reduction technique to reduce the thermal circuit complexity and large-step integration method to exploits the piecewise constant power input traces, which is typical in the power traces at the architecture level. Experimental results show that FEKIS runs 10X faster than the precise time-step integration method only and $1000X$ faster than the traditional numerical integration method with high accuracy.
Pu Liu, Sheldon X.-D. Tan, Wei Wu 0024, Murli Tirumala
ACM Great Lakes Symposium on VLSI2
2008 Parameterized transient thermal behavioral modeling for chip multiprocessors
abstract
In this paper, we propose a new architecture-level parameterized transient thermal behavioral modeling algorithm for emerging thermal related design and optimization problems for high-performance chip-multiprocessor (CMP) design. We propose a new approach, called ParThermPOF, to build the parameterized thermal performance models from the given architecture thermal and power information. The new method can include a number of parameters such as the locations of thermal sensors in a heat sink, different components (heat sink, heat spread, core, cache, etc.), thermal conductivity of heat sink materials, etc. The method consists of two steps: first, response surface method based on low-order polynomials is applied to build the parameterized models at each time point for all the given sampling nodes in the parameter space. Second, an improved generalized pencil-of-function (GPOF) method is employed to build the transfer-function based behavioral models for each time-varying coefficient of the polynomials generated in the first step. Experimental results on a practical quad-core microprocessor show that the generated parameterized thermal model matchs the given data very well. ParThermPOF is very suitable for design space exploration and optimization where both time and system parameters need to be considered.
Sheldon X.-D. Tan, Eduardo H. Pacheco, Murli Tirumala
ICCAD2
2008 Modeling and simulation for on-chip power grid networks by locally dominant Krylov subspace method
abstract
Fast analysis of power grid networks has been a challenging problem for many years. The huge size renders circuit simulation inefficient and the large number of inputs further limits the application of existing Krylov-subspace macromodeling algorithms. However, strong locality has been observed that two nodes geometrically far have very small electrical impact on each other because of the exponential attenuation. However, no systematic approaches have been proposed to exploit such locality. In this paper, we propose a novel modeling and simulation scheme, which can automatically identify the dominant inputs for a given observed node in a power grid network. This enables us to build extremely compact models by projecting the system onto the locally dominant Krylov subspace corresponding to those dominant inputs only. The resulting simulation can be very fast with the compact models if we only need to view the responses of a few nodes under many different inputs. Experimental results show that the proposed method can have at least 100X speedup over SPICE-like simulations on a number of large power grid networks up to 1M nodes.
Boyuan Yan, Sheldon X.-D. Tan, Gengsheng Chen, Lifeng Wu 0002
ICCAD2
2008 Large scale P/G grid transient simulation using hierarchical relaxed approach
Yici Cai, Zhu Pan, Xianlong Hong, Sheldon X.-D. Tan
Integr.5
2008 An efficient terminal and model order reduction algorithm
Pu Liu, Sheldon X.-D. Tan, Boyuan Yan, Bruce McGaughy
Integr.2
2008 Fast Variational Analysis of On-Chip Power Grids by Stochastic Extended Krylov Subspace Method
abstract
This paper proposes a novel stochastic method for analyzing the voltage drop variations of on-chip power grid networks, considering lognormal leakage current variations. The new method, calledStoEKS, applies Hermite polynomial chaos to represent the random variables in both power grid networks and input leakage currents. However, different from the existing orthogonal polynomial-based stochastic simulation method, extended Krylov subspace (EKS) method is employed to compute variational responses from the augmented matrices consisting of the coefficients of Hermite polynomials. Our contribution lies in the acceleration of the spectral stochastic method using the EKS method to fast solve the variational circuit equations for the first time. By using the reduction technique, the new method partially mitigates increased circuit-size problem associated with the augmented matrices from the Galerkin-based spectral stochastic method. Experimental results show that the proposed method is about two-order magnitude faster than the existing Hermite PC-based simulation method and many order of magnitudes faster than Monte Carlo methods with marginal errors. StoEKS is scalable for analyzing much larger circuits than the existing Hermit PC-based methods.
Ning Mi, Sheldon X.-D. Tan, Yici Cai, Xianlong Hong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Fast Decoupling Capacitor Budgeting for Power/Ground Network Using Random Walk Approach
abstract
This paper proposes a fast and practical decoupling capacitor (decap) budgeting algorithm to optimize the power ground (P/G) network design. The new method adopts a modified random walk process to partition the circuit. Then, by utilizing the isolation property of decaps, this new method avoids solving the large nonlinear programming problem in traditional decap optimization process. Also, this method integrates leakage currents optimization algorithm using a refined leakage model. Experimental results demonstrate that our proposed method achieves approximate a 10times speed up over the heuristic method based on sensitivity and only about 6% decap area deviation from the optimal budget using the programming method.
Yici Cai, Xianlong Hong, Sheldon X.-D. Tan
ASP-DAC6
2007 Passive Interconnect Macromodeling Via Balanced Truncation of Linear Systems in Descriptor Form
abstract
In this paper, we present a novel passive model order reduction (MOR) method via projection-based truncated balanced realization method, PriTBR, for large RLC interconnect circuits. Different from existing passive truncated balanced realization (TBR) methods where numerically expensive Lur'e or algebraic Riccati (ARE's) equations are solved, the new method performs balanced truncation on linear system in descriptor form by solving generalized Lyapunov equations. Passivity preservation is achieved by congruence transformation instead of simple truncations. For the first time, passive model order reduction is achieved by combining Lyapunov equation based TBR method with congruence transformation. Compared with existing passive TBR, the new technique has the same accuracy and is numerically reliable, less expensive. In addition to passivity-preserving, it can be easily extended to preserve structure information inherent to RLC circuits, like block structure, reciprocity and sparsity. PriTBR can be applied as a second MOR stage combined with Krylov-subspace methods to generate a nearly optimal reduced model from a large scale interconnect circuit while passivity, structure, and reciprocity are preserved at the same time. Experimental results demonstrate the effectiveness of the proposed method and show PriTBR and its structure-preserving version, SP-PriTBR, are superior to existing passive TBR and Krylov-subspace based moment-matching methods.
Boyuan Yan, Sheldon X.-D. Tan, Pu Liu, Bruce McGaughy
ASP-DAC2
2007 Practical Implementation of Stochastic Parameterized Model Order Reduction via Hermite Polynomial Chaos
abstract
This paper describes the stochastic model order reduction algorithm via stochastic Hermite polynomials from the practical implementation perspective. Comparing with existing work on stochastic interconnect analysis and parameterized model order reduction, we generalized the input variation representation using polynomial chaos (PC) to allow for accurate modeling of non-Gaussian input variations. We also explore the implicit system representation using sub-matrices and improved the efficiency for solving the linear equations utilizing block matrix structure of the augmented system. Experiments show that our algorithm matches with Monte Carlo methods very well while keeping the algorithm effective. And the PC representation of non-Gaussian variables gains more accuracy than Taylor representation used in previous work (Wang et al., 2004).
Yici Cai, Qiang Zhou 0001, Xianlong Hong, Sheldon X.-D. Tan
ASP-DAC5
2007 Simultaneous Switching Noise Consideration for Power/Ground Network Optimization
abstract
With the rapid development of semiconductor technology, the working frequency of chips increases dramatically. Thus simultaneous switching noise (SSN) must be considered for robust power/ground (P/G) network design. In this paper, we mainly focus on the SSN effects for P/G network optimization. We first point out the drawbacks of the P/G optimization process without considering the SSN, by analyzing the optimized P/G grids. Then we propose a random walk based technique to consider SSN by adding decoupling capacitor (decap) prior to the nonlinear optimization process. This additional decap allocation phase constructs good current return path for the switching current caused by clock buffers and then reduces the dynamic voltage drop. Experiment results show that the proposed method achieves 2X speed up over the original approach without adding decaps in advance while the decap budget overhead is acceptable.
Yici Cai, Xianlong Hong, Sheldon X.-D. Tan
CAD/Graphics5
2007 SBPOR: Second-Order Balanced Truncation for Passive Order Reduction of RLC Circuits
abstract
RLC circuits have been shown to be better formulated as second-order systems instead of first-order systems. The corresponding model order reduction techniques for second- order systems have been developed. However, existing techniques are mainly based on moment-matching concept. While suitable for the reduction of large-scale circuits, those approaches cannot generate reduced models as compact as desired. To achieve smaller models with better error control, a novel technique, SBPOR (Second-order Balanced truncation for Passive Order Reduction), is proposed in this paper, which is the first second-order balanced truncation method proposed for passive reduction of RLC circuits. SBPOR is superior to the pioneering work in the control community because second-order systems can be balanced via congruency transformation without any accuracy loss. In addition, compared with the first-order balanced truncation approaches, SBPOR is a better choice for RLC reduction. SBPOR preserves not only passivity but also the structure information inherent to RLC circuits, which is a special need for RLC reduction. In addition, SBPOR is computationally more efficient as it only needs to solve one linear matrix equation instead of two quadratic matrix equations.
Boyuan Yan, Sheldon X.-D. Tan, Pu Liu, Bruce McGaughy
DAC2
2007 Statistical model order reduction for interconnect circuits considering spatial correlations
Jeffrey Fan, Ning Mi, Sheldon X.-D. Tan, Yici Cai, Xianlong Hong
DATE3
2007 Stochastic extended Krylov subspace method for variational analysis of on-chip power grid networks
abstract
In this paper, we propose a novel stochastic method for analyzing the voltage drop variations of on-chip power grid networks with log-normal leakage current variations. The new-method, called StoEKS, applies Hermite polynomial chaos (PC) to represent the random variables in both power grid networks and input leakage currents. But different from the existing Hermit PC based stochastic simulation method, extended Krylov subspace method (EKS) is employed to compute variational responses using the augmented matrices consisting of the coefficients of Hermite polynomials. Our contribution lies in the combination of the statistical spectrum method with the extended Krylov subspace method to fast solve the variational circuit equations for the first time. Experimental results show that the proposed method is about two-order magnitude faster than the existing Hermite PC based simulation method and more order of magnitudes faster than Monte Carlo methods with marginal errors. StoEKS also can analyze much larger circuits than the exiting Hermit PC based methods.
Ning Mi, Sheldon X.-D. Tan, Pu Liu, Yici Cai, Xianlong Hong
ICCAD2
2007 Voltage drop reduction for on-chip power delivery considering leakage current variations
abstract
In this paper, we propose a novel on-chip voltage drop reduction technique for on-chip power delivery networks of VLSI systems in the presence of variational leakage current sources. The new method inserts decoupling capacitors (decaps) into the power grid networks to reduce the voltage fluctuation. The optimization is based on sensitivity-based conjugate gradientmethod and sequence of linear programming approach. Different from existing power grid noise reduction methods, the new approach considers the impacts of inter-die and intra-die variational leakage current sources due to unavoidable process variability during the decap optimization process for the first time. Leakage currents, which although are static in nature typically, can still add to the total voltage drops and dynamic voltage reduction thus must consider the leakage-induced voltage variations. The proposed algorithm exploits the relative constant variations for different decap configurations of power grid circuits to speed up the statistical optimization process. Decaps can be inserted in such a way that the resulting circuits have much higher probability to meet the voltage drop constraints in the presence of leakage current variations. Experimental results demonstrate the effectiveness of the proposed approach and show that the new method has 100X to 1,000X of speedup over the Monte Carlo based statistical decap optimization method.
Jeffrey Fan, Ning Mi, Sheldon X.-D. Tan
ICCD3
2007 Improving the reliability of on-chip data caches under process variations
abstract
On-chip caches take a large portion of the chip area. They are much more vulnerable to parameter variation than smaller units. As leakage current becomes a significant component of the total power consumption, the leakage current variations induced thermal and reliability problem to the on-chip caches become an important design concern. This paper studies the impact of process variations, particular the leakage variations, on the temperature and reliability of on-chip caches. Our statistical simulation shows that, under process variation, 85% of the caches see shortened lifetime, with average lifetime being 81.6% of the ideal cache. At runtime, unevenly distributed dynamic power and the corresponding thermal variation would further deteriorate the situation. To mitigate this problem, we propose a dynamic cache subarray permutation scheme that can alleviate the thermal stress on a high-leakage area to improve the reliability of the caches. Experiments on 17 Spec2k benchmarks show that our scheme can extend the cache lifetime by up to 20.3%, and reduce the peak temperature by 7 degrees on average and more on data-intensive applications.
Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002, Shih-Lien Lu
ICCD2
2007 Partitioning-based decoupling capacitor budgeting via sequence of linear programming
Jeffrey Fan, Sheldon X.-D. Tan, Yici Cai, Xianlong Hong
Integr.2
2007 TermMerg: An Efficient Terminal-Reduction Method for Interconnect Circuits
abstract
In this paper, a novel method to efficiently reduce the terminal number of general linear-interconnect circuits with a large number of input or output terminals considering delay uncertainty is proposed. Our new algorithm is motivated by the fact that terminal reduction can lead to a more compact order-reduced model and the observation that very large-scale integration interconnect circuits have many similar terminals in terms of their timing and delay metrics due to their closeness in structure or due to the mathematical discretization using meshing in finite-difference or finite-element scheme during the extraction process. The new method, called TermMerg ( Proc. ICCAD, p. 821, 2005), is based on the moments of the circuits as the metrics for the timing or delay. It then employs a singular-value-decomposition (SVD) method to determine the best number of clusters based on the low-rank approximation. After this, the -means clustering algorithm is used to cluster the moments of the terminals into the different clusters. The proposed method can work with any passive-model order reduction and ensure the passive models. In contrast, we show that singular value decomposition model order reduction (SVDMOR) does not generate passive models in general. Passivity enforcement in SVDMOR will significantly hamper the terminal-reduction effectiveness. Experimental results on a number of real industry interconnect circuits demonstrate the effectiveness of the proposed method and show also that the proposed method is more accurate than SVDMOR when the used moment matrix does not give good terminal correlations.
Pu Liu, Sheldon X.-D. Tan, Bruce McGaughy, Lifeng Wu 0002, Lei He 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Pattern-Based Iterative Method for Extreme Large Power/Ground Analysis
abstract
In this paper, we present a novel pattern-based method to simulate large-scaled power/ground (P/G) grids. This method takes advantage of both traditional direct simulation methods and iterative simulation methods. The new method explores the geometry characteristics of regular P/G grids and translates topology similarity to submatrix regularity, which is called “pattern” in this paper. Such pattern structures can reduce the memory usage dramatically. Further, a new type of preconditioner is constructed to optimize the simulation process. Experimental results show that the proposed approach is about 5$\times$faster than the previous iterative methods with much lower memory, and is superior to the macro model-based hierarchical method on the tested large cases with pattern structures.
Yici Cai, Sheldon X.-D. Tan, Jeffrey Fan, Xianlong Hong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2007 Efficient power modeling and software thermal sensing for runtime temperature monitoring
abstract
The evolution of microprocessors has been hindered by increasing power consumption and heat dissipation on die. An excessive amount of heat creates reliability problems, reduces the lifetime of a processor, and elevates the cost of cooling and packaging considerably. It is therefore imperative to be able to monitor the temperature variations across the die in a timely and accurate manner. Most current techniques rely on on-chip thermal sensors to report the temperature of the processor. Unfortunately, significant variation in chip temperature both spatially and temporally exposes the limitation of the sensors. We present a compensating approach to tracking chip temperature through an OS resident software module that generates live power and thermal profiles of the processor. We developed such a software thermal sensor (STS) in a Linux system with a Pentium 4 Northwood core. We employed highly efficient numerical methods in our model to minimize the overhead of temperature calculation. We also developed an efficient algorithm for functional unit power modeling. Our power and thermal models are calibrated and validated against on-chip sensor readings, thermal images of the Northwood heat spreader, and the thermometer measurements on the package. The resulting STS offers detailed power and temperature breakdowns of each functional unit at runtime, enabling more efficient online power and thermal monitoring and management at a higher level, such as the operating system.
Wei Wu 0024, Lingling Jin, Jun Yang 0002, Pu Liu, Sheldon X.-D. Tan
ACM Trans. Design Autom. Electr. Syst.5
2007 Minimum Decoupling Capacitor Insertion in VLSI Power/Ground Supply Networks by Semidefinite and Linear Programs
abstract
Nanometer-scale VLSI design demands reliable on-chip power/ground (P/G) supply. Decoupling capacitors effectively reduce P/G supply fluctuation at the cost of leakage increase and yield loss. Existing P/G supply network decoupling capacitor insertion techniques are based on sensitivity analysis and greedy optimization. In this paper, we propose a semidefinite program and a linear program for minimum decoupling capacitor insertion in a P/G supply network, which are global optimizations with theoretically guaranteed supply voltage degradation bounds. We also propose scalability improvement schemes which enable application of the proposed semidefinite and linear programs to practical industry designs. Our experimental results on industry designs verify that the proposed semidefinite program guarantees supply voltage degradation bound for all possible supply current sources, while the proposed linear program achieves the most accurate supply voltage degradation control for a given set of supply current sources.
Bao Liu 0001, Sheldon X.-D. Tan
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Efficient early stage resonance estimation techniques for C4 package
abstract
In this paper, we study the relationship between C4 package resonance effects and logical switching timing correlations, which has not been thoroughly investigated in the past. We show that improper logic designs with some special timing correlations can lead to adverse large voltage drops, which are due to resonance effects in the widely used C4 package. We first present the numerical analysis results on industry C4 package circuits to demonstrate resonance phenomenon. Then we propose a simple algorithm to compute the worst-case logical timing correlations among cells leading to resonance. Finally, we develop an efficient technique in early logic design stage to estimate the resonance risk. Experiment results demonstrate the effectiveness of the proposed method for the accurate prediction of the resonance effect in C4 package.
Yici Cai, Sheldon X.-D. Tan, Xianlong Hong
ASP-DAC3
2006 A systematic method for functional unit power estimation in microprocessors
abstract
We present a new method for mathematically estimating the active unit power of functional units in modern microprocessors such as the Pentium 4 family. Our method leverages the phasic behavior in power consumption of programs, and captures as many power phases as possible to form a linear system of equations such that the functional unit power can be solved. Our experiment results on a real Pentium 4 processor show that power estimations attained as such agree with the measured power very well, with deviations less than 5% only.
Wei Wu 0024, Lingling Jin, Jun Yang 0002, Pu Liu, Sheldon X.-D. Tan
DAC5
2006 Statistical Analysis of Power Grid Networks Considering Lognormal Leakage Current Variations with Spatial Correlation
abstract
As the technology scales into 90 nm and below, process-induced variations become more pronounced. In this paper, we propose an efficient stochastic method for analyzing the voltage drop variations of on-chip power grid networks, considering log-normal leakage current variations with spatial correlation. The new analysis is based on the Hermite polynomial chaos (PC) representation of random processes. Different from the existing Hermite PC based method for power grid analysis, which models all the random variations as Gaussian processes without considering spatial correlation. The new method focuses on the impacts of stochastic sub-threshold leakage currents, which are modeled as log-normal distribution random variables, on the power grid voltage variations. To consider the spatial correlation, we apply orthogonal decomposition to map the correlated random variables into independent variables. Our experiment results show that the new method is more accurate than the Gaussian-only Hermite PC method using the Taylor expansion method for analyzing leakage current variations, and two orders of magnitude faster than the Monte Carlo method with small variance errors. We also show that the spatial correlation may lead to large errors if not being considered in the statistical analysis.
Ning Mi, Jeffrey Fan, Sheldon X.-D. Tan
ICCD3
2006 Efficient decoupling capacitor planning via convex programming methods
abstract
Achieving P/G supply signal integrity is crucial to success of nanometer VLSI designs. Existing P/G network optimization techniques are dominated by sensitivity based approaches. In this paper, we propose two novel convex programming based approaches for decoupling capacitor insertion in a P/G network, i.e., a semidefinite program and a linear program, which are global optimizations with theoretically guaranteed supply voltage degradation bounds. We also propose a scalability improvement scheme which enables us to apply the proposed convex programs to industry designs. We present a simple illustrative example and experimental results on an industry design, which show that the proposed semidefinite program guarantees supply voltage degradation bound for all possible supply current sources, while the proposed linear program achieves the most accurate supply voltage degradation control for a given set of supply current sources.
Andrew B. Kahng, Bao Liu 0001, Sheldon X.-D. Tan
ISPD3
2006 High accurate pattern based precondition method for extremely large power/ground grid analysis
abstract
In this paper, we propose more accurate power/ground network circuit model, which consider both via and ground bounce effects to improve the performance estimation accuracy of on-chip power distribution networks. On top of this, a new precondition iterative method, which exploits geometry characters of power/ground networks, is developed to reduce memory usage and speed up the simulation. Experimental results show that the proposed method is about 5X faster than the incomplete LU decomposition (ILU) based preconditioned conjugate gradient iterative method and about half memory usage for simulating multi-layers large scale power/ground networks.
Yici Cai, Sheldon X.-D. Tan, Xianlong Hong
ISPD3
2006 Time-domain analysis methodology for large-scale RLC circuits and its applications
Zuying Luo, Yici Cai, Sheldon X.-D. Tan, Xianlong Hong, Zhu Pan, Jingjing Fu
Sci. China Ser. F Inf. Sci.3
2006 Partitioning-Based Approach to Fast On-Chip Decoupling Capacitor Budgeting and Minimization
abstract
This paper proposes a fast decoupling capacitance (decap) allocation and budgeting algorithm for both early stage decap estimation and later stage decap minimization in today's very large scale integration physical design. The new method is based on a sensitivity-based conjugate gradient (CG) approach. But several new techniques that significantly improve the efficiency of the optimization process were adopted. First, an efficient search step scheme to replace the time-consuming line search phase in the conventional CG method for decap budget optimization was proposed. Second, instead of optimizing an entire large circuit, the circuit is partitioned into a number of smaller subcircuits and optimized separately by exploiting the locality of adding decaps. Third, the time-domain merged adjoint method was applied to compute the sensitivity information and show that the partitioning-based merged adjoint method leads to better results than the flat merged adjoint method with the improved search scheme. Experimental results show that the proposed algorithm achieves at least ten times speed-up over similar decap allocation methods reported so far with similar budget quality, and a power grid circuit with about one million nodes can be optimized using the new method in half an hour on the latest Linux workstations
Jeffrey Fan, Zhenyu Qi 0002, Sheldon X.-D. Tan, Lifeng Wu 0002, Yici Cai, Xianlong Hong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 Fast Thermal Simulation for Runtime Temperature Tracking and Management
abstract
As the power density increases exponentially, the runtime regulation of operating temperature by dynamic thermal management (DTM) becomes necessary. This paper proposes two novel approaches to the thermal analysis at the chip architecture level for efficient DTM. The first method, i.e., thermal moment matching with spectrum analysis, is based on observations that the power consumption of architecture-level modules in microprocessors running typical workloads presents a strong nature of periodicity. Such a feature can be exploited by fast spectrum analysis in the frequency domain for computing steady-state response. The second method, i.e., thermal moment matching based on piecewise constant power inputs, is based on the observation that the average power consumption of architecture-level modules in microprocessors running typical workloads determines the trend of temperature variations. As a result, using piecewise constant average power inputs can further speed up the thermal analysis. To obtain transient temperature changes due to the initial condition and constant/average power inputs, numerically stable moment matching methods with enhanced pole searching are carried out to speed up online temperature tracking with high accuracy and low overhead. The resulting thermal analysis algorithm has a linear time complexity in runtime setting when the average power inputs are applied. Experimental results show that the resulting thermal analysis algorithms lead to 10times-100times speedup over the traditional integration-based transient analysis with small accuracy loss
Pu Liu, Lingling Jin, Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2006 Wideband passive multiport model order reduction and realization of RLCM circuits
abstract
This paper presents a novel compact passive modeling technique for high-performance RF passive and interconnect circuits modeled as high-order resistor-inductor-capacitor-mutual inductance circuits. The new method is based on a recently proposed general s-domain hierarchical modeling and analysis method and vector potential equivalent circuit model for self and mutual inductances. Theoretically, this paper shows that s-domain hierarchical reduction is equivalent to implicit moment matching at around s=0 and that the existing hierarchical reduction method by one-point expansion is numerically stable for general tree-structured circuits. It is also shown that hierarchical reduction preserves the reciprocity of passive circuit matrices. Practically, a hierarchical multipoint reduction scheme to obtain accurate-order reduced admittance matrices of general passive circuits is proposed. A novel explicit waveform-matching algorithm is proposed for searching dominant poles and residues from different expansion points based on the unique hierarchical reduction framework. To enforce passivity, state-space-based optimization is applied to the model order reduced admittance matrix. Then, a general multiport network realization method to realize the passivity-enforced reduced admittance based on the relaxed one-port network synthesis technique using Foster's canonical form is proposed. The resulting modeling algorithm can generate the multiport passive SPICE-compatible model for any linear passive network with easily controlled model accuracy and complexity. Experimental results on an RF spiral inductor and a number of high-speed transmission line circuits are presented. In comparison with other approaches, the proposed reduction is as accurate as passive reduced-order interconnect macromodeling algorithm in the high-frequency domain due to the enhanced multipoint expansion, but leads to smaller realized circuit models. In addition, under the same reduction ratio, realized models by the new method have less error compared with reduced circuits by time-constant-based reduction techniques in time domain.
Zhenyu Qi 0002, Hao Yu 0001, Pu Liu, Sheldon X.-D. Tan, Lei He 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2005 Relaxed hierarchical power/ground grid analysis
abstract
This paper proposes a novel hierarchical approach to the efficient analysis of large VLSI power/ground grids. Different from the existing hierarchical approach where sub-circuit equivalent models are sparsified with computation-intensive integer programming and the resulting modeling may lead to larger errors if the top circuit matrix has large condition number, the new approach employs an iterative (relaxation) procedure to explicitly compensate the errors and avoid introducing dense matrix caused by the circuit reduction. We also propose an efficient scheme for partitioning high performance center-bumped P/G grids. Experimental results demonstrate that the new algorithm is more accurate than the existing hierarchical method while delivering more speedup over the flat simulators.
Yici Cai, Zhu Pan, Sheldon X.-D. Tan, Xianlong Hong, Wenting Hou, Lifeng Wu 0002
ASP-DAC3
2005 VLSI on-chip power/ground network optimization considering decap leakage currents
abstract
In today's power/ground(P/G) network design, on-chip decoupling capacitors(decaps) are usually made of MOS transistors with source and drain connected together. The gate leakage current becomes worse as the gate oxide layer thickness continues to shrink below 20Å. As a result, decaps will become leaky due to the gate leakage from CMOS devices. In this paper, we take a first look at the leaky decaps in P/G network optimization. We propose a leakage current model for practical decaps and also present a new two-stage leakage-current-aware approach to efficiently optimize P/G networks in a more area efficient way.
Jingjing Fu, Zuying Luo, Xianlong Hong, Yici Cai, Sheldon X.-D. Tan, Zhu Pan
ASP-DAC5
2005 Wideband modeling of RF/Analog circuits via hierarchical multi-point model order reduction
abstract
This paper proposes a novel wideband modeling technique for high-performance RF passives and linear(ized) analog circuits. The new method is based on a recently proposed sdomain hierarchical modeling and analysis method [27]. Theoretically, we show that the s-domain hierarchical reduction is equivalent to implicit moment matching around s = 0, and that the existing hierarchical reduction method by one-point expansion is numerically stable for general tree-structured circuits. Practically, we propose a hierarchical multi-point reduction scheme for high-fidelity, wideband modeling of general passive or active linear circuits. A novel explicit waveform matching algorithm is proposed for searching the dominant poles and residues from different expansion points based on the unique hierarchical reduction framework. Experimental results with large analog circuits, on-chip spiral inductors are presented to validate the proposed method.
Zhenyu Qi 0002, Sheldon X.-D. Tan, Hao Yu 0001, Lei He 0001
ASP-DAC2
2005 A wideband hierarchical circuit reduction for massively coupled interconnects
abstract
We develop a realizable circuit reduction to generate the interconnect macro-model for parasitic estimation in wideband applications. The inductance is represented by VPEC (vector potential equivalent circuit) model, which not only enables the passive sparsification but also gives correct low-frequency response, whereas the recent circuit reduction intrinsically has inaccurate value and low-frequency response due to nodal-susceptance formulation. Applying hierarchical circuit-reduction enhanced by multi-point expansions, we can obtain an accurate high-order impedance function to capture the high-frequency response. The impedance function is further enforced passivity by convex programming, and realized by a Foster's synthesis. Experiments show that our method is as accurate as PRIMA in high frequency range, but leads to a realized circuit model with up to 10X times less complexity and up to 8X smaller simulation time. In addition, under the same reduction ratio, its error margin is less than that for the time-constant based reduction in both time-domain and frequency-domain simulations.
Hao Yu 0001, Lei He 0001, Zhenyu Qi 0002, Sheldon X.-D. Tan
ASP-DAC4
2005 Analysis of buffered hybrid structured clock networks
abstract
This paper presents a novel approach for fast transient analysis of buffered hybrid structured clock networks. The new method applies structure reduction and relaxed hierarchical analysis methods to reduce the circuit complexity and speedup the simulation. A simple controlled sources model is used for modeling clock buffers to deal with nonlinearity in the buffered clock trees. Our experiment results show that the proposed algorithm is about two orders of magnitude faster than HSPICE without loss on accuracy and stability. The relatively errors on delay times are within a few percent of the exact ones.
Qiang Zhou 0001, Yici Cai, Xianlong Hong, Sheldon X.-D. Tan
ASP-DAC5
2005 Partitioning-based approach to fast on-chip decap budgeting and minimization
abstract
This paper proposes a fast decoupling capacitance (decap) allocation and budgeting algorithm for both early stage decap estimation and later stage decap minimization in today's VLSI physical design. The new method is based on a sensitivity-based conjugate gradient (CG) approach. But it adopts several new techniques, which significantly improve the efficiency of the optimization process. First, the new approach applies the time-domain merged adjoint network method for fast sensitivity calculation. Second, an efficient search step scheme is proposed to replace the timeconsuming line search phase in conventional conjugate gradient method for decap budget optimization. Third, instead of optimizing an entire large circuit, we partition the circuit into a number of smaller sub-circuits and optimize them separately by exploiting the locality of adding decaps. Experimental results show that the proposed algorithm achieves at least 10X speed-up over the fastest decap allocation method reported so far with similar or even better budget quality and a power grid circuit with about one million nodes can be optimized using the new method in half an hour on the latest Linux workstations.
Zhenyu Qi 0002, Sheldon X.-D. Tan, Lifeng Wu 0002, Yici Cai, Xianlong Hong
DAC3
2005 A Study of the Scalability of On-Chip Routing for Just-in-Time FPGA Compilation
abstract
Just-in-time (JIT) compilation has been used in many applications to enable standard software binaries to execute on different underlying processor architectures. We previously introduced the concept of a standard hardware binary, using a just-in-time compiler to compile the hardware binary to a field-programmable gate array (FPGA). Our JIT compiler includes lean versions of technology mapping, placement, and routing algorithms, of which routing is the most computationally and memory expensive step. As FPGAs continue to increase in size, a JIT FPGA compiler must be capable of efficiently mapping increasingly larger hardware circuits. In this paper, we analyze the scalability of our lean on-chip router, the Riverside on-chip router (ROCR), for routing increasingly large hardware circuits. We demonstrate that ROCR scales well in terms of execution time, memory usage and circuit quality, and we compare the scalability of ROCR to the well known versatile place and route (VPR) timing-driven routing algorithm, comparing to both their standard routing algorithm and their fast routing algorithm. Our results show that on average ROCR executes 3 times faster using 18 times less memory than VPR. ROCR requires only 1% more routing resources, while creating a critical path 30% longer VPR's standard timing-driven router. Furthermore, for the largest hardware circuit, ROCR executes 3 times faster using 14 times less memory, and results in a critical path 2.6% shorter than VPR's fast timing-driven router.
Roman L. Lysecky, Frank Vahid, Sheldon X.-D. Tan
FCCM3
2005 Fast thermal simulation for architecture level dynamic thermal management
abstract
As power density increases exponentially, runtime regulation of operating temperature by dynamic thermal managements becomes necessary. This paper proposes a novel approach to the thermal analysis at chip architecture level for efficient dynamic thermal management. Our new approach is based on the observation that the power consumption of architecture level modules in microprocessors running typical workloads presents strong nature of periodicity. Such a feature can be exploited by fast spectrum analysis in frequency domain for computing steady state response. To obtain the transient temperature changes due to initial condition and constant power inputs, numerically stable moment matching approach is carried out. The total transient responses is the addition of the two simulation results. The resulting fast thermal analysis algorithm leads to at least 10/spl times/-100/spl times/ speedup over traditional integration-based transient analysis with small accuracy loss.
Pu Liu, Zhenyu Qi 0002, Lingling Jin, Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002
ICCAD6
2005 An efficient method for terminal reduction of interconnect circuits considering delay variations
abstract
This paper proposes a novel method to efficiently reduce the terminal number of general linear interconnect circuits with a large number of input and/or output terminals considering delay variations. Our new algorithm is motivated by the fact that VLSI interconnect circuits have many similar terminals in terms of their timing and delay metrics due to their closeness in structure or due to mathematic approximation using meshing in finite difference or finite element scheme during the extraction process. By allowing some delay tolerance or variations, we can reduce many similar terminals and keep a small number of representative terminals. After terminal reduction, traditional model order reduction methods can achieve more compact models and improve simulation efficiency. The new method, TermMerg, is based on the moments of the circuits as the metrics for the timing or delay. It then employs singular value decomposition (SVD) method to determine the optimum number of clusters based on the low-rank approximation. After this, the K-means clustering algorithm is used to cluster the moments of the terminals into different clusters. Experimental results on a number of real industry interconnect circuits demonstrate the effectiveness of the proposed method.
Pu Liu, Sheldon X.-D. Tan, Zhenyu Qi 0002, Bruce McGaughy, Lei He 0001
ICCAD2
2005 Efficient Thermal Simulation for Run-Time Temperature Tracking and Management
abstract
As power density increases exponentially, run-time regulation of operating temperature by dynamic thermal management becomes imperative. This paper proposes a novel approach to real-time thermal estimation at chip level for efficient dynamic thermal management in lieu of the thermal sensors, which are erroneous and having longer delays. Our new approach is based on the observation that the average power consumption of architecture level modules in microprocessors running typical workloads determines the trend of temperature variations. Such a feature can be exploited by applying fast moment matching technique in frequency domain. To obtain the transient temperature changes due to initial condition and constant power input pattern, numerically stable moment matching approach is carried out to speed up on-line temperature tracking with high accuracy and low overhead. The resulting fast thermal analysis algorithm has linear time complexity in run-time setting and leads to about two orders of magnitude speed-up over traditional integration-based transient analysis. The average maximum error under running typical benchmarks is only about 0.37/spl deg/C as compared to other well-accepted simulation tools.
Pu Liu, Zhenyu Qi 0002, Lingling Jin, Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002
ICCD6
2005 A general hierarchical circuit modeling and simulation algorithm
abstract
This paper proposes a new hierarchical circuit modeling and simulation technique in s-domain for linear analog and interconnect circuits. The new method is based on a graph-based symbolic hierarchical circuit decomposition scheme. It can derive the exact or approximate admittances in the reduced circuit matrix and compute the circuit characteristics in rational function forms for very large linear analog and interconnect circuits. We show that the exact symbolic expressions of a circuit can be obtained by finding the cancellation-free expressions from the same circuit with hierarchical definitions. Some theoretical results are characterized for the presence and conditions of cancellations in the symbolic expressions from the subcircuit reduction. A novel decancellation strategy based on a graph-based hierarchical decomposition process is proposed and canceling terms are removed both symbolically and numerically to obtain the order-reduced circuit models. The proposed method can be used for modeling and simulation of any passive or active linear circuit, which makes our method very attractive for modeling both analog circuits and resistance-capacitance-inductance interconnect circuits in both frequency and time domain. An example RC circuit is illustrated and experimental results with some large analog and interconnects circuits are presented to validate the proposed method. Our experimental results also show that subcircuit (multiple-node) reduction scheme in general is better than one-node reduction methods such as Y-/spl Delta/ transformation in terms of CPU time and memory usage.
Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2005 Hierarchical approach to exact symbolic analysis of large analog circuits
abstract
This paper proposes a novel approach to the exact symbolic analysis of very large analog circuits. The new method is based on determinant decision diagrams (DDDs) representing symbolic product terms. But instead of constructing DDD graphs directly from a flat circuit matrix, the new method constructs DDD graphs in a hierarchical way based on hierarchically defined circuit structures. The resulting algorithm can analyze much larger analog circuits exactly than before. The authors show that exact symbolic expressions of a circuit are cancellation-free expressions when the circuit is analyzed hierarchically. With this, the authors propose a novel symbolic decancellation process, which essentially leads to the hierarchical DDD graph constructions. The new algorithm partially avoids the exponential DDD construction time by employing more efficient DDD graph operations during the hierarchical construction. The experimental results show that very large analog circuits, which cannot be analyzed exactly before like /spl mu/A725 and other unstructured circuits up to 100 nodes, can be analyzed by the new approach for the first time. The new approach significantly improves the exact symbolic capacity and promises huge potentials for the applications of exact symbolic analysis.
Sheldon X.-D. Tan, Weikun Guo, Zhenyu Qi 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2004 A fast decoupling capacitor budgeting algorithm for robust on-chip power delivery
Jingjing Fu, Zuying Luo, Xianlong Hong, Yici Cai, Sheldon X.-D. Tan, Zhu Pan
ASP-DAC5
2004 Dynamic FPGA routing for just-in-time FPGA compilation
abstract
Just-in-time (JIT) compilation has previously been used in many applications to enable standard software binaries to execute on different underlying processor architectures. However, embedded systems increasingly incorporate Field Programmable Gate Arrays (FPGAs), for which the concept of a standard hardware binary did not previously exist, requiring designers to implement a hardware circuit for a single specific FPGA. We introduce the concept of a standard hardware binary, using a just-in-time compiler to compile the hardware binary to an FPGA. A JIT compiler for FPGAs requires the development of lean versions of technology mapping, placement, and routing algorithms, of which routing is the most computationally and memory expensive step. We present the Riverside On-Chip Router (ROCR) designed to efficiently route a hardware circuit for a simple configurable logic fabric that we have developed. Through experiments with MCNC benchmark hardware circuits, we show that ROCR works well for JIT FPGA compilation, producing good hardware circuits using an order of magnitude less memory resources and execution time compared with the well known Versatile Place and Route (VPR) tool suite. ROCR produces good hardware circuits using 13X less memory and executing 10X faster than VPR's fastest routing algorithm. Furthermore, our results show ROCR requires only 10% additional routing resources, and results in circuit speeds only 32% slower than VPR's timing-driven router, and speeds that are actually 10% faster than VPR's routability-driven router.
Roman L. Lysecky, Frank Vahid, Sheldon X.-D. Tan
DAC3
2004 Hierarchical approach to exact symbolic analysis of large analog circuits
abstract
This paper provides a novel approach to exact symbolic analysis of very large analog circuits. The new method is based on determinant decision diagrams (DDDs) to represent symbolic product terms. But instead of constructing DDD graphs directly from a flat circuit matrix, the new method constructs DDD graphs in a hierarchical way based on hierarchically defined circuit structures. The resulting algorithm can analyze much larger analog circuits exactly than before. Theoretically, we show that exact symbolic expressions of a circuit are cancellation-free expressions when the circuit is analyzed hierarchically. Practically we propose a novel hierarchical DDD graph construction algorithm. Our experimental results show that very large analog circuits, which can't be analyzed exactly before like μA725 and other unstructured circuits up to 100 nodes, can be analyzed by the new approach for the first time. The new approach significantly improves the exact symbolic capacity and promises huge potentials for the new applications of symbolic analysis in analog circuit design automation.
Sheldon X.-D. Tan, Weikun Guo, Zhenyu Qi 0002
DAC1
2004 Hierarchical Modeling and Simulation of Large Analog Circuits
abstract
This paper proposes a new hierarchical circuit modeling and simulation technique in s-domain for linear analog circuits. The new algorithm can perform circuit complexity reduction by deriving the exact or approximate admittances in rational form in the reduced circuit matrix and deriving the circuit characteristics for very large linear analog and interconnect circuits. We characterize some theoretical results regarding the conditions on the generations of canceling terms during the general hierarchical circuit analysis and propose an explicit de-cancellation scheme to remove canceling terms based on a new hierarchical symbolic analysis framework. The resulting algorithm can be used for modeling and simulation of linear analog and interconnect circuits in both frequency and time domain.
Sheldon X.-D. Tan, Zhenyu Qi 0002
DATE1
2004 A Fast Delay Analysis Algorithm for The Hybrid Structured Clock Network
abstract
This paper presents a novel approach to reducing the complexity of the transient linear circuit analysis for a hybrid structured clock network. Topology reduction is first used to reduce the complexity of the circuits and a preconditioned Krylov-subspace iterative method is then used to perform the nodal analysis on the reduced circuits. By proper choice of the simulation time step based on Elmore delay model, the delay of the clock signal between the clock source and the sink node and the skews between the sink nodes can be obtained efficiently and accurately. Our experimental results show that the proposed algorithm is two orders of magnitude faster than HSPICE without loss of accuracy and stability and the maximum error is within 0.4% of the exact delay time.
Yici Cai, Qiang Zhou 0001, Xianlong Hong, Sheldon X.-D. Tan
ICCD5
2004 Efficient approximation of symbolic expressions for analog behavioral modeling and analysis
abstract
Efficient algorithms are presented to generate approximate expressions for transfer functions and characteristics of large linear-analog circuits. The algorithms are based on a compact determinant decision diagram (DDD) representation of exact transfer functions and characteristics. Several theoretical properties of DDDs are characterized, and three algorithms, namely, based on dynamic programming, based on consecutive k-shortest path (SP), and based on incremental k-SP, are presented in this paper. We show theoretically that all three algorithms have time complexity linearly proportional to |DDD|, the number of vertices of a DDD, and that the incremental k-SP-based algorithm is fastest and the most flexible one. Experimental results confirm that the proposed algorithms are the most efficient ones reported so far, and are capable of generating thousands of dominant terms for typical analog blocks in CPU seconds on a modern computer workstation.
Sheldon X.-D. Tan, Chuanjin Richard Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Efficient DDD-based term generation algorithm for analog circuit behavioral modeling
abstract
An efficient approach to generating symbolic product terms for behavioral modeling of large linear analog circuits is presented. The approach is based on a compact determinant decision diagram (DDD) representation of transfer functions and characteristics of analog circuits. The new algorithm is based on the concept that a dominant term in a DDD graph can be found by searching the shortest path in the graph. But instead of traversing a whole DDD graph each time, we show that a shortest path can be found by just updating a small number of the newly added vertices after the first shortest path is found. Experimental results indicate that the new symbolic term generation algorithm outperforms both pure shortest path based algorithm and dynamic programming based algorithm, which is the fastest symbolic term generation algorithm published so far.
Sheldon X.-D. Tan, Chuanjin Richard Shi
ASP-DAC1
2003 A General S-Domain Hierarchical Network Reduction Algorithm
Sheldon X.-D. Tan
ICCAD1
2003 Balanced multi-level multi-way partitioning of analog integrated circuits for hierarchical symbolic analysis
Sheldon X.-D. Tan, Chuanjin Richard Shi
Integr.1
2003 Efficient very large scale integration power/ground network sizing based on equivalent circuit modeling
abstract
We present an efficient method of minimizing the area of power/ground (P/G) networks in integrated circuit layouts subject to reliability constraints. Instead of directly sizing the original P/G network extracted from a circuit layout, as done previously, the new method first constructs a reduced but electrically equivalent P/G network. Then the sequence of linear programming method is applied to optimize the reduced network. The solution of the original network is then backsolved from the optimized reduced network. The new method exploits the regularities in the P/G networks to reduce the complexities of P/G networks. Experimental results show that the sizes of reduced networks are typically significantly smaller than that of the original networks. The resulting algorithm is fast enough that P/G networks with more than one million branches can be sized in a few minutes on modern Sun workstations.
Sheldon X.-D. Tan, Chuanjin Richard Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Reliability-constrained area optimization of VLSI power/ground networks via sequence of linear programmings
abstract
This paper presents a new method of sizing the widths of the power and ground routes in integrated circuits so that the chip area required by the routes is minimized subject to electromigration and IR voltage drop constraints. The basic idea is to transform the underlying constrained nonlinear programming problem into a sequence of linear programs. Theoretically, we show (that the sequence of linear programs always converges to the optimum solution of the relaxed convex optimization problem. Experimental results demonstrate that the proposed sequence-of-linear-program method Is orders of magnitude faster than the best-known method based on conjugate gradients with constantly better solution qualities.
Sheldon X.-D. Tan, Chuanjin Richard Shi, Jyh-Chwen Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2001 Fast Power/Ground Network Optimization Based on Equivalent Circuit Modeling
abstract
This paper presents an efficient algorithm for optimizing the area of power or ground networks in integrated circuits subject to the reliability constraints. Instead of solving the original power/ground networks extracted from circuit layouts as previous methods did, the new method first builds the equivalent models for many series resistors in the original networks, then the sequence of linear programming method [9] is used to solve the simplified networks. The solutions of the original networks then are back solved from the optimized, simplified networks. The new algorithm simply exploits the regularities in the power/ground networks. Experimental results show that the complexities of simplified networks are typically significantly smaller than that of the original circuits, which renders the new algorithm extremely fast. For instance, power/ground networks with more than one million branches can be sized in a few minutes on modern SUN workstations.
Sheldon X.-D. Tan, Chuanjin Richard Shi
DAC1
2001 Compact representation and efficient generation of s-expandedsymbolic network functions for computer-aided analog circuit design
abstract
A graph-based approach is presented for the generation of exact symbolic network functions in the form of rational polynomials of the complex frequency variable s for analog integrated circuits. The approach employs determinant decision diagrams (DDDs) to represent the determinant of a circuit matrix and its cofactors. A notion of multiroot DDDs is introduced, where each root represents a symbolic expression for an individual coefficient of the powers of s in the numerator and denominator of a network function, and multiple roots share their common subgraphs. A DDD-based algorithm is presented for generating s-expanded network functions. We prove theoretically and validate experimentally that the algorithm constructs in O(kl|DDD|) time an s-expanded DDD with no more than kl|DDD| vertices, where k is the degree of the denominator s polynomial, l is the maximum number of devices that connect to a circuit node, and |DDD| is the number of DDD vertices representing the circuit-matrix determinant. For a practical circuit, |DDD| is often many orders-of-magnitude less than the number of product terms. In contrast, previous approaches require the time and space complexities proportional to the number of product terms, which grows exponentially with the size of a circuit. Experimental results have demonstrated that the new approach can produce exact s-expanded-symbolic network functions for /spl mu/A741 operational amplifiers in several CPU seconds on an UltraSparc-I workstation. The expressive power of multiroot s-expanded DDDs is so remarkable that in one instance, over 10/sup 35/ symbolic product terms have been represented by a multiroot DDD with less than 17 K vertices. The compactness of DDDs is further demonstrated in the context of symbolic noise evaluation, where potentially many transfer functions, each being used for a noise source in the circuit, can be represented by a single DDD with the size comparable to that for a few transfer functions. This provides a powerful tool for solving many symbolic analysis problems such as deriving interpretable symbolic expressions, dominant pole/zero estimation, and analog testability analysis. We have also demonstrated that repetitive numerical evaluation with the derived s-expanded symbolic expressions for frequency-domain simulation and small-signal noise analysis can be much faster than SPICE-like simulators and the resulting expressions for a circuit block can be used as behavioral models for high-level simulation.
Chuanjin Richard Shi, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 Symbolic circuit-noise analysis and modeling with determinant decision diagrams
abstract
Abstract | In this paper, a new symbolic noise analysis and modeling technique is presented. The new method exploits the sharing of symbolic expressions in the noise models by using a recently introduced graph, called determinant decision diagrams (DDDs), for symbolic determinant representations. With efcient DDD-based graph manipulations, we are able to generate the exact noise models for analog blocks. Symbolic noise analysis and modeling on real analog circuit examples are presented and compared with SPICE noise simulation. 1.
Sheldon X.-D. Tan, Chuanjin Richard Shi
ASP-DAC1
2000 Canonical symbolic analysis of large analog circuits withdeterminant decision diagrams
abstract
Symbolic analysis has many applications in the design of analog circuits. Existing approaches rely on two forms of symbolic-expression representation: expanded sum-of-product form and arbitrarily nested form. Expanded form suffers the problem that the number of product terms grows exponentially with the size of a circuit. Nested form is neither canonical nor amenable to symbolic manipulation. In this paper, we present a new approach to exact and canonical symbolic analysis by exploiting the sparsity and sharing of product terms. It consists of representing the symbolic determinant of a circuit matrix by a graph-called a determinant decision diagram (DDD)-and performing symbolic analysis by graph manipulations. We show that DDD construction, as well as many symbolic analysis algorithms, takes time almost linear in the number of DDD vertices. We describe an efficient DDD-vertex ordering heuristic and prove that it is optimum for ladder-structured circuits. For practical analog circuits, the numbers of DDD vertices are several orders of magnitude less than the numbers of product terms. The algorithms have been implemented and compared respectively to symbolic analyzers ISAAC and Maple-V in generating the expanded sum-of-product expressions, and SCAPP in generating the nested sequences of expressions.
Chuanjin Richard Shi, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 Hierarchical symbolic analysis of analog integrated circuits viadeterminant decision diagrams
abstract
A new method is proposed for hierarchical symbolic analysis of large analog integrated circuits. It consists of performing symbolic suppression of each subcircuit to its terminals in terms of subcircuit matrix determinants and cofactors, and applying Cramer's rule to symbolically solve the set of equations at the top level of the circuit hierarchy. An annotated, directed, and acyclic graph, called determinant decision diagram (DDD), is used to represent symbolic determinants of subcircuit matrices and cofactors used in subcircuit suppression, as well as symbolic determinants of the top-level circuit matrix and cofactors required in applying Cramer's rule. DDD enables us to systematically exploit the inherent sparsity of circuit matrices and the sharing of symbolic expressions. It is capable of representing a huge number of symbolic product terms in a canonical and highly compact manner. The proposed method is illustrated using a Cauer parameter low-pass filter. It has been implemented in a symbolic analyzer and compared to best-known hierarchical symbolic analyzer SCAPP and numerical simulator SPICE. Experimental results on several analog circuits including the /spl mu/A741 operational amplifier - a circuit with less structural regularities - are described.
Sheldon X.-D. Tan, Chuanjin Richard Shi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1999 Balanced Multi-Level Multi-Way Partitioning of Large Analog Circuits for Hierarchical Symbolic Analysis
abstract
Symbolic analysis of analog circuits is important in analog design automation. However, it is limited to the analysis of small analog circuits where exact symbolic expressions are required. In this paper, we present an efficient algorithm for partitioning large general analog circuits into smaller subcircuits so that symbolic analysis can be performed hierarchically. Experimental results have demonstrated that our method outperforms the best partitioning-based symbolic analyzer SCAPP.
Sheldon X.-D. Tan, Chuanjin Richard Shi
ASP-DAC1
1999 Reliability-Constrained Area Optimization of VLSI Power/Ground Networks via Sequence of Linear Programmings
abstract
This paper presents a new method for determining the widths of the power and ground routes in integrated circuits so that the area required by the routes is minimized subject to the reliability constraints. The basic idea is to transform the resulting constrained nonlinear programming problem into a sequence of linear programs. Theoretically, we show that the sequence of linear programs always converges to the optimum solution of the relaxed convex problem. Experimental results demonstrate that the sequence-of-linear-programming method is orders of magnitude faster than the best-known method based on conjugate gradients, with constantly better optimization solutions.
Sheldon X.-D. Tan, Chuanjin Richard Shi, Dragos Lungeanu, Jyh-Chwen Lee, Li-Pen Yuan
DAC1
1999 Interpretable Symbolic Small-Signal Characterization of Large Analog Circuits using Determinant Decision Diagrams
abstract
A new approach is proposed to generate interpretable symbolic expressions of small-signal characteristics for large analog circuits. The approach is based on a complete, exact, yet compact representation of symbolic expressions via determinant decision diagrams (DDDs). We show that two key tasks of generating interpretable symbolic expressions-term de-cancellation and term simplification-can be performed in linear time in terms of the number of DDD vertices. With the number of DDD vertices many-orders-of-magnitude less than the number of product terms, the proposed approach has been shown to be much more efficient than other start-of-the-art approaches.
Sheldon X.-D. Tan, Chuanjin Richard Shi
DATE1
1997 Symbolic analysis of large analog circuits with determinant decision diagrams
abstract
Symbolic analog-circuit analysis has many applications, and is especially useful for analog synthesis and testability analysis. We present a new approach to exact and canonical symbolic analysis by exploiting the sparsity and sharing of product terms. It consists of representing the symbolic determinant of a circuit matrix by a graph-called determinant decision diagram (DDD)-and performing symbolic analysis by graph manipulations. We showed that DDD construction and DDD-based symbolic analysis can be performed in time complexity proportional to the number of DDD vertices. We described a vertex ordering heuristic, and showed that the number of DDD vertices can be quite small-usually orders-of-magnitude less than the number of product terms. The algorithm has been implemented. An order-of-magnitude improvement in both CPU time and memory usage over existing symbolic analyzers ISAAC and Maple-V has been observed for large analog circuits.
Chuanjin Richard Shi, Sheldon X.-D. Tan
ICCAD2