Farid N. Najm

dblp:n/FaridNNajm · DBLP profile ↗
← Back
132ranked-venue papers
19as first author
3since 2021 · last 2025
0000-0001-5393-7794ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 132 · 19 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11Software engineering, systems software and programming languages · 3
YearPublicationVenuePosition
2025 Novel Partitioning-Based Approach for Electromigration Assessment With Neural Networks
abstract
Due to continuing technology scaling, electromigration (EM) remains a prominent reliability concern in integrated circuit design. Traditional empirical methods often result in over-design in very large scale integration (VLSI) due to model inaccuracy. Recently, researchers have focused on analyzing EM susceptibility by tracking hydrostatic stress evolution in metal lines, governed by computationally expensive partial differential equations (PDEs). In this paper, we propose a partitioning-based approach using neural networks to efficiently forecast the stress evolution along interconnect trees during the void nucleation and growth phases. This approach begins by decomposing the interconnect tree into subcomponents, providing computationally efficient analytical solutions for predicting stress evolution within each subtree. Subsequently, we employ a lightweight neural network to reassemble these components with their corresponding solutions to the original structure, ensuring accurate stress prediction. This divide-and-conquer strategy can accommodate various tree structures, with offshoots at arbitrary junctions, and holds substantial promise for using NN-based methods to solve mesh-free stress evolution on much larger interconnect trees than previously possible, with reduced computational overhead and heightened accuracy. The proposed approach eliminates the need for time discretization and grid meshing typically required in numerical methods. Numerical results confirm its advantages in accuracy and computational efficiency.
Tianshu Hou, Farid N. Najm, Ngai Wong 0001, Haibao Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Electromigration Assessment in Power Grids with Account of Redundancy and Non-Uniform Temperature Distribution
abstract
A recently proposed methodology for electromigration (EM) assessment in on-chip power/ground grid of integrated circuits has been validated by means of measurements, performed on dedicated test grids. IR drop degradation in the grid is used for defining the EM failure criteria. Physics-based models are involved for simulation of EM-induced stress evolution in interconnect structures, void formation and evolution, resistance increase of the voided segments, and consequent re-distribution of electric current in the redundant grid paths. A grid-like test structure, fabricated with a 65 nm technology and consisting of two metal layers, allowed to calibrate the voiding models by tracking voltage evolution in all grid nodes in experiment and in simulation. Good fit of the measured and simulated time-to-failure (TTF) probability distribution was obtained in both cases of uniform and non-uniform temperature distribution across the grid. The second test grid was fabricated with a 28 nm technology, consisted of 4 metal layers, and contained power and ground nets connected to "quasi-cells" with poly-resistors, which were specially designed for operating at elevated temperatures ~350°C. The existing current distributions resulted in different behavior of EM-induced failures in these nets: a gradual voltage evolution in power net, and sharp changes in ground net were observed in experiment, and successfully reproduced in simulations.
Armen Kteyan, Valeriy Sukharev, Alexander Volkov, Jun-Ho Choy, Farid N. Najm, Yong Hyeon Yi, Chris H. Kim, Stéphane Moreau
ISPD5
2022 Experimental Validation of a Novel Methodology for Electromigration Assessment in On-Chip Power Grids
abstract
A recently proposed theoretical methodology for the assessment of the electromigration (EM) induced IR-drop degradation in on-chip power/ground grids has been validated by means of measurements performed on real silicon. A voltage tapping technique was employed for the direct measurement of voltage variations at 162 nodes of the power net, stressed with 10 mA constant source current at an elevated temperature of 350 °C. A voltage drop between cathode and anode pads exceeding a specified threshold was considered as a failure. Times-to-failure (TTF) was measured on 19 packaged test grids and used for computing the mean TTF (MTTF). The EM-induced voltage degradation in this grid was also analyzed with an assessment methodology based on a simulation of stress evolution everywhere in the grid, resulting in a voiding in some of grid branches and corresponding resistance increase. A set of voiding compact models for different grid segments was developed and used in the simulations. The stochastic nature of the EM phenomenon was captured by introducing random distributions of atomic diffusivities and critical stresses across the grid and iterating them with Monte Carlo loops. A good fit between the measured voltage evolution kinetics at different grid nodes and that predicted by simulation, and the good agreement between measured and simulated failure distributions can be considered as the ever first experimental validation of this EM assessment methodology for on-chip power/ground (p/g) grids.
Valeriy Sukharev, Armen Kteyan, Farid N. Najm, Yong Hyeon Yi, Chris H. Kim, Jun-Ho Choy, Sofya Torosyan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Electromigration Checking Using a Stochastic Effective Current Model
abstract
Electromigration (EM) degradation evolves slowly towards failure, over a period of years. This is why EM checking methods use effective current models to represent the underlying circuit workload, which are typically constant (DC) currents over time. However, ignoring all input current variations around the mean can be risky, because low-frequency input variations can have a significant impact on EM, resulting in shorter than expected lifetimes. With the use of dark silicon and multimodal chip operation, such low-frequency changes in workload are becoming increasingly common in modern designs. Ignoring these variations can lead to false positives and must be avoided. We tackle this by developing a stochastic effective current model for the input current waveforms that is easy for users to specify and which allows stochastic estimation of the impact of input variability on the lifetime. User-provided guidance on the expected durations of various modes of operation is used to provide input current variances, which are then propagated to provide variances around the stress waveforms in the metal network, which gives a more realistic estimate of the EM lifetime. Variance propagation can be expensive for large systems, but a novel simulation-like framework will be presented that allows efficient variance propagation for large interconnect trees. This has revealed that the variance can be highly significant. Even when the standard deviation of the inputs is small, at around 20--30% of the mean, we see a 30--40% drop in the lifetimes.
Adam Issa, Valeriy Sukharev, Farid N. Najm
ICCAD3
2019 Power Grid Fixing for Electromigration-induced Voltage Failures
abstract
Electromigration (EM) is a major reliability concern in chip power grids in the wake of smaller feature sizes. EM degradation of grid metal lines can cause large voltage drops on the grid, leading to timing failures and logic errors. During the design process, modifications to the grid design may be required in order to protect from the risk of such EM-induced voltage drop failures. We consider this problem in light of recent efficient full-chip EM assessment techniques. We present a systematic approach that resizes the grid metal lines to meet a design target lifetime while requiring minimal increase in metal area of the grid.
Zahi Moudallal, Valeriy Sukharev, Farid N. Najm
ICCAD3
2019 Power Scheduling With Active RC Power Grids
abstract
Power gating is widely used in large chip design as a way to manage the total power dissipation and avoid overheating. It works by turning OFF the power supply to circuit blocks that are not required to operate in certain operational modes. Many authors have studied the scheduling of chip workload to manage total power and temperature. But power gating also has an impact on the supply voltage levels across the die, because voltage drop is generated in the grid depending on the combination of blocks that are ON. We consider the question of how to manage the chip workload so that supply voltage variations remain within specs. The worst case voltage drop is the result of two things: the power budgets that were allocated to the various circuit blocks during the design process and the combination of blocks that are turned ON in a given operational mode. In this paper, we propose a framework to manage this tradeoff between how many blocks are ON simultaneously and how big the power budgets of the individual blocks are, assuming resistive and capacitive (RC) elements in the power grid model. Subject to user guidance, we generate block-level circuit current constraints as well as an implicit binary decision diagram (BDD) that helps identify the safe working modes. If the blocks are designed to respect these constraints, then the BDD can be used during normal operation to check whether a candidate working mode is safe or not.
Zahi Moudallal, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Power Grid Electromigration Checking Using Physics-Based Models
abstract
Due to technology scaling, electromigration (EM) signoff has become increasingly difficult, mainly due to the use of inaccurate methods for EM assessment, such as the empirical Black's model. In this paper, we present a novel finite-differencebased approach for power grid EM checking using physics-based models, that can account for process, voltage, and temperature variations across the die. Our main contribution is to extend existing physical models for EM in metal branches to track EM degradation in multibranch interconnect trees. The extended model is represented as a homogeneous linear time invariant system. We also detect early failures and account for their impact on grid lifetime. We speed up our implementation by proposing a macromodeling-based filtering scheme and a predictor-based approach. Our results, for a number of IBM power grid benchmarks, confirm that Black's model is overly inaccurate. The lifetimes found using our physics-based approach are on average 2.75× longer than those based on a (calibrated) Black's model, as extended to handle mesh power grids. With a maximum runtime of 2.3 h among all the IBM benchmarks, our method appears to be suitable for very large scale integration circuits.
Sandeep Chatterjee, Valeriy Sukharev, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Fast physics-based electromigration assessment by efficient solution of linear time-invariant (LTI) systems
abstract
Electromigration (EM) is a key reliability concern in chip power/ ground (p/g) grids, which has been exacerbated by the high current levels and narrow metal lines in modern grids. EM checking is expensive due to the large sizes of modern p/g grids and is also inherently difficult due to the complex nature of the EM phenomenon. Traditional EM checking, based on empirical models, cannot capture the complexity of EM and better models are needed for accurate prediction. Thus, recent physics-based EM models have been proposed, which remain computationally expensive because they require solution of a system of partial differential equations (PDEs). In this paper, we propose a fast and scalable methodology for power grid EM verification, building on previous physics-based models. We first convert the PDE system to a succession of homogeneous linear time invariant (LTI) systems. Because these systems are found to be stiff, we numerically integrate them using optimized variable-step backward differentiation formulas (BDFs). Our method, for a number of IBM power grids and internal benchmarks, achieves an average speed-up of over 20x as compared to previously published work and has a runtime of only about 8 minutes for a 4 million node grid.
Sandeep Chatterjee, Valeriy Sukharev, Farid N. Najm
ICCAD3
2017 Power grid verification under transient constraints
abstract
Checking the power grid must begin early in the design. One way of doing this is using vectorless verification which, unlike standard simulation, only requires limited information about the currents drawn from the grid, in the form of DC local and global upper-bounds, or current constraints. We extend the standard vectorless verification to allow transient constraints, where circuit currents may be bounded by different values at different times. This is useful to check the validity of candidate sequences of chip operations, each having different current requirements. We show that this framework leads to a less pessimistic estimation of voltage drops.
Mohammad Fawaz, Farid N. Najm
ICCAD2
2017 Power scheduling with active power grids
abstract
Power-gating is widely used in large chip design as a way to manage the total power dissipation and avoid overheating. It works by turning OFF the power supply to circuit blocks that are not required to operate in certain operational modes. Many authors have studied the scheduling of chip workload to manage total power and temperature. But power-gating also has an impact on the supply voltage levels across the die, because voltage drop is generated in the grid depending on the combination of blocks that are ON. We consider the question of how to manage the chip workload so that supply voltage variations remain within specs. The worst-case voltage drop is the result of two things, the power budgets that were allocated to the various circuit blocks during the design process and the combination of blocks that are turned ON in a given operational mode. Intuitively, more blocks can be turned ON simultaneously if the blocks are constrained to have low current levels, and vice versa. In this paper, we propose a framework to manage this trade-off between how many blocks are ON simultaneously and how big the power budgets of the individual blocks are, assuming resistive and capacitive (RC) elements in the power grid model. Subject to user guidance, we generate block-level circuit current constraints as well as an implicit binary decision diagram (BDD) that helps identify the safe working modes. If the blocks are designed to respect these constraints, then the BDD can be used during normal operation to check whether a candidate working mode is safe or not.
Zahi Moudallal, Farid N. Najm
ICCAD2
2017 Fast Vectorless RLC Grid Verification
abstract
Checking the power distribution network of an integrated circuit must start early in the design process, when changes to the grid can be more easily implemented. Vectorless verification is a technique that achieves this goal by demanding limited information about the currents drawn from the grid. State of the art techniques that deal with RLC grids become prohibitive even for medium size grids. In this paper, we propose a novel technique that estimates the worst-case voltage fluctuations for RLC grids by carefully selecting the time step, in a way that significantly reduces the number of linear programs that need to be solved, and eliminates the need for other expensive computations, like dense matrix-matrix multiplications. Results show that our technique is accurate and scalable for large grids as it achieves over 19× speedup over existing methods.
Mohammad Fawaz, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Generating Current Constraints to Guarantee RLC Power Grid Safety
abstract
A critical task during early chip design is the efficient verification of the chip power distribution network. Vectorless verification, developed since the mid-2000s as an alternative to traditional simulation-based methods, requires the user to specify current constraints (budgets) for the underlying circuitry and checks if the corresponding voltage variations on all grid nodes are within a user-specified margin. This framework is extremely powerful, as it allows for efficient and early verification, but specifying/obtaining current constraints remains a burdensome task for users and a hurdle to adoption of this framework by the industry. Recently, the inverse problem has been introduced: Generate circuit current constraints that, if satisfied by the underlying logic circuitry, would guarantee grid safety from excessive voltage variations. This approach has many potential applications, including various grid quality metrics, as well as voltage drop-aware placement and floorplanning. So far, this framework has been developed assuming only resistive and capacitive (RC) elements in the power grid model. Inductive effects are becoming a significant component of the power supply noise and can no longer be ignored. In this article, we extend the constraints generation approach to allow for inductance. We give a rigorous problem definition and develop some key theoretical results related to maximality of the current space defined by the constraints. Based on this, we then develop three constraints generation algorithms that target the peak total chip power that is allowed by the grid, the uniformity of current distribution across the die area, and a combination of both metrics.
Zahi Moudallal, Farid N. Najm
ACM Trans. Design Autom. Electr. Syst.2
2016 Accurate verification of RC power grids
Mohammad Fawaz, Farid N. Najm
DATE2
2016 Fast physics-based electromigration checking for on-die power grids
abstract
Due to technology scaling, electromigration (EM) signoff has become increasingly difficult, mainly due to the use of inaccurate methods for EM assessment, such as the empirical Black's model. In this paper, we present a novel approach for EM checking using physics-based models of EM degradation, which effectively removes the inaccuracy, with negligible impact on run-time. Our main contribution is to extend the existing physical models for EMin metal branches to track the degradation in multi-branch interconnect trees. We also propose effective filtering and predictor-based schemes to speed up our implementation, with minimal impact on accuracy. Our results, for a number of IBM power grid benchmarks, confirm that Black's model is overly inaccurate. The lifetimes found using our physics-based approach are on average 3× longer than those based on a (calibrated) Black's model, such as currently used in industry. For the two largest IBM benchmarks (700K branches each), our runtime is comparable to that of the Black's based approach, requiring 3 hours for the largest grid.
Sandeep Chatterjee, Valeriy Sukharev, Farid N. Najm
ICCAD3
2016 A fast layer elimination approach for power grid reduction
abstract
Simulation and verification of the on-die power delivery network (PDN) is one of the key steps in the design of integrated circuits (ICs). With the very large sizes of modern grids, verification of PDNs has become very expensive and a host of techniques for faster simulation and grid model approximation have been proposed. These include topological node elimination, as in TICER and full-blown numerical model order reduction (MOR) as in PRIMA and related methods. However, both of these traditional approaches suffer from certain drawbacks that make them expensive and limit their scalability to very large grids. In this paper, we propose a novel technique for grid reduction that is a hybrid of both approaches-the method is numerical but also factors in grid topology. It works by eliminating whole internal layers of the grid at a time, while aiming to preserve the dynamic behavior of the resulting reduced grid. Effectively, instead of traditional node-by-node topological elimination we provide a numerical layer-by-layer block-matrix approach that is both fast and accurate. Experimental results show that this technique is capable of handling very large power grids and provides a 4.25× speed-up in transient analysis.
Abdul-Amir Yassine, Farid N. Najm
ICCAD2
2016 Generating voltage drop aware current budgets for RC power grids
abstract
Efficient verification of the chip power distribution network is a critical task in modern chip design. It should be done early in the design process where adjustments can be most easily incorporated. As an alternative to simulation based methods, vectorless verification is a class of techniques that requires user-specified current constraints (budgets), and checks if the corresponding worst-case voltage drops at all grid nodes are below user-specified thresholds. However, obtaining/specifying the current constraints remains a burdensome task for users. Recent literature has addressed the constraints generation problem by proposing the inverse problem: for a given grid, we would like to generate circuit current constraints which, if adhered to by the underlying logic, would guarantee grid safety. In this paper, we adopt the same framework. We develop an efficient algorithm for constraints generation that targets a key grid quality metric namely the uniformity of temperature distribution across the die area.
Zahi Moudallal, Farid N. Najm
ISCAS2
2016 Generating Current Budgets to Guarantee Power Grid Safety
abstract
Efficient and early verification of the chip power distribution network is a critical step in modern chip design. Vectorless verification, developed over the last decade as an alternative to simulation-based methods, requires user-specified current constraints (budgets) and checks if the corresponding worst-case voltage drops at all grid nodes are below user-specified thresholds. However, obtaining/specifying the current constraints remains a burdensome task for users. In this paper, we define and address the inverse problem: for a given grid, we would like to generate circuit current constraints which, if adhered to by the underlying logic, would guarantee grid safety. There are many potential applications for this approach, including various grid quality metrics, as well as voltage drop aware placement and floorplanning. We give a rigorous problem definition and develop some key theoretical results related to maximality of the current space defined by the constraints. Based on this, we then develop two algorithms for constraints generation that target the peak total chip power that is allowed by the grid and the uniformity of the temperature distribution. Finally, we develop a superior algorithm which targets a combination of both quality metrics.
Zahi Moudallal, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Verification of the Power and Ground Grids Under General and Hierarchical Constraints
abstract
As part of power distribution network verification, one should check if the voltage fluctuations exceed some critical threshold. The traditional simulation-based solution to this problem is intractable due to the large number of possible circuit behaviors. This approach also requires full knowledge of the details of the underlying circuitry, not allowing one to verify the power distribution network early in the design flow. Contrary to previous work on power distribution network verification, we consider the power and ground (P/G) grids together and describe an early verification approach under the framework of current constraints. Then, we present a solution technique in which tight lower and upper bounds on worst case voltage fluctuations are computed via linear programs. Experimental results indicate that the proposed technique results in errors in the range of a few millivolts. In addition to P/G grid verification techniques, we also provide very efficient solution technique to power (single) grid verification under hierarchical current constraints.
Mehmet Avci, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Generating circuit current constraints to guarantee power grid safety
abstract
Efficient and early verification of the chip power distribution network is a critical step in modern IC design. Vectorless verification, developed over the last decade as an alternative to simulation based methods, requires user-specified current constraints and checks if the corresponding worst-case voltage drops at all grid nodes are below user-specified thresholds. However, obtaining/specifying the current constraints remains a burdensome task for design teams. In this paper, we define and address the inverse problem: for a given grid, we will generate circuit current constraints which, if adhered to by the underlying logic, would guarantee grid safety. There are many potential applications for this approach, including various grid quality metrics, as well as power grid-aware placement and floorplanning. We give a rigorous problem definition and develop some key theoretical results related to maximality of the current space defined by the constraints. Based on this, we then develop two algorithms for constraints generation that target key quality metrics like the peak total power allowed by the grid and the uniformity of the temperature distribution.
Zahi Moudallal, Farid N. Najm
ASP-DAC2
2015 Physical Design Challenges in the Chip Power Distribution Network
abstract
The power supply and ground networks in large integrated circuits or, simply, the power grids, have become very large billion-node metal interconnect structures that often span all levels of the metal stack. The grid may be connected to about 2,000 C4 pads at the top layers and to hundreds of millions of gates and other circuitry at the bottom. It is not uncommon to reserve the top metal layers exclusively for the power grid. However, the extensive use of metal resources on lower metal layers for the grid has become a real bottleneck for signal routing. This adds time and cost to the overall chip design project and represents a problem for physical design. Yet there are reasons to believe that allocation of so much metal resources to the grid is ``overkill'' and that there is much room for improvement. The grid is over-designed because of lack of certainty about its safety from various concerns, like electromigration, IR drop, and inductive drop. There are also open problems in the grid design problem itself, which may be viewed as an optimization problem, albeit a very difficult one. In this talk, I will review developments in the verification of power grids that aim to provide certainty that the grid is safe, and indicate directions for possible ways that the grid may be automatically generated to suit various objectives.
Farid N. Najm
ISPD1
2015 Redundancy-Aware Power Grid Electromigration Checking Under Workload Uncertainties
abstract
Electromigration (EM) in on-die metal lines is becoming a significant problem in modern integrated circuits technology. Due to the high levels of current density on the die, the large number of metal lines, and the inherent conservatism in classical full-chip EM models, designers are finding it very hard to meet the area and design specs while guaranteeing EM reliability. The EM problem is most significant in power grid lines, because unlike signal and clock lines, they do not benefit from healing due to their mostly unidirectional currents. In this paper, we develop a new model, referred to as the mesh model, for power grid EM checking which takes into account the inherent redundancy of its mesh structure while determining the reliability. To implement the mesh model, we also develop a framework to estimate the change in statistics of an interconnect as its effective-EM current varies. In order to overcome the conservative assumptions that designers usually make about chip workloads, we also propose a novel vectorless mesh model technique to estimate the average minimum time-to-failure of a power grid under workload uncertainties. The results indicate that the series model, which is currently used in the industry, gives a pessimistic estimate of power grid MTF and reliability by a factor of 3-4. Finally, we exploit multithreading and grid locality to speedup our implementation by almost $6{\times }$ .
Sandeep Chatterjee, Mohammad Fawaz, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Redundancy-aware electromigration checking for mesh power grids
abstract
Electromigration (EM) is re-emerging as a significant problem in modern integrated circuits (IC). Especially in power grids, due to shrinking wire widths and increasing current densities, there is little or no margin left between the predicted EM stress and that allowed by the EM design rules. Statistical Electromigration Budgeting (SEB) estimates the reliability of the grid by considering it entirely as a series system. However, a power grid with its many parallel paths has much inherent redundancy. In this paper, we propose a new model to estimate the MTF and reliability of the power grid under the influence of EM, which accounts for these redundancies. We refer to this as the mesh model. To implement the mesh model, we also develop a framework to estimate the change in statistics of an interconnect as its effective-EM current varies. The proposed algorithm is quite fast and has an overall observed empirical complexity of 0(n1.4). The results indicate that the series model, which is currently used in the industry, gives a pessimistic estimate of power grid MTF and reliability by a factor of 3-4.
Sandeep Chatterjee, Mohammad Fawaz, Farid N. Najm
ICCAD3
2013 A vectorless framework for power grid electromigration checking
abstract
Electromigration (EM) in the on-die metal lines has re-emerged as a significant concern in modern VLSI circuits. The higher levels of temperature on die and the very large number of metal lines, coupled with the conservatism inherent in traditional EM checking strategies, have led to a situation where trying to guarantee EM reliability often leads to unacceptably conservative designs that may not meet the area or performance specs. Due to unidirectional currents, this problem is most significant in the power and ground grids. Thus, this work is aimed at reducing the pessimism in EM prediction for power/ground grids. There are two sources for the high pessimism: 1) the use of the traditional series model for EM checking and 2) pessimistic assumptions about the chip workload and the corresponding supply currents. To address this problem, we propose a framework for EM checking that allows users to specify conditions-of-use type constraints that help capture realistic chip workload and which includes the use of a novel mesh model for EM prediction in the grid, instead of the traditional series model.
Mohammad Fawaz, Sandeep Chatterjee, Farid N. Najm
ICCAD3
2012 Incremental power grid verification
abstract
Verification of the on-die power grid is a key step in the design of complex high-performance integrated circuits. For the very large grids in modern designs, incremental verification is highly desirable, because it allows one to skip the verification of a certain section of the grid (internal nodes) and instead, verify only the rest of the grid (external nodes). We propose an efficient approach for incremental verification in the context of vectorless constraints-based grid verification, under dynamic conditions. The traditional difficulty is that the dynamic case requires iterative analysis of both the internal and external sections. This has been previously overcome for simulation purposes, but we provide the first solution for verification, through two key contributions: 1) a bound on the internal nodes' voltages is developed that eliminates the need for iterative analysis, and 2) a multi-port Norton approach is used to construct a reduced macromodel for the internal section. As a result, we demonstrate significant reductions in runtime, with speed-ups in the range of 3-8x, with negligible impact on accuracy.
Farid N. Najm
DAC2
2012 Overview of vectorless/early power grid verification
abstract
The power distribution network of an integrated circuit must be checked throughout the design process to ensure that supply voltage fluctuations do not exceed certain critical thresholds. One way of doing this is by simulation, which requires knowledge of the circuit currents that load the grid. These currents are hard to specify. In many cases, and certainly during early power grid design, they may be simply unknown because the circuit itself may not yet be specified. Vectorless verification refers to the class of techniques, developed over the last 12 years, for verifying the grid in the absence of complete information about the circuit currents.
Farid N. Najm
ICCAD1
2012 Efficient Block-Based Parameterized Timing Analysis Covering All Potentially Critical Paths
abstract
In order for the results of timing analysis to be useful, they must provide insight and guidance on how the circuit may be improved so as to fix any reported timing problems. A limitation of many recent variability-aware timing analysis techniques is that, while they report delay distributions, or verify multiple corners, they do not provide the required guidance for re-design. We propose an efficient block-based parameterized timing analysis technique that can accurately capture circuit delay at every point in the parameter space, by reporting all paths that can become critical. Using an efficient pruning algorithm, only those potentially critical paths are carried forward, while all other paths are discarded during propagation. This allows one to examine local robustness to parameters in different regions of the parameter space, not by considering differential sensitivity at a point (that would be useless in this context) but by knowledge of the paths that can become critical at nearby points in parameter space. We give a formal definition of this problem and propose a technique for solving it, which improves on the state of the art, both in terms of theoretical computational complexity and in terms of runtime on various test circuits.
Khaled R. Heloue, Sari Onaissi, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2012 Maximum Circuit Activity Estimation Using Pseudo-Boolean Satisfiability
abstract
With lower supply voltages, increased integration densities and higher operating frequencies, power grid verification has become a crucial step in the very large-scale integration design cycle. The accurate estimation of maximum instantaneous power dissipation aims at finding the worst-case scenario where excessive simultaneous switching could impose extreme current demands on the power grid. This problem is highly input-pattern dependent and is proven to be NP-hard. In this paper, we capitalize on the compelling advancements in satisfiability (SAT) solvers to propose a pseudo-Boolean SAT-based framework that reports the input patterns maximizing circuit activity, and consequently peak dynamic power, in combinational and sequential circuits. The proposed framework is enhanced to handle unit gate delays and output glitches. In order to disallow unrealistic input transitions, we show how to integrate input constraints in the formulation. Finally, a number of optimization techniques, such as the use of gate switching equivalence classes, are described to improve the scalability of the proposed method. An extensive suite of experiments on ISCAS85 and ISCAS89 circuits confirms the robustness of the approach compared to simulation-based techniques and encourages further research for low-power solutions using Boolean SAT.
Hratch Mangassarian, Andreas G. Veneris, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 Power grid verification using node and branch dominance
abstract
The verification of power grids in modern integrated circuits must start early in the design process when adjustments can be most easily incorporated. This work describes a vectorless verification technique that deals with circuit uncertainty in the framework of current constraints. In such a framework, grid verification becomes a question of computing the worst-case voltage drops which, in turn, entails the solution of as many linear programs (LPs) as there are nodes. First, we extend grid verification to also check for the worst-case branch currents. We show that this would require as many LPs as there are branches. Second, we propose a starkly different approach to reduce the number of LPs in the verification problem. We achieve this by examining dominance relations among node voltage drops and among branch currents. This allows us to replace a group of LPs by one conservative and tight LP. Results show a dramatic reduction in the number of LPs thus making vectorless grid verification in the framework of current constraints practical and scalable.
Nahi H. Abdul Ghani, Farid N. Najm
DAC2
2011 Power grid correction using sensitivity analysis under an RC model
abstract
Verifying an RC model of the power grid requires one to check if the steady state voltage drops on all the nodes of the grid do not exceed a certain threshold. We propose an approach to correct the grid, in case some voltage drops violate the threshold condition, by making minor changes to the original design. Previous work has been done in [1] on the DC model of the grid and this paper deals with the transient model. Rather than directly reducing the steady state voltage drops below the threshold we work on reducing the first time step voltage drops. The method uses current constraints proposed in [2] to find the first time step voltage drop whose distance to the corresponding threshold is the largest. It then tries to estimate it as a function of the metal widths on the grid. A non-linear optimization problem is then formulated and the required metal line width changes that reduce the first time step voltage drops by a sufficient amount are then determined. The reduction of the first time step voltage drop by that amount will make the steady state voltage drops of all the nodes less than the threshold.
Pamela Al Haddad, Farid N. Najm
DAC2
2011 A fast approach for static timing analysis covering all PVT corners
abstract
The increasing sensitivity of circuit performance to process, temperature, and supply voltage (PVT) variations has led to an increase in the number of process corners that are required to verify circuit timing. Typically, designers attempt to reduce this computational load by choosing, based on experience, a subset of the available corners and running static timing analysis (STA) at only these corners. Although running a few corners, which are chosen beforehand, can lead to acceptable results in some cases, this is not always the case. Our results show that in the case of setup timing analysis, one can indeed bound circuit slacks across all corners by running a small number of corners. On the other hand, we show that this is not possible in the case of hold analysis. Instead, we present an alternative method for performing fast and accurate hold timing analysis which covers all corners. In this method a full timing run is performed for a small number of corners, and partial timing runs, which cover only the clock network, are performed for others. We then combine the results of the full and partial runs to find the worst-case hold slacks over all corners. Our results show that this method is accurate and can achieve much improved runtimes.
Sari Onaissi, Feroze Taraporevala, Farid N. Najm
DAC4
2011 Efficient RC power grid verification using node elimination
abstract
To ensure the robustness of an integrated circuit, its power distribution network (PDN) must be validated beforehand against any voltage drop on VDD nets. However, due to the increasing size of PDNs, it is becoming difficult to verify them in a reasonable amount of time. Lately, much work has been done to develop Model Order Reduction (MOR) techniques to reduce the size of power grids but their focus is more on simulation. In verification, we are concerned about the safety of nodes, including the ones which have been eliminated in the reduction process. This paper proposes a novel approach to systematically reduce the power grid and accurately compute an upper bound on the voltage drops at power grid nodes which are retained. Furthermore, a criterion for the safety of nodes which are removed is established based on the safety of other nearby nodes and a user specified margin.
Farid N. Najm
DATE2
2011 Fast Vectorless Power Grid Verification Under an RLC Model
abstract
As part of early system design, one must verify that the power grid provides the underlying logic circuitry with voltage levels that are within specified ranges. In this paper, we describe a vectorless verification approach that can be applied early in the design process. We adopt an RLC model of the grid in the framework of current constraints that capture uncertainty about circuit details and activity. With just a few linear programs and one linear system solve, our proposed approach provides tight conservative bounds on the maximum and minimum worst-case voltage drops at every node of the grid. Results show the accuracy and speed of our technique thus making it practical and scalable.
Nahi H. Abdul Ghani, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Managing verification error traces with bounded model debugging
abstract
Managing long verification error traces is one of the key challenges of automated debugging engines. Today, debuggers rely on the iterative logic array to model sequential behavior which drastically limits their application. This work presents bounded model debugging, an iterative, systematic and practical methodology to allow debuggers to tackle larger problems than previously possible. Based on the empirical observation that errors are excited in temporal proximity of the observed failures, we present a framework that improves performance by up to two orders of magnitude and solve 2.7x more problems than a conventional debugger.
Sean Safarpour, Andreas G. Veneris, Farid N. Najm
ASP-DAC3
2010 Early P/G grid voltage integrity verification
abstract
As part of power delivery network verification, one should check if the voltage fluctuations exceed some critical threshold. In this work, we consider the power and ground grids together and describe an early verification approach under the framework of current constraints where tight lower and upper bounds on worst-case voltage fluctuations are computed via linear programs. Experimental results indicate that the proposed technique results in errors in the range of a few mV.
Mehmet Avci, Farid N. Najm
ICCAD2
2010 Power grid correction using sensitivity analysis
abstract
Power grid voltage integrity verification requires one to check if all the voltage drops on the grid are less than a certain threshold. This paper addresses the problem of correcting the grid when some voltage drops exceed this threshold, by making minor modifications to the existing design. The method uses current constraints that capture the uncertainty about the underlying circuit behavior to find the maximum voltage drop on the grid, and then to estimate the voltage drop as a function of the metal widths on the grid. It formulates a non-linear optimization problem and finds the required change in widths that reduces the maximum voltage drop below the threshold while keeping the total area cost at a minimum.
Meric Aydonat, Farid N. Najm
ICCAD2
2010 Verification and Codesign of the Package and Die Power Delivery System Using Wavelets
abstract
As part of the design of large integrated circuits, one must verify that the power delivery network provides supply and ground voltages to the circuit that are within specified ranges. We introduce the concept of time-frequency description of circuit currents using wavelets, and use that to set up an optimization framework that finds the worst-case supply/ground voltage fluctuations. This framework allows for the quick determination of the impact of either the package or the die on the worst-case behavior, which enables their codesign. This approach has been applied to an industrial microprocessor design, resulting in realistic and nonobvious worst-case waveforms.
Imad A. Ferzli, Eli Chiprout, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Fast vectorless power grid verification using an approximate inverse technique
abstract
Power grid verification in modern integrated circuits is an integral part of early system design where adjustments can be most easily incorporated. In this work, we describe an early verification approach under the framework of current constraints where worst-case node voltage drops are computed via linear programs proportional to the grid size. We propose an efficient method based on a sparse approximate inverse technique to greatly reduce the size of such linear programs while ensuring a user-specified over-estimation margin (in volts) on the exact solution.
Nahi H. Abdul Ghani, Farid N. Najm
DAC2
2009 Clock skew optimization via wiresizing for timing sign-off covering all process corners
abstract
Manufacturing process variability impacts the performance of synchronous logic circuits by means of its effect on both clock network and functional block delays. Typically, variability in clock networks is either handled early in the design flow by assigning margins to clock network delays, or at a later stage through post-processing steps that only focus on achieving minimal skew, without regard to functional block variability. In this work, we present a technique that alters clock network lines so that the circuit meets its timing constraints at all process corners. This is done near the end of the design flow while considering delay variability in both the clock network and the functional blocks. Our method operates at the physical level and provides designers with the required changes in clock network line widths and/or lengths. This can be formulated as a Linear Programming (LP) problem, and thus can be solved efficiently. Empirical results for a set of ISCAS-89 benchmark circuits show that our approach can considerably reduce the effect of process variations on circuit performance.
Sari Onaissi, Khaled R. Heloue, Farid N. Najm
DAC3
2009 Quantifying robustness metrics in parameterized static timing analysis
abstract
Process and environmental variations continue to present significant challenges to designers of high-performance integrated circuits. In the past few years, while much research has been aimed at handling parameter variations as part of timing analysis, few proposals have actually included ways to interpret the results of this parameterized static timing analysis (PSTA) step. In this paper, we propose a new post-variational analysis metric that can be used to quantify the robustness of designs to parameter variations. In addition to helping designers diagnose if and when different nodes can fail, this metric can give insights on what to fix, by identifying nodes with small robustness values and proceeding to fix those nodes first. Inspired by the rich literature on design centering, to lerancing, and tuning (DCTT), we use distance as a measure for robustness. Our analysis thus determines the minimum distance from the nominal point in the parameter space to any timing violation, and works under the assumption that parameters are specified as ranges rather than statistical distributions. We demonstrate the usefulness of this distance-based robustness metric on circuit blocks extracted from a commercial 45nm microprocessor.
Khaled R. Heloue, Chandramouli V. Kashyap, Farid N. Najm
ICCAD3
2009 PSTA-based branch and bound approach to the silicon speedpath isolation problem
abstract
The lack of good "correlation" between pre-silicon simulated delays and measured delays on silicon (silicon data) has spurred efforts on so-called silicon debug. The identification of speed-limiting paths, or simply speedpaths, in silicon debug is a crucial step, required for both "fixing" failing paths and for accurate learning from silicon data. We propose using characterized, pre-silicon, variational timing models to identify speedpaths that can best explain the observed delays from silicon measurements. Delays of all logic paths are written as affine functions of process parameters, called hyperplanes, and a branch and bound approach is then applied to find the "best" path combinations. Our method has been tested on a set of ISCAS-89 circuits and the results show that it accurately identifies the speedpaths in most cases, and that this is achieved in a very efficient manner.
Sari Onaissi, Khaled R. Heloue, Farid N. Najm
ICCAD3
2009 Full-Chip Model for Leakage-Current Estimation Considering Within-Die Correlation
abstract
In this paper, we present an efficient technique for finding the mean and variance of the full-chip leakage of a candidate design, while considering logic structures and both die-to-die and within-die (WID) process variations, and taking into account the spatial correlation due to WID variations. Our model uses a ldquorandom-gaterdquo concept to capture high-level characteristics of a candidate chip design, which are sufficient to determine its leakage. These high-level characteristics include information about the process, the standard cell library, and expected design characteristics. We show empirically that, for large gate count, the set of all chip designs that share the same high-level characteristics have approximately the same leakage, with very small error. Therefore, our model can be used as either anearlyor alateestimator of leakage, with high accuracy. In its simplest form, we show that full-chip-leakage estimation reduces in finding the area under a scaled version of the WID channel length autocorrelation function, which can be done in constant time.
Khaled R. Heloue, Navid Azizi, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Low-Power Programmable FPGA Routing Circuitry
abstract
We consider circuit techniques for reducing field-programmable gate-array (FPGA) power consumption and propose a family of new FPGA routing switch designs that are programmable to operate in three different modes: high-speed, low-power, or sleep. High-speed mode provides similar power and performance to traditional FPGA routing switches. In low-power mode, speed is curtailed in order to reduce power consumption. Leakage is reduced by 28%-52% in low-power versus high-speed mode, depending on the particular switch design selected. Dynamic power is reduced by 28%-31% in low-power mode. Leakage power in sleep mode, which is suitable for unused routing switches, is 61%-79% lower than in high-speed mode. Each of the proposed switch designs has a different power/area/speed tradeoff. All of the designs require only minor changes to a traditional routing switch and involve relatively small area overhead, making them easy to incorporate into current commercial FPGAs. The applicability of the new switches is motivated through an analysis of timing slack in industrial FPGA designs. It is observed that a considerable fraction of routing switches may be slowed down (operate in low-power mode), without impacting overall design performance.
Jason Helge Anderson, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Parameterized timing analysis with general delay models and arbitrary variation sources
abstract
Many recent techniques for timing analysis under variability, in which delay is an explicit function of underlying parameters, may be described as parameterized timing analysis. The "max" operator, used repeatedly during block-based timing analysis, causes several complications during parameterized timing analysis. We introduce bounds on, and an approximation to, the max operator which allow us to develop an accurate, general, and efficient approach to parameterized timing, which can handle either uncertain or random variations. Applied to random variations, the approach is competitive with existing statistical static timing analysis (SSTA) techniques, in that it allows for nonlinear delay models and arbitrary distributions. Applied to uncertain variations, the method is competitive with existing multi-corner STA techniques, in that it more reliably reproduces overall circuit sensitivity to variations. Crucially, this technique can also be applied to the mixed case where both random and uncertain variations are considered. Our results show that, on average, circuit delay is predicted with less than 2% error for multi-corner analysis, and less than 1% error for SSTA.
Khaled R. Heloue, Farid N. Najm
DAC2
2008 Efficient block-based parameterized timing analysis covering all potentially critical paths
abstract
In order for the results of timing analysis to be useful, they must provide insight and guidance on how the circuit may be improved so as to fix any reported timing problems. A limitation of many recent variability-aware timing analysis techniques is that, while they report delay distributions, or verify multiple corners, they do not provide the required guidance for re-design. We propose an efficient block-based parameterized timing analysis technique that can accurately capture circuit delay at every point in the parameter space, by reporting all paths that can become critical. Using an efficient pruning algorithm, only those potentially critical paths are carried forward, while all other paths are discarded during propagation. This allows one to examine local robustness to parameters in different regions of the parameter space, not by considering differential sensitivity at a point (which would be useless in this context) but by knowledge of the paths that can become critical at nearby points in parameter space. We give a formal definition of this problem and propose a technique for solving it that improves on the state of the art, both in terms of theoretical computational complexity and in terms of run time on various test circuits.
Khaled R. Heloue, Sari Onaissi, Farid N. Najm
ICCAD3
2008 Early Analysis and Budgeting of Margins and Corners Using Two-Sided Analytical Yield Models
abstract
Manufacturing process variations lead to variability in circuit delay and, if not accounted for, can cause excessive timing yield loss. The familiar traditional approaches to timing verification, such as the use of process corners and predefined timing margins, cannot readily handle within-die variations. Recently, statistical static timing analysis (SSTA) has been proposed as a way to deal with variability. Although many powerful techniques have been proposed, the fact that SSTA requires a significant change of methodology has delayed its wide adoption. In this paper, we propose a framework whereby the familiar concepts of corners and margins, which are generally meaningful at the transistor or cell level, are elevated to the chip level in order to handle within-die variations. This is achieved by using high-level models, such as the generic path model or the generic circuit model with different classes of paths, to represent the behavior of typical designs. These models allow us to determine ldquoyield-specificrdquo margins (setup and hold margins) and virtual corners, which, if applied during standard (deterministic) timing analysis, would guarantee the desired yield. Our framework can be used at an early stage of circuit design and is consistent with traditional timing verification methodology.
Khaled R. Heloue, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2008 A Linear-Time Approach for Static Timing Analysis Covering All Process Corners
abstract
Manufacturing process variations lead to circuit timing variability and a corresponding timing yield loss. Traditional corner analysis consists of checking all process corners (combinations of process parameter extremes) to make sure that circuit timing constraints are met at all corners, typically by running static timing analysis (STA) at every corner. This approach is becoming too expensive due to the increase in the number of corners with modern processes. As an alternative, we propose a linear-time approach for STA which covers all process corners in a single pass. Our technique assumes a linear dependence of delays and slews on process parameters and provides estimates of the worst case circuit delay and slew. It exhibits high accuracy in practice, and if the circuit has gates and relevant process parameters, the complexity of the algorithm is O(mn).
Sari Onaissi, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Modeling and Estimation of Full-Chip Leakage Current Considering Within-Die Correlation
abstract
We present an efficient technique for finding the mean and variance of the full-chip leakage of a candidate design, while considering logic-structures and both die-to-die and within-die process variations, and taking into account the spatial correlation due to within-die variations. Our model uses a random concept to capture high-level characteristics of a candidate chip design, which are sufficient to determine its leakage. We show empirically that, for large gate count, the set of all chip designs that share the same high level characteristics have approximately the same leakage, with very small error. Therefore, our model can be used as either an early or a late estimator of leakage, with high accuracy. In its simplest form, we show that full-chip leakage estimation reduces to finding the area under a scaled version of the within-die channel length auto-correlation function, which can be done in constant time.
Khaled R. Heloue, Navid Azizi, Farid N. Najm
DAC3
2007 Maximum circuit activity estimation using pseudo-boolean satisfiability
Hratch Mangassarian, Andreas G. Veneris, Sean Safarpour, Farid N. Najm, Magdy S. Abadir
DATE4
2007 A geometric approach for early power grid verification using current constraints
abstract
The verification of power grids in modern integrated circuits must start, at design time, where circuit information is unknown but could be specified or inferred from design or architectural considerations. This work builds on previously proposed techniques to deal with circuit uncertainty in the framework of linear current constraints, but proposes a cost-controlled solution, by following a geometric approach, and transforming a problem that requires as many linear programs as there are power grid nodes, to another involving a user-limited number of solutions of one linear system.
Imad A. Ferzli, Farid N. Najm, Lars Kruse
ICCAD2
2007 Early power grid verification under circuit current uncertainties
abstract
As power grid safety becomes increasingly important in modern integrated circuits, so does the need to start power grid verification early in the design cycle and incorporate circuit uncertainty of the early stages into useful power grid information. This work adopts the framework of capturing circuit uncertainty via constraintson circuit currents, and follows a geometric approachto transform a problem whose solution requires as many linear programs as there are power grid nodes, to another involving a user-limited number of solutions of one linear system.
Imad A. Ferzli, Farid N. Najm, Lars Kruse
ISLPED2
2007 A Yield Model for Integrated Circuits and its Application to Statistical Timing Analysis
abstract
A model for process-induced parameter variations is proposed, combining die-to-die, within-die systematic, and within-die random variations. This model is put to use toward finding suitable timing margins and device file settings, to verify whether a circuit meets a desired timing yield. While this parameter model is cognizant of within-die correlations, it does not require specific variation models, layout information, or prior knowledge of intrachip covariance trends. The approach works with a "generic" critical path, leading to what is referred to as a "process-specific" statistical-timing-analysis technique that depends only on the process technology, transistor parameters, and circuit style. A key feature is that the variation model can be easily built from process data. The derived results are "full-chip," applicable with ease to circuits with millions of components. As such, this provides a way to do a statistical timing analysis without the need for detailed statistical analysis of every path in the design
Farid N. Najm, Noel Menezes, Imad A. Ferzli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 Variations-Aware Low-Power Design and Block Clustering With Voltage Scaling
abstract
We present a new methodology which takes into consideration the effect of within-die (WID) process variations on a low-voltage parallel system. We show that in the presence of process variations one should use a higher supply voltage than would otherwise be predicted to minimize the power consumption of a parallel systems. Previous analyses, which ignored WID process variations, provide a lower nonoptimal supply voltage which can underestimate the energy/operation by 8.2. We also present a novel technique to limit the effect of temperature variations in a parallel system. As temperatures increases, the scheme reduces the power increase by 43% allowing the system to remain at it's optimal supply voltage across different temperatures. To further limit the effect of variations, and allow for a reduced power consumption, we analyzed the effects of clustering. It was shown that providing different voltages to each cluster can provide a further 10% reduction in energy/operation to a low-voltage parallel system, and that the savings by clustering increase as technology scales.
Navid Azizi, Muhammad M. Khellah, Vivek De, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.4
2006 A family of cells to reduce the soft-error-rate in ternary-CAM
abstract
Modern integrated circuits require careful attention to the soft-error rate (SER) resulting from bit upsets, which are normally caused by alpha particle or neutron hits. These events, also referred to as single-event upsets (SEUs), will become more problematic in future technologies. This paper presents a ternary content-addressable memory (CAM) design with high immunity to SEU. Conventionally, error-correcting codes (ECC) have been used in SRAMs to address this issue, but these techniques are not immediately applicable to CAMs because they depend on processing the full contents of the memory word outside the array, which is not possible in a normal CAM access. We propose a family of TCAM cells that reduce the SER at the cost of some area increase. An SER reduction of up to 40% can be obtained with a 18% increase of area; another design reduces the SER by 16% with only a 5% increase in area.
Navid Azizi, Farid N. Najm
DAC2
2006 An adaptive FPGA architecture with process variation compensation and reduced leakage
abstract
Process induced threshold voltage variations bring about fluctuations in circuit delay, that affect the FPGA timing yield. We propose an adaptive FPGA architecture that compensates for these fluctuations. The architecture includes an additional characterizer circuit that classifies logic and routing blocks on each die according to their performance. Based on this classification, the architecture adaptively body-biases these resources by either speeding up the slow blocks or by slowing down the leaky ones. This procedure mitigates the effect of the variations and provides a better yield. We further diminish leakage by slowing down areas of the FPGA that have a positive slack. Overall, this architecture minimizes the timing variance of within-die and die-to-die Vth variations by up to 3.45X and reduces leakage power in the non-critical areas of the FPGA by 3X with no effect on frequency.
Georges Nabaa, Navid Azizi, Farid N. Najm
DAC3
2006 Handling inductance in early power grid verification
abstract
As part of integrated circuit design verification, one should check if the voltage drop on the power grid exceeds some critical threshold. One way to do this is by simulation, but that is computationally expensive and gets prohibitive for large circuits with a large variety of possible operational modes. Another limitation of a simulation-based approach is that it requires complete knowledge of the logic circuitry drawing current from the grid, thus precluding grid verification early in the design process. In this paper, we model the grid as an RLC circuit and we propose three verification techniques that can be applied in the early stages of the design process. These techniques do not require exact knowledge of the circuit currents. Instead, the currents drawn by the logic beneath the power grid are described by means of current constraints that capture the uncertainty about circuit details and activity. The first verification approach gives the exact worst-case voltage drop at every node of the grid, but it is slow. A second faster approach gives conservative bounds on the worst-case voltage drop at every node of the grid. The third approach is much faster; it is a conservative approach which simply checks if the grid voltage drop exceeds some pre-defined thresholds, without actually computing the worst-case voltage drop at every node.
Nahi H. Abdul Ghani, Farid N. Najm
ICCAD2
2006 A linear-time approach for static timing analysis covering all process corners
abstract
Manufacturing process variations lead to circuit timing variability and a corresponding timing yield loss. Traditional corner analysis consists of checking all process corners (combinations of process parameter extremes) to make sure that circuit timing constraints are met at all corners, typically by running static timing analysis (STA) at every corner. This approach is becoming too expensive due to the exponential increase in the number of corners with modern processes. As an alternative, we propose a linear-time approach for STA which covers all process corners in a single pass. Our technique assumes a linear dependence of delay on process parameters and provides tight bounds on the worst-case circuit delay. It exhibits high accuracy (within 1-3%) in practice and, if the circuit has m gates and n relevant process parameters, the complexity of the algorithm is O(mn).
Sari Onaissi, Farid N. Najm
ICCAD2
2006 Active leakage power optimization for FPGAs
abstract
Active leakage power dissipation is considered in field-programmable gate arrays (FPGAs) and two "no cost" approaches for active leakage reduction are presented. It is well known that the leakage power consumed by a digital CMOS circuit depends strongly on the state of its inputs. The authors' first leakage reduction technique leverages a fundamental property of basic FPGA logic elements [look-up tables (LUTs)] that allows a logic signal in an FPGA design to be interchanged with its complemented form without any area or delay penalty. This property is applied to select polarities for logic signals so that FPGA hardware structures spend the majority of time in low-leakage states. In an experimental study, active leakage power is optimized in circuits mapped into a state-of-the-art 90-nm commercial FPGA. Results show that the proposed approach reduces active leakage by 25%, on average. The authors' second approach to leakage optimization consists of altering the routing step of the FPGA computer-aided design (CAD) flow to encourage more frequent use of routing resources that have low leakage power consumptions. Such "leakage-aware routing" allows active leakage to be further reduced, without compromising design performance. Combined, the two approaches offer a total active leakage power reduction of 30%, on average.
Jason Helge Anderson, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 High-level current macro model for logic blocks
abstract
The authors present a frequency domain current macro-modeling technique for capturing the dependence of the block current waveform on its input vectors. The macro model is based on estimating the discrete cosine transform (DCT) of the current waveform and then taking the inverse transform to estimate the time domain current waveform. The DCT of a current waveform is very regular and closely resembles the DCT of a triangular or a trapezoidal wave. The authors use this fact and the relation between the DCT, discrete Fourier transform (DFT), discrete time Fourier transform (DTFT), and the Fourier transform (FT) to infer the template functions for the current macro model. These template functions are characterized by using various parameters like amplitude, phase decay factor, time period, etc. These parameters are modeled as functions of the input vector pair using regression. Regression is done on a set of current waveforms generated for each circuit using HSPICE. These template functions are used in an automatic characterization process to generate current macro models for various CMOS combinational circuits.
Srinivas Bodapati, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 Analysis and verification of power grids considering process-induced leakage-current variations
abstract
The ongoing trends in technology scaling imply a reduction in the transistor threshold voltage (V/sub th/). With smaller feature lengths and smaller parameters, variability becomes increasingly important, for ignoring it may lead to chip failure and assuming worst case renders almost any design nonachievable. This paper presents a methodology for the analysis and verification of the power grid of integrated circuits considering variations in leakage currents. These variations are large due to the exponential relation between leakage current and transistor threshold voltage and appear as random background noise on the nodes of the grid. We propose a lognormal distribution to model the grid voltage drops, derive bounds on the voltage-drop variances, and develop a numerical Monte Carlo method to estimate the variance of each node voltage on the grid. This model is used toward the solution of a statistical formulation of the power-grid-verification problem.
Imad A. Ferzli, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 Voltage-Aware Static Timing Analysis
abstract
Static timing analysis (STA) techniques allow a designer to check the timing of a circuit at different process corners, which typically include corner values of the supply voltages as well. Traditionally, however, this analysis only considers cases where the supplies are either all low or all high. As will be demonstrated, this may not yield the true maximum delay of a circuit because it neglects the possible mismatch between the supplies of successive gates on a path. A new methodology for timing analysis is proposed, where, in a first step, the critical paths of a circuit are identified under an assumption that all the supply nodes are independent of one another, thus allowing for mismatch between the supplies. Then, given these critical paths, the authors incorporate into the analysis the relationships between the supply node voltages by considering the power grid that they are tied to, and refine the worst case time delay values on a per-critical-path basis. This refinement is posed as a sequence of optimization problems where the operation of the circuit is abstracted in terms of current constraints. The authors present their technique and report on the implementation results using benchmark circuits tied to a number of test-case power grids
Dionysios Kouroussis, Rubil Ahmadi, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Dynamic-range estimation
abstract
It has been widely recognized that the dynamic-range information of an application can be exploited to reduce the datapath bitwidth of either processors or application-specific integrated circuits and, therefore, the overall circuit area, delay, and power consumption. While recent proposals of analytical dynamic-range-estimation methods have shown significant advantages over the traditional profiling-based method in terms of runtime, it is argued here that the rather simplistic treatment of input correlation and system nonlinearity may lead to significant error. In this paper, three mathematical tools, namely Karhunen-Loe/spl grave/ve expansion, polynomial chaos expansion, and independent component analysis are introduced, which enable not only the orthogonal decomposition of input random processes, but also the propagation of random processes through both linear and nonlinear systems with difficult constructs such as multiplications, divisions, and conditionals. It is shown that when applied to interesting nonlinear applications such as adaptive filters, polynomial filters, and rational filters, this method can produce complete accurate statistics of each internal variable, thereby allowing the synthesis of bitwidth with the desired trade off between circuit performance and signal-to-noise ratio.
Jianwen Zhu, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 Leakage power: trends, analysis and avoidance
abstract
Leakage power is emerging as a key challenge in IC design. Leakage is increasingly exponentially with each technology generation and is expected to become the dominant part of total power. Device threshold voltage scaling, shrinking device dimensions, and larger circuit sizes are causing this dramatic increase in leakage. As leakage varies exponentially with process parameters, yield of the chip is often directly influenced by leakage. Increasing amount of leakage is also critical for power constraint ICs. Traditionally, leakage has been considered as an important design variable in handheld devices and in standby circuit operation. However, this significant increase of leakage now warrants that it be considered as the key design variable in all IC designs.This tutorial presents a comprehensive review of leakage power issues in IC design. The tutorial is organized in four major parts. The first part provides an overview of technology and scaling trends which are causing the significant increase in leakage current. The device physics that leads to sub-threshold and gate leakage will be described, along with their dependence on circuit design variables. This part of the tutorial will also cover basic transistor and circuit techniques to minimize leakage, such as the stack effect.The second part of the tutorial will focus on circuit level leakage estimation and avoidance. Use of multiple threshold voltages has been very successful in controlling the leakage of the circuit. Comprehensive description of multiple-Vt techniques for leakage avoidance will be presented along with associated leakage estimation techniques. Multiple-threshold design (MTCMOS) will be described along with its leakage benefits and performance trade-offs. Multiple oxide technology options and associated impact on gate leakage will also be discussed.Third part of the tutorial focuses on chip level effects on leakage. Leakage is heavily dependent on local and global process variations and can vary by an order of magnitude over the technology spread. Leakage estimation techniques which consider both inter and intra-die process variations will be covered. This part of the tutorial also focuses on chip-level leakage minimization techniques. Leakage minimization techniques such as Adaptive Body Bias (ABB) and power supply control will be presented.The last part of the tutorial covers system and circuit architectures for leakage avoidance. In standby mode, the leakage of the circuit can be lowered by putting it a low-leakage state. Caches and memory circuits occupy large percentage of area in model chips. The leakage of caches and memories need to be carefully controlled. This section of the tutorial will cover topics including state assignment for leakage minimization, leakage-driven memory and cache circuits and architectures.The tutorial is intended for designers and CAD engineers interested in next generation design techniques and methodologies and emerging power challenges. Basic background of VLSI and CAD is useful though not needed.
David T. Blaauw, Anirudh Devgan, Farid N. Najm
ASP-DAC3
2005 Variations-aware low-power design with voltage scaling
abstract
We present a new methodology which takes into consideration the effect of Within-Die (WID) process variations on a low-voltage parallel system. We show that in the presence of process variations one should use a higher supply voltage than would otherwise be predicted to minimize the power consumption of a parallel systems. Previous analyses, which ignored WID process variations, provide a lower non-optimal supply voltage which can underestimate the energy/ operation by 8.2X. We also present a novel technique to limit the effect of temperature variations in a parallel system. As temperatures increases, the scheme reduces the power increase by 43% allowing the system to remain at it's optimal supply voltage across different temperatures.
Navid Azizi, Muhammad M. Khellah, Vivek De, Farid N. Najm
DAC4
2005 On the need for statistical timing analysis
abstract
Traditional corner analysis fails to guarantee a target yield for a given performance metric. However, recently proposed solutions, in the form of statistical timing analysis, which work by propagating delay distributions, do not conform to modern design methodology. Instead, new statistical techniques are needed to modify corner analysis in ways that overcome its weaknesses without violating usage models of timing tools in modern flows.
Farid N. Najm
DAC1
2005 A non-parametric approach for dynamic range estimation of nonlinear systems
abstract
It has been widely recognized that the dynamic range information of an application can be exploited to reduce the datapath bitwidth of either processors or ASICs, and therefore the overall circuit area, delay, and power consumption. Recent advances in analytical dynamic range estimation methods indicate that by systematically decomposing the system inputs into orthonormal random variables using a mathematical procedure called polynomial chaos expansion (PCE), output statistics of interest can be obtained for both linear and nonlinear systems. Despite its power for capturing both spatial and temporal correlation, the application of this method has been limited only to near-Gaussian inputs. In this paper, we propose the first algorithm with the capacity of handling both near-Gaussian and non-Gaussian input signals. Our method is based on the use of independent component analysis (ICA). Our experiments show that the new algorithm can reduce the original relative errors of 2nd order moments from 25% - 65% to 1% - 2%.
Jianwen Zhu, Farid N. Najm
DAC3
2005 Statistical timing analysis with two-sided constraints
abstract
Based on a timing yield model, a statistical static timing analysis technique is proposed. This technique preserves existing methodology by selecting a "device file setting" that takes into account within-die statistical variations, and with which to run traditional static timing analysis in order to meet the desired yield. Using process-specific "generic paths" representing critical paths in a given process technology, our approach can be used early in the design process, most importantly during the pre-placement phase. Within-die variations are taken care of using a simple model that assumes positive correlation, which leads to upper and lower bounds on the timing yield. Our approach also handles both setup and hold timing constraints.
Khaled R. Heloue, Farid N. Najm
ICCAD2
2005 Incremental partitioning-based vectorless power grid verification
abstract
To ensure reliable performance of a chip, design verification of the power grid is of critical importance. This paper builds on previous work that models the working behavior of the circuit in terms of abstracted current constraints and solves for worst-case voltage drop on the grid as a linear program. The main motivation is to allow the efficient verification of local power grid sections or blocks, enabling incremental design analysis of the grid. This approach substantially improves the computational time by reducing the problem size and the constraint set and replacing them by black box macromodels. This increase the capacity of the solver to handle industrial sized grids.
Dionysios Kouroussis, Imad A. Ferzli, Farid N. Najm
ICCAD3
2005 Power grid voltage integrity verification
abstract
Full-chip verification requires one to check if the power grid is safe, i.e., if the voltage drop on the grid does not exceed a certain threshold. The traditional simulation-based solution to this problem is computationally expensive, because of the large variety of possible circuit behaviors that would need to be simulated; it also has the disadvantage that it requires full knowledge of the details of the circuit attached to the grid, thereby precluding early verification of the grid. We propose a power grid verification technique that can be applied before the complete circuit has been designed and without exact knowledge of the circuit currents. We use current constraints, which are upper bound constraints on the currents that can be drawn from the grid, as a way to capture the uncertainty about the circuit details and activity. Based on this, we propose two solution approaches. One approach gives an upper-bound on the worst-case voltage drop at every node of the grid. Another, less expensive approach, applies a sufficient condition (thus, this becomes a conservative approach) to check if the drop on the grid exceeds a given voltage threshold
Maha Nizam, Farid N. Najm, Anirudh Devgan
ISLPED2
2005 Early power estimation for VLSI circuits
abstract
Early power estimation, a requirement for design exploration early in the design phase, must often be done based on a design specification that is available only at a high level of abstraction. One way of doing this is to use high-level estimation of circuit total capacitance and average activity. This paper addresses these problems and proposes a high-level area estimation technique based on the complexity of a Boolean network representation of the design. In addition to the high-level area estimation, the paper also proposes a high-level activity estimation methodology that is capable of handling correlated input streams. High-level power estimates based on the total capacitance and average activity estimates are also given.
Kavel M. Büyüksahin, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 A Case for Asymmetric-Cell Cache Memories
abstract
In this paper, we make the case for building high-performance asymmetric-cell caches (ACCs) that employ recently-proposed asymmetric SRAMs to reduce leakage proportionally to the number of resident zero bits. Because ACCs target memory value content (independent of cell activity and access patterns), they complement prior proposals for reducing cache leakage that target memory access characteristics. Through detailed simulation and leakage estimation using a commercial 0.13-/spl mu/m CMOS process model, we show that: 1) on average 75% of resident data cache bits and 64% of resident instruction cache bits are zero; 2) while prior research carefully evaluated the fraction of accessed zero bytes, we show that a high fraction of accessed zero bytes is neither a necessary nor a sufficient condition for a high fraction of resident zero bits; 3) the zero-bit program behavior persists even when we restrict our attention to live data, thereby complementing prior leakage-saving techniques that target inactive cells; and 4) ACCs can reduce leakage on the average by 4.3/spl times/ compared to a conventional data cache without any performance loss, and by 9/spl times/ at the cost of a 5% increase in overall cache access latency.
Andreas Moshovos, Babak Falsafi, Farid N. Najm, Navid Azizi
IEEE Trans. Very Large Scale Integr. Syst.3
2004 Interconnect capacitance estimation for FPGAs
Jason Helge Anderson, Farid N. Najm
ASP-DAC2
2004 Worst-case circuit delay taking into account power supply variations
abstract
Current Static Timing Analysis (STA) techniques allow one to verify the timing of a circuit at different process corners which only consider cases where all the supplies are low or high. This analysis may not give the true maximum delay of a circuit because it neglects the possible mismatch between drivers and loads. We propose a new approach for timing analysis in which we first identify the critical path(s) of a circuit using a power-supply-aware timing model. Given these critical paths, we then take into account how the power nodes of the gates on the critical path are connected to the power grid, and re-analyze for the worst-case time delay. This re-analysis is posed as an optimization problem where the complete operation of the entire circuit is abstracted in terms of current constraints. We present our technique and report on the implementation results using benchmark circuits tied to a number of test-case power grids.
Dionysios Kouroussis, Rubil Ahmadi, Farid N. Najm
DAC3
2004 Statistical timing analysis based on a timing yield model
abstract
Starting from a model of the within-die systematic variations using principal components analysis, a model is proposed for estimation of the parametric yield, and is then applied to estimation of the timing yield. Key features of these models are that they are easy to compute, they include a powerful model of within-die correlation, and they are "full-chip" models in the sense that they can be applied with ease to circuits with millions of components. As such, these models provide a way to do statistical timing analysis without the need for detailed statistical analysis of every path in the design.
Farid N. Najm, Noel Menezes
DAC1
2004 An analytical approach for dynamic range estimation
abstract
It has been widely recognized that the dynamic range information of an application can be exploited to reduce the datapath bitwidth of either processors or ASICs, and therefore the overall circuit area, delay and power consumption. While recent proposals of analytical dynamic range estimation methods have shown significant advantages over the traditional profiling-based method in terms of runtime, we argue that the rather simplistic treatment of input correlation may lead to significant error. We instead introduce a new analytical method based on a mathematical tool called Karhunen-Loeve Expansion (KLE), which enables the orthogonal decomposition of random processes. We show that when applied to linear systems, this method can not only lead to much more accurate result than previously possible, thanks to its capability to capture and propagate both spatial and temporal correlation, but also richer information than the value bounds previously produced, which enables the exploration of interesting trade-off between circuit performance and signal-to-noise ratio.
Jianwen Zhu, Farid N. Najm
DAC3
2004 Active leakage power optimization for FPGAs
abstract
We consider active leakage power dissipation in FPGAs and present a "no cost" approach for active leakage reduction. It is well-known that the leakage power consumed by a digital CMOS circuit depends strongly on the state of its inputs. Our leakage reduction technique leverages a fundamental property of basic FPGA logic elements (look-up-tables) that allows a logic signal in an FPGA design to be interchanged with its complemented form without any area or delay penalty. We apply this property to select polarities for logic signals so that FPGA hardware structures spend the majority of time in low leakage states. In an experimental study, we optimize active leakage power in circuits mapped into a state-of-the-art 90nm commercial FPGA. Results show that the proposed approach reduces active leakage by 25%, on average.
Jason Helge Anderson, Farid N. Najm, Tim Tuan
FPGA2
2004 Low-power programmable routing circuitry for FPGAs
abstract
We propose two new FPGA routing switch designs that are programmable to operate in three different modes: high-speed, low-power or sleep. High-speed mode provides similar power and performance to a traditional routing switch. In low-power mode, speed is curtailed in order to reduce power consumption. Our first switch design reduces leakage power consumption by 36-40% in low-power vs. high-speed mode (on average); dynamic power is reduced by up to 28%. Leakage power in sleep mode is 61% lower than in high-speed mode. A second switch design offers a 36% smaller area overhead and reduces leakage by 28-30% in low-power vs. high-speed mode. The proposed switch designs require only minor changes to a traditional routing switch, making them easy to incorporate into current FPGA interconnect. The applicability of the new switches is motivated through an analysis of timing slack in industrial FPGA designs. Specifically, we show that a considerable fraction of routing switches may be slowed down (operate in low-power mode), without impacting overall design performance.
Jason Helge Anderson, Farid N. Najm
ICCAD2
2004 Dynamic range estimation for nonlinear systems
abstract
It has been widely recognized that the dynamic range information of an application can be exploited to reduce the datapath bitwidth of either processors or ASICs, and therefore the overall circuit area, delay and power consumption. While recent advances in analytical dynamic range estimation can deliver results accurate enough to account for both spatial and temporal correlation, the reported methods are only valid for linear systems. In this paper, we use a powerful mathematical tool, called polynomial chaos, which enables not only the orthogonal decomposition of random processes, but also the propagation of random processes through nonlinear systems with difficult constructs such as multiplications, divisions and conditionals. We show that when applied to interesting nonlinear applications such as adaptive filters, polynomial filters and rational filters, this method can produce complete, accurate statistics of each internal variable, thereby allowing the synthesis of bitwidth with the desired tradeoff between circuit performance and signal-to-noise ratio.
Jianwen Zhu, Farid N. Najm
ICCAD3
2004 Power estimation techniques for FPGAs
abstract
The dynamic power consumed by a digital CMOS circuit is directly proportional to both switching activity and interconnect capacitance. In this paper, we consider early prediction of net activity and interconnect capacitance in field-programmable gate array (FPGA) designs. We develop empirical prediction models for these parameters, suitable for use in power-aware layout synthesis, early power estimation/planning, and other applications. We examine how switching activity on a net changes when delays are zero (zero delay activity) versus when logic delays are considered (logic delay activity) versus when both logic and routing delays are considered (routed delay activity). We then describe a novel approach for prelayout activity prediction that estimates a net's routed delay activity using only zero or logic delay activity values, along with structural and functional circuit properties. For capacitance prediction, we show that prediction accuracy is improved by considering aspects of the FPGA interconnect architecture in addition to generic parameters, such as net fanout and bounding box perimeter length. We also demonstrate that there is an inherent variability (noise) in the switching activity and capacitance of nets that limits the accuracy attainable in prediction. Experimental results show the proposed prediction models work well given the noise limitations.
Jason Helge Anderson, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Statistical estimation of leakage-induced power grid voltage drop considering within-die process variations
abstract
Transistor threshold voltages Vth have been reduced as part of on-going technology scaling. The smaller Vth values feature increased fluctuations due to process variations, with a strong within-die component. Correspondingly, given the exponential dependence of leakage on.Vth, circuit leakage currents are increasing significantly and have strong within-die statistical variations. With these currents loading the power grid, the grid develops large voltage drops, which is an unavoidable background level of noise on the grid. We develop techniques for estimation of the statistics of the leakage-induced power grid voltage drop based on given statistics of the circuit leakage currents.
Imad A. Ferzli, Farid N. Najm
DAC2
2003 A static pattern-independent technique for power grid voltage integrity verification
abstract
Design verification must include the power grid. Checking that the voltage on the power grid does not drop by more than some critical threshold is a very difficult problem, for at least two reasons: i) the obviously large size of the power grids for modern high-performance chips, and ii) the difficulty of setting up the right simulation conditions for the power grid that provide some measure of a realistic worst case voltage drop. The huge number of possible circuit operational modes or workloads makes it impossible to do exhaustive analysis. We propose a static technique for power grid verification, where static is in the sense of static timing analysis, meaning that it does not depend on, nor require, user-specified stimulus to drive a simulation. The verification is posed as an optimization problem under user-supplied current constraints. We propose that current constraints are the right kind of abstraction to use in order to develop a practical methodology for power grid verification. We present our verification approach, and report on the results of applying it to a number of test-case power grids.
Dionysios Kouroussis, Farid N. Najm
DAC2
2003 Timing Analysis in Presence of Power Supply and Ground Voltage Variations
Rubil Ahmadi, Farid N. Najm
ICCAD2
2003 Statistical Verification of Power Grids Considering Process-Induced Leakage Current Variations
Imad A. Ferzli, Farid N. Najm
ICCAD2
2003 ESTIMA: an architectural-level power estimator for multi-ported pipelined register files
abstract
We introduce an architectural-level power, area, and latency estimator for multi-ported, pipelined register files. Strengths of the proposed approach include the handling of pipelined operation and clock power, the simulation-based device size estimation, and the ability to handle user-specified timing constraints. The model proposed can be used as a stand-alone estimation and design exploration tool for register files and register-file type structures, or it can be incorporated into a high-level performance simulator to add power estimation capabilities.
Kavel M. Büyüksahin, Priyadarsan Patra, Farid N. Najm
ISLPED3
2003 Low-leakage asymmetric-cell SRAM
abstract
We introduce a novel family of asymmetric dual-V/sub t/ static random access memory cell designs that reduce leakage power in caches while maintaining low access latency. Our designs exploit the strong bias toward zero at the bit level exhibited by the memory value stream of ordinary programs. Compared to conventional symmetric high-performance cells, our cells offer significant leakage reduction in the zero state and, in some cases, also in the one state, albeit to a lesser extent. A novel sense amplifier, in combination with dummy bitlines, allows for read times to be on par with conventional symmetric cells. With one cell design, leakage is reduced by 7/spl times/ (in the zero state) with no performance degradation, but with a stability degradation of 6%. Another cell design reduces leakage by 2/spl times/ (in the zero state) with no performance or stability loss. An alternative cell design reduces leakage by 58/spl times/ (in the zero state) with a performance degradation of 1% and an area increase of 2.4% and no stability degradation.
Navid Azizi, Farid N. Najm, Andreas Moshovos
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Energy and peak-current per-cycle estimation at RTL
abstract
We present novel macromodeling techniques for estimating the energy dissipated and peak-current drawn in a logic circuit for every input vector pair (we call this the energy-per-cycle and peak-current-per-cycle, respectively). The macromodels are based on classifying the input vector pairs on the basis of their Hamming distances and using a different equation-based macromodel for every Hamming distance. The variables of our macromodel are the zero-delay transition counts at three logic levels inside the circuit. We present an automatic characterization process by which such macromodels can be constructed. The energy-per-cycle macromodel provides a transient energy waveform, and can also be used to estimate the moving average energy over any time window, whereas peak-current-per-cycle macromodel provides peak-current which can be used for studying IR drop problems. Some key features of this technique are: 1) the models are compact (linear in the number of inputs); 2) they can be used for any input sequence; and 3) the characterization is automatic and requires no user intervention. These approaches have been implemented and models have been built and tested for many circuits. The average errors observed in estimating the energy-per-cycle and peak-current-per-cycle are under 20%. The energy-per-cycle model can also be used to measure the long-term average power, with an observed error of under 10% on average.
Subodh Gupta, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
2002 High-level current macro-model for power-grid analysis
abstract
We present a frequency domain current macro-modeling technique for capturing the dependence of the block current waveform on its input vectors. The macro-model is based on estimating the Discrete Cosine Transform (DCT) of the current waveform as a function of input vector pair and then taking the inverse transform to estimate the time domain current waveform. The input vector pairs are partitioned according to Hamming distance and a current macro-model is built for each Hamming distance using regression. Regression is done on a set of current waveforms generated for each circuit, using HSPICE. The average relative error in peak current estimation using the current macro-model is less than 20%.
Srinivas Bodapati, Farid N. Najm
DAC2
2002 Power-aware technology mapping for LUT-based FPGAs
abstract
We present a new power-aware technology mapping technique for LUT-based FPGAs which aims to keep nets with high switching activity out of the FPGA routing network and takes an activity-conscious approach to logic replication. Logic replication is known to be crucial for optimizing depth in technology mapping; an important contribution of our work is to recognize the effect of logic replication on circuit structure and to show its consequences on power. In an experimental study, we examine the power characteristics of mapping solutions generated by several publicly available technology mappers. Results show that for a specific depth of mapping solution, the power consumption can vary considerably, depending on the technology mapping approach used. Furthermore, results show that our proposed mapping algorithm leads to circuits with substantially less power dissipation than previous approaches.
Jason Helge Anderson, Farid N. Najm
FPT2
2002 Low-leakage asymmetric-cell SRAM
abstract
We introduce a novel family of asymmetric dual-Vt SRAM cell designs that reduce leakage power in caches while maintaining low access latency. Our designs exploit the strong bias towards zero at the bit level exhibited by the memory value stream of ordinary programs. Compared to conventional symmetric high-performance cells, our cells offer significant leakage reduction in the zero state and in some cases also in the one state albeit to a lesser extend. A novel sense-amplifier, in coordination with dummy bitlines, allows for read times to be on par with conventional symmetric cells. With one cell design, leakage is reduced by 7X (in the zero state) with no performance degradation. An alternative cell design reduces leakage by 40X (in the zero state) with a performance degradation of 5%.
Navid Azizi, Andreas Moshovos, Farid N. Najm
ISLPED3
2002 High-level area estimation
abstract
Early power estimation requires one to estimate the area (gate count) of a design from a high-level description. We propose a method to do this that makes use of the concept of Boolean networks (BN) and introduces an invariant area complexity measure which captures the gate-count requirement of a design. The method can be adapted to be used at different points on the area/delay tradeoff curve, with different synthesizer/mapper tools, and different target gate libraries. The area model is experimentally verified and tested using a number of ISCAS and MCNC benchmark circuits and two different target cell libraries, on two different synthesis systems.
Kavel M. Büyüksahin, Farid N. Najm
ISLPED2
2002 A multigrid-like technique for power grid analysis
abstract
Modern submicron very large scale integration designs include huge power grids that are required to distribute large amounts of current, at increasingly lower voltages. The resulting voltage drop on the grid reduces noise margin and increases gate delay, resulting in a serious performance impact. Checking the integrity of the supply voltage using traditional circuit simulation is not practical, for reasons of time and memory complexity. The authors propose a novel multigrid-like technique for the analysis of power grids. The grid is reduced to a coarser structure, and the solution is mapped back to the original grid. Experimental results show that the proposed method is very efficient as well as suitable for both de and transient analysis of power grids.
Joseph N. Kozhaya, Sani R. Nassif, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2002 Estimation of state line statistics in sequential circuits
abstract
In this article, we present a simulation-based technique for estimation of signal statistics (switching activity and signal probability) at the flip-flop output nodes (state signals) of a general sequential circuit. Apart from providing an estimate of the power consumed by the flip-flops, this information is needed for calculating power in the combinational portion of the circuit. The statistics are computed by collecting samples obtained from fast RTL simulation of the circuit under input sequences that are either randomly generated or independently selected from user-specified pattern sets. An important advantage of this approach is that the desired accuracy can be specified up front by the user; with some approximation, the algorithm iterates until the specified accuracy is achieved. This approach has been implemented and tested on a number of sequential circuits and has been shown to handle very large sequential circuits that can not be handled by other existing methods, while using a reasonable amount of CPU time and memory (the circuit s38584.1, with 1426 flip-flops, can be analyzed in about 10 minutes).
Vikram Saxena, Farid N. Najm, Ibrahim N. Hajj
ACM Trans. Design Autom. Electr. Syst.2
2002 A technique for Improving dual-output domino logic
abstract
We present a technique, termed clock-generating (CG) domino, for improving dual-output domino logic that reduces area, clock load and power without increasing the delay. A delayed clock, generated from certain dual-output gates, is used to convert other dual-output gates to single output. Simulation results with ISCAS 85 benchmark circuits indicate an average reduction in area, clock load, and power of 17%, 20%, and 24%, respectively, over dual-output domino and a 48% power reduction for the largest circuit.
Sumant Ramprasad, Ibrahim N. Hajj, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.3
2001 Multigrid-Like Technique for Power Grid Analysis
abstract
Modern sub-micron VLSI designs include huge power grids that are required to distribute large amounts of current, at ever lower voltages. The resulting voltage drop on the grid reduces noise margin and increases gate delay, resulting in a serious performance impact. Checking the integrity of the supply voltage using traditional circuit simulation is not practical, for reasons of time and memory complexity. We propose a novel multigrid-like technique for the analysis of power grids. The grid is reduced to a coarser structure, and the solution is mapped back to the original grid. Experimental results show that the proposed method is very efficient as well as suitable for both DC and transient analysis of power grids.
Joseph N. Kozhaya, Sani R. Nassif, Farid N. Najm
ICCAD3
2001 Frequency-domain supply current macro-model
abstract
In order to perform block level analysis of the on-chip power distribution network, a high-level model is required that captures the dependence of the current waveform drawn by a logic block, per cycle, on its input vector pair. We present a frequency domain macro-modeling technique for capturing this dependence. The macro-model is based on estimating the Discrete Cosine Transform (DCT) of the current waveform and then taking the inverse transform to estimate the time domain current waveform.
Srinivas Bodapati, Farid N. Najm
ISLPED2
2001 Prelayout estimation of individual wire lengths
abstract
We present a novel technique for estimating individual wire lengths in a given standard-cell-based design during the technology mapping phase of logic synthesis. The proposed method is based on creating a black box model of the place and route tool as a function of a number of parameters, which are all available before layout. The place and route tool is characterized, only once, by applying it to a set of typical designs in a certain technology. We also propose a net bounding box estimation technique based on the layout style and net neighborhood analysis. We show that there is inherent variability in wire lengths obtained using commercially available place and route tools-wire length estimation error cannot be any smaller than a lower limit due to this variability. The proposed model works well within these variability limitations.
Srinivas Bodapati, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
2001 Power estimation for large sequential circuits
abstract
A power estimation approach is presented in which blocks of consecutive vectors are selected at random from a user-supplied realistic input vector set and the circuit is simulated for each block starting from an unknown state. This leads to two (upper and lower) bounds on the desired power value which can be quite tight (under 10% difference between the two in many cases). As a result, the power dissipation is obtained by simulating only a fraction of the potentially very large vector set.
Joseph N. Kozhaya, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
2000 High-level power estimation with interconnect effects
abstract
We extend earlier work on high-level average power estimation to include the power due to interconnect loading. The resulting technique is a combination of a RTL-level gate count prediction method and average interconnect estimation based on Rent's rule. The method can be adapted to be used with different place and route engines and standard cell libraries. For a number of benchmark circuits, the method is verified by extracting wire lengths from a layout of each circuit and then comparing the predicted (at RTL) power against that measured using SPICE. An average error of 14.4% is obtained for the average interconnect length, and an average error of 25.8% is obtained for average power estimation including interconnect effects.
Kavel M. Büyüksahin, Farid N. Najm
ISLPED2
2000 Analytical models for RTL power estimation of combinational andsequential circuits
abstract
In this paper, we propose a modeling technique that captures the dependence of the power dissipation of a (combinational or sequential) logic circuit on its input/output signal switching statistics. The resulting power macromodel consists of a quadratic or cubic equation in four variables, that can be used to estimate the power consumed in the circuit for any given input/output signal statistics. Given a low-level (typically gate-level) description of the circuit, we describe a characterization process that uses a recursive least squares (RLS) algorithm by which such an equation-based model can be automatically built. This approach has been implemented and models have been built and tested for many combinational and sequential benchmark circuits.
Subodh Gupta, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 Power modeling for high-level power estimation
abstract
In this paper, we propose a modeling approach that captures the dependence of the power dissipation of a combinational logic circuit on its input/output signal switching statistics. The resulting power macromodel, consisting of a single four-dimensional table, can be used to estimate the power consumed in the circuit for any given input/output signal statistics. Given a low-level (typically gate-level) description of the circuit, we describe a characterization process by which such a table model can be automatically built. The four dimensions of our table-based model are the average input signal probability, average input transition density, average spatial correlation coefficient, and average output zero-delay transition density. This approach has been implemented and models have been built for many benchmark circuits. Over a wide range of input signal statistics, we show that this model gives very good accuracy, with an rms error of about 4% and average error of about 6%. Except for one out of about 10 000 cases, the largest error observed was under 20%. If one ignores the glitching activity, then the rms error becomes under 1%, the average error becomes under 5%, and the largest error observed in all cases is under 18%.
Subodh Gupta, Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.2
1999 Power macro-models for DSP blocks with application to high-level synthesis
abstract
In this paper, we propose a modeling approach for the average power consumption of macro-blocks that are typically used in digital signal processing (DSP) systems, such as adders, multipliers and delay elements, in terms of their input/output signal switching statistics.The resulting power mucromodel, consisting of a quadratic or cubic equation in four variables, can be used to estimate the average power consumed in the macro-block for any given input/output signal statistics.This enables highlevel power estimation and allows one to compare the power performance of different competing DSP systems during high-level synthesis.This approach has been implemented and models have been built and tested for many macro-blocks.
Subodh Gupta, Farid N. Najm
ISLPED2
1999 Energy-per-cycle estimation at RTL
abstract
Abstract – We present a novel macromodeling technique for estimating the energy dissipated in a logic circuit for every input vector pair (we call this the energy-per-cycle). The macromodel is based on classifying the input vector pairs on the basis of their Hamming distances and using a different equation-based macromodel for every Hamming distance. The variables of our macromodel are the zero-delay transition counts at three logic levels inside the circuit. We present an automatic characterization process by which such macromodels can be constructed. This energy-per-cycle macromodel provides a transient energy waveform, and can also be used to estimate the moving average energy over any time window. This approach has been implemented and models have been built and tested for many circuits. The average error observed in estimating the energy-per-cycle is under 20%. The model can also be used to measure the long-term average power, with an observed error of under 10% on average. 1.
Subodh Gupta, Farid N. Najm
ISLPED2
1999 An optimization technique for dual-output domino logic
abstract
Dynamic logic circuits [2] are used in high-performance circuits due to their speed and area advantage over static CMOS circuits. One well-known dynamic logic family is the domino CMOS family, which, however, su ers from its inability to
Sumant Ramprasad, Ibrahim N. Hajj, Farid N. Najm
ISLPED3
1999 High-level area and power estimation for VLSI circuits
abstract
High-level power estimation, when given only a high-level design specification such as a functional or register-transfer level (RTL) description, requires high-level estimation of the circuit average activity and total capacitance. Considering that total capacitance is related to circuit area, this paper addresses the problem of computing the "area complexity" of multi-output combinational logic given only their functional description, i.e., Boolean equations, where area complexity refers to the number of gates required for an optimal multilevel implementation of the combinational logic. The proposed area model is based on transforming the multi-output Boolean function description into an equivalent single output function. The area model is empirical and results demonstrating its feasibility and utility are presented. Also, a methodology for converting the gate count estimates, obtained from the area model, into capacitance estimates is presented. High-level power estimates based on the total capacitance estimates and average activity estimates are also presented.
Mahadevamurty Nemani, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 Delay Estimation VLSI Circuits from a High-Level View
abstract
Estimation of the delay of a Boolean function from its functional description is an important step towards design exploration at the register transfer level (RTL). This paper addresses the problem of estimating the delay of certain optimal multi-level implementations of combinational circuits, given only their functional description. The proposed delay model uses a new complexity measure called the delay measure to estimate the delay. It has an advantage that it can be used to predict both, the minimum delay (associated with an optimum delay implementation) and the maximum delay (associated with an optimum area implementation) of a Boolean function without actually resorting to logic synthesis. The model is empirical and results demonstrating its feasibility and utility are presented. 1. Introduction Rapid increase in the design complexity and reduction in design time have resulted in a need for CAD tools that can help make important design decisions early in the design process. To do...
Mahadevamurty Nemani, Farid N. Najm
DAC2
1997 Power Macromodeling for High Level Power Estimation
abstract
A modeling approach is presentedthat captures the dependence of the power dissipationof a combinational logic circuit on its input/outputsignal switching activity.The resultingpower macromodel, consisting of a single three dimensionaltable, can be used to estimate the powerconsumed in the circuit for any given input/outputsignal statistics.Given a low-level (typically gate-level)description of the circuit, we describe a characterizationprocess by which such a table modelcan be automatically built.In contrast to otherproposed techniques, this can be done for any givenlogic circuit without any user intervention, and appliesto all possible input/output signal statistics;it does not require one to construct specialized analyticalequations for the power dissipation.Thethree dimensions of our table-based model are theaverage input signal probability, average input transitiondensity, and average output zero-delay transition density.This approach has been implemented and modelshave been built for many benchmark circuits.Overa wide range of input signal statistics, we show thatthis model gives very good accuracy, with an RMSerror of under about 6%.
Subodh Gupta, Farid N. Najm
DAC2
1997 Technology-Dependent Transformations for Low-Power Synthesis
abstract
We propose a methodology for applying gate-level logictransformations to optimize power in digital circuits.Statisticallysimulated switching information, gate delays,signal arrival patterns, and signal probabilities are consideredin reducing the switching activity-capacitance products.Power reduction up to 45.4% (acerage 12.4%) is achieved,with considerable improvements in area and delay, in pre-optimizedbenchamarks.Also the effect of transformationson the random pattern testability of the circuits is studied.
Rajendran Panda, Farid N. Najm
DAC2
1997 Accurate power estimation for large sequential circuits
abstract
A power estimation approach is presented in which blocks of consecutive vectors are selected at random from a user-supplied realistic input vector set and the circuit is simulated for each block starting from an unknown state. This leads to two (upper and lower) bounds on the desired power value which can be quite tight (under 10% difference between the two in many cases). As a result, the power dissipation is obtained by simulating only a fraction of the potentially very large vector set.
Joseph N. Kozhaya, Farid N. Najm
ICCAD2
1997 High-level area and power estimation for VLSI circuits
abstract
This paper addresses the problem of computing the area complexity of a multi-output combinational logic circuit, given only its functional description, i.e., Boolean equations, where area complexity is measured in terms of the number of gates required for an optimal multilevel implementation of the combinational logic. The proposed area model is based on transforming the given, multi-output Boolean function description into an equivalent single-output function. The model, is empirical, and results demonstrating its feasibility and utility are presented. Also, a methodology for converting the gate count estimates, obtained from the area model, into capacitance estimates is presented. High-level power estimates based on the total capacitance estimates and average activity estimates are also presented.
Mahadevamurty Nemani, Farid N. Najm
ICCAD2
1996 High-level power estimation and the area complexity of Boolean functions
abstract
Estimation of the area complexity of a Boolean function from its functional description is an important step towards a power estimation capability at the register transfer level (RTL). This paper addresses the problem of computing the area complexity of single-output Boolean functions given only their functional description, where area complexity is measured in terms of the number of gates required for an optimal implementation of the function. We propose an area model to estimate the area based on a new complexity measure called the average cube complexity. This model has been implemented, and empirical results demonstrating its feasibility and utility are presented.
Mahadevamurty Nemani, Farid N. Najm
ISLPED2
1996 Towards a high-level power estimation capability [digital ICs]
abstract
We present a power estimation technique for digital integrated circuits that operates at the register transfer level (RTL). Such a high-level power estimation capability Is required in order to provide early warning of any power problems before the circuit-level design has been specified. With such early warning, the designer can explore design trade-offs at a higher level of abstraction than previously possible, reducing design time and cost. Our estimator is based on the use of entropy as a measure of the average activity to be expected in the final implementation of a circuit, given only its Boolean functional description. This technique has been implemented and tested on a variety of circuits. The empirical results to be presented are very promising and demonstrate the feasibility and utility of this approach.
Mahadevamurty Nemani, Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1995 Feedback, Correlation, and Delay Concerns in the Power Estimation of VLSI Circuits
abstract
With the advent of portable and high-density microelectronic devices, the power dissipation of integrated circuits has become a critical concern.Accurate and ecient p o w er estimation during the design phase is required in order to meet the power speci cations without a costly redesign process.As an introduction to the other papers in this session, this paper gives a tutorial presentation of the issues involved in power estimation.
Farid N. Najm
DAC1
1995 Power Estimation in Sequential Circuits
abstract
A new method for power estimation in sequential circuits is presented that is based on a statistical estimation technique. By applying randomly generated input sequences to the circuit, statistics on the latch outputs are collected, by simulation, that allow efficient power estimation for the whole design. An important advantage of this approach is that the desired accuracy can be specified up-front by the user; the algorithm iterates until the specified accuracy is achieved. This has been implemented and tested on a number of sequential circuits and found to be much faster than existing techniques. We can complete the analysis of a circuit with 1,452 flip-flops and 19,253 gates in about 4.6 hours (the largest test case reported previously has 223 flip-flops). I. INTRODUCTION The dramatic decrease in feature size and the corresponding increase in the number of devices on a chip, combined with the growing demand for portable communication and computing systems, have made power consump...
Farid N. Najm, Shashank Goel, Ibrahim N. Hajj
DAC1
1995 Extreme Delay Sensitivity and the Worst-Case Switching Activity in VLSI Circuits
abstract
We observe that the switching activity at a circuit node, also called the transition density, can be extremely sensitive to the circuit internal delays.As a result, slight delay v ariations can lead to several orders of magnitude changes in the node activity.This has important implications for CAD in that, if the transition density is estimated by simulation, then minor inaccuracies in the timing models can lead to very large errors in the estimated activity.As a solution, we propose an efcient technique for estimating an upper bound on the transition density a t e v ery node.While it is not always very tight, the upper bound is robust, i n the sense that it is valid irrespective of delay v ariations and modeling errors.We will describe the technique and present experimental results based on a prototype implementation.
Farid N. Najm, Michael Y. Zhang
DAC1
1995 Power estimation techniques for integrated circuits
abstract
With the advent of portable and high-density microelectronic devices, the power dissipation of very large scale integrated (VLSI) circuits is becoming a critical concern. Accurate and efficient power estimation during the design phase is required in order to meet the power specifications without a costly redesign process. Recently, a variety of power estimation techniques have been proposed, most of which are based on: (1) the use of simplified delay models, and (2) modeling long-term behavior of logic signals with probabilities. The array of available techniques differ in subtle ways in the assumptions that they make, the accuracy that they provide, and the kinds of circuits that they apply to. In this tutorial, I will survey the many power estimation techniques that have been recently proposed and, in an attempt to make sense of all the variety, I will try to explain the different assumptions on which these techniques are based, and the impact of these assumptions on their accuracy and speed.
Farid N. Najm
ICCAD1
1995 Pattern independent maximum current estimation in power and ground buses of CMOS VLSI circuits: Algorithms, signal correlations, and their resolution
abstract
Currents flowing in the power and ground (P&G) buses of CMOS digital circuits affect both circuit reliability and performance by causing excessive voltage drops. Excessive voltage drops manifest themselves as glitches on the P&G buses and cause erroneous logic signals and degradation in switching speeds. Maximum current estimates are needed at every contact point in the buses to study the severity of the voltage drop problems and to redesign the supply lines accordingly. These currents, however, depend on the specific input patterns that are applied to the circuit. Since it is prohibitively expensive to enumerate all possible input patterns, this problem has, for a long time, remained largely unsolved. In this paper, we propose a pattern-independent, linear time algorithm (iMax) that estimates at every contact point, an upper bound envelope of all possible current waveforms that result by the application of different input patterns to the circuit. The algorithm is extremely efficient and produces good results for most circuits as is demonstrated by experimental results on several benchmark circuits. The accuracy of the algorithm can be further improved by resolving the signal correlations that exist inside a circuit. We also present a novel partial input enumeration (PIE) technique to resolve signal correlations and significantly improve the upper bounds for circuits where the bounds produced by iMax are not tight. We establish with extensive experimental results that these algorithms represent a good time-accuracy trade-off and are applicable to VLSI circuits.>
Harish Kriplani, Farid N. Najm, Ibrahim N. Hajj
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 Statistical Estimation of the Switching Activity in Digital Circuits
abstract
Abstract{Higher levels of integration have led to a generation of integrated circuits for which power dissipation and reliability are major design concerns. In CMOS circuits, both of these problems are directly related to the extent of circuit switching activity. The average number of transitions per second at a circuit node is a measure of switching activity that has been called the transition density. This paper presents a statistical simulation technique to estimate individual node transition densities. The strength of this approach is that the desired accuracy and con dence can be speci ed up-front by the user. Another key feature is the classi cation of nodes into two categories: regular- and low-density nodes. Regulardensity nodes are certi ed with user-speci ed percentage error and con dence levels. Low-density nodes are certi ed with an absolute error, with the same con dence. This speeds convergence while sacri cing percentage accuracy only on nodes which contribute little to power dissipation and have few reliability problems. I.
Michael G. Xakellis, Farid N. Najm
DAC2
1994 Improved Delay and Current Models for Estimating Maximum Currents in CMOS VLSI Circuits
abstract
Excessive voltage drops in power and ground (P&G) buses of CMOS VLSI circuits can severely degrade both design reliability and performance. Maximum current estimates are needed in the circuit to accurately determine the impact of these problems. In a previous paper by the authors (see Design Automation Conf., p. 2-7, June 8-12, 1992), a pattern-independent, linear time algorithm (iMax) is described that is very effective in estimating the maximum current waveforms at various contact points in the circuit. In the aforementioned paper, the algorithm was demonstrated for simple gate delay and current models. In this paper, we first derive expressions for modeling delays and current waveforms for a general gate and then describe how the algorithm can be extended under more general models.>
Harish Kriplani, Farid N. Najm, Ibrahim N. Hajj
ISCAS2
1994 Low-pass filter for computing the transition density in digital circuits
abstract
Estimating the power dissipation and the reliability of integrated circuits is a major concern of the semiconductor industry. Previously, we showed that a good measure of power dissipation and reliability is the extent of circuit switching activity, called the transition density (see ibid., vol. 12, no. 2, p. 310-23, 1993). However, the algorithm for computing the density in the afore-mentioned paper is very basic and does not take into account the effect of inertial delays of logic gates. Thus, as we will show in this paper, the transition density may be severely overestimated in high-frequency applications. To overcome this problem, we model the effect of gate delay on logic signals in the form of a conceptual low-pass filter module that does not allow unacceptably short logic pulses to propagate. Using a stochastic model of logic signals, we then derive the equations required to propagate the transition density through the filter. We will present experimental results that illustrate the validity and importance of these results.>
Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1994 A survey of power estimation techniques in VLSI circuits
abstract
With the advent of portable and high-density microelectronic devices, the power dissipation of very large scale integrated (VLSI) circuits is becoming a critical concern. Accurate and efficient power estimation during the design phase is required in order to meet the power specifications without a costly redesign process. In this paper, we present a review of the power estimation techniques that have recently been proposed.>
Farid N. Najm
IEEE Trans. Very Large Scale Integr. Syst.1
1993 Resolving Signal Correlations for Estimating Maximum Currents in CMOS Combinational Circuits
abstract
: Currents flowing in the power and ground (P&G) lines of CMOS digital circuits affect both circuit reliability and performance by causing excessive voltage drops. Maximum current estimates are therefore needed in the P&G lines to determine the severity of the voltage drop problems and to properly design the supply lines to eliminate these problems. These currents, however, depend on the specific input patterns that are applied to the circuit. Since it is prohibitively expensive to enumerate all possible inputs, this problem has, for a long time, remained largely unsolved. In [1], we proposed a pattern-independent, linear time algorithm (iMax) that estimates an upper bound envelope of all possible current waveforms that result from the application of different input patterns to the circuit. While the bound produced by iMax is fairly tight on many circuits, there can be a significant loss in accuracy due to correlations between signals internal to the circuit. In this paper, we present ...
Harish Kriplani, Farid N. Najm, Ping Yang 0001, Ibrahim N. Hajj
DAC2
1993 Transition density: a new measure of activity in digital circuits
abstract
Noting that a common element in most causes of runtime failure is the extent of circuit activity, i.e. the rate at which its nodes are switching, the author proposes a measure of activity, called the transition density, which may be defined as the average switching rate at a circuit node. An algorithm is also presented to propagate density values from the primary inputs to internal and output nodes. To illustrate the practical significance of this work, it is shown how the density values at internal nodes can be used to study circuit reliability by estimating the average power and ground currents; the average power dissipation; the susceptibility to electromigration failures; and the extent of hot-electron degradation. The density propagation algorithm has been implemented in a prototype density simulator which is used to assess the validity and feasibility of the approach experimentally. The results show that the approach is very efficient, and makes possible the analysis of VLSI circuits.>
Farid N. Najm
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1993 A Monte Carlo approach for power estimation
abstract
The authors investigate a power estimation technique for VLSI that combines the accuracy of simulation-based techniques with the speed of the probabilistic techniques. The resulting method is statistical in nature; it consists of applying randomly generated input patterns to the circuit and monitoring, with a simulator, the resulting power value. This is continued until a value of power is obtained with a desired accuracy, at a specified confidence level. The authors present the algorithm and experimental results, and discuss the superiority of the approach.>
Richard Burch, Farid N. Najm, Ping Yang 0001, Timothy N. Trick
IEEE Trans. Very Large Scale Integr. Syst.2
1992 Maximum Current Estimation in CMOS Circuits
Harish Kriplani, Farid N. Najm, Ibrahim N. Hajj
DAC2
1992 McPOWER: a Monte Carlo approach to power estimation
abstract
An alternative technique for power estimation that combines the accuracy of simulation-based techniques with the speed of the probabilistic technique is investigated. The resulting method is statistical in nature; it consists of applying randomly-generated input patterns to the circuit and monitoring, with a simulator, the resulting power value. This is continued until a value of power is obtained with a desired accuracy, at a specified confidence level. The algorithm and experimental results are presented and the superiority of this approach is discussed.>
Richard Burch, Farid N. Najm, Ping Yang 0001, Timothy N. Trick
ICCAD2
1991 Transition Density, A Stochastic Measure of Activity in Digital Circuits
abstract
Reliability assessment is an important part of the design process of digital integrated circuits. We observe that a common thread that runs through most causes of run-time failure is the extent of circuit activity, i.e., the rate at which its nodes are switching. We propose a new measure of activity, called the transition density, which may be de ned as the \\average switching rate " at a circuit node. Based on a stochastic model of logic signals, we rigorously de ne the transition density and present an algorithm to propagate it from the primary inputs to internal and output nodes. This algorithm may be thought ofasasimulation of the circuit, and has been implemented in a prototype density simulator. We present some results of this implementation to verify the theoretical results and assess the feasibility of the approach. In order to obtain the same density information by traditional means, the circuit would need to be simulated for thousands of input transitions. Thus this approach is very e cient and makes possible the analysis of VLSI circuits, which are traditionally too big to simulate for long input sequences. ACM/IEEE Design-Automation Conference, 1991. 1.
Farid N. Najm
DAC1
1991 An extension of probabilistic simulation for reliability analysis of CMOS VLSI circuits
abstract
The probabilistic simulation approach is extended to include the computation of the variance waveform of the power/ground current, in addition to its expected waveform. The focus is on the problem of estimating the median time-to-failure (MTF) due to electromigration (EM) in the power and ground buses of CMOS circuits. Theoretical results that quantify the relationship between the MTF and the statistics of the stochastic current are presented. This leads to a more accurate estimate of the MTF that requires both the expected and variance waveforms. A novel technique is then presented to compute the variance waveform for CMOS circuits, which has been incorporated into the probabilistic simulator CREST. Results of this implementation demonstrating efficiency and accuracy on a number of circuits are provided. The authors use these results to study the importance of the variance waveform by estimating its contribution to the MTF relative to that of the expected waveform.>
Farid N. Najm, Ibrahim N. Hajj, Ping Yang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1990 Probabilistic simulation for reliability analysis of CMOS VLSI circuits
abstract
A current-estimation approach to support the analysis of electromigration (EM) failures in power supply and ground buses of CMOS VLSI circuits is discussed. It uses the original concept of probabilistic simulation to efficiently generate accurate estimates of the expected current waveform required for electromigration analysis. Thus, the approach is pattern-independent and relieves the designer of the tedious task of specifying logical input waveforms. This approach has been implemented in the program CREST (current estimator) which has shown excellent accuracy and dramatic speedups compared with traditional approaches. The approach and its implementation are described, and the results of numerous CREST runs on real circuits are presented.>
Farid N. Najm, Richard Burch, Ping Yang 0001, Ibrahim N. Hajj
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1990 The complexity of fault detection in MOS VLSI circuits
abstract
Consideration is given to the fault detection problem for a single fault in a single MOS channel-connected subcircuit. The following three decision subproblems are identified: (1) decide if a test vector exists, (2) decide if an initializing vector exists, and (3) decide if a test pair is robust. It is proven that each of these problems is NP-complete. More importantly, it is proved that the first two remain NP-complete for the simplest subcircuit design styles, namely series/parallel nMOS or CMOS logic gates. The third subproblem is shown to be of linear complexity for a CMOS logic gate with a stuck-open fault. It is illustrated that a test pair that is not robust may contain a robust subtest pair, and a necessary and sufficient condition for this to happen in CMOS logic gates is given. This leads to a linear time algorithm for CMOS logic gates which tests for robustness and, if possible, derives a robust test pair from a possibly nonrobust pair. The implications of these complexity results on practical transistor-level test generation tools are discussed.>
Farid N. Najm, Ibrahim N. Hajj
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1989 Computation of bus current variance for reliability estimation of VLSI circuits
abstract
A novel technique for deriving the variance waveform for CMOS circuits is presented. Using this technique, the authors establish the importance of the variance waveform by showing that its contribution to the mean-time-to-failure estimate can be in the range of 100% to 200% relative to that of the expected waveform. The technique has been built into the probabilistic simulator CREST and has shown good agreement with SPICE, as well as excellent speedup.>
Farid N. Najm, Ibrahim N. Hajj, Ping Yang 0001
ICCAD1
1989 Electromigration median time-to-failure based on a stochastic current waveform
abstract
The estimation of the median time-to-failure (MTF) due to electromigration in the power and ground buses of VLSI circuits is addressed. In their previous work (Proc. 25th ACM/IEEE Design Autom. Conf., p.294-9, 1988), the authors presented a novel technique for MTF estimation based on a stochastic current waveform model. They derived the mean (or expected) waveform (not a time average) of such a current model and conjectured that it is the appropriate current waveform to be used for MTF estimation. The authors prove that conjecture and present new theoretical results which show the exact relationship between the MTF and the statistics of the stochastic current. This leads to a more accurate technique for deriving the MTF which requires the variance waveform of the current, in addition to its mean waveform. The authors then show how the variances of the bus branch currents can be derived from those of the gate currents, and describe several simplifying approximations that can be used to maintain efficiency and, therefore, make possible the analysis of VLSI circuits.>
Farid N. Najm, Ibrahim N. Hajj, Ping Yang 0001
ICCD1
1988 Pattern-Independent Current Estimation for Reliability Analysis of CMOS Circuits
Richard Burch, Farid N. Najm, Ping Yang 0001, Dale E. Hocevar
DAC2
1988 CREST-a current estimator for CMOS circuits
abstract
CREST is a pattern-independent current estimation approach developed to support electromigration analysis tools. It uses the powerful, original concept of probabilistic simulation to generate accurate estimates of the expected current waveforms efficiently. The original implementation of CREST is extended to circuits containing pass transistors, reconvergent fanout, and feedback, and heuristics to simulate circuits with large reconvergent fanout or feedback blocks efficiently are provided. The results of using CREST on several real circuits are presented.>
Farid N. Najm, Richard Burch, Ping Yang 0001, Ibrahim N. Hajj
ICCAD1