Lawrence T. Pileggi

dblp:p/LawrenceTPileggi · also Larry T. Pileggi, Lawrence T. Pillage · DBLP profile ↗
← Back
188ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-8605-8240ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 173 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 15 · 1 since 2021Artificial intelligence and machine learning · 10 · 2 since 2021Databases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Second-Order Optimization via Quiescence Trajectory Tracing
abstract
Circuit simulation has developed robust numerical methods to achieve fast DC operating point convergence in highly nonlinear systems. Similar challenges arise in nonconvex optimization, where second-order optimization methods often require restrictive step sizes to ensure a monotonically decreasing objective function. Moreover, in the presence of nonlinear objective functions with large Lipschitz constants, increasingly small step-sizes become a bottleneck to fast convergence. Building on established connections between optimization and circuit dynamics, we explore the application of fast DC circuit simulation methods to second-order optimization. Using a dynamic system representation of the trajectory of optimization variables, we exploit the quiescence of the dynamical system to determine the steady state that coincides with the critical point of the objective function. This optimization via quiescence uses a variation of the quasi-steady state analysis method in ACES to adaptively select large step-sizes that sequentially follow each optimization variable to a quasi-steady state until all state variables reach the actual steady state. The result is a second-order optimization method that utilizes large step-sizes and does not require a monotonically decreasing objective function to reach a critical point. Experimentally, we demonstrate the use of this fast DC circuit simulation for optimizing nonconvex problems in general unconstrained optimization problems including a power systems example and compare them to existing state-of-the-art second-order methods, including damped Newton-Raphson, Broyden–Fletcher–Goldfarb–Shanno (BFGS), and Symmetric Rank 1 (SR1).
Aayushya Agarwal, Ronald A. Rohrer, Lawrence T. Pileggi
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 FedECADO: A Dynamical System Model of Federated Learning
abstract
Federated learning harnesses the power of distributed optimization to train a unified machine learning model across separate clients. However, heterogeneous data distributions and computational workloads can lead to inconsistent updates and limit model performance. This work tackles these challenges by proposing FedECADO, a new algorithm inspired by a dynamical system representation of the federated learning process. FedECADO addresses non-IID data distribution through an aggregate sensitivity model that reflects the amount of data processed by each client. To tackle heterogeneous computing, we design a multi-rate integration method with adaptive step-size selections that synchronizes active client updates in continuous time. Compared to prominent techniques, including FedProx, FedExp, and FedNova, FedECADO achieves higher classification accuracies in numerous heterogeneous scenarios.
Aayushya Agarwal, Gauri Joshi, Lawrence T. Pileggi
ICML3
2024 An IP-Agnostic Foundational Cell Array Offering Supply Chain Security
abstract
Growing IC manufacturing complexity and reliance on third-party fabrication create supply chain fragility, contributing to chip shortages and IP security risks. General-purpose ICs can mitigate manufacturing security risks but rely on software-based configurations, which is not optimal for high-consequence applications. Our work proposes a novel IP-agnostic Foundational Cell Array (FC-Array) platform to overcome these challenges. Built on verified standard cells and industry-standard EDA tools, this platform preserves many advantages of an ASIC. By incorporating 3D split manufacturing, we provide semantically secure IP protection and a base wafer that can be stockpiled. Our tests demonstrate both power-efficient (100 MHz) and high-performance (1 GHz) options. In a post-place-and-route simulated 28nm design, our FC-Array shows a worst-case 1.85x increase in power consumption and a 2.61x increase in area compared to standard cell ASICs for equivalent timing performance.
Christopher Talbot, Deepali Garg, Lawrence T. Pileggi, Ken Mai
DAC3
2023 Shedding Light on Inconsistencies in Grid Cybersecurity: Disconnects and Recommendations
abstract
The operational, academic, and policy communities disagree on which threats against the power grid are likely and what damage would ensue. For instance, the feasibility and impact of MadIoT-style attacks is being actively debated. By surveying grid experts (N=18) we find that disagreements are not unique to MadIoT attacks but occur across multiple well-studied grid threats. Based on prior work and our survey, we hypothesize that the disagreements stem from inconsistencies in how grid threats are modeled. We identify five likely causes of modeling inconsistencies: 1) using unrealistic grid topologies, 2) assuming unrealistic capabilities for attackers, 3) exploring too few grid scenarios, 4) using incomplete simulators that omit relevant grid processes, and 5) using simulators that incorrectly model key grid processes. To check these hypotheses, we create a modeling framework and examine how these factors change our understanding of the feasibility and impact of grid threats. We use four diverse grid threats as case studies: MadIoT, False Data Injection Attacks, Substation Circuit Breaker Takeover, and Power Plant Takeover. We find that each of our hypothe-sized causes of modeling inconsistencies has a significant effect on modeling the outcomes of attacks. For example, we find that MadIoT attacks are much less feasible and require significantly more high-wattage IoT devices on realistic topologies than on topologies previously used to model them. In contrast, we find that Substation Circuit Breaker Takeover attacks are much more feasible in emergency scenarios and may require significantly fewer substations for failure than previous modeling suggested. We conclude with actionable recommendations for accurately assessing the impact of threats against the grid.
Brian Singer, Amritanshu Pandey, Shimiao Li, Lujo Bauer, Craig Miller, Lawrence T. Pileggi, Vyas Sekar
SP6
2021 Hardware Redaction via Designer-Directed Fine-Grained eFPGA Insertion
abstract
In recent years, IC reverse engineering and IC fabrication supply chain security have grown to become significant economic and security threats for designers, system integrators, and end customers. Many of the existing logic locking and obfuscation techniques have shown to be vulnerable to attack once the attacker has access to the design netlist either through reverse engineering or through an untrusted fabrication facility. We introduce soft embedded FPGA redaction, a hardware obfuscation approach that allows the designer substitute security-critical IP blocks within a design with a synthesizable eFPGA fabric. This method fully conceals the logic and the routing of the critical IP and is compatible with standard ASIC flows for easy integration and process portability. To demonstrate eFPGA redaction, we obfuscate a RISC-V control path and a GPS P-code generator. We also show that the modified netlists are resilient to SAT attacks with moderate VLSI overheads. The secure RISC-V design has 1.89x area and 2.36x delay overhead while the GPS design has 1.39x area and negligible delay overhead when implemented on an industrial 22nm FinFET CMOS process.
Prashanth Mohan, Oguz Atli, Joseph Sweeney, Onur O. Kibar, Lawrence T. Pileggi, Ken Mai
DATE5
2021 Top-down Physical Design of Soft Embedded FPGA Fabrics
abstract
In recent years, IC reverse engineering and IC fabrication supply chain security have grown to become significant economic and security threats for designers, system integrators, and end customers. Many of the existing logic locking and obfuscation techniques have shown to be vulnerable to attack once the attacker has access to the design netlist either through reverse engineering or through an untrusted fabrication facility. We introduce soft embedded FPGA redaction, a hardware obfuscation approach that allows the designer substitute security-critical IP blocks within a design with a synthesizable eFPGA fabric. This method fully conceals the logic and the routing of the critical IP and is compatible with standard ASIC flows for easy integration and process portability. To demonstrate eFPGA redaction, we obfuscate a RISC-V control path and a GPS P-code generator. We also show that the modified netlists are resilient to SAT attacks with moderate VLSI overheads. The secure RISC-V design has 1.89x area and 2.36x delay overhead while the GPS design has 1.39x area and negligible delay overhead when implemented on an industrial 22nm FinFET CMOS process.
Prashanth Mohan, Oguz Atli, Onur O. Kibar, Mohammed Zackriya V, Lawrence T. Pileggi, Ken Mai
FPGA5
2021 Adversarially robust learning for security-constrained optimal power flow
abstract
In recent years, the ML community has seen surges of interest in both adversarially robust learning and implicit layers, but connections between these two areas have seldom been explored. In this work, we combine innovations from these areas to tackle the problem of N-k security-constrained optimal power flow (SCOPF). N-k SCOPF is a core problem for the operation of electrical grids, and aims to schedule power generation in a manner that is robust to potentially $k$ simultaneous equipment outages. Inspired by methods in adversarially robust training, we frame N-k SCOPF as a minimax optimization problem -- viewing power generation settings as adjustable parameters and equipment outages as (adversarial) attacks -- and solve this problem via gradient-based techniques. The loss function of this minimax problem involves resolving implicit equations representing grid physics and operational decisions, which we differentiate through via the implicit function theorem. We demonstrate the efficacy of our framework in solving N-3 SCOPF, which has traditionally been considered as prohibitively expensive to solve given that the problem size depends combinatorially on the number of potential outages.
Priya L. Donti, Aayushya Agarwal, Neeraj Vijay Bedmutha, Lawrence T. Pileggi, J. Zico Kolter
NeurIPS4
2020 Modeling Techniques for Logic Locking
abstract
Logic locking is a method to prevent intellectual property (IP) piracy. However, under a reasonable attack model, SAT-based methods have proven to be powerful in obtaining the secret key. In response, many locking techniques have been developed to specifically resist this form of attack. In this paper, we demonstrate two SAT modeling techniques that can provide many orders of magnitude speed up in discovering the correct key. Specifically, we consider relaxed encodings and symmetry breaking. To demonstrate their impact, we model and attack a state-of-the-art logic locking technique, Full-Lock. We show that circuits previously unbreakable within 15 days of run time can be solved in seconds. Consequently, in assessing the strength of any given locking, it is imperative that these modeling techniques be considered. To remedy this vulnerability in the considered locking technique, we demonstrate an extended version, logic-enhanced Banyan locking, that is resistant to our proposed modeling techniques.
Joseph Sweeney, Marijn Heule, Lawrence T. Pileggi
ICCAD3
2020 Sensitivity Analysis of Locked Circuits
abstract
Globalization of integrated circuits manufacturing has led to increased security con- cerns, notably theft of intellectual property. In response, logic locking techniques have been developed for protecting designs, but many of these techniques have been shown to be vulnerable to SAT-based attacks. In this paper, we explore the use of Boolean sensi- tivity to analyze these locked circuits. We show that in typical circuits there is an inverse relationship between input width and sensitivity. We then demonstrate the utility of this relationship for deobfuscating circuits locked with a class of “provably secure” logic lock- ing techniques. We conclude with an example of how to resist this attack, although the resistance is shown to be highly circuit dependent.
Joseph Sweeney, Marijn Heule, Lawrence T. Pileggi
LPAR3
2020 From Virtual Characterization to Test-Chips: DFM Analysis Through Pattern Enumeration
abstract
As CMOS technology continues to scale down due to advances in lithography, the interaction of neighboring patterns is exacerbated. Every pattern printed on silicon is influenced by its neighbors given a technology-specific interaction range. As transistors and cells shrink to sizes of the same order of the interaction range, pattern dependencies at the 16-nm node (and below) make the prediction of functional and parametric yield challenging. Precharacterizing all combinations of layouts as a function of all possible neighboring patterns is impractical due to exponential complexity, whereas silicon characterization of all patterns is practically impossible. In this paper, we propose a virtual characterization vehicle (VCV) methodology that can exhaustively identify all uniquely occurring layout patterns in a standard cell library. VCVs can expose all cell neighboring arrangements as a function of a radius of influence while compiling the pattern frequency and also uncovering the arrangement of cells that create unique patterns. Effects that span multiple layers are also captured in our analysis. VCV results can be used to co-design libraries by suggesting favorable compositions as well as favorable layout styles. Finally, VCVs can be turned into test-chips that are guaranteed to cover all identified patterns with the aid of self-testing features. An extremely regular library was developed in 16-nm FinFET technology and is used to showcase our pattern analysis in which DFM quality is shown to be improved with respect to commercial libraries. Our silicon results demonstrate the feasibility of turning the VCV approach into real silicon test chips.
Mayler G. A. Martins, Samuel Nascimento Pagliarini, Mehmet Meric Isgenc, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 A Probabilistic Synapse With Strained MTJs for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are of interest for applications for which conventional computing suffers from the nearly insurmountable memory-processor bottleneck. This paper presents a stochastic SNN architecture that is based on specialized logic-in-memory synaptic units to create a unique processing system that offers massively parallel processing power. Our proposed synaptic unit consists of strained magnetic tunnel junction (MTJ) devices and transistors. MTJs in our synapse are dual purpose, used as both random bit generators and as general-purpose memory. Our neurons are modeled as integrate-and-fire components with thresholding and refraction. Our circuit is implemented using CMOS 28-nm technology that is compatible with the MTJ technology. Our design shows that the required area for the proposed synapse is only [Formula: see text]. When idle, the synapse consumes 675 pW. When firing, the energy required to propagate a spike is 8.87 fJ. We then demonstrate an SNN that learns (without supervision) and classifies handwritten digits of the MNIST database. Simulation results show that our network presents high classification efficiency even in the presence of fabrication variability.
Samuel Nascimento Pagliarini, Sudipta Bhuin, Mehmet Meric Isgenc, Ayan Kumar Biswas, Lawrence T. Pileggi
IEEE Trans. Neural Networks Learn. Syst.5
2020 Logic IP for Low-Cost IC Design in Advanced CMOS Nodes
abstract
Routing closure and design-for-manufacturability (DFM) challenges exacerbate nonrecurring engineering (NRE) costs, a steep barrier to entry for advanced sub-20-nm CMOS nodes, making low-volume fabrication of integrated circuits (ICs) almost intangible. For ICs in which the cost of design dominates the fabrication, we seek to trade some amount of chip area to lower NRE costs. To this end, we consider designing logic cell layouts for easier routing and DFM closure. We accustom a layout simplification and reuse approach to build standard cell libraries such that good pin access and layout regularity are ensured for all cells. Using a commercial 14-/16-nm technology, we developed two 100-cell logic cell libraries that are, respectively, 9 and 10.5 metal tracks tall. Our routing experiments on multiple designs show that taller cells can endure 20% higher placement density while reducing the number of design rule check (DRC) violations by four orders of magnitude compared with a commercial 7.5 track library. Silicon measurements show that taller cells can enable faster design closure in ICs with stringent performance requirements at the cost of a marginal power overhead. Finally, our logic cell design approach can make advanced CMOS nodes more affordable for low-volume design.
Mehmet Meric Isgenc, Mayler G. A. Martins, Mohammed Zackriya V, Samuel Nascimento Pagliarini, Lawrence T. Pileggi
IEEE Trans. Very Large Scale Integr. Syst.5
2019 Efficient SpMV Operation for Large and Highly Sparse Matrices using Scalable Multi-way Merge Parallelization
abstract
The importance of Sparse Matrix dense Vector multiplication (SpMV) operation in graph analytics and numerous scientific applications has led to development of custom accelerators that are intended to over-come the difficulties of sparse data operations on general purpose architectures. However, efficient SpMV operation on large problem (i.e. working set exceeds on-chip storage) is severely constrained due to strong dependence on limited amount of fast random access memory to scale. Additionally, unstructured matrix with high sparsity pose difficulties as most solutions rely on exploitation of data locality. This work presents an algorithm co-optimized scalable hardware architecture that can efficiently operate on very large (~billion nodes) and/or highly sparse (avg. degree <10) graphs with significantly less on-chip fast memory than existing solutions. A novel parallelization methodology for implementing large and high throughput multi-way merge network is the key enabler of this high performance SpMV accelerator. Additionally, a data compression scheme to reduce off-chip traffic and special computation for nodes with exceptionally large number of edges, commonly found in power-law graphs, are presented. This accelerator is demonstrated with 16-nm fabricated ASIC and Stratix® 10 FPGA platforms. Experimental results show more than an order of magnitude improvement over current custom hardware solutions and more than two orders of magnitude improvement over commercial off-the-shelf (COTS) architectures for both performance and energy efficiency.
Fazle Sadi, Joe Sweeney, Tze Meng Low, James C. Hoe, Lawrence T. Pileggi, Franz Franchetti
MICRO5
2018 ChangeDAR: Online Localized Change Detection for Sensor Data on a Graph
abstract
Given electrical sensors placed on the power grid, how can we automatically determine when electrical components (e.g. power lines) fail? Or, given traffic sensors which measure the speed of vehicles passing over them, how can we determine when traffic accidents occur? Both these problems involve detecting change points in a set of sensors on the nodes or edges of a graph. To this end, we propose ChangeDAR (Change Detection And Resolution), which detects changes in an online manner, and reports when and where the change occurred in the graph.
Bryan Hooi, Leman Akoglu, Dhivya Eswaran, Amritanshu Pandey, Marko Jereminov, Lawrence T. Pileggi, Christos Faloutsos
CIKM6
2018 GridWatch: Sensor Placement and Anomaly Detection in the Electrical Grid
Bryan Hooi, Dhivya Eswaran, Hyun Ah Song, Amritanshu Pandey, Marko Jereminov, Lawrence T. Pileggi, Christos Faloutsos
ECML/PKDD (1)6
2018 StreamCast: Fast and Online Mining of Power Grid Time Sequences
abstract
How can we efficiently forecast the power consumption of a location for the next few days? More challengingly, how can we forecast the power consumption if the temperature increases by 10° C, the number of appliances in the grid increase by 20%, and voltage levels increase by 5%? Such ‘what-if scenarios' are crucial for future planning, to ensure that the grid remains reliable even under extreme conditions. Our contributions are as follows: 1) Domain knowledge infusion: we propose a novel Temporal BIG model that extends the physics-based BIG model, allowing it to capture changes over time, trends, and seasonality, and temperature effects. 2) Forecasting: our StreamCast algorithm forecasts multiple steps ahead and outperforms baselines in accuracy. Our algorithm is online, requiring constant update time per new data point and bounded memory. 3) What-if scenarios and anomaly detection: our approach can handle scenarios in which the voltage levels, temperature, or number of appliances change. It also spots anomalies in real data, and provides confidence intervals for its forecasts, to assist in planning for various scenarios. Experimental results show that StreamCast has 27% lower forecasting error than baselines on real data, scales linearly, and runs in 4 minutes on a time sequence of 40 million points.
Bryan Hooi, Hyun Ah Song, Amritanshu Pandey, Marko Jereminov, Lawrence T. Pileggi, Christos Faloutsos
SDM5
2018 Application and Product-Volume-Specific Customization of BEOL Metal Pitch
Samuel Nascimento Pagliarini, Mehmet Meric Isgenc, Mayler G. A. Martins, Lawrence T. Pileggi
IEEE Trans. Very Large Scale Integr. Syst.4
2017 A Systems Approach to Computing in Beyond CMOS Fabrics: Invited
abstract
No abstract available.
Ameya Patil 0001, Naresh R. Shanbhag, Lav R. Varshney, Eric Pop, H.-S. Philip Wong, Subhasish Mitra, Jan M. Rabaey, Jeffrey A. Weldon, Lawrence T. Pileggi, Sasikanth Manipatruni, Dmitri E. Nikonov, Ian A. Young
DAC9
2017 PowerCast: Mining and Forecasting Power Grid Sequences
Hyun Ah Song, Bryan Hooi, Marko Jereminov, Amritanshu Pandey, Lawrence T. Pileggi, Christos Faloutsos
ECML/PKDD (2)5
2016 Re-thinking polynomial optimization: Efficient programming of reconfigurable radio frequency (RF) systems by convexification
abstract
Reconfigurable radio frequency (RF) system has emerged as a promising avenue to achieve high communication performance while adapting to versatile commercial wireless environment. In this paper, we propose a novel technique to optimally program a reconfigurable RF system in order to achieve maximum performance and/or minimum power. Our key idea is to adopt an equation-based optimization method that relies on general-purpose, non-convex polynomial performance models to determine the optimal configurations of all tunable circuit blocks. Most importantly, our proposed approach guarantees to find the globally optimal solution of the non-convex polynomial programming problem by solving a sequence of convex semi-definite programming (SDP) problems based on convexification. A reconfigurable RF front-end example designed for WLAN 802.11g demonstrates that the proposed method successfully finds the globally optimal configuration, while other traditional techniques often converge to local optima.
Fa Wang, Shihui Yin, Minhee Jun, Xin Li 0001, Tamal Mukherjee, Rohit Negi, Lawrence T. Pileggi
ASP-DAC7
2016 Extended statistical element selection: a calibration method for high resolution in analog/RF designs
abstract
In this paper we propose a high resolution digital calibration method for analog/RF circuits that is an extension of the statistical element selection (SES) approach. As compared to SES, the proposed ESES method provides wider calibration range to accommodate multiple variation sources and produces higher calibration yield for the same calibration resolution target. Two types of ESES-based calibration with application in analog/RF designs are demonstrated; current source calibration and phase/delay calibration. As compared to traditional calibration methods, the proposed ESES-based calibration incurs lower circuit overhead while achieving higher calibration resolution. ESES calibration is further applied to a wideband harmonic-rejection receiver design that achieves best-in-class harmonic-rejection performance after calibration.
Renzhi Liu, Jeffrey A. Weldon, Lawrence T. Pileggi
DAC3
2016 On the design of phase locked loop oscillatory neural networks: Mitigation of transmission delay effects
abstract
This paper introduces a novel design of phase locked loop (PLL) based oscillatory neural networks (ONNs) to mitigate the frequency clustering phenomenon caused by transmission delays in real systems. Theoretical analysis of the ONN reveals that transmission delays can produce frequency clustering that leads to synchronization and convergence failure. This paper describes the redesign of ONN dynamics and associated system-level architecture to achieve robustness. Specifically, we first demonstrate that using the phase information of zero-crossing points of inputs as the PLL error signal enables the ONN dynamical model to correctly synchronize under uniform transmission delays. A Type-II PLL based ONN architecture is shown via simulation to provide this property in hardware. Furthermore, to accommodate non-uniform transmission delays in hardware, a phase synchronization technique is proposed that is shown to provide the correct synchronization behavior.
Rongye Shi, Thomas C. Jackson, Brian Swenson, Soummya Kar, Lawrence T. Pileggi
IJCNN5
2016 A wideband RF receiver with extended statistical element selection based harmonic rejection calibration
Renzhi Liu, Lawrence T. Pileggi, Jeffrey A. Weldon
Integr.2
2015 Accurate passivity-enforced macromodeling for RF circuits via iterative zero/pole update based on measurement data
abstract
Passive macromodeling for RF circuit blocks is a critical task to facilitate efficient system-level simulation for large-scale RF systems (e.g., wireless transceivers). In this paper we propose a novel algorithm to find the optimal macromodel that minimizes the modeling error based on measurement data, while simultaneously guaranteeing passivity. The key idea is to attack the passive macromodeling problem by solving a sequence of convex semi-definite programming (SDP) problems. As such, the proposed method can iteratively find the optimal poles and zeros for macromodeling. Our experimental results with several commercial RF circuit examples demonstrate that the proposed macromodeling method reduces the modeling error by 1.31-2.74× over other conventional approaches.
Ying-Chih Wang, Shihui Yin, Minhee Jun, Xin Li 0001, Lawrence T. Pileggi, Tamal Mukherjee, Rohit Negi
ASP-DAC5
2015 A synthesis methodology for application-specific logic-in-memory designs
abstract
For deeply scaled digital integrated systems, the power required for transporting data between memory and logic can exceed the power needed for computation, thereby limiting the efficacy of synthesizing logic and compiling memory independently. Logic-in-Memory (LiM) architectures address this challenge by embedding logic within the memory block to perform basic operations on data locally for specific functions. While custom smart memories have been successfully constructed for various applications, a fully automated LiM synthesis flow enables architectural exploration that has heretofore not been possible. In this paper we present a tool and design methodology for LiM physical synthesis that performs co-design of algorithms and architectures to explore system level trade-offs. The resulting layouts and timing models can be incorporated within any physical synthesis tool. Silicon results shown in this paper demonstrate a 250x performance improvement and 310x energy savings for a data-intensive application example.
Huseyin Ekin Sumbul, Kaushik Vaidyanathan, Qiuling Zhu, Franz Franchetti, Lawrence T. Pileggi
DAC5
2015 Analog neuromorphic computing enabled by multi-gate programmable resistive devices
Vehbi Calayir, Mohamed Darwish, Jeffrey A. Weldon, Lawrence T. Pileggi
DATE4
2015 Enabling portable energy efficiency with memory accelerated library
abstract
Over the last decade, the looming power wall has spurred a flurry of interest in developing heterogeneous systems with hardware accelerators. The questions, then, are what and how accelerators should be designed, and what software support is required. Our accelerator design approach stems from the observation that many efficient and portable software implementations rely on high performance software libraries with well-established application programming interfaces (APIs). We propose the integration of hardware accelerators on 3D-stacked memory that explicitly targets the memory-bounded operations within high performance libraries. The fixed APIs with limited configurability simplify the design of the accelerators, while ensuring that the accelerators have wide applicability. With our software support that automatically converts library APIs to accelerator invocations, an additional advantage of our approach is that library-based legacy code automatically gains the benefit of memory-side accelerators without requiring a reimplementation. On average, the legacy code using our proposed MEmory Accelerated Library (MEALib) improves performance and energy efficiency for individual operations in Intel's Math Kernel Library (MKL) by 38x and 75x, respectively. For a real-world signal processing application that employs Intel MKL, MEALib attains more than 10x better energy efficiency.
Qi Guo 0001, Tze Meng Low, Nikolaos Alachiotis 0001, Berkin Akin, Lawrence T. Pileggi, James C. Hoe, Franz Franchetti
MICRO5
2014 Toward efficient programming of reconfigurable radio frequency (RF) receivers
abstract
Reconfigurable radio frequency (RF) system is an emerging component to mitigate the growing engineering cost for wireless chip design. In this paper, we propose a new methodology for efficient programming of reconfigurable RF receiver. The proposed method is facilitated by two novel techniques: two-phase relaxation search and Pareto-based search space reduction. Our numerical experiments demonstrate that the proposed methodology is more robust (i.e., close to global optimum) and/or efficient (i.e., with low computational cost) than other traditional algorithms based on either local relaxation or simulated annealing.
Jun Tao 0001, Ying-Chih Wang, Minhee Jun, Xin Li 0001, Rohit Negi, Tamal Mukherjee, Lawrence T. Pileggi
ASP-DAC7
2014 Detecting Reliability Attacks during Split Fabrication using Test-only BEOL Stack
abstract
Split fabrication, the process of splitting an IC into an untrusted and trusted tier, facilitates access to the most advanced semiconductor manufacturing capabilities available in the world without requiring disclosure of design intent. While obfuscation techniques have been proposed to prevent malicious circuit insertion or modifications in the untrusted tier, detecting a pernicious reliability attack induced in the offshore foundry is more elusive. We describe a methodology for exhaustive testing of components in the untrusted tier using a specialized test-only metal stack for selected sacrificial dies.
Kaushik Vaidyanathan, Bishnu Prasad Das, Lawrence T. Pileggi
DAC3
2014 Sub-20 nm design technology co-optimization for standard cell logic
abstract
Efficiency and manufacturability of standard cell logic is critical for an IC, as standard cells are at the heart of the nexus between technology definition, circuit design and physical synthesis. Conventional standard cell design techniques are increasingly ineffective as we scale to patterning restricted sub-20 nm CMOS nodes. To meet the constraints and leverage the features of future technology offerings, we propose a holistic design technology co-optimization (DTCO) for standard cell logic. In our holistic DTCO we co-optimize the standard cell architecture to balance manufacturability and efficiency at the cell level while taking into account block level considerations such as pin accessibility and power rail robustness. Our DTCO in a foundry 14 nm CMOS resulted in two standard cell architectures, namely, 10T_BiDir and 10T_UniDir. We evaluated these cell libraries with physically synthesized blocks and ring oscillator test structures in IBM 14SOI process. We observed that 10T_BiDir emerges as the preferred alternative at 14 nm CMOS, with 10T_UniDir promising better scalability to future nodes.
Kaushik Vaidyanathan, Lars Liebmann, Andrzej J. Strojwas, Lawrence T. Pileggi
ICCAD4
2013 Neurocomputing and associative memories based on ovenized aluminum nitride resonators
abstract
Neurocomputing has been regarded as an intriguing alternative to the von Neumann architecture for computing systems, especially for such applications as pattern recognition, image processing, and associative memory. However, implementations using CMOS technology have largely been considered impractical due to the required circuit complexity and corresponding power consumption. In this paper we propose a novel configuration for a recently-developed ovenized aluminum nitride (AlN) resonator that is used as a thermally-tunable analog impedance for implementation of artificial neurons and synapses. We demonstrate and elaborate on our building blocks for artificial neurons and synapses using such resonators. Localized impedance tuning via multiple heaters on a single device enables a compact DAC (digital-to-analog converter) for programming artificial synapses and a simple-yet-efficient means for implementing artificial neurons. We also show the functionality of our proposed circuits using two pattern recognition examples based on compact circuit simulation models for ovenized AlN resonators. The resonator device models are characterized from measurement data.
Vehbi Calayir, Augusto Tazzoli, Gianluca Piazza, Lawrence T. Pileggi
IJCNN5
2013 Fully-digital oscillatory associative memories enabled by non-volatile logic
abstract
Due to its brain-like parallel processing, neurocomputing has been regarded as intriguing alternative to traditional von Neumann architectures for such applications as image processing, pattern recognition, and associative memory. Associative memories based on neurocomputing attempt to mimic the human brain via a parallel network of coupled artificial neurons. Oscillatory neural networks (ONNs) have been proposed for such purposes; however, CMOS-based implementations would be inefficient due to the corresponding circuit complexity of oscillators and phase-locking mechanisms. In addition, programmability of the synaptic weights would require numerous reconfigurable, complex analog circuits that represent an impractical power and area overhead. In this paper we propose a fully-digital ONN architecture that is enabled by non-volatile logic. Using a newly proposed all-magnetic logic family, mLogic, we demonstrate the efficacy of digitizing the oscillators and phase relationships by exploiting the inherent storage. We perform a device-level simulation-based comparison of mLogic and 32nm CMOS for a fully-interconnected 60-neuron system, and show approximately 15× area improvement and 18× power improvement that would be achieved for a large system with 100k neurons.
Vehbi Calayir, Lawrence T. Pileggi
IJCNN2
2012 Design Automation Framework for Application-Specific Logic-in-Memory Blocks
abstract
This paper presents a design methodology forhardware synthesis of application-specific logic-in-memory(LiM) blocks. Logic-in-memory designs tightly integrate specializedcomputation logic with embedded memory, enablingmore localized computation, thus save energy consumption. Asa demonstration, we present an end-to-end design frameworkto automatically synthesize an interpolation based logic-in-memoryblock named interpolation memory, which combinesa seed table with simple arithmetic logic to efficiently evaluatefunctions. In order to support multiple consecutive seed dataaccess that is required in the interpolation operation, wesynthesize the physical memory into the novel rectangular accesssmart memory blocks. We evaluated a large designspace of interpolation memories in sub-20 nm commercialCMOS technology by using the proposed design framework.Furthermore, we implemented a logic-in-memory based computedtomography (CT) medical image reconstruction systemand our experimental results show that the logic-in-memorycomputing method achieves orders of magnitude of energysaving compared with the traditional in-processor computing.
Qiuling Zhu, Kaushik Vaidyanathan, Ofer Shacham, Mark Horowitz, Lawrence T. Pileggi, Franz Franchetti
ASAP5
2012 mLogic: ultra-low voltage non-volatile logic circuits using STT-MTJ devices
abstract
This paper introduces the design of logic circuits based exclusively on novel magnetoelectronic devices. Current signals are steered by 2x resistance change switching while operating with sub-100 mV voltage pulses for power and synchronization. The inherent memory of the devices results in fully pipelined nonvolatile logic. We demonstrate that co-optimization of the devices, circuits and logic can achieve ultra-low energy-per-operation for design examples.
Daniel D. Morris, David M. Bromberg, Jian-Gang Jimmy Zhu, Lawrence T. Pileggi
DAC4
2012 Statistical design and optimization for adaptive post-silicon tuning of MEMS filters
abstract
Large-scale process variations can significantly limit the practical utility of microelectro-mechanical systems (MEMS) for RF (radio frequency) applications. In this paper we describe a novel technique of adaptive post-silicon tuning to reliably design MEMS filters that are robust to process variations. Our key idea is to implement a number of redundant MEMS resonators to form an array and then optimally select a subset of these resonators to achieve the desired frequency response. Several new CAD algorithms and methodologies are proposed to optimize and configure the design variables of the proposed MEMS resonator array. A MEMS design example demonstrates that the proposed post-silicon tuning is able to reduce the ripple of the channel filter gain by 7x over other traditional approaches.
Fa Wang, Gökçe Keskin, Andrew Phelps 0001, Jonathan Rotner, Xin Li 0001, Gary K. Fedder, Tamal Mukherjee, Lawrence T. Pileggi
DAC8
2012 Polar format synthetic aperture radar in energy efficient application-specific logic-in-memory
abstract
In this paper we present a local interpolation-based variant of the well-known polar format algorithm used for synthetic aperture radar (SAR) image formation. We develop the algorithm to match the capabilities of the application-specific logic-in-memory processing paradigm, which off-loads lightweight computation directly into the SRAM and DRAM. Our proposed algorithm performs filtering, an image perspective transformation, and a local 2D interpolation and supports partial and low-resolution reconstruction. We implement our customized SAR grid interpolation logic-in-memory hardware in advanced 32nm silicon technology. Our high-level design tools allow to instantiate various optimized design choices to fit image processing and hardware needs of application designers. Our simulation results show that the logic-in-memory approach has the potential to enable substantial improvements in energy efficiency without sacrificing image quality.
Qiuling Zhu, Christian R. Berger, Eric L. Turner, Lawrence T. Pileggi, Franz Franchetti
ICASSP4
2012 Cost-effective smart memory implementation for parallel backprojection in computed tomography
abstract
As nanoscale lithography challenges mandate greater pattern regularity and commonality for logic and memory circuits, new opportunities are created to affordably synthesize more powerful smart memory blocks for specific applications.Leveraging the ability to embed logic inside the memory block boundary, we demonstrate the synthesis of smart memory archi tectures that exploits the inherent memory address patterns of the backprojection algorithm to enable efficient parallel image reconstruction at minimum hardware overhead.An end-to end design framework in sub-20nm CMOS technologies was constructed for the physical synthesis of smart memories and evaluation of the huge design space.Our experimental resultsshow that customizing memory for the computerized tomogra phy (CT) parallel backprojection can achieve more than 30% area and power savings while offering significant performance improvements with marginal sacrifice of image accuracy.
Qiuling Zhu, Lawrence T. Pileggi, Franz Franchetti
VLSI-SoC2
2011 Formal verification of phase-locked loops using reachability analysis and continuization
abstract
We present an approach for verifying locking of charge-pump phase-locked loops by performing reachability analysis on a behavioral model of the circuit. Bounded uncertain parameters in the behavioral model make it possible to represent all possible behaviors of more detailed models. The dynamics of the behavioral model is hybrid (i.e., discrete and continuous) due to the switching of charge pumps that drive the analog control circuits. A unique feature of phase-locked loops compared to most other hybrid systems is that they require thousands of switchings in the continuous dynamics to converge sufficiently close to a limit cycle. This makes reachability analysis a challenging task since switches in the dynamics are expensive to compute and result in conservative overapproximations. We solve this problem by overapproximating the effects of the switching conditions with uncertain parameters in linear continuous models, a method we call continuization. Using efficient reachability algorithms for discrete-time linear systems, locking is verified over the complete range of possible initial states of a charge-pump PLL designed in 32nm CMOS SOI technology in comparable time required for Monte Carlo simulations of the same behavioral model.
Matthias Althoff, Soner Yaldiz, Akshay Rajhans, Xin Li 0001, Bruce H. Krogh, Lawrence T. Pileggi
ICCAD6
2010 Reducing variability in chip-multiprocessors with adaptive body biasing
abstract
Body biasing has been demonstrated to be effective in addressing process variability in a variety of simple chip designs. Modern microprocessors implement dynamic voltage/frequency scaling, with significant implications for the use of body biasing. For a 16-core chip-multiprocessor implemented in a high-performance 22 nm technology, the body biases required to meet the frequency target at the lowest and highest voltage/frequency levels differ by an average of 0.7 V, implying that per-level biases are required to fully leverage body biasing. The need to make abrupt changes in the bias voltages when the voltage/frequency level changes affects the cost/benefit analysis of body biasing schemes. It is demonstrated that computing unique body biases for each voltage/frequency level at chip power-on offers the best tradeoff among a variety of methods in terms of area, performance, and power.
Alyssa Bonnoit, Lawrence T. Pileggi
ISLPED2
2010 Co-Optimization of Circuits, Layout and Lithography for Predictive Technology Scaling Beyond Gratings
abstract
The financial backbone of the semiconductor industry is based on doubling the functional density of integrated circuits every two years at fixed wafer costs and die yields. The increasing demands for 'computational' rather than 'physical' lithography to achieve the aggressive density targets, along with the complex device-engineering solutions needed to maintain the power density objectives, have caused a rapid escalation in systematic yield limiters that threaten scaling. Specifically, the traditional contract between design and manufacturing based solely on design rules is no longer sufficient to guarantee functional silicon and instead requires a convoluted set of restrictions that force complex modifications to the already costly design flows. In this paper, we claim that a far superior result can be achieved by moving the design-to-manufacturing interface from design rules to a higher level of abstraction based on a defined set of pre-characterized layout templates. We will demonstrate how this methodology can simplify optical proximity correction and lithography processes for sub-32 nm technology nodes, along with various digital block design examples for synthesized intellectual property (IP) cores. Furthermore, with a cost-per-good-die analysis we will show that this methodology will extend economical scaling to sub-32 nm technology nodes.
Tejas Jhaveri, Vyacheslav Rovner, Lars Liebmann, Lawrence T. Pileggi, Andrzej J. Strojwas, Jason Hibbeler
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2009 Creating an affordable 22nm node using design-lithography co-optimization
abstract
Achieving the required time-to-market with economically acceptable yield levels and maintaining them in volume production has become a daunting task for the advanced technology nodes. These difficulties are primarily attributable to the increase in process variability that is incurred while aggressively scaling technology nodes which are based on the same fundamental device architectures and process solutions. The introduction of a Metal Gate/High-K (MGHK) stack at the 32/28nm technology node will help in addressing the random variations due to random dopant fluctuations (RDF), but its benefit will be exhausted after a single process generation [3]. As a result, for the 22/20nm technology nodes, the only hope to limit RDF will be to adopt novel device architecture, such as FinFET and Ultra Thin Body or Fully Depleted SOI, that would reduce the dopant concentration in the channel.
Andrzej J. Strojwas, Tejas Jhaveri, Vyacheslav Rovner, Lawrence T. Pileggi
DAC4
2009 SRAM parametric failure analysis
abstract
With aggressive technology scaling, SRAM design has been seriously challenged by the difficulties in analyzing rare failure events. In this paper we propose to create statistical performance models with accuracy sufficient to facilitate probability extraction for SRAM parametric failures. A piecewise modeling technique is first proposed to capture the performance metrics over the large variation space. A controlled sampling scheme and a nested Monte Carlo analysis method are then applied for the failure probability extraction at cell-level and array-level respectively. Our 65nm SRAM example demonstrates that by combining the piecewise model and the fast probability extraction methods, we have significantly accelerated the SRAM failure analysis.
Jian Wang 0100, Soner Yaldiz, Xin Li 0001, Lawrence T. Pileggi
DAC4
2009 Integrating dynamic voltage/frequency scaling and adaptive body biasing using test-time voltage selection
abstract
Adaptive body biasing is a promising technique for addressing increasing process variability, but it also provides new opportunities for reducing power when combined with dynamic voltage/frequency scaling. Limitations of existing ABB/DVFS proposals are explored, and a new scheme, test-time voltage selection (TTVS), is presented. By delaying the mapping between frequency and supply voltage until test, variability information can be incorporated into the VDD selection process. For a 16-core chip-multiprocessor implemented in a high-performance predictive 22 nm technology, TTVS results in 18% power savings over independent ABB/DVFS and 11% power savings over the best of several previously proposed ABB/DVFS schemes.
Alyssa Bonnoit, Sebastian Herbert, Diana Marculescu, Lawrence T. Pileggi
ISLPED4
2009 Regular Analog/RF Integrated Circuits Design Using Optimization With Recourse Including Ellipsoidal Uncertainty
abstract
Long design cycles due to the inability to predict silicon realities are a well-known problem that plagues analog/RF integrated circuit product development. As this problem worsens for nanoscale IC technologies, the high cost of design and multiple manufacturing spins causes fewer products to have the volume required to support full-custom implementation. Design reuse and analog synthesis make analog/RF design more affordable; however, the increasing process variability and lack of modeling accuracy remain extremely challenging for nanoscale analog/RF design. We propose a regular analog/RF IC using metal-mask configurability design methodology Optimization with Recourse of Analog Circuits including Layout Extraction (ORACLE), which is a combination of reuse and shared-use by formulating the synthesis problem as an optimization with recourse problem. Using a two-stage geometric programming with recourse approach, ORACLE solves for both the globally optimal shared and application-specific variables. Furthermore, robust optimization is proposed to treat the design with variability problem, further enhancing the ORACLE methodology by providing yield bound for each configuration of regular designs. The statistical variations of the process parameters are captured by a confidence ellipsoid. We demonstrate ORACLE for regular Low Noise Amplifier designs using metal-mask configurability, where a range of applications share common underlying structure and application-specific customization is performed using the metal-mask layers. Two RF oscillator design examples are shown to achieve robust designs with guaranteed yield bound.
Yang Xu 0017, Kan-Lin Hsiung, Xin Li 0001, Lawrence T. Pileggi, Stephen P. Boyd
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2008 Automated Testability Enhancements for Logic Brick Libraries
abstract
Circuit fabrics composed of highly regular structures, called logic bricks, have been described recently for improving yield. An automated logic brick design flow based on a SAT formulation of the brick routing has been developed to minimize wire length and the number of vias while maintaining several design-for-manufacturability constraints. In this work, testability enhancements are imposed into a logic brick to reduce the likelihood of (i) feedback bridges to improve test and (ii) equivalent faults to improve diagnosis. This is accomplished by adding constraints to the SAT formulation of the logic brick routing that restricts certain wires from being routed in close proximity, thus making bridges between them unlikely. Application to several brick designs resulted in critical-area reductions for targeted bridges with little degradation in terms of additional wire length and via count.
Jason G. Brown, R. D. (Shawn) Blanton, Lawrence T. Pileggi
DATE4
2008 Digital Circuit Design Challenges and Opportunities in the Era of Nanoscale CMOS
abstract
Well-designed circuits are one key ldquoinsulatingrdquo layer between the increasingly unruly behavior of scaled complementary metal-oxide-semiconductor devices and the systems we seek to construct from them. As we move forward into the nanoscale regime, circuit design is burdened to ldquohiderdquo more of the problems intrinsic to deeply scaled devices. How this is being accomplished is the subject of this paper. We discuss new techniques for logic circuits and interconnect, for memory, and for clock and power distribution. We survey work to build accurate simulation models for nanoscale devices. We discuss the unique problems posed by nanoscale lithography and the role of geometrically regular circuits as one promising solution. Finally, we look at recent computer-aided design efforts in modeling, analysis, and optimization for nanoscale designs with ever increasing amounts of statistical variation.
Benton H. Calhoun, Yu Cao 0001, Xin Li 0001, Ken Mai, Lawrence T. Pileggi, Rob A. Rutenbar, Kenneth L. Shepard
Proc. IEEE5
2008 Defining Statistical Timing Sensitivity for Logic Circuits With Large-Scale Process and Environmental Variations
abstract
The large-scale process and environmental variations for today's nanoscale ICs require statistical approaches for timing analysis and optimization. In this paper, we demonstrate why the traditional concept of slack and critical path becomes ineffective under large-scale variations and propose a novel sensitivity framework to assess the ldquocriticalityrdquo of every path, arc, and node in a statistical timing graph. We theoretically prove that the path sensitivity is exactly equal to the probability that a path is critical and that the arc (or node) sensitivity is exactly equal to the probability that an arc (or a node) sits on the critical path. An efficient algorithm with incremental analysis capability is developed for fast sensitivity computation that has linear runtime complexity in circuit size. The efficacy of the proposed sensitivity analysis is demonstrated on both standard benchmark circuits and large industrial examples.
Xin Li 0001, Jiayong Le, Mustafa Celik, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2008 Quadratic Statistical MAX Approximation for Parametric Yield Estimation of Analog/RF Integrated Circuits
abstract
In this paper, we propose an efficient numerical algorithm for estimating the parametric yield of analog/RF circuits, considering large-scale process variations. Unlike many traditional approaches that assume normal performance distributions, the proposed approach is particularly developed to handle multiple correlated nonnormal performance distributions, thereby providing better accuracy than the traditional techniques. Starting from a set of quadratic performance models, the proposed parametric yield estimation conceptually maps multiple correlated performance constraints to a single auxiliary constraint by using a MAX operator. As such, the parametric yield is uniquely determined by the probability distribution of the auxiliary constraint and, therefore, can easily be computed. In addition, two novel numerical algorithms are derived from moment matching and statistical Taylor expansion, respectively, to facilitate efficient quadratic statistical MAX approximation. We prove that these two algorithms are mathematically equivalent if the performance distributions are normal. Our numerical examples demonstrate that the proposed algorithm provides an error reduction of 6.5 times compared to a normal-distribution-based method while achieving a runtime speedup of 10-20 times over the Monte Carlo analysis with 103samples.
Xin Li 0001, Yaping Zhan, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2007 Efficient Parametric Yield Extraction for Multiple Correlated Non-Normal Performance Distributions of Analog/RF Circuits
abstract
In this paper we propose an efficient numerical algorithm to estimate the parametric yield of analog/RF circuits with consideration of large-scale process variations. Unlike many traditional approaches that assume Normal performance distributions, the proposed approach is especially developed to handle multiple correlated non-Normal performance distributions, thereby providing better accuracy than other traditional techniques. Starting from a set of quadratic performance models, the proposed parametric yield extraction conceptually maps multiple correlated performance constraints to a single auxiliary constraint using a MAX(·) operator. As such, the parametric yield is uniquely determined by the probability distribution of the auxiliary constraint and, therefore, can be easily computed. In addition, a novel second-order statistical Taylor expansion is proposed for an analytical MAX(·) approximation, facilitating fast yield estimation. Our numerical examples in a commercial BiCMOS process demonstrate that the proposed algorithm provides 2--3x error reduction compared with a Normal-distribution-based method, while achieving orders of magnitude more efficiency than the Monte Carlo analysis with 104 samples.
Xin Li 0001, Lawrence T. Pileggi
DAC2
2007 Exact Combinatorial Optimization Methods for Physical Design of Regular Logic Bricks
abstract
As minimum feature sizes continue to scale down, increasing difficulties with subwavelength lithography have spurred research into more regular layout styles, such as Restrictive Design Rules (RDRs) [11] and regular logic fabrics [10]. In this paper we show that the simplicity and discreteness of regular fabrics give rise to powerful exact combinatorial optimization methods for the brick layout problem (the regular fabric equivalent of the cell layout problem). These methods are either inapplicable or intractable for less regular layout styles, such as the DRC-based approach of standard cell layout. Results from our prototype tool demonstrate that these optimization methods are quite practical for bricks of typical size found in large-scale designs.
Lawrence T. Pileggi
DAC2
2007 Parameterized Macromodeling for Analog System-Level Design Exploration
abstract
In this paper we propose a novel parameterized macromodeling technique for analog circuits. Unlike traditional macromodels that are only extracted for a small variation space, our proposed approach captures a significantly larger analog design space to facilitate system-level design exploration. Combining a novel piece-wise approximation algorithm and a new multi-point model-order-reduction approach, the proposed method generates compact macromodels covering the entire feasible design space. Our experiments demonstrate that using such models can achieve more than 60 x speed-up while incurring less than 4% overall error when varying design parameters by an order of magnitude.
Jian Wang 0100, Xin Li 0001, Lawrence T. Pileggi
DAC3
2007 Adaptive post-silicon tuning for analog circuits: concept, analysis and optimization
abstract
The well-known Pelgrom model [14] has demonstrated that the variation between two devices on the same die due to random mismatch is inversely proportional to the square root of the device area: σ ∼ 1/sqrt(Area). Based on the Pelgrom model, analog devices are sized to be large enough to average out random variations. Importantly, with CMOS scaling, variations due to random doping fluctuations are making it exceedingly difficult to control device mismatches by sizing alone; namely, the devices have to be made so large that the benefits of CMOS scaling are not realized for analog and RF circuits. In this paper we propose a novel post-silicon tuning methodology to reduce random mismatches for analog circuits in sub-90nm CMOS. A novel dynamic programming algorithm is incorporated into a fast Monte Carlo simulation flow for statistical analysis and optimization of the proposed tunable analog circuits. We apply the proposed postsilicon tuning methodology to several commonly-used analog circuit blocks. We demonstrate that with the post-silicon tuning, device mismatch exponentially decreases as area increases: σ ∼ exp(-α·Area).
Xin Li 0001, YuTsun Chien, Lawrence T. Pileggi
ICCAD4
2007 Robust Analog/RF Circuit Design With Projection-Based Performance Modeling
abstract
In this paper, a robust analog design (ROAD) tool for post-tuning (i.e., locally optimizing) analog/RF circuits is proposed. Starting from an initial design derived from hand analysis or analog circuit optimization based on simplified models, ROAD extracts accurate performance models via transistor-level simulation and iteratively improves the circuit performance by a sequence of geometric programming steps. Importantly, ROAD sets up all design constraints to include large-scale process and environmental variations, thereby facilitating the tradeoff between yield and performance. A crucial component of ROAD is a novel projection-based scheme for quadratic (both polynomial and posynomial) performance modeling, which allows our approach to scale well to large problem sizes. A key feature of this projection-based scheme is a new implicit power iteration algorithm to find the optimal projection space and extract the unknown model coefficients with robust convergence. The efficacy of ROAD is demonstrated on several circuit examples
Xin Li 0001, Padmini Gopalakrishnan, Yang Xu 0017, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 Asymptotic Probability Extraction for Nonnormal Performance Distributions
abstract
While process variations are becoming more significant with each new IC technology generation, they are often modeled via linear regression models so that the resulting performance variations can be captured via normal distributions. Nonlinear response surface models (e.g., quadratic polynomials) can be utilized to capture larger scale process variations; however, such models result in nonnormal distributions for circuit performance. These performance distributions are difficult to capture efficiently since the distribution model is unknown. In this paper, an asymptotic-probability-extraction (APEX) method for estimating the unknown random distribution when using a nonlinear response surface modeling is proposed. The APEX begins by efficiently computing the high-order moments of the unknown distribution and then applies moment matching to approximate the characteristic function of the random distribution by an efficient rational function. It is proven that such a moment-matching approach is asymptotically convergent when applied to quadratic response surface models. In addition, a number of novel algorithms and methods, including binomial moment evaluation, PDF/CDF shifting, nonlinear companding and reverse evaluation, are proposed to improve the computation efficiency and/or approximation accuracy. Several circuit examples from both digital and analog applications demonstrate that APEX can provide better accuracy than a Monte Carlo simulation with 104samples and achieve up to 10times more efficiency. The error, incurred by the popular normal modeling assumption for several circuit examples designed in standard IC technologies, is also shown
Xin Li 0001, Jiayong Le, Padmini Gopalakrishnan, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2006 Architecture-aware FPGA placement using metric embedding
abstract
Since performance on FPGAs is dominated by the routing architecture rather than wirelength, we propose a new ar-chitecture-aware approach to initial FPGA placement that models the relationship between performance and the routing grid, using concepts from graph embedding and metric geometry. Our approach, CAPRI, can be viewed as an embedding of a graph representing the netlist into a metric space that is representative of the FPGA. First, we develop an analytic metric of distance that models delays along the FPGA routing grid. We then embed a netlist into the defined metric space using matrix projections and online bipartite matching. Experimental comparisons with the popular FPGA tool, VPR, show that with CAPRI's initial solution, the resulting placements show median improvements of 10% in critical path delays for the larger MCNC benchmarks. Total placement runtime is also improved by 2x on average.
Padmini Gopalakrishnan, Xin Li 0001, Lawrence T. Pileggi
DAC3
2006 Projection-based statistical analysis of full-chip leakage power with non-log-normal distributions
abstract
In this paper we propose a novel projection-based algorithm to estimate the full-chip leakage power with consideration of both inter-die and intra-die process variations. Unlike many traditional approaches that rely on log-Normal approximations, the proposed algorithm applies a novel projection method to extract a low-rank quadratic model of the logarithm of the full-chip leakage current and, therefore, is not limited to log-Normal distributions. By exploring the underlying sparse structure of the problem, an efficient algorithm is developed to extract the non-log-Normal leakage distribution with linear computational complexity in circuit size. In addition, an incremental analysis algorithm is proposed to quickly update the leakage distribution after changes to a circuit are made. Our numerical examples in a commercial 90nm CMOS process demonstrate that the proposed algorithm provides 4x error reduction compared with the previously proposed log-Normal approximations, while achieving orders of magnitude more efficiency than a Monte Carlo analysis with 10 4 samples.
Xin Li 0001, Jiayong Le, Lawrence T. Pileggi
DAC3
2006 Design Methodology of Regular Logic Bricks for Robust Integrated Circuits
abstract
Regularity in IC design has been recognized as an effective means to combat variability in nanoscale technologies. One way to enforce design regularity is to implement ICs using a small library of regular logic bricks. In this paper we propose a methodology for the design and synthesis of such logic bricks. Since logic bricks are comprised of a limited set of logic primitives for manufacturability reasons, we propose a primitive-based direct mapping approach for generating optimized bricks that, in contrast to classical synthesis approaches, can provide direct control of implementation structures at abstract functional level based on the detection of natural decompositions that exist in the function. We demonstrate considerable improvement in the performance of logic bricks that are generated by the proposed method as compared with those produced by a commercial synthesis tool.
Kim Yaw Tong, Lawrence T. Pileggi
ICCD2
2006 IC thermal simulation and modeling via efficient multigrid-based approaches
abstract
The ever-increasing power consumption and packaging density of integrated systems creates on-chip temperatures and gradients that can have a substantial impact on performance and reliability. While it is conceptually understood that a thermal equivalent circuit can be constructed to characterize the temperature gradients across the chip, direct and iterative solutions of the corresponding three-dimensional (3-D) equations are often intractable for a full-chip analysis. Integrated circuit (IC)-specific multigrid (MG) techniques for fast chip level thermal steady-state and transient simulation are proposed. This approach avoids an explicit construction of the matrix problem that is intractable for most full-chip problems. Specific MG treatments are proposed to cope with the strong anisotropy of the full-chip thermal problem that is created by the vast difference in material thermal properties and chip geometries. Importantly, this paper demonstrates that only with careful thermal modeling assumptions and appropriate choices for grid hierarchy, MG operators, and smoothing steps across grid points can a full-chip thermal problem be accurately and efficiently analyzed. This paper further speeds up the large thermal transient simulations by incorporating reduced-order thermal models that can be efficiently extracted under the same MG framework. The experiments carried out in this work have shown that the proposed methodology provides sufficient efficiency in both runtime and memory usage.
Peng Li 0001, Lawrence T. Pileggi, Mehdi Asheghi, Rajit Chandra
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 Design methodology for IC manufacturability based on regular logic-bricks
abstract
Implementing logic blocks in an integrated circuit in terms of repeating or regular geometry patterns [6,7] can provide significant advantages in terms of manufacturability and design cost [2]. Various forms of gate and logic arrays have been recently proposed that can offer such pattern regularity to reduce design risk and costs [2,4,9,11,12]. In this paper, we propose a full-mask-set design methodology which provides the same physical design coherence as a configurable array, but with area and other design benefits comparable to standard cell ASICs. This methodology is based on a set of simple logic primitives that are mapped to a set of logic bricks that are defined by a restrictive set of RET(Resolution Enhancement Technique)-friendly geometry patterns. We propose a design methodology to explore trade-offs between the number of bricks and associated level of configurability versus the required silicon area. Results are shown to compare a design implemented with a small number of regular bricks to an implementation based on a full standard cell library in a 90nm CMOS technology.
V. Kheterpal, Vyacheslav Rovner, T. G. Hersan, D. Motiani, Y. Takegawa, Andrzej J. Strojwas, Lawrence T. Pileggi
DAC7
2005 OPERA: optimization with ellipsoidal uncertainty for robust analog IC design
abstract
As the design-manufacturing interface becomes increasingly complicated with IC technology scaling, the corresponding process variability poses great challenges for nanoscale analog/RF design. Design optimization based on the enumeration of process corners has been widely used , but can suffer from inefficiency and overdesign. In this paper we propose to formulate the analog and RF design with variability problem as a special type of robust optimization problem, namely robust geometric programming. The statistical variations in both the process parameters and design variables are captured by a pre-specified confidence ellipsoid. Using such optimization with ellipsoidal uncertainy approach, robust design can be obtained with guaranteed yield bound and lower design cost, and most importantly, the problem size grows linearly with number of uncertain parameters. Numerical examples demonstrate the efficiency and reveal the trade-off between the design cost versus the yield requirement. We will also demonstrate significant improvement in the design cost using this approach compared with corner-enumeration optimization.
Yang Xu 0017, Kan-Lin Hsiung, Xin Li 0001, Ivan Nausieda, Stephen P. Boyd, Lawrence T. Pileggi
DAC6
2005 Correlation-aware statistical timing analysis with non-gaussian delay distributions
abstract
Process variations have a growing impact on circuit performance for today's integrated circuit (IC) technologies. The Non-Gaussian delay distributions as well as the correlations among delays make statistical timing analysis more challenging than ever. In this paper, we present an efficient block-based statistical timing analysis approach with linear complexity with respect to the circuit size, which can accurately predict Non-Gaussian delay distributions from realistic nonlinear gate and interconnect delay models. This approach accounts for all correlations, from manufacturing process dependence, to re-convergent circuit paths to produce more accurate statistical timing predictions. With this approach, circuit designers can have increased confidence in the variation estimates, at a low additional computation cost.
Yaping Zhan, Andrzej J. Strojwas, Xin Li 0001, Lawrence T. Pileggi, David Newmark, Mahesh Sharma
DAC4
2005 Specification Test Compaction for Analog Circuits and MEMS
abstract
Testing a non-digital integrated system against all of its specifications can be quite expensive due to the elaborate test application and measurement setup required. We propose to eliminate redundant tests by employing /spl epsi/-SVM based statistical learning. The application of the proposed methodology to an operational amplifier and a MEMS accelerometer reveal that redundant tests can be statistically identified from a complete set of specification-based tests, with negligible error. Specifically, after eliminating five of eleven specification-based tests for an operational amplifier, the defect escape and yield loss is small at 0.6% and 0.9%, respectively. For the accelerometer, defect escape of 0.2% and yield loss of 0.1% occurs when the hot and cold tests are eliminated. For the accelerometer, this level of compaction would reduce test cost by more than half.
Sounil Biswas, Peng Li 0001, R. D. (Shawn) Blanton, Lawrence T. Pileggi
DATE4
2005 Modeling Interconnect Variability Using Efficient Parametric Model Order Reduction
abstract
Assessing IC manufacturing process fluctuations and their impacts on IC interconnect performance has become unavoidable for modern DSM designs. However, the construction of parametric interconnect models is often hampered by the rapid increase in computational cost and model complexity. In this paper we present an efficient yet accurate parametric model order reduction algorithm for addressing the variability of IC interconnect performance. The efficiency of the approach lies in a novel combination of low-rank matrix approximation and multi-parameter moment matching. The complexity of the proposed parametric model order reduction is as low as that of a standard Krylov subspace method when applied to a nominal system. Under the projection-based framework, our algorithm also preserves the passivity of the resulting parametric models.
Peng Li 0001, Frank Liu 0001, Xin Li 0001, Lawrence T. Pileggi, Sani R. Nassif
DATE4
2005 Defining statistical sensitivity for timing optimization of logic circuits with large-scale process and environmental variations
abstract
The large-scale process and environmental variations for today's nanoscale ICs are requiring statistical approaches for timing analysis and optimization. Significant research has been recently focused on developing new statistical timing analysis algorithms, but often without consideration for how one should interpret the statistical timing results for optimization. In this paper (Li et al., 2005) we demonstrate why the traditional concepts of slack and critical path become ineffective under large-scale variations, and we propose a novel sensitivity-based metric to assess the "criticality" of each path and/or arc in the statistical timing graph. We define the statistical sensitivities for both paths and arcs, and theoretically prove that our path sensitivity is equivalent to the probability that a path is critical, and our arc sensitivity is equivalent to the probability that an arc sits on the critical path. An efficient algorithm with incremental analysis capability is described for fast sensitivity computation that has a linear runtime complexity in circuit size. The efficacy of the proposed sensitivity analysis is demonstrated on both standard benchmark circuits and large industry examples.
Xin Li 0001, Jiayong Le, Mustafa Celik, Lawrence T. Pileggi
ICCAD4
2005 Parameterized interconnect order reduction with explicit-and-implicit multi-parameter moment matching for inter/intra-die variations
abstract
In this paper we propose a novel parameterized interconnect order reduction algorithm, CORE, to efficiently capture both inter-die and intra-die variations. CORE applies a two-step explicit-and-implicit scheme for multiparameter moment matching. As such, CORE can match significantly more moments than other traditional techniques using the same model size. In addition, a recursive Arnoldi algorithm is proposed to quickly construct the Krylov subspace that is required for parameterized order reduction. Applying the recursive Arnoldi algorithm significantly reduces the computation cost for model generation. Several RC and RLC interconnect examples demonstrate that CORE can provide up to 10/spl times/ better modeling accuracy than other traditional techniques, while achieving smaller model complexity (i.e. size). It follows that these interconnect models generated by CORE can provide more accurate simulation result with cheaper simulation cost, when they are utilized for gate-interconnect co-simulation.
Xin Li 0001, Peng Li 0001, Lawrence T. Pileggi
ICCAD3
2005 Projection-based performance modeling for inter/intra-die variations
abstract
Large-scale process fluctuations in nano-scale IC technologies suggest applying high-order (e.g., quadratic) response surface models to capture the circuit performance variations. Fitting such models requires significantly more simulation samples and solving much larger linear equations. In this paper, we propose a novel projection-based extraction approach, PROBE, to efficiently create quadratic response surface models and capture both inter-die and intra-die variations with affordable computation cost. PROBE applies a novel projection scheme to reduce the response surface modeling cost (i.e., both the required number of samples and the linear equation size) and make the modeling problem tractable even for large problem sizes. In addition, a new implicit power iteration algorithm is developed to find the optimal projection space and solve for the unknown model coefficients. Several circuit examples from both digital and analog circuit modeling applications demonstrate that PROBE can generate accurate response surface models while achieving up to 12/spl times/ speedup compared with the traditional methods.
Xin Li 0001, Jiayong Le, Lawrence T. Pileggi, Andrzej J. Strojwas
ICCAD3
2005 Performance-centering optimization for system-level analog design exploration
abstract
In this paper we propose a novel analog design optimization methodology to address two key aspects of top-down system-level design: (1) how to optimally compare and select analog system architectures in the early phases of design; and (2) how to hierarchically propagate performance specifications from system level to circuit level to enable independent circuit block design. Importantly, due to the inaccuracy of early-stage system-level models, and the increasing magnitude of process and environmental variations, the system-level exploration must leave sufficient design margin to ensure a successful late-stage implementation. Therefore, instead of minimizing a design objective function, and thereby converging on a constraint boundary, we apply a novel performance centering optimization. Our proposed methodology centers the analog design in the performance space, and maximizes the distance to all constraint boundaries. We demonstrate that this early-stage design margin, which is measured by the volume of the inscribed ellipsoid lying inside the performance constraints, provides an excellent quality measure for comparing different system architectures. The efficacy of our performance centering approach is shown for analog design examples, including a complete clock data recovery system design and implementation.
Xin Li 0001, Jian Wang 0100, Lawrence T. Pileggi, Tun-Shih Chen, Wanju Chiang
ICCAD3
2005 Temperature-Dependent Optimization of Cache Leakage Power Dissipation
abstract
Leakage power consists of an increasing portion of the total power consumption for modern IC designs. Due to the strong inter-dependency between leakage and temperature, it becomes imperative to consider the thermal effects while optimizing the leakage power. In this paper, we present a temperature-dependent optimization methodology for on-chip caches. By integrating fast yet accurate coupled thermal-leakage simulations into an optimization flow, we are able to optimally tradeoff between the cache performance and leakage power while considering realistic on-chip temperature distribution. Our analysis indicates that for future memory intensive designs, the lack of chip temperature information can cause a significant error in the leakage power estimation, thus leading to non-optimal cache designs. Our results further imply that the optimization of cache performance and leakage power shall be attacked as part of the whole system design task in which chip-level floor planning and its thermal impacts are fully addressed.
Peng Li 0001, Yangdong Deng, Lawrence T. Pileggi
ICCD3
2005 Compact reduced-order modeling of weakly nonlinear analog and RF circuits
abstract
A compact nonlinear model order-reduction method (NORM) is presented that is applicable for time-invariant and periodically time-varying weakly nonlinear systems. NORM is suitable for model order reduction of a class of weakly nonlinear systems that can be well characterized by low-order Volterra functional series. The automatically extracted macromodels capture not only the first-order (linear) system properties, but also the important second-order effects of interest that cannot be neglected for a broad range of applications. Unlike the existing projection-based reduction methods for weakly nonlinear systems, NORM begins with the general matrix-form Volterra nonlinear transfer functions to derive a set of minimum Krylov subspaces for order reduction. Moment matching of the nonlinear transfer functions by projection of the original system onto this set of minimum Krylov subspaces leads to a significant reduction of model size. As we will demonstrate as part of comparison with existing methods, the efficacy of model reduction for weakly nonlinear systems is determined by the achievable model compactness. Our results further indicate that a multipoint version of NORM can substantially improve the model compactness for nonlinear system reduction. Furthermore, we show that the structure of the nonlinear system can be exploited to simplify the reduced model in practice, which is particularly effective for circuits with sharp frequency selectivity. We demonstrate the practical utility of NORM and its extension for macromodeling weakly nonlinear RF communication circuits with periodically time-varying behavior.
Peng Li 0001, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2004 CHIME: coupled hierarchical inductance model evaluation
abstract
Modeling inductive effects accurately and efficiently is a critical necessity for design verification of high performance integrated systems. While several techniques have been suggested to address this problem, they are mostly based on sparsification schemes for the L or L-inverse matrix. In this paper, we introduce CHIME, a methodology for non-local inductance modeling and simulation. CHIME is based on a hierarchical model of inductance that accounts for all inductive couplings at a linear cost, without requiring any window size assumptions for sparsification. The efficacy of our approach stems from representing the mutual inductive couplings at various levels of hierarchy, rather than discarding some of them. A prototype implementation demonstrates orders of magnitude speedup over a full, flat model and significant accuracy improvements over a truncated model. Importantly, this hierarchical circuit simulation capability produces a solution that is as accurate as the hierarchically extracted circuits, thereby providing a "golden standard" against which simpler truncation based models can be validated.
Satrajit Gupta, Lawrence T. Pileggi
DAC2
2004 Routing architecture exploration for regular fabrics
abstract
In an effort to control the parameter variations and systematic yield problems that threaten the affordability of application-specific ICs, new forms of design regularity and structure have been proposed. For example, there has been speculation [6] that regular logic fabrics [1] based on regular geometry patterns [2] can offer tighter control of variations and greater control of systematic manufacturing failures. In this paper we describe a routing framework that accommodates arbitrary descriptions of regular and structured routing architectures. We further propose new regular routing architectures and explore the various performance vs. manufacturability trade-offs. Results demonstrate that a more regular, restricted routing architecture can provide a substantial advantage in terms of manufacturability and predictability while incurring a moderate performance penalty.
V. Kheterpal, Andrzej J. Strojwas, Lawrence T. Pileggi
DAC3
2004 STAC: statistical timing analysis with correlation
abstract
Current technology trends have led to the growing impact of both inter-die and intra-die process variations on circuit performance. While it is imperative to model parameter variations for sub-100nm technologies to produce an upper bound prediction on timing, it is equally important to consider the correlation of these variations for the bound to be useful. In this paper we present an efficient block-based statistical static timing analysis algorithm that can account for correlations from process parameters and re-converging paths. The algorithm can also accommodate dominant interconnect coupling effects to provide an accurate compilation of statistical timing information. The generality and efficiency for the proposed algorithm is obtained from a novel simplification technique that is derived from the statistical independence theories and principal component analysis (PCA) methods. The technique significantly reduces the cost for mean, variance and covariance computation of a set of correlated random variables.
Jiayong Le, Xin Li 0001, Lawrence T. Pileggi
DAC3
2004 A frequency relaxation approach for analog/RF system-level simulation
abstract
The increasing complexity of today's mixed-signal integrated circuits necessitates both top-down and bottom-up system-level verification. Time-domain state-space modeling and simulation approaches have been successfully applied for such purposes (e.g. Simulink); however, analog circuits are often best analyzed in the frequency domain. Circuit-level analyses, such as harmonic balance, have been successfully extended to the frequency domain [2], but these algorithms are impractical for simulating large systems with wide-band input and noise signals. In this paper we proposed a frequency-domain approach for analog/RF system-level simulation that is capable of capturing various second order effects (e.g. nonlinearity, noise, etc.) for both time-invariant and time-varying systems with wide-band inputs. The simulator directly evaluates the frequency domain response at each node via a relaxation scheme that is proven to be convergent under typical circuit conditions. Our experimental results demonstrate the accuracy and efficiency of the proposed simulator under various wide-band input and noise excitations.
Xin Li 0001, Yang Xu 0017, Peng Li 0001, Padmini Gopalakrishnan, Lawrence T. Pileggi
DAC5
2004 ORACLE: optimization with recourse of analog circuits including layout extraction
abstract
Long design cycles due to the inability to predict silicon realities is a well-known problem that plagues analog/RF integrated circuit product development. As this problem worsens for technologies below 100nm, the high cost of design and multiple manufacturing spins causes fewer products to have the volume required to support full custom implementation. Design reuse and analog synthesis make analog/RF design more affordable; however, the increasing process variability and lack of modeling accuracy remains extremely challenging for nanoscale analog/RF design. We propose an analog/RF circuit design methodology ORACLE, which is a combination of reuse and \emph{shared-use by formulating the synthesis problem as an \emph{optimization with recourse problem. Using a two-stage geometric programming with recourse approach, ORACLE solves for both the globally optimal shared and application-specific variables. Concurrently, we demonstrate ORACLE for novel metal-mask configurable designs, where a range of applications share common underlying structure and application-specific customization is performed using the metal-mask layers. We also include the silicon validation of the metal-mask configurable designs.
Yang Xu 0017, Lawrence T. Pileggi, Stephen P. Boyd
DAC2
2004 An Interconnect Channel Design Methodology for High Performance Integrated Circuits
abstract
On-chip communication is becoming a bottleneck for high performance designs. Conventional interconnect design methodology does not account for architectures and/or communication schemes that require storage buffers (first-in-first-out queues or FIFOs) in the interconnect channel. For example, FIFOs and flow-control are needed for Network-on-Chip, high performance ASICs and multiple clock domain designs. These IC implementation architectures require an efficient methodology to determine the size of the FIFOs in the channel since the FIFO sizes affect system performance. In this work we devised a methodology to size the FIFOs in an interconnect channel containing one or more FIFOs connected in series. We show that the sizing of the FIFOs in the channel is a function of system parameters such as data production rate and consumption rate, data burstiness, number of channel stages etc. and we also quantify their effect on performance. For a single clock design, we have developed an efficient algorithm which reduces the search space for the optimal sizing of the FIFOs in the channel.
Vikas Chandra, Anthony Xu, Herman Schmit, Lawrence T. Pileggi
DATE4
2004 Exploring Logic Block Granularity for Regular Fabrics
abstract
Driven by the economics of design and manufacturing nanoscale integrated circuits, an emphasis is being placed on developing new, regular logic fabrics that leverage the regularity and programmability of FPGAs, yet deliver a level of performance and density close to ASICs. One example of such a fabric is a Via-Patterned Gate Array (VPGA) according to Pillegi et al. (2002), which employs ASIC style global routing on top of an array of patternable logic blocks (PLBs). Previous works (Koorapaty et al., 2003; Koorapaty, 2003; Pileggi et al., 2003) showed that by employing even limited heterogeneity for the VPGA logic blocks, namely combining a 3-LUT with two 3-input Nand gates, one can achieve performance comparable to that provided by standard cells. Since the area cost for such heterogeneity id far less for FPGAs, we can explore new configurations of via-configurable logic blocks that offer greater heterogeneity and granularity to achieve even higher performance. In this paper, we present a new, more granular, via-patterned heterogeneous logic block architecture and compare it to a less granular LUT-based heterogeneous PLB. Our results show higher performance and more effective packing of the logic functions due to increased granularity.
Aneesh Koorapaty, V. Kheterpal, Padmini Gopalakrishnan, M. Fu, Lawrence T. Pileggi
DATE5
2004 A power aware system level interconnect design methodology for latency-insensitive systems
abstract
Latency-insensitive interconnects require first-in-first-out buffers (FIFO) for flow-control and storage. Interconnect delays are not scaling in proportion to the clock period and hence multiple stages of FIFOs will be needed for high performance interconnects. FIFOs in the interconnect are a significant contributor to the total power consumption. In this work, we propose a design methodology to synthesize a low power interconnect channel containing series connected FIFOs for latency-insensitive systems. Our approach is the first to consider and simultaneously optimize the channel clock frequency, voltage and the FIFO sizes to minimize the power consumption. For small problem size, we show that our approach finds solutions which are close to optimal. The power aware interconnect channel synthesis is affected by the system parameters like the data production rate and data consumption rate. The choice of optimal channel clock frequency, voltage and FIFO sizes can lead to power savings as high as 77.7%, 83.6% and 87% for a 3 stage, 4 stage and a 5 stage channel respectively.
Vikas Chandra, Herman Schmit, Anthony Xu, Lawrence T. Pileggi
ICCAD4
2004 Robust analog/RF circuit design with projection-based posynomial modeling
abstract
We propose a robust analog design tool (ROAD) for post-tuning analog/RF circuits. Starting from an initial design derived from hand analysis or analog circuit synthesis based on simplified models, ROAD extracts accurate posynomial performance models via transistor-level simulation and optimizes the circuit by geometric programming. Importantly, ROAD sets up all design constraints to include large-scale process variations to facilitate the tradeoff between yield and performance. A novel convex formulation of the robust design problem is utilized to improve the optimization efficiency and to produce a solution that is superior to other local tuning methods. In addition, a novel projection-based approach for posynomial fitting is used to facilitate scaling to large problem sizes. A new implicit power iteration algorithm is proposed to find the optimal projection space and extract the posynomial coefficients with robust convergence. The efficacy of ROAD is demonstrated on several circuit examples.
Xin Li 0001, Padmini Gopalakrishnan, Yang Xu 0017, Lawrence T. Pileggi
ICCAD4
2004 Asymptotic probability extraction for non-normal distributions of circuit performance
abstract
While process variations are becoming more significant with each new IC technology generation, they are often modeled via linear regression models so that the resulting performance variations can be captured via normal distributions. Nonlinear (e.g. quadratic) response surface models can be utilized to capture larger scale process variations; however, such models result in non-normal distributions for circuit performance which are difficult to capture since the distribution model is unknown. In this paper we propose an asymptotic probability extraction method, APEX, for estimating the unknown random distribution when using nonlinear response surface modeling. APEX first uses a binomial moment evaluation to efficiently compute the high order moments of the unknown distribution, and then applies moment matching to approximate the characteristic function of the random circuit performance by an efficient rational function. A simple statistical timing example and an analog circuit example demonstrate that APEX can provide better accuracy than Monte Carlo simulation with 10 samples and achieve orders of magnitude more efficiency. We also show the error incurred by the popular normal modeling assumption using standard IC technologies.
Xin Li 0001, Jiayong Le, Padmini Gopalakrishnan, Lawrence T. Pileggi
ICCAD4
2004 Efficient harmonic balance simulation using multi-level frequency decomposition
abstract
Efficient harmonic balance (HB) simulation provides a useful tool for the design of RF and microwave integrated circuits. For practical circuits that can contain strong nonlinearities, however, HB problems cannot be solved reliably or efficiently using conventional techniques. Various preconditioning techniques have been proposed to facilitate a robust and efficient analysis based on Krylov subspace linear solvers. In This work we introduce a multi-level frequency domain preconditioner based on a hierarchical frequency decomposition approach. At each Newton iteration, we recursively solve a set of smaller problems to provide an effective preconditioner for the large linearized HB problem. Compared to the standard single-level block diagonal preconditioner, our experiments indicate that our approach provides a more robust, memory efficient solution while offering a 2-9/spl times/ speedup for several strongly nonlinear HB problems in our experiments.
Peng Li 0001, Lawrence T. Pileggi
ICCAD2
2004 Efficient full-chip thermal modeling and analysis
abstract
The ever-increasing power consumption and packaging density of integrated systems creates on-chip temperatures and gradients that can have a substantial impact on performance and reliability. While it is conceptually understood that a thermal equivalent circuit can be constructed to characterize the temperature gradients across the chip, direct and iterative solutions of the corresponding 3D equations are often intractable for a full-chip analysis. Multigrid accelerated iterative methods can be applied to solve the equivalent circuit problem that is provably symmetric positive definite; however, explicitly building the matrix problem is intractable for most full-chip problems. In This work we present a multigrid iterative approach for the full-chip thermal analysis which does not require explicit construction of the equivalent circuit matrix. We propose specific multigrid treatments to cope with the strong anisotropy of the full-chip thermal problem that is created by the vast difference in material thermal properties and chip geometries. Importantly, we demonstrate that only with careful thermal modeling assumptions and appropriate choices for grid hierarchy, multigrid operators and smoothing steps across grid points, can we accurately and efficiently analyze a full-chip thermal problem. Experimental results demonstrate the efficacy of the proposed multigrid methodology. Our prototyped thermal simulator is able to solve a steady-state problem with more than 10 million unknowns in 125 CPU seconds with a peak memory usage of 231 mega bytes.
Peng Li 0001, Lawrence T. Pileggi, Mehdi Asheghi, Rajit Chandra
ICCAD2
2004 Toward an Integrated Design Methodology for Fault-Tolerant, Multiple Clock/Voltage Integrated Systems
abstract
This paper describes a communication-centric design methodology that addresses the fundamental challenges induced by the emergence of truly heterogeneous systems-on-chip (SoCs). For such systems, the globally asynchronous design paradigm seems to be the most promising (if not the only) solution for providing an underlying substrate for cost-effective and power efficient on-chip communication among diverse, mixed technology IPs. Additional challenges are related to reliability and error resilience of on-chip communication architectures. The proposed on-chip communication methodology targets all levels of abstraction, from circuit, to microarchitecture and system-level by seamlessly integrating solutions for robust and efficient globally asynchronous communication among diverse IPs.
Radu Marculescu, Diana Marculescu, Lawrence T. Pileggi
ICCD3
2004 Parasitics extraction with multipole refinement
abstract
Modern chip design pushes the performance of a given technology to its limits, therefore, it is necessary to find increasingly more accurate models for interconnect parasitics. The growing complexity of today's integrated systems, however, makes fast analysis crucial as well. We present a novel hierarchical potential evaluation technique which is able to represent detailed near-field and global far-field couplings with equal accuracy and efficiency by combining the best features of known hierarchical approaches in this field.
Michael W. Beattie, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 A frequency separation macromodel for system-level simulation of RF circuits
abstract
In this paper we propose a frequency-separation methodology to generate system-level macromodels for analog and RF circuits. The proposed macromodels are similar in form to those based on Volterra kernel calculations, but are much simpler in terms of characterization and overall model complexity, and can be derived from existing device models. This simplicity is realized by applying some basic assumptions on the form of the input excitations, and via separation of the nonlinearities from the dynamic behavior. In addition, by further separating the ideal model functionality, this macromodel is applicable to strongly nonlinear components such as mixers. While time-varying Volterra series models have been proposed for mixers with a fixed local oscillation (LO) signal, the proposed frequency separation model is completely general and can capture the variations of the LO input during a system-level simulation. The proposed macromodels are demonstrated in a system-level simulation tool based on Simulink for efficient evaluation of the entire RF system and associated components. A GSM receiver system in 0.25μm CMOS process is used to demonstrate the efficacy of these macromodels in our system-level simulation environment.
Xin Li 0001, Peng Li 0001, Yang Xu 0017, Robert Dimaggio, Lawrence T. Pileggi
ASP-DAC5
2003 Nonlinear distortion analysis via linear-centric models
abstract
An efficient distortion analysis methodology is presented for analog and RF circuits that utilizes linear-centric circuit models to generate individual distortion contributions due to the various circuit nonlinearities. The per-nonlinearity distortion results are obtained via a straightforward post-simulation step that is simpler and more efficient than the Volterra series based approaches and do not require the high order device model derivatives. For this reason the order of analysis can be significantly higher than that for a Volterra series implementation while fully accounting for all nonlinearity effects. The proposed methodology is not restricted to weakly nonlinear circuits, but can also analyze per-nonlinearity distortion for active switching mixers and switch capacitor circuits when they are modeled as periodically time-varying weakly nonlinear systems. While Volterra series have also been attempted for this same class of circuits, the requirement of device models for all of the high order model derivatives makes such analysis somewhat impractical. The proposed methodology provides important design insights regarding the relationships between design parameters and circuit linearity, hence the overall system performance. Circuit examples are used to demonstrate the efficacy of the proposed approach, and interesting insights are observed for RF switching mixers in particular.
Peng Li 0001, Lawrence T. Pileggi
ASP-DAC2
2003 Fast, cheap and under control: the next implementation fabric
abstract
No abstract available.
Abbas El Gamal, Ivo Bolsens, Andy Broom, Christopher Hamlin, Philippe Magarshack, Zvi Or-Bach, Lawrence T. Pileggi
DAC7
2003 Analog and RF circuit macromodels for system-level analysis
abstract
Design and validation of mixed-signal integrated systems require system-level model abstractions. Generalized Volterra series based models have been successfully applied for analog and RF component macromodels, but their complexity can sometimes limit their utility for time-varying systems and large circuits with complex device models or numerous parasitics. In this paper we propose simple and efficient analog and RF circuit macromodels that provide accurate model abstractions for large, complex time-varying circuits over frequency bands of interest. By starting with the system-level block diagram model structures and focusing on the narrow RF bands, the proposed macromodels can efficiently capture the nonlinear behavior as well as the impact of RLC coupling parasitics via compact reduced-order model forms. While the macromodel can trade accuracy for simplicity in terms of the number of frequency expansion points, we find that expansion about one frequency point provides the accuracy required for system-level analysis of most RF and narrow-band analog components. The macromodel form corresponds to block diagram structures that are easily incorporated into our system-level simulation tool based on Simulink.
Xin Li 0001, Peng Li 0001, Yang Xu 0017, Lawrence T. Pileggi
DAC4
2003 NORM: compact model order reduction of weakly nonlinear systems
abstract
This paper presents a compact Nonlinear model Order Reduction Method (NORM) that is applicable for time-invariant and time-varying weakly nonlinear systems. NORM is suitable for reducing a class of weakly nonlinear systems that can be well characterized by low order Volterra functional series. Unlike existing projection based reduction methods [6]-[8], NORM begins with the general matrix-form Volterra nonlinear transfer functions to derive a set of minimum Krylov subspaces for order reduction. Direct moment matching of the nonlinear transfer functions by projection of the original system onto this set of minimum Krylov subspaces leads to a significant reduction of model size. As we will demonstrate as part of our comparison with existing methods, the efficacy of model order for weakly nonlinear systems is determined by the extend to which models can be reduced. Our results further indicate that a multiple-point version of NORM can substantially reduce the model size and approach the ultimate model compactness that is achievable for nonlinear system reduction. We demonstrate the practical utility of NORM for macro-modeling weakly nonlinear RF circuits with time-varying behavior.
Peng Li 0001, Lawrence T. Pileggi
DAC2
2003 Exploring regular fabrics to optimize the performance-cost trade-off
abstract
While advances in semiconductor technologies have pushed achievable scale and performance to phenomenal limits for ICs, nanoscale physical realities dictate IC production based on what we can afford. We believe that IC design and manufacturing can be made more affordable, and reliable, by removing some design and implementation flexibility and enforcing new forms of design regularity. This paper discusses some of the trade-offs to consider for determination of how much regularity a particular IC or application can afford. A Via Patterned Gate Array is proposed as one such example that trades performance for cost by way of new forms of design regularity.
Lawrence T. Pileggi, Herman Schmit, Andrzej J. Strojwas, Padmini Gopalakrishnan, V. Kheterpal, Aneesh Koorapaty, Chetan Patel, Vyacheslav Rovner, Kim Yaw Tong
DAC1
2003 Heterogeneous Programmable Logic Block Architectures
Aneesh Koorapaty, Vikas Chandra, Kim Yaw Tong, Chetan Patel, Lawrence T. Pileggi, Herman Schmit
DATE5
2003 Noise Macromodel for Radio Frequency Integrated Circuits
abstract
Noise performance is a critical analog and RF circuit design constraint, and can impact the selection of the IC system-level architecture. It is therefore imperative that some model of the noise is represented at the highest levels of abstraction during the design process. In this paper we propose a noise macromodel for analog circuits and demonstrate it by way of implementation in a system level simulator based on MATLAB. We also explain our process of macromodel extraction via reformulation of frequency-domain noise analysis results, and the corresponding steps of model order reduction. The results demonstrate the efficacy of this macromodel for frequency domain system level simulation.
Yang Xu 0017, Xin Li 0001, Peng Li 0001, Lawrence T. Pileggi
DATE4
2003 Heterogeneous Logic Block Architectures for Via-Patterned Programmable Fabrics
Aneesh Koorapaty, Lawrence T. Pileggi, Herman Schmit
FPL2
2003 Bounding the efforts on congestion optimization for physical synthesis
abstract
In this era of Deep Sub-Micron (DSM) technologies, interconnects are becoming increasingly important as their effects strongly impact the integrated circuit (IC) functionality and performance. Moreover, logic block size is no longer determined exclusively by total cell area, and is often limited by wiring resources, yet synthesis optimization objectives are focused on minimizing the number and size of library cells. Methodologies that incorporate congestion within the logic synthesis have been proposed in the past. However, in [15] and [16] it was demonstrated that predicting the true congestion prior to layout is not possible, since different layout regions can have very different routing demands, and the effectiveness of any congestion minimization approach can only be evaluated after routing is completed within the assigned die size. In these works, congestion minimization efforts at the synthesis level are controlled by means of a global weighting factor in the technology mapping cost function. Nevertheless, due to the lack of accurate congestion models, during logic synthesis it is not possible to estimate a priori which values of the congestion minimization factor will yield a congestion-free synthesized netlist. In this paper, we derive practical bounds, which limit the search space for an optimal congestion minimization factor that produces a routable netlist within fixed floorplan constraints. Although we believe that a top-down single-pass congestion-aware logic synthesis is not going to work in general, the bounds obtained in this work can be used in a practical and robust congestion minimization methodology, which can be implemented into any commercial design flow.
Davide Pandini, Lawrence T. Pileggi, Andrzej J. Strojwas
ACM Great Lakes Symposium on VLSI2
2003 A fast simulation approach for inductive effects of VLSI interconnects
abstract
Modeling on-chip inductive effects for interconnects of multi-gigahertz microprocessors remains challenging. SPICE simulation of these effects is very slow because of the large number of mutual inductances. Meanwhile, ignoring the non-linear behavior of drivers in a fast linear circuit simulator results in large errors for the inductive effect. In this paper, a fast and accurate time-domain transient analysis approach is presented, which captures the non-linearity of circuit drivers, the effect of non-ideal ground and de-coupling capacitors in a bus structure. The proposed method models the non-linearity of drivers in conjunction with specific bus geometries. Linearized waveforms at each driver output are incorporated into an interconnect reduced-order simulator for fast transient simulation. In addition, non-ideal ground and de-coupling capacitor models enable accurate signal and ground bounce simulations. Results show that this simulation approach is up-to 68x faster than SPICE while maintaining 95% accuracy.
Xiaoning Qi, Goetz Leonhardt, Daniel Flees, Sangwoo Kim, Stephan Mueller, Hendrik T. Mau, Lawrence T. Pileggi
ACM Great Lakes Symposium on VLSI8
2003 Circuit Simulation of Nanotechnology Devices with Non-monotonic I-V Characteristics
Jiayong Le, Lawrence T. Pileggi, Anirudh Devgan
ICCAD2
2003 A Hybrid Approach to Nonlinear Macromodel Generation for Time-Varying Analog Circuits
Peng Li 0001, Xin Li 0001, Yang Xu 0017, Lawrence T. Pileggi
ICCAD4
2003 An architectural exploration of via patterned gate arrays
abstract
In this work we investigate the architecture of a Via Patterned Gate Array (VPGA) [1], focusing primarily on: 1) the optimal lookup table (LUT) size; and 2) a comparison the crossbar and switch block routing architectures. Unlike FPGAs, the routing architectures in a VPGA do not dominate the total area of the circuit. Therefore our results suggest that using smaller LUTs results in a much faster and smaller design. In the routing architecture comparison, our results also show that the switch block architecture is inferior to the crossbar architecture in terms of area utilization. As the number of routing tracks grows, the switch block architecture begins to dominate the total area of the design as in the case of the FPGAs.
Chetan Patel, Anthony Cozzie, Herman Schmit, Lawrence T. Pileggi
ISPD4
2003 Efficient per-nonlinearity distortion analysis for analog and RF circuits
abstract
An efficient distortion analysis methodology is presented for analog and RF circuits that utilizes linear-centric circuit models to generate individual distortion contributions due to each nonlinear component in a circuit. The per-nonlinearity distortion results are obtained via a straightforward post-simulation step that is simpler and more efficient than the Volterra series-based approaches and does not require high-order device-model derivatives. For this reason, the order of analysis can be significantly higher than that for a Volterra series-based implementation while fully accounting for all distortion effects using most existing device models. Moreover, the proposed methodology can also analyze per-nonlinearity distortion for active switching mixers and switch capacitor circuits when they are modeled as periodically time-varying weakly nonlinear systems. The proposed methodology provides important design insights regarding the relationships between design parameters and circuit linearity, hence, the overall system performance. Circuit examples are used to demonstrate the efficacy of the proposed approach, and interesting insights are observed for RF switching mixers in particular.
Peng Li 0001, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Global and local congestion optimization in technology mapping
abstract
In this era of deep submicrometer technologies, interconnects are becoming increasingly important as their effects strongly impact the integrated circuit (IC) functionality and performance. Moreover, logic block size is no longer determined exclusively by total cell area and is often limited by wiring area. However, synthesis optimization objectives are focused on minimizing the number and size of library cells. Methodologies that incorporate congestion within the logic synthesis objective function have been proposed in the past. Nevertheless, we will demonstrate that predicting the true congestion prior to layout is not possible, and the effectiveness of any congestion minimization approach can only be evaluated after routing is completed within the fixed die size. In this paper, we propose a practical, complete methodology which first performs congestion-aware technology mapping using a global weighting factor for the technology-dependent synthesis cost function and then applies incremental localized unmapping and remapping on layout congested areas. This complete approach addresses the problem that one global factor is not suited for all layout regions of the design, which might have very different routing demands. Most importantly, through the application of this methodology to industrial examples, we will show that any attempt at a purely top-down single-pass congestion-aware technology mapping is merely wishful thinking.
Davide Pandini, Lawrence T. Pileggi, Andrzej J. Strojwas
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 On the efficacy of simplified 2D on-chip inductance models
abstract
Full three-dimensional (3D) inductance models of on-chip interconnect contain an extremely large number of forward coupling terms. It is therefore desirable to use a two-dimensional (2D) approximation in which forward couplings are not included. Unlike capacitive coupling, however, truncating mutual inductance terms can result in loss of accuracy and even instability. This paper investigates whether ignoring forward couplings is an acceptable choice for all good IC designs or if full 3D models are necessary in certain on-chip interconnect configurations. We show that the significance of the forward coupling inductance depends on various aspects of the design.
Michael W. Beattie, Lawrence T. Pileggi
DAC3
2002 Modeling and analysis of regular symmetrically structured power/ground distribution networks
abstract
In this paper we propose a novel and efficient methodology for modeling and analysis of regular symmetrically-structured power/ ground distribution networks. The modeling of inductive effects is simplified by a folding technique which exploits the symmetry in the power/ground distribution. Furthermore, employment of susceptance [10,11] (inverse of inductance) models enables further simplification of the analysis, and is also shown to preserve the symmetric positive definiteness of the circuit equations. Experimental results demonstrate that our approach can provide up to 8x memory savings and up to10x speedup over the already efficient simulation based on the original sparse susceptance matrix without loss of accuracy. Importantly, this work demonstrates that by employing limited regularity, one can create excellent power/ground distribution designs that are dramatically simpler to analyze, and therefore amenable to more powerful global design optimization.
Lawrence T. Pileggi
DAC2
2002 A Linear-Centric Simulation Framework for Parametric Fluctuations
abstract
The relative tolerances for interconnect and device parameter variations have not scaled with feature sizes which have brought about significant performance variability. As we scale toward 10 nm technologies, this problem will only worsen. New circuit families and design methodologies will emerge to facilitate construction of reliable systems from unreliable nanometer scale components. Such methodologies require new models of performance which accurately capture the manufacturing realities. Recently, one step toward this goal was made via a new variational reduced order interconnect model that efficiently captures large scale fluctuations in global parameter values. Using variational calculus the linear interconnect systems are represented by analytical models that include the global variational parameters explicitly. In this work we present a framework which extends the previous work to a linear-centric simulation methodology with accurate nonlinear device models and their fluctuations. The framework is applied to generate path delay distributions under nonlinear and linear parameter fluctuations.
Emrah Acar, Sani R. Nassif, Lawrence T. Pileggi
DATE3
2002 A Linear-Centric Modeling Approach to Harmonic Balance Analysis
abstract
In this paper, we propose a new harmonic balance simulation methodology based on a linear-centric modeling approach. A linear circuit representation of the nonlinear devices and associated parasitics is used along with corresponding time and frequency domain inputs to solve for the nonlinear steady-state response via successive chord (SC) iterations. For our circuit examples, this approach is shown to be up to 60/spl times/ more run-time efficient than traditional Newton-Raphson (N-R) based iterative methods, while providing the same level of accuracy. This SC-based approach converges as reliably as the N-R approaches, including for circuit problems which cause alternative relaxation-based harmonic balance approaches to fail. The efficacy of this linear-centric methodology further improves with increasing model complexity, the inclusion of interconnect parasitics and other analyses that are otherwise difficult with traditional nonlinear models.
Peng Li 0001, Lawrence T. Pileggi
DATE2
2002 On-Chip Inductance Models: 3D or Not 3D?
abstract
Full 3D lumped partial inductance models usually contain a tremendous amount of forward coupling terms. To reduce the complexity of simulation and analysis, a simplified model that excludes the forward coupling terms is often adopted in practice. This paper addresses the question whether ignoring forward couplings is always an acceptable choice or if full 3D models are necessary in certain cases. We show that the significance of the forward coupling inductance depends on various aspects of the design.
Michael W. Beattie, Lawrence T. Pileggi
DATE3
2002 Congestion-Aware Logic Synthesis
abstract
In this era of Deep Sub-Micron (DSM) technologies, the impact of interconnects is becoming increasingly important as it relates to integrated circuit (IC) functionality and performance. In the traditional top-down IC design flow, interconnect effects are first taken into account during logic synthesis by way of wireload models. However, for technologies of 0.25 /spl mu/m and below, the wiring capacitance dominates the gate capacitance and the delay estimation based on fanout and design legacy statistics can be highly inaccurate. In addition, logic block size is no longer dictated solely by total cell area, and is often limited by wiring area resources. For these reasons, wiring congestion is an extremely important design factor, and should be taken into consideration at the earliest possible stages of the design flow. In this paper we propose a novel methodology to incorporate congestion minimization within logic synthesis, and present results for industrial circuits that validate our approach.
Davide Pandini, Lawrence T. Pileggi, Andrzej J. Strojwas
DATE2
2002 Window-Based Susceptance Models for Large-Scale RLC Circuit Analyses
abstract
Due to the increasing operating frequencies and the manner in which the corresponding integrated circuits and systems must be designed, the extraction, modeling and simulation of the magnetic couplings for final design verification can be a daunting task. In general, when modeling inductance and the associated return paths, one must consider the on-chip conductors as well as the system packaging. This can result in an RLC circuit size that is impractical for traditional simulators. In this paper we demonstrate a localized, window-based extraction and simulation methodology that employs the recently proposed susceptance (the inverse of inductance matrix) concept. We provide a qualitative explanation for the efficacy of this approach, and demonstrate how it facilitates pre-manufacturing simulations that would otherwise be intractable. A critical aspect of this simulation efficiency is owed to a susceptance-based circuit formation that we prove to be symmetric positive definite. This property, along with the sparsity of the susceptance matrix, enables the use of some advanced sparse matrix solvers. lye demonstrate this extraction and simulation methodology on some industrial examples.
Lawrence T. Pileggi, Michael W. Beattie, Byron Krauter
DATE2
2002 Modular, Fabric-Specific Synthesis for Programmable Architectures
Aneesh Koorapaty, Lawrence T. Pileggi
FPL2
2002 Throughput-driven IC communication fabric synthesis
abstract
As the scale of system integration continues to grow, the on-chip communication becomes the ultimate bottleneck of system performance and the primary determinant of system architecture. In this paper we propose a throughput-driven synthesis methodology for on-chip communication fabrics based on optimized bus models. Compared with traditional delay-driven, wire-by-wire planning methods, the throughput-driven methodology provides a feasible and accurate system-level solution to address delay and congestion problems simultaneously during earlyphase design planning. Unlike the conventional methods which are based on rather inaccurate RC models and simplistic delay metrics, in our methodology the communication fabrics are characterized in terms of realistic Partial Element Equivalent Circuits (PEEC) extracted from the multi-layer interconnects and transistor level transient analysis via SPICE-like tools. The characterized models facilitate a flexible interconnect fabric optimization engine that can be embedded into a system planner for throughput-driven synthesis. Furthermore, engineering trade-offs considering repeater area and interconnect power consumption are further considered as part of this methodology.
Lawrence T. Pileggi
ICCAD2
2002 Robust and passive model order reduction for circuits containing susceptance elements
abstract
Numerous approaches have been proposed to address the overwhelming modeling problems that result from the emergence of magnetic coupling as a dominant performance factor for ICsand packaging. Firstly, model order reduction (MOR) methods have been extended to robustly capture very high frequency behaviors for large RLC systems via methods such as PRIMA[8] with guaranteed passivity. In addition, new models of the magnetic couplings in terms of susceptance (inverse of inductance) have shown great promise for robust sparsification of otherwise intractable inductance coupling-matrix problems [3--5]. However, model order reduction via PRIMA for circuits that include susceptance elements does not guarantee passivity. Moreover, susceptance elements are incompatible with the path tracing algorithms that provide the fundamental runtime efficiency of RICE [10]. In this paper a novel MOR algorithm, SMOR, is proposed as an extension of ENOR [11] which exploits the matrix properties of susceptance-based circuits for runtime efficiency, and provides for a numerically stable, provably passive MOR using a new or-thonormalization strategy.
Lawrence T. Pileggi
ICCAD2
2002 Understanding and addressing the impact of wiring congestion during technology mapping
abstract
Traditionally, interconnect effects are taken into account during logic synthesis via wireload models, but their ineffectiveness for DSM technologies has been demonstrated and various physical synthesis approaches have been spawned to address the problem. Of particular interest is that logic block size is no longer dictated exclusively by total cell area, yet synthesis optimization objectives are aimed specifically at minimizing the number and size of cells. Methodologies that incorporate congestion within the logic synthesis objective function have been proposed in [9][10][11] and [15]; however, as we will demonstrate, predicting the true congestion prior to layout is not possible, and the efficacy of any approach can only be evaluated after routing is completed within the fixed die size. In this paper we propose a practical, complete methodology which first performs congestion-aware technology mapping using a global weighting factor for the cost function [15], and then applies incremental localized unmapping and remapping on congested areas. This complete approach addresses the problem that one global factor is not ideally suited for all regions of the designs. Most importantly, through the application of this methodology to industrial examples we will show that any attempt at a purely top-down single-pass congestion-aware technology mapping is merely wishful thinking.
Davide Pandini, Lawrence T. Pileggi, Andrzej J. Strojwas
ISPD2
2002 TETA: transistor-level waveform evaluation for timing analysis
abstract
Static timing analysis breaks down the longest path problem into waveform analysis of paths of logic stages that are comprised of nonlinear transistors and complex RLC loads. Runtime efficiency is of the utmost importance; however, the waveform evaluation of these logic stages cannot be accelerated via timing simulation algorithms that attempt to exploit temporal or spatial latency since the simulation problem is already a partitioned one. TETA was developed as a general purpose transistor-level waveform evaluation engine for providing accuracy-efficiency tradeoffs for these logic-stage waveform evaluation problems that are encountered during timing analysis. Of particular emphasis are the large RC(L) coupled logic stages which present the bottleneck for waveform evaluation along multiple stages of a digital circuit path. TETA applies a novel compaction scheme for the logic-stage transistor clusters and employs a novel nonlinear algebraic solution method to analyze the circuit. Importantly, stability of the waveform evaluation with TETA requires only stable single-input multi-output N-port interconnect models that are not necessarily passive. Waveform evaluators that use general transistor and piecewise linear device models require provably passive multi-input multi-output interconnect models that can be extremely inefficient for large coupled N-port problems. Furthermore, the methodology in TETA brings extra efficiency by avoiding extra matrix factorizations and enabling the use of device model tables without any loss of accuracy. Complex logic gates and nonlinear capacitors are handled without loss of generality.
Emrah Acar, Florentin Dartu, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2002 An analysis of the wire-load model uncertainty problem
abstract
Traditional integrated-circuit (IC) design methodologies have used wire-load models during logic synthesis to estimate the expected impact of the metal wiring on the gate delays. These models are based on wire-length statistics from legacy designs to facilitate a top-down IC design flow process. Recently, there has been increased concern regarding the efficacy of wire-load models as deep-submicrometer (DSM) interconnect parasitics begin to dominate the delay of digital IC logic gates. Some technology projections (Sylvester and Keutzer, 1998) have suggested that wire-load models will remain effective to block sizes on the order of 50 000 gates. This suggests that existing top-down synthesis methodologies will not have to be changed substantially since this is approximately the maximum size for which logic synthesis is effective. However, our analyses on production designs show that the problem is not quite so straightforward and the efficacy of synthesis using wire-load models depends upon technology data as well as specific characteristics of the design and the granularity of available physical information. We analyze these effects and dependencies in detail in this paper and draw some conclusions regarding the future challenges associated with top-down IC design and block synthesis, in particular, in the DSM design era.
Padmini Gopalakrishnan, Altan Odabasioglu, Lawrence T. Pileggi, Salil Raje
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2002 On-chip induction modeling: basics and advanced methods
abstract
Modeling magnetic interactions for on-chip interconnect has become an issue of great interest for integrated circuit design in recent years. This paper describes the basic concepts of magnetic interaction, loop and partial inductance, along with some of the high frequency effects such as skin and proximity effect. We also discuss and contrast options for stable and accurate window-based extraction of large-scale magnetic coupling. We analyze the required window sizes to consider the possibilities for pattern-matching style solutions, and propose three schemes for determining coupling values and window sizing for extraction via on-the-fly field solution.
Michael W. Beattie, Lawrence T. Pileggi
IEEE Trans. Very Large Scale Integr. Syst.2
2001 False Coupling Interactions in Static Timing Analysis
abstract
Neighboring line switching can contribute to a large portion of the delay of a line for today's deep submicron designs. In order to avoid excessive conservatism in static timing analysis, it is important to determine if aggressor lines can potentially switch simultaneously with the victim. In this paper, we present a comprehensive ATPG-based approach that uses functional information to identify valid interactions between coupled lines. Our algorithm accounts for glitches on aggressors that can be caused by static and dynamic hazards in the circuit. We present results on several benchmark circuits that show the value of considering functional information to reduce the conservatism associated with worst-case coupled line switching assumptions during static timing analysis.
Ravishankar Arunachalam, R. D. (Shawn) Blanton, Lawrence T. Pileggi
DAC3
2001 Inductance 101: Modeling and Extraction
abstract
Modeling magnetic interactions for on-chip interconnect has become an issue of great interest for inte-grated circuit design in recent years. This tutorial paper de-scribes the basic concepts of magnetic interaction, loop and partial inductance, along with some of the high frequency ef-fects such as skin and proximity effect.
Michael W. Beattie, Lawrence T. Pileggi
DAC2
2001 Modeling Magnetic Coupling for On-Chip Interconnect
abstract
As advances in IC technologies and operating frequencies make the modeling of on--chip magnetic interactions a necessity, it is apparent that extension of traditional inductance extraction approaches to full-chip scale problems is impractical. There are primarily two obstacles to performing inductance extraction with the same efficacy as full-chip capacitance extraction: 1) neglecting far-away coupling terms can generate an unstable inductance matrix approximation; and 2) the penetrating nature of inductance makes localized extraction via windowing extremely difficult. In this paper we propose and contrast three new options for stable and accurate window-based extraction of large-scale magnetic coupling. We analyze the required window sizes to consider the possibilities for pattern-matching style solutions, and propose three schemes for determining coupling values and window sizing for extraction via on-the-fly field solution.
Michael W. Beattie, Lawrence T. Pileggi
DAC2
2001 Min/max On-Chip Inductance Models and Delay Metrics
abstract
This paper proposes an analytical inductance extraction model for characterizing min/max values of typical on-chip global intercon-nect structures, and a corresponding delay metric that can be used to provide RLC delay prediction from physical geometries. The model extraction and analysis is efficient enough to be used within optimization and physical design exploration loops. The analytical min/max inductance approximations also provide insight into the effects caused by inductances.
Yi-Chang Lu, Mustafa Celik, Tak Young, Lawrence T. Pileggi
DAC4
2001 Efficient inductance extraction via windowing
abstract
We propose a new, efficient and accurate localized inductance modeling technique via windowing in a manner that is analogous to localized capacitance extraction. The stability and accuracy of this process is made possible by twice inverting the localized inductance models, and in the process exploiting properties of the magnetostatic interactions as modeled via the susceptance (inverse inductance). Application of these localized double-inverse inductance models to actual IC bus examples demonstrates the significant improvement in simulation efficiency and overall accuracy as compared to alternative methods of approximation and simplification.
Michael W. Beattie, Lawrence T. Pileggi
DATE2
2001 Overcoming wireload model uncertainty during physical design
abstract
The advent of deep sub-micron technologies has created a number of problems for existing design methodologies. Most prominent among them is the problem of timing closure, whereby design time is dramatically increased due to iterations between gate-level synthesis and physical design. It is well known that the heart of this problem lies in the use of wireload models based on wirelength statistics from legacy designs. Some technology projections in have suggested that wireload models will remain effective to block sizes on the order of 50k gates. This suggests that synthesis will not have to be changed much since this is approximately the maximum size for which logic synthesis is effective. However, our analyses on production designs show that the problem is not quite so straightforward, and the efficacy of synthesis using wireload models depends upon technology data as well as specific characteristics of the design. We analyze these effects and dependencies in detail in this paper, and draw some conclusions about the amount of physical information that is required for synthesis to be effective. Finally, we discuss the implications on hierarchical design flows, and propose a solution via physical prototyping.
Padmini Gopalakrishnan, Altan Odabasioglu, Lawrence T. Pileggi, Salil Raje
ISPD3
2001 RC(L) interconnect sizing with second order considerations via posynomial programming
abstract
There has been substantial work on interconnect sizing algorithms for delay and area optimization in terms of the Elmore delay. Recently, however, signal integrity issues have become of equal or greater importance than delay and area for deep submicron designs. Modeling signal integrity requires more than the Elmore delay approximation, especially when interconnect inductance effects are considered. This paper studies a new interconnect sizing formula?tion with signal attenuation and transition time constraints that cap?tures the same global optimality as the Elmore delay based approaches. With the signal attenuation (or the signal transition time) modeled by the second order central moment of the circuit response, we formulate a provably posynomial optimization prob?lem for RC trees such that the well studied algorithms for geometric programming can be applied with guaranteed convergence to a glo?bal minima. For RCL cases we demonstrate that this formulation remains convex and posynomial under reasonable conditions. Suffi?cient conditions are given in terms of the technology parameters and termination conditions.
Lawrence T. Pileggi
ISPD2
2001 Limitations and challenges of computer-aided design technology for CMOS VLSI
abstract
As manufacturing technology moves toward fundamental limits of silicon CMOS processing, the ability to reap the full potential of available transistors and interconnect is increasingly important. Design technology (DT) is concerned with the automated or semi-automated conception, synthesis, verification, and eventual testing of microelectronic systems. While manufacturing technology faces fundamental limits inherent in physical laws or material properties, design technology faces fundamental limitations inherent in the computational intractability of design optimizations and in the broad and unknown range of potential applications within various design processes. In this paper, we explore limitations to how design technology can enable the implementation of single-chip microelectronic systems that take full advantage of manufacturing technology with respect to such criteria as layout density performance, and power dissipation.
Randal E. Bryant, Kwang-Ting Cheng, Andrew B. Kahng, Kurt Keutzer, Wojciech Maly, A. Richard Newton, Lawrence T. Pileggi, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli
Proc. IEEE7
2001 Equipotential shells for efficient inductance extraction
abstract
To make three-dimensional (3-D) on-chip interconnect inductance extraction tractable, it is necessary to ignore parasitic couplings without compromising critical properties of the interconnect system. It is demonstrated that simply discarding faraway mutual inductance couplings can lead to an unstable approximate inductance matrix. In this paper, we describe an equipotential shell methodology, which generates a partial inductance matrix that is sparse yet stable and symmetric. We prove the positive definiteness of the resulting approximate inductance matrix when the equipotential shells are properly defined. Importantly, the equipotential shell approach also provably preserves the inductance of loops if they are enclosed entirely within the shells of their segments. Methods for sizing the shells to control the accuracy are presented. To demonstrate the overall efficacy for on-chip extraction, ellipsoid shells, which are a special case of the general equipotential shell approach, are presented and demonstrated for both on-chip and system-level extraction examples.
Michael W. Beattie, Byron Krauter, Lale Alatan, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2000 TACO: timing analysis with coupling
abstract
The impact of coupling capacitance on delay is usually estimated by scaling the coupling capacitances (often by a factor of 2) and modeling them as grounded. This simple approach has been shown to be overly pessimistic in some cases, while somewhat optimistic in others. This paper introduces TACO, a timing analysis methodology that produces tight bounds on worst- and best-case timing for circuits with dominant coupling capacitance. The methodology utilizes a coupled Ceff gate model for capturing the provably worst- and best-case delays as a function of the timing-window inputs to the gates.
Ravishankar Arunachalam, Karthik Rajagopal, Lawrence T. Pileggi
DAC3
2000 Design closure (panel session): hope or hype?
abstract
It's been one year since Richard Goering told us that the EDA RTL-to-GDSII world was about to change. What have we learned? Who's winning, and who's not? This panel, consisting of the leading large and upcoming players in this space, will deliver concrete data to differentiate leading approaches to achieving design closure. Does the solution lie in raw speed and RTL optimization, with the synthesis-place-route back end just a commodity? Does the solution lie in new metrics for design convergence, and symmetric multiprocessing platforms for efficiency? Does the solution lie in a holistic, unified architecture of data model and tools? Or does the solution lie in extensions and unifications of existing production-proven logic, timing, and layout optimization technologies? A hard-hitting panel session will reveal the answers!
Raúl Camposano, Jacob Greidinger, Patrick Groeneveld, Michael Jackson 0004, Lawrence T. Pileggi, Louis K. Scheffer
DAC5
2000 Impact of interconnect variations on the clock skew of a gigahertz microprocessor
abstract
Due to the large die sizes and tight relative clock skew margins, the impact of interconnect manufacturing variations on the clock skew in today's gigahertz microprocessors can no longer be ignored. Unlike manufacturing variations in the devices, the impact of the interconnect manufacturing variations on IC timing performance cannot be captured by worst/best case corner point methods. Thus it is difficult to estimate the clock skew variability due to interconnect variations. In this paper we analyze the timing impact of several key statistically independent interconnect variations in a context-dependent manner by applying a previously reported interconnect variational order-reduction technique. The results show that the interconnect variations can cause up to 25% clock skew variability in a modern microprocessor design.
Sani R. Nassif, Lawrence T. Pileggi, Andrzej J. Strojwas
DAC3
2000 Hierarchical Interconnect Circuit Models
abstract
The increasing size of integrated systems combined with deep submicron physical modeling details creates an explosion in RLC interconnect modeling complexity of unmanageable proportions. Interconnect extraction tools employ hierarchy to manage complexity, but this hierarchy is discarded via eliminating far away coupling terms when the equivalent ARC circuits are Formed. The increasing dominance of capacitance coupling along with the emergence of on chip inductance, however, makes the composite effect of far-away couplings increasingly evident. Even if newly enforced design rules and practices will ultimately obviate the need for modeling these couplings for design verification, some approximation of the "exact" solution is required to validate these rules. This paper proposes an efficient hierarchical equivalent circuit representation of interconnect parasitics that utilizes the efficient hierarchical long-distance modeling already existing within extractors. Results from a prototype simulator based on these hierarchical models demonstrate the simulation inaccuracy incurred when the faraway coupling terms are ignored. Such a form of interconnect modeling may provide the key to hierarchical modeling of electro-magnetic interactions between large components on future gigascale systems.
Michael W. Beattie, Satrajit Gupta, Lawrence T. Pileggi
ICCAD3
1999 IC Analyses Including Extracted Inductance Models
abstract
Article IC analyses including extracted inductance models Share on Authors: Michael W. Beattie Carnegie Mellon University, Dept. of ECE, 5000 Forbes Ave., Pittsburgh, PA Carnegie Mellon University, Dept. of ECE, 5000 Forbes Ave., Pittsburgh, PAView Profile , Lawrence T. Pileggi Carnegie Mellon University, Dept. of ECE, 5000 Forbes Ave., Pittsburgh, PA Carnegie Mellon University, Dept. of ECE, 5000 Forbes Ave., Pittsburgh, PAView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 915–920https://doi.org/10.1145/309847.310098Online:01 June 1999Publication History 16citation337DownloadsMetricsTotal Citations16Total Downloads337Last 12 Months1Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Michael W. Beattie, Lawrence T. Pileggi
DAC2
1999 Model Order-Reduction of RC(L) Interconnect Including Variational Analysis
abstract
Article Model order-reduction of RC(L) interconnect including variational analysis Share on Authors: Ying Liu Department of Electrical and Computer Engineering, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA Department of Electrical and Computer Engineering, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile , Lawrence T. Pileggi Department of Electrical and Computer Engineering, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA Department of Electrical and Computer Engineering, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile , Andrzej J. Strojwas Department of Electrical and Computer Engineering, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA Department of Electrical and Computer Engineering, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 201–206https://doi.org/10.1145/309847.309914Online:01 June 1999Publication History 44citation390DownloadsMetricsTotal Citations44Total Downloads390Last 12 Months14Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Lawrence T. Pileggi, Andrzej J. Strojwas
DAC2
1999 S2P: A Stable 2-Pole RC Delay and Coupling Noise Metric
abstract
The Elmore delay is the metric of choice for performance-driven design applications due to its simple, explicit form and ease with which sensitivity information can be calculated. However, for deep submicron technologies, the accuracy of the Elmore delay is insufficient. In this paper we formulate a delay model using a provably stable two pole waveform response that provides a unique mapping between four moments and a specific delay value. Unlike traditional moment matching, this two-pole model permits us to precharacterize the delays, and store them in a table, as a mapped function of three parameters. The model also provides an explicit expression for the peak noise induced on a coupled line as a function of the same three moments. The results indicate runtimes comparable to an Elmore delay calculation but with the accuracy of an AWE approximation.
Emrah Acar, Altan Odabasioglu, Mustafa Celik, Lawrence T. Pileggi
Great Lakes Symposium on VLSI4
1999 Electromagnetic parasitic extraction via a multipole method with hierarchical refinement
abstract
The increasing interconnect density and operating frequencies of system-on-a-chip (SOC) designs necessitates extraction of parasitic electromagnetic couplings beyond the localized confines of functional design blocks. In addition, SOC design styles and gridless variable-width routing make it increasingly difficult to use precharacterized library shapes for parasitic extraction. A comprehensive capacitance and inductance extraction solution requires a hierarchical data representation and fast runtime algorithms. We illustrate through examples that both the multipole method and hierarchical refinement, which are the two most successful approaches for parasitic extraction to date, work efficiently only under certain, limiting conditions. To improve this situation we present an approach which combines the best of both methods into a concurrent multipole refinement representation of the electromagnetic interaction which is efficient for arbitrary interconnect configurations. We use a generalized formulation of electromagnetic interactions to exploit the similarities in capacitance and inductance extraction for greater efficiency.
Michael W. Beattie, Lawrence T. Pileggi
ICCAD2
1999 Practical considerations for passive reduction of RLC circuits
abstract
Krylov space methods initiated a new era for RLC circuit model order reduction. Although theoretically well-founded, these algorithms can fail to produce useful results for some types of circuits. In particular controlling accuracy and ensuring passivity are required to fully utilize these algorithms in practice. In this paper we propose a methodology for passive reduction of RLC circuits based on extensions of PRIMA, that is both broad and practical. This work is made possible by uncovering the algebraic connections between this passive model order reduction algorithm and other Krylov space methods. In addition, a convergence criteria based on an error measure for PRIMA is presented as a first step towards intelligent order selection schemes. With these extensions and error criterion examples demonstrate that accurate approximations are possible well into the RF frequency range even with expansions about s=0.
Altan Odabasioglu, Mustafa Celik, Lawrence T. Pileggi
ICCAD3
1999 Error bounds for capacitance extraction via window techniques
abstract
The overwhelming size of the capacitance extraction problem forces designers to localize the capacitive coupling and determine a distance (a "window") outside of which the mutual capacitance between two wires is "small enough" to ignore. The primary difficulties with such approaches are determining how large the extraction windows have to be to capture all of the relevant mutual capacitances, and estimating the error incurred due to "windowing." This paper proposes solutions for both problems. We first show that the shift-truncate method and the windowing method yield opposite bounds for the exact values of the mutual and self capacitances. It is also shown that the capacitance matrices resulting from the application of these two localization methods are positive definite and, therefore, lead to stable approximations of the exact parasitics system. For the windowing method, we show that the original asymmetric capacitance matrix can be made symmetric while guaranteeing the positive definiteness and making the error bounds even tighter. In summary, we describe an adaptive window sizing methodology based on error values from the windowing and shift-truncate bounds. The proposed methodology is also potentially useful in identifying crosstalk problem zones for interconnect optimization and noise reduction, and for the generation of noise-reducing design rules.
Michael W. Beattie, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1999 Metrics and bounds for phase delay and signal attenuation in RC(L)clock trees
abstract
As IC clock frequencies approach the GHz range, the distribution of the clock signals becomes more critical in terms of controlling both skew and signal attenuation. Moreover, inductance effects are evident since RC transmission lines will overly attenuate these high-frequency clock signals. To facilitate accurate optimization of clock tree performance and skew requires simple metrics which capture these high-frequency effects. In this paper, we derive simple metrics and bounds for the phase delay and the attenuation of a periodic [RC(L)] tree response as a function of the fundamental frequency of the clock signal. These metrics are based on the first two moments of the impulse response, and are shown to further provide a mechanism for control of underdamped responses (reflections). An important result of this work is the clear demonstration that once the attenuation of the clock signal is controlled, the phase delay can be accurately captured in terms of the first-moment. Furthermore, the form of these metrics and their relationship to one another provides an excellent foundation for various forms of clock tree optimization.
Mustafa Celik, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 TETA: Transistor-Level Engine for Timing Analysis
abstract
TETA is an interconnect-centric waveform calculator that was optimized to achieve the utmost efficiency for analyzing logic stages comprised of transistors and large coupled RC(L) interconnect models. TETA applies a novel compaction for the transistor clusters and employs successive chord iterations to solve the resulting nonlinear equations. These algorithms permit the use of simple SIMO (single input multi-output) N-port interconnect models since macromodel passivity is not required. The successive chord analysis also enables TETA to avoid the N-port matrix factorization during nonlinear iterations and allows the use of simple table look-up models for MOS devices. Complex gates and nonlinear capacitors can be handled without loss of generality.
Florentin Dartu, Lawrence T. Pileggi
DAC2
1998 PRIMO: Probability Interpretation of Moments for Delay Calculation
abstract
Moments of the impulse response are widely used for interconnect delay analysis, from the explicit Elmore delay (first moment of the impulse response) expression, to moment matching methods which create reduced order transimpedance and transfer function approximations. However, the Elmore delay is fast becoming ineffective for deep submicron technologies, and reduced order transfer function delays are impractical for use as early-phase design metrics or as design optimization cost functions. This paper describes an approach for fitting moments of the impulse response to probability density functions so that delays can be estimated from probability tables. For RC trees it is demonstrated that the incomplete gamma function provides a provably stable approximation. The step response delay is obtained from a one-dimensional table lookup.
Rony Kay, Lawrence T. Pileggi
DAC2
1998 ftd: An Exact Frequency to Time Domain Conversion for Reduced Order RLC Interconnect Models
abstract
Recursive convolution provides an exact solution for interfacing reduced-order frequency domain representations with discrete time domain models of piecewise linear voltage waveforms. The state-space method is more efficient, but not exact, and can sometimes produce large time domain errors. This paper presents a new algorithm, ftd (frequency to time domain), for incorporating linear frequency domain macro-models into time domain simulators. ftd provides accuracy equivalent to recursive convolution with efficiency that is superior to the state-space methods.
Lawrence T. Pileggi, Andrzej J. Strojwas
DAC2
1998 Determination of worst-case aggressor alignment for delay calculation
abstract
Increases in delay due to coupling can have a dramatic impact on IC performance for deep submicron technologies. To achieve maximum performance there is a need for analyzing logic stages with large complex coupled interconnects. In timing analysis, the worst-case delay of gates along a critical path must include the effect of noise due to switching of nearby aggressor gates. In this paper, we propose a new waveform iteration strategy to compute the delay in the presence of coupling and to align aggressor inputs to determine the worst-case victim delay. We demonstrate the application of our methodology at both the transistor-level and celllevel. In addition, we prove that the waveforms generated in our methodology converge under typical timing analysis conditions. 1.
Paul D. Gross, Ravishankar Arunachalam, Karthik Rajagopal, Lawrence T. Pileggi
ICCAD4
1998 h-gamma: an RC delay metric based on a gamma distribution approximation of the homogeneous response
abstract
Recently a probability interpretation of moments wm proposed as a compromise between the Elmore delay and higher or&r mommt matching for RC timing estimation [5].By modeling RC impufses as tinze-sht~ed incomplete gamma distribution functions, the delays could be obtained via table 100tip using a gamma integral table and the first three moments of the imptise response.However, while this approximation worti well for many examples, it struggles with responses when the metal resistance becomes dominant, andproduces results with impractical h.meshz~values.In this paper the probability interpretation is etiended to the circuit homogeneous response, without requiring the time shi$ parameter The gamma distribution is used to characterize the normalized homogeneous portion of the step response.For a generalized RC interconnect model (RC tree or mesh), the stability of the holtlogetleol{s-gml?ta distribution model is guaranteed.It is demonstrated that when a table model is carefilly constructed the hgamma approxinrationprovides for exellent improvemeti over the Elmore delay in terms of accur~, with ve~little additiond cost in terms of CPU time.
Emrah Acar, Lawrence T. Pileggi
ICCAD3
1998 Timing metrics for physical design of deep submicron technologies
abstract
Performance-driven physical design is becoming more important as advances in IC technologies enable gigahertz operating frequencies. These same IC technologies, however, exhibit dominant interconnect resistance, non-negligible coupling capacitance, and even the potential for inductance effects, which makes the performance modeling and prediction more difficult. In this tutorial paper we will overview some of the existing timing metrics that are suitable for use during physical design, and introduce new metrics and directions for future work.
Lawrence T. Pileggi
ISPD1
1998 EWA: efficient wiring-sizing algorithm for signal nets and clock nets
abstract
The wire sizing problem under inequality Elmore delay constraints is known to be posynomial, hence convex under an exponential variable transformation. Due to their efficiency and ease of implementation, one-wire-at-a-time downhill improvement heuristics are often applied to solve such problems. Unfortunately, when there are complex boundary constraints, the solutions from such heuristics can be far away from the global minimum. There are formal methods for solving convex programs, but they are too costly in terms of runtime for some applications. Some optimization techniques can be quite efficient but they solve less desirable formulations, such as minimum weighted sum of area and critical path delays. This paper proposes an efficient wire-sizing algorithm (EWA) that is able to trade solution accuracy for time efficiency while providing an upper bound on the distance from the optimal solution. EWA solves the practical problem of minimizing the total wiring area or the capacitance of an interconnect RC tree subject to hard constraints on the Elmore delay. The implementation is simple and efficiency is comparable to the available heuristics. No restrictions are placed on the circuit or wire widths. Furthermore, it is shown that the optimal wire width assignment for a minimum wiring area objective satisfies all the delay constraints as equalities when minimum wire width constraints are not active. It follows that EWA can be applied for problems with equality delay constraints such as clock trees. Moreover, this and other properties are general enough to permit extensions to higher order delay models and can be used to enhance other optimization methods.
Rony Kay, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 PRIMA: passive reduced-order interconnect macromodeling algorithm
abstract
This paper describes an algorithm for generating provably passive reduced-order N-port models for RLC interconnect circuits. It is demonstrated that, in addition to macromodel stability, macromodel passivity is needed to guarantee the overall circuit stability once the active and passive driver/load models are connected. The approach proposed here, PRIMA, is a general method for obtaining passive reduced-order macromodels for linear RLC systems. In this paper, PRIMA is demonstrated in terms of a simple implementation which extends the block Arnoldi technique to include guaranteed passivity while providing superior accuracy. While the same passivity extension is not possible for MPVL, comparable accuracy in the frequency domain for all examples is observed.
Altan Odabasioglu, Mustafa Celik, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1998 Analytic termination metrics for pin-to-pin lossy transmission lines with nonlinear drivers
abstract
A significant percentage of the critical nets in high-performance systems are of the pin-to-pin type. To optimally design these nets such that signal integrity is preserved, efficient analytical metrics for transmission line termination are a valuable part of a system-level designer's toolset. Using the symbolic moment-based expressions in this paper, proper termination can be determined via a single-step procedure, without any preprocessing steps and/or time-domain simulations. Driver nonlinearities and effects of nonzero rise-time are also considered in the proposed termination methodology.
Rohini Gupta, John Willis, Lawrence T. Pileggi
IEEE Trans. Very Large Scale Integr. Syst.3
1997 Bounds for BEM Capacitance Extraction
abstract
In this paper we prove that simply discarding conductors beyonda certain spacing during BEM capacitance extraction willresult in a lower bound on the self-capacitance calculations andan upper bound on the mutual capacitance calculations that liewithin that spacing. We prove that a potential-shift and truncatescheme can yield bounds opposite to those for the truncate onlycase; namely, an upper bound on the self capacitance and a lowerbound on the mutual capacitance that lies within the chosenspacing. The ease with which the upper and lower bounds arecalculated is shown, and their utility for selection of an optimalwindow size is described. A metal shell is also presented here thatresults in bounds similar to those of shift-truncate. We furtherpropose a new potential-shift function that yields increased approximationaccuracy compared to shift-truncate in many cases.
Michael W. Beattie, Lawrence T. Pileggi
DAC2
1997 Calculating Worst-Case Gate Delays Due to Dominant Capacitance Coupling
abstract
In this paper we develop a gate level model that allows us to determine the best and worst case delay when there is dominant interconnect coupling. Assuming that the gate input windows of transition are known, the model can predict the worst and best case noise, as well as the worst and best case impact on delay. This is done in terms of a Ceff based gate model under general RC interconnect loading conditions. I. INTRODUCTION As IC dimensions scale to the deep submicron range, their multi-level interconnects are constructed such that the coupling capacitance becomes the dominant component of load capacitance. This effect is largely the result of the increased ratio between the lateral and the vertical capacitance of the line. The increased number of metal layers is the other source of coupling capacitance problems, since there is a reduced likelihood of a nearby "ground plan." The lateral capacitance is increased by the relative increase in the metal thickness with respect to line sp...
Florentin Dartu, Lawrence T. Pileggi
DAC2
1997 SPIE: Sparse Partial Inductance Extraction
abstract
Extracting the inductance of complex interconnect topologiesis a formidable task, and simulating the resulting dense partialinductance matrix is even more difficult. Furthermore, it is wellknown that simply discarding smallest terms to sparsify the inductancematrix can render the partial inductance matrix indefiniteand result in an unstable circuit model. In this paper, wedescribe a methodology for incrementally generating a sparsepartial inductance matrix based on using moments about s=0 todetermine when a sufficient number of mutual inductances havebeen captured. The minimally required mutual inductances areextracted for a provably stable model.
Zhijiang He, Mustafa Celik, Lawrence T. Pileggi
DAC3
1997 A hierarchical decomposition methodology for multistage clock circuits
abstract
This paper describes a novel methodology to automate the design of the interconnect distribution for multistage clock circuits. We introduce two key ideas. First, a hierarchical decomposition of the layout divides the problem into a set of local Steiner-wired latch clusters (to minimize and balance local capacitance) fed globally by a balanced binary tree (to maximize performance). Second, we recast the global clock distribution problem as a simultaneous optimization of clock topology, clock segment routing, wire sizing and buffering. The hierarchical decomposition reduces the problem complexity and allows use of more aggressive optimization techniques. Integration of the geometric and electrical optimizations likewise allows more aggressive performance goals. Experiments with an industrial design comprising over 16,000 latches demonstrate the efficiency of the approach: a complete clock distribution solution met a 200-MHz cycle time specification with only 310 ps of skew, met strict current density constraints, exhibited good delay matching across uniform wire width and device variations, and was completed in under 10 CPU hours.
Gary Ellis, Lawrence T. Pileggi, Rob A. Rutenbar
ICCAD2
1997 PRIMA: passive reduced-order interconnect macromodeling algorithm
abstract
This paper describes PRIMA, an algorithm for generating provably passive reduced order N-port models for RLC interconnect circuits. It is demonstrated that, in addition to requiring macromodel stability, macromodel passivity is needed to guarantee the overall circuit stability once the active and passive driver/load models are connected. PRIMA extends the block Arnoldi technique to include guaranteed passivity. Moreover, it is empirically observed that the accuracy is superior to existing block Arnoldi methods. While the same passivity extension is not possible for MPVL, we observed comparable accuracy in the frequency domain for all examples considered. Additionally a path tracing algorithm is used to calculate the reduced order macromodel with the utmost efficiency for generalized RLC interconnects.
Altan Odabasioglu, Mustafa Celik, Lawrence T. Pileggi
ICCAD3
1997 CMOS Gate Delay Models for General RLC Loading
abstract
Gate and cell level timing analysis remains popular yet inherently incompatible with RC and RCL interconnect loads. The Ceff concept was proposed in Qian et. al. (1994) to model the interaction of empirical gate/cell delay models and RC loads. The most efficient Ceff model works in terms of precharacterizing the parameters of a time varying Thevenin voltage source model (in series with a fixed resistor) over a wide range of effective capacitance load values. In this paper we generalize this Thevenin equivalent Ceff model to enable future technologies which may include reduced supply voltages and RCL loads, without further complicating the Ceff algorithm or iterations.
Ravishankar Arunachalam, Florentin Dartu, Lawrence T. Pileggi
ICCD3
1997 Clustering and Load Balancing for Buffered Clock Tree Synthesis
abstract
Buffers in clock trees introduce two additional sources of skew: the first source of skew is the effect of process variations on buffer delays. The second source of skew is the imbalance in buffer loading. We propose a buffered clock tree synthesis methodology whereby we first apply a clustering algorithm to obtain clusters of approximately equal capacitance loading. We drive each of these clusters with identical buffers. A sensitivity based approach is then used for equalizing the Elmore delay from the buffer output to all of the clock nodes. The skew due to load imbalance is minimized concurrently by matching a higher-order model of the load by wire sizing and wire lengthening. We demonstrate how this algorithm can be used recursively to generate low-skew buffered clock trees.
Ashih D. Mehta, Yao-Ping Chen, Noel Menezes, Martin D. F. Wong, Lawrence T. Pileggi
ICCD5
1997 EWA: exact wiring-sizing algorithm
abstract
The wire sizing problem under inequality Elmore delay constraints is known to be posynomial, hence convex under an exponential variable-transformation. There are formal methods for solving convex programs. In practice heuristics are often applied because they provide good approximations while offering simpler implementation and better efficiency. There are methods for solving related problems, which are comparable to heuristics from efficiency point of view, but they solve a less desirable formulation in terms of the objective function and constraints. In this paper the EWA algorithm is described. It solves the problem of minimizing the wiring area or capacitance of an interconnect tree subject to constraints on the Elmore delay. EWA is simple to implement and its efficiency is comparable to the available heuristics. No restrictions are placed on the circuit or wire widths, e.g., non-monotone wire widths assignment solutions are feasible. We prove that the optimal wire width assignment for a minimum wiring area objective satisfies all the delay constraint as equalities, when minimum wire width constraints are relaxed. It follows that EWA can be applied also for problems with equality delay constraints such as clock trees. This and other described properties are general enough to permit extensions to higher order delay models in the future.
Rony Kay, Gennady Bucheuv, Lawrence T. Pileggi
ISPD3
1997 Transmission line synthesis via constrained multivariable optimization
abstract
The design of system level interconnects to meet signal integrity objectives is a challenging problem. This paper formulates the transmission line synthesis problem as a constrained multidimensional optimization task for performance-driven routing of the complete net, taking into account factors like loading conditions on the line, loss in the line, and rise time of the input signal. Different design variables, such as width or resistivity of the interconnect, resistive source, or far-end termination, etc., can all be considered concurrently. A novel termination metric is described that is based upon forcing the impulse response waveform to be symmetric using the first three moments of the distributed system response. This metric provides an efficient means to trade off between signal rise time and ringing without requiring a time-domain simulation. Several examples are presented to demonstrate the efficacy of the proposed methodology.
Rohini Gupta, Byron Krauter, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1997 The Elmore delay as a bound for RC trees with generalized input signals
abstract
The Elmore delay is an extremely popular timing-performance metric which is used at all levels of electronic circuit design automation, particularly for resistor-capacitor (RC) tree analysis. The widespread usage of this metric is mainly attributable to it being a delay measure that is a simple analytical function of the circuit parameters. The only drawback to this delay metric is the uncertainty of its accuracy and the restriction to it being an estimate only for the step response delay. In this paper, we prove that the Elmore delay measure is an absolute upper bound on the actual 50% delay of an RC tree response. Moreover, we prove that this bound holds for input signals other than steps and that the actual delay asymptotically approaches the Elmore delay as the input signal rise time increases. A lower bound on the delay is also developed using the Elmore delay and the second moment of the impulse response. The utility of this bound is for understanding the accuracy and the limitations of the Elmore metric as we use it as a performance metric for design automation.
Rohini Gupta, Bogdan Tutuianu, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1997 A sequential quadratic programming approach to concurrent gate and wire sizing
abstract
With an ever-increasing portion of the delay in high-speed CMOS chips attributable to the interconnect, interconnect-circuit design automation continues to grow in importance. By transforming the gate and multilayer wire sizing problem into a convex programming problem for the Elmore delay approximation, we demonstrate the efficacy of a sequential quadratic programming (SQP) solution method. For cases where accuracy greater than that provided by the Elmore delay approximation is required, we apply SQP to the gate and wire sizing problem with more accurate delay models. Since efficient calculation of sensitivities is of paramount importance during SQP, we describe an approach for efficient computation of the RC circuit delay sensitivities.
Noel Menezes, Ross Baldick, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1997 Moment-sensitivity-based wire sizing for skew reduction in on-chip clock nets
abstract
Sensitivity-based methods for wire sizing have been shown to be effective in reducing clock skew in routed nets. However, lack of efficient sensitivity computation techniques and excessive space and time requirements often limit their utility for large clock nets. Furthermore, most skew reduction approaches work in terms of the Elmore delay model and, therefore, fail to balance the signal slopes at the clocked elements. In this paper, we extend the sensitivity-based techniques to balance the delays and signal-slopes by matching several moments instead of just the Elmore delay. As sensitivity computation is crucial to our approach, we present a new path-tracing algorithm to compute moment sensitivities for RC trees. Finally, to improve the runtime statistics of sensitivity-based methods, we also present heuristics to allow for efficient handling of large nets by reducing the size of the sensitivity matrix.
Satyamurthy Pullela, Noel Menezes, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1996 RC-Interconnect Macromodels for Timing Simulation
abstract
Most timing simulators obtain their efficiency over circuit simulation in terms of explicit integration algorithms that have dif-ficulty handling the stiff RC circuit models which characterize interconnect-dominated paths. In this paper we describe a reduced-order N-port interconnect macromodel for timing simula-tion. This macromodel is shown to improve the timing simulation efficiency dramatically since it alleviates the stiff circuit problem. Moreover, through its compatibility with the simple timing simula-tion transistor models, it is shown that this macromodel does not suffer from the dramatic increase in complexity with an increase in the number of ports like circuit simulation. I.
Florentin Dartu, Bogdan Tutuianu, Lawrence T. Pileggi
DAC3
1996 A Sparse Image Method for BEM Capacitance Extraction
abstract
Boundary element methods (BEM) are often used for complex 3-D capacitance extraction because of their efficiency, ease of data preparation, and automatic handling of open regions.BEM capacitance extraction, however, yields a dense set of linear equations that makes solving via direct matrix methods such as Gaussian elimination prohibitive for large problem sizes.Although iterative, multipole-accelerated techniques have produced dramatic improvements in BEM capacitance extraction, accurate sparse approximations of the electrostatic potential matrix are still desirable for the following reasons.First, the corresponding capacitance models are sufficient for a large number of analysis and design applications.Moreover, even when the utmost accuracy is required, sparse approximations can be used to precondition iterative solution methods.In this paper, we propose a definition of electrostatic potential that can be used to formulate sparse approximations of the electrostatic potential matrix in both uniform and multilayered planar dielectrics.Any degree of sparsity can be obtained, and unlike conventional techniques which discard the smallest matrix terms, these approximations are provably positive definite for the troublesome cases with a uniform dielectric and without a groundplane.
Byron Krauter, E. Aykut Dengi, Lawrence T. Pileggi
DAC4
1996 An Explicit RC-Circuit Delay Approximation Based on the First Three Moments of the Impulse Response
abstract
Due to its simplicity, the ubiquitous Elmore delay, or first moment of the impulse response, has been an extremely popular delay metric for analyzing RC trees and meshes. Its inaccuracy has been noted however and it has been demonstrated that higher order moments can be mapped to dominant pole approximations (e.g. AWE) in the general case. The first three moments can be mapped to a two pole approximation, but stability is an issue and even a stable model results in a transcendental equation that must be iteratively evaluated to determine the delay. We describe an explicit delay approximation based on the first three moments of the impulse response. We begin with the development of a provably stable two pole transfer function/impedance model based on the first three moments (about s=0) of the impulse response. Then, since the model form is known, we evaluate the delay (any waveform percentage point) in terms of an analytical approximation that is consistently within a fraction of 1 percent of the exact solution for this model. The result is an accurate, explicit delay expression that will be an effective metric for high speed interconnect circuit models.
Bogdan Tutuianu, Florentin Dartu, Lawrence T. Pileggi
DAC3
1996 Performance computation for precharacterized CMOS gates with RC loads
abstract
For efficiency, the performance of digital CMOS gates is often expressed in terms of empirical models. Both delay and short-circuit power dissipation are sometimes characterized as a function of load capacitance and input signal transition time. However, gate loads can no longer be modeled by purely capacitive loads for high performance CMOS due to the RC metal interconnect effects. This paper presents a methodology for interfacing empirical gate models to reduced order RC interconnect models in terms of a nonlinear iteration procedure. The delay and power are calculated with errors on the same order as those for the original empirical equations. Moreover, a linear equivalent gate model is generated which accurately captures the delays at the interconnect fan-out nodes.
Florentin Dartu, Noel Menezes, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1996 Domain characterization of transmission line models and analyses
abstract
Electronic system design requires noise and delay analyses of lossy, low-loss, frequency dependent, and coupled transmission lines, along with lumped parasitic elements. Numerous models and techniques have been developed for simulating these interconnect circuits, however, no single method represents the optimal approach for all transmission line problems. Accuracy and efficiency considerations would suggest the use of different models and analyses for the various types of interconnects. Such an approach to simulating transmission lines requires a measure of goodness for selecting the appropriate model and analysis in terms of the circuit parameters. Toward this goal, this paper presents an efficient methodology for characterizing transmission line models into solution domains. This provides an efficient technique for automatically selecting the type of interconnect model and analysis. Specifically, for a specified acceptable modeling error, analysis domains are established for automatic model selection between the classical method of characteristics, a lumped modeling approach and the state-based convolution technique within a simulation and analysis tool. In general, however, this approach can be applied to any combination of models and analyses, if a measure of accuracy and efficiency can be established.
Rohini Gupta, Seok-Yoon Kim, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1996 Post-processing of clock trees via wiresizing and buffering for robust design
abstract
Achieving near-zero skew for a large clock tree is generally at the expense of adding large amounts of metal interconnect. This added metal can significantly increase the total clock-net capacitance, thereby increasing the power dissipation in proportion. In addition, it is shown that it becomes increasingly difficult to control clock-signal skew due to metal-wiring process variations as the total clock-net capacitance increases. In this paper we demonstrate that buffer insertion can be used to reduce the total capacitance, hence the power, while generating a design which is as robust as one with no intermediate buffering. Moreover, delays are reduced substantially as well. Given an initial feasible route with constraints. On wire widths, wire sizing and buffer insertion are performed concurrently at each iteration of optimal buffer location search. Statistical process variations of the buffers, their loads, and the metal interconnect parameters are considered as part of the robust design process.
Satyamurthy Pullela, Noel Menezes, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1995 The Elmore Delay as a Bound for RC Trees with Generalized Input Signals
abstract
The Elmore delay is an extremely popular delay metric, particularly for RC tree analysis.The widespread usage of this metric is mainly attributable to it being the most accurate delay measure that is a simple analytical function of the circuit parameters.The only drawbacks to this delay metric are the uncertainty as to whether it is an optimistic or a pessimistic estimate, and the restriction to step response delay estimation.In this paper, we prove that the Elmore delay is an absolute upper bound on the 50% delay of an RC tree response.Moreover, we prove that this bound holds for input signals other than steps, and that the actual delay asymptotically approaches the Elmore delay as the input signal rise time increases.A lower bound on the delay is also developed using the Elmore delay and the second moment of the impulse response.The utility of this bound is for understanding the accuracy and the limitations of the Elmore delay metric as we use it for design automation.
Rohini Gupta, Byron Krauter, Bogdan Tutuianu, John Willis, Lawrence T. Pileggi
DAC5
1995 Transmission Line Synthesis
abstract
RLC transmission line synthesis is a difficult problem because minimal net delay cannot be achieved or accurately predicted unless: 1) a termination scheme is chosen and properly implemented, and 2) output driver transition rates are constrained at or below the net's capability.This paper describes a method to concurrently find, for any RLC net, an optimal termination value, a maximum source transition rate, and an approximate net delay using the first few response moments.When optimal termination is accomplished and transition rates do not exceed the capabilities of the net, the resulting delay metrics are an interesting extension of the popular Elmore delay metric for RC interconnects.The task of physical RLC interconnect design is facilitated by the ease with which these first few moments are calculated for generalized RLC lines.
Byron Krauter, Rohini Gupta, John Willis, Lawrence T. Pileggi
DAC4
1995 Simultaneous Gate and Interconnect Sizing for Circuit-Level Delay Optimization
abstract
Abstract—With delays due to the physical interconnect domi-nating the overall logic path delays, circuit-level delay optimiza-tion must take interconnect effects into account. Instead of sizing only the gates along the critical paths for delay reduction, the trade-off possible by simultaneously sizing gate and interconnect must also be considered. We show that for optimal gate and interconnect sizing, it is imperative that the interaction between the driver and the RC interconnect load be taken into account. We present an iterative sensitivity-based approach to simulta-neous gate and interconnect sizing in terms of a gate delay model which captures this interaction. During each iteration, the path delay sensitivities are efficiently calculated and used to size the components along a path. I.
Noel Menezes, Satyamurthy Pullela, Lawrence T. Pileggi
DAC3
1995 Constrained multivariable optimization of transmission lines with general topologies
abstract
The design of system level interconnects to meet signal integrity objectives is a challenging problem. This paper formulates the transmission line synthesis problem as a constrained multi-dimensional optimization of the complete net, taking into account factors like loading conditions on the line loss in the line and rise-time of the input signal. Different design variables such as width or resistivity of the interconnect, resistive source or far-end termination, etc. can all be considered concurrently. The termination metric is based upon forcing the impulse response waveform to be symmetric using the first three exact moments of the distributed system. An efficient means to trade-off between signal rise-time and ringing is presented and no time-domain simulations are needed. Several examples are presented to demonstrate the efficacy of the proposed methodology.
Rohini Gupta, Lawrence T. Pileggi
ICCAD2
1995 Generating sparse partial inductance matrices with guaranteed stability
abstract
This paper proposes a definition of magnetic vector potential that can be used to evaluate sparse partial inductance matrices. Unlike the commonly applied procedure of discarding the smallest matrix terms, the proposed approach maintains accuracy at middle and high frequencies and is guaranteed to be positive definite for any degree of sparsity (thereby producing stable circuit solutions). While the proposed technique is strictly based upon potential theory (i.e. the invariance of potential differences on the zero potential reference choice), the technique is, nevertheless, presented and discussed in both circuit and magnetic terms. The conventional and the proposed sparse formulation techniques are contrasted in terms of eigenvalues and circuit simulation results on practical examples.
Byron Krauter, Lawrence T. Pileggi
ICCAD2
1995 A sequential quadratic programming approach to concurrent gate and wire sizing
abstract
With an ever-increasing portion of the delay in highspeed CMOS chips attributable to the interconnect, interconnect-circuit design automation continues to grow in importance. By transforming the gate and multilayer wire sizing problem into a convex programming problem for the Elmore delay approximation, we demonstrate the efficacy of a sequential quadratic programming (SQP) solution method. For cases where accuracy greater than that provided by the Elmore delay approximation is required we apply SQP to the gate and wire sizing problem with more accurate delay models. Since efficient calculation of sensitivities is of paramount importance during SQP, we describe an approach for efficient computation of the accurate delay sensitivities.
Noel Menezes, Ross Baldick, Lawrence T. Pileggi
ICCAD3
1995 Coping with RC(L) interconnect design headaches
abstract
Physical interconnect effects have a dominant impact on today's deep submicron IC designs. In this tutorial paper we will describe the technology trends which have brought about this interconnect dominance, then consider some of the modeling and analysis approximations available for both pre- and post-layout interconnect design. This coverage will not be an exhaustive summary, but one that is primarily focused on moment-based analysis techniques, from the Elmore delay, to the more recent advances in moment-matching approximations, and the corresponding nonlinear driver/load interfaces. Future modeling, analysis, and design challenges will be considered throughout this paper.
Lawrence T. Pileggi
ICCAD1
1994 A Gate-Delay Model for high-Speed CMOS Circuits
abstract
As signal speeds increase and gate delays decrease for high-perform ance digital integrated circuits, the gate delay m odeling problem becom es increasingly m ore difficult.With scaling, increasing interconnect resistances and decreasing gateoutput im pedances m ake it m ore difficult to em pirically characteriz e gate-delay m odels.Moreover, the single-input-switching assum ption for the em pirical m odels is incom patible with the inevitable sim ultaneous switching for todays high-speed logic paths.In this paper a new em pirical gate delay m odel is proposed.Instead of building the em pirical equations in term s of capacitance loading and input-signal transition tim e, the m odels are generated in term s of param eters which com bine the benefits of em pirically derived k -factor m odels and switch-resistor m odels to efficiently: 1) handle capacitance shielding due to m etal interconnect resistance, 2) m odel the RC interconnect delay, and 3) provide tighter bounds for sim ultaneous switching.
Florentin Dartu, Noel Menezes, Jessica Qian, Lawrence T. Pileggi
DAC4
1994 OTTER: Optimal Termination of Transmission Lines Excluding Radiation
abstract
Article Free Access Share on OTTER: optimal termination of transmission lines excluding radiation Authors: Rohini Gupta Department of Electrical and Computer Engineering, The University of Texas at Austin, Austin, TX Department of Electrical and Computer Engineering, The University of Texas at Austin, Austin, TXView Profile , Lawrence T. Pillage Department of Electrical and Computer Engineering, The University of Texas at Austin, Austin, TX Department of Electrical and Computer Engineering, The University of Texas at Austin, Austin, TXView Profile Authors Info & Claims DAC '94: Proceedings of the 31st annual Design Automation ConferenceJune 1994 Pages 640–645https://doi.org/10.1145/196244.196600Published:06 June 1994Publication History 13citation195DownloadsMetricsTotal Citations13Total Downloads195Last 12 Months4Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Rohini Gupta, Lawrence T. Pileggi
DAC2
1994 RC interconnect synthesis-a moment fitting approach
Noel Menezes, Satyamurthy Pullela, Florentin Dartu, Lawrence T. Pileggi
ICCAD4
1994 Domain Characterization of Transmission Line Models for Efficient Simulation
abstract
Electronic system design requires simulating combinations of lossy, low-loss, frequency dependent, and coupled transmission lines. Accuracy and efficiency considerations would suggest the use of different models and analyses for the various types of interconnects. Toward this goal this paper presents an efficient methodology for characterizing transmission line models into solution domains. Specifically, given an acceptable percentage modeling error, analysis domains are established for automatic model selection between method of characteristics and a lumped modeling approach.>
Rohini Gupta, Seok-Yoon Kim, Lawrence T. Pileggi
ICCD3
1994 Enhancing the stability of asymptotic waveform evaluation for digital interconnect circuit applications
abstract
Asymptotic Waveform Evaluation (AWE) has been demonstrated as an efficient approach for interconnect circuit simulation/analysis. However, since it is based upon moment-matching, it is prone to yielding unstable approximations for stable circuits. This paper describes a systematic approach for alleviating the inherent instability associated with AWE and moment-matching methods as they apply to passive, digital, interconnect circuits. The efficiency and accuracy of this approach are demonstrated on several large, RLC interconnect-circuit problems.>
Demos F. Anastasakis, Nanda Gopal, Seok-Yoon Kim, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
1994 Time-domain macromodels for VLSI interconnect analysis
abstract
This paper presents a method of obtaining time-domain macromodels of VLSI interconnection networks for circuit simulation. The goal of this work is to include interconnect parasitics in a circuit simulation as efficiently as possible, without significantly compromising accuracy. Stability issues and enhancements to incorporate transmission line interconnects are also discussed. A unified circuit simulation framework, incorporating different classes of interconnects and based on the proposed macromodels, is described. The simplicity and generality of the macromodels is demonstrated through examples employing RC- and RLC-interconnects.>
Seok-Yoon Kim, Nanda Gopal, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1994 Modeling the "Effective capacitance" for the RC interconnect of CMOS gates
abstract
With finer line widths and faster switching speeds, the resistance of on-chip metal interconnect is having a dominant impact on the timing behavior of logic gates. Specifically, the gates are switching faster and the interconnect delays are getting longer due to scaling. This results in a trend in which the RC interconnect delay is beginning to comprise a larger portion of the overall logic stage delay. This shift in relative delay dominance from the gate to the RC interconnect is increased by resistance shielding. That is, as the gate "resistance" gets smaller and the metal resistance gets larger, the gate no longer "sees" the total net capacitance and the gate delay may be significantly less than expected. This trend complicates the timing analysis of digital circuits, which relies upon simple, empirical gate delay equations for efficiency. In this paper, we develop an analytical expression for the "effective load capacitance" of a pc interconnect. In addition, when there is significant shielding, the response waveforms at the gate output may have a large exponential tail. We show that this waveform tail can strongly influence the delay of the RC interconnect. Therefore, we propose an extension of the effective capacitance equation that captures the complete waveform response accurately, with a two-piece gate-output-waveform approximation.>
Jessica Qian, Satyamurthy Pullela, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1994 RICE: rapid interconnect circuit evaluation using AWE
abstract
This paper describes the Rapid Interconnect Circuit Evaluator (RICE) software developed specifically to analyze RC and RLC interconnect circuit models of virtually any size and complexity. RICE focuses specifically on the passive interconnect problem by applying the moment-matching technique of Asymptotic Waveform Evaluation (AWE) and application-specific circuit analysis techniques to yield large gains in run-time efficiency over circuit simulation without sacrificing accuracy. Moreover, this focus of AWE on passive interconnect problems permits the use of moment-matching techniques that produce stable, pre-characterized, reduced-order models for RC and RLC interconnects. RICE is demonstrated to be as accurate as a transient circuit simulation with hundreds or thousands of times the efficiency. The use of RICE is demonstrated on several VLSI interconnect and off-chip microstrip models.>
Curtis L. Ratzlaff, Lawrence T. Pileggi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1993 Reliable Non-Zero Skew Clock Trees Using Wire Width Optimization
abstract
Recognizingthat routing constraints and process variations make non-zero skew inevitable, this paper describes a novel methodology for constructing reliable low-skew clock trees.The algorithm efficiently calculates clock-tree delay sensitivities to achieve a target delay and a target skew.Moreover, the sensitivities also show that wires should be widened as opposed to lengthened to reduce skew since the former improves reliability while the latter reduces it.This paper introduces the concept of designing reliable clock nets with process-insensitive skew.
Satyamurthy Pullela, Noel Menezes, Lawrence T. Pileggi
DAC3
1993 Evaluation of Parts by Mixed-Level DC-Connected Components in Logic Simulation
abstract
Article Evaluation of parts by mixed-level DC-connected components in logic simulation Share on Authors: Dah-Cherng Yuan View Profile , Lawrence T. Pillage View Profile , Joseph T. Rahmeh View Profile Authors Info & Claims DAC '93: Proceedings of the 30th international Design Automation ConferenceJuly 1993 Pages 367–372https://doi.org/10.1145/157485.164934Online:01 July 1993Publication History 1citation129DownloadsMetricsTotal Citations1Total Downloads129Last 12 Months1Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Dah-Cherng Yuan, Lawrence T. Pileggi, Joseph T. Rahmeh
DAC2
1993 An efficient methodology for extraction and simulation of transmission lines for application specific electronic modules
abstract
Physical interconnect introduces new challenges for parameter extraction and delay calculation for application specific electronic module (ASEM) design automation. Efficiency dictates the precharacterization of extracted electrical parameters in the same manner as application specific integrated circuits (ASICs). However, ASEM interconnect is dominated by frequency dependent LC propagation which makes precharacterization difficult for all possible configurations. Moreover, simulating the transient behavior of the ASEM interconnect for noise and delay analysis requires the combined use of a variety of models and techniques for efficiently handling lossy, low-loss, frequency dependent, and coupled transmission lines together with lumped parasitic elements. We propose to use conformal mapping to generate "abstracted" models for the electrical parameters of various RLC interconnect cross-sections, including the frequency dependence caused by ground plane proximity and skin effects. Along with precharacterized lumped parasitic elements and nonlinear driver and load models, these models are simulated using a generalized time-domain macromodeling approach that can combine different types of transmission line analysis in one simulation environment. An automatic selection mechanism is derived for determination of the best time-domain macromodel for a particular distributed segment.
S. Y. Kim, Emre Tuncer, Rohini Gupta, Byron Krauter, Thomas L. Savarino, Dean P. Neikirk, Lawrence T. Pileggi
ICCAD7
1992 On the Stability of Moment-Matching Approximations in Asymptotic Waveform Evaluation
Demos F. Anastasakis, Nanda Gopal, Seok-Yoon Kim, Lawrence T. Pileggi
DAC4
1992 ETA: electrical-level timing analysis
abstract
A timing analyzer which performs timing analysis considering electrical-level details such as input signal slope, gate input distinction, charge sharing, and interconnect, while also taking into account such high-level concerns as path sensitization, is described. To achieve the greatest efficiency, ETA operates in two phases: (1) logic and delay precharacterization, and (2) longest path analysis. During the precharacterization phase, each gate is analyzed to get its Boolean function and its load at the transistor level. During the longest path analysis phase, paths from each primary input are enumerated and examined separately. Each potentially longest path is tested for sensitizability at the gate level until a sensitizable longest path is found. The circuit examples given demonstrate the importance of an accurate delay calculation in correctly finding the longest statistically sensitizable path.>
Ronn B. Brashear, Douglas R. Holberg, M. Ray Mercer, Lawrence T. Pileggi
ICCAD4
1992 AWE macromodels of VLSI interconnect for circuit simulation
abstract
The results of linear asymptotic waveform evaluation (AWE) and of nonlinear circuit simulation are combined for the purpose of efficiently incorporating accurate interconnect information in the overall circuit description. A simple macromodel based on the y-parameter description of a complex interconnect network is discussed. The model makes possible the reduction of large, stiff interconnect configurations into compact representations that pose minimal problems to conventional circuit simulation techniques. The macromodel can be incorporated directly into a conventional circuit simulation with no modification of the original simulation software. The techniques are suitable for development as a software library that can be used to enhance existing simulators so that they can handle large VLSI interconnect configurations very efficiently. In addition, the ability to handle elements at the port level makes approach extremely attractive for any linear(ized) macromodels or for application such as mixed-mode simulation.>
Seok-Yoon Kim, Nanda Gopal, Lawrence T. Pileggi
ICCAD3
1991 RICE: Rapid Interconnect Circuit Evaluator
abstract
Article RICE: Rapid interconnect circuit evaluator Share on Authors: Curtis L. Ratzlaff The University of Texas at Austin, Department of Electrical and Computer Engineering, Austin, TX The University of Texas at Austin, Department of Electrical and Computer Engineering, Austin, TXView Profile , Nanda Gopal The University of Texas at Austin, Department of Electrical and Computer Engineering, Austin, TX The University of Texas at Austin, Department of Electrical and Computer Engineering, Austin, TXView Profile , Lawrence T. Pillage The University of Texas at Austin, Department of Electrical and Computer Engineering, Austin, TX The University of Texas at Austin, Department of Electrical and Computer Engineering, Austin, TXView Profile Authors Info & Claims DAC '91: Proceedings of the 28th ACM/IEEE Design Automation ConferenceJune 1991 Pages 555–560https://doi.org/10.1145/127601.127732Online:01 June 1991Publication History 127citation652DownloadsMetricsTotal Citations127Total Downloads652Last 12 Months7Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Curtis L. Ratzlaff, Nanda Gopal, Lawrence T. Pileggi
DAC3
1991 Evaluating RC-Interconnect Using Moment-Matching Approximations
abstract
The authors describes a relation for specifying the 'optimal' number of lumped RC sections needed to approximate a distributed RC element for an estimated digital-signal bandwidth. The bandwidth approximation also aids in determining the order of the AWE (asymptotic waveform evaluation) approximation for the driving-point and transfer function models. Since moving to arbitrarily high orders of approximation to meet the bandwidth requirements is complicated by moment-matching instability problems, a constrained mapping from moments to dominant time constants is used which guarantees stability for RC interconnect models.>
Nanda Gopal, Dean P. Neikirk, Lawrence T. Pileggi
ICCAD3
1990 DC Parameterized Piecewise-Function Transistor Models for Bipolar and MOS Logic Stage Delay Evaluation
abstract
A novel technique is presented for analyzing nonlinear active devices driving RC trees which accounts for nonlinear behavior in such a way that the accuracy obtained is typically within 10% of SPICE. This is achieved with a tremendous savings in calculation time as compared to SPICE, using exclusively DC parameterized transistor models. These models have been included in a prototype program for analyzing bipolar and CMOS logic-stage delay models. These techniques are currently being extended to provide best-case and worst-case delay approximations in terms of the DC parameterized piecewise-linear and piecewise-quadratic Thevenin models.>
Douglas R. Holberg, Santanu Dutta, Lawrence T. Pileggi
ICCAD3
1990 Asymptotic waveform evaluation for timing analysis
abstract
Asymptotic waveform evaluation (AWE) provides a generalized approach to linear RLC circuit response approximations. The RLC interconnect model may contain floating capacitors, grounded resistors, inductors, and even linear controlled sources. The transient portion of the response is approximated by matching the initial boundary conditions and the first 2q-1 moments of the exact response to a lower-order q-pole model. For the case of an RC tree model, a first-order AWE approximation reduces to the RC tree methods.>
Lawrence T. Pileggi, Ronald A. Rohrer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1989 AWEsim: Asymptotic Waveform Evaluation for Timing Analysis
abstract
Most timing analyzers rely upon a linear approximate interconnect model, typically an RC tree, to estimate efficiently the propagation delays for digital MOS integrated circuits. RC tree methods are adequate to analyze a large class of MOS circuits, but are not sufficient in general for high speed, dynamic and precharge MOS circuits. In addition bipolar logic and board level digital systems can have interconnect models which may not be compatible with RC tree topologies. In this paper we describe AWEsim, a variable refinement waveform estimator for generalized linear RLC approximate interconnect models.
Lawrence T. Pileggi, Xueqing Huang, Ronald A. Rohrer
DAC1
1989 Efficient Final Placement Based on Nets-as-Points
abstract
Deterministic optimization programs are coming to be considered as viable alternatives for the placement of very large sea-of-gate, gate array and standard cell designs. A nets-as-points placement program has been described which provides competitive results in comparison with non-deterministic placement, and at a fraction of the run time. A new pseudo Steiner tree model for the gate placement about the netpoints, along with a virtual channel snap-to-grid procedure, provides results superior to the original nets-as-points placement program without requiring iterative improvement.
Lawrence T. Pileggi, Ronald A. Rohrer
DAC2
1988 A Quadratic Metric with a Simple Solution Scheme for Initial Placement
Lawrence T. Pileggi, Ronald A. Rohrer
DAC1