EDBT 2026 Demo / reviewers in the wild / expert
Brian Cline
dblp:63/1307
· DBLP profile ↗
28ranked-venue papers
3as first author
3since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Design-Aware Partitioning-Based 3-D IC Design Flow With 2-D Commercial Toolsabstract3-D ICs can continue to improve power, performance, area, and cost beyond traditional Moore’s law scaling limitations by leveraging the third dimension and short vertical interconnects. Several recent studies present methodologies to implement 3-D ICs, but most of these studies implement each tier separately after partitioning a design into multiple tiers, resulting in inaccurate buffer insertion, which becomes more severe in advanced technology nodes. In this article, we present a new methodology called “Cascade2D flow” which utilizes design and microarchitecture insight for tier partitioning and implements 3-D ICs using 2-D commercial tools. By modeling vertical interconnects with sets of anchor cells and dummy wires, Cascade2D flow places, and routes and optimizes multiple tiers simultaneously in the 2-D version of a 3-D IC called “cascade2D design,” which enables accurate buffer insertion. Two flavors of 3-D ICs—monolithic 3-D (M3D) and face-to-face-bonded (F2F-bonded) 3-D ICs—of a commercial in-order, 32-bit application processor at foundry 28 nm, 14/16 nm, and predictive 7-nm technology nodes are implemented using this new methodology. We investigate the power, performance and area improvements of 3-D ICs over the 2-D counterparts to examine the efficacy of the methodology. Our new methodology outperforms the state-of-the-art 3-D IC design flows in the both flavors of 3-D ICs with up to$4\times $better power savings. In the best case, 3-D ICs from Cascade2D flow show 25% better performance at iso-power and 20% lower power at iso-performance. Kyungwook Chang, Saurabh Sinha 0001, Brian Cline, Greg Yeric, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Power Delivery and Thermal-Aware Arm-Based Multi-Tier 3D Architectureabstract3D integration is becoming a cost-effective way to incorporate more CPU cores and memory to improve the performance of computing systems. Meanwhile, due to the higher power density, power delivery and thermal issues become more significant in multi-tier 3DICs. In this paper, we explore and evaluate multiple design options for an Arm Neoverse-based 3D architecture focusing on power and thermals at 7nm process and sub-10$\mu $m pitch. Using a rapid voltage-drop and thermal analysis methodology, we model a system with a 32-core CPU layer and up to 4 layers of system-level caches, and quantity the trade-offs between performance, cost, voltage-drop, and temperature. A 3-layer configuration shows a good balance with 17% IPC gain and 17% lower cost, while incurring 15mV worse voltage drop and 8.5°C higher temperature compared with 2D. Our studies suggest that the co-optimization of system architecture, technology, and physical design is key for high-performance 3D systems. Lingjun Zhu, Tuan Ta, Rossana Liu, Rahul Mathur, Shidhartha Das, Ankit Kaul, Alejandro Rico, Doug Joseph, Brian Cline, Sung Kyu Lim |
ISLPED | 10 |
| 2021 | High-Performance Logic-on-Memory Monolithic 3-D IC Designs for Arm Cortex-A ProcessorsabstractMonolithic 3-D IC (M3-D) is a promising solution to improve the performance and energy-efficiency of modern processors. But, designers are faced with challenges in design tools and methodologies, especially for power and thermal verifications. We developed a new physical design flow that optimally places and routes cache modules in one tier and logic gates in the other. Our tool also builds high-quality clock and power delivery networks targeting logic-on-memory M3-D designs. Finally, we developed a sign-off analysis tool flow to evaluate power, performance, area (PPA), thermal, and voltage-drop quality for given M3-D designs. Using our complete register transfer level (RTL)-to-Graphic Design System (GDS) tool flow, we designed commercial quality 2-D and M3-D implementation of Arm Cortex-A7 and Cortex-A53 processors in a commercial 28-nm technology. Experimental results show that our 3-D processors offer 20% (A7) and 21% (A53) performance gain, compared with their 2-D commercial counterparts. The voltage-drop degradation of our 3-D Cortex-A7 and Cortex-A53 processors is less than 3% of the supply voltage, while temperature increase is 10.71 °C and 13.04 °C, respectively. Lingjun Zhu, Lennart Bamberg, Sai Pentapati, Kyungwook Chang, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Brian Cline, Saurabh Sinha 0001, Alberto García Ortiz, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2019 | Enhanced 3D Implementation of an Arm® Cortex®-A MicroprocessorabstractHigh-density 3D techniques (such as wafer bonding and monolithic-3D) show tremendous promise in reducing interconnect lengths and relieving 2D congestion. We propose an enhanced 3D implementation methodology and use it to design an Arm Cortex-A microprocessor in 3D. The methodology is fully integrated and tested using commercial EDA tools and incorporates all physical IP needed to implement modern microprocessors. The resulting 3D implementation consists of two parts, 1) a multi-tier co-placement approach for enhanced placement quality, 2) integration of 3D SRAMs for improved microprocessor PPA. Compared to the 2D baseline, our implementations show an overall area reduction of 8.5% and can either achieve an 18% peak frequency uplift at iso-power or a 41% power reduction at near iso-performance (-3% frequency). Mudit Bhargava, Saurabh Sinha 0001, Brian Cline |
ISLPED | 5 |
| 2019 | System-Level Power Delivery Network Analysis and Optimization for Monolithic 3-D ICsabstractAs 2-D scaling reaches its limit, monolithic 3-D (M3D) IC is a leading contender for continuing equivalent scaling. Although M3D shows power and performance benefits over 2-D designs, designing a power delivery network (PDN) for M3D is challenging. In this paper, for the first time, we present a system-level PDN model of M3D designs focusing on both resistive (IR) and inductive (Ldi/dt) components of power supply integrity. In addition, we present frequency- and time-domain analyses of M3D PDNs. We show that the additional resistance in M3D PDNs, while being worse for resistive drops, improves resiliency against ac current noise showing 35.9% peak impedance reduction compared to 2-D PDNs during worst case resonant oscillations. Then, we present methodologies to improve power supply integrity of M3D designs based on the observations. Our optimization methodologies offer up to 32.6% and 17.0% static and dynamic voltage drop reduction compared to the baseline M3D designs, respectively, showing 9.0% lower dynamic voltage drop compared to the 2-D counterparts. Kyungwook Chang, Shidhartha Das, Saurabh Sinha 0001, Brian Cline, Greg Yeric, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Accurate processor-level wirelength distribution model for technology pathfinding using a modernized interpretation of rent's ruleabstractFaithful system-level modeling is vital to design and technology pathfinding, and requires accurate representation of interconnects. In this study, Rent's rule is modernized to cater to advanced technology and design, and applied to derive a priori wirelength distribution models. Furthermore, a priori interconnect branching models are proposed to capture design constraints and their handling by the Electronic-Design-Automation tools. These interconnect branching models are embedded into the wirelength distribution models and validated against a suite of state-of-the-art commercial designs across technology nodes. Novel design-specific critical-path models are presented which capture trends in technology and microarchitecture, providing a reliable framework for future technology and design benchmarking. Divya Prasad, Saurabh Sinha 0001, Brian Cline, Azad Naeemi |
DAC | 3 |
| 2017 | Standard cell library design and optimization methodology for ASAP7 PDK: (Invited paper)abstractStandard cell libraries are the foundation for the entire back-end design and optimization flow in modern application-specific integrated circuit designs. At 7nm technology node and beyond, standard cell library design and optimization is becoming increasingly difficult due to extremely complex design constraints, as described in the ASAP7 process design kit (PDK). Notable complexities include discrete transistor sizing due to FinFETs, complicated design rules from lithography and restrictive layout space from modern standard cell architectures. The design methodology presented in this paper enables efficient and high-quality standard cell library design and optimization with the ASAP7 PDK. The key techniques include exhaustive transistor sizing for cell timing optimization, transistor placement with generalized Euler paths and back-end design prototyping for library-level explorations. Nishi Shah, Andrew Evans, Saurabh Sinha 0001, Brian Cline, Greg Yeric |
ICCAD | 5 |
| 2017 | DTCO for DSA-MP Hybrid Lithography with Double-BCP Materials in Sub-7nm NodeabstractIn the sub-7nm technology nodes, as the mask cost for printing the dense via layers increases dramatically with conventional lithography technique, industries are actively looking for some alternative techniques, such as E-beam, EUV or nanoimprint. In recent years, directed self-assembly (DSA) has been demonstrated to be a promising candidate to reduce the number of masks due to its pitch multiply ability by grouping several vias in the same DSA guiding patterns. A novel idea has been proposed recently to further reduce the mask cost by applying double block copolymer (double-BCP) materials for DSA lithography. However, it is also discovered that new challenges are introduced for the DSA pattern assignment and decomposition problem with double-BCP DSA. This paper discusses the design technology co-optimization (DTCO) for BCP materials and DSA-MP hybrid lithography. We show that it is necessary to consider the pitch-range of BCP materials for the DSA pattern assignment and decomposition. We first propose a pitch-range optimization method for the BCP material to minimize the potential DSA decomposition conflicts. We then solve the double-BCP pattern assignment and decomposition by formulating it to a maximum weighted independent set problem, and obtain the optimal solution with integer linear programming (ILP). We also propose a bounded approximation algorithm to solve the ILP problem more efficiently. The experimental results demonstrate that our pitch-range optimization method can quickly find the pitch-range that can minimize the decomposition conflicts. With the optimized pitch-ranges, double-BCP outperforms single-BCP by reducing more than 500 conflicts for very dense design. In addition, our proposed approximation algorithm solves the problem 38× faster than the ILP formulation. Jiaojiao Ou, Brian Cline, Greg Yeric, David Z. Pan |
ICCD | 3 |
| 2017 | Frequency and time domain analysis of power delivery network for monolithic 3D ICsabstractAs 2D scaling reaches its limit, monolithic 3D IC (M3D) is a leading contender to continue equivalent scaling. Although M3D shows power and performance benefits over 2D designs, designing a power delivery network (PDN) for M3D is challenging. In this paper, for the first time, we present a system-level PDN model of M3D designs focusing on both resistive (IR) and inductive (Ldi/dt) components of power-supply integrity. In addition, we present frequency- and time-domain analysis of the M3D PDN. We show that the additional resistance in the M3D PDN, while being worse for resistive drops, improves resiliency against current noise showing 35.9% peak impedance reduction during worst-case resonant oscillations. Kyungwook Chang, Shidhartha Das, Saurabh Sinha 0001, Brian Cline, Greg Yeric, Sung Kyu Lim |
ISLPED | 4 |
| 2017 | Redundant Local-Loop Insertion for Unidirectional RoutingabstractAs the semiconductor manufacturing technology continues to scale down to sub-10 nm, unidirectional layout style has become the mainstream for lower metal layers with tight pitches. Conventional redundant via (RV) insertion for yield improvement has become obsolete because unidirectional routing patterns forbid off-track routing, i.e., wire bending, for the metal coverage of RVs. To enhance the yield, redundant local-loop insertion (RLLI) is a new way of inserting RVs due to its compatibility with the unidirectional layout style. This paper proposes the first global optimization engine for RLLI considering advanced manufacturing constraints. Our key contributions include bounded timing impact analysis and evaluation for the local-loop structure, net-based local-loop candidate generation and pruning, an integer linear programming (ILP) formulation and scalable iterative relaxation/linear programming solving (IRLS) with incremental search scheme. Experimental results demonstrate that with bounded timing impact (within 1%), the ILP formulation obtains highest insertion rate while the IRLS with incremental search scheme achieves scalable solutions with competitive solution qualities. Yibo Lin, Meng Li 0004, Jiaojiao Ou, Brian Cline, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Impact and Design Guideline of Monolithic 3-D IC at the 7-nm Technology NodeabstractMonolithic 3-D (M3D) IC is one of the potential technologies to break through the challenges of continued circuit power and performance scaling. In this paper, for the first time, we demonstrate the power benefits of M3D and present design guideline in a 7-nm FinFET technology node. The predictive 7-nm process design kit (PDK) and the standard cell library using both high-performance (HP) and low-standby-power (LSTP) device technologies are developed based on NanGate 45-nm PDK using accurate dimensional, material, and electrical parameters from publications and a commercial-grade tool flow. We implement full-chip M3D designs utilizing industry-standard physical design tools, and gauge the impact of M3D technology on performance, power, and area metrics. We also provide the design guidelines as well as a new partitioning methodology to improve M3D design quality. This paper shows that M3D designs outperform 2-D counterparts by 16% and 16.5% on average in terms of isoperformance total power reduction with 7-nm HP and LSTP cell library, respectively. This demonstrates the power benefits of M3D technology in both HP and low-power future generation devices. Kyungwook Chang, Kartik Acharya, Saurabh Sinha 0001, Brian Cline, Greg Yeric, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Match-making for monolithic 3D IC: finding the right technology nodeabstractMonolithic 3D IC (M3D) has the potential to provide a break-through in the power and performance scaling challenges. We, for the first time, present a comprehensive study of M3D on a commercial design across multiple technology nodes. The performance and power impact of M3D is investigated using a commercial, in-order, 32-bit application processor, implemented on foundry 28nm and 14/16nm process nodes, as well as a predictive 7nm node. We study the factors across the technology nodes that affect the efficiency of M3D, and propose a roadmap for optimum technology and design interaction that will enable the full entitlement of M3D. Kyungwook Chang, Saurabh Sinha 0001, Brian Cline, Greg Yeric, Sung Kyu Lim |
DAC | 3 |
| 2016 | Near-threshold computing in FinFET technologies: opportunities for improved voltage scalabilityabstractIn recent years, operating at near-threshold supply voltages has been proposed to improve energy efficiency in circuits, yet decreased efficacy of dynamic voltage scaling has been observed in recent planar technologies. However, foundries have introduced a shift from planar to FinFET fabrication processes. In this paper, we study 7nm FinFET's ability to voltage scale and compare it to planar technologies across three dynamic voltage scaling scenarios. The switch to FinFET allows for a return to strong voltage scalability. We find up to 8.6× higher energy efficiency at NT compared to nominal supply voltage (vs. 4.8× gain in 20nm planar). Nathaniel Ross Pinckney, Lucian Shifren, Brian Cline, Saurabh Sinha 0001, Supreet Jeloka, Ronald G. Dreslinski, Trevor N. Mudge, Dennis Sylvester, David T. Blaauw |
DAC | 3 |
| 2016 | Cascade2D: A design-aware partitioning approach to monolithic 3D IC with 2D commercial toolsabstractMonolithic 3D IC (M3D) can continue to improve power, performance, area and cost beyond traditional Moore's law scaling limitations by leveraging the third-dimension and fine-grained monolithic inter-tier vias (MIVs). Several recent studies present methodologies to implement M3D designs, but most, if not all of these studies implement top and bottom tier separately after partitioning, which results in inaccurate buffer insertion. In this paper, we present a new methodology called ‘Cascade2D’ that utilizes design and micro-architecture insight to partition and implement an M3D design using 2D commercial tools. By modeling MIVs with sets of anchor cells and dummy wires, we implement and optimize both top and bottom tier simultaneously in a single 2D design. M3D designs of a commercial, in-order, 32-bit application processor at the foundry 28nm, 14/16nm and predictive 7nm technology nodes are implemented using this new methodology and we investigate the power, performance and area improvements over 2D designs. Our new methodology consistently outperforms the state-of-the-art M3D design flow with up to 4× better power savings. In the best case scenario, M3D designs from the Cascade2D flow show 25% better performance at iso-power and 20% lower power at isoperformance. Kyungwook Chang, Saurabh Sinha 0001, Brian Cline, Raney Southerland, Michael Doherty, Greg Yeric, Sung Kyu Lim |
ICCAD | 3 |
| 2016 | Four-tier Monolithic 3D ICs: Tier Partitioning Methodology and Power Benefit StudyabstractMonolithic 3D IC is an emerging technology to continuously satisfy demands for power reduction under challenges posed by traditional device scaling. In this paper, for the first time, we study power benefits of 4-tier monolithic 3D ICs compared with 2-tier monolithic 3D and 2D ICs. We present a tier partitioning methodology that significantly extends the capability of a state-of-the-art flow built for 2-tier monolithic 3D ICs. We develop two complete RTL-to-GDSII design flows to achieve this goal and offer quantitative comparisons. In addition, we study impacts of inter-tier via usage on 2-tier and 4-tier monolithic 3D ICs. Our experiments show that poorly controlled inter-tier via usage results in up to 6.05% degradation in total power savings. Thus, we develop an effective strategy to achieve inter-tier via configurations to optimize power metrics. Experiments show that 4-tier monolithic 3D ICs outperform 2-tier and 2D IC by 15% and 50% in terms of power and 25% and 75% in area under the same performance. Kwang Min Kim, Saurabh Sinha 0001, Brian Cline, Greg Yeric, Sung Kyu Lim |
ISLPED | 3 |
| 2015 | Power benefit study of monolithic 3D IC at the 7nm technology nodeabstractMonolithic 3D IC (M3D) is one potential technology to help with the challenges of continued circuit power and performance scaling. In this paper, for the first time, the power benefits of monolithic 3D IC (M3D) using a 7nm FinFET technology are investigated. The predictive 7nm Process Design Kit (PDK) and standard cell library for both high performance (HP) and low standby power (LSTP) device technologies are built based on NanGate 45nm PDK using accurate dimensional, material, and electrical parameters from publications and a commercial-grade tool flow. In addition, we implement full-chip M3D GDS layouts using both 7nm HP and LSTP cells and industry-standard physical design tools, and evaluate the resulting full-chip power, performance, and area metrics. Our study first shows that 7nm HP M3D designs outperform 7nm HP 2D designs by 16.8% in terms of iso-performance total power reduction. Moreover, 7nm LSTP M3D designs reduce the total power consumption by 14.3% compared to their 2D counterparts. This convincingly demonstrates the power benefits of M3D technologies in both high performance as well as low power future generation devices. Kyungwook Chang, Kartik Acharya, Saurabh Sinha 0001, Brian Cline, Greg Yeric, Sung Kyu Lim |
ISLPED | 4 |
| 2015 | Self-Aligned Double Patterning Aware Pin Access and Standard Cell Layout Co-OptimizationabstractSelf-aligned double patterning (SADP) is being considered for use at the 10-nm technology node and below for routing layers with pitches down to ~50 nm because it has better line edge roughness and overlay control compared to other multiple patterning candidates. To date, most of the SADP-related literature has focused on enabling SADP-legal routing in physical design tools while few attempts have been made to address the impact SADP routing has on local, standard cell (SC) I/O pin access. At the same time, via layers are used to connect the local SADP routing layers to the I/O pins on lower metal layers. Due to the high via density on the Via-1 layer, the litho-etch-litho-etch (LELE)-aware Via-1 design becomes a necessity to achieve legal pin access at the SC level. In this paper, we present the first study on SADP-aware pin access and layout optimization at the SC level. Accounting for SADP-specific and Via-1 design rules, we propose a coherent framework that uses depth first search, mixed integer linear programming, and backtracking method to enable LELE friendly Via-1 design and simultaneously optimize SADP-based local pin access and within-cell connections. Our experimental results show that, compared with the conventional approach, our framework effectively improves pin access of the SCs and maximizes the pin access flexibility for routing. Brian Cline, Greg Yeric, Bei Yu 0001, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | DFM is dead - Long live DFMabstractFor many years, a key aspect of Design-for-Manufacturability (DFM) has been adjustment of polygons in standard cell layout. Similarly, radically restricted design rules and unidirectional layout have been proposed as DFM-friendly design styles with the ability to dramatically improve yield. This paper looks at the history of such approaches over the last few technology nodes and shows that as we approach 10nm and beyond, both of these techniques have essentially run their course as lithography restrictions and related effects increasingly dictate key aspects of standard cell layout. In their place, new approaches to manufacturability and yield are increasingly important. Robert C. Aitken, David Pietromonaco, Brian Cline |
ICCD | 3 |
| 2014 | Physical design and FinFETsabstractFinFETs have recently overtaken bulk CMOS transistors as the device of choice for systems-on-chip. This paper provides some background on FinFETs together with their associated manufacturing processes and shows how they influence physical design of standard cells as well as place & route and timing closure for larger blocks. Robert C. Aitken, Greg Yeric, Brian Cline, Saurabh Sinha 0001, Lucian Shifren, Imran Iqbal, Vikas Chandra |
ISPD | 3 |
| 2014 | Self-aligned double patterning aware pin access and standard cell layout co-optimizationabstractSelf-Aligned Double Patterning (SADP) is being considered for use at the 10$nm$ technology node and below for routing layers with pitches down to ~50nm because it has better LER and overlay control compared to other multiple patterning candidates. To date, most of the SADP-related literature has focused on enabling SADP-legal routing in physical design tools while few attempts have been made to address the impact SADP routing has on local, standard cell (SC) I/O pin access. In this paper, we present the first study on SADP-aware pin access and layout optimization at the SC level. Accounting for SADP-specific design rules, we propose a coherent framework that uses Mixed Integer Linear Programming (MILP) and branch and bound method to simultaneously optimize SADP-based local pin access and within-cell connections. Our experimental results show that, compared with the conventional approach, our framework effectively improves pin access of the standard cells and maximizes the pin access flexibility for routing. Brian Cline, Greg Yeric, Bei Yu 0001, David Z. Pan |
ISPD | 2 |
| 2012 | Exploring sub-20nm FinFET design with predictive technology modelsabstractPredictive MOSFET models are critical for early stage design-technology co-optimization and circuit design research. In this work, Predictive Technology Model files for sub-20nm multi-gate transistors have been developed (PTM-MG). Based on MOSFET scaling theory, the 2011 ITRS roadmap and early stage silicon data from published results, PTM for FinFET devices are generated for 5 technology nodes corresponding to the years 2012-2020 on the ITRS roadmap. Saurabh Sinha 0001, Greg Yeric, Vikas Chandra, Brian Cline, Yu Cao 0001 |
DAC | 4 |
| 2012 | Design benchmarking to 7nm with FinFET predictive technology modelsabstractThe coming ten years promise great changes in silicon technology, with the end of planar bulk CMOS and the rise of interconnect parasitics to true significance. With such shifts in the underlying technology, the simple extrapolation of performance metrics may lead to pronounced prediction errors in design pathfinding. In this work, we utilize newly developed Predictive Technology Models for FinFETs aligned to the 2011 ITRS. Together with predictive interconnect models, we project performance and power landscape for the technology nodes from 20nm to 7nm. We present an overview of models, assess the advantage of FinFET over bulk CMOS devices, benchmark the scaling of critical design metrics, and illustrate major design barriers toward the 7nm node. Saurabh Sinha 0001, Brian Cline, Greg Yeric, Vikas Chandra, Yu Cao 0001 |
ISLPED | 2 |
| 2010 | Mechanical Stress Aware Optimization for Leakage Power ReductionabstractProcess-induced mechanical stress is used to enhance carrier transport and achieve higher drive currents in current complementary metal-oxide-semiconductor technologies. This paper explores how to fully exploit the layout dependence of stress enhancement and proposes a circuit-level, block-based, stress-enhanced optimization algorithm that uses stress-optimized layouts in conjunction with dual-Vthassignment to achieve optimal power-performance tradeoffs. We begin by studying how channel stress and drive current depend on layout parameters such as active area length and contact placement, while considering all layout-dependent sources of mechanical stress in a 65 nm industrial process. We then investigate the three main layout properties that impact mechanical stress in this process and discuss how to improve stress-based performance enhancement in standard cell libraries. While varying the stress-altering layout properties of a number of standard cells in a 65 nm industrial library, we show that ?dual-stress? standard cell layouts (analogous to ?dual-Vth?) can be designed to achieve drive current differences up to ~ 14% while incurring less than half the leakage penalty of dual-Vth. Therefore, when the flexibility of ?dual-stress? assignment is combined with dual-Vthassignment (within the proposed joint optimization framework), simulation results for a set of benchmark circuits show that leakage is reduced by ~ 24% on average, for iso-delay, when compared to dual-Vthassignment. Since mobility enhancement does not incur the exponential leakage penalty associated withVthassignment, our optimization technique is ideal for leakage power reduction. However, our framework can also be used to achieve higher performance circuits for iso-leakage and our joint optimization framework can be used to reduce delay on average by ~ 5%. In both cases, the proposed method only incurs a small area penalty (<0.5%). Vivek Joshi, Brian Cline, Dennis Sylvester, David T. Blaauw, Kanak Agarwal 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Leakage power reduction using stress-enhanced layoutsabstractIn recent years, process-induced mechanical stress has emerged as a useful manufacturing technique that enhances carrier transport and increases drive currents. This improvement in current has helped to compensate the decline of device scaling factors in parameters such as tox, Vth, and Vdd. In this work, we propose stress as a means to achieve optimal power-performance trade-off by combining stress-based, performance-enhanced standard cell assignment with dual-Vth assignment. We study how stress-induced performance enhancements are affected by layout properties and improve standard cell layouts so that performance gains are maximized. We then develop a circuit-level, block-based, stress-enhanced optimization algorithm that includes all layout-dependent sources of mechanical stress. By combining the two performance enhancement techniques (stress-based and dual-Vth) for a set of benchmark circuits, we find that our stress-aware optimization, decreases leakage by ~24% on average, for iso-delay, when compared to dual-Vth assignment. Similarly, for iso-leakage, our optimization algorithm reduces delay on average by 5%. In both cases, the proposed method only incurs a small area penalty (< 0.5%). Vivek Joshi, Brian Cline, Dennis Sylvester, David T. Blaauw, Kanak Agarwal 0001 |
DAC | 2 |
| 2008 | Transistor-Specific Delay Modeling for SSTAabstractSSTA has received a considerable amount of attention in recent years. However, it is a general rule that any approach can only be as accurate as the underlying models. Thus, variation models are an important research topic, in addition to the development of statistical timing tools. These models attempt to predict fluctuations in parameters like doping concentration, critical dimension (CD), and ILD thickness, as well as their spatial correlations. Modeling CD variation is a difficult problem because it contains a systematic component that is context dependent as well as a probabilistic component that is caused by exposure and defocus variation. Since these variations are dependent on topology, modern-day designs can potentially contain thousands of unique CD distributions. To capture all of the individual CD distributions within statistical timing, a transistor-specific model is required. However, statistical CD models used in industry today do not distinguish between transistors contained within different standard cell types (at the same location in a die), nor do they distinguish between transistors contained within the same standard cell. In this work we verify that the current methodology is error-prone using a 90 nm industrial library and lithography recipe (with industrial OPC) and propose a new SSTA delay model that on average reduces error of standard deviation from 11.8% to 4.1% when the total variation (sigma/mu) is 4.9% - a 2.9X reduction. Our model is compatible with existing SSTA techniques and can easily incorporate other sources of variation such as random dopant fluctuation and line-edge roughness. Brian Cline, Kaviraj Chopra, David T. Blaauw, Andres Torres, Savithri Sundareswaran |
DATE | 1 |
| 2008 | STEEL: a technique for stress-enhanced standard cell library designabstractMobility degradation and device scaling limitations have led process engineers to develop new techniques that introduce mechanical stress in MOSFET channels, which results in enhanced carrier transport. New fabrication steps strive to increase carrier mobility which, consequently, increases both Ionand Ioffin CMOS devices. However, most stress-enhancement techniques are dependent on layout parameters and their effects can be exploited within standard cell library design. In this work, we propose a new standard cell library design methodology that shares VDDand VSSsource/drain connections across standard cell boundaries. Such sharing allows for increased channel stress in both the corresponding device as well as its neighboring devices. Using an industrial 65 nm process and standard cell library, we show that our standard cell design methodology can be seamlessly integrated into current, state-of-the-art digital IC design flows. The new shared source/ drain technique improves critical path delay by 11% on average over a number of benchmarks for only a ~35% increase in leakage. Furthermore, stress-enhanced standard cell libraries offer a superior power/ delay tradeoff compared to dual-Vthacross a wide range of operating points with reduced manufacturing costs. Specifically, our stress-enhanced library (with a single Vth) consumes ~2.5X less leakage than its dual-Vthcounterpart. Brian Cline, Vivek Joshi, Dennis Sylvester, David T. Blaauw |
ICCAD | 1 |
| 2008 | Stress aware layout optimizationabstractProcess-induced mechanical stress is used to enhance carrier transport and achieve higher drive currents in current CMOS technologies. In this paper, we study how stress-induced performance enhancements are affected by layout properties and suggest guidelines for improving layouts so that performance gains are maximized. All MOS devices in this work include STI and nitride stress liners as sources of stress. Additionally, the PMOS devices incorporate the stress effects caused by the embedded SiGe S/D layer common in today's processes. First, we study how stress and drive current depend on layout parameters such as active area length and contact placement. We develop an intuition for the drive current dependency on these parameters and propose simple guidelines to improve a layout while considering mechanical stress effects. We then use these guidelines to improve the standard cell layouts in a 65nm industrial library. Experimental results show that we can enhance NMOS and PMOS drive currents by ~5% and ~12%, respectively, while only increasing NMOS leakage current by 1.48X and PMOS leakage current by 3.78X. By applying our guidelines to a 3-input NOR gate and a 3-input NAND gate, we are able to achieve a ~13.5% PMOS drive current improvement in the NOR gate and a ~7% NMOS drive current improvement in the NAND gate, without increasing cell area in either case Vivek Joshi, Brian Cline, Dennis Sylvester, David T. Blaauw, Kanak Agarwal 0001 |
ISPD | 2 |
| 2006 | Analysis and modeling of CD variation for statistical static timingabstractStatistical static timing analysis (SSTA) has become a key method for analyzing the effect of process variation in aggressively scaled CMOS technologies. Much research has focused on the modeling of spatial correlation in SSTA. However, the vast majority of these works used artificially generated process data to test the proposed models. Hence, it is difficult to determine the actual effectiveness of these methods, the conditions under which they are necessary, and whether they lead to a significant increase in accuracy that warrants their increased runtime and complexity. In this paper, we study 5 different correlation models and their associated SSTA methods using 35420 critical dimension (CD) measurements that were extracted from 23 reticles on 5 wafers in a 130nm CMOS process. Based on the measured CD data, we analyze the correlation as a function of distance and generate 5 distinct correlation models, ranging from simple models which incorporate one or two variation components to more complex models that utilize principle component analysis and Quad-trees. We then study the accuracy of the different models and compare their SSTA results with the result of running STA directly on the extracted data. We also examine the trade-off between model accuracy and run time, as well as the impact of die size on model accuracy. We show that, especially for small dies (< 6.6mm x 5.7mm), the simple models provide comparable accuracy to that of the more complex ones, while incurring significantly less runtime and implementation difficulty. The results of this study demonstrate that correlation models for SSTA must be carefully tested on actual process data and must be used judiciously. Brian Cline, Kaviraj Chopra, David T. Blaauw, Yu Cao 0001 |
ICCAD | 1 |