EDBT 2026 Demo / reviewers in the wild / expert
Puneet Gupta 0001
dblp:06/1383-1
· DBLP profile ↗
117ranked-venue papers
20as first author
20since 2021 · last 2026
0000-0002-6188-1134ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 110 · 20 first-author · 20 since 2021Software engineering, systems software and programming languages · 13 · 3 since 2021Security and privacy · 3 · 1 since 2021Theory of computation · 3Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DART: Dynamic Repair for Interconnect Fault Tolerance in Hybrid Bonding
Partho Bhoumik, Puneet Gupta 0001, Krishnendu Chakrabarty |
VTS | 3 |
| 2026 | CATCH: A Cost Analysis Tool for Co-Optimization of Chiplet-Based Heterogeneous SystemsabstractWith the increasing prevalence of chiplet systems in high-performance computing applications, the number of design options has increased dramatically. Instead of chips defaulting to a single die per package, now there are viable 3D stacking options to integrate multiple dies in a package either through vertical stacking or horizontal integration on a substrate along with a plethora of choices regarding configurations and processes. For chiplet-based designs, high-impact decisions such as those regarding the number of chiplets, the design partitions, the interconnect types, and other factors must be made early in the development process. In this work, we describe an open-source tool, CATCH, that can be used to guide these early design choices. We also present case studies showing some of the insights we can draw by using this tool. We look at case studies on optimal chip size, defect density, test cost, IO types, assembly processes, and substrates. Additionally, we include the cost breakdown for a specific case based on a prior work. Alexander Graening, Jonti Talukdar, Saptadeep Pal, Krishnendu Chakrabarty, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | ChipletPart: Cost-Aware Partitioning for 2.5D SystemsabstractIndustry adoption of chiplets has been growing as chiplets are a cost-effective option for making large, high-performance systems. Consequently, partitioning large systems into chiplets is increasingly important. In this work, we introduce ChipletPart —a cost-driven 2.5D system partitioner that addresses the unique constraints of chiplet systems, including complex objective functions, limited reach of inter-chiplet I/O transceivers, and the assignment of heterogeneous manufacturing technologies to different chiplets. ChipletPart integrates a sophisticated chiplet cost model with a genetic algorithm (GA)-based technology assignment and partitioning methodology, along with a simulated annealing (SA)-based chiplet floorplanner. Our results show that ChipletPart : (i) reduces chiplet cost by up to 58% (20% geometric mean) compared to state-of-the-art min-cut partitioners, which often yield floorplan-infeasible solutions; (ii) generates partitions with up to 47% (6% geometric mean) lower cost compared to the prior work Floorplet ; (iii) reduces chiplet cost up to 48% (30% geometric mean) compared to Chipletizer , while consistently producing I/O-feasible chiplet solutions across all testcases; and (iv) for the testcases we study, heterogeneous integration reduces cost by up to 43% (15% geometric mean) compared to homogeneous implementations. Additionally, we explore Bayesian optimization (BO) for finding low cost and floorplan-feasible chiplet solutions with technology assignments. On some testcases, our BO framework achieves better system cost (up to 5.3% improvement) with higher runtime overhead (up to 4×) compared to our GA-based framework. We also present case studies that show how changes in packaging and inter-chiplet signaling technologies can affect partitioning solutions. Finally, ChipletPart , the underlying chiplet cost model, and our chiplet testcase generator are available as open-source tools for the community. Alexander Graening, Puneet Gupta 0001, Andrew B. Kahng, Bodhisatta Pramanik, Zhiang Wang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | YAP: Yield Modeling and Simulation for Advanced PackagingabstractThree-dimensional integration technologies present a promising path forward for extending Moore’s law, facilitating high-density interconnects between chips and supporting multi-tier architectural designs. Cu-Cu hybrid bonding has emerged as a favored technique for the integration of chiplets at high interconnect density. This paper introduces YAP, a yield model for wafer-to-wafer (W2W) and die-to-wafer (D2W) hybrid bonding process. The model accounts for key failure mechanisms that contribute to yield loss, including overlay errors, particle defects, Cu recess variations, excessive wafer surface roughness, and Cu density. We also develop an open-source yield simulator and compare the accuracy of the near-analytical yield model with the simulation results. The results demonstrate that YAP achieves virtually identical accuracy while offering over 10,000x faster runtime. YAP enables the co-optimization of packaging technologies, assembly design rules, and overall design methodologies. We used YAP to examine the impact of bonding pitch, compare W2W and D2W hybrid bonding for varying chiplet sizes, and explore the benefits of tighter process controls, such as improved particle defect density. Puneet Gupta 0001 |
DAC | 2 |
| 2025 | FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingabstractWafer-scale systems are an emerging technology that tightly integrates high-end accelerator chiplets with high-speed wafer-scale interconnects, enabling low-latency and high-bandwidth connectivity.This makes them a promising platform for deep neural network (DNN) training.However, current network-on-wafer topologies, such as 2D Meshes, lack the flexibility needed to support various parallelization strategies effectively.In this paper, we propose Fred, a wafer-scale fabric architecture tailored to the unique communication needs of DNN training.Fred creates a distributed on-wafer topology with tiny microswitches, providing nonblocking connectivity for collective communications between arbitrary groups of accelerators and enabling in-switch collective support.Our results show that for sample parallelization strategies, Fred can improve the average end-to-end training time of ResNet-152, Transformer-17B, GPT-3, and Transformer-1T by 1.76×, 1.87×, 1.34×, and 1.4×, respectively, compared to a baseline wafer-scale Mesh. Saeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta 0001, Tushar Krishna |
ISCA | 4 |
| 2025 | Learned Approximate Computing: Algorithm Hardware Co-OptimizationabstractApproximate hardware trades acceptable error for improved performance and previous literature focuses on optimizing this tradeoff in the hardware. We show in this article that the application and the hardware can be co-optimized to achieve the best-quality-performance tradeoff. We propose LAC: learned approximate computing to optimize the algorithm and approximate hardware at the same time to maximize quality of output. Our approach allows automatic selection of approximate computing hardware while achieving similar quality as dedicated training for a single hardware configuration. Our improved training algorithm allows simultaneous hardware selection and application optimization without additional runtime overhead. Multihardware setup chooses a separate approximate hardware for each part of an application which allows for more hardware configurations and further improves quality. Egor Glukhov, Tianmu Li, Vaibhav Gupta, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | A Comparative Analysis of Low Temperature and Room Temperature Circuit OperationabstractLow-temperature (LT) conditions can potentially lead to lower power consumption and enhanced performance in circuit operations by reducing the transistor leakage current, increasing carrier mobility, reducing wear-out, and reducing interconnect resistance. We develop PROCEED-LT, a pathfinding framework to co-optimize devices and circuits over a wide performance range. Our results demonstrate that circuit operations at LT (−196 °C) reduce power compared to room temperature (RT, 85 °C) by$15\times $to over$23.8\times $depending on performance level. Alternatively, LT improves performance by$2.4\times $(high-power, high-performance)$- 7.0\times $(low-power, low-performance) at the same power point. These gains are further improved in low-activity circuits and when using multivoltage configurations. Meanwhile, we highlight the need for improvement in$V_{\text {th}}$variation to leverage benefits at cryogenic temperatures. Ali H. Hassan, Rhesa Muhammad Ramadhan, Yingheng Li, Chih-Kong Ken Yang, Sudhakar Pamarti, Puneet Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | Achieving DRAM-Like PCM by Trading Off Capacity for LatencyabstractPhase Change Memory (PCM) is considered one of the most promising scalable non-volatile main memory alternatives to DRAM. It provides$\sim$4x-5x cost per bit advantage over DRAM, thus enabling cost-effective dense main memory solution. However, PCM accesses are slower than DRAM, which leads to significantly poorer overall system performance (upto 80% higher execution time for memory intensive applications based on our analysis). To use PCM as a viable DRAM replacement, the performance gap between the two memory technologies has to be bridged, primarily by improving PCM read latency. In this work we propose an optimized PCM architecture, PCM-Duplicate, that trades off capacity to improve PCM read latency. In PCM-Duplicate, every row in the PCM subarray has a duplicate row. During memory read, both the rows are activated simultaneously. As a result, the bitline discharges through two PCM cells. This reduces the discharge time significantly, bringing down the overall sensing latency by$ \gt $3x compared to baseline PCM. While the overall PCM density benefit over DRAM halves, it still provides 2x more capacity than DRAM while having almost comparable read latency. PCM-Duplicate can either be used as low-cost DRAM main memory alternative or it can be used to replace the DRAM-based last level cache used in today's hybrid main memory systems for the slower PCM memories. Both these system options not only improve main memory capacity but also allow main memory based persistence by replacing DRAM and making the entire main memory non-volatile. Irina Alam, Puneet Gupta 0001 |
IEEE Trans. Computers | 2 |
| 2024 | SCIMITAR: Stochastic Computing In-Memory In-Situ Tracking ARchitecture for Event-Based CamerasabstractEvent-based cameras offer low latency and high-dynamic range imaging data in a sparse format that is well-suited for high-speed object tracking. Processing this sparse data in the same way as traditional camera data requires a great deal of unnecessary computation, making it difficult to take advantage of the high-effective frame rate for real-time processing. In this work, we propose an accelerator for high-speed object tracking on event-based camera data. SCIMITAR combines digital in-memory stochastic computing, in-situ stochastic stream generation, and multiple optimizations for utilizing input sparsity. SCIMITAR provides unparalleled performance with latency and energy that scale with sparsity. We demonstrate SCIMITAR performance on an object tracking application using circuit-level simulations of custom-designed compute-in-memory (CIM) macros and digital circuits. We achieve a frame processing rate of 26k frames/s with 100 regions-of-interest per frame and equivalent or better than state-of-the-art tracking accuracy. The accelerator achieves a peak throughput of 71 TOP/S and energy efficiency of 733 to 1702 TOP/S/W demonstrated on a range of event-based vision datasets, which is$5\times $higher than other CIM solutions. Wojciech Romaszkan, Jiyue Yang, Alexander Graening, Vinod Kurian Jacob, Jishnu Sen, Sudhakar Pamarti, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | DeepFlow: A Cross-Stack Pathfinding Framework for Distributed AI SystemsabstractOver the past decade, machine learning model complexity has grown at an extraordinary rate, as has the scale of the systems training such large models. However, there is an alarmingly low hardware utilization (5–20%) in large scale AI systems. The low system utilization is a cumulative effect of minor losses across different layers of the stack, exacerbated by the disconnect between engineers designing different layers spanning across different industries. To address this challenge, in this work we designed a cross-stack performance modelling and design space exploration framework. First, we introduce CrossFlow, a novel framework that enables cross-layer analysis all the way from the technology layer to the algorithmic layer. Next, we introduce DeepFlow (built on top of CrossFlow using machine learning techniques) to automate the design space exploration and co-optimization across different layers of the stack. We have validated CrossFlow’s accuracy with distributed training on real commercial hardware and showcase several DeepFlow case studies demonstrating pitfalls of not optimizing across the technology-hardware-software stack for what is likely the most important workload driving large development investments in all aspects of computing stack. Newsha Ardalani, Saptadeep Pal, Puneet Gupta 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Chiplets: How Small is too Small?abstractAs chiplet systems increase in popularity, it is important to revisit the tradeoffs for converting a monolithic design to a chiplet system. Chip yield, reusability, performance binning, and floorplanning push us toward smaller chiplets. Meanwhile, inter-chiplet interconnect and assembly overheads push us toward larger chips both in terms of power and cost. This work explores the impacts of these considerations on the minimum chiplet size that makes sense. We examine the case of a large design that could be built as a single monolithic system on chip (SoC) or as a system of chiplets and show that optimal chiplet size depends on a wide range of parameters. Our analysis indicates that the smallest chiplet sizes that are viable cost-wise depends both on technology node and on type of logic. The optimal point appears to be 50-150mm2in 40nm and 40-80mm2in 7nm for microprocessor type logic. For random logic, the optimal point increases beyond 200mm2in both cases. This makes the case for chipletization weaker in all but the largest SoCs. Alexander Graening, Saptadeep Pal, Puneet Gupta 0001 |
DAC | 3 |
| 2023 | PhotoFourier: A Photonic Joint Transform Correlator-Based Neural Network AcceleratorabstractThe last few years have seen a lot of work to address the challenge of low-latency and high-throughput convolutional neural network inference. Integrated photonics has the potential to dramatically accelerate neural networks because of its low-latency nature. Combined with the concept of Joint Transform Correlator (JTC), the computationally expensive convolution functions can be computed instantaneously (time of flight of light) with almost no cost. This ‘free’ convolution computation provides the theoretical basis of the proposed PhotoFourier JTC-based CNN accelerator. PhotoFourier addresses a myriad of challenges posed by on-chip photonic computing in the Fourier domain including 1D lenses and high-cost optoelectronic conversions. The proposed PhotoFourier accelerator achieves more than 28× better energy-delay product compared to state-of-art photonic neural network accelerators. Shurui Li 0002, Hangbo Yang, Chee Wei Wong, Volker J. Sorger, Puneet Gupta 0001 |
HPCA | 5 |
| 2023 | ReFOCUS: Reusing Light for Efficient Fourier Optics-Based Photonic Neural Network AcceleratorabstractIn recent years, there has been a significant focus on achieving low-latency and high-throughput convolutional neural network (CNN) inference. Integrated photonics offers the potential to substantially expedite neural networks due to its inherent low-latency properties. Recently, on-chip Fourier optics-based neural network accelerators have been demonstrated and achieved superior energy efficiency for CNN acceleration. By incorporating Fourier optics, computationally intensive convolution operations can be performed instantaneously through on-chip lenses at a significantly lower cost compared to other on-chip photonic neural network accelerators. This is thanks to the complexity reduction offered by the convolution theorem and the passive Fourier transforms computed by on-chip lenses. However, conversion overhead between optical and digital domains and memory access energy still hinder overall efficiency. Shurui Li 0002, Hangbo Yang, Chee Wei Wong, Volker J. Sorger, Puneet Gupta 0001 |
MICRO | 5 |
| 2023 | DRDebug: Automated Design Rule DebuggingabstractDesign rule checking (DRC) is an important step in the physical design flow that checks if a design meets the manufacturing constraints or design rules imposed by the process technology. It allows the foundry to ensure high acceptable manufacturing yield. Design rule verification is one of the most challenging steps because of the sheer size of the rule decks and the lack of standardization among these design rule manuals (DRMs). One way of efficiently discovering missed rule checks is by comparing the rule deck with that of another mature or well-established process. In this work, we develop two complementary techniques for comparing process design rule decks and automatically establishing a one-to-one correspondence between rules from two different process design kits (PDKs). The first approach, random layout generation (RLG), creates random layouts of different shapes and sizes. The generated layout is checked using both rule decks. The rules are then matched based on the violations generated. The second approach, based on rule language processing (RLP), matches rules based on the similarity between rule commands. Rules are directly matched based on the layer names and keywords present in the DRC commands. The two approaches are complementary and together they can correctly match more than 80% of the rules in two DRMs. Irina Alam, Tianmu Li, Sean Brock, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | REX-SC: Range-Extended Stochastic Computing Accumulation for Neural Network AccelerationabstractDeep learning has grown in capability and size in recent years, prompting research on alternative computing methods to cope with the increased compute cost. Stochastic computing (SC) promises higher compute efficiency with its compact compute units, but accuracy issues have prevented wide adoption, and accuracy-improving techniques have sacrificed runtime or training performance. In this work, we propose extended range SC—Range-Extended SC Accumulation to deal with the accuracy issues of SC. By modifying the functionality of OR-based SC accumulation, we increase SC computation accuracy without sacrificing the performance benefits. Our approach achieves a$2\times $reduction in stream length for the same accuracy compared to SC with OR-based accumulation and an up to$3.6\times $improvement in energy compared to SC with binary addition. With proper modeling, our approach improves training performance for SC-based neural networks and makes training SC models practical for large datasets like ImageNet. Tianmu Li, Wojciech Romaszkan, Sudhakar Pamarti, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | LAC: Learned Approximate ComputingabstractApproximate hardware trades acceptable error for improved performance and previous literature focuses on optimizing this trade-off in the hardware. We show in this paper that the application (i.e., the software) can be optimized for better accuracy without losing any performance benefits of the approximate hardware. We propose LAC: learned approximate computing as a method of tuning the application parameters to compensate for hardware errors. Our approach showed improvements across a variety of standard signal/image processing applications delivering an average improvement of 5.82db in PSNR and 0.23 in SSIM of the outputs. This translates to up to 87% power reduction and 83% area reduction for similar application quality. LAC allows the same approximate hardware to be used for multiple applications. Vaibhav Gupta, Tianmu Li, Puneet Gupta 0001 |
DATE | 3 |
| 2022 | COMET: On-die and In-controller Collaborative Memory ECC Technique for Safer and Stronger Correction of DRAM ErrorsabstractDRAM manufacturers have started adopting on-die error correcting coding (ECC) to deal with increasing error rates. The typical single error correcting (SEC) ECC on the memory die is coupled with a single-error correcting, double-error detecting (SECDED) ECC in the memory controller. Unfortunately, the on-die SEC can miscorrect double-bit errors (which would have been safely detected but uncorrected errors in conventional in-controller SECDED) resulting in triple bit errors more than 45% of the time. These are then miscorrected in the memory controller >55% of the time resulting in silent data corruption. We introduce COllaborative Memory ECC Technique (COMET), a novel method to efficiently design either the on-die or the in-controller ECC code, that, for the first time, will eliminate silent data corruption when a double-bit error happens within the DRAM. Further, we propose a collaboration mechanism between the on-die and in-controller ECC decoders that corrects most of the double-bit errors without adding any additional redundancy bits to either of the two codes. Overall, COMET can eliminate all double-bit error induced silent data corruptions and correct almost all (99.9997%) double-bit errors with negligible area, power, and performance impact. Irina Alam, Puneet Gupta 0001 |
DSN | 2 |
| 2022 | SASCHA - Sparsity-Aware Stochastic Computing Hardware Architecture for Neural Network AccelerationabstractStochastic computing (SC) has recently emerged as a promising method for efficient machine learning acceleration. Its high compute density, affinity with dense linear algebra primitives, and approximation properties have an uncanny level of synergy with the deep neural network computational requirements. However, there is a conspicuous lack of works trying to integrate SC hardware with sparsity awareness, which has brought significant performance improvements to conventional architectures. In this work, we identify why common sparsity-exploiting techniques are not easily applicable to SC accelerators and propose a new architecture—SASCHA—sparsity-aware SC hardware architecture for the neural network acceleration that addresses those issues. SASCHA encompasses a set of techniques that make utilizing sparsity in inference practical for different types of SC computation. At 90% weight sparsity, SASCHA can be up to$6.5\times $faster and$5.5\times $more energy-efficient than comparable dense SC accelerators with a similar area without sacrificing the dense network throughput. SASCHA also outperforms sparse fixed-point accelerators by up to$4\times $in terms of latency. To the best of our knowledge, SASCHA is the first SC accelerator architecture oriented around sparsity. Wojciech Romaszkan, Tianmu Li, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Designing a 2048-Chiplet, 14336-Core Waferscale ProcessorabstractWaferscale processor systems can provide the large number of cores, and memory bandwidth required by today’s highly parallel workloads. One approach to building waferscale systems is to use a chiplet-based architecture where pre-tested chiplets are integrated on a passive silicon-interconnect wafer. This technology allows heterogeneous integration and can provide significant performance and cost benefits. However, designing such a system has several challenges such as power delivery, clock distribution, waferscale-network design, design for testability and fault-tolerance. In this work, we discuss these challenges and the solutions we employed to design a 2048-chiplet, 14,336-core waferscale processor system. Saptadeep Pal, Irina Alam, Nick Cebry, Haris Suhail, Shi Bu, Subramanian S. Iyer, Sudhakar Pamarti, Rakesh Kumar 0002, Puneet Gupta 0001 |
DAC | 10 |
| 2021 | GEO: Generation and Execution Optimized Stochastic Computing Accelerator for Neural NetworksabstractStochastic computing (SC) has seen a renaissance in recent years as a means for machine learning acceleration due to its compact arithmetic and approximation properties. Still, SC accuracy remains an issue, with prior works either not fully utilizing the computational density or suffering from significant accuracy losses. In this work, we propose GEO - Generation and Execution Optimized Stochastic Computing Accelerator for Neural Networks, which optimizes stream generation and execution components of SC, and bridges the accuracy gap between stochastic computing and fixed-point neural networks. It improves accuracy by coupling controlled stream sharing with training and balancing OR and binary accumulations. GEO further optimizes the SC execution through progressive shadow buffering and architectural optimizations. GEO can improve accuracy compared to state-of-the-art SC by 2.2-4.0% points while being up to 4.4X faster and 5.3X more energy efficient. GEO eliminates the accuracy gap between SC and fixed-point architectures while delivering up to 5.6X higher throughput and 2.6X lower energy. Tianmu Li, Wojciech Romaszkan, Sudhakar Pamarti, Puneet Gupta 0001 |
DATE | 4 |
| 2020 | ACOUSTIC: Accelerating Convolutional Neural Networks through Or-Unipolar Skipped Stochastic ComputingabstractAs privacy and latency requirements force a move towards edge Machine Learning inference, resource constrained devices are struggling to cope with large and computationally complex models. For Convolutional Neural Networks, those limitations can be overcome by taking advantage of enormous data reuse opportunities and amenability to reduced precision. To do that however, a level of compute density unattainable for conventional binary arithmetic is required. Stochastic Computing can deliver such density, but it has not lived up to its full potential because of multiple underlying precision issues. We present ACOUSTIC: Accelerating Convolutions through Or-Unipolar Skipped sTochastIc Computing, an accelerator framework that enables fully stochastic, high-density CNN inference. Leveraging split-unipolar representation, OR-based accumulation and novel computation-skipping approach, ACOUSTIC delivers server-class parallelism within a mobile area and power budget - a 12mm2accelerator can be as much as 38.7x more energy efficient and 72.5x faster than conventional fixed-point accelerators. It can also be up to 79.6x more energy efficient than state-of-the-art stochastic accelerators. At the lower-end ACOUSTIC achieves 8x-120X inference throughput improvement with similar energy and area when compared to recent mixed-signal/neuromorphic accelerators. Wojciech Romaszkan, Tianmu Li, Tristan Melton, Sudhakar Pamarti, Puneet Gupta 0001 |
DATE | 5 |
| 2020 | Reverse Engineering for 2.5-D Split Manufactured ICsabstractIntegrated circuit (IC) split manufacturing has been shown to be one of the most effective protection schemes to prevent reverse engineering from malicious foundries. Among the existing split manufacturing approaches, the 2.5-D split manufacturing using silicon interposer has much less fabrication and testing costs compared to layer splitting approaches. In this article, we propose a Boolean satisfiability (SAT)-based attack to reconstruct the wire connections of the 2.5-D split manufacturing netlists. Our SAT-based attack can fully reconstruct the missing wires between modules with 100% correctness and therefore the functionality of the chip can be completely reverse engineered. In addition, we show that the runtime of attack is significantly reduced compared to existing 2.5-D split manufacturing SAT attacks by applying grouping hints obtained from a satisfiability modulo theories (SMTs)-based grouping algorithm, which is purely depending on the circuit functionality, so no physical defensive mechanisms can prevent such attack. In our experiments, we show that our grouping algorithm can speed up the existing SAT attack runtime by more than 1000× and can successfully reverse engineer reasonable size benchmarks even when the split nets contains more than one fanouts and the total cut size is close to 1000. Wei-Che Wang, Yizhang Wu, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | 3PXNet: Pruned-Permuted-Packed XNOR Networks for Edge Machine LearningabstractAs the adoption of Neural Networks continues to proliferate different classes of applications and systems, edge devices have been left behind. Their strict energy and storage limitations make them unable to cope with the sizes of common network models. While many compression methods such as precision reduction and sparsity have been proposed to alleviate this, they don’t go quite far enough. To push size reduction to its absolute limits, we combine binarization with sparsity in Pruned-Permuted-Packed XNOR Networks (3PXNet), which can be efficiently implemented on even the smallest of embedded microcontrollers. 3PXNets can reduce model sizes by up to 38X and reduce runtime by up to 3X compared with already compact conventional binarized implementations with less than 3% accuracy reduction. We have created the first software implementation of sparse-binarized Neural Networks, released as open source library targeting edge devices. Our library is complete with training methodology and model generating scripts, making it easy and fast to deploy. Wojciech Romaszkan, Tianmu Li, Puneet Gupta 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | SLATE: A Secure Lightweight Entity Authentication Hardware PrimitiveabstractLightweight cryptography has become more and more important in recent years because of the rise of the Internet of Things (IoT) and usage of smart mobile devices. In this paper, we propose a novel secure lightweight entity authentication hardware primitive called SLATE, where its area is about 50% to more than 3X smaller than existing lightweight ciphers and strong physical unclonable functions (PUFs), respectively. Even though the authentication of SLATE is done through challenge response pair (CRP) verification similar to strong PUFs, the source of the key for SLATE must be coming from any existing secret key storage used for any ciphers. A main advantage of SLATE over most existing strong PUFs being an entity authentication primitive is that SLATE is resistant to known attacks to strong PUFs or logic obfuscations, such as model building attacks and Boolean satisfiability (SAT) attacks. Furthermore, we show that the implementation cost of SLATE with a 176-bit key and 244 CRPs is only 663 gate equivalents (GEs). Compared with lightweight ciphers and existing secure strong PUFs, we show that SLATE is a practical security primitive for resource constrained systems for its extremely small footprint and security. Finally, we show that SLATE is information theoretically secure when valid CRPs are communicated through insecure channels. Wei-Che Wang, Yair Yona, Yizhang Wu, Suhas N. Diggavi, Puneet Gupta 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | Design Space Exploration for Chiplet-Assembly-Based ProcessorsabstractRecent advancements in 2.5-D integration technologies have made chiplet assembly a viable system design approach. Chiplet assembly is emerging as a new paradigm for heterogeneous design at lower cost, design effort, and turnaround time and enables low-cost customization of hardware. However, the success of this approach depends on identifying a minimum chiplet set which delivers these benefits. We develop the first microarchitectural design space exploration framework for chiplet assembly-based processors which enables us to identify the minimum set of chiplets to design and manufacture. Since chiplet assembly makes heterogeneous technology and cost-effective application-dependent customization possible, we show the benefits of using multiple systems built from multiple chiplets to service diverse workloads (up to 35% improvement in energy-delay product over a single best system) and advantages of chiplet assembly approaches over system-on-chip (SoC) methodology in terms of total cost (up to 72% improvement in cost) while satisfying the energy and performance constraints of individual applications. Saptadeep Pal, Daniel Ruelas-Petrisko, Rakesh Kumar 0002, Puneet Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Architecting Waferscale Processors - A GPU Case StudyabstractIncreasing communication overheads are already threatening computer system scaling. One approach to dramatically reduce communication overheads is waferscale processing. However, waferscale processors [1], [2], [3] have been historically deemed impractical due to yield issues [1], [4] inherent to conventional integration technology. Emerging integration technologies such as Silicon-Interconnection Fabric (Si-IF) [5], [6], [7], where pre-manufactured dies are directly bonded on to a silicon wafer, may enable one to build a waferscale system without the corresponding yield issues. As such, waferscalar architectures need to be revisited. In this paper, we study if it is feasible and useful to build today's architectures at waferscale. Using a waferscale GPU as a case study, we show that while a 300 mm wafer can house about 100 GPU modules (GPM), only a much scaled down GPU architecture with about 40 GPMs can be built when physical concerns are considered. We also study the performance and energy implications of waferscale architectures. We show that waferscale GPUs can provide significant performance and energy efficiency advantages (up to 18.9x speedup and 143x EDP benefit compared against equivalent MCM-GPU based implementation on PCB) without any change in the programming model. We also develop thread scheduling and data placement policies for waferscale GPU architectures. Our policies outperform state-of-art scheduling and data placement policies by up to 2.88x (average 1.4x) and 1.62x (average 1.11x) for 24 GPM and 40 GPM cases respectively. Finally, we build the first Si-IF prototype with interconnected dies. We observe 100% of the inter-die interconnects to be successfully connected in our prototype. Coupled with the high yield reported previously for bonding of dies on Si-IF, this demonstrates the technological readiness for building a waferscale GPU architecture. Saptadeep Pal, Daniel Ruelas-Petrisko, Matthew Tomei, Puneet Gupta 0001, Subramanian S. Iyer, Rakesh Kumar 0002 |
HPCA | 4 |
| 2019 | Context-Aware Resiliency: Unequal Message Protection for Random-Access MemoriesabstractA common way to protect data stored in DRAM and related memory systems is through the use of an error-correcting code such as the extended Hamming code. Traditionally, these error-correcting codes provide equal protection guarantees to all messages. In this paper, we focus on unequal message protection (UMP), in which a subset of messages is deemed as special, and is afforded additional error-correction protection while maintaining the same number of redundancy bits as the baseline code. UMP is a powerful approach when the special messages are chosen based on the knowledge of data patterns in context. Our objective is to construct deterministic, algebraic codes with guaranteed UMP properties, derive their cardinality bounds using novel combinatorial techniques, and to demonstrate their efficacy for realistic memory benchmarks. We first introduce a UMP alternative to the single-bit parity-check code, and then we generalize to a broader UMP code family, including a UMP alternative to the extended Hamming code, offering full double-error correction protection to special messages. Our UMP constructions, applied to main memory in high-performance computing applications, could lead to significant system-level benefits such as less frequent checkpoints in supercomputers and decreased risk of catastrophic failure from erroneous special messages. Clayton Schoeny, Frederic Sala, Mark Gottscho, Irina Alam, Puneet Gupta 0001, Lara Dolecek |
IEEE Trans. Inf. Theory | 5 |
| 2018 | A Case for Packageless ProcessorsabstractDemand for increasing performance is far outpacing the capability of traditional methods for performance scaling. Disruptive solutions are needed to advance beyond incremental improvements. Traditionally, processors reside inside packages to enable PCB-based integration. We argue that packages reduce the potential memory bandwidth of a processor by at least one order of magnitude, allowable thermal design power (TDP) by up to 70%, and area efficiency by a factor of 5 to 18. Further, silicon chips have scaled well while packages have not. We propose packageless processors - processors where packages have been removed and dies directly mounted on a silicon board using a novel integration technology, Silicon Interconnection Fabric (Si-IF). We show that Si-IF-based packageless processors outperform their packaged counterparts by up to 58% (16% average), 136%(103% average), and 295% (80% average) due to increased memory bandwidth, increased allowable TDP, and reduced area respectively. We also extend the concept of packageless processing to the entire processor and memory system, where the area footprint reduction was up to 76%. Saptadeep Pal, Daniel Ruelas-Petrisko, Adeel Ahmad Bajwa, Puneet Gupta 0001, Subramanian S. Iyer, Rakesh Kumar 0002 |
HPCA | 4 |
| 2018 | Error Correction and Detection for Computing Memories Using System Side InformationabstractError correction and detection are the core components of all modern memory systems. Current computing memory systems use simple coding schemes to simultaneously meet the resiliency and latency requirements. In this paper, we review our recent results on context-aware coding for computing memories, an approach that explicitly takes into account various intrinsic side information for improved robustness to faults. We discuss both error correction and detection, codes' theoretical properties, and provide examples of how these solutions can be implemented in practice. We explicitly describe the special case of the error localization codes. We also discuss promising future directions and connections with classical information theoretic concepts. Clayton Schoeny, Irina Alam, Mark Gottscho, Puneet Gupta 0001, Lara Dolecek |
ITW | 4 |
| 2018 | Assessing Layout Density Benefits of Vertical Channel DevicesabstractVertical channel devices have been considered as promising candidates for sub-5 nm regime for the reduced area and large driving current. Several styles of layout designs and fabrication details of vertical channel devices have been proposed. However, due to the fast-changing manufacturing constraints for the advanced devices, the most efficient layout structures are still yet to be explored. In this paper, we study the efficiency in terms of cell area of in-bound power vertical channel device layout, which is potentially the most compact vertical layout style. We develop and implement an efficient vertical layout generation framework for in-bound power layout to provide a quick evaluation of cell area given design rules and choices of folding strategies. The results are compared to vertically stacked lateral channel devices and out-bound power vertical channel devices. Both cell-level and chip-level comparisons show that in-bound power layout is more area-efficient than lateral devices and vertical out-bound power layouts. Wei-Che Wang, Charles Zhao, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Design and Analysis of Stability-Guaranteed PUFsabstractThe lack of stability is one of the limitations that constrain physical unclonable function (PUF) from being put in widespread practical use. In this paper, we propose a weak PUF and a strong PUF that are both completely stable. These PUFs are called locally enhanced defectivity physical unclonable function (LEDPUF). An LEDPUF is a pure functional PUF that does not require any kinds of correction schemes as conventional parametric PUFs do. The source of randomness of an LEDPUF is extracted from locally enhance defectivity without affecting other parts of the chip. In this paper, we construct a weak LEDPUF by forming arrays of directed self-assembly random connections, and the strong LEDPUF is implemented by using the weak LEDPUF as the key of a keyed-hash message authentication code. Our simulation and statistical results show that the entropy of the weak LEDPUF bits is close to ideal, and the inter-chip Hamming distances of both weak and strong LEDPUFs are about 50%, which means that these LEDPUFs are not only stable but also unique. We develop a new unified framework for evaluating the security of PUFs, based on password security, by using information theoretic tools of guesswork. The guesswork model allows us to quantitatively compare, with a single unified metric, PUFs with varying levels of stability, bias, and available side information. In addition, it generalizes other measures to evaluate the security level, such as min-entropy and mutual information. We evaluate the guesswork-based security of some measured static random access memory and ring oscillator PUFs as an example and compare them with an LEDPUF to show that the stability has a more severe impact on the PUF security than biased responses. Furthermore, we find the guesswork of two new problems: guesswork under the probability of attack failure and the guesswork of strong PUFs that are used for authentication. Wei-Che Wang, Yair Yona, Suhas N. Diggavi, Puneet Gupta 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Hybrid VC-MTJ/CMOS non-volatile stochastic logic for efficient computingabstractIn this paper, we propose a non-volatile stochastic computing (SC) scheme using voltage-controlled magnetic tunnel junction (VC-MTJ) and negative differential resistance (NDR). The proposed design includes a VC-MTJ based true stochastic bit stream generator and VC-MTJ and NDR based stochastic adder, multiplier, register, which are experimentally demonstrated using 60nm VC-MTJ and CMOS NDR connected on die. These components are then used to realize FIR filter and AdaBoost (machine-learning algorithm). 3X–37X energy advantage is shown for the proposed SC compared with CMOS binary arithmetic ASIC and SC designs. Shaodi Wang, Saptadeep Pal, Tianmu Li, Andrew Pan, Cecile Grezes, Pedram Khalili Amiri, Kang L. Wang, Puneet Gupta 0001 |
DATE | 8 |
| 2017 | Context-aware resiliency: Unequal message protection for random-access memoriesabstractA common way to protect data stored in DRAM and related memory systems is through the use of a single-error-correcting/double-error-detecting (SECDED) code. Traditionally, these error-correcting codes provide equal protection guarantees to all messages. In a recent work, we demonstrated enhanced error correction capabilities for SECDED codes by taking into account contextual side-information about the data. This paper is concerned with a closely related scenario: unequal message protection (UMP), where a subset of special messages is afforded additional error-correction ability. UMP is relevant to computing systems where certain messages are critical and failures cannot be tolerated. We study practical UMP constructions where messages are guaranteed either one or two bit-error-correction. We provide upper and lower bounds on the number of special messages. We introduce an explicit and practical code construction based on BCH subcodes and demonstrate the efficacy of our technique on data from the AxBench and SPEC CPU2006 benchmark suites. Clayton Schoeny, Frederic Sala, Mark Gottscho, Irina Alam, Puneet Gupta 0001, Lara Dolecek |
ITW | 5 |
| 2017 | Mask Assignment and DSA Grouping for DSA-MP Hybrid Lithography for Sub-7 nm Contact/Via HolesabstractDirected self assembly (DSA) is a very promising candidate for the sub-7 nm technology nodes. To print such small dimensions, multiple patterning (MP) is likely to be used to print the guiding templates for DSA. Therefore, algorithms are required to perform the DSA grouping at the same time as the mask assignment. In this paper, we present an optimal integer linear program (ILP) to solve this problem for two schemes of hybrid DSA-MP process. Scalable heuristic algorithms are also proposed to solve the same problem. In comparison to the ILP, the proposed heuristics are 4×-213× faster, and result in an increase of total number of violations by 4%-29%. Yasmine Badr, Andres Torres, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Benchmarking of Mask Fracturing HeuristicsabstractAggressive resolution enhancement techniques such as inverse lithography (ILT) often lead to complex, nonrectilinear mask shapes which make mask writing extremely slow and expensive. To reduce shot count of complex mask shapes, mask writers allow overlapping shots, due to which the problem of fracturing mask shapes with minimum shot count is NP-hard. The need to account for e-beam proximity effect makes mask fracturing even more challenging. Although a number of fracturing heuristics have been proposed, there has been no systematic study to analyze the quality of their solutions. In this paper, we first propose a method to generate tight upper and lower bounds for actual ILT mask shapes by formulating mask fracturing as an integer linear program and solving it using branch and price. Since the integer program requires significant computational resources to compute reasonable bounds, we propose a new method to generate benchmarks with known optimal solutions, that can be used to evaluate the suboptimality of mask fracturing heuristics. To make the generated benchmark shapes realistic, we further propose a novel automated benchmark generation method that takes any ILT shape as input and returns a benchmark shape which looks similar to the input shape and for which the optimal fracturing solution is known. Using these methods, we compare the suboptimality of four mask fracturing heuristics. Our results show that even a state-of-the-art prototype (version of) capability within a commercial EDA tool for e-beam mask shot decomposition can be suboptimal by as much as 2.6× for real ILT shapes and by 6.0× for generated benchmarks. Tuck-Boon Chan, Puneet Gupta 0001, Kwangsoo Han, Abde Ali Kagalwalla, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Assessing Benefits of a Buried Interconnect Layer in Digital DesignsabstractIn sub-15 nm technology nodes, local metal layers have witnessed extremely high congestion leading to pin-access-limited designs, and hence affecting the chip area and related performance. In this paper, we assess the benefits of adding a buried interconnect layer below the device layers for the purpose of reducing cell area, improving pin access, and reducing chip area. After adding the buried layer to a projected 7 nm standard cell library, results show ~9%-13% chip area reduction and 126% pin access improvement. This shows that buried interconnect, as an integration primitive, is very promising as an alternative method to density scaling. Liheng Zhu, Yasmine Badr, Shaodi Wang, Subramanian S. Iyer, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Low-Cost Memory Fault Tolerance for IoT DevicesabstractIoT devices need reliable hardware at low cost. It is challenging to efficiently cope with both hard and soft faults in embedded scratchpad memories. To address this problem, we propose a two-step approach: FaultLink and Software-Defined Error-Localizing Codes (SDELC). FaultLink avoids hard faults found during testing by generating a custom-tailored application binary image for each individual chip. During software deployment-time, FaultLink optimally packs small sections of program code and data into fault-free segments of the memory address space and generates a custom linker script for a lazy-linking procedure. During run-time, SDELC deals with unpredictable soft faults via novel and inexpensive Ultra-Lightweight Error-Localizing Codes (UL-ELCs). These require fewer parity bits than single-error-correcting Hamming codes. Yet our UL-ELCs are more powerful than basic single-error-detecting parity: they localize single-bit errors to a specific chunk of a codeword. SDELC then heuristically recovers from these localized errors using a small embedded C library that exploits observable side information (SI) about the application’s memory contents. SI can be in the form of redundant data (value locality), legal/illegal instructions, etc. Our combined FaultLink+SDELC approach improves min-VDD by up to 440 mV and correctly recovers from up to 90% (70%) of random single-bit soft faults in data (instructions) with just three parity bits per 32-bit word. Mark Gottscho, Irina Alam, Clayton Schoeny, Lara Dolecek, Puneet Gupta 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2017 | A Word Line Pulse Circuit Technique for Reliable Magnetoelectric Random Access MemoryabstractA word line pulse (WLP) circuit scheme is proposed toward the implementation of magnetoelectric random access memory (MeRAM). The circuit improves the write error rate (WER) and cell area efficiency by generating a better write pulse compared to conventional bitline pulse (BLP) techniques in terms of the pulse slew rate and amplitude. For the voltage-controlled magnetic anisotropy-induced precessional switching of the magnetic tunnel junction (MTJ), the write pulse shape has a large impact on the switching probability. Typically, a square shape pulse results in higher switching probability compared to that of a triangular shape pulse with long rise and falling edges, since the square shape pulse causes a more stable precessional trajectory of the free layer magnetization by providing a relatively constant in-plane-dominant effective field. Compared to the BLP scheme, the WLP can generate a better square shape pulse by eliminating discharge paths under the pulse condition, using the gain of the access transistor, and effectively diminishing the capacitive loading which needs to be driven. A macrospin compact model of voltage-controlled MTJ shows that the WLP can improve WER by${10}^{7}$times and allow MeRAM to have four-time improvement in area efficiency of driver circuits compared to the BLP. Hochul Lee, Shaodi Wang, Farbod Ebrahimi, Puneet Gupta 0001, Pedram Khalili Amiri, Kang L. Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Hardware Reliability margining for the dark silicon eraabstractHardware reliability margin should be derived from the worst-case aging scenario, which typically occurs when the circuits are operating at peak performance state with the highest operating voltage and frequency. However, as integrated circuits enter the “dark silicon” era, it is impossible to power up all circuits throughout the entire lifetime. Reliability margining in absence of architecture-level power/thermal constraints can be overly pessimistic. In this work, we propose a margining scheme that employs the power/thermal contexts and system management policies to derive the actual worst-case workload pattern for different reliability phenomena. Our experiment results show that at 60% dark ratio, conventional margining approach can overestimate the aging degradation due to EM and BTI by up to 3-7X and 18% respectively. Our margining method is able to eliminate these over-pessimism and results in about 20% delay margin and 40%-60% metal width margin reduction. Liangzhen Lai, Puneet Gupta 0001 |
ASP-DAC | 2 |
| 2016 | MTJ variation monitor-assisted adaptive MRAM writeabstractSpin-transfer torque random access memory (STT-RAM) and magnetoelectric random access memory (MeRAM) are promising non-volatile memory technologies. But STT-RAM and Me RAM both suffer from high write error rate due to thermal fluctuation of magnetization. Temperature and wafer-level process variation significantly exacerbate these problems. In this paper, we propose a design that adaptively selects optimized write pulse for STT-RAM and MeRAM to overcome ambient process and temperature variation. To enable the adaptive write, we design specific MTJ-based variation monitor, which precisely senses process and temperature variation. The monitor is over 10X faster, 5X more energy-efficient, and 20X smaller compared with conventional thermal monitors of similar accuracy. With adaptive write, the write latency of STT-RAM and MeRAM cache are reduced by up to 17% and 59% respectively, and application run time is improved by up to 41%. Shaodi Wang, Hochul Lee, Cecile Grezes, Pedram Khalili Amiri, Kang L. Wang, Puneet Gupta 0001 |
DAC | 6 |
| 2016 | Multi-story power distribution networks for GPUs
Liangzhen Lai, Mark Gottscho, Puneet Gupta 0001 |
DATE | 4 |
| 2016 | X-Mem: A cross-platform and extensible memory characterization tool for the cloudabstractEffective use of the memory hierarchy is crucial to cloud computing. Platform memory subsystems must be carefully provisioned and configured to minimize overall cost and energy for cloud providers. For cloud subscribers, the diversity of available platforms complicates comparisons and the optimization of performance. To address these needs, we present X-Mem, a new open-source software tool that characterizes the memory hierarchy for cloud computing. Mark Gottscho, Sriram Govindan, Bikash Sharma, Mohammed Shoaib, Puneet Gupta 0001 |
ISPASS | 5 |
| 2016 | Efficient Layout Generation and Design Evaluation of Vertical Channel DevicesabstractVertical gate-all-around (VGAA) structure has been shown to be one of the most promising devices for the scaling beyond 10 nm for its reduced area, large driving current, and good gate control. Moreover, emerging devices such as heterojunction tunneling FETs are more amenable to vertical fabrication. However, past studies of vertical channel devices focused more on regular memory architectures and simple standard cells like inverters. Since naïve migration of regular FinFET layouts to vertical FETs yields little benefits, we identify several vertical efficient layout structures and propose novel layout generation heuristics for vertical channel devices. We also compare VGAA with symmetric and asymmetric source/drain architectures and different contact placement strategies. The layout efficiencies of several VGAA structures, vertical double-gate, lateral gate-all-around (LGAA), and FinFET are presented in our experiments. Routing congestion estimation on both cell-level and chip-level after placement and routing are also presented. We observe that even though most vertical channel standard cells have more diffusion gaps than lateral cells do, they still benefit from vertical architectures in area because of the vertically aligned top contacts. For asymmetric architectures, the area is larger than symmetric architectures because of the extra diffusion gaps needed, but our experiments indicate that for both symmetric and asymmetric architectures, vertical channel devices are likely to have a density advantage over lateral channel devices. Wei-Che Wang, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | MEMRES: A Fast Memory System Reliability SimulatorabstractWith scaling technology, emerging nonvolatile devices, and data-intensive applications, memory faults have become a major reliability concern for computing systems. With various hardware and software approaches proposed to address this issue, a comprehensive evaluation is required to understand the effectiveness of these solutions. Considering the complex nature of various memory faults as well as interactions between various correction mechanisms, we propose MEMRES, a fast main memory system reliability simulator. It enables memory fault simulation with error-correcting code (ECC) algorithms and modern memory reliability management, including memory page retirement, mirroring, scrubbing, and hardware sparing. MEMRES is computationally efficient in obtaining memory failure probabilities in the presence of multiple failure mechanisms and complex correction scheme, allowing the optimization of memory system reliability, the prediction of emerging memory reliability, and designing a reliability enhancement technique. The accuracy of MEMRES is verified by an existing analytical model and an existing memory fault simulator. We performed a case study on spin-transfer torque random access memory (STT-RAM)-based main memory, and the results indicate that in-memory ECC can significantly mitigate the write error rate of STT-RAM, demonstrating the capability of handling emerging memory system. Shaodi Wang, Henry Chaohong Hu, Hongzhong Zheng, Puneet Gupta 0001 |
IEEE Trans. Reliab. | 4 |
| 2016 | An Evaluation Framework for Nanotransfer Printing-Based Feature-Level Heterogeneous Integration in VLSI CircuitsabstractWe develop an evaluation framework to assess the potential benefits of feature-level heterogeneous integration (HGI) in nanoscale VLSI circuits. We study, for the first time, the impact of HGI on circuit delay, layout area, and power by comparing the integration of 15-nm InGaAs and Ge FinFETs via nanotransfer printing with the baseline Si-only FinFET technology. To properly account for the performance, power, and area tradeoffs, we perform comprehensive evaluations, including synthesis, placement, and routing of digital circuit benchmarks. We show the circuits designed with an HGI exhibit lower delay and power due to improved device performance at the cost of larger area induced by misalignment errors. We also demonstrate that the HGI misalignment area penalties can be drastically reduced using posttransfer fin trimming. Our findings provide substantial motivation for industry to explore HGI as a technology route for the post-Si era. Greg Leung, Shaodi Wang, Andrew Pan, Puneet Gupta 0001, Chi On Chui |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | PROCEED: A Pareto Optimization-Based Circuit-Level Evaluator for Emerging DevicesabstractEvaluation of novel devices in the context of circuits is crucial to identifying and maximizing their value. We propose a new framework, Pareto optimization-based circuit-level evaluator for emerging device (PROCEED), that uses comprehensive performance, power, and area metrics for accurate device-circuit coevaluation through optimization of digital circuit benchmarks. The PROCEED assesses technology suitability over a wide operating region (megahertz to gigahertz) by leveraging available circuit knobs (threshold voltage assignment, power management, sizing, and so on). It improves the benchmark accuracy by 3x to 115x compared with the existing methods while offering orders of magnitude improvements in runtime over full physical design implementation flows. To illustrate the PROCEED's capabilities, we deploy it to assess emerging technologies, including novel tunneling field-effect transistors, compared with conventional silicon CMOS. As a further illustration, we extend PROCEED to evaluate future heterogeneous integration of varied devices onto the same silicon substrate. Shaodi Wang, Andrew Pan, Chi On Chui, Puneet Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Evaluating and exploiting impacts of dynamic power management schemes on system reliabilityabstractHardware reliability has been a major concern for nano-scale computing systems. Different hardware design choices, application workloads and software management schemes can jointly affect the system's resilience. In this paper, we first develop a hardware evaluation platform based on an embedded/mobile development board and standard Linux kernel. We demonstrate the use of our platform to evaluate the system's power and radiation-induced soft error rate in presence of system power management schemes and with different application workloads and various hardware design configurations. We also propose system/cloud-based virtual sensing to capture varying ambient conditions for reliability evaluation. New reliability management policies are proposed and implemented in Linux kernel to exploit the flexibility in different existing power management schemes. We demonstrate that our policies can achieve the system reliability target under varying application workloads and ambient conditions. Experiments show that our policies are efficient and with less than 3% additional power overhead compared to the optimal schemes characterized offline. Liangzhen Lai, Vikas Chandra, Puneet Gupta 0001 |
CASES | 3 |
| 2015 | Mask assignment and synthesis of DSA-MP hybrid lithography for sub-7nm contacts/viasabstractIntegrating Directed Self Assembly (DSA) and Multiple Patterning (MP) is an attractive option for printing contact and via layers for sub-7nm process nodes. In the DSA-MP hybrid process, an optimized decomposition algorithm is required to perform the MP mask assignment while considering the DSA advantages and limitations. In this paper, we present an optimal Integer Linear Programming (ILP) formulation for the simultaneous DSA grouping and MP decomposition problem for contacts and vias. Then we propose a heuristic and develop an efficient algorithm for solving the same problem. In comparison to the optimal ILP results, the proposed algorithm is 197x faster and results in 16.3% more violations. The proposed algorithm produces 56% fewer violations than the sequential approaches which perform DSA grouping followed by MP decomposition and vice versa. Yasmine Badr, Andres Torres, Puneet Gupta 0001 |
DAC | 3 |
| 2015 | Effective model-based mask fracturing for mask cost reductionabstractThe use of aggressive resolution enhancement techniques like multiple patterning and inverse lithography (ILT) has led to expensive photomasks. Growing mask write time has been a key reason for the cost increase. Moreover, due to scaling, e-beam proximity effects can no longer be ignored. Model-based mask fracturing has emerged as a useful technique to address these critical challenges by allowing overlapping shots and compensating for proximity effects during fracturing itself. However, it has been shown recently that heuristics for model-based mask fracturing can be suboptimal by more than 1:6x on average for ten real ILT shapes, highlighting the need for better heuristics. In this work, we propose a new model-based mask fracturing method that significantly outperforms all the previously reported heuristics. The number of e-beam shots of our method is 23% less than a state-of-the-art prototype version of capability within a commercial EDA tool for e-beam mask shot decomposition (PROTO-EDA) for ten ILT mask shapes. Moreover, our method has an average runtime of less than 1:4s per shape. Abde Ali Kagalwalla, Puneet Gupta 0001 |
DAC | 2 |
| 2015 | Cyberphysical-system-on-chip (CPSoC): a self-aware MPSoC paradigm with cross-layer virtual sensing and actuation
Santanu Sarma, Nikil Dutt, Puneet Gupta 0001, Nalini Venkatasubramanian, Alexandru Nicolau |
DATE | 3 |
| 2015 | DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale EraabstractFault-Tolerant Voltage-Scalable (FTVS) SRAM cache architectures are a promising approach to improve energy efficiency of memories in the presence of nanoscale process variation. Complex FTVS schemes are commonly proposed to achieve very low minimum supply voltages, but these can suffer from high overheads and thus do not always offer the best power/capacity trade-offs. We observe on our 45nm test chips that the “fault inclusion property” can enable lightweight fault maps that support multiple runtime supply voltages. Based on this observation, we propose a simple and low-overhead FTVS cache architecture for power/capacity scaling. Our mechanism combines multilevel voltage scaling with optional architectural support for power gating of blocks as they become faulty at low voltages. A static (SPCS) policy sets the runtime cache VDD once such that a only a few cache blocks may be faulty in order to minimize the impact on performance. We describe a Static Power/Capacity Scaling (SPCS) policy and two alternate Dynamic Power/Capacity Scaling (DPCS) policies that opportunistically reduce the cache voltage even further for more energy savings. This architecture achieves lower static power for all effective cache capacities than a recent more complex FTVS scheme. This is due to significantly lower overheads, despite the inability of our approach to match the min-VDD of the competing work at a fixed target yield. Over a set of SPEC CPU2006 benchmarks on two system configurations, the average total cache (system) energy saved by SPCS is 62% (22%), while the two DPCS policies achieve roughly similar energy reduction, around 79% (26%). On average, the DPCS approaches incur 2.24% performance and 6% area penalties. Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001 |
ACM Trans. Archit. Code Optim. | 5 |
| 2015 | ViPZonE: Hardware Power Variability-Aware Virtual Memory Management for Energy SavingsabstractHardware variability is predicted to increase dramatically over the coming years as a consequence of continued technology scaling. In this paper, we apply the Underdesigned and Opportunistic Computing (UnO) paradigm by exposing system-level power variability to software to improve energy efficiency. We present ViPZonE, a memory management solution in conjunction with application annotations that opportunistically performs memory allocations to reduce DRAM energy. ViPZonE's components consist of a physical address space with DIMM-aware zones, a modified page allocation routine, and a new virtual memory system call for dynamic allocations from userspace. We implemented ViPZonE in the Linux kernel with GLIBC API support, running on a real x86-64 testbed with significant access power variation in its DDR3 DIMMs. We demonstrate that on our testbed, ViPZonE can save up to 27.80 percent memory energy, with no more than 4.80 percent performance degradation across a set of PARSEC benchmarks tested with respect to the baseline Linux software. Furthermore, through a hypothetical “what-if” extension, we predict that in future non-volatile memory systems which consume almost no idle power, ViPZonE could yield even greater benefits, demonstrating the ability to exploit memory hardware variability through opportunistic software. Mark Gottscho, Luis Angel D. Bathen, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001 |
IEEE Trans. Computers | 5 |
| 2014 | Comprehensive die-level assessment of design rules and layoutsabstractCo-development of design rules and layout methodologies is the key to successful adoption of a technology. In this work, we propose Chip-level Design Rule Evaluator (ChipDRE), the first framework for systematic evaluation of design rules and their interaction with layouts, performance, margins and yield at the chip scale (as opposed to standard cell-level). A “good chips per wafer” metric is used to unify area, performance, variability and yield. The framework uses a generated virtual standard-cell library coupled with a mix of physical design, semi-empirical, and machine-learning-based models to estimate area and delay at the chip level. The result is a unified design-quality estimate that can be computed fast enough to allow using ChipDRE to optimize a large number of complex design rules. For instance, a study of well-to-active spacing rule reveals a non-monotone dependence of rule value to chip area (although the dependence to cell area is monotone) due to delay changes coming from well-proximity effect. Rani S. Ghaida, Yasmine Badr, Mukul Gupta, Puneet Gupta 0001 |
ASP-DAC | 5 |
| 2014 | EUV-CDA: Pattern shift aware critical density analysis for EUV mask layoutsabstractDespite the use of mask defect avoidance and mitigation techniques, finding a usable defective mask blank remains a challenge for Extreme Ultraviolet Lithography (EUVL) at sub-10nm node due to dense layouts and low CD tolerance. In this work, we propose a pattern shift-aware metric called critical density, which can quickly evaluate the robustness of EUV layouts to mask defects (300-1300x faster than Monte Carlo, with average mask yield root mean square error (RMSE) ranging from 0.08%-6.44%), thereby enabling design-level mask defect mitigation techniques. Our experimental results indicate that reducing layout regularity improves the ability of layouts to tolerate mask defects via pattern shift. Abde Ali Kagalwalla, Michale Lam, Kostas Adam, Puneet Gupta 0001 |
ASP-DAC | 4 |
| 2014 | Accurate and inexpensive performance monitoring for variability-aware systemsabstractDesigning reliable integrated systems has become a major challenge with shrinking geometries, increasing fault rates and devices which age substantially in their usage life. The proposed research is motivated by the observation that many of the in-field failures are delay failures and several variability signatures are also delay-related. The origins of temporal delay fluctuations include manufacturing variability, voltage/temperature changes, negative or positive bias temperature instability-related Vth degradation, etc. Since the actual delay changes depend on process variations as well as workload, on-chip monitoring may be the best way of predicting them. There is a need to monitor circuit performance during manufacturing as well as at runtime to predict achievable performance and warn against impending failures. Adaptive mechanisms in hardware and/or software can optimize the trade-off between errors, energy and performance based on the feedback from runtime circuit performance monitors. This paper presents approaches for automated synthesis of design-dependent performance monitors. These monitors can be used to predict impending delay failures relatively inexpensively. For low-overhead monitoring, we propose multiple design-dependent ring oscillators (DDROs) as smart canary structures which can reliably predict achievable chip frequency but with margins for local variations. Early silicon results indicate that DDROs can reduce delay monitoring error by 35% compared to conventional ring oscillators. To further improve the prediction (albeit at a higher overhead), we propose in-situ slack monitors (SlackProbe) which can match local variations as well at overheads much smaller than monitoring all sequential elements. SlackProbe reduces the number of monitors required by over 15X with 5% additional delay margin in several commercial processor benchmarks. Finally, we show an example of software testbed that demonstrates a variability-aware system that utilizes the hardware monitors and operates with both hardware and software adaptation. Liangzhen Lai, Puneet Gupta 0001 |
ASP-DAC | 2 |
| 2014 | PROCEED: A pareto optimization-based circuit-level evaluator for emerging devicesabstractEvaluation of novel devices in a circuit context is crucial to identifying and maximizing their value. We propose a new framework, PROCEED, and metrics for accurate device-circuit co-evaluation through proper optimization of digital circuit benchmarks. PROCEED assesses technology suitability over a wide operating region (MHz to GHz) by leveraging available circuit knobs (Vtassignment, power management, sizing, etc.) and improves accuracy by 3X to 115X compared to existing methods while offering orders of magnitude improvements in runtime over full physical design implementation flows. To illustrate PROCEED's capabilities, we deploy it to assess novel tunneling transistors (TFETs) compared to conventional CMOS. Shaodi Wang, Andrew Pan, Chi On Chui, Puneet Gupta 0001 |
ASP-DAC | 4 |
| 2014 | Multi-Layer Memory ResiliencyabstractWith memories continuing to dominate the area, power, cost and performance of a design, there is a critical need to provision reliable, high-performance memory bandwidth for emerging applications. Memories are susceptible to degradation and failures from a wide range of manufacturing, operational and environmental effects, requiring a multi-layer hardware/software approach that can tolerate, adapt and even opportunistically exploit such effects. The overall memory hierarchy is also highly vulnerable to the adverse effects of variability and operational stress. After reviewing the major memory degradation and failure modes, this paper describes the challenges for dependability across the memory hierarchy, and outlines research efforts to achieve multi-layer memory resilience using a hardware/software approach. Two specific exemplars are used to illustrate multilayer memory resilience: first we describe static and dynamic policies to achieve energy savings in caches using aggressive voltage scaling combined with disabling faulty blocks; and second we show how software characteristics can be exposed to the architecture in order to mitigate the aging of large register files in GPGPUs. These approaches can further benefit from semantic retention of application intent to enhance memory dependability across multiple abstraction levels, including applications, compilers, run-time systems, and hardware platforms. Nikil Dutt, Puneet Gupta 0001, Alexandru Nicolau, Abbas BanaiyanMofrad, Mark Gottscho, Majid Namaki-Shoushtari |
DAC | 2 |
| 2014 | Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant CachesabstractComplicated approaches to fault-tolerant voltage-scalable (FTVS) SRAM cache architectures can suffer from high overheads. We propose static (SPCS) and dynamic (DPCS) variants of power/capacity scaling, a simple and low-overhead fault-tolerant cache architecture that utilizes insights gained from our 45nm SOI test chip. Our mechanism combines multi-level voltage scaling with power gating of blocks that become faulty at each voltage level. The SPCS policy sets the runtime cache VDD statically such that almost all of the cache blocks are not faulty. The DPCS policy opportunistically reduces the voltage further to save more power than SPCS while limiting the impact on performance caused by additional faulty blocks. Through an analytical evaluation, we show that our approach can achieve lower static power for all effective cache capacities than a recent complex FTVS work. This is due to significantly lower overheads, despite the failure of our approach to match the min-VDD of the competing work at fixed yield. Through architectural simulations, we find that the average energy saved by SPCS is 55%, while DPCS saves an average of 69% of energy with respect to baseline caches at 1 V. Our approach incurs no more than 4% performance and 5% area penalties in the worst case cache configuration. Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001 |
DAC | 5 |
| 2014 | Benchmarking of mask fracturing heuristicsabstractAggressive resolution enhancement techniques such as inverse lithography (ILT) often lead to complex, non-rectilinear mask shapes which make mask writing extremely slow and expensive. To reduce shot count of complex mask shapes, mask writers allow overlapping shots, due to which the problem of fracturing mask shapes with minimum shot count is NP-hard. The need to correct for e-beam proximity effect makes mask fracturing even more challenging. Although a number of fracturing heuristics have been proposed, there has been no systematic study to analyze the quality of their solutions. In this work, we propose a new method to generate benchmarks with known optimal solutions that can be used to evaluate the suboptimality of mask fracturing heuristics. We also propose a method to generate tight upper and lower bounds for actual ILT mask shapes by formulating mask fracturing as an integer linear program and solving it using branch and price. Our results show that a state-of-the-art prototype [version of] capability within a commercial EDA tool for e-beam mask shot decomposition can be suboptimal by as much as 3.7× for generated benchmarks, and by as much as 3.6× for actual ILT shapes. Tuck-Boon Chan, Puneet Gupta 0001, Kwangsoo Han, Abde Ali Kagalwalla, Andrew B. Kahng, Emile Sahouria |
ICCAD | 2 |
| 2014 | Efficient layout generation and evaluation of vertical channel devicesabstractVertical gate-all-around (VGAA) has been shown to be one of the most promising devices for the scaling beyond 10nm for its reduced delay, large driving current, and good gate control. Moreover, emerging devices such as heterojunction tunneling FETs are more amenable to vertical fabrication. However, past studies of vertical channel devices focused more on regular memory architectures and simple standard cells like inverter. Since naive migration of regular FinFET layouts to vertical FETs yields little benefits, we identify several vertical efficient layout structures and propose novel layout generation heuristics for vertical channel devices. We also compare VGAA with symmetric and asymmetric source/drain architectures. The layout efficiencies of several VGAA structures, vertical double gate (VDG), lateral gate-all-around (LGAA), and FinFET are presented in our experiments. We observe that even though most vertical channel standard cells have more diffusion gaps than lateral cells do, they still benefit from vertical architectures in area because of the elimination of diffusion contacts. For asymmetric architectures, the area is larger than symmetric architectures because of the extra diffusion gaps needed, but our experiments indicate that for both symmetric and asymmetric architectures, vertical channel devices are likely to have a density advantage over lateral channel devices assuming that current drive strengths of both are similar. Wei-Che Wang, Puneet Gupta 0001 |
ICCAD | 2 |
| 2014 | Pattern-restricted design at 10nm and beyondabstractManufacturing has been incapable of keeping up with Moore's law without significantly increasing process variability and imposing massive geometric restrictions on design. This paper highlights the design impact of variability and geometric constraints - including traditional design rules and pattern-scale constraints - and describes our approach for evaluating and enforcing pattern-scale restrictions on design. Rani S. Ghaida, Yasmine Badr, Puneet Gupta 0001 |
ICCD | 3 |
| 2014 | Statistical timing and power analysis of VLSI considering non-linear dependence
Lerong Cheng, Wenyao Xu, Fengbo Ren, Fang Gong, Puneet Gupta 0001, Lei He 0001 |
Integr. | 5 |
| 2014 | SlackProbe: A Flexible and Efficient In Situ Timing Slack Monitoring MethodologyabstractIn situ monitoring is an accurate way to monitor circuit delay or timing slack, but usually incurs significant overhead. We observe that most existing slack monitoring methods focus exclusively on monitoring path endpoints, which is not cost efficient from power and area perspectives. In this paper, we first propose SlackProbe methodology, which inserts timing slack monitors like probes at a selected set of nets, including intermediate nets along critical paths. SlackProbe can be used to detect impending delay failures due to various reasons (process variations, ambient fluctuations, circuit aging, etc.) and can be used with various preventive actions (e.g., voltage/frequency scaling, clock stretching/time borrowing, etc.). Then we perform thorough analysis of the potential benefits and caveats of SlackProbe over conventional approaches in terms of number of monitors required, monitoring efficiency and observability, delay margin, and design perturbation. Experimental results on commercial processors show that with 5% extra timing margin, SlackProbe can reduce the number of monitors by 12-16X as compared to the number of monitors inserted at path ending pins. SlackProbe can also improve the monitoring efficiency by up to 1.9X and improve the monitoring observability by up to 32%, as compared to endpoint monitoring. Liangzhen Lai, Vikas Chandra, Robert C. Aitken, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Synthesis and Analysis of Design-Dependent Ring Oscillator (DDRO) Performance MonitorsabstractWith CMOS technology scaling, circuit performance has become more sensitive to manufacturing and environmental variations. Hence, there is a need to measure or monitor circuit performance during manufacturing and at runtime. Since each circuit may have different sensitivities to process variations, previous works have focused on the synthesis of circuit performance monitors that are specific to a given design. We develop a systematic approach for the synthesis of multiple design-dependent monitors, as well as the corresponding calibration and delay estimation methods. Our approach synthesizes design-dependent ring oscillators (DDROs) using standard-cell library gates and conventional physical implementation flows. Our delay estimation method limits the memory usage overhead by clustering critical paths with similar delay sensitivities. Experimental results show that our delay estimation method using multiple DDROs reduces overestimation (timing margin) by up to 25% compared to using a single monitor. Furthermore, our silicon measurement results for monitoring an industrial microprocessor implemented in a 45-nm silicon-on-insulator process show that DDRO can reduce the mean delay estimation error by 35% compared to inverter-based ring oscillators. Tuck-Boon Chan, Puneet Gupta 0001, Andrew B. Kahng, Liangzhen Lai |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Variability-aware memory management for nanoscale computingabstractAs the semiconductor industry continues to push the limits of sub-micron technology, the ITRS expects hardware (e.g., die-to-die, wafer-to-wafer, and chip-to-chip) variations to continue increasing over the next few decades. As a result, it is imperative for designers to build variation-aware software stacks that may adapt and opportunistically exploit said variations to increase system performance/responsiveness as well as minimize power consumption. The memory subsystem is one of the largest components in today's computing system, a main contributor to the overall power consumption of the system, and therefore one of the most vulnerable components to the effects of variations (e.g., power). This paper discusses the concept of variability-aware memory management for nanoscale computing systems. We show how to opportunistically exploit the hardware variations in on-chip and off-chip memory at the system level through the deployment of variation-aware software stacks. Nikil Dutt, Puneet Gupta 0001, Alexandru Nicolau, Luis Angel D. Bathen, Mark Gottscho |
ASP-DAC | 2 |
| 2013 | Reliable on-chip systems in the nano-era: lessons learnt and future trendsabstractReliability concerns due to technology scaling have been a major focus of researchers and designers for several technology nodes. Therefore, many new techniques for enhancing and optimizing reliability have emerged particularly within the last five to ten years. This perspective paper introduces the most prominent reliability concerns from today's points of view and roughly recapitulates the progress in the community so far. The focus of this paper is on perspective trends from the industrial as well as academic points of view that suggest a way for coping with reliability challenges in upcoming technology nodes. Jörg Henkel, Lars Bauer, Nikil Dutt, Puneet Gupta 0001, Sani R. Nassif, Muhammad Shafique 0001, Mehdi Baradaran Tahoori, Norbert Wehn |
DAC | 4 |
| 2013 | Role of design in multiple patterning: technology development, design enablement and process controlabstractMultiple-patterning optical lithography is inevitable for technology scaling beyond the 22nm technology node. Multiple patterning imposes several counter-intuitive restrictions on layout and carries serious challenges for design methodology. This paper examines the role of design at different stages of the development and adoption of multiple patterning: technology development, design enablement, and process control. We discuss how explicit design involvement can enable timely adoption of multi-patterning with reduced costs both in design and manufacturing. Rani S. Ghaida, Puneet Gupta 0001 |
DATE | 2 |
| 2013 | SlackProbe: a low overhead in situ on-line timing slack monitoring methodologyabstractIn situ monitoring is an accurate way to monitor circuit delay or timing slack, but usually incurs significant overhead. We observe that most existing slack monitoring methods exclusively focus on monitoring path ending registers, which is not cost efficient from power and area perspectives. Liangzhen Lai, Vikas Chandra, Robert C. Aitken, Puneet Gupta 0001 |
DATE | 4 |
| 2013 | Towards analyzing and improving robustness of software applications to intermittent and permanent faults in hardwareabstractAlthough a significant fraction of emerging failure and wearout mechanisms result in intermittent or permanent faults in hardware, their impact (as distinct from transient faults) on software applications has not been well studied. In this paper, we develop a distinguishing application characteristic, referred to as similarity from fundamental circuit-level understanding of the failure mechanisms. We present a mathematical definition and a procedure for similarity computation for practical software applications and experimentally verify the relationship between similarity and fault rate. Leveraging dependence of application robustness on the similarity metric, we present example architecture independent code transformations to reduce similarity and thereby the worst-case fault rate with minimal performance degradation. Our experimental results with arithmetic unit faults show as much as 74% improvement in the worst case fault rate on benchmark kernels, with less than 10% runtime penalty. Joseph Sloan, Lucas Francisco Wanner, Salma Hosni Emam Mohamed Elmalaki, Mani Srivastava 0001, Puneet Gupta 0001 |
ICCD | 6 |
| 2013 | Layout Decomposition and Legalization for Double-Patterning TechnologyabstractThe use of multiple-patterning (MP) optical lithography for sub-20 nm technologies has inevitably become slow to adopt the next generation of lithography systems. The biggest technical challenge of MP is failure to reach a manufacturable layout-coloring solution, especially in dense layouts. This paper offers a post layout solution for the removal of conflicts, i.e., patterns that cannot be assigned to different masks without violating spacing rules. The proposed method essentially consists of three steps: 1) layout coloring; 2) exposure layers; 3) geometric rules definition; and 4) layout legalization using compaction and MP rules as constraints. The method is general and can be used for different MP technologies, including lithography-etch, lithography-etch double-patterning (DP), triple patterning/MP (i.e., multiple litho-etch steps), and self-aligned DP (SADP). For demonstration purposes, we apply the proposed method in this paper to remove conflicts in DP. We offer anO(n) layout-coloring heuristic algorithm for DP, which is up to 80× faster than the integer linear program-based approach. The conflict-removal problem is formulated as a linear program, which permits an extremely fast runtime (less than 1 min in real time for macro layouts). The method was tested on standard cells and macro layouts from a commercial 22-nm library designed without any MP awareness. For many cells, the method removes all conflicts without any area increase. For some complex cells and macros, the method still removes all conflicts but with a modest 6% average increase in area. Rani S. Ghaida, Kanak Agarwal 0001, Sani R. Nassif, Lars Liebmann, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2013 | Underdesigned and Opportunistic Computing in Presence of Hardware VariabilityabstractMicroelectronic circuits exhibit increasing variations in performance, power consumption, and reliability parameters across the manufactured parts and across use of these parts over time in the field. These variations have led to increasing use of overdesign and guardbands in design and test to ensure yield and reliability with respect to a rigid set of datasheet specifications. This paper explores the possibility of constructing computing machines that purposely expose hardware variations to various layers of the system stack including software. This leads to the vision of underdesigned hardware that utilizes a software stack that opportunistically adapts to a sensed or modeled hardware. The envisioned underdesigned and opportunistic computing (UnO) machines face a number of challenges related to the sensing infrastructure and software interfaces that can effectively utilize the sensory data. In this paper, we outline specific sensing mechanisms that we have developed and their potential use in building UnO machines. Puneet Gupta 0001, Yuvraj Agarwal, Lara Dolecek, Nikil Dutt, Rajesh K. Gupta 0001, Rakesh Kumar 0002, Subhasish Mitra, Alexandru Nicolau, Tajana Rosing, Mani Srivastava 0001, Steven Swanson, Dennis Sylvester |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Hardware Variability-Aware Duty Cycling for Embedded SensorsabstractInstance and temperature-dependent power variation has a direct impact on quality of sensing for battery-powered long-running sensing applications. We measure and characterize the active and leakage power for an ARM Cortex M3 processor and show that, across a temperature range of 20 -60, there is a 10% variation in active power, and a variation in leakage power. We introduce variability-aware duty cycling methods and a duty cycle (DC) abstraction for TinyOS which allows applications to explicitly specify the lifetime and minimum DC requirements for individual tasks, and dynamically adjusts the DC rates so that the overall quality of service is maximized in the presence of power variability. We show that variability-aware duty cycling yields a improvement in total active time over schedules based on worst case estimations of power, with an average improvement of across a wide variety of deployment scenarios based on the collected temperature traces. Conversely, datasheet power specifications fail to meet required lifetimes by 7%-15%, with an average 37 days short of the required lifetime of 1 year. Finally, we show that a target localization application using variability-aware DC yields a 50% improvement in quality of results over one based on worst case estimations of power consumption. Lucas Francisco Wanner, Charwak Apte, Rahul Balani, Puneet Gupta 0001, Mani Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | VaMV: Variability-aware Memory VirtualizationabstractPower consumption variability of both on-chip SRAMs and off-chip DRAMs is expected to continue to increase over the next decades. We opportunistically exploit this variability through a novel Variability-aware Memory Virtualization (VaMV) layer that allows programmers to partition their application's address space (through annotations) into virtual address regions and create mapping policies for each region. Each policy has different requirements (e.g., power, fault-tolerance) and is exploited by our dynamic memory management module (VaMVisor), which adapts to the underlying hardware, prioritizes the memory resources according to their characteristics (e.g., power consumption), and selectively maps data to the best-fitting memory resource (e.g., high-utilization data to low-power memory space). Our experimental results on embedded benchmarks show that VaMV is capable of reducing dynamic power consumption by 63% on average while reducing total execution time by an average of 34% by exploiting: 1) SRAM voltage scaling, 2) DRAM power variability, and 3) Efficient dynamic policy-driven variability-aware memory allocation. Luis Angel D. Bathen, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001 |
DATE | 4 |
| 2012 | A methodology for the early exploration of design rules for multiple-patterning technologiesabstractDouble/Multiple-patterning (DP/MP) lithography in a multiple litho-etch steps process is a favorable solution for technology scaling to the 20nm node and below. Mask-assignment conflicts represent the biggest challenge for MP and limiting them through design rules is crucial for the adoption of MP technology. In this paper, we offer a methodology for the early evaluation and exploration of layout and MP rules intended for speeding up the rules-development cycle. Using a novel wiring-estimation method, we create layout estimates with fine-grained congestion prediction. MP-conflicts are then predicted using a machine-learning approach. In this work, we demonstrate the use of the method for double-patterning lithography in litho-etch-litho-etch process; the methodology is more general, however, and can be applied for other multiple-patterning technologies including tripe/multiple-patterning with multiple litho-etch steps, self-aligned double patterning (SADP), and directed self-assembly. Results of testing the methodology on standard-cell layouts show an 81% accuracy in DP-conflicts prediction. The methodology was then used to explore DP and layout rules and investigate their effects on DP-compatibility and layout area. The methodology allows for rules optimization; for example, pushing the minimum tip-to-side same-color spacing rule value from 1.7x to 1.5x the minimum side-to-side spacing design rule (i.e., from 110nm down to 90nm) would more than double the number of DP-compatible cells in the library. Rani S. Ghaida, Tanaya Sahu, Parag Kulkarni, Puneet Gupta 0001 |
ICCAD | 4 |
| 2012 | Impact of range and precision in technology on cell-based designabstractWith the introduction of non-planar CMOS technologies in commercial designs, the effects of the range and precision allowed in a technology is an important. The limited range and precision (i.e. granularity) in a technology, and consequently, in a standard cell design, may result in significant penalties in the power and delay performance in a design. In this work, the impact of the range and precision is examined by providing a new framework for estimating the power suboptimality incurred by a design relative to a given library. Methods that predict the suboptimality well, both qualitatively and quantitatively, and the implications on standard cell library design are explored. While no other methods for estimating suboptimality are known, compared to a method derived from literature, our method provides a nearly 2x better estimate for vth assignment and 10x improvement for gate sizing. John Lee 0002, Puneet Gupta 0001 |
ICCAD | 2 |
| 2012 | DRE: A Framework for Early Co-Evaluation of Design Rules, Technology Choices, and Layout MethodologiesabstractDesign rules have been the primary contract between technology developers and designers and are likely to remain so to preserve abstractions and productivity. While current approaches for defining design rules are largely unsystematic and empirical in nature, this paper offers a novel framework for early and systematic evaluation of design rules and layout styles in terms of major layout characteristics of area, manufacturability, and variability. The framework essentially creates a virtual standard-cell library and performs the evaluation based on the virtual layouts. Due to the focus on the exploration of rules at an early stage of technology development, we use first-order models of variability and manufacturability (instead of relying on accurate simulation) and layout topology/congestion-based area estimates (instead of explicit and slow layout generation). Such a framework can be used to co-evaluate and cooptimize design rules, patterning technologies, layout methodologies, and library architectures. Rani S. Ghaida, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Design-Aware Mask InspectionabstractMask inspection has become a major bottleneck in the manufacturing flow taking up as much as 40% of the total mask manufacturing time. In this paper, we explore techniques to improve the reticle inspection flow by increasing its design awareness. We develop an algorithm to locate nonfunctional features in a postoptical proximity correction layout without using any design information. Using this, and the timing information of the design (if available), the smallest defect size that could cause the design to fail is assigned to each reticle feature. The criticality of various reticle features is then used to partition the reticle such that each partition is inspected at a different pixel size and sensitivity so that the false and nuisance defect count is reduced without missing any critical defect. We also develop an analytical model to estimate the false and nuisance defect count. Using those models, our simulation results show that this design-aware mask inspection can reduce the false and nuisance defect count for a critical polysilicon layer from 80 defects down to 49 defects, leading to substantial reduction in defect review load. We also develop a model to estimate first pass yield (FPY) and show that our method can improve the FPY for a polysilicon layer from 11% to 30%. Apart from the polysilicon layer, the potential benefit of this approach is analyzed for active, contact and all the metal/via layers. Abde Ali Kagalwalla, Puneet Gupta 0001, Christopher J. Progler, Steve McDonald |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | ECO cost measurement and incremental gate sizing for late process changesabstractChanges in the manufacturing process parameters may create timing violations in a design, making it necessary to perform an engineering change order (ECO) to correct these problems. We present a framework for performing incremental gate sizing for process changes late in the design cycle, and a method for creating initial designs that are robust to late process changes. This includes a method for measuring and estimating ECO cost and for transforming these costs into linear programming optimization problems. In the case of ECOs, the method reduces ECO costs on average, by 89% in changed area compared to a leading commercial tool. Furthermore, the robust initial designs are, on average, 55% less likely to need redesign in the future. John Lee 0002, Puneet Gupta 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2012 | Discrete sizing for leakage power optimization in physical design: A comparative studyabstractWhile sizing has been studied for over three decades, the absence of a common framework with which to compare methods has made progress difficult to measure. In this article, we compare popular sizing techniques in which gates are chosen from a discrete standard cell library and slew and interconnect effects are accounted for. The difference between sizing methods reduces from roughly 53% to 8% between best and worst case after slew propagation is taken into account. In our benchmarks, no one sizing technique consistently outperforms the others. Santiago Mok, John Lee 0002, Puneet Gupta 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2012 | AppAdapt: Opportunistic Application Adaptation in Presence of Hardware VariationabstractIn this work, we propose a method to reduce the impact of process variations by adapting the application's algorithm at the software layer. We introduce the concept of hardware signatures as the measured post manufacturing hardware characteristics that can be used to drive software adaptation across different die. Using H.264 encoding as an example, we demonstrate significant yield improvements (as much as 30% points at 0% hardware overdesign), a reduction in overdesign (by as much as 8% points at 80% yield) as well as application quality improvements (about 2.0 dB increase in average peak-signal-to-noise ratio at 70% yield). Further, we investigate implications of limited information exchange (i.e., signature quantization) on yield and quality. We conclude that hardware-signature-based application adaptation is an easy and inexpensive (to implement), better informed (by actual application requirements) and effective way to manage yield-cost-quality tradeoffs in application-implementation design flows. Aashish Pant, Puneet Gupta 0001, Mihaela van der Schaar |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Underdesigned and Opportunistic ComputingabstractVariation in the specifications of microelectronic chips across parts and over time has been a great source of concern for the integrated circuit chip designers because of the ever-increasing guard-bands that the circuit and system designers must rely upon to ensure working parts and systems. This prompts us to look for solutions that can mitigate the effect of performance and power variability through innovations in sys-tem software. In this paper, we outline a novel, flexible hardware-software stack and interface that use adaptation in software to relax variation-induced guard-bands in hardware design. Puneet Gupta 0001, Rajesh K. Gupta 0001 |
Asian Test Symposium | 1 |
| 2011 | On the efficacy of NBTI mitigation techniquesabstractNegative Bias Temperature Instability (NBTI) has become an important reliability issue in modern semiconductor processes. Recent work has attempted to address NBTI-induced degradation at the architecture level. However, such work has relied on device-level analytical models that, we argue, are limited in their flexibility to model the impact of architecture-level techniques on NBTI degradation. In this paper, we propose a flexible numerical model for NBTI degradation that can be adapted to better estimate the impact of architecture-level techniques on NBTI degradation. Our model is a numerical solution to the reaction-diffusion equations describing NBTI degradation that has been parameterized to model the impact of dynamic voltage scaling, averaging effects across logic paths, power gating, and activity management We use this model to understand the effectiveness of different classes of architecture-level techniques that have been proposed to mitigate the effects of NBTI. We show that the potential benefits from these techniques are, for the most part, smaller than what has been previously suggested, and that guardbanding may still be an efficient way to deal with aging. Tuck-Boon Chan, John Sartori, Puneet Gupta 0001, Rakesh Kumar 0002 |
DATE | 3 |
| 2011 | Variability-aware duty cycle scheduling in long running embedded sensing systemsabstractInstance and temperature-dependent leakage power variability is already a significant issue in contemporary embedded processors, and one which is expected to increase in importance with scaling of semiconductor technology. We measure and characterize this leakage power variability in current microprocessors, and show that variability aware duty cycle scheduling produces 7.1× improvement in sensing quality for a desired lifetime. In contrast, pessimistic estimations of power consumption leave 61% of the energy untapped, and datasheet power specifications fail to meet required lifetimes by 14%. Finally, we introduce a duty cycle abstraction for TinyOS that allows applications to explicitly specify lifetime and minimum duty cycle requirements for individual tasks, and dynamically adjusts duty cycle rates so that overall quality of service is maximized in the presence of power variability. Lucas Francisco Wanner, Rahul Balani, Sadaf Zahedi, Charwak Apte, Puneet Gupta 0001, Mani Srivastava 0001 |
DATE | 5 |
| 2011 | A framework for double patterning-enabled designabstractWhile the next generation of lithography systems is still under development, extending optical lithography using double patterning (DP) is the only solution to continue technology scaling. The biggest technical challenge of DP is the presence of mask-assignment conflicts in dense layers. In this paper, we propose a framework for DP conflict removal for standard cells. First, we offer an O(n) algorithm for mask assignment (up to 200× faster than the ILP-based approach) that guarantees a conflict-free solution if one exists. We then formulate the problem of conflict removal as a linear program (LP), which permits an extremely fast run-time (less than 10 seconds in real time for typical cells). The framework removes DP conflicts and legalizes the layout across all layers simultaneously while minimizing layout perturbation. For cells from a commercial 22nm library designed without any DP awareness, our method usually removes all DP conflicts without any area increase; for some complex cells, the method still removes all conflicts with a modest 6.7% average increase in area. The method is more general, however, and can also be applied for macro layouts and the interconnect layers in complete designs as we demonstrate in the paper. Rani S. Ghaida, Kanak Agarwal 0001, Sani R. Nassif, Lars Liebmann, Puneet Gupta 0001 |
ICCAD | 6 |
| 2011 | Physically Justifiable Die-Level Modeling of Spatial Variation in View of Systematic Across Wafer VariabilityabstractModeling spatial variation is important for statistical analysis. Most existing works model spatial variation as spatially correlated random variables. We discuss process origins of spatial variability, all of which indicate that spatial variation comes from deterministic across-wafer variation, and purely random spatial variation is not significant. We analytically study the impact of across-wafer variation and show how it gives an appearance of correlation. We have developed a new die-level variation model considering deterministic across-wafer variation and derived the range of conditions under which ignoring spatial variation altogether may be acceptable. Experimental results show that for statistical timing and leakage analysis, our model is within 2% and 5% error from exact simulation result, respectively, while the error of the existing distance-based spatial variation model is up to 6.5% and 17%, respectively. Moreover, our new model is also faster than the spatial variation model for statistical timing analysis and faster for statistical leakage analysis. Lerong Cheng, Puneet Gupta 0001, Costas J. Spanos, Kun Qian 0014, Lei He 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | On confidence in characterization and application of variation modelsabstractIn this paper we study statistics of statistics. Statistical modeling and analysis have become the mainstay of modern design-manufacturing flows. Most analysis techniques assume that the statistical variation models are reliable. However, due to limited number of samples (especially in the case of lot-to-lot variation), calibrated models have low degree of confidence. The problem is further exacerbated when production volumes are low (¿ 65 lots) causing additional loss of confidence in the statistical analysis (since production only sees a small snapshot of the entire distribution). The problem of confidence in statistical analysis is going to be further worsened with advent of 450mm wafers. We mathematically derive the confidence intervals for commonly used statistical measures (mean, variance, percentile corner) and analysis (SPICE corner extraction, statistical timing). Our estimates are within 2% of simulated confidence values. Our experiments (with variability assumptions derived from test silicon data from a 45nm industrial process) indicate that for moderate characterization volumes (10 lots) and low-to-medium production volumes (15 lots), a significant guardband (e.g., 34.7% of standard deviation for single parameter corner, 38.7% of standard deviation for SPICE corner, and 52% of standard deviation for 95%-tile point of circuit delay) is needed to ensure 95% confidence in the results. The guardbands are non-negligible for all cases when either production or characterization volume is not large. We also study the interesting one production lot case which may be common for prototyping as well as for academic designs. The proposed methods require are not runtime-intensive (always within 10s) as they require Monte-Carlo simulations on closed form expressions. Lerong Cheng, Puneet Gupta 0001, Lei He 0001 |
ASP-DAC | 2 |
| 2010 | Eyecharts: constructive benchmarking of gate sizing heuristicsabstractDiscrete gate sizing is one of the most commonly used, flexible, and powerful techniques for digital circuit optimization. The underlying problem has been proven to be NP-hard [1]. Several (suboptimal) gate sizing heuristics have been proposed over the past two decades, but research has suffered from the lack of any systematic way of assessing the quality of the proposed algorithms. We develop a method to generate benchmark circuits (called eyecharts) of arbitrary size along with a method to compute their optimal solutions using dynamic programming. We evaluate the suboptimalities of some popular gate sizing algorithms. Eyecharts help diagnose the weaknesses of existing gate sizing algorithms, enable systematic and quantitative comparison of sizing algorithms, and catalyze further gate sizing research. Our results show that common sizing methods (including commercial tools) can be suboptimal by as much as 54% (V t -assignment), 46% (gate sizing) and 49% (gate-length biasing) for realistic libraries and circuit topologies. Puneet Gupta 0001, Andrew B. Kahng, Amarnath Kasibhatla |
DAC | 1 |
| 2010 | Software adaptation in quality sensitive applications to deal with hardware variabilityabstractIn this work, we propose a method to reduce the impact of process variations by adapting the application's algorithm at the software layer. We introduce the concept of hardware signatures as the measured post manufacturing hardware characteristics that can be used to drive software adaptation across different die. Using H.264 encoding as an example, we demonstrate significant yield improvements (as much as 40% points at 0% over-design), a reduction in over-design (by as much as 10% points at 80% yield) as well as application quality improvements (about 2.6dB increase in average PSNR at 80% yield). Further, we investigate implications of limited information exchange (i.e. signature measurement granularity) on yield and quality. We show that our proposed technique for determining optimal signature measurement points results in an improvement in PSNR of about 1.3dB over naive sampling for the H.264 encoder. We conclude that hardware-signature based application adaptation is an easy and inexpensive (to implement), better informed (by actual application requirements) and e ffective way to manage yield-cost-quality tradeoffs in application-implementation design flows. Aashish Pant, Puneet Gupta 0001, Mihaela van der Schaar |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Design dependent process monitoring for back-end manufacturing cost reductionabstractShort-loop process monitoring structures (usually simple device I - V, C - V measurements made after M1 fabrication) are commonly put in wafer scribe-lines. These test structures are almost always design independent and measured/monitored by the foundry to keep track of process deviations. We propose a design-dependent process monitoring strategy which can accurately predict design performance based on simple Ieff-based delay and Ioff-based leakage power estimates. We show that our strategy works much better (0.99 correlation vs. 0.87) compared to conventional design-independent monitors. Further, we use the predicted delay and leakage power for early yield estimation for pruning bad wafers to save test and back-end manufacturing costs We show that wafer pruning based on our approach can achieve upto 98% of the maximum achievable benefit/profit. We design the measurement and prediction schemes so as to minimize data as well as computation that needs to be kept track of during wafer fabrication. Such design-dependent process monitoring can help target process control/optimization effort, enable quicker yield ramp besides saving test and manufacturing costs. Tuck-Boon Chan, Aashish Pant, Lerong Cheng, Puneet Gupta 0001 |
ICCAD | 4 |
| 2010 | Design-aware mask inspectionabstractMask inspection has become a major bottleneck in the manufacturing flow taking up as much as 30% of the total manufacturing time. In this work we explore techniques to improve the reticle inspection flow by increasing its design awareness. We develop an algorithm to locate non-functional features in a post-OPC layout with 100% accuracy without using any design information. Using this, and timing information of the design (if available), we assign a minimum size defect to each reticle feature that could cause the design to fail. The criticality of various reticle features is then used to partition the reticle such that each partition is inspected at a different pixel size and sensitivity so that the false+nuisance defect count is reduced without missing any critical defect. Up to 4X improvement in false+nuisance defect count is observed with our technique resulting in up to 55% improvement in first pass yield coming from reduction in nuisance defects and substantial reduction in defect review load. Abde Ali Kagalwalla, Puneet Gupta 0001, Christopher J. Progler, Steve McDonald |
ICCAD | 2 |
| 2010 | Incremental gate sizing for late process changesabstractCircuit design often runs in parallel with the development of the manufacturing process that will be used to fabricate it. However, as the manufacturing process matures, its models may undergo substantial changes as the design nears production. These changes may cause the design itself to fail its specifications, and in these cases it is necessary to perform an Engineering Change Order (ECO) to correct these problems. We present a new framework to perform incremental gate sizing for process changes late in the design cycle. This includes a method to measure and estimate ECO cost, transform these costs into a linear programming optimization problem, and solve the problem to find the ECO. This method performs well, compared to a leading commercial physical design tool, reducing ECO costs by 18% to 99% in changed area, and 1% to 96% in number of pins with unnecessary pin timing changes. John Lee 0002, Puneet Gupta 0001 |
ICCD | 2 |
| 2010 | Evaluating Statistical Power OptimizationabstractIn response to the increasing variations in integrated-circuit manufacturing, the current trend is to create designs that take these variations into account statistically. In this paper, we quantify the difference between the statistical and deterministic optima of leakage power while making no assumptions about the delay model. We develop a framework for deriving a theoretical upper bound on the suboptimality that is incurred by using the deterministic optimum as an approximation for the statistical optimum. We show that for the mean power measure, the deterministic optima is an excellent approximation, and for the mean plus standard deviation measures, the optimality gap increases as the amount of inter-die variation grows, for a suite of benchmark circuits in a 45 nm technology. For large variations, we show that there are excellent linear approximations that can be used to approximate the effects of variation. Therefore, the need to develop special statistical power optimization algorithms is questionable. Jason Cong, Puneet Gupta 0001, John Lee 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Accounting for non-linear dependence using function driven component analysisabstractMajority of practical multivariate statistical analyses and optimizations model interdependence among random variables in terms of the linear correlation among them. Though linear correlation is simple to use and evaluate, in several cases non-linear dependence between random variables may be too strong to ignore. In this paper, We propose polynomial correlation coefficients as simple measure of multivariable non-linear dependence and show that need for modeling non-linear dependence strongly depends on the end function that is to be evaluated from the random variables. Then, we calculate the errors in estimation which result from assuming independence of components generated by linear de-correlation techniques such as PCA and ICA. The experimental result shows that the error predicted by our method is within 1% error compared to the real simulation. In order to deal with non-linear dependence, we further develop a target function driven component analysis algorithm (FCA) to minimize the error caused by ignoring high order dependence and apply such technique to statistical leakage power analysis and SRAM cell noise margin variation analysis. Experimental results show that the proposed FCA method is more accurate compared to the traditional PCA or ICA. Lerong Cheng, Puneet Gupta 0001, Lei He 0001 |
ASP-DAC | 2 |
| 2009 | On the futility of statistical power optimizationabstractIn response to the increasing variations in integrated-circuit manufacturing, the current trend is to create designs that take these variations into account statistically. In this paper we try to quantify the difference between the statistical and deterministic optima of leakage power while making no assumptions about the delay model. We develop a framework for deriving a theoretical upper-bound on the suboptimality that is incurred by using the deterministic optimum as an approximation for the statistical optimum. On average, the bound is 2.4% for a suite of benchmark circuits in a 45 nm technology. We further give an intuitive explanation and show, by using solution rank orders, that the practical suboptimality gap is much lower. Therefore, the need for statistical power modeling for the purpose of optimization is questionable. Jason Cong, Puneet Gupta 0001, John Lee 0002 |
ASP-DAC | 2 |
| 2009 | Physically justifiable die-level modeling of spatial variation in view of systematic across wafer variabilityabstractModeling spatial variation is important for statistical analysis. Most existing works model spatial variation as spatially correlated random variables. We discuss process origins of spatial variability, all of which indicate that spatial variation comes from deterministic across-wafer variation, and purely random spatial variation is not significant. We analytically study the impact of across-wafer variation and show how it gives an appearance of correlation. We have developed a new dielevel variation model considering deterministic across-wafer variation and derived the range of conditions under which ignoring spatial variation altogether may be acceptable. Experimental results show that our model is within 1% error from exact simulation result while the error of the existing distance-based spatial variation model is up to 8%. Moreover, our new model is also 10X faster than the spatial variation model for Monte-Carlo analysis. Lerong Cheng, Puneet Gupta 0001, Costas J. Spanos, Kun Qian 0014, Lei He 0001 |
DAC | 2 |
| 2009 | A framework for early and systematic evaluation of design rulesabstractAbstract—Design rules have been the primary contract be-tween technology and design and are likely to remain so to preserve abstractions and productivity. While current approaches for defining design rules are largely unsystematic and empirical in nature, this paper offers a novel framework for early and systematic evaluation of design rules and layout styles in terms of major layout characteristics of area, manufacturability, and variability. Due to the focus on co-exploration in early stages of technology development, we use first order models of variability and manufacturability (instead of relying on accurate simulation) and layout topology/congestion-based area estimates (instead of explicit and slow layout generation). The framework is used to efficiently co-evaluate several debatable rules (evaluation for a 104-cell library takes 20 minutes). Results show that: a) diffusion-rounding mainly from diffusion power-straps is a dominant source of variability, b) cell-area overhead of fixed gate-pitch implementation compared to 1D-poly implementation is tolerable (5%) given the improvement in variability, and c) 1D-poly restriction, which improves manufacturability and variability, has almost no area overhead compared to 2D-poly. In addition, we explore gate-spacing rules using our evaluation framework. This exploration yields almost identical values as those of a commercial 65nm process, which serves as a validation for our approach. I. Rani S. Ghaida, Puneet Gupta 0001 |
ICCAD | 2 |
| 2009 | Efficient Additive Statistical Leakage EstimationabstractNominal power estimation is quick but gives minimal information. Statistical power analysis can provide information on yield, chip robustness, etc., but current methods are unnecessarily slow and complex. This is primarily because existing leakage-power models, which model leakage power as lognormal distribution and calculate chip leakage power based on Wilkinson's approach, are not directly additive. Hence, for each incremental change of the circuit, the covariances between each pair of circuit elements need to be recalculated, which is inefficient. In this paper, we proposed a simple additive polynomial leakage-variation model. With additivity, we can calculate chip leakage power and leakage power after incremental change very efficiently. Experimental results show that our method is five times faster than the existing Wilkinson's approach while having no accuracy loss in mean estimation and about 1% accuracy loss in standard-deviation and 99%-percentile-point estimations. Lerong Cheng, Puneet Gupta 0001, Lei He 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Investigation of diffusion rounding for post-lithography analysisabstractDue to aggressive scaling of device feature size to improve circuit performance in the sub-wavelength lithography regime, both diffusion and poly gate shapes are no longer rectilinear. Diffusion rounding occurs most notably where the diffusion shapes are not perfectly rectangular, including common L and T-shaped diffusion layouts to connect to power rails. This paper investigates the impact of the non-rectilinear shape of diffusion (i.e., sloped diffusion or diffusion rounding) on circuit performance (delay and leakage). Simple weighting function models for Ionmiddot and Ioffto account for the diffusion rounding effects are proposed, and compared with TCAD simulation. Our experiments show that diffusion rounding has an asymmetric characteristic for Ioff due to the differing significance of source/drain junctions on device threshold voltage. Therefore, we can model Ionmiddot and Ioffas a function of slope angle and direction. The proposed models match well with TCAD simulation results, with less than 2% and 6% error in Ionmiddot and Ioff, respectively. Puneet Gupta 0001, Andrew B. Kahng, Saumil Shah, Dennis Sylvester |
ASP-DAC | 1 |
| 2008 | Bounded-lifetime integrated circuitsabstractIntegrated circuits with bounded lifetimes can have many business advantages. We give some simple examples of methods to enforce tunable expiration dates for chips using nanometer reliability mechanisms. Puneet Gupta 0001, Andrew B. Kahng |
DAC | 1 |
| 2007 | Line-End Shortening is Not Always a FailureabstractLine-end shortening (LES) has always been considered a catastrophic failure in circuits. However, we find that a device with some LES can continue to function correctly. Such devices have large drive current and reduced capacitance at the expense of much higher leakage current. In this paper, we investigate the power and performance characteristics of devices with LES. Our simulations show that LES does not always cause catastrophic failure of device functionality. However, in this regime LES can lead to parametric failures, aspects of which we investigate . Puneet Gupta 0001, Andrew B. Kahng, Saumil Shah, Dennis Sylvester |
DAC | 1 |
| 2007 | Self-Compensating Design for Reduction of Timing and Leakage Sensitivity to Systematic Pattern-Dependent VariationabstractCritical dimension (CD) variation caused by defocus is largely systematic with dense lines ldquosmilingrdquo through focus while isolated lines ldquofrown.rdquo In this paper, we propose a new design methodology that allows explicit compensation of focus-dependent CD variation, in particular, either within a cell (self-compensated cells) or across cells in a critical path (self-compensated design). By creating iso and dense variants for each library cell, we can achieve designs that are more robust to focus variation. Optimization with a mixture of dense and iso cell variants is possible, both for area and leakage power in timing constraints (critical delay), with the latter an interesting complement to existing leakage-reduction techniques, such as dual-Vth. We implement both a heuristic and mixed-integer linear-programming (MILP) solution methods to address this optimization and experimentally compare their results. Results indicate that designing with a self-compensated cell library incurs 12% area penalty and 6% leakage increase over a baseline library while compensating for focus-dependent CD variation (i.e., the design meets timing constraints across a large range of focus variation). We observe 27% area penalty and 7% leakage increase at the worst case defocus condition using only single-pitch cells. The area penalty of circuits after using both the heuristic and MILP optimization approaches is reduced to 3% while maintaining timing. We also apply the optimization to leakage, which traditionally shows very large variability due to its exponential relationship with gate CD. We conclude that a mixed iso/dense library that is combined with a sensitivity-based optimization approach yields much better area/timing/leakage tradeoffs than using a self-compensated cell library alone. Self-compensated designs show 25% less leakage power on average at the worst defocus condition compared to a design employing a conventional library for the benchmarks studied. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Detailed Placement for Enhanced Control of Resist and Etch CDsabstractSubresolution assist feature (SRAF) and etch-dummy-insertion techniques have been absolutely essential for process-window enhancement and CD control in photo and etch processes. However, as focus levels change during lithography manufacturing, CDs at a given ldquolegalrdquo pitch can fail to achieve manufacturing tolerances. Placed standard-cell layouts may not have the ideal whitespace distribution to allow for an optimal assist-feature insertion. This paper first describes a novel dynamic-programming-based technique for Assist-Feature Correctness (AFCorr) in detailed placement of standard-cell designs. At the same time, etch-dummy features are used in the mask data preparation flow to reduce CD skew between resist and etch processes and to improve the printability of layouts. However, etch-dummy rules conflict with the SRAF insertion because each of the two techniques requires specific design rules. We further present a novel SRAF-aware etch-dummy-insertion method (SAEDM) which optimizes the etch-dummy insertion to make the layout more conducive to the assist-feature insertion after the etch-dummy features have been inserted. Since placement of cells can create forbidden-pitch violations of resist process and can increase etch skew, the placer must also generate etch-dummy-correct placement. This can be solved by Etch-dummy Correctness (EtchCorr), which is an intelligent whitespace management for etch-dummy-corrected placement, an extension of the AFCorr methodology. These methods for enhanced resist and etch CD controls are validated on industrial test cases with respect to wafer printability, database complexity, and device performance. For benchmark designs, we validate the four methodologies: 1) AFCorr; 2) SAEDM; 3) AFCorr SAEDM; and 4) AFCorr EtchCorr SAEDM. The AFCorr placement perturbation achieves a significant reduction in forbidden pitches between polysilicon shapes. Using 1) flow, forbidden-pitch count of photo process is reduced by 76%-100% for 130 nm and by 87%-100% for 90 nm. Our novel Corr design-perturbation technique, which combines the AFCorr and EtchCorr methods, facilitates additional SRAF and etch-dummy insertions and, thus, reduces the CD skew between the photo and etch processes. After Corr with SAEDM, edge-placement-error count is also reduced by 91%-100% in the resist CD and by 72%-98% in the etch CD. Our methods provide a substantial improvement in CD control with negligible timing, area, and CPU overhead. The advantages of such correctness methods are expected to increase in future technology nodes. Puneet Gupta 0001, Andrew B. Kahng, Chul-Hong Park |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Standard cell library optimization for leakage reductionabstractScaling device geometries have caused leakage-power consumption to be one of the major challenges of deep sub-micron design and a major source for parametric yield loss. We propose a library optimization approach involving generation of additional variants for each cell master, by biasing gate-lengths of devices. We employ transistor-level gate-length assignment to exploit asymmetries in standard cell circuit topology as well slack distribution of the design. The enhanced library is used by a power optimizer to reduce design leakage without violating any timing constraints. Such transistor-level optimization of cell libraries offers significantly better leakage-delay tradeoff than simple cell-level biasing (CLB) proposed previously. Experimental results on benchmarks show transistor-level biasing (TLB) can improve the CLB leakage optimization results by 8-17%. There is a corresponding improvement in design leakage distribution as well. Saumil Shah, Puneet Gupta 0001, Andrew B. Kahng |
DAC | 2 |
| 2006 | Wafer Topography-Aware Optical Proximity CorrectionabstractDepth of focus is the major contributor to lithographic process margin. One of the major causes of focus variation is imperfect planarization of fabrication layers. Presently, optical proximity correction (OPC) methods are oblivious to the predictable nature of focus variation arising from wafer topography. As a result, designers suffer from manufacturing yield loss as well as loss of design quality through unnecessary guardbanding. In this paper, the authors propose a novel flow and method to drive OPC with a topography map of the layout that is generated by chemical-mechanical polishing simulation. The wafer topography variations result in local defocus, which the authors explicitly model in the OPC insertion and verification flows. In addition, a novel topography-aware optical rule check to validate the quality of result of OPC for a given topography is presented. The experimental validation in this paper uses simulation-based experiments with 90-nm foundry libraries and industry-strength OPC and scattering bar recipes. It is found that the proposed topography-aware OPC (TOPC) can yield up to 67% reduction in edge placement errors. TOPC achieves up to 72% reduction in worst case printability with little increase in data volume and OPC runtime. The electrical impact of the proposed TOPC method is investigated. The results show that TOPC can significantly reduce timing uncertainty in addition to process variation Puneet Gupta 0001, Andrew B. Kahng, Chul-Hong Park, Kambiz Samadi, Xu Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Gate-length biasing for runtime-leakage controlabstractLeakage power has become one of the most critical design concerns for the system level chip designer. While lowered supplies (and consequently, lowered threshold voltage) and aggressive clock gating can achieve dynamic power reduction, these techniques increase the leakage power and, therefore, causes its share of total power to increase. Manufacturers face the additional challenge of leakage variability: Recent data indicate that the leakage of microprocessor chips from a single 180-nm wafer can vary by as much as 20/spl times/. Previously proposed techniques for leakage-power reduction include the use of multiple supply and gate threshold voltages, and the assignment of input values to inactive gates, such that leakage is minimized. The additional design space afforded by the biasing of device gate lengths to reduce chip leakage power and its variability is studied. It is well known that leakage power decreases exponentially and delay increases linearly with increasing gate length. Thus, it is possible to increase gate length only marginally to take advantage of the exponential leakage reduction, while impairing performance only linearly. From a design-flow standpoint, the use of only slight increases in gate length preserves both pin and layout compatibility; therefore, the authors' technique can be applied as a postlayout enhancement step. The authors apply gate-length biasing only to those devices that do not appear in critical paths, thus assuring zero or negligible degradation in chip performance. To highlight the value of the technique, the multithreshold voltage technique, which is widely used for leakage reduction, is first applied and then gate-length biasing is used to show further reduction in leakage. Experimental results show that gate-length biasing reduces leakage by 24%-38% for the most commonly used cells, while incurring delay penalties of under 10%. Selective gate-length biasing at the circuit level reduces circuit leakage by up to 30% with no delay penalty. Leakage variability is reduced significantly by up to 41%, which may lead to substantial improvements in the manufacturing yield and the product cost. The use of gate-length biasing for leakage optimization of cell instances is also assessed, in which: 1) not all timing arcs are timing critical and/or 2) the rise and fall transitions are not both timing critical at the same time. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Detailed placement for improved depth of focus and CD controlabstractSub-resolution assist features (SRAFs) provide an absolutely essential technique for critical dimension (CD) control and process window enhancement in subwavelength lithography. However, as focus levels change during manufacturing, CDs at a given "legal" pitch can fail to achieve manufacturing tolerances required for adequate yield. Furthermore, adoption of off-axis illumination (OAI) and SRAF techniques to enhance resolution at minimum pitch worsens printability of patterns at other pitches. This paper describes a novel dynamic programming-based technique for Assist-Feature Correctness (AFCorr) in detailed placement of standard-cell designs. For benchmark designs in 130nm and 90nm technologies, AFCorr achieves improved depth of focus and substantial improvement in CD control with negligible timing, area, or CPU overhead. The advantages of AFCorr are expected to increase in future technology nodes. Puneet Gupta 0001, Andrew B. Kahng, Chul-Hong Park |
ASP-DAC | 1 |
| 2005 | Advanced Timing Analysis Based on Post-OPC Extraction of Critical DimensionsabstractProcess variations have become a bottleneck for predictable and high-yielding IC design and fabrication. Linewidth variation (∆L) due to defocus in a chip is largely systematic after the layout is completed, i.e., dense lines "smile" through focus while isolated (iso) lines "frown". In this paper, we propose a design flow that allows explicit compensation of focus variation, either within a cell (self-compensated cells) or across cells in a critical path (self-compensated design). Assuming that iso and dense variants are available for each library cell, we achieve designs that are more robust to focus variation. Design with a self-compensated cell library incurs ~11-12% area penalty while compensating for focus variation. Across-cell optimization with a mix of dense and iso cell variants incurs ~6-8% area overhead compared to the original cell library, while meeting timing constraints across a large range of focus variation (from 0 to 0.4um). A combination of original and iso cells provides an even better self-compensating design option, with only 1% area overhead. Circuit delay distributions are tighter with self-compensated cells and self-compensated design than with a conventional design methodology. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
DAC | 1 |
| 2005 | Layout-aware scan chain synthesis for improved path delay fault coverageabstractPath delay fault testing has become increasingly important due to higher clock rates and higher process variability caused by shrinking geometries. Achieving high-coverage path delay fault testing requires the application of scan justified test vector pairs, coupled with careful ordering of the scan flip-flops and/or insertion of dummy flip-flops in the scan chain. Previous works on scan synthesis for path delay fault testing using scan shifting have focused exclusively on maximizing fault coverage and/or minimizing the number of dummy flip-flops, but have disregarded the scan wirelength overhead. In this paper we propose a layout-aware coverage-driven scan chain ordering methodology and give exact and heuristic algorithms for computing the achievable tradeoffs between path delay fault coverage and both dummy flip-flop and wirelength costs. Experimental results show that our scan chain ordering methodology yields significant improvements in path delay coverage with a very small increase in wirelength overhead compared to previous layout-driven approaches, and similar coverage with up to 25 times improvement in wirelength compared to previous layout-oblivious coverage-driven approaches. Puneet Gupta 0001, Andrew B. Kahng, Ion I. Mandoiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Routing-aware scan chain orderingabstractScan chain insertion can have a large impact on routability, wirelength, and timing of the design. We present a routing-driven methodology for scan chain ordering with minimum wirelength objective. A routing-based approach to scan chain ordering, while potentially more accurate, can result in TSP (Traveling Salesman Problem) instances which are asymmetric and highly nonmetric; this may require a careful choice of solvers. We evaluate our new methodology on recent industry place-and-route blocks with 1200 to 5000 scan cells. We show substantial wirelength reductions for the routing-based flow versus the traditional placement-based flow. In a number of our test cases, over 86% of scan routing overhead is saved. Even though our experiments are, so far, timing oblivious, the routing-based flow also improves evaluated timing, and practical timing-driven extensions appear feasible. Puneet Gupta 0001, Andrew B. Kahng, Stefanus Mantik |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2004 | Toward a methodology for manufacturability-driven design rule explorationabstractResolution enhancement techniques (RET) such as optical proximity correction (OPC) and phase-shift mask (PSM) technology are deployed in modern processes to increase the fidelity of printed features, especially critical dimensions (CD) in polysilicon. Even given these exotic technologies, there has been momentum towards less exibility in layout, in order to ensure printability. However, there has not been a systematic study of the performance and manufacturability impact of such a move towards restrictive design rules. In this paper we present a design ow that evaluates the application of various restricted design rule (RDR) sets in deep submicron ASIC designs in terms of circuit performance and parametric yield. Using such a framework, process and design engineers can identify potential solutions to maximize manufacturability by selectively applying RDRs while maintaining chip performance. In this work we focus attention on the device layer which is the most difficult design layer to manufacture. We quantify the performance, manufacturability and mask cost impact of several common design rules. For instance, we find that small increases in the minimum allowable poly line end extension beyond active provide high levels of immunity to lithographic defocus conditions. Also, modification of the minimum field poly to diffusion spacing can provide good manufacturability, while a single pitch single orientation design rule can reduce gate 3σ uncertainty. Both of these improve in data volume as well, with little to no performance penalties. Reductions in data volume and worst-case edge placement error are on the order of 20-30% and 30-50% respectively compared to a standard baseline design rule set. Luigi Capodieci, Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester, Jie Yang 0010 |
DAC | 2 |
| 2004 | Toward a systematic-variation aware timing methodologyabstractVariability of circuit performance is becoming a very important issue for ultra-deep sub-micron technology. Gate length variation has the most direct impact on circuit performance. Since many factors contribute to the variability of gate length, recent studies have modeled the variability using Gaussian distributions. In reality, the through-pitch and through-focus variations of gate length are systematic. In this paper, we propose a timing methodology which takes these systematic variations into account and we show that it can reduce the timing uncertainty by up to 40%. Puneet Gupta 0001, Fook-Luen Heng |
DAC | 1 |
| 2004 | Selective gate-length biasing for cost-effective runtime leakage controlabstractWith process scaling, leakage power reduction has become one of the most important design concerns. Multi-threshold techniques have been used to reduce runtime leakage power without sacrificing performance. In this paper, we propose small biases of transistor gate-length to further minimize power in a manufacturable manner. Unlike multi-V th techniques, gate-length biasing requires no additional masks and may be performed at any stage in the design process.Our results show that gate-length biasing effectively reduces leakage power by up to 25% with less than 4% delay penalty. We show the feasibility of the technique in terms of manufacturability and pin-compatibility for post-layout power optimization. We also show up to 54% reduction in leakage uncertainty due to inter-die process variation in circuits when biased gate-lengths, versus only unbiased one, are used. Circuits selectively biased show much less sensitivity to both intra and inter die variations. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
DAC | 1 |
| 2003 | Routing-aware scan chain orderingabstractScan chain insertion can have a large impact on routability, wirelength and timing of the design. We present a routing-driven methodology for scan chain ordering with minimum wirelength objective. A routing-based approach to scan chain ordering, while potentially more accurate, can result in TSP (Traveling Salesman Problem) instances which are asymmetric and highly non-metric; this may require a careful choice of solvers. We evaluate our new methodology on recent industry place-and-route blocks with 1200 to 5000 scan cells. We show substantial wirelength reductions for the routing-based flow, versus the traditional placement-based flow: in a number of our test cases, over 86% of scan routing overhead is saved. Even though our experiments are so far timing-oblivious, the routing-based flow does also improve evaluated timing, and practical timing-driven extensions appear feasible. Puneet Gupta 0001, Andrew B. Kahng, Stefanus Mantik |
ASP-DAC | 1 |
| 2003 | Performance-impact limited area fill synthesisabstractChemical-mechanical planarization (CMP) and other manufacturing steps in very deep-submicron VLSI have varying effects on device and interconnect features, depending on the local layout density. To improve manufacturability and performance predictability, area fill features are inserted into the layout to improve uniformity with respect to density criteria. However, the performance impact of area fill insertion is not considered by any fill method in the literature. In this paper, we first review and develop estimates for capacitance and timing overhead of area fill insertions. We then give the first formulations of the Performance Impact Limited Fill (PIL-Fill) problem with the objective of either minimizing total delay impact (MDFC) or maximizing the minimum slack of all nets (MSFC), subject to inserting a given prescribed amount of fill. For the MDFC PIL-Fill problem, we describe three practical solution approaches based on Integer Linear Programming (ILP-I and ILP-II) and the Greedy method. For the MSFC PIL-Fill problem, we describe an iterated greedy method that integrates call to an industry static timing analysis tool. We test our methods on layout testcases obtained from industry. Compared with the normal fill method [3], our ILP-II method for MDFC PIL-Fill problem achieves between 25-% and 90% reduction in terms of total weighted edge delay (roughly, a measure of sum of node slacks) impact while maintaining identical quality of the layout density control; and our iterated greedy method for MSFC PIL-Fill problem also shows significant advantage with respect to the minimum slack of nets on post-fill layout. Yu Chen 0005, Puneet Gupta 0001, Andrew B. Kahng |
DAC | 2 |
| 2003 | A cost-driven lithographic correction methodology based on off-the-shelf sizing toolsabstractAs minimum feature sizes continue to shrink, patterned features have become significantly smaller than the wavelength of light used in optical lithography. As a result, the requirement for dimensional variation control, especially in critical dimension (CD) 3σ, has become more stringent. To meet these requirements, resolution enhancement techniques (RET) such as optical proximity correction (OPC) and phase shift mask (PSM) technology are applied. These approaches result in a substantial increase in mask costs and make the cost of ownership (COO) a key parameter in the comparison of lithography technologies. No concept of function is injected into the mask flow; that is, current OPC techniques are oblivious to the design intent, and the entire layout is corrected uniformly with the same effort. We propose a novel minimum cost of correction (MinCorr) methodology to determine the level of correction for each layout feature such that prescribed parametric yield is attained with minimum total RET cost. We highlight potential solutions to the MinCorr problem and give a simple mapping to traditional performance optimization. We conclude with experimental results showing that substantial RET costs may be saved while maintaining a given desired level of parametric yield. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester, Jie Yang 0010 |
DAC | 1 |
| 2003 | Manufacturing-Aware Physical Design
Puneet Gupta 0001, Andrew B. Kahng |
ICCAD | 1 |
| 2003 | Layout-Aware Scan Chain Synthesis for Improved Path Delay Fault Coverage
Puneet Gupta 0001, Andrew B. Kahng, Ion I. Mandoiu |
ICCAD | 1 |