Sandeep Gupta 0001

dblp:g/SandeepKGupta · also Sandeep K. Gupta 0001 · DBLP profile ↗
← Back
167ranked-venue papers
7as first author
11since 2021 · last 2024
0000-0002-2585-9378ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 158 · 7 first-author · 11 since 2021Software engineering, systems software and programming languages · 14 · 2 since 2021Computer networks · 8
YearPublicationVenuePosition
2024 Challenges and Unexplored Frontiers in Electronic Design Automation for Superconducting Digital Logic
abstract
Positioned as a highly promising post-CMOS computing technology, superconductor electronics (SCE) offer the potential for unparalleled performance and energy efficiency gains compared to end-of-roadmap CMOS circuits. However, achieving very large-scale integration poses numerous challenges. These challenges span from the modeling and analysis of superconducting devices and logic gates to the intricate design of complex SCE circuits and systems. Addressing power and clock distribution issues, minimizing adverse effects of flux trappings, and mitigating stray electromagnetic fields in sensitive SCE circuitry are key challenges that need attention. Verification and testing of SCE circuits also remain open problems. Moreover, scaling the minimum feature sizes of SCE circuits, currently set at 150nm, presents critical scaling and physical design challenges that must be overcome. This review aims to delve into these issues, providing detailed insights while exploring existing or potential solutions to overcome them.
Sasan Razmkhah, Robert Aviles, Mingye Li, Sandeep Gupta 0001, Peter A. Beerel, Massoud Pedram
DATE4
2024 A Novel Multi-Objective Optimization Framework for Analog Circuit Customization
abstract
Prior research has developed an approach called Analog Mixed-signal Parameter Search Engine (AMPSE) [1] to reduce the cost of design of analog/mixed-signal (AMS) circuits. In this paper, we propose an adaptive sampling method (AS) to identify a range of Pareto-optimal versions of a given AMS circuit with different combinations of metric values to enable parameter-search based methods like AMPSE to efficiently serve multiple users with diverse requirements. As AMS circuit simulation has high run-time complexity, our method uses a surrogate model to estimate the values of metrics for the circuit, given the values of its parameters. In each iteration, we use a mix of uniform and adaptive sampling to identify parameter value combinations, use the surrogate model to identify a subset of these samples to simulate, and use the simulation results to retrain the model. Our method is more effective and has lower complexity compared with prior methods [2]–[4] because it works with any surrogate model, uses a low-complexity yet effective strategy to identify samples for simulation, and uses an adaptive annealing strategy to balance exploration vs. exploitation. Experimental results demonstrate that, at lower complexity, our method discovers better Pareto-optimal designs compared to prior methods. The benefits of our method, relative to prior methods, increase as we move from AMS circuits with low simulation complexities to those with higher simulation complexities. For an AMS circuit with very high simulation complexity, our method identifies designs that are superior to the version of the circuit optimized by experienced designers.
Mutian Zhu, Mohsen Hassanpourghadi, Mike Shuo-Wei Chen, Anthony Levi, Sandeep Gupta 0001
DATE6
2024 Predictive Testing for Aging in SRAMs and Mitigation
abstract
We develop a method to estimate lifetime performance, yield, and power for Static Random-Access Memories (SRAMs) that captures the combination of process variations and aging. Using this method, we design and validate predictive tests to detect future aging failures. We use the results of predictive tests to reconfigure dynamic voltage and frequency scaling (DVFS) to reduce aging failures at minimal energy and latency overheads.
Yunkun Lin, Mingye Li, Sandeep Gupta 0001
ITC3
2024 Built in self test (BIST) for RSFQ circuits
abstract
In the era beyond the end of physical scaling of CMOS, growing attention is being paid to Superconducting electronics (SCE), especially Rapid Single Flux Quantum (RSFQ) logic due to its high-performance and low power consumption. In [1]–[3], static and delay fault models, corresponding automatic test pattern generator (ATPG) for testing these faults, and a scan architecture are developed. However, test pattern application involves moving patterns and responses via long wires from the test equipment at room temperature to the chip under test in liquid helium, which severely reduces the test clock frequency. At the same time, due to the high clock frequency of this technology, at-speed test is necessary for testing delay faults.In this paper, we present a scan-based BIST for RSFQ circuits which performs at-speed self-test including pseudo random pattern generation and response compression. We show that existing designs of pattern generators cannot be directly used for RSFQ and present new designs. Based on the scan architecture in [3], we design a new control strategy for at-speed self-test. We demonstrate that our new architecture supports testing at low overheads.
Mingye Li, Yunkun Lin, Sandeep Gupta 0001
VTS3
2024 Systematic Generation of Memristor-Transistor Single-Phase Combinational Logic Cells
abstract
The objective of our research is to create efficient methods and tools for the quick and thorough assessment of emerging digital circuit devices, facilitating the adoption of promising ones. In this work, we develop methods and tools for hybrid technology that combines memristors with MOS transistors and demonstrates their effectiveness. Although several types of memristor-transistor logic have been proposed, 15 years of research has created a small set of logic cells. We propose a systematic method for generating new and efficient memristor-transistor single-phase combinational logic cells. At the core of our approach is a cell enumerator, which enables us to explore a wide range of cell designs, including non-intuitive ones, and a data-driven inductive learning method, which identifies new properties of such cells and scales up our explorations. In conjunction with other completely new tools, these create a comprehensive and definitive library of logic cells. Our new cells provide significant improvements or significantly distinct Pareto-optimal alternatives for the few logic functions for which prior research has created cells. Importantly, our methods enable us to discover a previously unknown synergistic operation between memristors and transistors that occurs for specific cell topologies. We harness this synergy to develop a method for adding memristors to low-area pass-transistor circuits such that they produce strong output voltages and low power, including for patterns that cause ratioed operation. We have also developed a new memristor-transistor logic family, namely controlled-AND (cAND)/controlled-OR (cOR), which includes many of the best cells. We have also developed a constructive method for designing such cells.
Baishakhi Rani Biswas, Claire Yuan, Sandeep Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Design for testability (DFT) for RSFQ circuits
abstract
Superconducting electronics (SCE), especially Rapid Single Flux Quantum (RSFQ) logic, is being developed due to its high-performance and low power. In [1] –[3], we developed new static and delay fault models and an efficient automatic test pattern generator (ATPG) for testing both delay and static faults in RSFQ logic. However, test pattern application involves moving patterns and responses via long wires from the test equipment at room temperature to the chip under test in liquid helium. Due to the high cost associated with large numbers of such wires, testing is extremely expensive in absence of design for testability.We present a scan architecture for RSFQ circuits which enables the application of a large number of test patterns. Due to the unique characteristic of RSFQ, this scan architecture includes completely new scan cell design and a new scan control strategy. The on-chip test control logic enables scan chain to shift in test patterns from the test equipment at room temperature via a small number of wires, apply the pattern to the chip under test in parallel and at speed, and shift out the corresponding test response for checking. We demonstrate that our new scan architecture supports testing at low overheads.
Mingye Li, Yunkun Lin, Sandeep Gupta 0001
VTS3
2022 Fault-coverage Maximizing March Tests for Memory Testing
abstract
Every well-known march test for memories was generated to efficiently achieve 100% coverage of a target set of fault types. The question we pursue is: What to do if 100% coverage of the given target set cannot be achieved under tight constraints on test cost? We first study an obvious option: Remove some fault types from the given target set until a new or well-known test can cover 100% of the remaining fault types under the given test cost constraint. We find that this approach leaves significant room for improvement. We then pursue a different option and develop a new method which uses the original target set of fault types and generates a march test that maximizes the fault coverage under the given tight constraint on test cost. Our method generates fault-coverage maximizing tests for a wide range of target sets of fault types. A comparison with well-known march tests with equal lengths demonstrates that our new march tests provide significantly higher coverage for various sets of fault types. Importantly, our new march tests provide graceful decrease in fault coverage as we tighten constraints on test length. Hence our method and new march tests enable tradeoffs between test quality and test cost and provide a new direction of memory test research focused on fault-coverage-maximization.
Feng Yun, Yunkun Lin, Lou Yunfei, Vaibhav Gera, Boxuan Li, Vennela Chowdary Nekkanti, Aditya Rajendra Pharande, Kunal Sheth, Meghana Thommondru, Guizhong Ye, Sandeep Gupta 0001
ITC12
2022 Memristor-Specific Failures: New Verification Methods and Emerging Test Problems
abstract
We study two types of memristor-specific logic failures in Memristor Ratioed Logic (MRL), an extensively studied Memristor-CMOS hybrid logic design style. Cascading failures have been previously observed as causing significant voltage degradation, and hence logic values errors, in certain MRL logic circuits. Here, we present the first systematic study of this type of logical error and identify its key properties as a function of circuit structure and patterns applied. We then propose a method to generate patterns that cause the worst-case output voltage for a given MRL circuit and hence facilitate pre-fabrication verification. We then present the first study of another type of logic failure for voltage controlled memristor devices, namely a race when memristors in series, and with the same polarity, switch states from Ron to Roff. We show that such a race can cause non-deterministic behavior depending on circuit structure, values of memristor parameters, and the initial states of memristors when a pattern is applied. We then generate patterns and initial states that excite such race and hence potentially cause logic errors to enable verification.
Baishakhi Rani Biswas, Sandeep Gupta 0001
VTS2
2022 Methods for testing path delay and static faults in RSFQ circuits
abstract
Superconducting electronics (SCE), especially Rapid Single Flux Quantum (RSFQ) logic, is being developed due to its high-performance and low power. In [1] [2], we developed new static and delay fault models and an automatic test pattern generator (ATPG) for path delay faults in RSFQ logic. Here we develop a method for selecting path delay faults by identifying the subset of paths for which the delay can exceed the clock period under the main cause of delay faults for RSFQ, namely extreme process variations. We show that this dramatically reduces the number of delay tests required due to the characteristics of gate-level pipelined design, a necessary requirement for RSFQ. We also extend our method to be the first ATPG to generate tests for RSFQ-specific static fault models derived in [1]. We demonstrate that our new ATPG achieves very high coverage of static and delay faults with small numbers of patterns.
Mingye Li, Sandeep Gupta 0001
VTS3
2021 HW-BCP: A Custom Hardware Accelerator for SAT Suitable for Single Chip Implementation for Large Benchmarks
abstract
Boolean Satisfiability (SAT) has broad usage in Electronic Design Automation (EDA), artificial intelligence (AI), and theoretical studies. Further, as an NP-complete problem, acceleration of SAT will also enable acceleration of a wide range of combinatorial problems.
Soowang Park, Jae-Won Nam, Sandeep Gupta 0001
ASP-DAC3
2021 From Specification to Silicon: Towards Analog/Mixed-Signal Design Automation using Surrogate NN Models with Transfer Learning
abstract
We propose a complete analog mixed-signal circuit design flow from specification to silicon with minimum human-in-the-loop interaction, and verify the flow in a 12nm FinFET CMOS process. The flow consists of three key elements: neural network (NN) modeling of the parameterized circuit component, a search algorithm based on NN models to determine its sizing, and layout automation. To reduce the required training data for NN model creation, we utilize transfer learning to improve the NN accuracy from a relatively small amount of post-layout/silicon data. To prove the concept, we use a voltage-controlled oscillator (VCO) as a test vehicle and demonstrate that our design methodology can accurately model the circuit and generate designs with a wide range of specifications. We show that circuit sizing based on the transfer learned NN model from silicon measurement data yields the most accurate results.
Juzheng Liu, Shiyu Su, Meghna Madhusudan, Mohsen Hassanpourghadi, Samuel Saunders, Rezwan A. Rasul, Jiang Hu 0001, Arvind K. Sharma, Sachin S. Sapatnekar, Ramesh Harjani, Anthony Levi, Sandeep Gupta 0001, Mike Shuo-Wei Chen
ICCAD14
2020 Data-driven fault model development for superconducting logic
abstract
Superconducting technology is being seriously explored for certain applications. We propose a new clean-slate method to derive fault models from large numbers of simulation results. For this technology, our method identifies completely new fault models - overflow, pulse-escape, and pattern-sensitive - in addition to the well-known stuck-at faults.
Mingye Li, Sandeep Gupta 0001
ITC3
2020 Aging-resilient SRAM design: an end-to-end framework
abstract
The performance of transistors degrades due to aging. Bias temperature instability (BTI) is the most prominent aging mechanism in nano-scale CMOS technologies. Aging degradation causes lifetime failures and lowers the quality of shipped chips. We have developed an end-to-end SRAM design framework to maximize the aging resilience under the given constraints. Specifically, we analyze the impact of aging in SRAM peripheral circuits, including address decoder, precharge, write circuit and sense amplifiers (SAs). We explore the efficiency of error-correcting codes (ECC) to combat aging by quantifying the area and delay overheads of ECC and estimating the lifetime yield and DPPM of SRAMs with ECC, respectively. We also calculate the soft error resilience when ECC is used to repair aging failures. After comparing approaches based on cell sizing and ECC in terms of overheads, lifetime yield and DPPM, we can choose either one or a combination of these approaches to identify the optimal design against aging under the given constraints. We integrate our methods into an existing SRAM compiler, CACTI [1], to provide the end-to-end capability to designers.
Xuan Zuo, Sandeep Gupta 0001
VTS2
2019 Multi-cell characterization: Developing robust cells and abstraction for Rapid Single Flux Quantum (RSFQ) Logic
abstract
RSFQ, a Josephson-junction based technology, is becoming attractive due to its low energy and high speed. Researchers have designed cells and built circuits via composition of cells. Some designs have been fabricated and their functionalities and performance verified. This hierarchical approach relies on an abstraction for characterization of cells and composition of cells to design circuits. Researchers have developed such abstractions and methods for characterization of RSFQ cells. However, some instances of cells that are certified as being robust during cell characterization fail when incorporated into circuits. This motivated this research, which we view as the first step in the development of a systematic methodology for verification of cells and circuits in emerging technologies (such as RSFQ) leading to development of more robust abstractions and methods. In this paper, we present a new method for characterization of RSFQ cells to expose a much larger set of vulnerabilities, a systematic approach for identifying the root causes of these vulnerabilities to guide the refinement of cells designs, and a new way to extend test generation approaches to perform design validation at the circuit level. We demonstrate that our new methods and tools expose a large number of vulnerabilities and help identify root causes leading to refined cell designs which almost completely eliminate these vulnerabilities. Finally, we describe our extensions of ATPG for circuit level verification and use it to verify that our refined cells can indeed be composed to create error-free circuits.
Sandeep Gupta 0001
ITC2
2019 Cache Design for Yield-per-Area Maximization: Switchable Spare Columns with Disabling (SSC-Disable)
abstract
Modern SRAMs have high failure rates due to defects and variations, especially during low-voltage operation in low power modes. At high failure rates, a direct application of spare rows and columns approaches incurs high overheads. Since a recent paper [1] says that error correcting codes (ECCs) are the only cost-effective approach under high failure rates, we analyzed the strengths of ECC. This showed that, at high overheads, ECC provides one key advantage: the ability to correct failures in cells in different locations in different rows. To find a low overhead solution, we turn to the Divided Wordline/Bitline (DWL+DBL) [2] approach, a spares-based approach designed to provide exactly this capability. Since DWL+DBL also has high overheads, we develop our ideas for simplifying ECC as well as DWL+DBL and derive our new approach called Switchable Spare Columns (SSC), where a spare column can replace failing cells in different locations in different rows at lower overheads. To further reduce overheads, we propose SSC-Disable, which combines SSC with cache block disabling [3]. We then develop a method to find globally efficient cache designs with SSC or SSC-Disable. This method maximizes yield-per-area (YPA) under user-specified constraints on delay and power overheads. We show that the proposed SSC-Disable approach significantly improves yield compared to state-of-the-art ECC-based approaches.
Soowang Park, Sandeep Gupta 0001
VTS2
2019 Automatic Test Pattern Generation for timing verification and delay testing of RSFQ circuits
abstract
Rapid Single Quantum Flux (RSFQ) logic, based on Josephson Junctions (JJs), is seeing a resurgence as a way for providing high performance in the era beyond the end of physical scaling of CMOS. Since it uses fabrication processes with large feature sizes, the defect density for RSFQ is dramatically lower than its CMOS counterpart. Hence, process variations and other RSFQ-specific non-idealities become the major causes of chip failures. Because of the nature of its quantized pulse-based operation, even highly-distorted pulses are interpreted logically correctly by cells, but the timings is affected. Therefore, timing verification and delay testing increase in importance in RSFQ. In this paper, we address several radically new phenomena in RSFQ technology, especially the existence of single-pattern delay tests and the need to propagate delayed values via multiple pipeline stages. We then characterize cells under process variations and identify delay excitation conditions, sensitization conditions, and conditions for propagation of the logic errors caused by process variations. We then propose a completely new ATPG paradigm which utilizes these new phenomena to select target delay subpaths and generate test patterns that are guaranteed to excite the worst-case delay along each target delay sub-path. Finally, we present Monte Carlo simulation results for benchmark circuits with process variations to demonstrate the effectiveness of the vectors generated by our new ATPG.
Sandeep Gupta 0001
VTS2
2019 A New Method for Software Test Data Generation Inspired by D-algorithm
abstract
Test generation for digital hardware is highly automated, scalable (in practice), and provides high test quality. In contrast, current software automatic test data generation approaches suffer from low test quality or high complexity. While mutation-oriented constraint-based test data generation for software was proposed to generate high quality test data for real program bugs, all existing approaches require symbolic analysis for the whole program, and hence are not scalable even for unit testing, i.e., testing the lowest-level software modules. We propose a new method inspired by hardware D-algorithm and divide and conquer for software test data generation. To reduce runtime complexity and improve scalability, we combine global structural analysis and a sequence of small reusable symbolic analyses of parts of the program, instead of symbolically executing each mutated version of the entire program. We also propose a multi-pass test generation system to further reduce runtime complexity and compact test data. We compare our tools with one of the best software test generation tools (EvoSuite[20], which won the SBST 2017 tool competition) and demonstrate that our approach generates higher quality unit tests in a scalable manner and provides a compact set of tests.
Sandeep Gupta 0001, William G. J. Halfond
VTS2
2019 Collaborative circuit designs using the CRAFT repository
Adam Brinckman, Ewa Deelman, Sandeep Gupta 0001, Jarek Nabrzyski, Soowang Park, Rafael Ferreira da Silva, Ian J. Taylor, Karan Vahi
Future Gener. Comput. Syst.3
2017 Keynote address tribute to Professor Mel Breuer: Contributions to CAD and Test
abstract
This keynote is a tribute to the late Prof. Mel Breuer, entitled Contributions to CAD and Test. It is organized by Sandeep Gupta. A panel of three prominent speakers gives this keynote. The three speakers are Miron Abramovici, Magdy Abadir and Sridhar Narayanan.
Sandeep Gupta 0001, Miron Abramovici, Magdy Abadir, Sridhar Narayanan
VTS1
2017 Asymmetric sizing: An effective design approach for SRAM cells against BTI aging
abstract
The rate of aging of ICs is increasing with the continued reduction in feature sizes of devices. Bias temperature instability (BTI) is considered to be the major reliability hazard in nano-scale CMOS and causes stability degradation of SRAM cells. Some of the SRAM cells functioning properly at fabrication may fail during their desired lifetime due to aging. This will cause large aging quality loss. This paper addresses one key characteristic of aging, namely differential aging. This occurs due to the characteristics of data typically stored in SRAMs. After carefully studying the impact of differential aging on SRAM cells, we propose an asymmetric sizing approach for SRAM cells to maximize the probability of correct operation after m months of usage considering process variations. Our experiment results show that the asymmetric design can achieve much better aging quality loss (90× better) with optimal lifetime yield per area compared to the symmetric SRAM cell designs.
Xuan Zuo, Sandeep Gupta 0001
VTS2
2016 Using hardware testing approaches to improve software testing: Undetectable mutant identification
abstract
Over four decades of R&D has delivered near-universal automation of test generation for digital hardware. In contrast, software testing has limited automation and hence suffers from low test quality and high cost. One of the important reasons for this difference is that hardware ATPG is fault oriented. We note that the notion of a mutation in software testing is very similar to the notion of a fault in hardware testing, but current research on mutant oriented test generation for software is not extensive and the application is impeded by the scalability and undetectable mutant problems. This paper represents the first step in our identification of the similarities between software testing and hardware testing, and applying important insights from hardware testing to improve existing software mutant oriented testing. In particular, we propose the first approach for local analysis in software testing to identify mutants that are undetectable and demonstrate that our approach is effective and much more scalable than the state of the art.
Sandeep Gupta 0001
VTS2
2016 SRAM yield-per-area optimization under spatially-correlated process variation
abstract
Spatial correlation of process variation increases the probability that nearby transistors have similar variation values. However, the impact of spatial correlation on yield enhancement techniques, such as spare rows and columns for SRAM arrays, has not been well understood. In this paper we find that the clustering of failing cells caused by spatially-correlated process variation makes the SRAM array yield much less sensitive to the number of spare rows and columns. We then propose a coarse-to-fine search method to optimize yield-per-area (YPA) of SRAM systems under correlated process variation. We show that the proposed method can identify the optimal design efficiently.
Jizhe Zhang, Sandeep Gupta 0001
VTS2
2016 Process variation oriented delay testing of SRAMs
abstract
With continuing technology scaling, process variation is increasing and becoming an important cause of memory failure [1]. Most previous memory tests that were developed for delay faults, including GALPAT [2][3][4] and WCGD [3], focus on defects. In this paper we show that a different test strategy is necessary for process variation-induced delay faults (VIDFs) because variations are widespread. In particular, we determine that address dependent variation-induced failure mechanisms are likely to escape previous memory tests. We then propose a general approach to cover VIDFs in various versions of SRAM designs. We use our new method to generate O(n) tests and use extensive simulations to demonstrate that our new tests achieve nearly perfect coverage of VIDFs for different versions of SRAM design. Then we efficiently integrate our new tests for variations with tests for delay defects and clearly demonstrate the efficiency and effectiveness of our new combined memory tests compared to previously known memory tests for delay faults.
Xuan Zuo, Sandeep Gupta 0001
VTS2
2015 PPB: Partially-working processors binning for maximizing wafer utilization
abstract
Hardware redundancy, such as spare processors and cores, has been added to chip multi-processors (CMPs) to improve yield while sustaining all functionalities of CMPs. During post-silicon testing, spares processors and cores are used for repair. Even after repair, some CMPs may have processors with insufficient number of cores; in such CMPs some processors are disabled and such chips are sold at lower prices to improve yield per area. Despite binning on the number of processors, substantial functional resources are wasted in disabled components. In this work, we propose a new utility function and a new repair algorithm which enable utilization of every working core on a CMP. We demonstrate the benefits of the proposed approach for benchmarks from ISPASS and Nvidia CUDA SDK using GPGPU-sim to compute the instructions per cycle (IPC). Results show that our design and repair approaches provide above 50% IPC per wafer area even with 10x the current defect density.
Da Cheng, Sandeep Gupta 0001
VTS2
2015 A multi-layered methodology for defect-tolerance of datapath modules in processors
abstract
Technology scaling increases circuits' susceptibility to manufacturing imperfections and dramatically decreases processor yields. Traditional defect-tolerance approaches add explicit redundant circuitry to improve yield and hence are very expensive for datapath modules in processors. We propose a multi-layered methodology to develop new and efficient defect-tolerance approaches for processors. Specifically, we develop a microarchitecture layer approach for arithmetic logic units (ALU), a circuit layer approach for multipliers, and an ISA layer approach for floating-point units (FPU). We demonstrate that our three approaches improve performance-per-fabricated-die-area of a modern processor core by 3.5%, 2.4%, and at least 9%, and hence collectively provide significant gains.
Hsunwei Hsiung, Sandeep Gupta 0001
VTS2
2014 A Resizing Method to Minimize Effects of Hardware Trojans
abstract
Due to emerging threats of hardware Trojan insertion, many techniques to detect Trojans and protect the original design have been developed. However, an intelligent adversary is expected to be aware of every state-of-the-art detection technique and develop new countermeasures to make a majority of these obsolete. In this paper, we propose and analyze a new attack scenario that targets every known non-destructive detection method and imposes negligible impact on every measurable circuit parameter. We start by introducing our key ideas for hiding the delay impact of the Trojan via gate resizing. The gate resizing problem is then formulated and implemented. Our method redesigns the circuit at minimal cost without affecting the functionality of the circuit, can be applied in conjunction with any other attack scenario, and maintains its benefit independent of the type and functionality of Trojan circuitry. Finally, via extensive experiments on benchmarks we demonstrate that our method greatly increases the difficulty of detecting Trojans via any combination of delay, current, and power measurements and has very small area overhead.
Byeongju Cha, Sandeep Gupta 0001
ATS2
2014 Optimal Redundancy Designs for CNFET-Based Circuits
abstract
Substantial imperfections in carbon nanotube (CNT) field-effect transistors (CNFETs) are one key obstacle to the demonstration of large-scale CNFET circuits. In this paper, we first categorize transistors based on the impact of resizing on yield improvement and delay penalty for logic circuits. Then we propose an approach to size transistors in different categories by using redundant CNTs to improve yield/area with user-specified limit on delay penalty. We then propose a hybrid redundancy approach for memory arrays by optimally combining redundant CNTs approach with the traditional spare columns (rows) approach. Experimental results show that the proposed approach provides significant improvements in yield/area for logic circuits at very low increase in delays. For SRAM, spare columns (rows) approach becomes ineffective when it is applied alone since spare columns (rows) themselves have very low yield. The proposed hybrid approach for memory array provides 18% improvement in yield/area compared to a redundant-CNTs-only approach as well as reduces delay penalty on address decoder from 19.2% to 15.7%.
Da Cheng, Sandeep Gupta 0001
ATS4
2014 SRAM Array Yield Estimation under Spatially-Correlated Process Variation
abstract
In this paper we propose a systematic method to estimate SRAM array yield considering spatially-correlated process variation. Although many parameter variations have been shown to be spatially-correlated, this phenomenon has been neglected in SRAM yield estimation due to its small impact for small SRAMs. However, its impact on array yield is significant for SRAM arrays of realistic sizes. In this paper we show that three ways in which the spatially-correlated term can be simplified to use existing approaches all lead to high levels of inaccuracy. In particular even the closest yield estimate for 32KB SRAM has more than 20% error. To rectify this we develop a new two-stage SRAM array yield estimation framework that accurately considers spatial correlation. In our approach, first the correlated variation terms are sampled at the array level to remove cell-to-cell dependency. Then we estimate cell yield using our innovative simulation-reusing scheme which dramatically decreases the number of SPICE-like circuit simulations. To the best of our knowledge, this is the first method that can efficiently and accurately estimate SRAM array yield considering spatial correlation.
Jizhe Zhang, Sandeep Gupta 0001
ATS2
2014 An energy-aware fault tolerant scheduling framework for soft error resilient cloud computing systems
abstract
For modern high performance systems, aggressive technology and voltage scaling has drastically increased their susceptibility to soft errors. At the grand scale of cloud computing, it is clear that soft error induced failures will occur far more frequently, but it is unclear as to how to effectively apply current error detection and fault tolerance techniques in scale. In this paper, we focus on energy-aware fault tolerant scheduling in public, multi-user cloud systems, and explore the three-way tradeoff between reliability (in terms of soft error resiliency), performance and energy. Through a systematically optimized resource allocation, error detection approach selection, virtual machine placement, spatial/temporal redundancy augmentation and task scheduling process, the cloud service provider can achieve high error coverage and fault tolerance confidence while minimizing global energy costs under user deadline constraints. Our scheduling algorithm includes a static scheduling phase that operates on task graph based workload inputs prior to execution, and a light-weight dynamic scheduler that migrates tasks during execution in case of excessive reexecutions. All schedules are evaluated on a runtime simulation engine that (1) mimics the performance fluctuations in cloud systems, and (2) supports the injection of arbitrary fault patterns. Compared to current virtual machine or task replication techniques, we are able to reduce overall application failure rates by over 50% with approximately 76% total energy overhead.
Sandeep Gupta 0001, Yanzhi Wang 0001, Massoud Pedram
DATE2
2014 Optimizing redundancy design for chip-multiprocessors for flexible utility functions
abstract
Yield of chip-multiprocessors (CMP) can be improved by adding spares to the design. The optimal spare configuration has previously been derived for certain evaluation metrics, such as yield per area, performance-averaged yield, and so on. However, all previous approaches are limited to the scenario where only those chips which have the full-configuration (i.e., have n working processors, where n is the number of processors in the CMP's specifications) can be sold. In the meantime, yield problems have forced vendors of high-volume CMPs to sell chips with different numbers of working processors. The purpose of this paper is to extend our recent framework for the optimal redundancy design to systematically capture such flexibilities. We first define utility function for CMPs in terms of two functions: (i) the number-of-processors-binning (NPB) function, which captures the range of the number of (enabled) working processors over which a chip can be sold, and (ii) the value function, which captures how the value of a chip to the user depends upon the number of processors enabled in the chip. We use case studies to identify the relationships between utility functions and optimal spare configurations. Then we extend our branch and bound algorithm and incorporate a dynamic programming approach to develop the first framework for optimal and e-optimal (i.e., within e of optimal, typically at much lower area overhead) redundancy designs for different utility functions. We demonstrate that even in an era of high defect density which leads to extremely low yield, we are able to combine approaches in a way that provide figure of merit that is up to 83.8% of that of the ideal case, i.e., for a process with zero defect density.
Da Cheng, Sandeep Gupta 0001
ITC2
2014 Maximizing Yield per Area of Highly Parallel CMPs Using Hardware Redundancy
abstract
The manufacturing yield of chip multiprocessors (CMPs) has become a significant problem as more transistors are integrated onto a single die and the defect rate keeps increasing for “end-of-Moore” nano-scale CMOS technologies. Since such CMP designs usually have significant structural symmetry, adding spare copies to these should be an effective method for increasing yield per area, as is the case for memories. However, a systematic approach to add spare copies to optimize CMP yield per area has never been developed, primarily due to the lack of: 1) a general model of CMP architectures and 2) a practically-useable model for computing areas of chip versions with different configurations of spare copies. This paper develops such models and, in conjunction with a systematic approach for enumerating a wide range of spare configurations, uses these to compute the area overhead and yield for each configuration. In particular, this paper proposes a general spare cores sharing technique to maximize yield per area of any CMP by efficiently traversing the design space for adding spare cores. Experimental results show that the advantage of the proposed approach over traditional approaches increases with continued technology scaling. Specifically, the proposed approach achieves \(2\times \) yields per area over previous approaches for 32 nm and 22 nm technologies. Also, the obtained yield per area values provided by our approach are around 70% of that obtained for the ideal scenario where defect density is zero and no redundancy is added.
Da Cheng, Sandeep Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 A New March Test for Process-Variation Induced Delay Faults in SRAMs
abstract
Process variations are growing with technology scaling towards nano-scale. This brings new challenges to the design of memory modules, which are often the first circuits to be fabricated using a new technology and usually designed with critical timing. We observed that several delay faults, which are dependent on address transitions, may escape traditional march tests. This paper presents a new march test WT that targets such delay faults. Through Monte Carlo simulations and analytical studies on SRAM designs using an industrial 65nm process, we have demonstrated that WT provides a faulty-chip-coverage that is close to 100%. Most importantly, this is the first march test with test length that targets address-dependent delay faults and hence the first delay test which can be used in practice.
Da Cheng, Hsunwei Hsiung, Bin Liu 0004, Ramesh Govindan, Sandeep Gupta 0001
Asian Test Symposium7
2013 Interplay of Failure Rate, Performance, and Test Cost in TCAM under Process Variations
abstract
As process variations grow with technology scaling, failure rates increase and are predicted to be so high as to render devices unusable for computing domain. In order to continue to benefit from scaling, three dimensions can be explored: increasing operational margins, testing, and the resilience of applications. The solutions in each dimension bring out different trade-offs in yield, failure rate, test cost, and performance. We use a ternary content-addressable memory (TCAM) as a case study to better understand these trade-offs. We develop two new delay tests for TCAMs, and a new method to estimate yield, failure rate, test cost, and performance of TCAM under process variations when various tests are used. Our results show that with little test overhead and negligible yield loss, our new tests can significantly decrease the failure rates of TCAMs shipped to customers.
Hsunwei Hsiung, Da Cheng, Bin Liu 0004, Ramesh Govindan, Sandeep Gupta 0001
Asian Test Symposium5
2013 Trojan detection via delay measurements: a new approach to select paths and vectors to maximize effectiveness and minimize cost
abstract
One of the growing issues in IC design is how to establish trustworthiness of chips fabricated by untrusted vendors. Such process, often called Trojan detection, is challenging since the specifics of hardware Trojans inserted by intelligent adversaries are difficult to predict and most Trojans do not affect the logic behavior of the circuit unless they are activated. Also, Trojan detection via parametric measurements becomes increasingly difficult with increasing levels of process variations. In this paper we propose a method that maximizes the resolution of each path delay measurement, in terms of its ability to detect the targeted Trojan. In particular, for each Trojan, our approach accentuates the Trojan's impact by generating a vector that sensitizes the shortest path passing via the Trojan's site. We estimate the minimum number of chips to which each vector must be applied to detect the Trojan with sufficient confidence for a given level of process variations. Finally, we demonstrate the significant improvements in effectiveness and cost provided by our approach under high levels of process variations. Experimental results on several benchmark circuits show that we can achieve dramatic reduction in test cost using our approach compared to classical path delay testing.
Byeongju Cha, Sandeep Gupta 0001
DATE2
2013 Using explicit output comparisons for fault tolerant scheduling (FTS) on modern high-performance processors
abstract
Soft errors and errors caused by intermittent faults are a major concern for modern processors. In this paper we provide a drastically different approach for fault tolerant scheduling (FTS) of tasks in such processors. Traditionally in FTS, error detection is performed implicitly and concurrently with task execution, and associated overheads are incurred as increases in software run-time or hardware area. However, such embedded error detection (EED) techniques, e.g., watchdog processor assisted control flow checking, only provide approximately 70% error coverage [1, 2]. We propose the idea of utilizing straightforward explicit output comparison (EOC) which provides nearly 100% error coverage. We construct a framework for utilizing EOC in FTS, identify new challenges and tradeoffs, and develop a new off-line scheduling algorithm for EOC. We show that our EOC based approach provides higher error coverage and an average performance improvement of nearly 10% over EED-based FTS approaches, without increasing resource requirements. In our ongoing research we are identifying a richer set of ways of applying EOC, by itself and in conjunction with EED, to obtain further improvements.
Sandeep Gupta 0001, Melvin A. Breuer
DATE2
2013 Gate delay modeling for pre- and post-silicon timing related tasks for ultra-low power CMOS circuits
abstract
Power is increasingly the primary design constraint for chip designers and one of the main techniques for addressing this concern is aggressive voltage scaling. Device variability increases with voltage scaling and significantly affects gate delays at low voltages. Although existing delay models for near- and sub-threshold circuits capture the effects of variability on gate delays, they do not capture advanced delay phenomenon such as multiple input switching (MIS; also known as near-simultaneous transitions) at inputs of a gate. As a result, most of these gate delay models often grossly underestimate worst case delays, leading to selection of non-critical paths and generation of delay-inferior vectors for post-silicon timing related tasks. In this paper we present extensive experimental results to demonstrate that MIS has significant impact (around 30-40%) on delays of near-and sub-threshold nominal gates. We develop our model which guarantees that the minimum and maximum delay values it computes are guaranteed to bound the corresponding delay values in silicon. We show that our model has practical run-time complexity and works equally well for super-, near- and sub-threshold circuits. In particular, via extensive experimentations we show that our model never underestimates the delay and tightly bounds the actual delays. We also illustrate trade-offs between tightness of such bounds, their impact on validation cost, and runtime complexity.
Prasanjeet Das, Sandeep Gupta 0001
ICCD2
2013 Extending pre-silicon delay models for post-silicon tasks: Validation, diagnosis, delay testing, and speed binning
abstract
All post-silicon tasks - validation, diagnosis, delay testing, and speed-binning - must be carried out by applying vectors to actual chips, and capturing and analyzing responses. Yet, vectors used must be generated and analyzed using pre-silicon models of the circuit. Three comprehensive industrial studies demonstrate that existing approaches for generating such vectors are inadequate, and one major weakness is that existing delay models either do not capture process variations or do not capture advanced delay phenomenon that significantly affect delays. Hence, existing models underestimate the worst case delay leading to selection of non-critical paths and generation of vectors that do not invoke worst case delays. In this paper, we propose a simple notion of bounding approximation and show how it can extend any existing delay model to also capture process variations and to eliminate any underestimation. The main question we investigate is how best to use this approach to select paths and generate or evaluate vectors for post silicon tasks. In particular, we study whether it is better to use this approach to bound simple pin-to-pin delay models or more advanced delay models. At the level of timing analysis, bounded versions of pin-to-pin delay models have lower run-time complexity but looser bounds. However, we conduct path selection for delay testing and vector generation for delay validation and show that bounded versions of more advanced delay models are significantly more efficient in terms of validation cost and runtime complexity.
Prasanjeet Das, Sandeep Gupta 0001
VTS2
2012 Efficient Trojan Detection via Calibration of Process Variations
abstract
In this paper we present an efficient method to detect hardware Trojans under high levels of process variations, by measuring delays for vectors generated using a minimally delay-invasive Trojan model. Our method focuses on various sources of process variations and significantly reduces the effects of variations on delays by calibrating delays measured on each fabricated chip. Using test structures and additional measurements on these, our approach significantly reduces the number of chips that we need to test to detect minimally-invasive Trojans with a desired level of confidence. Our approach for tackling high levels of process variations can be used in conjunction with any other type of parametric Trojan detection strategy to significantly improve its efficiency.
Byeongju Cha, Sandeep Gupta 0001
Asian Test Symposium2
2012 Salvaging chips with caches beyond repair
abstract
Defect density and variabilities in values of parameters continue to grow with each new generation of nano-scale fabrication technology. In SRAMs, variabilities reduce yield and necessitate extensive interventions, such as the use of increasing numbers of spares to achieve acceptable yield. For most microprocessor chips, the number of SRAM bits is expected to grow 2× for every generation. Consequently, microprocessor chip yields will be seriously undermined if no defect-tolerance approach is used. In this paper, we show the limits of the traditional spares-based defect-tolerance approaches for SRAMs. We then propose and implement a software-based approach for improving cache yield. We demonstrate that our approach can significantly increase microprocessor chip yields (normalized with respect to chip area) compared to the traditional approaches, for upcoming fabrication technologies. In particular, we demonstrate that our approach dramatically increases effective computing capacity, measured in MIPS-per-unit-chip-area. Our approach does not require any hardware design changes and hence can be applied to improve yield of any modern microprocessor chip, incurs low performance penalty only for the chips with unrepaired defects in SRAMs, and adapts without requiring any design changes as the yield improves for a particular design and fabrication technology.
Hsunwei Hsiung, Byeongju Cha, Sandeep Gupta 0001
DATE3
2012 Towards systematic roadmaps for networked systems
abstract
Networked systems have benefited from unprecedented growth in hardware capabilities, but, as we move closer to the end of the Moore's law era, future networked systems are likely to be more constrained by hardware capabilities than they have been in the past. We take the position that the networking community should, in response to this development, proactively and systematically develop networking roadmaps, which attempt to predict how trends in hardware capabilities will impact networked systems. In this paper, we discuss a possible methodology for developing networking roadmaps, and present two case studies that illustrate the methodology and reveal how increasing hardware unreliability can affect the performance of routing and transport protocols.
Bin Liu 0004, Hsunwei Hsiung, Da Cheng, Ramesh Govindan, Sandeep Gupta 0001
HotNets5
2012 A design flow to maximize yield/area of physical devices via redundancy
abstract
This paper deals with using redundancy to maximize the number of “workable” die one can produce from a silicon wafer. When redundant modules are used to enhance yield, several issues need to be addressed, such as power, performance degradation, testability, area, and partitioning the original logic design into modules. The focus of this paper is on the long ignored issue of partitioning and clustering to form modules that are to be replicated. For this purpose we propose a design flow with two phases. The first phase consists of a partitioning process that generates all combinational logic blocks (CLBs) of a given logic circuit. CLB partitioning addresses design and test constraints such as timing closure and testing complexity, by using redundancy at finer levels of granularity. In the second phase we carry out an overall optimization of the generated CLBs to find the optimal level of granularity for replication to maximize yield/area. Using a real design (OpenSPARC T2) and defect densities projected in the near future, the experimental results show that the output of our design flow outperforms the traditional redundant design with spare core, e.g. we achieved 1.1 to 13.3 times better yield/area as a function of defect density.
Mohammad Mirza-Aghatabar, Melvin A. Breuer, Sandeep Gupta 0001
ITC3
2011 Yield-per-Area Optimization for 6T-SRAMs Using an Integrated Approach to Exploit Spares and ECC to Efficiently Combat High Defect and Soft-Error Rates
abstract
Memories constitute increasing proportions of most digital systems and memory-intensive chips lead the migration to new nanometer fabrication processes. With each process generation, process variations and defect rates are increasing, at the same time, cells are becoming more susceptible to soft errors with technology shrink. SRAMs will thus require increasing numbers of spares and stronger error correcting codes (ECCs), incurring higher area overheads and access-time penalties. Our overall objective is to develop new systematic approaches for designing defect-tolerant 6T-SRAMs optimized in terms of yield-per-area under high defect rates and high soft error rates, for given soft-error resilience and access-time requirements. In this paper, we analyze the key tradeoffs associated with using different numbers of spares and ECCs with different strengths. In addition to considering the usual role of each -- i.e., spares to combat defects and ECC to combat soft errors -- we also consider the ability of ECC to combat those defects which cannot be masked using available spares. We develop a new model that captures the benefits -- yield and resilience to soft errors -- of spares and ECC in an integrated manner. We also characterize area and access time overheads of the spares and the ECC scheme. We then integrate above into a framework to design 6T-SRAMs that optimizes yield-per-area. We demonstrate that the proposed approach provides dramatic improvements in yield and yield-per-area without compromising resilience to soft errors.
Jae Chul Cha, Sandeep Gupta 0001
Asian Test Symposium2
2011 On Generating Vectors for Accurate Post-Silicon Delay Characterization
abstract
In this paper, we propose a new method to generate vectors for post-silicon delay characterization, especially for exposing delay marginalities during post-silicon validation and speed binning during testing. Our method generates vectors that are guaranteed to excite the worst-case delays of fabricated chips without introducing any pessimism. It embodies several innovations, including a resilient gate delay model that captures multiple input switching effects and process variations, new conditions that vectors must satisfy to invoke the maximum delay of a target path, and a new approach to generate multiple vectors (vector-spaces) that are collectively guaranteed to invoke the worst-case delay of the target path. We present experimental results for benchmark circuits to demonstrate the effectiveness of our method for post-silicon validation and describe how the generated vectors can be adapted for speed binning.
Prasanjeet Das, Sandeep Gupta 0001
Asian Test Symposium2
2011 A new circuit simplification method for error tolerant applications
abstract
Starting from a functional description or a gate level circuit, the goal of the multi-level logic optimization is to obtain a version of the circuit that implements the original function at a lower cost. For error tolerant applications - images, video, audio, graphics, and games - it is known that errors at the outputs are tolerable provided that their severities are within application-specified thresholds. In this paper, we perform application level analysis to show that significant errors at the circuit level are tolerable. Then we develop a multi-level logic synthesis algorithm for error tolerant applications that minimizes the cost of the circuit by exploiting the budget for approximations provided by error tolerance. We use circuit area as the cost metric and use a test generation algorithm to select faults that introduce errors of low severities but provide significant area reductions. Selected faults are injected to simplify the circuit for the experiments. Results show that our approach provides significant reductions in circuit area even for modest error tolerance budgets.
Doochul Shin, Sandeep Gupta 0001
DATE2
2011 A novel software-based defect-tolerance approach for application-specific embedded systems
abstract
Traditional approaches for improving yield are based on the use of hardware redundancy (HR), and their benefits are limited for high defect densities due to increasing layout complexities and diminishing return effects. This research is based on an observation that completely correct operation of user programs can be guaranteed while using chips with one or more unrepairable memory modules if software-level techniques satisfy two condistions: (1) defects only affect a few memory cells rather than cause malfunction for the entire memory module, and (2) either we do not use any part of the memory affected by the un-repaired defect, or we do use the affected part, but only in a manner that does not excite the un repaired defect to cause errors. This paper proposes a software based defect-tolerance (SBDT) approach in combination with HR to utilize defective memory chips for application-specific systems. The proposed approach requires known and fixed program and information about defective locations for each memory module, hence this paper focuses on SoCs and other application-specific systems built around processors, such as DSP and graphics processors. We model an application program and defective memory copies as described next.
Da Cheng, Sandeep Gupta 0001
ICCD2
2010 HYPER: A Heuristic for Yield/Area imProvEment Using Redundancy in SoC
abstract
In this paper we present an efficient heuristic to significantly enhance yield/area of die in technologies where the inherent yield is low. Our technique makes use of a judicious use of redundancy and switching circuitry. Though our presentation is focused on pipeline (linear) structures, our techniques can be extended to apply to more general structures. The time complexity of our procedure is O(n3) for an n-stage pipeline.
Mohammad Mirza-Aghatabar, Melvin A. Breuer, Sandeep Gupta 0001
Asian Test Symposium3
2010 Algorithms to maximize yield and enhance yield/area of pipeline circuitry by insertion of switches and redundant modules
abstract
Increasing yield is important, especially for nano-scale technologies. Also, pipelines are an important aspect of many SoC architectures. In this paper we present new approaches to improve the yield and yield/area of pipeline architectures by using (1) an appropriate number of redundant copies for each module, and (2) sufficient steering logic resources. We present an optimal algorithm of time complexity O(n3) that adds redundant modules to an n-stage pipeline so as to maximize yield. Experimental results indicate that for parameter values of interests, this algorithm also improves the yield/area of the pipeline, especially when the yield for some modules is low.
Mohammad Mirza-Aghatabar, Melvin A. Breuer, Sandeep Gupta 0001
DATE3
2010 Approximate logic synthesis for error tolerant applications
abstract
Error tolerance formally captures the notion that - for a wide variety of applications including audio, video, graphics, and wireless communications - a defective chip that produces erroneous values at its outputs may be acceptable, provided the errors are of certain types and their severities are within application-specified thresholds. All previous research on error tolerance has focused on identifying such defective but acceptable chips during post-fabrication testing to improve yield. In this paper, we explore a completely new approach to exploit error tolerance based on the following observation: If certain deviations from the nominal output values are acceptable, then we can exploit this flexibility during circuit design to reduce circuit area and delay as well as to increase yield. The specific metric of error tolerance we focus on is error rate, i.e., how often the circuit produces erroneous outputs. We propose a new logic synthesis approach for the new problem of identifying how to exploit a given error rate threshold to maximally reduce the area of the synthesized circuit. Experiment results show that for an error rate threshold within 1%, our approach provides 9.43% literal reductions on average for all the benchmarks that we target.
Doochul Shin, Sandeep Gupta 0001
DATE2
2010 Design and test of latch-based circuits to maximize performance, yield, and delay test quality
abstract
The performance benefits of latch-based circuits have been known for some time. These benefits are due to the timing flexibility and skew-tolerance enabled by the ability of combinational logic blocks to borrow time from each other across the intervening level-sensitive latches. It has also been known that, by accommodating higher levels of process variations and small delay-defects, time borrowing can enhance yield at high clock frequencies. The main roadblock was that conventional scan-based delay testing approaches cannot be adapted from flip-flop-based (FF-based) circuits to latch-based circuits in a manner that can harvest above benefits. Recently, a scan-based delay testing approach has been proposed for latch-based circuits which holds the promise of harvesting the abovementioned performance and yield benefits. In this paper, we investigate two main questions. First, can this new scan-based delay testing approach provide high coverage of delay faults for all latch-based circuits - independent of the pervasiveness of time borrowing? Second, how do we design the circuit and develop tests so as to harvest maximal performance and yield benefits? We prove that the above delay testing approach for latch-based circuits obtains the maximum path delay fault coverage possible for any scan-based test methodology and this test quality is always greater than (or equal to) that obtainable for the corresponding FF-based circuit. We derive the conditions to satisfy during design and test development to guarantee maximal performance and yield benefits of latch-based designs vs. their FF-based counterparts. Hence, we show for the first time that it is possible for latch-based circuits to provide higher performance and yield and also to certify the higher performance via high delay test quality.
Kun Young Chung, Sandeep Gupta 0001
ITC2
2009 SIRUP: Switch Insertion in RedUndant Pipeline Structures for Yield and Yield/Area Improvement
abstract
Except for regular arrays, yield enhancement for high performance VLSI systems is usually addressed at the physical layers rather than at the architectural level. In addition, pipelines are prevalent in many SoC architectures. In this paper we present new architectural approaches and results to improve the yield and yield/area of pipelines by using redundancy and steering logic . We present a procedure of time complexity O(n) that finds the minimal number of switches to insert within an n-stage redundant pipeline of order q to improve yield. Experimental results indicate that for parameter values of interests, this procedure also improves the yield/area of the pipeline, especially when the yields for some modules are low.
Mohammad Mirza-Aghatabar, Melvin A. Breuer, Sandeep Gupta 0001
Asian Test Symposium3
2009 Tolerance of performance degrading faults for effective yield improvement
abstract
To provide a new avenue for improving yield for nano-scale fabrication processes, we introduce a new notion: performance degrading faults (pdef). A fault is said to be a pdef if it cannot cause a functional error at system outputs but may result in system performance degradation. In a processor, a fault is a pdef if it causes no error in the execution of user programs but may reduce performance, e.g., decrease the number of instructions executed per cycle. By identifying faulty chips that contain pdef's that degrade performance within some limits and binning these chips based on the their resulting instruction throughput, effective yield can be improved in a radically new manner that is completely different from the current practice of performance binning on clock frequency. To illustrate the potential benefits of this notion, we analyze the faults in the branch prediction unit of a processor. Experimental results show that every stuck-at fault in this unit is a pdef. Furthermore, 97% of these faults induce almost no performance degradation.
Tong-Yu Hsieh, Melvin A. Breuer, Murali Annavaram, Sandeep Gupta 0001, Kuen-Jong Lee
ITC4
2009 Modeling and test generation for worst-case performance evaluation of MAC protocols for wireless ad hoc networks
abstract
We propose a novel framework to critically analyze a given MAC protocol for wireless ad hoc networks with respect to its correctness criteria and performance metrics. The framework is composed of wanted state generation and test scenario generation algorithms. The wanted state generation algorithm generates a set of conditions that meet our study objective. The test generation algorithm then generates complete test scenarios that satisfy our objective, e.g., minimize the value of a particular performance metric. The core of our search engine utilizes novel algorithms that use combinations of goal-oriented backward and forward search and implications as well as heuristics that enable the generation of worst case scenarios in manageable complexity for our practical purposes. We demonstrate the effectiveness of our approach by using our framework to analyze the worst case performance, in terms of throughput and fairness of IEEE 802.11 for ad hoc networks. For all topologies, the worst case scenarios generated by our framework show the worst performance among all scenarios that we generate. The scenarios generated by our framework include the scenarios typically used for performance evaluation of IEEE 802.11 protocol. The case study of IEEE 802.11 shows that the complexity of our novel algorithms are quite practical.
Shamim Begum, Ahmed Helmy, Sandeep Gupta 0001
MASCOTS3
2009 Efficient Scheduling of Path Delay Tests for Latch-Based Circuits
abstract
In many high-speed parts of chips, latch-based circuits are used to enable time borrowing, where a block may take longer time than its nominal delay to complete its computation. This enables such circuits to attain high performance and yield. In [1] and [2], we focused on maximizing path delay fault coverage and proposed the first structural delay testing approach and the associated design-for-testability (DFT) for such circuits. This approach provides dramatically higher coverage of path delay faults. In this paper, we focus on minimizing test application cost for delay testing latch-based circuits while ensuring that maximum coverage is achieved. We show that conventional test scheduling methods may not be applicable due to the unique characteristics of latch-based circuits with time borrowing. We then formulate the minimization problem and propose two heuristic approaches. The experimental results show that, for many example circuits, the proposed approaches achieve overall test application costs that are within 5% of the corresponding lower-bounds.
Kun Young Chung, Sandeep Gupta 0001
VTS2
2009 Threshold Testing: Improving Yield for Nanoscale VLSI
abstract
Yields for digital very-large-scale-integration chips have been declining in the recent years, and the decline is accelerating as the technology moves deep into nanoscale. Recently, we have proposed the notion of error tolerance to improve yields for a wide range of high-performance digital applications, including audio, speech, video, graphics, visualization, games, and wireless communication. Error tolerance systematically codifies the fact that chips used in such applications can be acceptable despite having defects that produce erroneous outputs, provided that the errors are guaranteed to be of certain types and have severities within thresholds specified by the application. In this paper, we propose a new testing approach called threshold testing to practically exploit the notion of error tolerance for applications where errors with absolute numerical magnitudes lower than an application-specified threshold are acceptable. We propose a new automatic test pattern generator (ATPG) for threshold testing for single stuck-at faults. This test generator embodies several completely new techniques, including new approaches for directing the search for a test vector, new types of objectives, new types of necessary conditions, and new approaches to identify and exploit these conditions. We demonstrate that threshold testing can enhance yield and that it is practical in terms of test generation effort and test application costs. We also propose threshold fault simulators and ATPG for bridging and transition delay faults. We use these tools to show that the stuck-at-fault model is indeed a suitable model for threshold testing. This opens the way for developing low-cost tools for threshold testing that will provide high threshold coverage for realistic faults and defects and hence help provide higher yields in future nanoscale processes at low costs.
Sandeep Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2008 A Re-design Technique for Datapath Modules in Error Tolerant Applications
abstract
Scaling CMOS into nano-scale is decreasing yields. The concept of error tolerance has been proposed to reverse this trend by developing new test techniques for chips used in many applications, such as audio, video, graphics, games, and error-correcting codes for wireless communication. In such chips, manufacturing defects that induce errors with severities within specified thresholds (determined via analysis of applications) are deemed acceptable. In this paper, we develop an approach to re-design datapath modules to exploit acceptable errors to improve yield. Under the manufacturing yield model (Ym= Ypara* Yfunc), parametric yield (Ypara) improvement due to decrease in delay is more important than functional yield (Yfunc) improvement due to decrease in area. So our re-designing technique mainly focuses on reducing delay of datapath modules to improve parametric yield. In particular, we propose multiple approaches and apply them to improve yield of a wide range of adder architectures by exploiting error tolerance in these applications. Experiment results show that even for small thresholds on error severity, we can obtain significant improvements in manufacturing yield.
Doochul Shin, Sandeep Gupta 0001
ATS2
2008 A Multi-valued Algebra for Capacitance Induced Crosstalk Delay Faults
abstract
Capacitive crosstalk can slowdown transitions which can propagate to outputs and cause erroneous operation. Test generation methods such as XGEN and XGEN-E were proposed to generate tests for such failures. However, a drawback of these test generation methods is that a large proportion of faults are aborted. In this paper, we systematically derive a multi-valued algebra. We first show that a composite value system must be derived considering all operations performed by the algorithm that will use the value system. In particular, we identify that for our test generation algorithm it is imperative to consider its timing operations and their impact on what composite values can exist in our algebra. We also identify the fact that even some key procedures - namely the backtrace procedure as well as the search procedure - need to be modified to work with the composite value system. We derive a new 57-valued algebra and modify the key ATPG procedures to obtain a test generation methodology which we call XGEN-M ('M' for multi-valued). We present experimental results that demonstrate the superiority of XGEN-M compared to previous methods.
Arani Sinha, Sandeep Gupta 0001, Melvin A. Breuer
ATS2
2008 Multi-Vector Tests: A Path to Perfect Error-Rate Testing
abstract
The importance of testing approaches that exploit error tolerance to improve yield has previously been established. Error rate, defined as the percentage of vectors for which the value at a circuit's output deviates from the corresponding error-free value, has been identified as a key metric for severity. In error-rate testing every chip that has an error rate greater than or equal to a threshold specified by the application is unacceptable for the application and discarded; all other chips are acceptable. The objective of error-rate testing is to reject every unacceptable chip while accepting all (or a maximum number) of the acceptable chips. We previously showed that it is not always possible to generate a test set that detects all unacceptable faults, i.e., faults that cause an error rate greater than or equal to the threshold error rate, without detecting some of the acceptable faults, i.e., faults that cause an error rate less than the threshold. In this paper, we introduce the new notion of multi-vector testing and prove that this notion enables us to detect all unacceptable faults without detecting any of the acceptable faults. We derive an upper bound on the size of such a test for a general case. As this universal bound can be large in some cases, we use a structural approach and find much tighter upper bounds for special classes of circuits. Experiments on benchmark circuits show that the required test-sizes for arbitrary circuits are much lower than our universal bounds, and practically useful.
Shideh Shahidi, Sandeep Gupta 0001
DATE2
2008 Characterization of granularity and redundancy for SRAMs for optimal yield-per-area
abstract
Memories are significant proportions of most digital systems and memory-intensive chips continue to lead the migration to new nano-fabrication processes. As these processes have increasingly higher defect rates, especially when they are first adopted, such early migration necessitates the use of increasing levels of redundancy to obtain high yield (per area). We show that as we move into nanometer processes with high defect rates, the level of redundancy needed to optimize yield-per-area is sufficiently high so as to significantly influence design tradeoffs. We then report a first step towards considering the overheads of redundancy during design optimization by characterizing the tradeoffs between the granularity of a design and the level of redundancy that optimizes the yield-per-area of static RAMs (SRAMs). Starting with physical layouts of cells and the desired memory size, we derive probabilities of failure at a range of abstractions - transistor level, cell level, and system level. We then estimate optimal memory granularity, i.e., the size of memory blocks, as well as the optimal number of spare rows and columns that maximize yield-per-area. In particular, we demonstrate the non-monotonic nature of these tradeoffs and present efficient designs for large SRAMs. Our ongoing research is characterizing several other specific tradeoffs, for SRAMs as well as logic blocks.
Jae Chul Cha, Sandeep Gupta 0001
ICCD2
2008 Data Partitioning and Placement Schemes for Matrix Multiplications on a PIM Architecture
abstract
Data intensive applications require massive data transfers between storage and processing units. VLSI scaling has increased the sizes of dynamic memories as well as speeds and capabilities of processing units to a point where, for many such applications, storage and computational processing capabilities are no longer the main limiting factors. Despite this fact, most current architectures fail to meet the performance requirements for such data intensive applications. In this paper, we describe a PIM architecture that harnesses the benefits of VLSI scaling to accelerate matrix operations that constitute the core of many data-intensive applications. We then present data partitioning and placement schemes that are efficient in terms of the computational complexities and internode communication cost. Such approaches are evaluated and analyzed under various computing environments. We also discuss on how to apply such partitioning and placement schemes to each matrix when chains of matrix operations are given as a task.
Jae Chul Cha, Sandeep Gupta 0001
ISPDC2
2008 On Accelerating Path Delay Fault Simulation of Long Test Sequences
abstract
In this paper, we propose an approach to accelerate path delay fault simulation of long test sequences. Several key ideas, namely judicious selection of path delay faults to be simulated, extraction of a compact set of necessary conditions to detect selected faults at primary inputs, and an on-demand selective simulation of input vectors based on their satisfaction of these necessary conditions, are proposed. We demonstrate the benefits of our methodology via experiments on benchmark circuits, with one large test case (S9234) showing a 114X speed-up over a traditional approach.
I-De Huang, Yi-Shing Chang, Suriyaprakash Natarajan, Ramesh Sharma, Sandeep Gupta 0001
ITC5
2008 An Industrial Case Study of Sticky Path-Delay Faults
abstract
Sticky path-delay faults are path delay faults that are neither robustly nor non-robustly testable, but cannot be proven functionally unsensitizable. Better characterization of delay test quality requires a proper analysis of sticky path-delay faults. Furthermore, careful elimination of sticky path-delay faults contributes significantly to test development productivity and reduction of delay test cost. We present an industrial case study that shows the following, (a) On average, even after designers have removed false paths using automated tools and manual overrides, about 8% of path-delay faults with slack less than 10% of the clock period can be sticky, (b) Our approach, which extends a previously proposed technique, identifies a large subset of sticky path-delay faults that cannot cause functional failures and hence can be eliminated from further consideration. This significantly refines the delay test quality assessment and test development effort, (c) Our approach significantly reprioritizes (reorders) the remaining paths for test generation thereby improving the quality of the target path list.
I-De Huang, Yi-Shing Chang, Sandeep Gupta 0001, Sreejit Chakravarty
VTS3
2008 An Efficient Data-Distribution Mechanism in a Processor-In-Memory (PIM) Architecture Applied to Motion Estimation
abstract
In general, the main purpose of using processor-in-memory (PIM) modules is to dramatically increase the data-level parallelism (DLP) and avoid the limited issue rate of current systems (even when they include SIMD extensions) caused by the limited data bandwidth and functional units. Our approach is to divide the PIM module into hundreds of smaller pieces so that each of these smaller PIMs can execute motion estimation for a group of macro blocks in a parallel fashion. We also design the logic in each PIM to execute in a highly pipelined fashion so that even more parallelism can be exploited. The main contribution of this paper is the presentation of architectural techniques that can be used in the PIM module to overcome the addressing and data sharing overhead when these smaller PIMs are used. Our architectural techniques have been applied to motion estimation. Indeed, it has been reported that motion estimation takes the majority of the execution time of MPEG encoding and it has been researched by many because of its importance in MPEG encoding. With our paradigm and techniques, the host processor can be relieved from the most computationally demanding and data-intensive portions of the workload, which should therefore yield a significant performance gain. Indeed, we observed (when 512 of these smaller PIMs were used) a reduction in the number of memory accesses by a factor of up to 2,034 times. At the same time, the performance improved by a multiplicative factor as high as 439 times.
Jung-Yup Kang, Sandeep Gupta 0001, Jean-Luc Gaudiot
IEEE Trans. Computers2
2007 On Generating Vectors That Invoke High Circuit Delays - Delay Testing and Dynamic Timing Analysis
abstract
In this paper, we propose an approach to generate vectors that invoke high delays. We first identify properties of different types of paths, especially sticky paths, i.e., paths that are functionally sensitizable but not even non-robustly testable. In particular, we show that it is impossible to guarantee detection of sticky path-delay faults. We then identify logic and timing conditions that are necessary to cover a target path and develop a new logic-and-timing implication procedure to exploit these conditions. We incorporate this procedure in a new ATPG that also prioritizes the order in which these conditions are used to generate high quality vectors. We use this ATPG to identify paths that cannot or need not be tested and to generate high quality vectors for all other paths. Experimental results demonstrate that the vectors we generate invoke much higher delays than previously generated vector sets, especially for circuits with many sticky paths.
I-De Huang, Sandeep Gupta 0001
ATS2
2007 Improving Timing-Independent Testing of Crosstalk Using Realistic Assumptions on Delay Faults
abstract
Test generation methodology previously developed for crosstalk targets in the presence of manufacturing defects and process variations results in low coverage. In this paper, under a realistic assumption about the nature of manufacturing defects, we show that by incorporating two new concepts, namely, non- criticality and delay-superiority, significantly higher coverage of targets and lower test generation and test application costs are achieved.
Shahdad Irajpour, Sandeep Gupta 0001, Melvin A. Breuer
ATS2
2007 Performance Analysis of Wireless MAC Protocols Using a Search Based Framework
abstract
Previously, we have developed a framework to perform systematic analysis of CSMA/CA based wireless MAC protocols. The framework first identifies protocol states that meet our study objective of minimizing a given performance metric. It then applies search techniques and heuristics to construct sequences of protocol events in a given topology that satisfy our objective. In this paper, we demonstrate that our framework can easily be extended to evaluate performance of new protocols by evaluating two completely different variants, namely MAC protocols for (i) quality of service (QoS), and (ii) power control. In each case, we identify previously unknown problems with the protocol. In particular, we generate scenarios where throughput of a lower priority class can be as high as 5 times compared to the throughput of a higher priority class, thus contradicting the basic notion of QoS. Traditional performance evaluation approaches typically evaluate average performance but do not capture the worst cases, nor do they expose the protocol breaking points. Thus this paper demonstrates the usefulness of a systematic approach in evaluating the protocol breaking points.
Shamim Begum, Sandeep Gupta 0001, Ahmed Helmy
MASCOTS2
2006 Test Generation for Weak Resistive Bridges
abstract
An approach for testing weak resistive bridge targets is developed. The approach is based on defining and generating tests for a set of surrogates associated with each target. Either no or limited timing information is used during test generation. Experimental results show much higher coverage of targets and much lower complexity compared to those for crosstalk targets
Shahdad Irajpour, Sandeep Gupta 0001, Melvin A. Breuer
ATS2
2006 Diagnosis of delay faults due to resistive bridges, delay variations and defects
abstract
In this paper, we present the first diagnosis algorithm for combinational circuit blocks that considers all combinations of multiple gate/wire delay variations/defects and any single resistive bridge. One key component of the proposed algorithm is a new path-oriented effect-cause procedure to identify all possible suspects of above type that might have caused the timing errors observed during test. The second key component is an efficient data structure to represent the suspects. The third key component is a new algorithm to analyze passing tests to vindicate some of the suspects identified. This algorithm exploits the newly identified concept of fixed-but-unknown delays, i.e., the fact that during the period of diagnosis the values of all delay parameters for every gate and wire in the particular circuit under test remain fixed, although at values unknown to us. The final set of suspects reported by the algorithm is guaranteed to contain all possible causes of the observed timing errors. Experimental results on benchmark circuits show the effectiveness of the proposed approach
Sandeep Gupta 0001, Melvin A. Breuer
ATS2
2006 A theory of Error-Rate Testing
abstract
We have entered an era where chip yields are decreasing with scaling. A new concept called intelligible testing has been previously proposed with the goal of reversing this trend for classes of systems which do not require completely error-free operation. Such error tolerant applications include audio, speech, graphics, video, and digital communications. Analyses of such applications have identified error rate as one of the key metrics for error severity, where error rate is defined as the percentage of vectors for which the value at outputs deviates from the corresponding error-free value. In error-rate testing, every fault with an error rate less than a threshold specified by the application is called an acceptable fault; all other faults are called unacceptable. The objective of error-rate testing is to detect every unacceptable fault while detecting none of the acceptable faults. In this paper we develop a theory of error-rate testing. First we study fanout-free circuits with primitive gates and identify new relationships between error rates and fault equivalence and dominance, develop a new test generation procedure, and prove that in such circuits it is possible to detect every unacceptable fault without detecting any acceptable fault. We then analyze more general circuits, including those containing complex gates and fanouts, and show that the above result may not hold for such circuits. We then use a modified version of a classical test generator and a classical fault simulator to obtain empirical data that show that even in arbitrary circuits, it is possible to detect every unacceptable fault while detecting only a fraction of acceptable faults.
Shideh Shahidi, Sandeep Gupta 0001
ICCD2
2006 Low-Cost Scan-Based Delay Testing of Latch-Based Circuits with Time Borrowing
abstract
Classical test approaches typically provide abysmally low path delay fault coverage for high-speed latch-based circuits where time borrowing may occur. Furthermore, none of the classical design-for-testability (DFT) approaches can be used to improve coverage. In [Chung, 2003] we proposed the first structural testing approach that can provide high robust path delay fault coverage for such circuits. However, that approach suffered from high DFT overheads since it required a fully-reconfigurable scan circuitry. In this paper we propose an approach that can provide even higher path delay fault coverage for such circuits using dramatically fewer scan configurations. The proposed test generation approach can also provide high path delay fault coverage under any given set of scan chain configurations. We demonstrate the benefits of the proposed approach via extensive experiments.
Kun Young Chung, Sandeep Gupta 0001
VTS2
2006 LT-RTPG: a new test-per-scan BIST TPG for low switching activity
abstract
A new built-in self-test (BIST) test pattern generator (TPG) design, called low-transition random TPG (LT-RTPG), is presented. An LT-RTPG is composed of a linear feedback shift register (LFSR), a /spl kappa/-input AND gate, and a T flip-flop. When used to generate test patterns for test-per-scan BIST, it decreases the number of transitions that occur during scan shifting and, hence, decreases switching activity during testing. Various properties of LT-RTPGs are identified and a methodology for their design is presented. Experimental results demonstrate that LT-RTPGs designed using the proposed methodology decrease switching activity during BIST by significant amounts while providing high fault coverage.
Seongmoon Wang, Sandeep Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 Selection of Paths for Delay Testing
abstract
In this paper, we propose a new approach to efficiently identify paths for delay testing. We use a realistic delay model and several new concepts (timing threshold, settling times [14], and timing blocking line) and algorithms, to identify a set of paths that is guaranteed to include all paths that may potentially cause a timing error if the accumulated values of additional delays along circuit paths is upper bounded by a desired limit, ... The first phase of the proposed approach identifies a small subset of all possible paths in the circuit for further analysis. Since this phase only requires breadth-first static timing analysis (forward and backward), its complexity is independent of the number of paths in the circuit as well as the number of all possible two-vector sequences that may be applied to the circuit. We then use new conditions for functional sensitization that help identify paths that may be functionally sensitizable and have the potential of causing timing errors if accumulated values of additional delays along any path is upper bounded by ... The results show that without any search, the proposed approach identifies a near minimal number of paths at low complexity.
I-De Huang, Sandeep Gupta 0001
Asian Test Symposium2
2005 Threshold testing: Covering bridging and other realistic faults
abstract
In the recent years, yields for digital VLSI chips have been declining and the decline is expected to accelerate. We have recently proposed a new testing approach called threshold testing, with the goal of providing acceptable yields in future processes for a wide range of high performance digital applications, including audio, speech, video, graphics, visualization, games, and wireless communication. The motivation of this paper is to answer the following question: Do threshold tests generated for stuck-at faults provide as high a threshold coverage for realistic faults as the classical coverage for realistic faults provided by classical stuck-at test sets? Using a combination of analysis and experiments, we show that the stuck-at fault model is indeed a suitable model for threshold testing. This opens the way for developing low cost tools for threshold testing that will provide high threshold coverage for realistic faults, and hence help provide higher yields in future processes at low costs. We also present a threshold automatic test pattern generator (ATPG) for bridging faults.
Sandeep Gupta 0001
Asian Test Symposium2
2005 A Methodology to Compute Bounds on Crosstalk Effects in Arbitrary Interconnects
abstract
In this paper, we present a methodology that uses the moments of a generic crosstalk pulse signal to derive upper bounds on the amplitude of crosstalk pulse in arbitrary interconnects. We apply the proposed methodology to identify vectors that invoke crosstalk pulses with severities less than thresholds at which circuit may malfunction. These vectors are then excluded from the test set, reducing test application time. Case studies that consider interconnects with rich topologies, including cases where we consider process variations or treat lengths of nets as variables, clearly demonstrate the effectiveness of the methodology.
Wichian Sirisaengtaksin, Sandeep Gupta 0001
Asian Test Symposium2
2005 TCP vs. TCP: a systematic study of adverse impact of short-lived TCP flows on long-lived TCP flows
abstract
This paper describes systematical development of TCP adversarial scenarios where we use short-lived TCP flows to adversely influence long-lived TCP flows. Our scenarios are interesting since, (a) they point out the increased vulnerabilities of recently proposed scheduling, AQM and routing techniques that further favor short-lived TCP flows and (b) they are more difficult to detect when intentionally found to target long-lived TCP flows. We systematically exploit the ability of TCP flows in slow-start to rapidly capture greater proportion of bandwidth compared to long-lived TCP flows in congestion avoidance phase, to a point where they drive long-lived TCP flows into timeout. We use simulations, analysis and experiments to systematically study the dependence of the severity of impact on long-lived TCP flows on key parameters of short-lived TCP flows-including their locations, durations and numbers, as well as the intervals between consecutive flows. We derive characteristics of pattern of short-lived flows that exhibit extreme adverse impact on long-lived TCP flows. Counter to common beliefs, we show that targeting bottleneck links does not always cause maximal performance degradation for the long-lived flows. In particular, our approach illustrates the interactions between TCP flows and multiple bottleneck links and their sensitivities to correlated losses in the absence of 'non-TCP friendly' flows and paves the way for a systematic synthesis of worst-case congestion scenarios. While randomly generated sequences of short-lived TCP flows may provide some reductions (up to 10%) in the throughput of the long-lived flows, the scenarios we generate cause much greater reductions (>85%) for several TCP variants and for different packet drop policies (DropTail, RED).
Shirin Ebrahimi-Taghizadeh, Ahmed Helmy, Sandeep Gupta 0001
INFOCOM3
2005 Multiple tests for each gate delay fault: higher coverage and lower test application cost
abstract
Different tests for a single gate delay fault can detect different ranges of delay fault sizes. It is of interests to determine whether, for most faults, a single test covers all the ranges of delay fault sizes covered collectively by all tests for the fault. Using an enhanced gate delay fault simulation algorithm, we show that for a considerable number of gate delay faults in benchmark circuits, multiple tests collectively provide more comprehensive coverage than any single test. We then present many key implications of this observation, especially in the areas of test generation and test set compaction
Shahdad Irajpour, Sandeep Gupta 0001, Melvin A. Breuer
ITC2
2004 Efficient Identification of Crosstalk Induced Slowdown Targets
abstract
This paper deals with filter development in XIDEN, a "pruning " tool used to identify crosstalk targets that can potentially create Boolean errors. XIDEN employs multiple tools to adoptively estimate and/or extract electrical parameters required by its filters, and uses a novel approach to construct an efficient sequence of extractors and filters that are applied to a circuit. Thus, an initially enormous collection of targets can usually be reduced to a very small set of targets via a vectorless process. This process flow is much more efficient than using ATPG without pruning to identify targets that represent faults. The XIDEN framework, including the filters that capture the effect of crosstalk-induced pulses, has been previously presented. In this paper, our focus is on filters associated with crosstalk induced slowdown targets. To accurately compute timing information associated with signal transitions, we have enhanced the XIDEN framework with enhanced static timing analysis procedures that take into consideration single and multiple capacitive crosstalk couplings in a circuit.
Melvin A. Breuer, Sandeep Gupta 0001, Shahin Nazarian
Asian Test Symposium2
2004 Modeling and Testing Crosstalk Faults in Inter-Core Interconnects that Include Tri-State and Bi-Directional Nets
abstract
In this paper, we carry out a systematic simulation study of crosstalk effects in inter-core interconnects that include tri-state and bi-directional nets and have arbitrary topologies. These studies allow us to develop a new crosstalk fault model for such interconnects. We also develop a framework to generate compact tests for such interconnects. We present experimental results that show that the proposed approach provides high fault coverage using a small number of tests.
Wichian Sirisaengtaksin, Sandeep Gupta 0001
Asian Test Symposium2
2004 Modeling and Simulation for Crosstalk Aggravated by Weak-Bridge Defects between On-Chip Interconnects
abstract
This paper presents comprehensive analytic models that consider both the case of a weak bridge, and the combination of a weak bridge and crosstalk between two interconnects. Our models capture the induced signal delay and pulse as a function of the parameters of the circuit and input signals. Our results are compared with HSPICE and shown to be accurate. A simulator is developed that implements our models and accurately captures timing (delay) characteristics of a circuit. We contrast our results with others, and show the benefits of this new model as well as the ability to predict the range of resistance that leads to delay errors.
Sandeep Gupta 0001, Melvin A. Breuer
Asian Test Symposium2
2004 Timing-Independent Testing of Crosstalk in the Presence of Delay Producing Defects Using Surrogate Fault Models
abstract
All previous approaches for generating tests for crosstalk slow-downs are timing-dependent, i.e., they use the nominal values of gate and wire delays while generating tests. None of these methodologies can be used when a crosstalk slow-down must be considered in the presence of process variations and delay producing defects, since such variations and defects change the delay values from the nominal. We present the first timing-independent approach to generate tests for crosstalk slow-downs. The framework is based upon defining a set of surrogates for each crosstalk slow-down target and generating a test for each surrogate. The timing-independent conditions that a test for each surrogate must satisfy are presented. Under the pin-to-pin delay model, we prove that a set of two-vector sequences that covers every surrogate for a crosstalk slow-down target is guaranteed to detect the target, even in the presence of arbitrary delay variations and delay producing defects. We present a method to identify the crosstalk targets and surrogates for which tests must be generated as well as a test generator that we have implemented. We present extensive experimental results using combinational parts of ISCAS 89 circuits.
Shahdad Irajpour, Sandeep Gupta 0001, Melvin A. Breuer
ITC2
2004 Designing Reconfigurable Multiple Scan Chains for Systems-on-Chip
abstract
We propose a framework for designing reconfigurable multiple scan chains for systems-on-chip to minimize test application time. Multiple scan chain design problem defined in this paper involves (1) designing a suitable reconfigurable scan chain architecture, (2) partitioning wrapper cells and core internal scan registers into multiple scan chains, (3) ordering the wrapper cells and the internal scan registers within each scan chain, and (4) identifying sessions in which to activate bypass control signals. We demonstrate significant reductions in test application times and hardware costs over prior heuristics.
Md. Saffat Quasem, Sandeep Gupta 0001
VTS2
2004 A framework for systematic evaluation of multicast congestion control protocols
abstract
Congestion control is a major requirement for multicast to be deployed in the current Internet. Due to the complexity and conflicting tradeoffs, the design and testing of multicast congestion control protocols is difficult. In this paper, we present a novel framework for systematic testing of multicast congestion control protocols. In our framework, we first design an appropriate model for the studied protocols based on the protocols specifications and correctness conditions, and then we develop an automated search engine to generate all possible error scenarios and filter these errors to come up with a selected set of scenarios that we evaluate in more detailed simulations. Our methodology helps in identifying the potential problems of the studied protocols and points to possible improvements. We hope that this will provide a valuable tool to expedite the development and standardization of such protocols.
Karim Seada, Ahmed Helmy, Sandeep Gupta 0001
IEEE J. Sel. Areas Commun.3
2004 The STRESS method for boundary-point performance analysis of end-to-end multicast timer-suppression mechanisms
abstract
The advent of multicast and the growth and complexity of the Internet has complicated network protocol design and evaluation. Evaluation of Internet protocols usually uses random scenarios or scenarios based on designers' intuition. Such an approach may be useful for average case analysis but does not cover boundary-point (worst or best case) scenarios. To synthesize boundary-point scenarios, a more systematic approach is needed. In this paper, we present a method for automatic synthesis of worst and best case scenarios for protocol boundary-point evaluation. Our method uses a fault-oriented test generation (FOTG) algorithm for searching the protocol and system state space to synthesize these scenarios. The algorithm is based on a global finite state machine (FSM) model. We extend the algorithm with timing semantics to handle end-to-end delays and address performance criteria. We introduce the notion of a virtual LAN to represent delays of the underlying multicast distribution tree. Our algorithms utilize implicit backward search using branch and bound techniques and start from given target events. As a case study, we use our method to evaluate variants of the timer suppression mechanism, used in various multicast protocols, with respect to two performance criteria: overhead of response messages and response time. Simulation results for reliable multicast protocols show that our method provides a scalable way for synthesizing worst case scenarios automatically. Results obtained using stress scenarios differ dramatically from those obtained through average case analyses. We hope for our method to serve as a model for applying systematic evaluation to other multicast protocols.
Ahmed Helmy, Sandeep Gupta 0001, Deborah Estrin
IEEE/ACM Trans. Netw.2
2003 An Efficient PIM (Processor-In-Memory) Architecture for Motion Estimation
abstract
Motion estimation is the most time consuming stage of MPEG family encodings and it reportedly absorbs up to 90% of the total execution time of MPEG processing. Therefore, we propose a hardware/software co-design paradigm that uses a PIM module to efficiently execute motion estimation operations. We use a PIM module to reduce the memory access penalty caused by a large number of memory accesses. We segment the PIM module into small pieces so that each smaller PIM module can execute the operations in parallel fashion. However, in order to execute the operations in parallel, there are critical overheads that involve replicating a huge amount of data to many of these smaller PIM modules. Not only do these replications require a huge amount of additional memory accesses but also calculations when generating addresses. Therefore, we also present an efficient data distribution mechanism to effectively support parallel executions among these smaller PIM modules. With our paradigm, the host processor can be relieved from computationally-intensive and data-intensive workloads of motion estimation. We observed up to 2034/spl times/ improvement in reduction of the number of memory accesses and up to 439/spl times/ performance improvement for the execution of motion estimation operations when using our computing paradigm.
Jung-Yup Kang, Sandeep Gupta 0001, Saurabh Shah, Jean-Luc Gaudiot
ASAP2
2003 A Test Generation Approach for Systems-on-Chip that Use Intellectual Property Cores
abstract
In this paper, we propose a hierarchical automatic test pattern generation (ATPG) framework that can generate custom tests for full-scan systems-on-chip (SOCs) containing intellectual property (IP) cores without revealing much IP. The proposed ATPG is shown to be correct and complete and its average and worst case complexities are shown to be comparable with those of classical ATPG. The proposed ATPG reduces DFT overheads and test application costs. It also enables utilization of a range of test methodologies at the SOC level.
Sandeep Gupta 0001
Asian Test Symposium2
2003 Designing Multiple Scan Chains for Systems-on-Chip
abstract
We propose a branch-and-bound framework for designing non-reconfigurable multiple scan chains for systems-on-chip to minimize test application time. Multiple scan chain design problem defined in this paper involves (1) partitioning wrapper cells and core internal scan registers into multiple scan chains, and (2) ordering the wrapper cells and the registers in each scan chain. We design multiple scan chains with test application times within a few percentage of the corresponding optimal in practical run-times. We also demonstrate significant improvements in test application times over prior heuristics for designing multiple scan chains.
Md. Saffat Quasem, Sandeep Gupta 0001
Asian Test Symposium2
2003 An Enhanced Test Generator for Capacitance Induced Crosstalk Delay Faults
abstract
Capacitive crosstalk can give rise to slowdown of signals that can propagate to a circuit output and create a functional error A test generation methodology, called XGEN, was developed to generate tests for such failures. Two drawbacks of XGEN are: (i) it is not complete because of restricted propagation conditions, and (ii) a constrained logic value system is used. In this paper, we relax the propagation conditions to increase the solution space. This increases the likelihood of finding a test. We also present a nine-valued algebra that distinguishes between hazardous values and non-hazardous values. Finally, we use the relation between arrival time and required time ranges to selectively turn off the timing computation procedure which is computationally expensive. Other drawbacks of previous versions of XGEN are: (i) a simplified pin-to-pin delay model was used, and (ii) crosstalk computation could not handle timing ranges. We have addressed both of those issues.
Arani Sinha, Sandeep Gupta 0001, Melvin A. Breuer
Asian Test Symposium2
2003 Structural Delay Testing of Latch-based High-speed Pipelines with Time Borrowing
abstract
High-speed circuits use latch-based pipelines in some of their most delay-criticalparts. The use of latches not only allows attainment of high clock rate but also enables attainment of high yield at desired clock rate by permitting unintentional time borrowing. In this paper, we first demonstrate that none of the existing design-for-testability (OFT) techniques can be used to simplijj delay testing of such circuits. We then demonstrate that this leads to very high test generation and test application times. In many circuits, very low path delay fault coverage is obtained. We then propose a systematic test approach and associated DFT that significantly reduces the test generation and test application costs, and, for many circuits, significantly increases path delay fault coverage.
Kun Young Chung, Sandeep Gupta 0001
ITC2
2003 Path-Delay Fault Simulation for Circuits with Large Numbers of Paths for Very Large Test Sets
abstract
We propose an exact non-enumerative path-delay fault simulation technique for combinational circuits using very large test sets. We focus on combinational circuits that contain very large numbers of path delay faults, for example, c6288 that has about 2/spl times/10/sup 20/ possible path-delay faults. Enumerative fault simulators simply cannot handle such circuits. Existing non-enumerative fault simulators work for small test sets only. The proposed method uses the mathematical principle of inclusion and exclusion and handles the large number of tests using a novel encoding technique.
Nabil M. Abdulrazzaq, Sandeep Gupta 0001
VTS2
2003 Generating Complete and Optimal March Tests for Linked Faults in Memories
abstract
We show that no published march test detects all march-test detectable instances of linked faults in memories. We present necessary and sufficient conditions for detection of single cell linked faults. We identify the set of faults that are undetectable by march tests. We also present sets of faults that dominate all march-test detectable instances of linked multiple cell faults along with the necessary and sufficient conditions for their detection. Using a test generator that takes these conditions as input, we generate the first march tests that detect all march-test detectable linked faults. By considering the subsets of linked faults targeted by the well-known March A and March B tests, we also prove that these well-known tests are optimal for the corresponding sets of target faults.
Sultan M. Al-Harbi, Sandeep Gupta 0001
VTS2
2003 Test Generation for Maximizing Ground Bounce Considering Circuit Delay
abstract
In this paper, we focus on the aspect of ground bounce due to the combination of current produced by gates (signals) switching and the flow of this current through pin electronics. We present a branch-and-bound test generation procedure to obtain high quality 2-vector tests that produce a large amount of ground bounce. We present a framework that accurately captures the relationship between a test and the associated relative size of the maximum amount of ground bounce while taking into account gate delay. Experimental results show that our search procedure can efficiently and effectively find a test that produces the maximum value of ground bounce. We also discuss a binary search based approach that allows our search to cover a larger portion of the search space and find a good test in a reduced amount of CPU time.
Yi-Shing Chang, Sandeep Gupta 0001, Melvin A. Breuer
VTS2
2003 Analyzing Crosstalk in the Presence of Weak Bridge Defects
abstract
An extensive simulation study of various combinations of resistive bridges and crosstalk has been performed and several notable properties that have significant implications for test development have been discovered. Scenarios have been identified where a combination of a bridge at one site and a crosstalk at a separate site in its transitive fanout (or vice versa) can cause slowdown/speed-up whose magnitude significantly exceeds the sum of the slow-down/speed-up, caused by each effect in isolation. It has also been identified that a test vector generated for crosstalk may in fact be invalidated due to the presence of a weak bridge at the crosstalk site. The properties discovered, provide the motivation for a more analytical study that will eventually lead to the proposed framework for test development.
Shahdad Irajpour, Shahin Nazarian, Sandeep Gupta 0001, Melvin A. Breuer
VTS4
2002 Enhanced Crosstalk Fault Model and Methodology to Generate Tests for Arbitrary Inter-core Interconnect Topology
abstract
In this paper we develop a new fault model for capacitive crosstalk in inter-core interconnects. We also develop a framework to generate compact tests for interconnects with arbitrary topologies. Experimental results show that the proposed approach can significantly reduce test application time for large interconnects. We are in the process of extending the framework to interconnects that include tri-state as well as bi-directional nets.
Wichian Sirisaengtaksin, Sandeep Gupta 0001
Asian Test Symposium2
2002 Accurate and Efficient Static Timing Analysis with Crosstalk
abstract
We have developed an accurate and efficient methodology to perform static timing analysis (STA) in combinational logic blocks in the presence of multiple crosstalk-induced noise effects. The crosstalk model used is more accurate because it considers skew, input transition times, and driver strengths. This crosstalk model is enhanced to handle timing ranges for performing STA. The methodology also uses more accurate delay models for gates. The presence of one or more coupling capacitances can create cyclic timing dependencies, even in an otherwise acyclic circuit. We have developed an approach to partition the circuit into minimal timing-iterative subcircuits (TISs) that encapsulate the cyclic timing dependencies. When used in conjunction with our levelization procedure, iterative timing analysis is confined within individual TISs. We have demonstrated that the maximum arrival time values computed by the proposed STA using integrated delay models are much closer to detailed circuit simulation results than an STA that uses the 3C/sub c/ delay model.
I-De Huang, Sandeep Gupta 0001, Melvin A. Breuer
ICCD2
2002 An ATPG for Threshold Testing: Obtaining Acceptable Yield in Future Processes
abstract
When VLSI scaling reaches closer to the limits of laws of physics and to the limits of fabrication processes, yields will decrease, especially at desired speed. However, for a large class of applications, chips need not be perfect to be acceptable. In this paper, we describe the notion of threshold testing that can help improve effective yield for future processes. We then develop an ATPG and demonstrate that significant increase in effective yield can be attained at negligible increase in test application cost.
Sandeep Gupta 0001
ITC2
2002 XIDEN: Crosstalk Target Identification Framework
abstract
An efficient crosstalk target identification framework called XIDEN has been developed that is used prior to the computationally expensive processes of crosstalk validation and test generation. XIDEN is mainly composed of a set of extractors and filters that together identify the prime crosstalk targets. These prime targets include all error producing targets, i.e. targets that can potentially create Boolean errors. A methodology has been developed to determine the sequence of extractors and filters to identify a small set of targets with low computational cost. The effects of process variation and extraction accuracy as well as complexity are considered in XIDEN. After performing a training process on sample circuits, XIDEN produces a set of effective extractor filter sequences to be used for production circuits.
Shahin Nazarian, Hang Huang, Suriyaprakash Natarajan, Sandeep Gupta 0001, Melvin A. Breuer
ITC4
2002 Test Generation for Crosstalk-Induced Faults: Framework and Computational Results
Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer
J. Electron. Test.2
2002 TA-PSV - Timing Analysis for Partially Specified Vectors
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
J. Electron. Test.2
2002 Analytical models for crosstalk excitation and propagation in VLSI circuits
abstract
The authors develop a general methodology to analyze crosstalk effects that are likely to cause errors in deep submicron high-speed circuits. They focus on crosstalk due to capacitive coupling between a pair of lines. Closed form equations are derived that quantify the severity of these effects and describe qualitatively the dependence of these effects on the values of circuit parameters, the rise/fall times of the input transitions, and the skew between the transitions. For noise propagation, they present a new way for predicting the output waveform produced by an inverter due to a nonsquare wave pulse at its input. To expedite the computation of the response of a logic gate to an input pulse, the authors have developed a novel way of modeling such gates by an equivalent inverter. The results of their analysis provide conditions that must be satisfied by a sequence of vectors used for validation of designs as well as post-manufacturing testing of devices in the presence of significant crosstalk. They present data to demonstrate accuracy of their results, including example runs of a test generator that uses these results.
Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 DS-LFSR: a BIST TPG for low switching activity
abstract
A test pattern generator (TPG) for built-in self-test (BIST), which can reduce switching activity during test application, is proposed. The proposed TPG, called dual-speed LFSR (DS-LFSR), consists of two linear feedback shift registers (LFSRs), a slow LFSR and a normal-speed LFSR. The slow LFSR is driven by a slow clock whose speed is 1/dth that of the normal clock, which drives the normal-speed LFSR. The use of DS-LFSR reduces the frequency of transitions at the circuit inputs driven by the slow LFSR, leading to a reduction in switching activity during test application. A procedure is presented to design a DS-LFSR so as to achieve high fault coverage by ensuring that patterns generated by it are unique and uniformly distributed. A new gain function and a method to compute its value for each circuit input are proposed to select inputs to be driven by the slow LFSR. Also, a procedure to increase the number of inputs driven by the slow LFSR by combining compatible inputs is presented to further decrease the switching activity. Finally, DS-LFSRs are designed for the ISCAS85 and ISCAS89 benchmark circuits and shown to provide a 13% to 70% reduction in the numbers of load-capacitance weighted transitions with no loss of fault coverage (for stuck-at as well as transition delay faults) and at very slight area overheads.
Seongmoon Wang, Sandeep Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 An automatic test pattern generator for minimizing switching activity during scan testing activity
abstract
An automatic test pattern generation (ATPG) technique is proposed that reduces switching activity during testing of sequential circuits that have full scan. The objective is to permit safe and inexpensive testing of low-power circuits and bare dies that would otherwise require expensive heat removal equipment for testing at high speed. The approach works with standard scan designs that are commonly used and typically have significantly lower overhead than enhanced scan designs. The proposed ATPG exploits all possible "don't cares" that occur during scan shifting, test application, and response capture to minimize switching activity in the circuit under test. An ATPG that minimizes the number of state inputs that are assigned specific binary values has been developed. Don't cares at state inputs are assigned binary values that cause the minimum number of transitions during scan shifting and don't cares at primary inputs during scan shifting and capture are used to block gates that may have transitions during scan shifting. The proposed technique has been implemented and the generated tests are compared with those generated by a simple PODEM implementation for full scan versions of ISCAS89 benchmark circuits.
Seongmoon Wang, Sandeep Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2001 A New Gate Delay Model for Simultaneous Switching and Its Applications
abstract
We present a new model to capture the delay phenomena associ-ated with simultaneous to-controlling transitions. The proposed delay model accurately captures the effect of the targeted delay phe-nomena over a wide range of transition times and skews. It also cap-tures the effects of more variables than table lookup methods can handle. The model helps improve the accuracy of static timing anal-ysis, incremental timing refinement, and timing-based ATPG.
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
DAC2
2001 Exact fault simulation for systems on Silicon that protects each core's intellectual property
abstract
We present a fault simulation approach for multicore systems on silicon (SOC) (a) that provides exact fault coverage for the entire SOC, (b) does so without revealing any intellectual property (IP) of core vendors, and (c) whose run time is comparable to that required by the existing approaches that require all IP to be revealed. This fault simulator assumes a full scan SOC design and is first in a suite of simulation, test generation, and DFT tools that are currently under development. The proposed approach allows flexibility in selection of a test methodology for SOC, reduces test application cost and area and performance overheads, and allows more comprehensive testing.
Md. Saffat Quasem, Sandeep Gupta 0001
DATE2
2001 Crosstalk test generation on pseudo industrial circuits: a case study
abstract
In this paper, we present data that validates the viability of a university prototype crosstalk ATPG system, XGEN, on real designs. We remodeled Intel circuits and performed test generation using actual parasitic data. A crosstalk ATPG implementation flow was developed based on Intel tools. Validation results are shown for the modified circuits. Critical issues for preserving accurate timing information and capturing crosstalk effects are discussed.
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
ITC3
2001 Switch-level delay test of domino logic circuits
abstract
We address the testing of delay faults in domino circuits that contain complex gates. The different ways in which these faults can cause errors are demonstrated. We identify structures in both the evaluate and the precharge logic that should be tested for delay faults. We propose conditions to generate delay tests for them, and outline extensions to handle mixed static-domino circuits. Testability results are reported for benchmark circuits that are mapped to domino gates and in-house domino circuits.
Suriyaprakash Natarajan, Sandeep Gupta 0001, Melvin A. Breuer
ITC2
2001 An Efficient Methodology for Generating Optimal and Uniform March Tests
abstract
A large number of march tests that provide different fault coverages have been published and a few methodologies have been presented for automatically generating march tests. This paper presents a new methodology for generating optimal and uniform march tests. The new methodology uses a compact representation of faults, generates necessary and sufficient conditions for their detection, and generates tests using the conditions along with the properties of march tests. The methodology is demonstrated as being more efficient than those previously presented. It has been used to (a) generate new optimal tests that are uniform, which are desired to simplify BIST architecture, (b) prove the optimally of some well-known tests such as March C-, and (c) generate a complete set of optimal march tests for different combinations of faults. The proposed approach hence provides memory manufacturers with an optimal test to cover the types of faults that are likely to occur in their memories.
Sultan M. Al-Harbi, Sandeep Gupta 0001
VTS2
2001 Test Generation for Maximizing Ground Bounce for Internal Circuitry with Reconvergent Fan-out
abstract
Due to technology scaling and increasing clock rate, problems due to noise effects lead to an increase in design and test efforts and a decrease in circuit performance. This paper addresses the problem of efficiently and effectively generating two vector tests to produce the maximum ground bounce in a circuit. We have developed a branch and bound procedure that can find a good quality test for maximum ground bounce in a rather short time. Comparison of results with SPICE simulations confirms the quality of tests obtained by our procedure.
Yi-Shing Chang, Sandeep Gupta 0001, Melvin A. Breuer
VTS2
2001 Path delay fault diagnosis in combinational circuits with implicitfault enumeration
abstract
A new methodology involving effect-cause analysis has been demonstrated for the diagnosis of path delay faults. The paper illustrates a structural representation, called the suspect circuit, of all the possible path delay faults in a faulty circuit. This representation has been used to design efficient algorithms that enable us to manipulate the suspect faults without having to enumerate them explicitly. Procedures for removing fault-free paths from the list of suspect faults have been implemented to improve the diagnostic resolution. Moreover, efficient data structures are used to complement the procedures and reduce the memory footprint of the algorithms. Results indicate that the diagnostic resolution obtained is very high and includes all possible causes of the observed delay faults.
Pankaj Pant, Yuan-Chieh Hsu, Sandeep Gupta 0001, Abhijit Chatterjee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2001 Introducing redundant computations in RTL data paths for reducing BIST resources
abstract
The need for considering BIST requirements during the scheduling and assignment stages of behavioral synthesis has been demonstrated in previous research and techniques for reducing BIST resources of a data path during these stages of synthesis have been developed. However, the degree of freedom that can be exploited during scheduling and assignment to minimize these resources is often limited by the data and control dependencies of a behavior. In this paper, we propose transformation of a behavior before scheduling and assignment, namely introducing redundant computations such that the resulting data path is testable using few BIST resources. The transformation makes use of spare capacity of modules to add redundancy that enables test paths to be shared among the modules. A technique for identifying potential BIST resource sharing problems in a behavior and resolving them by redundant computations is presented. Introduiction of redundant computations is performed without compromising the latency and functional resource requirement of the behavior.
Ishwar Parulkar, Sandeep Gupta 0001, Melvin A. Breuer
ACM Trans. Design Autom. Electr. Syst.2
2000 A new framework for static timing analysis, incremental timing refinement, and timing simulation
abstract
In this paper we present a framework that enables the computation of tight ranges of signal arrival, transition, and required times for rising and falling transitions at each circuit line, given an input sequence consisting of two partially specified vectors. At one extreme, when the vectors are completely unspecified, this framework becomes identical to static timing analysis (STA). At the other extreme, when the vectors are completely specified, this framework performs timing simulation (TS). Our key motivation for developing this framework was to reduce the amount of search required by a test generator that uses timing information. During test generation for a target fault, values are specified incrementally and this framework enables refinement of timing windows. We demonstrate that this approach significantly improves test generation efficiency. In this mode, the ATPG is said to be performing incremental timing refinement (ITR).
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
Asian Test Symposium2
2000 Test generation for crosstalk-induced faults: framework and computational result
abstract
Due to technology scaling and increasing clock frequency, problems due to noise effects are leading to an increase in design/debugging efforts and a decrease in circuit performance. This paper addresses the problem of efficiently and accurately generating two-vector tests for crosstalk-induced effects, such as pulses, signal speedup and slowdown, in digital combinational circuits. We have developed a mixed-signal test generator, called XGEN, that incorporates classical static values as well as dynamic signals such as transitions and pulses, and timing information such as signal arrival times, rise/fall times and gate delay. In this paper, we first discuss the general framework of the test generation algorithm followed by computational results. A comparison of our results with SPICE simulations confirms the accuracy of this approach.
Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer
Asian Test Symposium2
2000 BIST TPG for SRAM cluster interconnect testing at board level
abstract
A Built-In Self-Test (BIST) methodology and a test pattern generation (TPG) architecture for testing static random access memory (SRAM) interconnect at board level via IEEE 1149.1 Boundary Scan (BS) Architecture are presented. Due to the expense and complexity of BS circuitry the widely-used SRAMs on most modern telecommunication circuit boards seldom contain BS architecture. (We call such non-boundary scan ICs cluster-ICs.) Hence, a methodology that tests the large numbers of board-level interconnects at the control, address, and data lines of cluster SRAMs is necessary. This is especially essential for board-level interconnect BIST which is used not only for manufacturing testing but also for system testing after integration. Newly identified prohibited conditions, which enable re-arrangement and merger of tests, are incorporated into test conditions for SRAM cluster interconnects. These improvements have been exploited to develop an efficient test procedure that is suitable for BIST. The proposed BIST methodology generates TPGs that (i) guarantee the avoidance of multi-driver conflicts when testing via BSA, (ii) guarantee the detection of all testable SRAM cluster interconnect faults, (iii) have low area overhead, and (iv) have short test lengths.
Chen-Huan Chiang, Sandeep Gupta 0001
Asian Test Symposium2
2000 Systematic Performance Evaluation of Multipoint Protocols
Ahmed Helmy, Sandeep Gupta 0001, Deborah Estrin, Alberto Cerpa
FORTE2
2000 Systematic testing of multicast routing protocols: analysis of forward and backward search techniques
abstract
We present a new methodology for developing systematic and automatic test generation algorithms for multipoint protocols. These algorithms attempt to synthesize network topologies and sequences of events that stress the protocol's correctness or performance. This problem can be viewed as a domain-specific search problem that suffers from the state space explosion problem. One goal of this work is to circumvent the state space explosion problem utilizing knowledge of network and fault modeling, and multipoint protocols. The two approaches investigated are based on forward and backward search techniques. We use an extended finite state machine (FSM) model of the protocol. The first algorithm uses forward search to perform reduced reachability analysis. Using domain-specific information for multicast routing over LAN, the algorithm complexity is reduced from exponential to polynomial in the number of routers. This approach, however, does not fully automate topology synthesis. The second algorithm, the fault-oriented test generation, uses backward search for topology synthesis and uses backtracking to generate event sequences instead of searching forward from initial states. Using these algorithms, we have conducted studies for correctness of the multicast routing protocol PIM.
Ahmed Helmy, Deborah Estrin, Sandeep Gupta 0001
ICCCN3
2000 Test generation for path-delay faults in one-dimensional iterative logic arrays
abstract
We propose a test generation method for path-delay faults in combinational iterative logic arrays (ILAs). The number of paths as well as the number of critical paths in ILAs can grow exponentially with the number of stages. Existing path-delay test generation techniques explicitly target each selected path and cannot generate tests for ILAs with reasonable numbers of stages, e.g., 16 and 32. The proposed method overcomes this difficulty by implicitly targeting all testable paths and can generate tests for ILAs of arbitrary size and guarantees coverage of all testable faults. The proposed method also drastically decreases the test data volume to be stored in the high-speed memories in the probe unit of the tester by generating tests in the form of a small number of expressions. This is of great benefit since the ability to store large volumes of test data is a significantly greater limiting factor than the time required to apply the tests. Finally, for most ILAs, this method produces a compact set of tests.
Nabil M. Abdulrazzaq, Sandeep Gupta 0001
ITC2
2000 Systematic Testing of Protocol Robustness: Case Studies on Mobile IP and MARS
abstract
Systematic testing of robustness by evaluation of synthesized scenarios STRESS is a methodology developed for the systematic testing of protocols, and includes algorithms for generating topologies and event sequences that rigorously test the correctness or performance of a given protocol. In this paper, we apply the STRESS method to mobile IP (MIP) and the multicast address resolution server (MARS) protocol for supporting IP-multicast over ATM. For each protocol, we develop a protocol model and analyze its robustness. We also analyze complexity of the STRESS test generation algorithms. In the process, we identify the limitations of the existing STRESS models and algorithms, and propose extensions to carry out our case studies. With the aid of STRESS, we were able to identify several protocol behaviors that lead to error or performance degradation. For MIP we identified such behaviors with the crash of a home agent or the loss of a registration message. For MARS undesired behavior was detected with the crash of the MARS server or source and the selective loss of a join or leave message. The complexity of forward search was found to be O(n/sup 2/) for both MIP and MARS. Incorporating the fault model in the search was found to affect the number of states searched. The crash of a home agent in MIP, and the crash of a data source or server in MARS, were both found to increase the number of states searched. However, the asymptotic complexity was not affected.
Shamim Begum, Meeta Sharma, Ahmed Helmy, Sandeep Gupta 0001
LCN4
2000 A Framework to Minimize Test Escape and Yield Loss during IDDQ Testing: A Case Study
abstract
We describe a new framework to minimize test escape (TE) and yield loss (YL) during I/sub DDQ/ testing. The proposed framework defines the concept of critical severity of a fault S/sub k/', that provides a link between the fault magnitude and violation of one or more specifications. The framework provides various strategies to select the values of I/sub DDQ/ threshold and provides a mechanism to compute test escape and/or yield loss. The framework is illustrated using an SRAM as a case study. The results demonstrate various trade-offs that can be explored using the framework.
Hugo Cheung, Sandeep Gupta 0001
VTS2
2000 BIST TPG for Combinational Cluster Interconnect Testing at Board Level
Chen-Huan Chiang, Sandeep Gupta 0001
J. Electron. Test.2
2000 Novel Test Pattern Generators for Pseudoexhaustive Testing
abstract
Pseudoexhaustive testing of a combinational circuit involves applying all possible input patterns to all its individual output cones. The testing ensures detection of all detectable multiple stuck-at faults in the circuit and all detectable combinational faults within individual cones. Test pattern generators based on coding theory principles are not tailored to a specific circuit as they do not utilize any structural information. They usually generate test sets that are several orders of magnitude larger than the minimum size pseudoexhaustive test set required for a specific circuit. In this paper, we describe hardware efficient test pattern generators that employ knowledge of the circuit output cone structures for generating minimal test sets. Using our techniques, we have designed generators that generate minimum size test sets for the ISCAS benchmark circuits.
Rajagopalan Srinivasan, Sandeep Gupta 0001, Melvin A. Breuer
IEEE Trans. Computers2
1999 Validation and test generation for oscillatory noise in VLSI interconnects
abstract
Inductance of on-chip interconnects gives rise to signal overshoots and undershoots that can cause logic errors. By considering technology trends, we show that in 0.13 /spl mu/m technology such noise in local interconnects embedded in combinational logic can exceed the threshold voltage. We show the impact of such noise on different kinds of circuits. The magnitude of this noise can increase due to process variations. We present an algorithm for generating vectors for validation and manufacturing test to detect logic-value errors caused by inductance induced oscillation. To facilitate the vector generation method, we have derived analytical expressions, as functions of rise and fall times for (i) the magnitude of overshoots and undershoots, and (ii) the settling time, i.e., the time required for the circuit response to settle to a bound close to the final value.
Arani Sinha, Sandeep Gupta 0001, Melvin A. Breuer
ICCAD2
1999 Test generation for crosstalk-induced delay in integrated circuits
abstract
Due to technology scaling and increasing clock frequency, problems due to noise effects lead to an increase in design/debugging efforts and a decrease in circuit performance. This paper shows how crosstalk coupling between lines can affect the propagation delay of signals in integrated circuits. A model is presented to evaluate the effect of parasitic coupling crosstalk. Conditions for the creation of the worst-case coupling and propagation of a delayed signal are presented. A test pattern generation algorithm utilizing the above conditions is presented and applied to several example circuits.
Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer
ITC2
1999 Switch-level delay test
abstract
Gate-level models are usually used to generate tests for circuits containing non-primitive CMOS gates. It is shown that tests generated using these models and classical conditions for robust path delay testing can fail to detect delay faults in such circuits. A new delay-independent, switch-level delay test methodology, called /spl tau/-robust testing, is proposed that defines new entities called targets and proposes conditions to generate tests for each target. It is proven that, under the assumed delay model, a circuit that passes a test set containing a /spl tau/-robust test for every target is guaranteed to operate correctly at the desired speed. The effectiveness of the proposed methodology is demonstrated by (a) illustrating the difference between the delays excited by classical robust and /spl tau/-robust tests via circuit simulation, and (b) generation of /spl tau/-robust tests for benchmark circuits and comparison of /spl tau/-robust coverage of classical robust and /spl tau/-robust test sets.
Suriyaprakash Natarajan, Sandeep Gupta 0001, Melvin A. Breuer
ITC2
1999 LT-RTPG: a new test-per-scan BIST TPG for low heat dissipation
abstract
A new BIST TPG design, called low-transition random TPG (LT-RTPG), that is comprised of an LFSR, a k-input AND gate, and a T flip-flop, is presented. When used to generate test patterns for test-per-scan BIST, it decreases the number of transitions that occur during scan shifting and hence decreases the heat dissipated during testing. Various properties of LT-RTPGs are studied and a methodology for their design is presented. Experimental results demonstrate that LT-RTPGs designed using the proposed methodology decrease the heat dissipated during BIST by significant amounts while attaining high fault coverage, especially for circuits with moderate to large number of scan inputs.
Seongmoon Wang, Sandeep Gupta 0001
ITC2
1999 Test Generation for Ground Bounce in Internal Logic Circuitry
abstract
Ground bounce in internal circuitry is becoming an important design validation and test issue. In this paper a new circuit model for ground bounce in internal circuitry is proposed. Based on this model an algorithm for generating test patterns that maximize ground bounce in combinational logic is presented. Our algorithm is also applicable to other test problems such as delay testing in the presence of excessive ground bounce.
Yi-Shing Chang, Sandeep Gupta 0001, Melvin A. Breuer
VTS2
1998 BIST TPG for Combinational Cluster (Glue Logic) Interconnect Testing at Board Level
abstract
Due to the expense and complexity of boundary scan architecture (BSA) circuitry, non-boundary scan ICs (cluster ICs) are still used in modern circuit boards. A novel BIST architecture and a TPG design methodology to program this architecture are presented for testing inter-IC interconnects among combinational cluster ICs via IEEE 1149.1 BSA at board level. The two main issues that are tackled are: (1) safe testing, i.e., avoidance of conflicts at all multi-driver nets; and (2) obtaining complete coverage of all detectable faults in the interconnect. We have proposed a realistic model of interconnects that contain combinational cluster ICs, identified test requirements that must be satisfied by the TPG to achieve the above objectives, and developed a new TPG architecture that satisfies these test requirements. We have theoretically proven that the proposed technique guarantees safe and comprehensive testing of cluster interconnects and also demonstrated this fact via its applications to a few example circuits.
Chen-Huan Chiang, Sandeep Gupta 0001
Asian Test Symposium2
1998 An Automatic Test Pattern Generator for At-Speed Robust Path Delay Testing
abstract
At-speed robust path delay testing is more desirable than the commonly employed slow-fast delay testing due to lower test application time, lower area overhead, and increased possibility of detecting unmodeled faults that cause errors only during at-speed circuit operation. However, it has been observed in practice that at-speed application of tests generated by existing test generators can lead to their invalidation. In this paper, we first present techniques to modify tests generated by existing test generators to avoid invalidation during at-speed testing. We then present a new procedure to generate tests suitable for at-speed delay testing of combinational circuits. Experimental results show that (a) at-speed application of test sets generated by existing generators leads to significant test invalidation, where the degree of invalidation is approximately proportional to the degree of compactness of the test set, and (b) the at-speed robust path delay tests generated by the proposed test generator are significantly shorter than those obtained by modifying the tests generated by existing generators.
Yuan-Chieh Hsu, Sandeep Gupta 0001
Asian Test Symposium2
1998 Introducing Redundant Computations in a Behavior for Reducing BIST Resources
abstract
The degree of freedom that can be exploited during scheduling and assignment to minimize BIST resources is often limited by the data dependencies of a behavior. We propose transformation of a behavior by introducing redundant computations such that the resulting data path requires few BIST resources. The transformation makes use of sparecapacity of modules to add redundancy that enables test paths to be shared among the modules. A technique is presented for introducing redundant computations that reduce the BIST resource requirements of a data path without compromising the latency and functional resource constraints.
Ishwar Parulkar, Sandeep Gupta 0001, Melvin A. Breuer
DAC2
1998 Scheduling and Module Assignment for Reducing Bist Resources
abstract
Built-in self-test (BIST) techniques modify functional hardware to give a data path the capability to test itself. The modification of data path registers into registers (BIST resources) that can generate pseudo-random test patterns and/or compress test responses, incurs an area overhead penalty. We show how scheduling and module assignment in high-level synthesis affect BIST resource requirements of a data path. A scheduling and module assignment procedure is presented that produces schedules which, when used to synthesize data paths, result in a significant reduction in BIST area overhead and hence total area.
Ishwar Parulkar, Sandeep Gupta 0001, Melvin A. Breuer
DATE2
1998 Fault-oriented Test Generation for Multicast Routing Protocol Design
Ahmed Helmy, Deborah Estrin, Sandeep Gupta 0001
FORTE3
1998 Test generation in VLSI circuits for crosstalk noise
abstract
This paper addresses the problem of efficiently and accurately generating two-vector tests for crosstalk induced effects, such as pulses, signal speedup and slowdown, in digital combinational circuits. These effects are becoming more prevalent due to short signal switching times and deep submicron circuitry. These noise effects can propagate through a circuit and create a logic error in a latch or at a primary output. We first present a new way for predicting the output waveform produced by an inverter due to a non-square wave pulse at its input. Our modeling technique captures such properties as the amplitude of a pulse and its rise/fall times and the delay through a device. To expedite the computation of the response of a logic gate to an input pulse, we have developed a novel way of modeling such gates by an equivalent inverter. We have developed a mixed-signal test generator that incorporates classical PODEM-like static values as well as dynamic signals such as transitions and pulses, and timing information such as signal arrival times, rise/fall times, and gate delay. We also present a new analog cost function that is used to guide the search process. Comparison of results with SPICE simulations confirms the accuracy of this approach. This paper focuses primarily on crosstalk induced pulses, but these results have been extended to deal with speedup and slowdown effects.
Sandeep Gupta 0001, Melvin A. Breuer
ITC2
1998 A new path-oriented effect-cause methodology to diagnose delay failures
abstract
A new methodology to diagnose delay failures is described. Key characteristics of the methodology are: (a) path-oriented diagnosis, (b) effect-cause reasoning, and (c) utilization of information obtained from the passing vectors. Two new representations are developed to make manageable the complexity of a path-oriented methodology. The results of diagnosis are: (a) proven to include all possible causes of observed delay errors, and (b) empirically found to have very high resolution.
Yuan-Chieh Hsu, Sandeep Gupta 0001
ITC2
1998 A Methodology for Transforming Memory Tests for In-System Testing of Direct Mapped Cache Tags
abstract
While any efficient test developed for off-line testing of memory chips can be easily adapted for in-system testing of single level memory systems, no efficient methodology is known to transform such a test for in-system testing of multilevel memory systems that have one or more levels of cache. The main challenge is in transforming the known test to test the tags of the cache (testing of the data part of the cache is relatively straightforward). In this paper we present a general methodology to transform march tests for in-system testing of tags of direct mapped caches. The transformation has been used to obtain new versions of March B and March X tests. It is shown that the new versions of tests detect the same sets of faults in the cache tags as their original versions detect in memory chips. Finally, it is demonstrated that the proposed version of March B has significantly lower time complexity than previously proposed tests and can be applied without any modification of the memory system hardware.
Sultan M. Al-Harbi, Sandeep Gupta 0001
VTS2
1998 Allocation Techniques for Reducing BIST Area Overhead of Data Paths
Ishwar Parulkar, Sandeep Gupta 0001, Melvin A. Breuer
J. Electron. Test.2
1998 Estimation of BIST Resources During High-Level Synthesis
Ishwar Parulkar, Sandeep Gupta 0001, Melvin A. Breuer
J. Electron. Test.2
1998 ATPG for Heat Dissipation Minimization During Test Application
abstract
A automatic test pattern generator (ATPG) algorithm is proposed that reduces switching activity (between successive test vectors) during test application. The main objective is to permit safe and inexpensive testing of low power circuits and bare die that might otherwise require expensive heat removal equipment for testing at high speeds, Three new cost functions, namely transition controllability, observability, and test generation costs, have been defined. It has been shown, for a fanout free circuit under test, that the transition test generation cost for a fault is the minimum number of transitions required to test a given stuck-at fault. The proposed algorithm has been implemented and the generated tests are compared with those generated by a standard PODEM implementation for the larger ISCAS85 benchmark circuits. The results clearly demonstrate that the tests generated using the proposed ATPG can decrease the average number of (weighted) transitions between successive test vectors by a factor of 2 to 23.
Seongmoon Wang, Sandeep Gupta 0001
IEEE Trans. Computers2
1998 Bounds on pseudoexhaustive test lengths
abstract
Pseudoexhaustive testing involves applying all possible input patterns to the individual output cones of a combinational circuit. Based on our new algebraic results, we have derived both generic (cone-independent) and circuit-specific (cone-dependent) bounds on the minimal length of a test required so that each cone in a circuit is exhaustively tested. For any circuit with five or fewer outputs, and where each output has k or fewer inputs, we show that the circuit can always be pseudoexhaustively tested with just 2/sup k/ patterns. We derive a tight upper bound on pseudoexhaustive test length for a given circuit by utilizing the knowledge of the structure of the circuit output cones. Since our circuit-specific bound is sensitive to the ordering of the circuit inputs, we show how the bound can be improved by permuting these inputs.
Rajagopalan Srinivasan, Sandeep Gupta 0001, Melvin A. Breuer
IEEE Trans. Very Large Scale Integr. Syst.2
1997 ATPG for Heat Dissipation Minimization During Scan Testing
abstract
An ATPG technique is proposed that reduces heat dissipationduring testing of sequential circuits that have full-scan. The objectiveis to permit safe and inexpensive testing of low power circuitsand bare die that would otherwise require expensive heat removalequipment for testing at high speeds. The proposed ATPG exploitsall don't cares that occur during scan shifting, test application, andresponse capture to minimize switching activity in the circuit undertest. Furthermore, an ATPG that maximizes the number of state inputsthat are assigned don't care values, has been developed. Theproposedtechniquehas beenimplemented and usedto generatetestsfor full scan versions of ISCAS 89 benchmark circuits. These testsdecrease the average number of transitions during test by 19% to89%, when comparedwith those generatedby a simple PODEM implementation.
Seongmoon Wang, Sandeep Gupta 0001
DAC2
1997 BIST TPG for faults in system backplanes
abstract
A built-in self-test (BIST) methodology to test system backplanes by using BIST functionality in each of its constituent boards is presented. Since the configurations of systems change frequently, at the system level, the proposed methodology employs a simple test schedule which can be easily changed whenever the system configuration is changed. Since the boards used in such systems are designed for use in a wide variety of systems, the proposed methodology defines the test objectives to be achieved by a board's BIST circuit in terms of the board's edge pin connections, independent of the configurations of the systems in which the board may be used. It is shown that the combination of the proposed test schedule and the availability, on each board in the system, of any BIST circuit that satisfies the proposed test objectives, guarantees safe testing of faults in backplanes. A programmable test architecture and an algorithm to program the architecture to obtain BIST that satisfies the test objectives is also presented. Finally, the applicability and effectiveness of the methodology is demonstrated via its application to multiple configurations of an example system that uses a VME backplane.
Chen-Huan Chiang, Sandeep Gupta 0001
ICCAD2
1997 Analytic Models for Crosstalk Delay and Pulse Analysis Under Non-Ideal Inputs
abstract
In this paper we develop a general methodology to analyze crosstalk to obtain insight into effects that are likely to cause errors in deep submicron high speed circuits. We focus on crosstalk due to capacitive coupling between a pair of lines. We first consider the case where crosstalk noise manifests as a pulse and characterize the maximum amplitude, width, energy and timing of this pulse. Closed form equations quantifying the dependence of these pulse attributes on the values of circuit parameters and the rise time of the input transition are derived. We also consider how crosstalk causes slowdown (speedup), i.e. increases (decreases) the rise/fall times of signals on coupled lines, when their inputs have transitions in the opposite (same) directions. Expressions relating the slowdown (speedup) to circuit parameters, the rise/fall times of the input transitions, and the skew between the transitions are derived. We show that crosstalk effects can be significantly aggravated by variations in the fabrication process. New design corners are identified for validation of designs that have significant crosstalk effects. Finally, the results of our analysis provide conditions that must be satisfied by a sequence of vectors used for validation of designs as well as post-manufacturing testing of devices in the presence of significant crosstalk.
Melvin A. Breuer, Sandeep Gupta 0001
ITC3
1997 DS-LFSR: A New BIST TPG for Low Heat Dissipation
abstract
A test pattern generator (TPG) for built-in self-test (BIST), which can reduce heat dissipation during test application, is proposed. The proposed TPG, called dual-speed LFSR (DS-LFSR), consists of two linear feedback shift registers (LFSRs), a slow LFSR and a normal-speed LFSR. The slow LFSR is driven by a slow clock whose speed is width that of the normal clock which drives the normal-speed LFSR, The use of DS-LFSR lowers the transition density at the circuit inputs driven by the slow LFSR, leading to a reduction in heat dissipation during test application. A procedure is presented to design a DS-LFSR so as to achieve high fault coverage by ensuring that patterns generated by it are unique and uniformly distributed. A new gain function, and a method to compute its value for each circuit input, is proposed to select inputs to be driven by the slow LFSR. Also, a procedure to increase the number of inputs driven by the slow LFSR by combining compatible inputs is presented to further decrease the heat dissipation, Finally, DS-LFSRs are designed for the ISCAS85 and ISCAS89 benchmark circuits and shown to provide 13% to 70% reduction in the numbers of transitions with no loss of fault coverage and at very slight area overheads.
Seongmoon Wang, Sandeep Gupta 0001
ITC2
1997 Analysis of Ground Bounce in Deep Sub-Micron Circuits
abstract
Ground bounce occurs in integrated circuits and can cause signal distortion and increase gate delay. This can result in improper circuit operation. In the past, the switching of input/output buffers was the primary cause of the ground bounce. In designs employing deep sub-micron technology high operating frequency, and short rise/fall times, ground bounce due to switching in internal circuitry becomes a potential problem. In this paper experiments based on realistic assumptions are performed to explore the properties of ground bounce. Experiments indicate that (1) ground bounce is generated in gates, irrespective of whether outputs switch from 0 to 1 or from 1 to 0, (2) ground bounce is reduced when the load capacitance increases, and (3) ground bounce decreases when the number of gates that switch is held constant while the number of gates that don't switch increases. These conclusions are different from what has been found when input/output buffers switch and lead to new design, verification and test issues.
Yi-Shing Chang, Sandeep Gupta 0001, Melvin A. Breuer
VTS2
1997 High Quality Robust Tests for Path Delay Faults
abstract
Detailed circuit simulations have demonstrated that a classical two-pattern robust test for a path delay fault may not excite the worst case delay of the target path. We have developed a new definition of robust test that maintains the desirable properties of classical robust tests while incorporating two additional considerations, namely side-fan-in transitions and pre-initialization, which are shown to have a significant impact on the delay of the target path. The associated test generation problem was formulated as a constrained optimization problem, and an ATPG system developed to generate three-pattern robust tests that excite the worst case delay of the target path. The ATPG works on a gate level model that is augmented to capture the necessary switch level details. Experimental results show that the quality of robust delay tests varies dramatically and that the proposed high quality robust delay tests are needed for improving test quality.
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
VTS2
1997 BIST TPGs for Faults in Board Level Interconnect via Boundary Scan
abstract
In this paper we present a new BIST test pattern generator architecture and a methodology to program this architecture to generate tests for any given inter-chip interconnect circuitry via IEEE 1149.1 boundary-scan architecture. The test architecture uses two test pattern generators, a C-TPG that generates test patterns for the control cells in the boundary scan chain and a D-TPG that generates rest patterns for the data cells. The other main component of the test architecture is a lookup table which is programmed to select, for each boundary scan cell, a specific C-TPG or D-TPG stage whose content is shifted into that cell. This test architecture provides a complete BIST solution for interconnect testing. The proposed BIST TPG design procedure uses the notions of incompatibility and conditional incompatibility and generates TPG designs that (i) guarantee that no circuit damage can occur due to multi-driver conflicts, (ii) guarantee the detection of all interconnect faults, (iii) have low area overhead, and (iv) have low test length. The proposed procedure is used to obtain TPG designs that require significantly less test time and area than other TPG designs, for eight interconnect circuits extracted from industrial boards.
Chen-Huan Chiang, Sandeep Gupta 0001
VTS2
1996 A Satisfiability-Based Test Generator for Path Delay Faults in Combinational Circuts
abstract
This paper describes a new formulation to generate robust tests for path delay faults in combinational circuits based on Boolean satisfiability. Conditions to detect a target path delay fault are represented by a Boolean formula. Unlike the technique described in [16], which extracts the formula for each path delay fault, the proposed formulation needs to extract the formula only once for each circuit cone. Experimental results show tremendous time saving on formula extraction compared to other satisfiability-based ATPG algorithms. This also leads to low test generation time, especially for circuits that have many paths but few outputs. The proposed for-mulation has also been modied to generate other types of tests for path delay faults.
Chih-Ang Chen, Sandeep Gupta 0001
DAC2
1996 Lower Bounds on Test Resources for Scheduled Data Flow Graphs
abstract
Abstract|Lower bound estimations of resources at various stages of high-level synthesis are essential to guide synthesis algorithms towards optimal solutions.In this paper we present l o w er bounds on the number of test resources (i.e.test pattern generators, signature analyzers and CBILBO registers) required to test a synthesized data path using built-in self-test (BIST).The estimations are performed on scheduled data ow graphs and provide a practical way of selecting or modifying module assignments and schedules such that the resulting synthesized data path requires a small number of test resources to test itself.I.
Ishwar Parulkar, Sandeep Gupta 0001, Melvin A. Breuer
DAC2
1996 Process-Aggravated Noise (PAN): New Validation and Test Problems
abstract
How the trends in circuit design are increasing the significance of noise effects, such as crosstalk and ground bounce, is demonstrated. Further aggravated by process variations, these process aggravated noise (PAN) effects must be considered as an integral part of the design validation methodology. It is shown that the validation of a design in the presence of PAN effects requires simulation at new design corners that are not typically considered during the validation of digital CMOS circuits. Furthermore, it is shown that the choice of a new design corner depends not only on the particular PAN effect under consideration, but also on the nature of the circuit. We demonstrate the need for automating the process of selecting design corners as well as test sequences for validation and outline strategies for their automation.
Melvin A. Breuer, Sandeep Gupta 0001
ITC2
1996 Delay Fault Testing: How Robust are Our Models?
Sandeep Gupta 0001, Slawomir Pilarski, Sudhakar M. Reddy, Jacob Savir, Prab Varma
VTS1
1996 Fault macromodeling and a testing strategy for opamps
Chen-Yang Pan, Kwang-Ting Cheng, Sandeep Gupta 0001
J. Electron. Test.3
1996 BIST Test Pattern Generators for Two-Pattern Testing-Theory and Design Algorithms
abstract
Testing for delay and CMOS stuck-open faults requires two-pattern tests, and typically a large number of two pattern tests are needed. Built-in self-test (BIST) schemes are attractive for comprehensive testing of such faults. BIST test pattern generators (TPGs) for two-pattern testing, should be designed to ensure high transition coverage. In this paper, necessary and sufficient conditions to ensure complete/maximal transition coverage for linear feedback shift register (LFSR) and cellular automata (CA) have been derived. The theory developed here identifies all LFSR/CA TPGs that maximize transition coverage under any given TPG size constraint. It is shown that LFSRs with primitive feedback polynomials with large number of terms are better for two-pattern testing. Also, CA are shown to be better TPGs than LFSRs for two pattern testing, independent of their feedback rules. Based on the necessary sufficient conditions, efficient algorithms to design optimal TPGs for two-pattern testing have been developed. Experiments on benchmark circuits indicate that TPGs designed using the procedures outlined in this paper obtain high robust path delay fault coverage in short test lengths.
Chih-Ang Chen, Sandeep Gupta 0001
IEEE Trans. Computers2
1996 Utilization of On-Line (Concurrent) Checkers During Built-In-Self-Test and Vice Versa
abstract
Concurrent checkers are commonly used in computer systems to detect computational errors on-line, which enhances reliability. Using the coding theory framework developed earlier by the authors, it is shown in the following that concurrent checkers, already available within the circuit, can be utilized very effectively during off-line testing. Specifically, test time as well as fault escape probability can both be reduced simultaneously. The proposed combined scheme can be implemented with simple modification of existing hardware. Also shown is a novel use of BIST hardware for concurrent checking. Specifically proposed is a novel, dual use of concurrent checkers and built-in self-test hardware, yielding mutual advantage.
Sandeep Gupta 0001, Dhiraj K. Pradhan
IEEE Trans. Computers1
1996 A Simulator for At-Speed Robust Testing of Path Delay Faults in Combinational Circuits
abstract
Conditions are derived for robust testing of a path delay fault via a sequence of vectors applied at-speed. A simulator has been developed that uses the above conditions, along with the knowledge of paths that are robustly tested by the previous vectors, to determine the fault coverage obtained by such testing. The results demonstrate that the existing fault simulators can overestimate robust path delay fault coverage by 5-15%.
Yuan-Chieh Hsu, Sandeep Gupta 0001
IEEE Trans. Computers2
1995 Data Path Allocation for Synthesizing RTL Designs with Low BIST Area Overhead
abstract
Article Data path allocation for synthesizing RTL design with low BIST area overhead Share on Authors: Ishwar Parulkar Department of Electrical Engineering - Systems, University of Southern California, Los Angeles, CA Department of Electrical Engineering - Systems, University of Southern California, Los Angeles, CAView Profile , Sandeep Gupta Department of Electrical Engineering - Systems, University of Southern California, Los Angeles, CA Department of Electrical Engineering - Systems, University of Southern California, Los Angeles, CAView Profile , Melvin A. Breuer Department of Electrical Engineering - Systems, University of Southern California, Los Angeles, CA Department of Electrical Engineering - Systems, University of Southern California, Los Angeles, CAView Profile Authors Info & Claims DAC '95: Proceedings of the 32nd annual ACM/IEEE Design Automation ConferenceJanuary 1995 Pages 395–401https://doi.org/10.1145/217474.217561Online:01 January 1995Publication History 37citation177DownloadsMetricsTotal Citations37Total Downloads177Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Ishwar Parulkar, Sandeep Gupta 0001, Melvin A. Breuer
DAC2
1995 Test embedding with discrete logarithms
abstract
When using Built-In Self Test (BIST) for testing VLSI circuits, a major concern is the generation of proper test patterns that detect the faults of interest. Usually a linear feedback shift register (LFSR) is used to generate test patterns. We first analyze the probability that an arbitrary pseudo-random test sequence of short length detects all faults. The term short is relative to the probability of detecting the fault having the fewest test patterns. We then show how to guide the search for an initial state (seed) for a LFSR with a given primitive feedback polynomial so that all the faults of interest are detected by a minimum length test sequence. Our algorithm is based on finding the location of test patterns in the sequence generated by this LPSR. This is accomplished using, the theory of discrete logarithms. We then select the shortest subsequence that includes test patterns for all the faults of interest, hence resulting in 100% fault coverage.>
Mody Lempel, Sandeep Gupta 0001, Melvin A. Breuer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 Clock Grouping: A Low Cost DFT Methodology for Delay Testing
abstract
Article Free Access Share on Clock grouping: a low cost DFT methodology for delay testing Authors: Wen-Chang Fang Electrical Engineering Systems, University of Southern California, Los Angeles CA Electrical Engineering Systems, University of Southern California, Los Angeles CAView Profile , Sandeep K. Gupta Electrical Engineering Systems, University of Southern California, Los Angeles CA Electrical Engineering Systems, University of Southern California, Los Angeles CAView Profile Authors Info & Claims DAC '94: Proceedings of the 31st annual Design Automation ConferenceJune 1994 Pages 94–99https://doi.org/10.1145/196244.196291Published:06 June 1994Publication History 10citation181DownloadsMetricsTotal Citations10Total Downloads181Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Wen-Chang Fang, Sandeep Gupta 0001
DAC2
1994 Random pattern testable logic synthesis
Chen-Huan Chiang, Sandeep Gupta 0001
ICCAD2
1994 A comprehensive fault macromodel for opamps
abstract
In this paper, a comprehensive macromodel for transistor level faults in an operational amplifier is developed. With the observation that faulty behavior at output may result from interfacing error in addition to the faulty component, parameters associated with input and output characteristics are incorporated. Test generation and fault classification are addressed for stand-alone opamps. A high fault coverage is achieved by a proposed testing strategy. Transistor level short/bridging faults are analyzed and classified into catastrophic faults and parametric faults. Based on the macromodels for parametric faults, fault simulation is performed for an active filter. We found many parametric faults in the active filter cannot be detected by traditional functional testing. A DFT scheme alone with a current testing strategy to improve fault coverage is proposed.
Chen-Yang Pan, Kwang-Ting Cheng, Sandeep Gupta 0001
ICCAD3
1994 ATPG for Heat Dissipation Minimization During Test Application
abstract
A new ATPG algorithm has been proposed that reduces average heat dissipation (between successive test vectors) during test application. The objective is to permit safe and inexpensive testing of low power circuits and bare dies that would otherwise require expensive heat removal equipment for testing at high speeds. Three new functions, namely transition controllability, observability and test generation costs, have been defined. It has been shown that the transition test generation cost is the minimum number of transitions required to test the corresponding stuck-at fault in fanout free circuits. This cost function is used for target fault selection while the other two functions are used to guide the backtrace and objective selection procedures of PODEM. The tests generated by the proposed ATPG decrease heat dissipation during test application by a factor of 2-23 for benchmark circuits.
Seongmoon Wang, Sandeep Gupta 0001
ITC2
1994 Test embedding with discrete logarithms
abstract
When using Built-In Self Test (BIST) for testing VLSI circuits, a major concern is the generation of proper test patterns that detect the faults of interest. Usually a linear feedback shift register (LFSR) is used to generate test patterns. In this paper the authors show how to select a feedback polynomial and an initial state (seed) for the LFSR so that all the faults of interest are detected by a minimum length test sequence. Minimality is per polynomial. The algorithm is based on finding the location of test patterns in the sequence generated by an LFSR with a specific feedback polynomial. This is based on the theory of discrete logarithms. The authors then select the shortest subsequence that includes test patterns for all the faults of interest, hence 100% fault coverage.>
Mody Lempel, Sandeep Gupta 0001, Melvin A. Breuer
VTS2
1994 Weighted random robust path delay testing of synthesized multilevel circuits
abstract
Importance of delay testing is growing especially for high speed circuits. Delay testing using automatic test equipment is expensive. Built-in self-test can significantly reduce the cost of comprehensive delay testing by replacing the test equipment. It was found that several multilevel, synthesized, robust path delay testable circuits require impractically long pseudo-random test sequences. Weighted random testing techniques have been developed for robust path delay testing. The proposed technique is successfully applied to these circuits and 100% robust path delay fault coverage obtained using only 1-2 sets of weights.>
Sandeep Gupta 0001
VTS2
1993 An Efficient Partitioning Strategy for Pseudo-Exhaustive Testing
abstract
Pseudo-exhaustive testing involves applying all possible input patterns to individual output cones of a circuit.Circuits with output cones driven
Rajagopalan Srinivasan, Sandeep Gupta 0001, Melvin A. Breuer
DAC2
1993 Novel Test Pattern Generators for Pseudo-Exhaustive Testing
abstract
Pseudo-exhaustive testing involves applying all possible input patterns to individual output cones of a circuit. The testing ensures detection of all combinational faults within individual cones. Test pattern generators based on coding theory principles are not tailored for circuit-under-test and generate inefficient pseudo-exhaustive test sets. We shall describe novel hardware efficient test pattern generators that employ knowledge of the circuit output cone structures for generating minimal test sets. Using our techniques, we have designed generators that generate minimum test sets for partitioned versions of the combinational benchmark circuits.>
Rajagopalan Srinivasan, Sandeep Gupta 0001, Melvin A. Breuer
ITC2
1992 Can Concurrent Checkers Help BIST?
abstract
Concurrent checkers are commonly used in computer systems to detect computational errors on-line, which enhances reliability. Using the coding theory framework developed earlier by the authors, concurrent checkers already available within the circuit are shown to be significant help to off-line testing. Specifically, test time can be reduced while improving the fault escape probability. The proposed combined scheme can be implemented with simple modification of existing hardware. Specifically proposed is a novel, dual use of concurrent checkers and BIST hardware, yielding mutual advantage.
Sandeep Gupta 0001, Dhiraj K. Pradhan
ITC1
1992 Recent advances in BIST
abstract
The author briefly reviews various recent results in data compression for BIST. His primary focus is on MISR compression, considering the practical design issues that interest a designer. Issues related to the choice of error model, feedback polynomial, and suitable test lengths are discussed.>
Sandeep Gupta 0001
VTS1
1991 Aliasing and Diagnosis Probability in MISR and STUMPS Using a General Error Model
abstract
A number of methods have been proposed to study aliasing in MISR compression. However, most of the methods can compute aliasing probability only for specific test lengths and/or specific error models. Recently, a GLFSR structure [15] was introduced which admits coding theory formulation. The conventional signature analyzers such as LFSR and MISR form special cases of this GLFSR structure. Using this formulation, a general result is now presented which computes the exact aliasing probability for MISRs with primitive feedback polynomials, for any test length and for any error model. The framework is then extended to study the probability of correct diagnosis when faulty signature is used to identify the faulty CUT in the STUMPS environment. Specifically, the results in [7, 15, 161 are extended by proposing two new error models, a general error model which subsumes all the commonly used models, and a fixed magnitude error model which is shown to be useful for fault diagnosis. It is shown how statistical simulation can be used to determine the general error model, for a given CUT. Aliasing for some benchmark circuits, for various error models and test lengths is studied.
Mark G. Karpovsky, Sandeep Gupta 0001, Dhiraj K. Pradhan
ITC2
1991 A New Framework for Designing and Analyzing BIST Techniques and Zero Aliasing Compression
abstract
A general framework for shift register-based signature analysis is presented, and a mathematical model for this framework-based on coding theory-is developed. There are two key features of this formulation, first, it allows for uniform treatment of LFSR, MISR, and multiple MISR-based signature analyzer. In addition, using this formulation, a new compression scheme for multiple output CUT is proposed. This scheme, referred to as multiinput LFSR, has the potential to achieve better aliasing than other schemes such as the multiple MISR scheme of comparable hardware complexity. Several results on aliasing are presented, and certain known results are shown to be direct consequences of the formulation. Also developed are error models that take into account the circuit topology and the effect of faults at the outputs. Using these models, exact closed-form expressions for aliasing probability are developed. A closed-form aliasing expression for MISR under an independent error model is provided.>
Dhiraj K. Pradhan, Sandeep Gupta 0001
IEEE Trans. Computers2
1990 Aliasing Probability for Multiple Input Signature Analyzer
abstract
Single and multiple multiple-input-signature-register (MISR) aliasing probability expressions are presented for arbitrary test lengths. A framework, based on algebraic codes, is developed for the analysis and synthesis of MISR-based test response compressors for BIST. This framework is used to develop closed-form expressions for the aliasing probability of MISR for arbitrary test length. An error model, based on q-ary symmetric channel, is proposed using more realistic assumptions. Results are presented that provide the weight distributions for q-ary codes (q=2/sup m/, where the circuit under test has m outputs). These results are used to compute the aliasing probability for the MISR compression technique for arbitrary test lengths. This result is extended to compression using two different MISRs. It is shown that significant improvements can be obtained by using two signature analyzers instead of one. The weight distribution of a class of codes of arbitrary length is also given.>
Dhiraj K. Pradhan, Sandeep Gupta 0001, Mark G. Karpovsky
IEEE Trans. Computers2
1988 Concurrent Control of Multiple BIT Structures
abstract
A generic control graph for activating common built-in test structures is derived and its microprogrammed and hardwired implementations described. Three designs for activating multiple BIT structures concurrently are also presented along with simulation results of area/test time tradeoffs. Two designs for this generic controller are presented. The first design augments the classical microprogrammed controller with some circuitry used for counting. This design is most useful when a microprogrammable controller already exists for normal control operations. The second design is a hardwire unit which efficiently implements the generic controller. A complete circuit-level implementation is described. The controller can be easily made self-testing by modifying an existing register to support signature analysis. In addition, the controller is quite simple and can be incorporated on-chip. >
Sandeep Gupta 0001, Melvin A. Breuer, Jung-Cheun Lien
ITC1
1988 A New Framework for Designing and Analyzing BIST Techniques: Computation of Exact Aliasing Probability
abstract
A coding theory framework is developed for analysis and synthesis of compression techniques in the built-in self test (BIST) environment. Using this framework, exact expressions are derived for the linear feedback shift register aliasing probability. These are shown to be more accurate than earlier ones. Also shown is that there exist compression techniques for which the aliasing probability can be reduced to zero asymptotically. An error model is presented that incorporates the effects of faults on output response. It is shown that the coding theory framework correlates well with this proposed error model. A signature analysis technique is presented, which achieves smaller aliasing probability than other recently proposed schemes.>
Sandeep Gupta 0001, Dhiraj K. Pradhan
ITC1