VLDB 2026 Research / reviewers in the wild / expert
Sandip Kundu
dblp:84/1396
· DBLP profile ↗
155ranked-venue papers
26as first author
12since 2021 · last 2026
0000-0001-8221-3824ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 147 · 26 first-author · 10 since 2021Software engineering, systems software and programming languages · 29 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3Security and privacy · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sky to Edge-Cloud: Heterogeneous Computation Offloading for Energy-Efficient Drone ComputingabstractOffloading computation from edge devices to an edge cloud can save energy and enhance performance, but managing workloads from heterogeneous devices such as CPUs, GPUs, and NPUs remains challenging. This paper presents a prototype heterogeneous offloading system using an FPGA-based edge cloud capable of handling diverse computation models. Using drone computing as a use-case, we demonstrate that applications such as depth estimation, object detection, and Simultaneous Localization and Mapping (SLAM) can efficiently offload CPU, GPU, and DSP tasks to a reconfigurable accelerator. By intelligently mapping heterogeneous kernels to the reconfigurable FPGA edge cloud, our system reduces drone energy by up to 90%, underscoring its heterogeneous kernel offloading capability and real-time performance and energy gains. Zhehang Zhang, Bharadwaj Madabhushi, Sandip Kundu, Russell Tessier |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Managing Computation Offloading from Edge Devices to a Reconfigurable Edge CloudabstractEdge cloud computing is becoming increasingly vital to meet the computational demands of billions of interconnected devices, many of which operate under strict power and latency constraints. Modern 5G networks and advanced low-latency communication technologies enable rapid data transfer between edge devices and the cloud, making computational offloading both feasible and efficient. While traditional edge platforms have primarily relied on multi-core microprocessors, the growing architectural diversity of edge devices necessitates new approaches that support heterogeneous edge computing. This paper presents a novel edge computing framework that harnesses the capabilities of FPGAs to address the complexities of heterogeneous edge environments. A dynamically reconfigurable edge cloud enables seamless support for architecturally diverse edge devices. Central to our solution is an intelligent Offload Management System (OMS) that makes real-time decisions about whether to offload tasks or execute them locally, based on resource availability, energy efficiency, and deadline constraints. We validate our approach using an experimental setup featuring quad-core ARM Cortex-A76 and Qualcomm Snapdragon processors found in edge devices paired with an AMD ZCU104 FPGA board as an edge cloud node. Our results demonstrate how multiple edge devices can collaboratively utilize shared cloud resources via intelligent offloading. We specifically assess machine learning workloads processed with deep learning processing unit (DPU)-based models, showcasing the significant potential of a reconfigurable edge cloud in servicing edge devices. Zhehang Zhang, Bharadwaj Madabhushi, Sandip Kundu, Russell Tessier |
ASAP | 3 |
| 2024 | Memory Scraping Attack on Xilinx FPGAs: Private Data Extraction from Terminated ProcessesabstractFPGA-based hardware accelerators are becoming increasingly popular due to their versatility, customizability, energy efficiency, constant latency, and scalability. FPGAs can be tailored to specific algorithms, enabling efficient hardware implementations that effectively leverage algorithm parallelism. This can lead to significant performance improvements over CPUs and GPU s, particularly for highly parallel applications. For example, a recent study found that Stratix 10 FPGAs can achieve up to 90% of the performance of a TitanX Pascal GPU while consuming less than 50% of the power. This makes FPGAs an attractive choice for accelerating machine learning (ML) workloads. However, our research finds privacy and security vulnerabilities in existing Xilinx FPGA-based hardware acceleration solutions. These vulnerabilities arise from the lack of memory initialization and insufficient process isolation, which creates potential avenues for unauthorized access to private data used by processes. To illustrate this issue, we conducted experiments using a Xilinx ZCU104 board running the PetaLinux tool from Xilinx. We found that PetaLinux does not effectively clear memory locations associated with a terminated process, leaving them vulnerable to memory scraping attack (MSA). This paper makes two main contributions. The first contribution is an attack methodology of using the Xilinx debugger from a different user space. We find that we are able to access process IDs, virtual address spaces, and pagemaps of one user from a different user space because of lack of adequate process isolation. The second contribution is a methodology for characterizing terminated processes and accessing their private data. We illustrate this on Xilinx ML application library. Bharadwaj Madabhushi, Sandip Kundu, Daniel E. Holcomb |
DATE | 2 |
| 2024 | Highly Efficient Self-checking Matrix Multiplication on Tiled AMX AcceleratorsabstractGeneral Matrix Multiplication (GEMM) is a computationally expensive operation that is used in many applications such as machine learning. Hardware accelerators are increasingly popular for speeding up GEMM computation, with Tiled Matrix Multiplication (TMUL) in recent Intel processors being an example. Unfortunately, the TMUL hardware is susceptible to errors, necessitating online error detection. The Algorithm-based Error Detection (ABED) technique is a powerful technique to detect errors in matrix multiplications. In this article, we consider implementation of an ABED technique that integrates seamlessly with the TMUL hardware to minimize performance overhead. Unfortunately, rounding errors introduced by floating-point operations do not allow a straightforward implementation of ABED in TMUL. Previously an error bound was considered for addressing rounding errors in ABED. If the error detection threshold is set too low, it will a trigger false alarm, while a loose bound will allow errors to escape detection. In this article, we propose an adaptive error threshold that takes into account the TMUL input values to address the problem of false triggers and error escapes and provide a taxonomy of various error classes. This threshold is obtained from theoretical error analysis but is not easy to implement in hardware. Consequently, we relax the threshold such that it can be easily computed in hardware. While ABED ensures error-free computation, it does not guarantee full coverage of all hardware faults. To address this problem, we propose an algorithmic pattern generation technique to ensure full coverage for all hardware faults. To evaluate the benefits of our proposed solution, we conducted fault injection experiments and show that our approach does not produce any false alarms or detection escapes for observable errors. We conducted additional fault injection experiments on a Deep Neural Network (DNN) model and find that if a fault is not detected, it does not cause any misclassification. Chandra Sekhar Mummidi, Victor da Cruz Ferreira, Sudarshan Srinivasan, Sandip Kundu |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | A Novel Fault-Tolerant Architecture for Tiled Matrix MultiplicationabstractGeneral matrix multiplication (GEMM) is common to many scientific and machine-learning applications. Convolution, the dominant computation in Convolutional Neural Networks (CNNs), can be formulated as a GEMM problem. Due to its widespread use, a new generation of processors features GEMM acceleration in hardware. Intel recently announced an Advanced Matrix Multiplication (AMX®) instruction set for GEMM, which is supported by 1kB AMX registers and a Tile Multiplication unit (TMUL) for multiplying tiles (sub-matrices) in hardware. Silent Data Corruption (SDC) is a well-known problem that occurs when hardware generates corrupt output. Google and Meta recently reported findings of SDC in GEMM in their data centers. Algorithm-Based Fault Tolerance (ABFT) is an efficient mechanism for detecting and correcting errors in GEMM, but classic ABFT solutions are not optimized for hardware acceleration. In this paper, we present a novel ABFT implementation directly on hardware. Though the exact implementation of Intel TMUL is not known, we propose two different TMUL architectures representing two design points in the area-power-performance spectrum and illustrate how ABFT can be directly incorporated into the TMUL hardware. This approach has two advantages: (i) an error can be concurrently detected at the tile level, which is an improvement over finding such errors only after performing the full matrix multiplication; and (ii) we further demonstrate that performing ABFT at the hardware level has no performance impact and only a small area, latency, and power overhead. Sandeep Bal, Chandra Sekhar Mummidi, Victor da Cruz Ferreira, Sudarshan Srinivasan, Sandip Kundu |
DATE | 5 |
| 2023 | NUMAlloc: A Faster NUMA Memory AllocatorabstractThe NUMA architecture accommodates the hardware trend of an increasing number of CPU cores. It requires the cooperation of memory allocators to achieve good performance for multithreaded applications. Unfortunately, existing allocators do not support NUMA architecture well. This paper presents a novel memory allocator – NUMAlloc, that is designed for the NUMA architecture. is centered on a binding-based memory management. On top of it, proposes an “origin-aware memory management” to ensure the locality of memory allocations and deallocations, as well as a method called “incremental sharing” to balance the performance benefits and memory overhead of using transparent huge pages. According to our extensive evaluation, NUMAlloc has the best performance among all evaluated allocators, running 15.7% faster than the second-best allocator (mimalloc), and 20.9% faster than the default Linux allocator with reasonable memory overhead. NUMAlloc is also scalable to 128 threads and is ready for deployment. Hanmei Yang, Wei Wang 0054, Sandip Kundu, Bo Wu 0002, Hui Guan 0001, Tongping Liu |
ISMM | 5 |
| 2023 | ACTION: Adaptive Cache Block Migration in Distributed Cache ArchitecturesabstractChip multiprocessors (CMP) with more cores have more traffic to the last-level cache (LLC) . Without a corresponding increase in LLC bandwidth, such traffic cannot be sustained, resulting in performance degradation. Previous research focused on data placement techniques to improve access latency in Non-Uniform Cache Architectures (NUCA) . Placing data closer to the referring core reduces traffic in cache interconnect. However, earlier data placement work did not account for the frequency with which specific memory references are accessed. The difficulty of tracking access frequency for all memory references is one of the main reasons why it was not considered in NUCA data placement. In this research, we present a hardware-assisted solution called ACTION ( A daptive C ache Block Migra tion ) to track the access frequency of individual memory references and prioritize placement of frequently referred data closer to the affine core. ACTION mechanism implements cache block migration when there is a detectable change in access frequencies due to a shift in the program phase. ACTION counts access references in the LLC stream using a simple and approximate method and uses a straightforward placement and migration solution to keep the hardware overhead low. We evaluate ACTION on a 4-core CMP with a 5x5 mesh LLC network implementing a partitioned D-NUCA against workloads exhibiting distinct asymmetry in cache block access frequency. Our simulation results indicate that ACTION can improve CMP performance by up to 7.5% over state-of-the-art (SOTA) D-NUCA solutions. Chandra Sekhar Mummidi, Sandip Kundu |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | A Highly-Efficient Error Detection Technique for General Matrix Multiplication using Tiled Processing on SIMD ArchitectureabstractGeneral Matrix Multiplication (GEMM) is instrumental in myriads of scientific, high-performance computing, and machine learning applications such as computer vision, recommendation models, and weather forecasts. It is vital to make them fail-safe in safety-critical and high-precision applications. Companies like Meta and Google have recently reported sporadic silent errors in GEMM computations traced to hardware sources. Silent errors are hard to detect, requiring specialized solutions to detect them. Hardware redundancy approaches such as double or triple modular redundancy effectively detect or correct such errors, but they have a large area and power overhead. Algorithm-based Fault Tolerance (ABFT) has been shown to offer an effective alternative at a far lower overhead. Modern CPUs feature advanced vector extensions (AVX) capable of executing SIMD instructions. This paper describes a new ABFT approach designed to take advantage of the AVX feature. Our core algorithm relies on the classical tile-based outer-product approach but enhances standard check-sum calculation using a tile vector. The implementation parameters are fine-tuned to fit the available number of AVX registers. Our results indicate that we can achieve 100% error detection in GEMM at an overhead of just 0.21% for the integer data type. Unfortunately, due to rounding errors, addition of floating-point numbers is not an associative operation, creating difficulties for ABFT. To mitigate the impact of rounding errors, we introduce the concept of relative error checking and perform error analysis for various error classes to show that the proposed approach totally eliminates false positive errors. Chandra Sekhar Mummidi, Sandeep Bal, Brunno F. Goldstein, Sudarshan Srinivasan, Sandip Kundu |
ICCD | 5 |
| 2022 | A Software Approach Towards Defeating Power Management Side Channel LeakageabstractHardware Trojans are malicious, undesired, intentional modifications introduced in an Integrated Circuit (IC) which can be leveraged by a knowledgeable adversary to compromise the security of the IC. Trojans might be designed to modify the functionality of an IC, disable the security of a chip, access secret information or even destroy a system. In this paper, we propose PMU-Trojan, a hardware Trojan for leaking confidential information, such as, cryptographic secret key covertly to an adversary. For information leakage by hardware Trojan, we exploit a backdoor created by Power Management Unit (PMU) in Multi Processor System on Chip (MPSoC). PMU is a system block that initiates the voltage and the frequency changes to facilitate flexible power management and energy efficiency. It transmits voltage level change request to power supply. In this paper, we leverage this facility as an information side-channel to leak information to power-supply co-tenants. While the proposed approach can be generalized for any kind of secret information leakage, for the purpose of illustration, in this work, we focus on leaking Advanced Encryption Standard (AES) key. We demonstrate the working principle in Linux environment where a co-tenant thread monitors the change in voltage level and receives side-channel information from a thread affected by PMU-Trojan. The proposed Trojan defeats the traditional Trojan detection and suppression methods due to low information bit rate spread over long duration by a Trojan unit dissipating power at mere pico-Watts level. We propose a novel technique to defeat power management side channel by dynamically tuning processor power limit. The proposed software based solution towards suppressing PMU-Trojan is demonstrated on Intel computing platform using RAPL (Running Average Power Limit) interface. Md. Nazmul Islam, Sandip Kundu |
IOLTS | 2 |
| 2022 | Reinforcement Learning based Multi-Attribute Slice Admission Control for Next-Generation Networks in a Dynamic Pricing EnvironmentabstractNext-generation networks will provide intelligent infrastructure and management using machine learning. In real-world applications, demand for resources and performance within a service class may vary over time. Infrastructure providers choose which requests to accept with the goal of long-term profit maximization – a process known as slice admission control. In this paper, we envision a dynamic system with varying service requests attributes in urgency, duration, and amount of resources (i.e., computing, network, and storage). Further, we develop a dynamic pricing model that is responsive to demand and supply resulting in demand and supply reserve shaping. Then, we propose a solution to the slice admission control problem by using a reinforcement learning approach with Deep-Q Networks, where the state of the system is modeled using an array of parameters, similar to the input matrix in computer vision. Results show that our computer vision-inspired approach is capable of learning how the better policy to navigate this complex environment by selecting service requests that maximize the provider’s long-term profit. Victor da Cruz Ferreira, Haitham H. Esmat, Beatriz Lorenzo, Sandip Kundu, Felipe M. G. França |
VTC Spring | 4 |
| 2022 | NN-Lock: A Lightweight Authorization to Prevent IP Threats of Deep Learning ModelsabstractThe prevalent usage and unparalleled recent success of Deep Neural Network (DNN) applications have raised the concern of protecting their Intellectual Property (IP) rights in different business models to prevent the theft of trade secrets. In this article, we propose a lightweight, generic, key-based DNN IP protection methodology, NN-Lock , to defend against unauthorized usage of stolen DNN models. NN-Lock utilizes SBox, a cryptographic primitive, with good security properties to encrypt each parameter of a trained DNN model with the secret keys derived from a master key through a key-scheduling algorithm. The method ensures that only an authorized user with a correct master key can accurately use the locked DNN model. Evaluation results of NN-Lock on a Google Coral edge device for various DNN architectures on several datasets show that for an incorrect master key, the accuracy of a locked model is that of a random classifier. The dense network of encrypted parameters makes the method robust against the model fine-tuning attack and a novel approximation attack using the Genetic Algorithm, which achieves reasonable success against another recent IP protection scheme called HPNN Chakraborty et al. 2020 . The security evaluation of NN-Lock against other families of attacks demonstrates its soundness in practical scenarios. NN-Lock does not modify any internal structure of a DNN model, making it scalable for all of the existing DNN implementations without adversely affecting their performance. Manaar Alam, Sayandeep Saha, Debdeep Mukhopadhyay, Sandip Kundu |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2021 | MILR: Mathematically Induced Layer Recovery for Plaintext Space Error Correction of CNNsabstractThe increased use of Convolutional Neural Networks (CNN) in mission-critical systems has increased the need for robust and resilient networks in the face of both naturally occurring faults as well as security attacks. The lack of robustness and resiliency can lead to unreliable inference results. Current methods that address CNN robustness require hardware modification, network modification, or network duplication. This paper proposes MILR a software-based CNN error detection and error correction system that enables recovery from single and multi-bit errors. The recovery capabilities are based on mathematical relationships between the inputs, outputs, and parameters(weights) of the layers; exploiting these relationships allows the recovery of erroneous parameters (iveights) throughout a layer and the network. MILR is suitable for plaintext-space error correction (PSEC) given its ability to correct whole-weight and even whole-layer errors in CNNs. Jonathan Ponader, Kyle Thomas, Sandip Kundu, Yan Solihin |
DSN | 3 |
| 2020 | Building a portable deeply-nested implicit information flow trackingabstractDynamic Information Flow Tracking has been successfully used to prevent a wide range of attacks and detect illegal access to sensitive information. Most proposed solutions only track the explicit information flow where the taint is propagated through data dependencies. However, recent evasion attacks exploit implicit flows, that use control flow in the application, to manipulate the data thus making the malicious activity undetectable. We propose NIFT - a nested implicit flow tracking mechanism that extends explicit propagation to instructions affected by a control dependency. Our technique generates taint instructions at compile time which are executed by specialized hardware to propagate taint implicitly even in cases of deeply-nested branches. In addition, we propose a restricted taint propagation for data executed in conditional branches that affects only immediate instructions instead of all instructions inside the branch scope. Our technique efficiently locates implicit flows and resolves them with negligible performance overhead. Moreover, it mitigates the over-tainting problem. Leandro Santiago de Araújo, Leandro A. J. Marzulo, Tiago A. O. Alves, Felipe M. G. França, Israel Koren, Sandip Kundu |
CF | 6 |
| 2020 | Towards Adversarial Attack Resistant Deep Neural Networks
Tiago A. O. Alves, Sandip Kundu |
ESANN | 2 |
| 2020 | Reduced Fault Coverage as a Target for Design Scaffolding SecurityabstractThe hardware design process adds, at each level, scaffolding logic to aid test, debug and engineering changes. Fault injection attacks can place a design in a non-functional state where the scaffolding is utilized for obtaining information that is not intended for the user. To counter such security threats, this paper suggests to include in the design logic that will identify invalid state patterns, and reset or hold a sufficient part of the state to prevent useful computations from occurring. To support such a solution, the paper also suggests that the fault coverage of a scan-based test set in the presence of single stuck-at faults can be used for designing the reset and hold logic. Such a test set relies on the use of non-functional states for fault detection. Without using the reset and hold logic, the test set achieves a high fault coverage (over 80% in the experiments reported in this paper). When the reset and hold logic is used, the fault coverage of the same test set is reduced below a preselected target (30% in the experiments reported in this paper). The reduction in the fault coverage occurs because much of the non-functional state space is no longer available. This also eliminates the security risk associated with these states. Irith Pomeranz, Sandip Kundu |
IOLTS | 2 |
| 2020 | Weightless Neural Networks as Memory Segmented Bloom Filters
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França |
Neurocomputing | 8 |
| 2019 | Efficient Testing of Physically Unclonable Functions for UniquenessabstractPhysically unclonable functions (PUFs) have emerged as lightweight hardware security primitives for implementing secure authentication. Strong PUFs rely on random manufacturing process variation to create unique Boolean mappings from input (challenge) to output (output). For secure authentication, challenge to response mappings are required to be unique for each device. However, uniqueness is not guaranteed by design or manufacturing. Testing for uniqueness and weeding out non-unique parts are the only way to ensure uniqueness of devices. Uniqueness testing can be expensive in time as the challenge-responses of the N th device under-test, must be proven to be different from previously tested N - 1 devices, or the device must be discarded. To reduce the time complexity of uniqueness testing, Multi-Index hashing (MIH) was proposed for online testing in high volume manufacturing. Database search using MIH was shown to be fast, but it suffers from high memory cost. In this paper, we address the memory problem of MIH based uniqueness testing by proposing alternative MIH strategies. Our results indicate that the proposed search strategies can significantly reduce the memory cost without sacrificing performance, requiring ≈ 3.35× less memory with just a 17% performance overhead when testing the uniqueness of 1 million PUFs. Leandro Santiago de Araújo, Vinay C. Patil, Leandro A. J. Marzulo, Felipe M. G. França, Sandip Kundu |
ATS | 5 |
| 2019 | Memory Efficient Weightless Neural Network using Bloom Filter
Leandro Santiago de Araújo, Letícia Dias Verona, Fábio Medeiros Rangel, Fabrício Firmino de Faria, Daniel Sadoc Menasché, Wouter Caarls, Maurício Breternitz, Sandip Kundu, Priscila M. V. Lima, Felipe M. G. França |
ESANN | 8 |
| 2019 | MLPrivacyGuard: Defeating Confidence Information based Model Inversion Attacks on Machine Learning SystemsabstractAs services based on Machine Learning (ML) applications find increasing use, there is a growing risk of attack against such systems. Recently, adversarial machine learning has received a lot of attention, where an adversary is able to craft an input or manipulate an input to cause an ML system to misclassify. Another attack of concern is when an adversary with access to a ML model can reverse engineer attributes of a target class, creating a privacy concern, which is the subject of this paper. Such attacks use non-sensitive data obtainable by the adversary and the confidence levels returned by the ML model to infer sensitive attributes of the target user. Model Inversion attacks may be classified as white-box, where the ML model is known to the attacker, or black-box, where the adversary does not know the internals of the model. If the attacker has access to non-sensitive data of a target user, they can infer sensitive data by applying gradient ascent on the confidence returned by the model. Therefore, a black-box attack can be mounted by numerical approximations of the gradient to perform the gradient ascent. In this work, we present MLPrivacyGuard, a countermeasure against black-box model inversion attack is presented. This countermeasure consists of adding controlled noise to the output of the confidence function. It is important to preserve the accuracy of prediction/classification for the real users of the model while preventing attackers to infer sensitive data. This involves a trade-off between misclassification error and the effectiveness of defense. Based on experimental results, we demonstrate that when noise is injected with a long-tailed distribution, the objectives of low misclassification error with a strong defense can be attained as model inversion attacks are neutralized because numerical approximation of gradient ascent is unable to converge. Tiago A. O. Alves, Felipe M. G. França, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Enabling IC Traceability via Blockchain Pegged to Embedded PUFabstractGlobalization of IC supply chain has increased the risk of counterfeit, tampered, and re-packaged chips in the market. Counterfeit electronics poses a security risk in safety critical applications like avionics, SCADA systems, and defense. It also affects the reputation of legitimate suppliers and causes financial losses. Hence, it becomes necessary to develop traceability solutions to ensure the integrity of supply chain, from the time of fabrication to the end of product-life, which allows a customer to verify the provenance of a device or a system. In this article, we present an IC traceability solution based on blockchain. A blockchain is a public immutable database that maintains a continuously growing list of data records secured from tampering and revision. Over the lifetime of an IC, all ownership transfer information is recorded and archived in a blockchain. This safe, verifiable method prevents any party from altering or challenging the legitimacy of the information being exchanged. However, a chain of sales record is not enough to ensure provenance of an IC. There is a need for clone-proof method for securely binding the identity of an IC to the blockchain information. In this article, we propose a method of IC supply chain traceability via blockchain pegged to embedded physically unclonable function (PUF). The blockchain provides ownership transfer record, while the PUF provides unique identification for an IC allowing it to be linked uniquely to a blockchain. Our proposed solution automates hardware and software protocols using blockchain-powered Smart Contract that allows supply chain participants to authenticate, track, trace, analyze, and provision chips throughout their entire life cycle. Md. Nazmul Islam, Sandip Kundu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | PMU-Trojan: On exploiting power management side channel for information leakageabstractHardware Trojans are malicious, undesired, intentional modifications introduced in an Integrated Circuit (IC) which can be leveraged by a knowledgeable adversary to compromise the security of the IC. Trojans might be designed to modify the functionality of an IC, access sensitive information or even disable or destroy a system. In this paper, we propose PMU-Trojan, a hardware Trojan for leaking confidential information, such as, cryptographic secret key covertly to an adversary. For information leakage by hardware Trojan, we exploit a backdoor created by Power Management Unit (PMU) in Multi Processor System on Chip (MPSoC). PMU is a system block responsible for initiating voltage and the frequency changes to facilitate flexible power management and energy efficiency. It transmits voltage level change request to power supply. In this paper we leverage this facility as an information side-channel to leak information to power-supply co-tenants. While the proposed approach is general and can be applied for any kind of secret information leakage, for the purpose of illustration, in this study, we focus on leaking Advanced Encryption Standard (AES) key. We demonstrate the working principle of this system in Linux environment where a co-tenant thread monitors the voltage level and receives side-channel information from a thread affected by the Trojan. This scheme also defeats Differential Power Analysis (DPA) based Trojan detection due to low information bit rate spread over long duration by a Trojan unit dissipating power at mere pico-Watts level. Md. Nazmul Islam, Sandip Kundu |
ASP-DAC | 2 |
| 2018 | Adaptive and polymorphic VLIW processor to optimize fault tolerance, energy consumption, and performanceabstractBecause most traditional homogeneous and heterogeneous processors have a fixed design that limits its runtime adaptability, they are not able to cope with the varying application behavior when one considers the axes of fault tolerance, performance, and energy consumption altogether. In this context, we propose a new dynamically adaptive processor design that is capable of delivering the best trade-off among these three axes according to the application at hand, or be tuned to optimize a specific metric. This is achieved by extending a polymorphic processor that can change its issue-width during runtime with specific mechanisms for fault tolerance, energy optimization, and performance enhancement. They are controlled by an optimization algorithm that evaluates and chooses which is the best configuration according to given requirements. Considering a metric that weighs all three axes, the proposed adaptive processor delivers a result that is 94.88% of the oracle processor on average, while a static configuration (defined at design time without runtime adaptation) only achieves 28.24% at most, which means that dynamic adaptation is required to cope with different application behaviors as there is not one specific configuration that fits all applications. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Sandip Kundu, Israel Koren, Antonio Carlos Schneider Beck |
CF | 3 |
| 2017 | A guide to graceful aging: How not to overindulge in post-silicon burn-in for enhancing reliability of weak PUFabstractSRAM-based Weak PUFs have become popular in tamper sensitive key storage and device ID generation. Weak PUFs rely on intrinsic process variations to produce repeatable and unique start-up behavior. However, noise in the system compromises repeatability of SRAM start-up behavior. To obviate this problem, a number of solutions such as fuzzy extraction and error correcting codes have been proposed to generate a stable key from error-prone PUF cells. Unfortunately, the overhead from these techniques grows superlinearly with increasing error rate. Recently, it was suggested that the start-up error rate can be reduced significantly by accelerating device aging, which leads to reduced overhead for error correction. To accelerate aging, devices are subjected to temperature and voltage stress in a burn-in chamber. Unfortunately, burn-in accrues significant production cost. In this paper, we present a method to reduce the cumulative burn-in time by quantifying the minimum burn-in requirement for each device. We propose a low-cost proxy to measure the degree of process variation of each device at birth and use previously proposed device aging model to determine the burn-in requirements. Our results show that this procedure reduces cumulative burn-in cost without compromising the resultant reliability of Weak PUFs. Md. Nazmul Islam, Vinay C. Patil, Sandip Kundu |
ISCAS | 3 |
| 2017 | An analytical model for predicting the residual life of an IC and design of residual-life meterabstractIntegrated Circuits age differently based on their operating conditions. A device that operates under high voltage or temperature stress ages faster. In many safety critical applications, such as in automotive systems, avionics and cyber-physical infrastructure, users would like to know the residual life of its components to ensure timely replacements. This motivates design of a residual-life meter (RLM). A residual-life meter may also prevent re-entry of recycled ICs from discarded systems into the supply chain as new parts - a problem that is deemed as a threat to critical infrastructure. Design of lifespan meter requires analytical models for predicting residual life of an IC based on its history of operating conditions. A major contribution of this paper is development of an analytical model for determining the residual life of an IC. There are a number of aging mechanisms that impact the lifetime of an IC such as oxide breakdown, hot electron effect, negative bias temperature instability, electromigration and others. While the proposed approach is general, for the purpose of illustration, in this study, we focus on residual life modeling based on Electromigration. The model relies on current temperature as input which is obtained from on-chip temperature sensor. This obviates any need for workload characterization. Next, this paper presents three alternative circuit implementations of the proposed residual-life meter. Results show that an RLM can be implemented at an area of ~ 350μm2, and frequency of 3.3 GHz in CMOS 45nm technology. The RLM relies on non-volatile storage for data retention and is immune to power shut-offs. Prior aging sensors rely on sample size of single path, which is statistically insignificant. This work provides an alternative approach of analytical modeling, which is statistically comprehensive while retaining the simplicity of a single meter design. Md. Nazmul Islam, Sandip Kundu |
VTS | 2 |
| 2017 | Physical Design Obfuscation of Hardware: A Comprehensive Investigation of Device and Logic-Level TechniquesabstractThe threat of hardware reverse engineering is a growing concern for a large number of applications. A main defense strategy against reverse engineering is hardware obfuscation. In this paper, we investigate physical obfuscation techniques, which perform alterations of circuit elements that are difficult or impossible for an adversary to observe. The examples of such stealthy manipulations are changes in the doping concentrations or dielectric manipulations. An attacker will, thus, extract a netlist, which does not correspond to the logic function of the device-under-attack. This approach of camouflaging has garnered recent attention in the literature. In this paper, we expound on this promising direction to conduct a systematic end-to-end study of the VLSI design process to find multiple ways to obfuscate a circuit for hardware security. This paper makes three major contributions. First, we provide a categorization of the available physical obfuscation techniques as it pertains to various design stages. There is a large and multidimensional design space for introducing obfuscated elements and mechanisms, and the proposed taxonomy is helpful for a systematic treatment. Second, we provide a review of the methods that have been proposed or in use. Third, we present recent and new device and logic-level techniques for design obfuscation. For each technique considered, we discuss feasibility of the approach and assess likelihood of its detection. Then we turn our focus to open research questions, and conclude with suggestions for future research directions. Arunkumar Vijayakumar, Vinay C. Patil, Daniel E. Holcomb, Christof Paar, Sandip Kundu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2017 | Guest Editorial Securing IoT Hardware: Threat Models and Reliable, Low-Power Design SolutionsabstractIt is well understood that for Internet of Things (IoT), security of underlying hardware is the key to safe and reliable operation. IoT service stack relies on security of network, software, and firmware, all of which, in turn, depend on functionality provided by the underlying hardware. The hardware may be compromised or attacked by multiple threat actors. The designer may create a backdoor that leaks vital information such as encryption key used in secure channel; the manufacturer may tamper the design by inserting hardware Trojans or introducing artifacts with known reliability vulnerabilities. Either of these actors may enable writing into protected memory areas that may store secure hash of trusted code base, allowing malware to boot directly on the hardware. Today’s designs integrate IP blocks from multiple vendors; manufactured, tested, and repaired by different companies spanning across the globe. Consequently, there are many entry points for the hardware to be compromised. For a trusted hardware design, protection and security of intellectual property cores are of paramount importance. This special section aims to publish novel solutions for security problems related to hardware used in IoT.•Secured IoT Hardware: Induction of any form of third-party intervention in the hardware design methodology may raise grave security concern for IoT hardware. Securing IoT hardware can be in the form of protecting intellectual property cores against false claim of ownership/piracy/counterfeit. The first form of security measure requires anti-piracy methodologies such as digital watermarking, hardware metering, computational forensic engineering, and obfuscation that can nullify the false claim of ownership or detect unauthorized pirated designs. The second form of threat, which is formally called “hardware Trojan,” is an act of deliberate insertion into a design (such as intellectual property core, hardware) by a rogue designer or vendor, and also requires detection/correction strategies as a security measure. Both hardware threats discussed above may occur in any of the design abstraction levels (behavioral, register transfer, layout, etc.). Handling the threats higher in the abstraction level provides more assurance against possible attacks, however, it requires a more sophisticated approach. Further more, the level at which protective measure is applied often dictates the preprocessing or postprocessing style of the approach. These calls for novel technique that embeds hardware security measure a higher abstraction level for protection of IoT devices.•Reliable IoT Hardware: Due to multiple factors affecting reliability of hardware used in IoT devices, these devices are always at a risk of malfunctioning. For example, a manufacturer may deliberately change the width of a metal line for causing premature electromigration defect, possibly triggering a timed Trojan. Multiple trigger mechanisms may be used to attack hardware such as: 1) reducing device dimensions; 2) scaling supply voltage; and 3) modulating frequency of operation. Methodologies should incorporate techniques that provide resiliency/tolerance against such faults at higher abstraction levels to assure greater reliability from the beginning of design flow.•Low-Cost IoT Hardware:Another design aspect of hardware for IoT devices is performance and power. Consumer demand drives integration of multiple functionalities, often achieved by integrating dedicated IP cores and general purpose processors working in tandem. This creates a unique challenge in maintaining security and integrity of data passing through various IP blocks. Standard solutions involving redundancy, diversity, and check run up against power, performance, and latency constraints. Anirban Sengupta 0003, Sandip Kundu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Modeling Residual Lifetime of an IC Considering Spatial and Inter-Temporal Temperature VariationsabstractIn today's CMOS technology, reliability of integrated circuits is a significant concern. Higher temperatures during circuit operation accelerates the device aging. Power dissipation in an IC exhibits both spatial and inter-temporal variations, hence various blocks age differently. However, endusers only care about the reliability of the IC as a whole - specifically, they are only concerned about the residual life of the IC. Several aging monitoring sensors have been proposed to track aging of an IC. However, current solutions are built around detection of aging of a single critical path. Reliability modeling of ICs considering spatial and inter-temporal temperature variations have not received adequate attention in the literature. In this paper, we propose an analytical model to predict the residual lifetime of an IC, considering spatial and intertemporal temperature variations. This model can dynamically re-compute residual lifetime with changes in operating temperature of various blocks. The most substantial integrated circuit aging mechanisms are electromigration, time dependent dielectric breakdown, hot carrier injection and negative bias temperature instability. For this study, we focus on residual lifetime modeling based on electromigration. Unlike prior work, we consider the effect of both spatial and inter-temporal temperature variations and propose a generalized model for calculating the residual lifetime. Md. Nazmul Islam, Sandip Kundu |
ATS | 2 |
| 2016 | Managing Reliability of Integrated Circuits: Lifetime Metering and Design for HealingabstractSummary form only given. In nanoscale CMOS devices, manufacturing control, process variations and device reliability have emerged as dominant concerns. Reliability problems manifest over lifetime of ICs in multiple ways, from gradual degradation in performance to increasing leakage current to catastrophic failures. Each of these effects are modulated by workloads executing on ICs and ambient conditions which influence circuit operating conditions such as temperature and supply voltage noise, which in term affect the rate of aging. The rate at which a device ages determine its remaining lifetime. There are 3 papers in this session that deal with (i) modeling rate of aging for predicting lifetime, (ii) monitoring IC health through proxy monitors to predict lifetime and (iii) healing of IC against degradation due to NBTI. Sandip Kundu |
ATS | 1 |
| 2016 | Improving performance per Watt of non-monotonic Multicore Processors via bottleneck-based online program phase classificationabstractHeterogeneous architectures offer the promise of higher performance/Watt compared to symmetric multi-cores. Recent works have proposed the use of non-monotonic (NM) heterogeneous architectures with diverse core types where each core has unique power and performance characteristics. However, the power and performance benefits achieved by NM architectures is highly dependent on assignment of application to the most suitable core type for all program phases. In this paper we propose a novel online program phase detection technique that is based on the frequency of cache misses and processor stalls which correspond to core resource bottlenecks. We track performance monitors to formulate a Bottleneck Type Vector (BTV) that help direct the application to most appropriate core type for execution. We compare the proposed BTV-based core assignment method to prior online core assignment approaches and demonstrate as much as 22% improvement in average performance/Watt using Instructions per Second (IPS) as the performance metric. Sudarshan Srinivasan, Israel Koren, Sandip Kundu |
ICCD | 3 |
| 2016 | Guest Editorial: IEEE Transactions on Computers and IEEE Transactions on Nanotechnology Joint Special Section on Defect and Fault Tolerance in VLSI and Nanotechnology SystemsabstractThe papers in this special issue focus on defect and fault tolerance in VLSI and nanotechnology systems. With the increasing demand for ever smaller, portable, energy-efficient and high-performance electronic systems, scaling of CMOS technology continues. As CMOS scaling approaches physical limits, continued innovation in materials, manufacturing processes, device structures and design paradigms have been necessary. High-k oxide and metal-gate stack were introduced to address oxide leakage; thin body undoped channels, and multiple-gate structures were introduced to mitigate subthreshold leakage; restricted design rules were introduced to improve layout efficiency; yet CMOS technology continues to be challenged in the areas of device aging and reliability. While CMOS is expected to be the dominant semiconductor technology for the foreseeable future, for reasons that are both technological and financial, alternatives to CMOS technology are attracting attention from the researchers. Cristiana Bolchini, Sandip Kundu, Salvatore Pontarelli |
IEEE Trans. Computers | 2 |
| 2016 | Managing Test Coverage Uncertainty due to Random Noise in Nano-CMOS: A Case-Study on an SRAM ArrayabstractStatic random access memories (SRAM) are a major constituent in high performance microprocessors and systems-on-a-chip. With scaling of technology, manufacturing process variations in SRAMs are of significant concern. SRAM cells that are marginal due to process variations suffer from stability issues where random thermal noise and random telegraph noise (RTN) become determinant in memory state. This introduces uncertainty in testing, where passing a test does not necessarily mean that SRAM is good. A marginal cell may pass testing, but fail during operation. To address this problem, we introduce a new metric for probabilistic fault coverage. We present simulation methods for measuring such coverage and propose N-detect and multilevel wordline (WL) voltage techniques to increase the detection probability of marginal cells. We demonstrate the benefit of these testing approaches on a sample SRAM array, designed in 32 nm predictive technology models. To simulate marginal cells, we oversample the tail of the process variations distribution. Transient noise simulations show that thermal noise can lead to a fault coverage loss of 20% and RTN can lead to a maximum fault coverage loss of up to 75%. N-detect technique achieves close to a 100% fault coverage at n = 100 for thermal noise and n = 70 for RTN. Multilevel WL technique requires 20-30 mV boost in WL voltage during read test to achieve near ideal fault coverage at minimal yield loss. A combination of these techniques is shown to be effective in increasing fault coverage with a tradeoff between test time and yield loss. Vikram B. Suresh, Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Abstraction-Guided Simulation Using Markov Analysis for Functional VerificationabstractThis paper presents a novel abstraction-guided simulation approach for functional verification. The results of Markov analysis of the abstract model of the design under verification are used as the guidance of simulation on the concrete design. The results of the Markov analysis can offer the information about how hard it is to reach each abstract state from the initial state, and how hard it is to reach certain target states from each abstract state. Such information is able to guide the simulation in two aspects: 1) in exploring abstract state space and 2) in exercising target state. Assuming a good abstract model, experimental results show that the simulation using Markov analysis as guidance is highly efficient in both aspects. Huawei Li 0001, Tao Lv 0001, Xiaowei Li 0001, Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | Exploring Heterogeneity within a Core for Improved Power EfficiencyabstractAsymmetric multi-core processors (AMPs) comprise cores with different sizes of micro-architectural resources yielding very different performance and energy characteristics. Since the computational demands of workloads vary from one task to the other, AMPs can often provide a higher power efficiency than symmetric multi-cores. Furthermore, as the computational demands of a task change during its course of execution, reassigning the task from one core to another, where it can run more efficiently, can further improve the overall power efficiency. However, too frequent re-assignments of tasks to cores may result in high overhead. To greatly reduce this overhead, we propose a morphable core architecture that can dynamically adapt its resource sizes, operating frequency and voltage to assume one of four possible core configurations. Such a morphable architecture allows more frequent task to core configuration re-assignments for a better match between the current needs of the task and the available resources. To make the online morphing decisions we have developed a runtime analysis scheme that uses hardware performance counters. Our results indicate that the proposed morphable architecture controlled by the runtime management scheme, can improve the throughput/Watt of applications by 31 percent over executing on a static out-of-order core while the previously proposed big/little morphable architecture achieves only a 17 percent improvement. Sudarshan Srinivasan, Nithesh Kurella, Israel Koren, Sandip Kundu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | A novel modeling attack resistant PUF design based on non-linear voltage transfer characteristics
Arunkumar Vijayakumar, Sandip Kundu |
DATE | 2 |
| 2015 | Online mechanism for reliability and power-efficiency management of a dynamically reconfigurable coreabstractPrevious studies have shown that the best way to achieve high throughput/Watt of a single threaded application is by running it on an asymmetric multicore processor (AMP). AMPs feature cores that are tuned for specific workload characteristics. To increase efficiency, the core that offers the best power-performance trade-off for the executing thread is chosen. To reduce the overhead of thread migration, we have previously proposed a morphable core that can morph into multiple core types. In this study, apart from power-performance efficiency, we also consider the reliability of the different core types as indicated by their vulnerability to soft-errors. We show that the best core type for power-efficiency may not be the best for reliability. Accordingly, we develop a multi-objective thread migration strategy to determine the best core type considering power efficiency and reliability. To support runtime decision making, we have developed online estimators for reliability and power efficiency based on performance monitoring counters. In keeping with the existing literature, we use the architectural vulnerability factor (AVF) as the metric for reliability and instructions-per-second2/Watt as the metric for power efficiency. For the multi-objective optimization we use a Cobb-Douglas production function. Our results indicate that the proposed runtime mechanism for reliability and power-efficiency improves, on the average, the throughput/Watt of applications by 24% and reduces the Soft-Error Rate (SER) by 12% compared to the best static execution. Sudarshan Srinivasan, Israel Koren, Sandip Kundu |
ICCD | 3 |
| 2015 | A Hardware Framework for Yield and Reliability Enhancement in Chip MultiprocessorsabstractDevice reliability and manufacturability have emerged as dominant concerns in end-of-road CMOS devices. An increasing number of hardware failures are attributed to manufacturability or reliability problems. Maintaining an acceptable manufacturing yield for chips containing tens of billions of transistors with wide variations in device parameters has been identified as a great challenge. Additionally, today’s nanometer scale devices suffer from accelerated aging effects because of the extreme operating temperature and electric fields they are subjected to. Unless addressed in design, aging-related defects can significantly reduce the lifetime of a product. In this article, we investigate a micro-architectural scheme for improving yield and reliability of homogeneous chip multiprocessors (CMPs). The proposed solution involves a hardware framework that enables us to utilize the redundancies inherent in a multicore system to keep the system operational in the face of partial failures. A micro-architectural modification allows a faulty core in a CMP to use another core’s resources to service any instruction that the former cannot execute correctly by itself. This service improves yield and reliability but may cause loss of performance. The target platform for quantitative evaluation of performance under degradation is a dual-core and a quad-core chip multiprocessor with one or more cores sustaining partial failure. Simulation studies indicate that when a large, high-latency, and sparingly used unit such as a floating-point unit fails in a core, correct execution may be sustained through outsourcing with at most a 16% impact on performance for a floating-point intensive application. For applications with moderate floating-point load, the degradation is insignificant. The performance impact may be mitigated even further by judicious selection of the cores to commandeer depending on the current load on each of the candidate cores. The area overhead is also negligible due to resource reuse. Abhisek Pan, Rance Rodrigues, Sandip Kundu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Does the Sharing of Execution Units Improve Performance/Power of Multicores?abstractSeveral studies and recent real-world designs have promoted sharing of underutilized resources between cores in a multicore processor to achieve better performance/power. It has been argued that when utilization of such resources is low, sharing has a negligible impact on performance while offering considerable area and power benefits. In this article, we investigate the performance and performance/watt implications of sharing large and underutilized resources between pairs of cores in a multicore. We first study sharing of the entire floating-point datapath (including reservation stations and execution units) by two cores, similar to AMD’s Bulldozer. We find that while this architecture results in power savings for certain workload combinations, it also results in significant performance loss of up to 28%. Next, we study an alternative sharing architecture where only the floating-point execution units are shared, while the individual cores retain their reservation stations. This reduces the highest performance loss to 14%. We then extend the study to include sharing of other large execution units that are used infrequently, namely, the integer multiply and divide units. Subsequently, we analyze the impact of sharing hardware resources in Simultaneously Multithreaded (SMT) processors where multiple threads run concurrently on the same core. We observe that sharing improves performance/watt at a negligible performance cost only if the shared units have high throughput. Sharing low-throughput units reduces both performance and performance/watt. To increase the throughput of the shared units, we propose the use of Dynamic Voltage and Frequency Boosting (DVFB) of only the shared units that can be placed on a separate voltage island. Our results indicate that the use of DVFB improves both performance and performance/watt by as much as 22% and 10%, respectively. Rance Rodrigues, Israel Koren, Sandip Kundu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2014 | A runtime support mechanism for fast mode switching of a self-morphing core for power efficiencyabstractAsymmetric multicore processors (AMPs) consist of cores executing the same ISA, but differing in micro-architectural resources, performance, and power consumption. As the computational bottleneck of a workload shifts from one resource to the next, during its course of execution, reassigning it to the core where it runs most efficiently can improve the overall energy efficiency. Simulation studies show that the performance bottlenecks can shift frequently, often within a few thousands cycles. With frequent core hooping, the overhead of thread migration becomes significant. To mitigate this overhead, we propose a morphable core that can assume one of four possible configurations to address the dominant performance bottlenecks, while retaining the same cache and registers. This way the architectural state remains intact while the morphable core is reconfigured in resources and frequency. We then implement a runtime scheme to decide the best configuration to run on and switch configuration as necessary. Simulation results indicate that on the average, the proposed scheme results in performance/watt improvement of 41%. Sudarshan Srinivasan, Nithesh Kurella, Israel Koren, Rance Rodrigues, Sandip Kundu |
PACT | 5 |
| 2014 | Online error detection and recovery in dataflow executionabstractThe processor industry is well on its way towards manycore processors that comprise of large number of simple cores. The shift towards multi and manycores calls for new programming paradigms suitable for exploiting the inherent parallelism in applications. Dataflow execution was shown to be a good option for programming in such environments. It is well-known that as CMOS technology continues to scale, it becomes more prone to transient and permanent hardware faults. In this paper we present a novel mechanism for error detection and recovery that focuses on transient errors in dataflow execution. Due to the inherently parallel nature of dataflow, our solution is completely distributed and synchronizes only cores that have data dependencies between them, as opposed to prior work on error recovery that in general rely on global synchronization of the system. We evaluate the proposed solution via a software implementation on top of a dataflow runtime. Experimental results show that error detection overhead is highly related to the pressure on the memory bus. In memory bound applications, performance is found to deteriorate, while for other benchmarks, the observed overhead is less than 23%. We find no comparable previous work to contrast these results. Tiago A. O. Alves, Sandip Kundu, Leandro A. J. Marzulo, Felipe M. G. França |
IOLTS | 2 |
| 2014 | A low-power instruction replay mechanism for design of resilient microprocessorsabstractThere is a growing concern about the increasing rate of defects in computing substrates. Traditional redundancy solutions prove to be too expensive for commodity microprocessor systems. Modern microprocessors feature multiple execution units to take advantage of instruction level parallelism. However, most workloads do not exhibit the level of instruction level parallelism that a typical microprocessor is resourced for. This offers an opportunity to reexecute instructions using idle execution units. But, relying solely on idle resources will not provide full instruction coverage and there is a need to explore other alternatives. To that end, we propose and evaluate two instruction replay schemes within the same core for online testing of the execution units. One scheme (RER) reexecutes only the retired instructions, while the other (REI) reexecutes all the issued instructions. The complete proposed solution requires a comparator and minor modifications to control logic, resulting in negligible hardware overhead. Both soft and hard error detection are considered and the performance and energy impact of both schemes are evaluated and compared against previously proposed redundant execution schemes. Results show that even though the proposed schemes result in a small performance penalty when compared to previous work, the energy overhead is significantly reduced. Rance Rodrigues, Arunachalam Annamalai, Sandip Kundu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2014 | Performance-driven dynamic thermal management of MPSoC based on task reschedulingabstractHigh level of integration has led to the advent of Multiprocessor System-on-Chip (MPSoC) which consists of multiple processor cores and accelerators on the same die. A MPSoC programming model is based on a task graph where tasks are assigned to cores to maximize performance. To address thermal hotspots in MPSoCs, coarse-grain power management techniques based on Dynamic Frequency Scaling (DFS) are widely used. DFS is reactive in nature and has detrimental effects on performance. We propose an alternative solution based on dynamic task rescheduling where a temperature prediction scheme is built into the scheduler. The temperature look-ahead scheme is used for task reassignment or delay insertion in scheduling. Since temperature prediction and task assignment are done at runtime, both must be simple and extremely fast. To that end, we propose a heuristic solution based on a limited branch-and-bound search and compare results against an optimal Integer Linear Programming (ILP)-based solution. The proposed approach is shown to be superior to frequency scaling, and the resulting schedule length is within 5% to 10% of the optimal solution as obtained from ILP formulation. Kunal P. Ganeshpure, Sandip Kundu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Globally Constrained Locally Optimized 3-D Power Delivery NetworksabstractDesign of power delivery network (PDN) is a constrained optimization problem. An ideal PDN must limit voltage drop that results from switching circuits' transients, satisfy current density constraints that arise from electromigration limits, yet use only minimal metal resources so that design density targets can be met. It should also provide an efficient thermal conduit to address heat flux. Furthermore, an ideal PDN should be a regular structure to facilitate design productivity and manufacturability, yet be resilient to address varying power demands across its distribution area. In 3-D ICs, these problems are further constrained by the need to minimize through-silicon via (TSV) area and bridge power lines of different dimensions across tiers, while addressing varying power demands in lateral and vertical directions. In this paper, we propose an unconventional power grid optimization solution that allows us to resize each tier individually by applying tier-specific constraints and yet be optimal in a multitier network, where each tier is locally resized while globally constrained. Tier-specific constraints are derived from electrical and thermal targets of 3-D PDNs. Two resizing algorithms are presented that optimize 3-D PDNs standalone or 3-D PDNs together with TSVs. We demonstrate these solutions on a three-tier setup where significant area savings can be achieved. Aida Todri, Sandip Kundu, Patrick Girard 0001, Alberto Bosio, Luigi Dilillo, Arnaud Virazel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | An opportunistic prediction-based thread scheduling to maximize throughput/watt in AMPsabstractThe importance of dynamic thread scheduling is increasing with the emergence of Asymmetric Multicore Processors (AMPs). Since the computing needs of a thread often vary during its execution, a fixed thread-to-core assignment is sub-optimal. Reassigning threads to cores (thread swapping) when the threads start a new phase with different computational needs, can significantly improve the energy efficiency of AMPs. Although identifying phase changes in the threads is not difficult, determining the appropriate thread-to-core assignment is a challenge. Furthermore, the problem of thread reassignment is aggravated by the multiple power states that may be available in the cores. To this end, we propose a novel technique to dynamically assess the program phase needs and determine whether swapping threads between core-types and/or changing the voltage/frequency levels (DVFS) of the cores will result in higher throughput/Watt. This is achieved by predicting the expected throughput/Watt of the current program phase at different voltage/frequency levels on all the available core-types in the AMP. We show that the benefits from thread swapping and DVFS are orthogonal, demonstrating the potential of the proposed scheme to achieve significant benefits by seamlessly combining the two. We illustrate our approach using a dual-core High-Performance (HP)/Low-Power (LP) AMP with two power states and demonstrate significant throughput/Watt improvement over different baselines. Arunachalam Annamalai, Rance Rodrigues, Israel Koren, Sandip Kundu |
PACT | 4 |
| 2013 | On dynamic polymorphing of a superscalar core for improving energy efficiencyabstractThe computational needs of a program change over time. Sometimes a program exhibits low instruction level parallelism (ILP), while at other times the inherent ILP may be higher; sometimes a program stalls due to a large number of cache misses, while at other times it may exhibit high cache throughput. Asymmetric Multicore Processors (AMP) have been proposed to allow matching the computing needs of a thread to a core where it executes most efficiently. Some of the recent works focus on AMPs consisting of a monolithic large out-of-order (OOO) core and a small in-order (InO) core. Dynamic swapping of threads between these cores is then facilitated to improve energy efficiency of the threads without impacting performance too negatively. Swapping decisions are made at coarse grain instruction granularities to mitigate the impact of migration overhead. This excludes many opportunities for swap at a fine granular level. In this paper we consider a single superscalar OOO core that can morph itself dynamically into an InO core at runtime. In order to determine when to morph from OOO to InO and vice-versa, we rely on certain hardware performance monitors. Using these performance monitors we estimate the energy-delay-squared product (ED2P) for both modes of operation, which is then used to make morphing decisions. The morphing hardware support is simple and is already available in certain Intel processors to facilitate debug. The proposed scheme has low migration overhead, that enables fine-grain morphing to achieve more energy efficient computing by trading a small loss of performance for much greater energy reduction. Sudarshan Srinivasan, Rance Rodrigues, Arunachalam Annamalai, Israel Koren, Sandip Kundu |
ICCD | 5 |
| 2013 | Managing test coverage uncertainty due to thermal noise in nano-CMOS: A case-study on an SRAM arrayabstractFrom system-on-a-chip to high performance processors, SRAM is a critical component. In highly scaled CMOS devices, process variation is a major concern as it affects SRAM stability which often sets the floor on supply voltage and the ceiling on operating temperature of a semiconductor chip. Consequently, low-voltage and high temperature testing are often part of manufacturing test flow. In this paper, we show that for marginal cells, thermal noise is a major corrupting factor that affects the outcome of testing. A cell with large process variation which should ordinarily fail during memory test may pass due to impact of thermal noise at high temperature. To address this uncertainty during testing, we propose a stochastic metric for test coverage. We also propose application of N-detect and Multi-level Word Line (WL) techniques to improve test coverage based on this stochastic metric. Simulation studies on 32nm PTM models indicate varying probability of faulty bit detection across the spectrum of random thermal noise that lead to erroneous test results. Multiple accesses to each bit cell during test increases the fault coverage from -10% to near ideal 100%. Boosting WL voltage during read test and scaling it below nominal voltage during write test accelerates fault detection. Simulation of a 1KB SRAM array test case shows an improvement in fault coverage from -88% to 100% by increasing the number of detects to 100. Vikram B. Suresh, Sandip Kundu |
ICCD | 2 |
| 2013 | A Study of Tapered 3-D TSVs for Power and Thermal Integrityabstract3-D integration presents a path to higher performance, greater density, increased functionality and heterogeneous technology implementation. However, 3-D integration introduces many challenges for power and thermal integrity due to large switching currents, longer power delivery paths, and increased parasitics compared to 2-D integration. In this work, we provide an in-depth study of power and thermal issues while incorporating the physical design characteristics unique to 3-D integration. We provide a qualitative perspective of the power and thermal dissipation issues in 3-D and study the impact of Through Silicon Vias (TSVs) size for their mitigation. We investigate and discuss the design implications of power and thermal issues in the presence of decoupling capacitors, TSV/on-die/package parasitics, various resonance effects and power gating. Our study is based on a ten-tier system utilizing existing 3-D technology specifications. Based on detailed power distribution and heat dissipation models, we present a comprehensive analysis of TSV tapering for alleviating power and thermal integrity issues in 3-D ICs. Aida Todri, Sandip Kundu, Patrick Girard 0001, Alberto Bosio, Luigi Dilillo, Arnaud Virazel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Scalable Thread Scheduling in Asymmetric Multicores for Power EfficiencyabstractThe emergence of asymmetric multicore processors(AMPs) has elevated the problem of thread scheduling in such systems. The computing needs of a thread often vary during its execution (phases) and hence, reassigning threads to cores(thread swapping) upon detection of such a change, can significantly improve the AMP's power efficiency. Even though identifying a change in the resource requirements of a workload is straightforward, determining the thread reassignment is a challenge. Traditional online learning schemes rely on sampling to determine the best thread to core in AMPs. However, as the number of cores in the multicore increases, the sampling overhead may be too large. In this paper, we propose a novel technique to dynamically assess the current thread to core assignment and determine whether swapping the threads between the cores will be beneficial and achieve a higher performance/Watt. This decision is based on estimating the expected performance and power of the current program phase on other cores. This estimation is done using the values of selected performance counters in the host core. By estimating the expected performance and power on each core type, informed thread scheduling decisions can be made while avoiding the overhead associated with sampling. We illustrate our approach using an 8-core high performance/low-power AMP and show the performance/Watt benefits of the proposed dynamic thread scheduling technique. We compare our proposed scheme against previously published schemes based on online learning and two schemes based on the use of an oracle, one static and the other dynamic. Our results show that significant performance/Watt gains can be achieved through informed thread scheduling decisions in AMPs. Rance Rodrigues, Arunachalam Annamalai, Israel Koren, Sandip Kundu |
SBAC-PAD | 4 |
| 2012 | On Testing Prebond Dies with Incomplete Clock Networks in a 3D IC Using DLLs
Michael Buttrick, Sandip Kundu |
J. Electron. Test. | 2 |
| 2012 | A Pattern Generation Technique for Maximizing Switching Supply Currents Considering Gate DelaysabstractComputation of peak supply current is central to power rail design and analysis of power supply switching noise. Traditionally, peak switching current from all CMOS gates is added together to compute peak supply current. This approach can be improved significantly if temporal and Boolean relationships are taken into consideration. Previously, it was shown that worst case switching current in a subset of gates may imply that some other gates may not have the worst case switching condition due to logical relationship between input patterns of a gate. In this paper, we also take integer gate delays into consideration to show that gate switching events may be spaced out in time leading to lower peak current. Further, it is found that taking gate delays into account actually simplifies the size of individual problem instances to be solved, leading to both a faster and more accurate solution. Finally, we compare peak current waveform generated by the proposed solver against SPICE simulation to demonstrate effectiveness of the proposed solution. Kunal P. Ganeshpure, Alodeep Sanyal, Sandip Kundu |
IEEE Trans. Computers | 3 |
| 2012 | A Wavelet-Based Spatio-Temporal Heat Dissipation Model for Reordering of Program Phases to Produce Temperature Extremes in a ChipabstractLocalized heating leads to generation of thermal hotspots that affect the performance and reliability of an integrated circuit (IC). Functional workloads determine the locations and temperatures of hotspots on a die. In this paper, we present a systematic approach for developing a synthetic workload to maximize the temperature of a target hotspot. Our approach is based on the observation that hotspot temperature is determined not only by the current activity in that region, but also by the past activities in the surrounding regions. Accordingly, we develop a wavelet-based canonical spatio-temporal heat dissipation model for program traces, and use a novel integer linear programming formulation to rearrange program phases to generate target worst case hotspot temperature. Program phase behavior is rooted in the static structure of programs. In this case, the initial set of program phases is extracted from the SPEC 2000 benchmark. We apply this formulation to target another well-known problem of maximizing the temperature between a pair of coordinates in an IC. Experimental results show that by taking the spatio-temporal effect into account, we can raise the temperature of a hotspot higher than what is otherwise possible. Hotspot temperature maximization is important in design verification and testing. Sudarshan Srinivasan, Kunal P. Ganeshpure, Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | Improving performance per watt of asymmetric multi-core processors via online program phase classification and adaptive core morphingabstractAsymmetric multi-core processors (AMPs) have been shown to outperform symmetric ones in terms of performance and performance/watt. Improved performance and power efficiency are achieved when the program threads are matched to their most suitable cores. Since the computational needs of a program may change during its execution, the best thread to core assignment will likely change with time. We have, therefore, developed an online program phase classification scheme that allows the swapping of threads when the current needs of the threads justify a change in the assignment. The architectural differences among the cores in an AMP can never match the diversity that exists among different programs and even between different phases of the same program. Consider, for example, a program (or a program phase) that has a high instruction-level parallelism (ILP) and will exhibit high power efficiency if executed on a powerful core. We can not, however, include such powerful cores in the designed AMP, since they will remain underutilized most of the time, and they are not power efficient when the programs do not exhibit a high degree of ILP. Thus, we must expect to see program phases where the designed cores will be unable to support the ILP that the program can exhibit. We, therefore, propose in this article a dynamic morphing scheme. This scheme will allow a core to gain control of a functional unit that is ordinarily under the control of a neighboring core during periods of intense computation with high ILP. This way, we dynamically adjust the hardware resources to the current needs of the application. Our results show that combining online phase classification and dynamic core morphing can significantly improve the performance/watt of most multithreaded workloads. Rance Rodrigues, Arunachalam Annamalai, Israel Koren, Sandip Kundu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2012 | Test Pattern Generation for Multiple Aggressor Crosstalk Effects Considering Gate Leakage Loading in Presence of Gate DelaysabstractDecreasing process geometries and increasing operating frequencies have made VLSI circuits more susceptible to signal integrity related failures. Capacitive crosstalk on long signal nets is of particular concern. A typical long net is capacitively coupled with multiple aggressors and also tend to have multiple fan-outs. Gate leakage current that originates in fan-out receivers, terminates in the driver causing a shift in driver output voltage. This effect becomes more prominent as gate oxide is scaled more aggressively. Thus, in nano-scale CMOS circuits, noise margin gets eroded by both aggressor crosstalk noise as well as gate leakage loading from fan-outs. In this paper, we present an automatic test pattern generation solution which uses 0-1 integer linear programming to maximize the cumulative voltage noise at a given victim net because of crosstalk and loading in conjunction with propagating the fault effect to an observation point. The target ISCAS benchmark circuits are assumed to have unit gate delays. Results demonstrate both the viability of a solution as well as a need to consider both sources of noise for signal integrity analysis. Pattern pairs generated by this technique are useful for both manufacturing test application as well as signal integrity verification. Alodeep Sanyal, Kunal P. Ganeshpure, Sandip Kundu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Performance Per Watt Benefits of Dynamic Core Morphing in Asymmetric MulticoresabstractThe trend toward multicore processors is moving the emphasis in computation from sequential to parallel processing. However, not all applications can be parallelized and benefit from multiple cores. Such applications lead to under-utilization of parallel resources, hence sub-optimal performance/watt. They may however, benefit from powerful uniprocessors. On the other hand, not all applications can take advantage of more powerful uniprocessors. To address competing requirements of diverse applications, we propose a heterogeneous multicore architecture with a Dynamic Core Morphing (DCM) capability. Depending on the computational demands of the currently executing applications, the resources of a few tightly coupled cores are morphed at runtime. We present a simple hardware-based algorithm to monitor the time-varying computational needs of the application and when deemed beneficial, trigger reconfiguration of the cores at fine-grain time scales to maximize the performance/watt of the application. The proposed dynamic scheme is then compared against a baseline static heterogeneous multicore configuration and an equivalent homogeneous configuration. Our results show that dynamic morphing of cores can provide performance/watt gains of 43% and 16% on an average, when compared to the homogeneous and baseline heterogeneous configurations, respectively. Rance Rodrigues, Arunachalam Annamalai, Israel Koren, Sandip Kundu, Omer Khan |
PACT | 4 |
| 2011 | An Architecture to Enable Lifetime Full Chip Testability in Chip MultiprocessorsabstractSummary form only given. Technology scaling has led to a tremendous increase in the packing density of transistors. However, these small transistors are susceptible to certain impediments that were not present earlier. Manufacturability suffers due to trailing lithography technology which does not scale well with transistor technology. Increased leakage current has reduced effectiveness of burn-in tests. Infant mortality cannot therefore, be completely kept under check. Even during operation, reliability is affected due to CMOS wear-out mechanisms such as time-dependent dielectric breakdown (TDDB), hot carrier injection (HCI), negative bias temperature instability (NBTI), electro migration (EM), and stress induced voiding (SIV). Rance Rodrigues, Israel Koren, Sandip Kundu |
PACT | 3 |
| 2011 | Efficient BDD-based Fault Simulation in Presence of Unknown ValuesabstractUnknown (X) values, originating from memories, clock domain boundaries or A/D interfaces, may compromise test signatures and fault coverage. Classical logic and fault simulation algorithms are pessimistic w.r.t. the propagation of X values in the circuit. This work proposes efficient hybrid logic and stuck-at fault simulation algorithms which combine heuristics and local BDDs to increase simulation accuracy. Experimental results on benchmark and large industrial circuits show significantly increased fault coverage and low runtime. The achieved simulation precision is quantified for the first time. Michael A. Kochte, Sandip Kundu, Kohei Miyase, Xiaoqing Wen, Hans-Joachim Wunderlich |
Asian Test Symposium | 2 |
| 2011 | An Online Mechanism to Verify Datapath Execution Using Existing Resources in Chip MultiprocessorsabstractWith scaling of process technology, transistor and interconnect reliability has emerged as a growing concern for modern microprocessors. Traditional solutions for reliable operation rely on double or triple modular redundancies. However, chip multiprocessors (CMP) provide unique opportunity for low-cost data path verification for reliable operation. A recent paper presents a fault recovery scheme based on outsourcing instructions from identified faulty cores to fault free cores capable of executing them. The communication between the cores is managed via an inter-core queue (ICQ). However, no faulty core identification mechanism was presented. In this paper, we extend this research to enable self-test of the data path execution in a multicore processor. Specifically, whenever instructions are retired locally on a core (local), they are also dispatched for execution on another nearby (remote) core for execution verification via ICQ. Results obtained from local and remote cores are compared. If a fault is detected, the instruction may be re-executed on both local and remote cores to distinguish between hard and soft faults. In this study, we present results on frequency of coverage and latency between first execution and its verification. We also report performance impact of execution verification on the remote core. Results indicate that the proposed scheme is capable of remotely verifying ~80% integer ALU instructions and >;98% of other instruction types with very small impact on performance of just ~1% on the tester core and incurs less than 1% area overhead. Rance Rodrigues, Sandip Kundu |
Asian Test Symposium | 2 |
| 2011 | On testing prebond dies with incomplete clock networks in a 3D IC using DLLsabstract3D integration of ICs is an emerging technology where multiple silicon dies are stacked vertically. The manufacturing itself is based on wafer-to-wafer bonding, die-to-wafer bonding or die-to-die bonding. Wafer-to-wafer bonding has the lowest yield as a good die may be stacked against a bad die, resulting in a wasted good die. Thus the latter two options are preferred to keep yield high and manufacturing costs low. However, these methods require dies to be tested separately before they are stacked. A problem with testing dies separately is that the clock network of a prebond die may be incomplete before stacking. In this paper we present a solution to address this problem. The solution is based on on-die DLL implementations that are only activated during testing prebond unstacked dies to synchronize disconnected clock regions. A problem with using DLLs in testing is that they cannot be turned on or off within a single cycle. Since scan-based testing requires that test patterns be scanned in at a slow clock frequency before fast capture clocks are applied [1], on-product clock generation (OPCG) must be used. The proposed solution addresses the above problems. Furthermore, we show that a higher-speed DLL is better suited to not only high frequency system clocks, but lower power as well due to a smaller variable delay line. Michael Buttrick, Sandip Kundu |
DATE | 2 |
| 2011 | Modeling manufacturing process variation for design and testabstractFor process nodes 22nm and below, a multitude of new manufacturing solutions have been proposed to improve the yield of devices being manufactured. With these new solutions come an increasing number of defect mechanisms. There is a need to model and characterize these new defect mechanisms so that (i) ATPG patterns can be properly targeted, (ii) defects can be properly diagnosed and addressed at design or manufacturing level. This presentation reviews currently available defect modeling and test solutions and summarizes open issues faced by the industry today. It also explores the topic of creating special test structures to expose manufacturing process parameters which can be used as input to software defect models to predict die specific defect locations for better targeting of test. Sandip Kundu, Aswin Sreedhar |
DATE | 1 |
| 2011 | On design of test structures for lithographic process corner identificationabstractLithographic process variations, such as changes in focus, exposure, resist thickness introduce distortions to line shapes on a wafer. Large distortions may lead to line open and bridge faults and the locations of such defects vary with lithographic process corner. Based on lithographic simulation, it is easily verified that for a given layout, changing one or more of the process parameters shifts the defect location. Thus, if the lithographic process corner of a die is known, test patterns can be better targeted for both hard and parametric defects. In this paper, we present design of control structures such that preliminary testing of these structures can uniquely identify the manufacturing process corner. If the manufacturing process corner is known, we can easily attain highest possible fault coverage for lithography related defects during manufacturing test. Parametric defects such as delay defects are notorious to test because such defects may affect paths that are subcritical under nominal conditions and not ordinarily targeted for test. Adoption of the proposed approach can easily flag such paths for delay tests. Aswin Sreedhar, Sandip Kundu |
DATE | 2 |
| 2011 | Physically unclonable functions for embeded security based on lithographic variationabstractPhysically unclonable functions (PUF) are designed on integrated circuits (IC) to generate unique signatures that can be used for chip authentication. PUFs primarily rely on manufacturing process variations to create distinction between chips. In this paper, we present novel PUF circuits designed to exploit inherent fluctuations in physical layout due to photolithography process. Variations arising from proximity effects, density effects, etch effects, and non-rectangularity of transistors is leveraged to implement lithography-based physically unclonable functions (litho-PUFs). We show that the uniqueness level of these PUFs are adjustable and are typically much higher than traditional ring-oscillator or tri-state buffer based approaches. Aswin Sreedhar, Sandip Kundu |
DATE | 2 |
| 2011 | Stress aware switching activity driven low power design of critical paths in nanoscale CMOS circuitsabstractRecent research on low power circuits make use of dual VT and strained silicon technologies. Low VT assignment on performance critical paths improves performance at the expense of leakage current. Strained silicon devices on the other hand have greater leakage but provide much larger ON current. The increased drive strength can be used to reduce gate sizes. In this work we show that selective use of strained silicon gates in high-activity critical paths reduce overall power due to decreased dynamic power arising from smaller gate sizes, that more than offset increased leakage due to strained silicon gate. Similarly, we find that up-sizing gates in low-activity critical paths reduce total power. Even though up-sized gates dissipate more dynamic power, lower switching activity, coupled with lower leakage compared to low VT devices make this a better choice. An optimization algorithm based on slack distribution between low and high switching activity gates is presented. Simulation results show an overall power benefit of 5.4% while improving performance by 15% even though overall leakage is 1% higher than original circuit. Sudarshan Srinivasan, Bharath Phanibhushana, Arunkumar Vijayakumar, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 4 |
| 2011 | Task model for on-chip communication infrastructure design for multicore systemsabstractWith technology scaling, Multiprocessor System on Chip (MPSoC) which consist of multiple processors connected via a Network on Chip (NoC) have become prevalent. Applications are mapped to MPSoC's by representing it in the form of a task graph. Task scheduling involves mapping task to processor cores so as to meet the deadline. For a given deadline, slack at each node is defined by the amount of time by which a task execution can be delayed without missing the deadline. With increase in the number of cores and high application parallelism, NoC is becoming a bottleneck due to the presence of large number of concurrent communications. Increasing network resources (links and routers) reduces the communication time but the area and power goes up. In this paper we present an application aware heuristic to synthesize a minimal network connecting a set of cores in an MPSoC in the presence of hard deadlines. Our approach is based on modeling communication between a pair of processors as tasks known as “Comtasks”. The network is generated by “scheduling” these comtasks onto a set of routers so as to obtain a network with minimum area which fills up the available slack. Moreover, we also identify the set of overlapping comtasks and generate a minimal network to allow the maximum set of overlapping comtasks to execute concurrently. We compared our approach with a greedy network generation heuristic and the results show 80% benefit in the router area. Bharath Phanibhushana, Kunal P. Ganeshpure, Sandip Kundu |
ICCD | 3 |
| 2011 | On graceful degradation of microprocessors in presence of faults via resource bankingabstractReliability and manufacturability have emerged as dominant concerns for today's multi-billion transistor chips. In this paper, we investigate how to degrade a chip multiprocessor (CMP) gracefully in presence of faults, by keeping its architected functionality intact at the expense of some loss of performance. The proposed solution involves banking resources and functional units that allow partial shutdown to isolate faulty regions. Within a processor chip, resources can be classified into three broad classes, namely the large memory structures such as caches, TLB, register file, the small memory structures namely the reorder buffer, issue queue, load-store buffer etc. and the datapath and control logic. The large arrays are usually well protected by ECC. Recent research has suggested that large datapath units such as FPU and integer division units are good candidates for execution outsourcing to other working cores in CMP. In this paper, we focus on relatively small but critically important integer ALU unit and small array structures. Outsourcing ALU operations incur large performance penalty and small arrays are critical for control operations. The proposed solution is based on banking of array structures and execution units as well as reducing instruction fetch, issue and retire widths that allow individual bank(s) to be disabled for operation. Reduction in resource size diminishes performance but retains functionality. In this paper we present performance degradation data for shutting down one or more banks for various small array structures and ALU units. Simulation confirms gradual degradation with diminishing resource sizes. On average over all considered structures, performance loss of just 6% was observed for single bank failures. Rance Rodrigues, Sandip Kundu |
IOLTS | 2 |
| 2011 | On graceful degradation of chip multiprocessors in presence of faults via flexible pooling of critical execution unitsabstractReliability and manufacturability have emerged as dominant concerns for today's multi-billion transistor chips. In this paper, we investigate how to degrade a chip multiprocessor (CMP) gracefully in presence of faults, by keeping its architected functionality intact at the expense of some loss of performance. The proposed solution involves sharing critical execution resources among cores to survive faults. Recent research has suggested that large datapath units such as FPU and integer division units are good candidates for execution outsourcing to other working cores in CMP. In this paper, we focus on relatively small but critically important integer ALU unit. Outsourcing ALU operations incur large performance penalty and better solutions need to be in place to ensure survivability with minimal performance loss. We propose the provisioning of a shared ALU among a set of cores that can act as a spare for any constituent core in the group. This solution works well for single ALU failures, but leads to resource contention when multiple ALUs fail. Simulation case studies on MediaBench and MiBench benchmarks show that the proposed solution allows the CMP to remain functionally intact with no performance penalty for single ALU failures and no more than 1.5% performance loss on average for failure of single ALU in each core. Rance Rodrigues, Sandip Kundu |
IOLTS | 2 |
| 2011 | Lithography aware critical area estimation and yield analysisabstractAdvancing technology nodes have increased the impact of lithography variation on design yield and performance. Although lithographic distortion is emerging as the dominant mode of failure, particulate defect still remains an important source of defect. In this work, we present a novel critical area analysis technique that deals with non-deterministic line edges that arise due to statistical lithography variations. Previous critical area analysis techniques are based on fixed printed structures with variable defect sizes. The proposed solution integrates lithographic variations within critical area analysis, using available tools while keeping the computational complexity in check. The key optimizations include weighted stratified sampling of important lithographic parameters and Monte Carlo based Probability of Failure (POF) calculation that is scalable to large chips. Simulation results on ISCAS benchmarks show that inclusion of lithography induced physical variations can increase critical area by as much as 2X, while overall yield due to particulate defects in large designs may decrease by more than 5%; validating both the need for an integrated solution and its feasibility on large designs. Priyamvada Vijayakumar, Vikram B. Suresh, Sandip Kundu |
ITC | 3 |
| 2011 | Hardware/Software Codesign Architecture for Online Testing in Chip MultiprocessorsabstractAs the semiconductor industry continues its relentless push for nano-CMOS technologies, long-term device reliability and occurrence of hard errors have emerged as a major concern. Long-term device reliability includes parametric degradation that results in loss of performance as well as hard failures that result in loss of functionality. It has been reported in the ITRS roadmap that effectiveness of traditional burn-in test in product life acceleration is eroding. Thus, to assure sufficient product reliability, fault detection and system reconfiguration must be performed in the field at runtime. Although regular memory structures are protected against hard errors using error-correcting codes, many structures within cores are left unprotected. Several proposed online testing techniques either rely on concurrent testing or periodically check for correctness. These techniques are attractive, but limited due to significant design effort and hardware cost. Furthermore, lack of observability and controllability of microarchitectural states result in long latency, long test sequences, and large storage of golden patterns. In this paper, we propose a low-cost scheme for detecting and debugging hard errors with a fine granularity within cores and keeping the faulty cores functional, with potentially reduced capability and performance. The solution includes both hardware and runtime software based on codesigned virtual machine concept. It has the ability to detect, debug, and isolate hard errors in small noncache array structures, execution units, and combinational logic within cores. Hardware signature registers are used to capture the footprint of execution at the output of functional modules within the cores. A runtime layer of software (microvisor) initiates functional tests concurrently on multiple cores to capture the signature footprints across cores to detect, debug, and isolate hard errors. Results show that using targeted set of functional test sequences, faults can be debugged to a fine-granular level within cores. The hardware cost of the scheme is less than three percent, while the software tasks are performed at a high-level, resulting in a relatively low design effort and cost. Omer Khan, Sandip Kundu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2010 | A model to exploit power-performance efficiency in superscalar processors via structure resizingabstractPower consumption has become a major cause of concern spanning from data centers to handheld devices. Traditionally, improvement in power-performance efficiency of a modern superscalar processor came from technology scaling. However, that is no longer the case. Many of the current systems deploy coarse grain voltage and/or frequency scaling for power management. These techniques are attractive, but limited due to their granularity of control and effectiveness in nano-CMOS technologies. This paper proposes a novel architecture-level mechanism to exploit intra thread variations for power-performance efficiency in modern superscalar processors. This class of processors implements several buffering/queuing structures to support speculative out-of-order execution for performance enhancement. Applications may not need full capabilities of such structures at all times. A mechanism that collaboratively adapts a finite set of key hardware structures to the changing program behavior can allow the processor to operate with heterogeneous power-performance capabilities. We present a novel offline regression based empirical model to estimate structure resizing for a selected set of structures. It is shown that using a few processor runtime events, the system can dynamically estimate structure resizing to exploit power-performance efficiency. Results show that using the proposed empirical model, a selective set of key structures can be resized at runtime to deliver 35% power-performance efficiency over a baseline design, with only 5% loss of performance. Omer Khan, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | A self-adaptive scheduler for asymmetric multi-coresabstractAsymmetric chip multiprocessors are imminent in the multi-core era primarily due their potential for power-performance efficiency. In order for software to fully realize this potential, the scheduling of threads to cores must be automated to adapt to the changing program behavior. However, strict system abstraction layers limit the controllability and observability of low level hardware details, thereby, limiting the state-of-the-art systems to rely on manual or static mapping of threads to cores in an asymmetric multi-core. In this paper, we propose a self-adaptive scheduler that exploits program behavior at runtime by matching computational demands of threads to the capabilities of cores. We present a novel empirical model to predict the selection of an appropriate core (based on optimizing throughput, power or performance per watt) for changing program phases within threads. Thread migration is initiated when an optimal mapping of threads to cores is predicted. Results show that our predictive schedulers for the three target optimizations are within 10% of the ideal scheduler. Omer Khan, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | A mask double patterning technique using litho simulation by wavelet transformabstractOptical Lithography is a key to semiconductor device scaling. As technology continues to scale, the fundamental limits of lithography are pushed to an extreme. Today 193nm light is used to print features in 45nm technology node. As the minimum feature size on the mask is less than half the wavelength of light used for the lithography process, diffraction at the mask edges dominates the errors in mask printing. Mask transfer fidelity issues are countered using a number of Resolution Enhancement Techniques (RET) which include Optical Proximity Correction (OPC), Phase Shift Masking (PSM), Sub Resolution Assist Features (SRAF) and Dual Patterning Lithography (DPL). DPL reduces wafer throughput but has become a necessity in current and upcoming technology nodes. It involves splitting patterns in a mask into two masks that are exposed separately. DPL CAD problem is a pattern coloring problem to minimize mask edge placement error (EPE). EPE results from interaction of near field waves and any geometric solution that does not consider interaction of fields, suffers from inaccuracies. Previous publications were mostly focused on a rule based geometric solution. In this paper we investigate a method to implement DPL using fast lithography simulation taking into account not only the bad, but also the beneficial effects of having polygons in the mask close to one another in the final partitioned layout. We present results on metal layer 2 of the ISCAS-85 benchmarks. Results show that even though our model based solution is slower, unlike many previous approaches, the final output meets the objectives of reducing EPE. Rance Rodrigues, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | TURBONFS: turbo nand flash searchabstractNAND flash memories are popular due to their density and lower cost. However, due to serial access, NAND flash memories have slow read and write speeds. As the flash sizes increase to 64GB and beyond, searches through flash memories become painfully slow. In this paper, we present a hardware design enhancement to speed-up search through flash memories. The basic idea is to generate a small signature for every memory block and store them in a signature block(s). When a search is initiated, signature block is searched which produces reference of possible blocks where data might be contained, reducing the total number of read operations. The additional hardware has no impact on read access times or sequential write times. The additional step of signature generation increases the random write times by an average of 8-9%. Simulation experiments were performed for flash memory of size up to 16Gb. Simulation results show that the performance of searches improve by 4000X by using the proposed technique. The area overhead is estimated to be 5% of the memory area. The approach is easily extendible to NOR flashes and SSDs. Shruti Vyas, Aswin Sreedhar, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 3 |
| 2010 | A study on performance benefits of core morphing in an asymmetric multicore processorabstractMulticore architectures are designed so as to provide an acceptable level of performance per unit power for the majority of applications. Consequently, we must occasionally expect applications that could have benefited from a more powerful core in terms of either lower execution time and/or lower energy consumed. Fusing some of the resources of two (or more) cores to configure a more powerful core for such instances is a natural approach to deal with those few applications that have very high performance demands. However, a recent study has shown that fusing homogeneous cores is unlikely to benefit applications. In this paper we study the potential performance benefits of core morphing in a heterogeneous multicore processor that can be reconfigured at runtime. We consider as an example a dual core processor with one of the two cores being designed to target integer intensive applications while the other is better suited to floating-point intensive applications. These two cores can be fused into a single powerful core when an application that can benefit from such fusion is executing. We first discuss the design principles of the two individual cores so that the majority of the benchmarks that we consider execute in a satisfactory way. We then show that a small subset of the considered applications can greatly benefit from core morphing even in the case where two applications that could have been executed in parallel on the two cores are run, for some percentage of time, on the single morphed core. Our results indicate that a performance gain of up to 100% is achievable at a small hardware overhead of less than 1%. Anup Das 0001, Rance Rodrigues, Israel Koren, Sandip Kundu |
ICCD | 4 |
| 2010 | Modeling the impact of process variation on resistive bridge defectsabstractRecent research has shown that tests generated without taking process variation into account may lead to loss of test quality. At present there is no efficient device-level modeling technique that models the effect of process variation on resistive bridges. This paper presents a fast and accurate technique to model the effect of process variation on resistive bridge defects. The proposed model is implemented in two stages: firstly, it employs an accurate transistor model (BSIM4) to calculate the critical resistance of a bridge; secondly, the effect of process variation is incorporated in this model by using three transistor parameters: gate length (L), threshold voltage (Vth) and effective mobility (μeff), where each follow Gaussian distribution. Experiments are conducted on a 65-nm gate library (for illustration purposes), and results show that on average the proposed modeling technique is more than 7 times faster and in the worst case, error in bridge critical resistance is 0.8% when compared with HSPICE. S. Saqib Khursheed, Shida Zhong, Robert C. Aitken, Bashir M. Al-Hashimi, Sandip Kundu |
ITC | 5 |
| 2010 | Shadow checker (SC): A low-cost hardware scheme for online detection of faults in small memory structures of a microprocessorabstractAt various stages of a product life, faults arise from different sources. During product bring up, logic errors are dominant. During production, manufacturing defects are main concerns while during operation, the concern shifts to aging defects. No matter what the source is, debugging such defects may permit logic, circuit or physical design changes to eliminate them in future. Within a processor chip, there are three broad categories of structures, namely the large memory structures such as caches, small memory structures such as reorder buffer, issue queue, and load-store buffers and the data-path. Most control functions and data steering operations are based on small memory structures and they are hard to debug. In this paper, we propose a lightweight hardware scheme, called shadow checker to detect faults in these critical units. The entries in these units are tested by means of a shadow entry that mimics intended operation. A mismatch traps an error. The shadow checker shadows an entry for a few thousand cycles before moving on to shadow another. This scheme can be employed to test chips during silicon debug, manufacturing test as well as during regular operation. We ran experiments on 13 SPEC2000 benchmarks and found that our scheme detects 100% of inserted faults. Rance Rodrigues, Sandip Kundu, Omer Khan |
ITC | 2 |
| 2010 | Thread Relocation: A Runtime Architecture for Tolerating Hard Errors in Chip MultiprocessorsabstractAs the semiconductor industry continues its relentless push for nano-CMOS technologies, device reliability and occurrence of hard errors have emerged as a dominant concern in multicores. Although regular memory structures are protected against hard errors using error correcting codes or spare rows and columns, many of the structures within the cores are left unprotected. Even if the location of hard errors is known a priori, disabling faulty cores results in a substantial performance loss. Several proposed techniques use microarchitectural redundancy to allow defective cores to continue operation. These techniques are attractive, but limited due to either added cost of additional redundancy that offers no benefits to an error-free core, or limited coverage, due to the natural redundancy offered by the microarchitecture. We propose to exploit the intercore redundancy in chip multiprocessors for hard-error tolerance. Our scheme combines hardware reconfiguration to ensure reduced functionality of cores, and a runtime layer of software (microvisor) to manage mapping of threads to cores. Microvisor observes the changing phase behavior of threads and initiates thread relocation to match the computational demands of threads to the capabilities of cores. Our results show that in the presence of degraded cores, microvisor mitigates performance losses by an average of two percent. Omer Khan, Sandip Kundu |
IEEE Trans. Computers | 2 |
| 2010 | An Efficient Technique for Leakage Current Estimation in Nanoscaled CMOS Circuits Incorporating Self-Loading EffectsabstractWith the scaling of CMOS technology, subthreshold, gate, and reverse biased junction band-to-band-tunneling leakage have increased dramatically. Together, they account for more than 25 percent of power consumption in the current generation of leading edge designs. Different sources of leakage can affect each other by interacting through resultant intermediate node voltages. This is called the loading effect. In this paper, we propose a pattern dependent steady-state leakage estimation technique that incorporates loading effect and accounts for all three major leakage components, namely the gate, band-to-band-tunneling, and subthreshold leakage and accounts for transistor stack effect. By observing a recursive relationship between gate leakage and loading effect, we further refine our leakage estimation technique by developing a compact leakage model that supports iteration over node voltages based on Newton-Raphson method. The proposed estimation technique based on the compact model improves performance and capacity over SPICE. We report a speedup of 18,000X over SPICE simulation on smaller circuits, where SPICE simulation is feasible. Results also show that loading effect is a significant factor in leakage and worsens with technology scaling. Alodeep Sanyal, Ashesh Rastogi, Sandip Kundu |
IEEE Trans. Computers | 4 |
| 2010 | On ATPG for Multiple Aggressor Crosstalk FaultsabstractCrosstalk faults have emerged as a significant mechanism of circuit failure due to decreasing process geometries and increasing operation frequencies. Long signal nets are highly susceptible to crosstalk faults because they tend to have a higher coupling capacitance to overall capacitance ratio. Moreover, a typical long net also has multiple aggressors. In generating patterns to create maximal crosstalk induced delay on a victim net, it may be impossible to activate all aggressors logically or simultaneously to constructively induce maximum noise at the victim. Therefore, pattern generation must focus on activating a maximal subset of aggressors, weighted by actual coupling capacitance value, in close temporal proximity of the victim net transition. This max-satisfiability problem is constrained by fault effect propagation condition which involves determining an input signal assignment so as to propagate the fault effect at the victim to the primary output. In this paper, we present Automatic Test Pattern Generation (ATPG) solutions for multiple aggressor crosstalk faults for zero and unit delay models and compare the magnitude of crosstalk induced delay at the victim net. Our solution involves a combination of 0-1 Integer Linear Programming (ILP), for maximal aggressor excitation. Fault effect propagation is solved independently by using traditional stuck-at fault ATPG or by generating additional ILP constraints thus forming a integrated ILP formulation with error propagation. The effect of gate delays is summed by circuit transformation. The proposed technique was applied to ISCAS85 benchmark circuits. Results indicate that the percentage of total capacitance that can be switched varies from 75-100% for zero delay and 30-80% for variable delay case while achieving propagation of the fault effect to primary output. Kunal P. Ganeshpure, Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | A self-adaptive system architecture to address transistor agingabstractAs semiconductor manufacturing enters advanced nanometer design paradigm, aging and device wear-out related degradation is becoming a major concern. Negative Bias Temperature Instability (NBTI) is one of the main sources of device lifetime degradation. The severity of such degradation depends on the operation history of a chip in the field, including such characteristics as temperature and workloads. In this paper, we propose a system level reliability management scheme where a chip dynamically adjusts its own operating frequency and supply voltage over time as the device ages. Major benefits of the proposed approach are (i) increased performance due to reduced frequency guard banding in the factory and (ii) continuous field adjustments that take environmental operating conditions such as actual room temperature and the power supply tolerance into account. The greatest challenge in implementing such a scheme is to perform calibration without a tester. Much of this work is performed by a hypervisor like software with very little hardware assistance. This keeps both the hardware overhead and the system complexity low. This paper describes the entire system architecture including hardware and software components. Our simulation data indicates that under aggressive wear-out conditions, scheduling interval of days or weeks is sufficient to reconfigure and keep the system operational, thus the run time overhead for such adjustments is of no consequence at all. Omer Khan, Sandip Kundu |
DATE | 2 |
| 2009 | Hardware/software co-design architecture for thermal management of chip multiprocessorsabstractThe sustained push for performance, transistor count, and instruction level parallelism has reached a point where chip level power density issues are at the forefront of design constraints. Many high performance computing platforms are integrating several homogeneous or heterogeneous processing cores on the same die to fit small form factors. Due to the design limitations of using expensive cooling solutions, such complex chip multiprocessors require an architectural solution to mitigate thermal problems. Many of the current systems deploy Dynamic Voltage and Frequency Scaling (DVFS) to address thermal emergencies, either within the Operating System or hardware. These techniques have certain limitations in terms of response lag, scalability, cost and being reactive. In this paper, we present an alternative thermal management system to address these limitations, based on hardware/software co-design architecture. The results show that in the 65 nm technology, a predictive, targeted, and localized response to thermal events improves a quad-core performance by an average of 50% over conventional chip-level DVFS. Omer Khan, Sandip Kundu |
DATE | 2 |
| 2009 | A study on placement of post silicon clock tuning buffers for mitigating impact of process variationabstractOptical shrink for process migration, manufacturing process variation, temperature and voltage changes lead to clock skew as well as path delay variations in a manufactured chip. Such variations end up degrading the performance of manufactured chips. Since, such variations are hard to predict in pre-silicon phase, tunable clock buffers have been used in several designs. These buffers are tuned to improve maximum operating clock frequency of a design. Previously, we have presented an algorithmic approach that uses delay measurements on a few selected patterns to determine which buffers should be targeted for tuning. In this paper, a study on impact of tunable buffer placement on performance is reported. Greatest benefit from tunable buffer placement is observed, when the clock tree is designed by the proposed tuning system assuming random delay perturbations during design. Accordingly, we present a clock tree synthesis procedure which offer very good protection against process variation as borne out by the results. Kelageri Nagaraj, Sandip Kundu |
DATE | 2 |
| 2009 | Improving yield and reliability of chip multiprocessorsabstractAn increasing number of hardware failures can be attributed to device reliability problems that cause partial system failure or shutdown. In this paper we propose a scheme for improving reliability of a homogeneous chip multiprocessor (CMP) that also serves to improve manufacturing yield. Our solution centers on exploiting the natural redundancy that already exists in multi-core systems by using services from other cores for functional units that are defective in a faulty core. A micro-architectural modification allows a core on a CMP to use another core as a coprocessor to service any instruction that the former cannot execute correctly. This service is accessed to improve yield and reliability, but at the cost of some loss of performance. In order to quantify this loss we have used a cycle-accurate simulator to simulate the performance of a dual-core system with one or two cores sustaining partial failure. Our results indicate that when a large and sparingly-used unit such as a floating point arithmetic unit fails in a core, even for a floating point intensive benchmark, we can continue to run each faulty core with help from companion cores with as little as 10% impact to performance and less than 1% area overhead. Abhisek Pan, Omer Khan, Sandip Kundu |
DATE | 3 |
| 2009 | On linewidth-based yield analysis for nanometer lithographyabstractLithographic variability and its impact on printability is a major concern in today's semiconductor manufacturing process. To address sub-wavelength printability, a number of resolution enhancement techniques (RET) have been used. While RET techniques allow printing of sub-wavelength features, the feature width itself becomes highly sensitive to process parameters, which in turn detracts from yield due to small perturbations in manufacturing parameters. Yield loss is a function of random variables such as depth-of-focus and exposure dose. In this paper, we present a first order canonical dose/focus model that takes into account both the correlated and independent randomness of the effects of lithographic variation. A novel tile-based yield estimation technique for a given layout, based on a statistical model for process variability is presented. Another novel contribution of this paper is the computation of global and local line-yield probabilities. The key issues addressed in this paper are (i) layout error modeling, (ii) avoidance of mask simulation for chip layouts, (iii) avoidance of full Monte-Carlo simulation for variational lithography modeling, (iv) building a methodology for yield estimation based on existing commercial tools. Numerical results based on our approach are shown for 45nm ISCAS85 layouts. Aswin Sreedhar, Sandip Kundu |
DATE | 2 |
| 2009 | A process variation tolerant self-compensating FinFET based sense amplifier designabstractWith the emerging nanoscale devices, SIA roadmap identifies FinFET as a candidate for post-planar end-of-roadmap CMOS device. Lithography related CD variations, fluctuations in dopant density, oxide thickness and parametric variations of devices are identified as a major challenge to the classical bulk type MOSFET in ITRS. Yield loss due to device and process variation has never been so critical to cause failure in circuits. Due to growth in size of embedded SRAMs as well as usage of sense amplifier based signaling techniques, process variation in sense amplifiers lead to significant loss of yield. In this paper, we present a FinFET based Process Variation Tolerant Sense Amplifier design, that exploits the backgate of FinFET devices for dynamic compensation against process variation. Results from statistical simulation show that the proposed dynamic compensation is highly effective in restoring yield at a level comparable to that of sense amplifiers without significant process variations. Aarti Choudhary, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | Reducing temperature variability by routing heat pipesabstractA significant increase in power density in modern nano-electronic VLSI circuits has lead to increased localized heating and generation of hot spots. These temperature effects can lead to reliability and performance problems. This paper presents a novel design time temperature aware methodology which consists of using additional routing known as Heat Pipes, to transfer heat from hot to cold regions. In order to evaluate the effect of Heat Pipes, a thermal model to simulate effect of metal interconnect on heat distribution is also developed. Results show a 5% to 7% decrease in temperature variation through-out and 2 to 3 degree reduction in hotspot temperature as a result of Heat Pipes. Kunal P. Ganeshpure, Ilia Polian, Sandip Kundu, Bernd Becker 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Process variation mitigation via post silicon clock tuningabstractManufacturing process corner, voltage and temperature (PVT) conditions lead to variation in path delays and clock skews. Such variations end up degrading the performance of manufactured chips. Since, such variations are hard to predict in pre-silicon phase, tunable clock buffers have been used in several microprocessor designs. These buffers are tuned to maximize operating clock frequency of a design. In this paper, we report a study on using measured delays on selected patterns to determine which buffers should be targeted for tuning. Based on statistical simulation studies, it is found that the proposed approach can improve clock frequency by as much as 9% in the average case. Kelageri Nagaraj, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | On process variation tolerant low cost thermal sensor design in 32nm CMOS technologyabstractThermal management has emerged as an important design issue in a range of designs from portable devices to server systems. Internal thermal sensors are an integral part of such a management system. Process variations in CMOS circuits cause accuracy problems for thermal sensors which can be fixed by calibration tables. Stand-alone thermal sensors are calibrated to fix such problems. However, calibration requires going through temperature steps in a tester, increasing test application time and cost. Consequently, calibrating thermal sensors in typical digital designs including mainstream desktop and notebook processors increases the cost of the processor. This creates a need for design of thermal sensors whose accuracy does not vary significantly with process variations. Other qualities desired from thermal sensors include low area requirement so that many of them maybe integrated in a design as well as low power dissipation, such that the sensor itself does not become a significant source of heat. In this paper, we present a process variation tolerant thermal sensor design with (i) active compensation circuitry and (ii) signal dithering based self calibration technique to meet the above requirements in 32nm technology. Results show that we achieve ±3ºC temperature accuracy, with a relatively small design. This compares well with designs that are currently used. Spandana Remarsu, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | A study on impact of aggressor de-rating in the context of multiple crosstalk effects in circuitsabstractCapacitive crosstalk induced signal integrity effects have been studied for over a decade. A typical victim net has multiple aggressors. In worst-case analysis of crosstalk effects, it is customary to assume that (i) all aggressors can switch at the same time and (ii) aggressors themselves are not subject to other crosstalk effects. Further refinements of the worst-case analysis consider (a) Boolean filtering of the aggressors to take logical relationship among them into account and (b) timing filtering to exclude aggressors that cannot switch in the same timing window where the victim node is switching. However, even further refinement is possible by relaxing supposition (ii) above that assumes that aggressors are not subject to noise themselves. In this paper, we present a simulation study that considers multiple crosstalk effects where the aggressors of one net can be victim themselves with signals switching in their neighborhood. The simulations are performed on an innovative compact model that permits circular reasoning in an event-driven, non-zero gate delay, and dynamic simulation framework. Results indicate that when crosstalk on aggressors is also considered while processing crosstalk on a victim, the impact is often mitigated substantially. Alodeep Sanyal, Abhisek Pan, Sandip Kundu |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | Predictive Thermal Management for Chip Multiprocessors Using Co-designed Virtual Machines
Omer Khan, Sandip Kundu |
HiPEAC | 2 |
| 2009 | Optical lithography simulation using wavelet transformabstractOptical lithography is an indispensible step in the process flow of design for manufacturability (DFM). Optical lithography simulation is a compute intensive task and simulation performance, or lack thereof can be a determining factor in time to market. Thus, the efficiency of lithography simulation is of paramount importance. Coherent decomposition is a popular simulation technique for aerial imaging simulation. In this paper, we propose an approximate simulation technique based on the 2D wavelet transform and use a number of optimization methods to further improve polygon edge detection. Results show that the proposed method suffers from an average error of less than 5% when compared with the coherent decomposition method. The benefits of the proposed method are (i) >10X increase in performance and more importantly (ii) it allows very large circuits to be simulated while some commercial tools are severely capacity limited. Approximate simulation is quite attractive for layout optimization where it may be used in a loop and may even be acceptable for final layout verification. Rance Rodrigues, Aswin Sreedhar, Sandip Kundu |
ICCD | 3 |
| 2009 | Statistical timing analysis based on simulation of lithographic processabstractThe length of poly-gate printed on silicon depends on exposure dose, depth of focus, photo-resist thickness and planarity of the surface. In sub-wavelength lithography, polygate length also varies with layout topology. Poly-gate length determines the effective channel length of a transistor, which determines its performance. Since the sources of error are hard to control, statistical analysis can be used to measure the impact on circuit timing characteristics. Typical lithography-aware methodologies consider only systematic variation such as across chip linewidth variation (ACLV). In this paper we propose a statistical technique for timing yield prediction, based on variational lithography modeling of physical circuit layout. By statistically varying lithographic process parameters we estimate the difference in timing yield estimation of a design. Our simulation results show that if manufacturing process parameters follow a Gaussian distribution, resulting transistors follow a skewed normal distribution, where a greater number of them will have shorter channel length. This led us to investigate whether statistical static timing analysis (SSTA) is overly pessimistic. The baseline delay model assumed for SSTA in out approach is a Gaussian delay model fitted to skew normal distribution data obtained from statistical litho simulation. Our experiments showed that even after re-centering Gaussian delay model to fit the channel length data with minimum error, it is still overly pessimistic and significantly underestimates circuit performance. Aswin Sreedhar, Sandip Kundu |
ICCD | 2 |
| 2009 | An Improved Soft-Error Rate Measurement TechniqueabstractSoft errors caused by ionizing radiation have emerged as a major concern for current generation of CMOS technologies, and the trend is expected to get worse. The measurement unit for failures due to soft errors is failure in time (FIT) that represents the number of failures encountered per billion hours of device operation. FIT rate measurement is time consuming and calls for accelerated testing. To improve effectiveness of soft-error rate (SER) testing, the patterns must be targeted toward detecting node failures that are most likely. In this paper, we present a technique for identifying soft-error-susceptible sites based on efficient electrical analysis that treats soft errors as Boolean errors but uses analog strengths to decide whether such errors can propagate to the next stage. Next, we present pattern generation techniques for manifestable soft errors such that each pattern targets a maximal set of soft errors. These patterns maximize the likelihood of detecting a soft error when it occurs. The pattern generators target scan architecture. It is well known that scan test time is dominated by scan shifts, when no useful testing is being done. To improve efficiency of scan-based testing, we extend the functionality of the existing built-in logic block observation (BILBO) architecture to support test-per-clock operation. Such targeted pattern generation and test application improve SER characterization time by an order of magnitude. Alodeep Sanyal, Kunal P. Ganeshpure, Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | A Design-for-Debug (DfD) for NoC-Based SoC Debugging via NoCabstractThis paper presents design-for-debug (DfD) methods for the reuse of network-on-chip (NoC) as a debug data path in an NoC-based system-on-chip (SoC). We propose on-chip core debug supporting logics which can support transaction-based debug. A debug interface unit is also presented to enable debug data transfer through an NoC between an external debugger and a core-under-debug (CUD). The proposed approach supports debug of designs with multiple clock domains. It also supports collection of trace signatures to facilitate debug of long pattern sequences. Experimental results show that single and multiple stepping through transactions are feasible with moderately low area overhead. We also present simulation result to verify proper operation of the debug components. Hyunbean Yi, Sungju Park, Sandip Kundu |
ATS | 3 |
| 2008 | On Modeling and Testing of Lithography Related Open Faults in Nano-CMOS CircuitsabstractScaling of transistor feature size over time has been facilitated by corresponding improvement in lithography technology. However, in recent times the wavelength of the optical light source used for photolithography has not scaled in the same rate as that of the minimum feature size of the transistor. In fact, starting with 180 nm devices, the wavelength of optical source has remained the same (at 193 nm) due to difficulties in finding a flicker-free, high energy, coherent light source with compatible improvement in lens material for focusing this light. Consequently, upcoming technology nodes (65 nm, 45 nm, 32 nm and 22 nm) will be using a light source with wavelength much greater than the feature size. This creates a peculiar problem where line width on manufactured devices is a function of relative spacing between adjacent lines. Despite numerous restriction on layout rules, interconnects may still suffer from constriction due to this peculiarity also known as forbidden pitch problem. A small manufacturing variation turns the constrictions to open faults. Gate leakage current is a significant concern for present and upcoming technology nodes. Due to gate leakage, an open fault is not truly an open circuit. Our simulation studies show that the leakage current steers the floating input of a gate to certain meta- stable states. This property actually makes it easier to detect open faults either through side channel excitation or by stuck-at tests. The major contributions of this paper are (i) lithographic simulation based identification of potential open fault sites, (ii) identification of meta-stable input states for these open inputs, (iii) length calculation for side channel signals for definitive detection of open faults. Together, they provide a complete CAD framework for testing lithography related open faults. Aswin Sreedhar, Alodeep Sanyal, Sandip Kundu |
DATE | 3 |
| 2008 | A framework for predictive dynamic temperature management of microprocessor systemsabstractThe sustained push for performance, transistor count and instruction level parallelism has reached a point where IC thermal issues are at the forefront of design constraints. Many of the current systems deploy dynamic voltage and frequency scaling (DVFS) to address thermal emergencies. DVFS has certain limitations in terms of response lag, scalability and being reactive. On the other hand, several hardware based control theoretic schemes have been proposed to deliver optimal performance, but such schemes come at high cost and lack flexibility and scalability. In this paper, we present an alternative thermal monitoring and management system that utilizes software and hardware components, based on virtual machine concept. The proposed scheme delivers targeted, localized, and preemptive thermal management at low cost, adapts well to a multitasking environment, while delivering maximum performance under thermal stress. Omer Khan, Sandip Kundu |
ICCAD | 2 |
| 2008 | Modeling and analysis of non-rectangular transistors caused by lithographic distortionsabstractSub-wavelength lithography causes the shape of transistors to differ from idealized rectangles. Several researchers have proposed transistor simulation models to characterize the non-ideal shapes of transistors for various regions of transistor operation. It has been shown that the effective channel length of a non-rectangular gate (NRG) transistor may be different for ON and OFF currents. In this paper we present a composite post-litho non-rectangular transistor model that not only accounts for DC behavior of a transistor, but also accounts for parasitic capacitances across a range of voltages. Parameters of this model can be fitted to real silicon data. The proposed model has been validated by TCAD device simulations. Results show that a composite model that accounts for both DC currents and parasitic capacitances is no more complex than models optimized for DC currents only. Further, the proposed model integrates readily with available SPICE simulators. We have also presented cell library characterization data to illustrate the benefit of using a delay accurate transistor model. Aswin Sreedhar, Sandip Kundu |
ICCD | 2 |
| 2008 | A Built-In Self-Test Scheme for Soft Error Rate CharacterizationabstractSoft errors caused by ionizing radiation have emerged as a major concern for current generation of CMOS technologies and the trend is expected to get worse. Soft error rate (SER) measurement, expressed as number of failures encountered per billion hours of device operation, is time consuming and involves significant test cost. The cost stems from having to connect a device-under-test to a tester for extended period of time. While built-in self-test (BIST) mechanisms have been around for over a decade, that minimizes the use of a tester; they have not been applied to measure or characterize soft error rate. This is because traditional BIST methods cannot distinguish between a soft failure and a hard failure and have no provision for counting the number of errors. In this paper, we propose a BIST design for soft error rate (SER) characterization, which obviates those issues. The proposed BIST based SER measurement scheme can be further accelerated by improved controllability and observability while unlike traditional BIST schemes, a test by test failure detection capability enables higher diagnostic resolution for single event based transient errors. We further propose to integrate this chip-level BIST-based SER characterization system with a distributed on-line scheme using a network controller that tests multiple chips in parallel and completely eliminates the need for a tester. The hardware overhead of the proposed architecture is small and it becomes insignificant for larger design. Alodeep Sanyal, Syed M. Alam, Sandip Kundu |
IOLTS | 3 |
| 2008 | An Automatic Post Silicon Clock Tuning System for Improving System Performance based on Tester MeasurementsabstractOptical shrink for process migration, manufacturing process variation and dynamic voltage control leads to clock skew as well as path delay variation in a manufactured chip. Since such variations are difficult to predict in pre-silicon phase, tunable clock buffers have been used in several microprocessor designs. The buffer delays are tuned to improve maximum operating clock frequency of a design. This however shifts the burden of finding tuning settings for individual clock buffers to the test process. In this paper, we describe a process for using Boolean tester measurements for determining the settings of the tunable buffers. The results show that frequency improvements of 10% or more are possible by appropriate setting of tunable clock buffers. Kelageri Nagaraj, Sandip Kundu |
ITC | 2 |
| 2008 | Statistical Yield Modeling for Sub-wavelength LithographyabstractPhotolithography is at the heart of semiconductor manufacturing process. To support continued scaling of transistors, lithographic resolution must continue to improve. At today's volume manufacturing process, a light source of 193 nm wavelength is used to print devices with 45 nm feature size. To address sub-wavelength printability, a number of resolution enhancement techniques (RET) have been used. While RET techniques allow printing of sub-wavelength features, the feature length itself becomes highly sensitive to process parameters, which in turn detracts from yield due to small perturbations in manufacturing parameters. Yield loss is a function of random variables such as depth-of-focus, exposure dose, lens aberration and resist thickness. The loss-of-yield is also a function of systematic components such as specific layout structure and out-of-band radiation from optical source. In this paper, we present a yield modeling technique for a given layout, based on a statistical model for process variability. The key issues addressed in this paper are (i) layout error modeling, (ii) avoidance of mask simulation for chip layouts, (iii) avoidance of full Monte-Carlo simulation for variational lithography modeling, (iv) building a methodology for yield estimation based on existing commercial tools. Results based on our approach show that yield sensitivity increases at smaller feature sizes. Aswin Sreedhar, Sandip Kundu |
ITC | 2 |
| 2008 | On Composite Leakage Current Maximization
Ashesh Rastogi, Kunal P. Ganeshpure, Alodeep Sanyal, Sandip Kundu |
J. Electron. Test. | 4 |
| 2008 | On Detection of Resistive Bridging Defects by Low-Temperature and Low-Voltage TestingabstractTest application at reduced power supply voltage (low-voltage testing) or reduced temperature (low-temperature testing) can improve the defect coverage of a test set, particularly of resistive short defects. Using a probabilistic model of two-line nonfeedback short defects, we quantify the coverage impact of low-voltage and low-temperature testing for different voltages and temperatures. Effects of statistical process variations are not considered in the model. When quantifying the coverage increase, we differentiate between defects missed by the test set at nominal conditions and undetectable defects (flaws) detected at non nominal conditions. In our analysis, the performance degradation of the device caused by lower power supply voltage is accounted for. Furthermore, we describe a situation in which defects detected by conventional testing are missed by low-voltage testing and quantify the resulting coverage loss. Experimental results suggest that test quality is improved even if no cost increase is allowed. If multiple test applications are acceptable, a combination of low voltage and low temperature turns out to provide the best coverage of both hard defects and flaws. Piet Engelke, Ilia Polian, Michel Renovell, Sandip Kundu, Bharath Seshadri, Bernd Becker 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | On Estimating Impact of Loading Effect on Leakage Current in Sub-65nm Scaled CMOS Circuits Based on Newton-Raphson MethodabstractDifferent sources of leakage can affect each other by interacting through resulting intermediate node voltages. This is known as the loading effect, In this paper, we propose a pattern dependent steady state leakage estimation technique that incorporates loading effect and addresses the three dominant sources of leakage, namely the sub-threshold, gate oxide and band-to-band tunneling leakages. We have developed a compact leakage model that supports iteration over node voltages based on Newton-Raphson method. The proposed estimation technique based on the compact model improves performance and capacity over SPICE. We report a speed up of 18,000X over SPICE. Results show that loading effect is a significant factor in leakage and worsens with technology scaling. Ashesh Rastogi, Sandip Kundu |
DAC | 3 |
| 2007 | Interactive presentation: Automatic test pattern generation for maximal circuit noise in multiple aggressor crosstalk faultsabstractDecreasing process geometries and increasing operating frequencies have made VLSI circuits more susceptible to signal integrity related failures. Capacitive crosstalk is one of the causes of such kind of failures. Crosstalk fault results from switching of neighboring lines that are capacitively coupled. Long nets are more susceptible to crosstalk faults because they tend to have a higher coupling capacitance to overall capacitance ratio. A typical long net has multiple aggressors. In generating patterns to create maximal crosstalk noise, it may not be possible to activate all aggressors at the same time. Therefore, pattern generation must focus on activating a maximal subset of aggressors weighted by actual coupling capacitance values. This is a variant of max-satisfiability problem. Unlike a traditional max-satisfiability problem, here we must deal with signal propagation to an observable output. In this paper, the authors present a novel solution that combines 0-1 integer linear program (ILP) with traditional stuck-at fault ATPG. The maximal aggressor activation is formulated as a linear programming problem while the fault effect propagation is treated as an ATPG problem. The problems are separated by min-cut circuit partitioning technique based on Kernighan-Lin-Fiduccia-Mattheyses (KLFM) method. This proposed technique was applied to ISCAS 85 benchmark circuits. Results indicated that 75-100% of the aggressors could be switched for generating crosstalk noise while satisfying requirement of sensitizing a path to the output Kunal P. Ganeshpure, Sandip Kundu |
DATE | 2 |
| 2007 | On modeling impact of sub-wavelength lithography on transistorsabstractAs the VLSI technology marches beyond 65 and 45 nm process technologies, variation in gate length has a direct impact on leakage and performance of CMOS transistors. Due to sub-wavelength lithography, the shape of the transistor often differs from idealized rectangles. In silicon, the effective channel length of a transistor varies across its width. This is a modeling problem. The average effective channel length is different for ON current and OFF currents, making it difficult, if not impossible for a single Leff to accurately represent both. In this paper, we report an accurate post-litho non-rectangular transistor modeling methodology. We further studied the impact of focus and dose variations in lithographic process on transistor parameters. The resulting transistor models were applied for standard cell characterization in successive steps of lithographic simulation of layout and device characterization. Results show that the new models can improve the accuracy of estimation of leakage current by 40% or more over a nominal model that is primarily tuned for ON current. Aswin Sreedhar, Sandip Kundu |
ICCD | 2 |
| 2007 | Accelerating Soft Error Rate Testing Through Pattern SelectionabstractIn this paper we propose a technique for increasing the rate of failure due to soft errors by carefully choosing the patterns for soft-error detection. It is well known that all circuit nodes are not equally vulnerable to soft error. We propose a metric for measuring vulnerability of a node to soft-error. The pattern selection approach constructs a test set to maximize the node vulnerability metric. In order to facilitate scan based application of these tests, we propose a test-per-clock DFT scheme that allows counting of such errors. The test set thus derived is applied repeatedly to accelerate the soft error rate measurement. Acceleration reported for this technique over random pattern testing on ISCAS-85 benchmarks ranges from 5X to infinity. Alodeep Sanyal, Kunal P. Ganeshpure, Sandip Kundu |
IOLTS | 3 |
| 2007 | On Derating Soft Error Probability Based on Strength FilteringabstractSoft errors caused by ionizing radiation have emerged as a major concern for current generation of CMOS technologies and the trend is expected to get worse. A significant fraction of soft errors in semiconductor has been reported to never lead to a system failure. System level soft-error rate (SER) analysis shows that soft-error in internal circuit nodes frequently fail to propagate to an observable point due to Boolean filtering and latching window filtering. A previous study shows that when soft-error is viewed as an analog signal distortion rather than a digital error, it often disappears during signal propagation due to error signal attenuation. This has been termed as electrical filtering. Electrical filtering in system level soft-error rate analysis is expensive because it involves circuit level simulation. In this paper, we present an electrical filtering technique that treats soft-errors as digital errors, but uses analog strengths to decide whether such errors can propagate. We call this techniquestrengthfiltering. Strength filtering does not involve SPICE simulation, hence it is computationally efficient. Used as a pre-processing step, strength filtering improves the efficiency of system level soft-error rate analysis. Experimental results on ISCAS-85 benchmark circuits show that an average of ~38% of the soft errors have no potential impact on the system level behavior and therefore, can be filtered out to improve both accuracy and efficiency of soft rate estimation process. Alodeep Sanyal, Sandip Kundu |
IOLTS | 2 |
| 2007 | A Study on Impact of Leakage Current on Dynamic PowerabstractScaling of CMOS technologies has led to dramatic increase in sub-threshold, gate and reverse biased junction band-to-band-tunneling (BTBT) leakage. Leakage current has now become comparable to the switching current. Traditionally, dynamic power and leakage power are computed separately. Dynamic power computation does not include leakage from non-switching nodes. In this paper, we show that in upcoming 45nm technology, leakage from non-switching nodes can account for as much as 38% of total dynamic current. Hence leakage from non-switching nodes can not be neglected during dynamic power computation. To facilitate this study on large benchmark circuits on which spice level simulation is impractical, we created a compact simulation model for modeling various pattern dependent leakage currents to allow leakage computation at gate level. Using a simulation based experiment we compare leakage and switching currents on ISCAS-85 benchmark circuits. The experiments are based on Berkeley predictive technology model for 45nm technology. The results firmly establish the need to consider leakage from non-switching nodes during dynamic power computation. Ashesh Rastogi, Kunal P. Ganeshpure, Sandip Kundu |
ISCAS | 3 |
| 2007 | On ATPG for multiple aggressor crosstalk faults in presence of gate delaysabstractCrosstalk faults have emerged as a significant mechanism for circuit failure. Long signal nets are of particular concern because they tend to have a higher coupling capacitance to overall capacitance ratio. A typical long net also has multiple aggressors. In generating patterns to create maximal crosstalk noise on a net, it may not be possible to activate all aggressors logically or simultaneously. Therefore, pattern generation must focus on activating a maximal subset of aggressors switching around the same time the victim net switches. This is a well-known problem. In this paper, we present a novel solution assuming a unit delay model for the gates, combining 0-1 integer linear program (ILP) with traditional stuck-at fault ATPG. The maximal aggressor activation is formulated as a linear programming problem while the fault effect propagation is treated as an ATPG problem and the gate delays are subsumed by a circuit transformation. The proposed technique was applied to ISCAS 85 benchmark circuits. Results indicate that percentage of total capacitance that can be switched varies from 30-80%. Kunal P. Ganeshpure, Sandip Kundu |
ITC | 2 |
| 2007 | Guest Editorial: Special Section on "Autonomous Silicon Validation and Testing of Microprocessors and Microprocessor-Based Systems"abstractThe five papers in this special section focus on autonomous silicon validation and testing of microprocessors and microprocessor-based systems. The papers cover several important aspects of the technical area: from first silicon validation and debug to manufacturing and production testing, including both software-based and hardware-based techniques and design-for-test methods that range from pure functional instruction-based to structural scan-based tests. The scope of the papers extends from embedded uniprocessors to multicore microprocessors. Dimitris Gizopoulos, Robert C. Aitken, Sandip Kundu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2006 | A design for failure analysis (DFFA) technique to ensure incorruptible signaturesabstractFast failure analysis is a key enabler in shortening the time between design tape out and product introduction in the market. With faster detection of manufacturability issues, problems associated with parametric variations, model approximations or physical design rules can be fixed faster either at the process control level or at the mask level. Failure analysis can be accelerated with additional hardware support for design-for-testability (DFT) and design-for-failure-analysis (DFFA). In this paper, we focus on one such DFFA technique deployed in the industry, identify its shortcomings and offer improvements to fix deficiencies Sandip Kundu |
DATE | 1 |
| 2006 | A Pattern Generation Technique for Maximizing Power Supply CurrentsabstractMax-current analysis is essential in power rail design and power supply switching noise analysis. Traditionally, maximum current from all CMOS gates are added together to compute maximum current level. This approach ignores all Boolean relationships. The problem of finding the input vector pair that will cause worst case current draw from the power rails when Boolean relationships are considered is an NP-hard problem. In this paper, we propose a Current Maximizing Pattern Generation (CMPG) algorithm which greatly reduces the computational complexity by using a parameterized branch-and-bound heuristic that prunes the search space by looking for a lower as well as an upper bound for maximum switching currents. When allowed to proceed indefinitely, the CMPG algorithm converges on an exact solution instead of finding upper and lower bounds. When coupled with a switch level SAT solver, CMPG can generate patterns for cell library characterization. Kunal P. Ganeshpure, Alodeep Sanyal, Sandip Kundu |
ICCD | 3 |
| 2006 | Power Droop TestingabstractCircuit activity is a function of input patterns. When circuit activity changes abruptly, it can cause sudden drop or rise in power supply voltage. This change is known as power droop and is an instance of power supply noise. Although power droop may cause an IC to fail, such failures cannot currently be screened during testing as it is not covered by conventional fault models. In this paper we present a technique for screening such failures. We propose a heuristic method to generate test sequences which create worst-case power drop by accumulating the high-frequency and low-frequency effects. The generated patterns need to be sequential even for scan designs. We employ a dynamically constrained version of the classical D-algorithm for test generation, i.e., the algorithm generates new constraints on-the-fly depending on previous assignments. The obtained patterns can be used for manufacturing testing as well as for early silicon validation. A prototype ATPG is implemented to demonstrate the feasibility of the approach and test sequences are generated for ISCAS circuits. Ilia Polian, Alexander Czutro, Sandip Kundu, Bernd Becker 0001 |
ICCD | 3 |
| 2006 | An Improved Technique for Reducing False Alarms Due to Soft ErrorsabstractA significant fraction of soft errors in modern microprocessors has been reported to never lead to a system failure. Any concurrent error detection scheme that raises alarm every time a soft error is detected is not well heeded because most of these alarms are false and responding to them will affect system performance negatively. This paper improves state of the art in detecting and preventing false alarms. Existing techniques are enhanced by a methodology to handle soft errors on address bits. Furthermore, we demonstrate benefit of false alarm identification in implementing a roll-back recovery system by first calculating the optimum check pointing interval for a roll-back recovery system and then showing that the optimal number of check-points decreases by orders of magnitude when exclusion techniques are used even if the implementation of exclusion technique is not perfect Sandip Kundu, Ilia Polian |
IOLTS | 1 |
| 2005 | On Detection of Resistive Bridging Defects by Low-Temperature and Low-Voltage TestingabstractResistive defects are gaining importance in very-deepsubmicron technologies, but their detection conditions are not trivial. Test application can be performed under reduced temperature and/or voltage in order to improve detection of these defects. This is the first analytical study of resistive bridge defect coverage of CMOS ICs under low-temperature and mixed low-temperature, low-voltage conditions. We extend a resistive bridging fault model in order to account for temperature-induced changes in detection conditions. We account for changes in both the parameters of transistors involved in the bridge and the resistance of the short defect itself. Using a resistive bridging fault simulator, we determine fault coverage for low-temperature testing and compare it to the numbers obtained at nominal conditions. We also quantify the coverage of flaws, i.e. defects that are redundant at nominal conditions but could deteriorate and become earlylife failures. Finally, we compare our results to the case of low-voltage testing and comment on combination of these two techniques. Sandip Kundu, Piet Engelke, Ilia Polian, Bernd Becker 0001 |
Asian Test Symposium | 1 |
| 2005 | Path-oriented transition fault test generation considering operating conditionsabstractWe describe a test generation procedure for path-oriented transition faults that takes into account the fact that operating conditions may change during circuit operation. A path-oriented transition fault is detected through the longest sensitizable path that goes through the fault site. The operating conditions we consider are junction temperature and power supply voltage. Since path delays change with operating conditions, the longest path through a fault site may be different under different conditions. We show that test generation using nominal delays is not sufficient for covering the complete range of operating conditions, even if N-detection test generation is used. Therefore, operating conditions need to be addressed explicitly during test generation. However, since temperature and voltage are continuous variables and represent an infinite number of values in the range, test generation must concentrate on a small selected set of operating conditions. We discuss the selection of these conditions and demonstrate that N-detection test generation with multiple operating conditions is effective in covering the range of operation conditions almost completely. Bharath Seshadri, Irith Pomeranz, Sudhakar M. Reddy, Sandip Kundu |
ETS | 4 |
| 2005 | Transient fault characterization in dynamic noisy environmentsabstractTechnology trends are increasing the frequency of serious transient (soft) faults in digital systems. For example, ICs are becoming more susceptible to cosmic radiation, and are being embedded in applications with dynamic noisy environments. We propose a generic framework for representing such faults and characterizing them on-line. We formally define the impact of a transient fault in terms of three basic parameters: frequency, observability and severity. We distinguish fault modes in systems whose noise environment changes dynamically. Based on these ideas, the problem of designing on-line architectures for transient fault characterization is formulated and analyzed for several optimization goals. Finally, experiments are described that determine transient fault impact and the corresponding tests for various simulated fault modes of the ISCAS-89 benchmark circuits. Ilia Polian, John P. Hayes, Sandip Kundu, Bernd Becker 0001 |
ITC | 3 |
| 2005 | Resistive Bridge Fault Model Evolution from Conventional to Ultra Deep Submicron TechnologiesabstractWe present three resistive bridging fault models valid for different CMOS technologies. The models are partitioned into a general framework (which is shared by all three models) and a technology-specific part. The first model is based on Shockley equations and is valid for conventional but not deep submicron CMOS. The second model is obtained by fitting SPICE data. The third resistive bridging fault model uses Berkeley predictive technology model and BSIM4; it is valid for CMOS technologies with feature sizes of 90nm and below, accurately describing non-trivial electrical behavior in that technologies. Experimental results for ISCAS circuits show that the test patterns obtained for the Shockley model are still valid for the fitted model, but lead to coverage loss under the predictive model. Ilia Polian, Sandip Kundu, Jean-Marc Gallière, Piet Engelke, Michel Renovell, Bernd Becker 0001 |
VTS | 2 |
| 2005 | On modeling crosstalk faultsabstractTraditionally, digital testing of integrated semiconductor circuits has focused on manufacturing defects. There is another class of failures that happens due to circuit marginalities. Circuit-marginality failures are on the rise due to shrinking process geometries, diminishing supply voltage, sharper signal-transition rates, and aggressive styles in circuit design. There are many different marginality issues that may render a circuit nonoperational. Capacitive cross coupling between interconnects is known to be a leading cause for marginality-related failures. In this paper, we present novel techniques to model and prioritize capacitive crosstalk faults. Experimental results are provided to show effectiveness of the proposed modeling technique on large industrial designs. Sandip Kundu, Sujit T. Zachariah, Yi-Shing Chang, Chandra Tirumurti |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | A Modeling Approach for Addressing Power Supply Switching Noise Related Failures of Integrated CircuitabstractPower density of high-end microprocessors has been increasing by approximately 80% per technology generation, while the voltage is scaling by a factor of 0.8. This leads to 225% increase in current per unit area in successive generation of technologies. The cost of maintaining the same IR drop becomes too high. This leads to compromise in power delivery and power grid becomes a performance limiter. Traditional performance related test techniques with transition and path delay fault models focus on testing the logic but not the power delivery. In this paper we view power grid as performance limiter and develop a fault model to address the problem of vector generation for delay faults arising out of power delivery problems. A fault extraction methodology applied to a microprocessor design block is explained. Chandra Tirumurti, Sandip Kundu, Susmita Sur-Kolay, Yi-Shing Chang |
DATE | 2 |
| 2004 | Static statistical timing analysis for latch-based pipeline designsabstractA latch-based timing analyzer is an essential tool for developing high-speed pipeline designs. As process variations increasingly influence the timing characteristics of DSM designs, a timing analyzer capable of handling process-induced timing variations for latch-based pipeline designs becomes in demand. In this work, we present a static statistical timing analyzer, STAP, for latch-based pipeline designs. Our analyzer propagates statistical worst-case delays as well as critical probabilities across the pipeline stages. We present an efficient method to handle correlations due to re-convergent fanouts. We also demonstrate the impact of not including the analysis of reconvergent fanouts in latch-based pipeline designs. Comparing to a Monte-Carlo based timing analyzer, our experiments show that STAP can accurately evaluate the critical probability that a design violates the timing constraints under a given statistical timing model. The runtime comparison further demonstrates the efficiency of our STAP. Rob A. Rutenbar, Li-C. Wang, Kwang-Ting Cheng, Sandip Kundu |
ICCAD | 4 |
| 2004 | Trends in manufacturing test methods and their implicationsabstractDriven by market applications in the areas of computing, networking, storage, optical, wireless, portable, and consumer electronics, semiconductor chips today are as diverse as ever. Confluence of multiple applications and rapid integration has also driven the heterogeneity of chips. Test methods have evolved with the products. However, the basic goals in testing remain the same: quality of product, recurring and non-recurring costs and time to market. In this paper we try to catalog some commonly used test methods, identify their associated DFT requirements and trends in terms of tester requirements. Given the diversity of semiconductors chips today such as various PLDs, volatile and non-volatile memories, analog, mixed signal, FPGA, ASIC, SOC, MEMs and processors, it is impossible for a paper of this nature to be fully comprehensive. So we limit our focus on processor, ASIC and SOCs. Sandip Kundu, Rajesh Galivanche |
ITC | 1 |
| 2004 | Masking of Unknown Output Values during Output Response Compression byUsing Comparison UnitsabstractA circuit may produce unknown output values during simulation of a test set, e.g., due to an unknown initial state or due to the existence of tristate elements. Unknown output values in the output response of a circuit make it impossible to determine a single unique signature for the fault-free circuit when built-in self-test is used for testing the circuit. We consider the problem of synthesizing a logic block that replaces unknown output values in the output response of a circuit with a known constant. The logic block is constructed from building blocks called comparison units. The synthesis procedure ensures that the built-in self-test scheme will be able to detect all the faults detectable by the test set applied to the circuit while allowing a single unique signature to be computed. Two variations of the synthesis procedure are considered, a two-dimensional version suitable for synchronous sequential circuits without scan and for scan circuits with multiple scan chains and a one-dimensional version suitable for scan circuits with a single scan chain. Irith Pomeranz, Sandip Kundu, Sudhakar M. Reddy |
IEEE Trans. Computers | 2 |
| 2004 | Pitfalls of hierarchical fault simulationabstractCertain circuit structures, such as self-loop, asynchronous reset, and clock division, may not be visible in a hierarchical (mixed) simulation system. Since the simulator does not know about their existence, it cannot cope with them like it normally would in a flat circuit. If this leads to a logic-simulation problem, users can usually discover them easily during the validation process. However, if it only causes fault-simulation inaccuracy, it is hard to find the problem. In this paper, we show examples illustrating their existence. The examples negate an assumption that has been used in many papers on mixed-mode simulation. The examples have been abstracted from real industrial designs of microprocessors. Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | On the characterization and efficient computation of hard-to-detect bridging faultsabstractWe investigate a characterization of hard-to-detect bridging faults. For circuits with large numbers of lines (or nodes), this characterization can be used to select target faults for test generation efficiently, when it is impractical to target all the bridging faults (or all the realistic bridging faults). We demonstrate that the faults selected based on the proposed characterization are indeed hard-to-detect by performing the following experiments. 1) We show that the fault coverage of a given test set, with respect to the selected subset of bridging faults, is lower and more sensitive to the test set than the fault coverage obtained with respect to a random subset of bridging faults of the same size, with respect to the complete set of bridging faults, and when possible, with respect to a subset of realistic bridging faults of the same size. 2) We demonstrate that a test set generated for the selected subset of bridging faults detects other bridging faults more effectively than when a test set is derived for a randomly selected subset of bridging faults of the same size. We also describe an efficient procedure for selecting hard-to-detect bridging faults according to the proposed characterization. This procedure avoids enumeration of all the faults in order to select the hard-to-detect ones. This is important for large circuits where even enumeration of all the bridging faults may not be feasible. Irith Pomeranz, Sudhakar M. Reddy, Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2003 | Circuit and Platform Design Challenges in Technologies beyond 90nm
Bill Grundmann, Rajesh Galivanche, Sandip Kundu |
DATE | 3 |
| 2003 | On the Characterization of Hard-to-Detect Bridging Faults
Irith Pomeranz, Sudhakar M. Reddy, Sandip Kundu |
DATE | 3 |
| 2003 | On Modeling Cross-Talk Faults
Sujit T. Zachariah, Yi-Shing Chang, Sandip Kundu, Chandra Tirumurti |
DATE | 3 |
| 2003 | On-chip Compression of Output Responses with Unknown Values Using LFSR Reseeding
Masao Naruse, Irith Pomeranz, Sudhakar M. Reddy, Sandip Kundu |
ITC | 4 |
| 2002 | On output response compression in the presence of unknown output valuesabstractA circuit may produce unknown output values during simulation of an input sequence due to an unknown initial state or due to the existence of tri-state elements. For circuits tested using BIST, unknown output values make it impossible to determine a single unique signature for the fault free circuit. To accommodate unknown output values in a BIST scheme, we describe a procedure for synthesizing a minimal logic block that replaces unknown output values by a known constant. The proposed procedure ensures that the BIST scheme will be able to detect all the faults detectable by the input sequence applied to the circuit while allowing a single unique signature to be obtained. Irith Pomeranz, Sandip Kundu, Sudhakar M. Reddy |
DAC | 2 |
| 2001 | Fast Statistical Timing Analysis By Probabilistic Event PropagationabstractWe propose a new statistical timing analysis algorithm, which produces arrival-time random variables for all internal signals and primary outputs for cell-based designs with all cell delays modeled as random variables. Our algorithm propagates probabilistic timing events through the circuit and obtains final probabilistic events (distributions) at all nodes. The new algorithm is deterministic and flexible in controlling run time and accuracy. However, the algorithm has exponential time complexity for circuits with reconvergent fanouts. In order to solve this problem, we further propose a fast approximate algorithm. Experiments show that this approximate algorithm speeds up the statistical timing analysis by at least an order of magnitude and produces results with small errors when compared with Monte Carlo methods. Jing-Jia Liou, Kwang-Ting Cheng, Sandip Kundu, Angela Krstic |
DAC | 3 |
| 2001 | Test Challenges in Nanometer Technologies
Sandip Kundu, Sujit T. Zachariah, Sanjay Sengupta, Rajesh Galivanche |
J. Electron. Test. | 1 |
| 2000 | Performance sensitivity analysis using statistical method and its applications to delayabstractThe performance of deep submicron designs can be affected by various parametric variations, manufacturing defects, noise or modeling errors that are all statistical in nature. We propose a statistical framework for analyzing the performance sensitivity of designs to various timing related defects/noise/variations. The core engine of our approach is a highly efficient statistical timing analysis tool. We describe the application of our framework for delay fault modeling and analysis of resistive opens and shorts and as well as interconnect crosstalk. We present experimental results demonstrating the accuracy of our statistical framework as compared to SPICE (for a given set of input patterns) and nominal worst-case analysis. Experimental results for analysis of resistive opens and shorts are also included. Jing-Jia Liou, Angela Krstic, Kwang-Ting Cheng, Deb Aditya Mukherjee, Sandip Kundu |
ASP-DAC | 5 |
| 1999 | On Detecting Bridges Causing Timing FailuresabstractHigh resistance bridges (resistive bridges) are becoming more common. Such bridges cause speed failures. Published experimental results show that current tests are not good at detecting such defects. The following results, pertinent to resolving the above issue, are presented: mechanisms by which bridges cause speed failure; identification of non-classical transition tests for such bridges; and the usefulness and pitfalls of using low voltage testing to detect such bridges. Sreenivas Mandava, Sreejit Chakravarty, Sandip Kundu |
ICCD | 3 |
| 1998 | Timing Analysis and Optimization: From Devices to Systems (Abstract of Embedded Tutorial)abstractTiming analysis spans the entire design process from RTL synthesis to timing sign-off including schematic design, logic synthesis, floor plan, place and route, clock distribution etc. In this embedded tutorial we will cover timing analysis from devices to systems. At the transistor level we will cover static, dynamic and interconnect analysis. At the gate level we will cover path analysis including false paths and static and dynamic sensitization aspects. The new Delay Calculator Language (DCL) will also be covered. We will tie the statistical properties at each level to put them under proper perspective. Anirudh Devgan, Sandip Kundu |
ASP-DAC | 2 |
| 1998 | IDDQ Defect Detection in Deep Submicron CMOS ICsabstractTraditional testing is based on applying signals to circuit inputs and observing logic stale of outputs. However, with the introduction of CMOS technology a new test technique emerged, that is based on applying signal inputs and observing standby electrical current. This technique is known as IDDQ testing. In most CMOS circuits, when all inputs are frozen in place and circuit transient activities settle down, the current drops by 3-5 orders of magnitude below normal levels. During testing, if there is a deviation from this behavior a defect is detected. While parametric test of electric current has always been popular (including in bipolar and NMOS technologies) to detect a short between power supply lines, the promise of IDDQ test is that it will also allow detection of bridging between signal nets and transistor stuck-on faults in CMOS circuits. However, in reality, the resolution of the size of the defect is limited by the magnitude of stand-by current. Since normal process variation can cause a shift in stand-by current by two orders of magnitude, choosing a single threshold for IDDQ defect detection is impractical. Many ideas have been proposed to overcome this problem, most of which are based on choosing an adaptive cut-off point for a wafer, or a lot. In this paper, we suggest a technique, that in concept, is akin to setting an individual cutoff point for a die. The concept has been validated based on actual data. Sandip Kundu |
Asian Test Symposium | 1 |
| 1998 | GateMaker: a transistor to gate level model extractor for simulation, automatic test pattern generation and verificationabstractHierarchy is key to managing design complexity. A hierarchical design system needs to maintain many views of the same design entity. Some of the examples might be physical view for placement routing and extraction; transistor schematic view for circuit simulation, timing characterization and noise analysis; a gate level schematic view for timing, verification, logic simulation, fault simulation and automatic lest pattern generation (ATPG); a register transfer level (RTL) view for specification and high level simulation etc. In order to achieve highest system performance, multiple design iterations are necessary, each iteration involving both forward and backward pass through hierarchy, with manual changes at any level of the hierarchy. This poses an essential challenge of keeping all views of same design entity in sync. In this paper we describe an automatic tool called GateMaker, that has been developed to extract a gate level schematic model from a transistor level schematic model for the purposes of logic simulation, fault simulation and automatic test pattern generation. This eliminates a manual process and offers manifold advantages that will be discussed in this paper. Sandip Kundu |
ITC | 1 |
| 1997 | Timing analysis and optimization: from devices to systems (tutorial)
Anirudh Devgan, Leon Stok, Sandip Kundu |
ICCAD | 3 |
| 1996 | Self-Checking Comparator with One Periodic OutputabstractIn this paper we propose a new self-checking comparator with one periodic output. The comparator can be used as a two-rail checker or as an equality checker. Two different input patterns are sufficient to detect all the faults considered. Sandip Kundu, Egor S. Sogomonyan, Michael Gössel, Steffen Tarnick |
IEEE Trans. Computers | 1 |
| 1995 | Panel: New Research Problems in the Emerging Test Technology
Vishwani D. Agrawal, Bernard Courtois, Fumiyasu Hirose, Sandip Kundu, Yinghua Min, Parimal Pal Chaudhuri |
Asian Test Symposium | 4 |
| 1994 | Microprocessor Testing: Which Technique is Best? (Panel)abstractNo abstract available. Jacob A. Abraham, Sandip Kundu, Janak H. Patel, Manuel A. d'Abreu, Bulent I. Dervisoglu, Marc E. Levitt, Hector R. Sucar, Ron G. Walther |
DAC | 2 |
| 1994 | Incremental synthesis
Daniel Brand, Anthony D. Drumm, Sandip Kundu, Prakash Narain |
ICCAD | 3 |
| 1994 | Multifault Testable Circuits Based on Binary Parity DiagramsabstractWe introduce a new class of binary graphs called binary parity diagrams (BPDs). Ordered binary parity diagram of a function is canonical. Importance of canonical forms in circuit verification is well known. Circuits derived on the basis of these diagrams are multifault testable. In this paper we focus on multifault testability property.> Sandip Kundu |
ICCD | 1 |
| 1994 | An incremental algorithm for identification of longest (shortest) paths
Sandip Kundu |
Integr. | 1 |
| 1994 | An efficient technique for obtaining unate implementation of functions through input encoding
Sandip Kundu |
Integr. | 1 |
| 1994 | Highly Reliable Symmetric NetworksabstractWe generalize directed loop networks to loop-symmetric networks in which there are N nodes and in which each node has in-degree and out-degree k, subject to the condition that 2/sup k/ does not exceed N. We show that by proper selection of links one can obtain generalized loop networks with optimal or close to optimal diameter and connectivity. The optimized diameter is less than k/spl lsqb/N/sup 1/k//spl rsqb/, where /spl lsqb/x/spl rsqb/ indicates the ceiling of x. We also show that these networks are rather compact in that the diameter is not more than twice the average distance. Roughly 1/2(k/spl minus/1)N/sup 1/k/ nodes can be removed such that the network of remaining nodes is still strongly connected, if all remaining nodes have at least one incoming and one outgoing link left. > Leendert M. Huisman, Sandip Kundu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1994 | Diagnosing scan chain faultsabstractTesting screens for good chips. However, when test fall out is high (low yield) it becomes necessary to diagnose faults so that the manufacturing process or physical design can be filed to improve yield. Several scan based diagnostic schemes are used in industry. They work when the scan chain itself is fault free. In this paper we describe a diagnosis system that can diagnose faults in a scan chain.> Sandip Kundu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1993 | Design of Scan-Based Path-Delay-Testable Sequential CircuitsabstractSeveral techniques that produce robust path delay testable designs for arbitrary combinational logic functions have been recently proposed. Unfortunately, these often results in circuits with longer delays. Moreover, such designs also assume extra hardware in the form of "holding" latches in the scan-chain in order to apply arbitrary vector-pairs, as required for high coverage delay testing. This results in high area overheads. In this paper, we present a scan-based design for testability technique for arbitrary sequential circuits. This new scheme allows a delay vs. area tradeoff while enforcing delay fault testability. At one end of the spectrum, a fully robust path delay testable circuit is guaranteed with very low additional hardware overheads in the form of extra latches. At the other end, full robust path delay testability is ensured without compromising the overall circuit delay, and with hardware overheads.> Ankan K. Pramanick, Sandip Kundu |
ITC | 2 |
| 1993 | On diagnosis of faults in a scan-chainabstractTesting screens for good chips. However, when test fall out is high (low yield) it becomes necessary to diagnose faults so that the manufacturing process or physical design can be fixed to improve yield. Several scan based diagnostic schemes are used in industry. They work when the scan chain itself is fault free. This paper describes a diagnosis system that can diagnose faults in a scan chain.> Sandip Kundu |
VTS | 1 |
| 1993 | Testability preserving Boolean transforms for logic synthesisabstractSynthesis proceeds through local transformations with various objectives. If testability is a concern, these transformations are limited to those that preserve or enhance testability. Such transformations are called testability preserving transformations. The fault model used is pivotal to the analysis of any such transformation. In this paper, the authors chose single-path-propagating hazard-free robust delay fault testability to qualify them. This model was chosen because it disambiguates results of delay testing which are often inconclusive (the presence of a fault can neither be ascertained nor be denied) and ensures stuck-at fault testability as well. Unfortunately, only a few transformations are known to obey testability requirements. This limitation is a serious handicap in attaining other synthesis goals such as area and performance optimization. In this paper the authors establish a relationship between testability properties of logic transformations and their Boolean duals, the application of which enlarges the existing number of testability preserving transforms. They demonstrate further that some of the new transformations thus achieved may actually enhance testability.> Sandip Kundu, Ankan K. Pramanick |
VTS | 1 |
| 1992 | A Small Test Generator for Large DesignsabstractWe report an automatic test pattern generator that can handle designs with one million gates or more on medium size workstations. Run times and success rates, i.e. the fraction of faults that are resolved, are comparable to or better than those reported previously in the literature. No preprocessing is required and the amount of memory needed is less than 100 bytes per gate. The low memory requirements and high performance have been achieved by working with a larger but simpler search space, by simplifying decision making and backtracking and by using only implication techniques that are fast and that require no preprocessing. Sandip Kundu, Leendert M. Huisman, Indira Nair, Vijay S. Iyengar, Lakshmi N. Reddy |
ITC | 1 |
| 1991 | Design of robustly testable combinational logic circuitsabstractAn integrated approach to the design of combinational logic circuits in which all path delay faults and multiple line stuck-at, transistor stuck-open faults are detectable by robust tests is proposed. Robustly testable static CMOS primitive logic circuit designs are presented for any arbitrary combinational logic function. They require no special gates, and fan-in and fan-out constraints do not affect the designs. Extra controllable inputs or additional hardware to achieve testability was not used. It is demonstrated that the method guarantees the design of CMOS logic circuits in which all path delay faults are locatable.> Sandip Kundu, Sudhakar M. Reddy, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1990 | Robust tests for parity trees
Sandip Kundu, Sudhakar M. Reddy |
J. Electron. Test. | 1 |
| 1990 | On Symmetric Error Correcting and All Unidirectional Error Detecting CodesabstractThe design of binary block codes that are capable of correcting up to t symmetric errors and detecting all unidirectional errors is discussed. A class of systematic t-symmetric-error-correcting/all-unidirectional-error-detecting (t-SyEC/AUED) codes are proposed. When t=0 the proposed codes become Berger codes. For t=1, the proposed codes are shown to be of asymptotically optimal order. Methods to construct nonsystematic t-SyEC/AUED codes for t=2 and 3 are also presented.> Sandip Kundu, Sudhakar M. Reddy |
IEEE Trans. Computers | 1 |
| 1989 | Design of TSC checkers for implementation in CMOS technologyabstractA fault model for CMOS digital circuits includes FET stuck-open and FET stuck-on faults, in addition to line stuck-at faults used conventionally. It has been shown that delays in CMOS circuits, under test, may invalidate tests derived by neglecting such delays. This necessitates reinvestigation of existing totally self-checking (TSC) checker designs from a new perspective. It was shown earlier that TSC checkers derived on the basis of the line stuck-at fault model for constant weight codes may not be self-testing for the CMOS fault model, violating one of the conditions of TSC circuits. A design procedure for constructing a self-testing circuit for the CMOS fault model using at most four levels is suggested. Thus previous designs can be adapted for the CMOS fault model without any penalty. The new design also makes it possible to meet arbitrary fan-in restrictions.> Sandip Kundu, Sudhakar M. Reddy |
ICCD | 1 |
| 1989 | Design of multioutput CMOS combinational logic circuits for robust testabilityabstractThe author proposes a testable design for multioutput functions using parity gates that always produces a realization with robust tests. The use of parity gates allows more logic sharing among various outputs than would have been possible otherwise. The solution presented here has the ability to accommodate any fan-in restriction and grow in number of levels. The new design is well suited for multioutput circuits.> Sandip Kundu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1988 | On the design of robust multiple fault testable CMOS combinational logic circuitsabstractTests that detect modeled faults independent of the delays in the circuit under test are called robust tests. An integrated approach to the design of combinational logic circuits in which all single stuck-open faults and path delay faults are detectable by robust tests was presented by the authors earlier. It is shown that the earlier design actually results in circuits in which all multiple stuck-at and stuck-open and multipath delay faults are robustly testable. The tests to detect such faults are presented.> Sandip Kundu, Sudhakar M. Reddy, Niraj K. Jha |
ICCAD | 1 |
| 1988 | Robust Tests for Parity TreesabstractTests to detect line stuck-at and transistor stuck-open faults in CMOS binary parity trees are considered. The main objective was to derive bounds on test sequences and test sequences that detect all single faults, even in the presence of circuit delays. It is shown that the desired robustness of test requires tests whose lengths are proportional to the depth of the trees under test, in contrast to earlier used tests of constant length, in which robustness was not required. The results derived give exact bounds for complete trees, but similar results for incomplete trees need further investigations.> Sandip Kundu, Sudhakar M. Reddy |
ITC | 1 |