Nirmal Saxena

dblp:03/1575 · also Nirmal R. Saxena · DBLP profile ↗
← Back
35ranked-venue papers
13as first author
5since 2021 · last 2025
0009-0001-1248-0721ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 10 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSecurity and privacy · 2Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Invited Paper: Hardware-Software Co-Design for Highly Optimized, Customized, and Reliable AI Systems
abstract
Over the past decade, AI has been rapidly integrated into our daily life, coming in every shape and size and working across systems from big clouds to IoT. As a result, AI systems are increasingly requiring enhancements in model efficiency, hardware acceleration, and memory systems to satisfy stringent constraints on efficiency, reliability, and security. However, advancing across these fronts is challenging as compute demand outpaces Moore’s-law efficiency, hardening into an AI compute wall and an AI energy wall. Breaking through requires a unified AI co-design loop that co-optimizes algorithms and hardware, including efficient AI-to-hardware mapping, so that ongoing goals (accuracy, sparsity, latency) align with concrete hardware choices (precision modes, interconnects, memory hierarchies) and AI-specific execution and memory-reuse patterns. This paper details the principal co-design challenges, presents complementary strategies, and outlines a practical roadmap toward highly optimized, efficient, reliable, and secure AI systems.
Jörg Henkel, Mehdi Baradaran Tahoori, Heba Khdr, Hassan Nassar, Vincent Meyers, Deming Chen, Selin Yildirim, Yingbing Huang, Nirmal Saxena, Saurabh Hukerikar, Srivi Dhruvanarayan
ICCAD9
2023 DrGPU: A Top-Down Profiler for GPU Applications
abstract
GPUs have become common in HPC systems to accelerate scientific computing and machine learning applications. Efficiently mapping these applications to rapid evolutions of GPU architectures for high performance is a well-known challenge. Various performance inefficiencies exist in GPU kernels that impede applications from obtaining bare-metal performance. While existing tools are able to measure these inefficiencies, they mostly focus on data collection and presentation, requiring significant manual efforts to understand the root causes for actionable optimization. Thus, we develop DrGPU, a novel profiler that performs top-down analysis to guide GPU code optimization. As its salient feature, DrGPU leverages hardware performance counters available in commodity GPUs to quantify stall cycles, decompose them into various stall reasons, pinpoint root causes, and provide intuitive optimization guidance. With the help of DrGPU, we are able to analyze important GPU benchmarks and applications and obtain nontrivial speedups --- up to 1.77X on V100 and 2.03X on GTX 1650.
Yueming Hao, Rob Van der Wijngaart, Nirmal Saxena, Yuanbo Fan, Xu Liu 0001
ICPE4
2022 Runtime Fault Diagnostics for GPU Tensor Cores
abstract
Tensor cores in NVIDIA GPUs are important computational engines that accelerate diverse AI deep neural networks and algorithms for perception, mapping, localization and path planning in autonomous drive systems. The occurrence of random hardware faults, particularly permanent faults in these computational units have potentially catastrophic consequences. This paper describes software-based runtime diagnostics for tensor cores that leverage universal test patterns (UTP) to achieve high diagnostic fault coverage with very low latency enabling GPU-based systems to meet the ASIL B functional safety targets outlined in the ISO 26262 standard.
Saurabh Hukerikar, Nirmal Saxena
ITC2
2022 Error Model (EM) - A New Way of Doing Fault Simulation
abstract
This paper introduces the concept of error model (EM) that replaces a traditional fault-model-based simulation. EM is a temporal simulation of fault symptoms in an application processor. This paper shows that the application resiliency metrics, such as fault coverage, derived through EM are more comprehensive and accurate than those derived through empirical models like single stuck-at faults. In addition, EM can be used on high-level simulation models (behavioral RTL, emulation or in some cases in-silicon). EM approach gives greater than three orders of performance improvement over gate netlist models using stuck-at fault simulation. This paper shows that the coverage metrics, for a billion-logic-gate GPU design, obtained through in-silicon EM closely match the corresponding coverage metrics estimated from a low-level netlist with single stuck-at fault simulation.
Nirmal Saxena, Atieh Lotfi
ITC1
2021 Characterizing and Mitigating Soft Errors in GPU DRAM
abstract
GPUs are used in high-reliability systems, including high-performance computers and autonomous vehicles. Because GPUs employ a high-bandwidth, wide-interface to DRAM and fetch each memory access from a single DRAM device, implementing full-device correction through ECC is expensive and impractical. This challenge is compounded by worsening relative rates of multi-bit DRAM errors and increasing GPU memory capacities. This paper first presents high-energy neutron beam testing results for the HBM2 memory on a compute-class GPU. These results uncovered unexpected intermittent errors that we determine to be caused by cell damage from the high-intensity beam. As these errors are an artifact of the testing apparatus, we provide best-practice guidance on how to identify and filter them from the results of beam testing campaigns. Second, we use the soft error beam testing results to inform the design and evaluation of system-level error protection mechanisms by reporting the relative error rates and error patterns from soft errors in GPU DRAM. We observe locality in the multi-bit errors, which we attribute to the underlying structure of the HBM2 memory. Based on these error patterns, we propose several novel ECC schemes to decrease the silent data corruption risk by up to five orders of magnitude relative to SEC-DED ECC, while also reducing the number of uncorrectable errors by up to 7.87 ×. We compare novel binary and symbol-based ECC organizations that differ in their design complexity, hardware overheads, and permanent error correction abilities, ultimately recommending two promising organizations. These schemes replace SEC-DED ECC with no additional redundancy, likely no performance impacts, and modest area and complexity costs.
Michael B. Sullivan 0001, Nirmal Saxena, Mike O'Connor, Donghyuk Lee, Paul Racunas, Saurabh Hukerikar, Timothy Tsai 0002, Siva Kumar Sastry Hari, Stephen W. Keckler
MICRO2
2020 On the Measurement of Safe Fault Failure Rates in High-Performance Compute Processors
abstract
Accurate determination of the safe-fault failure rate of complex digital designs is an exascale problem. We present a novel measurement methodology and results which could have a profound impact on the performance and availability for GPUs in safety critical systems. We extend our analysis with a methodology for in-the-field verification.
Richard Bramley, Yanxiang Huang, Guangshan Duan, Nirmal Saxena, Paul Racunas
ITC4
2019 Resiliency of automotive object detection networks on GPU architectures
abstract
Safety is the most important aspect of an autonomous driving platform. Deep neural networks (DNNs) play an increasingly critical role in localization, perception, and control in these systems. The object detection and classification inference are of particular importance to construct a precise picture of a vehicle's surrounding objects. Graphics Processing Units (GPU) are well-suited to accelerate such DNN-based inference applications since they leverage data and thread-level parallelism in GPU architectures. Understanding the vulnerability of such DNNs to random hardware faults (including transient and permanent faults) in GPU-based systems is essential to meet the safety requirements of auto safety standards such as the ISO 26262, as well as to influence the design of hardware and software-based safety features in current and future generations of GPU architectures and GPU-based automotive platforms. In this paper, we assess the vulnerability of object detection and classification DNNs to permanent and transient faults using fault injection experiments and accelerated neutron beam testing respectively. We also evaluate the effectiveness of chip-level safety mechanisms in GPU architectures, such as ECC and parity, in detecting these random hardware faults. Our studies demonstrate that such object detection networks tend to be vulnerable to random hardware faults, which cause incorrect or mispredicted object detection outcomes. The neutron beam experiments show that existing chip-level protections successfully mitigate all silent data corruption events caused by transient faults. For permanent faults, while ECC and parity are effective in some cases, our results suggest the need for exploring other complementary detection methods, such as periodic online and offline diagnostic testing.
Atieh Lotfi, Saurabh Hukerikar, Keshav Balasubramanian, Paul Racunas, Nirmal Saxena, Richard Bramley, Yanxiang Huang
ITC5
2018 Low Overhead Tag Error Mitigation for GPU Architectures
abstract
Cache structures on modern GPUs or CPUs occupy a large area and are frequently accessed. This increases their vulnerability to transient errors. With some area and energy overhead, these structures are often protected by ECC or parity checking. However, in deference to the energy efficiency and scalability challenges in high-performance computing, it is crucial to minimize any unnecessary overhead while maintaining the desired reliability. This paper evaluates the reliability of unprotected tag SRAM structures in modern GPUs, and studies the use of a low-overhead tag error mitigation mechanism. The proposed mechanism exploits Galois-based hash functions for set-index calculation to mitigate some pathological address strides that cause false hit events. Extensive analysis on a modern GPU indicates that the hash-based mechanism yields 10x reduction in false hit probability (with 2% improvement in hit rate) for write-through data caches when compared to a baseline cache indexing scheme.
Atieh Lotfi, Nirmal Saxena, Richard Bramley, Paul Racunas, Philip P. Shirvani
DSN2
2008 How Many Test Patterns are Useless?
abstract
Studies by previous researchers using production test data reported that not all the production test patterns applied detected defective chips. Researchers found that 70% to 90% of their production test patterns seemed useless because these patterns detected no defective chips and they could therefore be removed without impacting test quality. Previous researchers qualitatively explained this finding by a lack of correlation between test metrics and defect coverage. Notwithstanding the lack of correlation between test metrics and defect coverage, in this paper we develop a simple statistical model that relates the expected number of useless patterns to the production yield, the defect coverage characteristics, and the number of tested chips. This model demonstrates that for practical values of production yield, defect coverage and number of chips tested, a significant fraction of test patterns will be useless. We validated this statistical model by comparing its results with actual production testing data.
François-Fabien Ferhani, Nirmal Saxena, Edward J. McCluskey, Phil Nigh
VTS2
2004 Efficient Design Diversity Estimation for Combinational Circuits
abstract
Redundant systems are designed using multiple copies of the same resource (e.g., a logic network or a software module) in order to increase system dependability: Design diversity has long been used to protect redundant systems against common-mode failures. The conventional notion of diversity relies on "independent" generation of "different" implementations of the same logic function. In a recent paper, we presented a metric to quantify diversity among several designs. The problem of calculating the diversity metric is NP-complete (i.e., can be of exponential complexity). In this paper, we present efficient techniques to estimate the value of the design diversity metric. For datapath designs, we have formulated very fast techniques to calculate the value of the metric by taking advantage of the regularity in the datapath structures. For general combinational logic circuits, we present an adaptive Monte-Carlo simulation technique for estimating accurate bounds on the value of the metric.
Subhasish Mitra, Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Computers2
2002 A Design Diversity Metric and Analysis of Redundant Systems
abstract
Redundant systems are designed using multiple copies of the same resource (e.g., a logic network or a software module) in order to increase system dependability. Design diversity has long been used to protect redundant systems from common-mode failures. The conventional notion of diversity relies on "independent" generation of "different" implementations. This concept is qualitative and does not provide a basis for comparing the reliabilities of two diverse systems. In this paper, for the first time, we present a metric to quantify diversity among several designs and illustrate its effectiveness using several examples. Applications of this metric in analyzing reliability and availability of diverse redundant systems, and deriving simple relationships between diversity, system failure rate, and mission time are also demonstrated.
Subhasish Mitra, Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Computers2
2001 Techniques for Estimation of Design Diversity for Combinational Logic Circuits
abstract
Design diversity has long been used to protect redundant systems against common-mode failures. The conventional notion of diversity relies on "independent" generation of "different" implementations of the same logic function. This concept is qualitative and does not provide a basis to compare the reliabilities of two diverse systems. In a recent paper, we presented a metric to quantify diversity among several designs. The problem of calculating the diversity metric is NP-complete and can be of exponential complexity. In this paper we present techniques to estimate the value of the design diversity metric. For datapath designs, we have formulated very fast techniques to calculate the value of the metric by exploiting the regularity in the datapath structures. For general combinational logic circuits, we present an adaptive Monte-Carlo simulation technique for estimating bounds on the value of the metric. The adaptive Monte-Carlo simulation technique provides accurate estimates of the design diversity metric; the number of simulations used to reach this estimate is polynomial (instead of exponential) in the number of circuit inputs. Moreover, the number of simulations can be tuned depending on the desired accuracy.
Subhasish Mitra, Nirmal Saxena, Edward J. McCluskey
DSN2
2000 A Reliable LZ Data Compressor on Reconfigurable Coprocessors
abstract
Data compression techniques based on the Lempel-Ziv (LZ) algorithm are widely used in a variety of applications, especially in communications and data storage. However, since the LZ algorithm involves a considerable amount of parallel comparisons, it may be difficult to achieve a very high throughput using software approaches on general-purpose processors. In addition, error propagation due to single-bit transient errors during LZ compression causes a data integrity problem. We present an implementation of LZ data compression on reconfigurable hardware with concurrent error detection for high performance and reliability. Our approach achieves 100 Mbps throughput using four Xilinx 4036XLA FPGA chips. We also present an inverse comparison technique for LZ compression to guarantee data integrity with less area overhead than traditional systems based on duplication. The resulting execution time overhead and compression ratio degradation due to concurrent error detection is also minimized.
Wei-Je Huang, Nirmal Saxena, Edward J. McCluskey
FCCM2
2000 An ACS Robotic Control Algorithm with Fault Tolerant Capabilities
abstract
This paper demonstrates that an adaptive computing system (ACS) is good platform for implementing robotic control algorithms. We show that an ACS can be used to provide both good performance and high dependability. An example of an FPGA-implemented dependable control algorithm is presented. The flexibility of ACS is exploited by choosing the best precision for our application. This reduces the amount of required hardware and improves performance. Results obtained from a WILDFORCE emulation platform showed that even using 0.35 /spl mu/m technology, an FPGA-implemented control algorithm has comparable performance with the software-implemented control algorithm in a 0.25 /spl mu/m microprocessor. Different voting schemes are used in conjunction with multi-threading and combinational redundancy to add fault tolerance to the robotic controller. Error-injection experiments demonstrate that robotic control algorithms with fault tolerance techniques are orders of magnitude less vulnerable to faults compared to algorithms without any fault tolerant features.
Shu-Yi Yu, Nirmal Saxena, Edward J. McCluskey
FCCM2
2000 Fault Escapes in Duplex Systems
abstract
Hardware duplication techniques are widely used for concurrent error detection in dependable systems to ensure high availability and data integrity. These techniques are vulnerable to common-mode failures (CMFs). Use of duplex systems with diverse implementations of the two modules has been proposed in the past for protection against CMFs. In this paper, we define a category of faults, called non-self-testable faults that undermine the data integrity of dependable systems. These faults produce identical errors at the outputs of the two modules of a duplex system and can potentially be caused by CMFs. The main contributions of this paper are: (1) techniques that identify non-self-testable faults in duplex systems, and (2) design methods that reduce the number of non-self-testable faults by test point insertion. We show that our algorithm for identifying non-self-testable faults runs orders of magnitude faster than exact techniques with minimal loss of accuracy. Also, there is a significant reduction in the number of test points required for duplex systems with diverse implementations compared to duplex systems with identical implementations. Thus, we can detect common-mode failures in diverse duplex systems using very few test points. These results are especially useful for systems with user-programmable logic elements that enhance the practicality of using diverse designs in duplex systems.
Subhasish Mitra, Nirmal Saxena, Edward J. McCluskey
VTS2
2000 Common-mode failures in redundant VLSI systems: a survey
abstract
This paper presents a survey of CMF (common-mode failures) in redundant systems with emphasis on VLSI (very large scale integration) systems. The paper discusses CMF in redundant systems, their possible causes, and techniques to analyze reliability of redundant systems in the presence of CMF. Current practice and results on the use of design diversity techniques for CMF are reviewed. By revisiting the CMF problem in the context of VLSI systems, this paper augments earlier surveys on CMF in nuclear and power-supply systems. The need for quantifiable metrics and effective models for CMF in VLSI systems is re-emphasized. These metrics and models are extremely useful in designing reliable systems. For example, using these metrics and models, system designers and synthesis tools can incorporate diversity in redundant systems to maximize protection against CMF.
Subhasish Mitra, Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Reliab.2
2000 Software-implemented EDAC protection against SEUs
abstract
In many computer systems, the contents of memory are protected by an error detection and correction (EDAC) code. Bit-flips caused by single event upsets (SEU) are a well-known problem in memory chips; EDAC codes have been an effective solution to this problem. These codes are usually implemented in hardware using extra memory bits and encoding/decoding circuitry. In systems where EDAC hardware is not available, the reliability of the system can be improved by providing protection through software. Codes and techniques that can be used for software implementation of EDAC are discussed and compared. The implementation requirements and issues are discussed, and some solutions are presented. The paper discusses in detail how system-level and chip-level structures relate to multiple error correction. A simple solution is presented to make the EDAC scheme independent of these structures. The technique in this paper was implemented and used effectively in an actual space experiment. We have observed that SEU corrupt the operating system or programs of a computer system that does not have any EDAC for memory, forcing the system to be reset frequently. Protecting the entire memory (code and data) might not be practical in software. However this paper demonstrates that software-implemented EDAC is a low-cost solution that provides protection for code segments and can appreciably enhance the system availability in a low-radiation space environment.
Philip P. Shirvani, Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Reliab.2
1999 A design diversity metric and reliability analysis for redundant systems
abstract
Design diversity has long been used to protect redundant systems against common-mode failures. The conventional notion of diversity relies on "independent" generation of "different" implementations. This concept is qualitative and does not provide a basis to compare the reliabilities of two diverse systems. In this paper, for the first time, we present a metric to quantify diversity among several designs. Based on this metric, we derive analytical reliability models that show a simple relationship between design diversity, system failure rate, and mission time. In addition, we present simulation results to demonstrate the effectiveness of design diversity in Duplex and Triple Modular Redundant (TMR) systems. For independent multiple-module failures, we show that, mere use of different implementations does not always guarantee higher reliability compared to redundant systems with identical implementations-it is important to analyze the reliability of redundant systems using our metric. For common-mode failures and design faults, there is a significant gain in using different implementations-however, as our analysis shows, the gain diminishes as the mission time increases. Our simulation results also demonstrate the usefulness of diversity for enhancing the self-testing properties of redundant systems.
Subhasish Mitra, Nirmal Saxena, Edward J. McCluskey
ITC2
1999 Finite state machine synthesis with concurrent error detection
abstract
A new synthesis technique for designing finite state machines with on-line parity checking is presented. The output logic and the next-state logic of the finite state machines are checked independently. By checking parity on the present state instead of the next state, this technique allows detection of errors in bistable elements (that were hitherto not detected by many previous techniques) while requiring no changes in the original machine specifications. This paper also examines design choices with respect to parity prediction circuits. Two such examined choices are the multi-parity-group and the single-parity-group techniques. A new state encoding technique based on the JEDI program is developed for the synthesis of the next-state logic with an additional parity output. Synthesis results produced by our proposed procedure for the MCNC'89 FSM benchmark circuits show on average a 25% reduction in literal counts compared to previous techniques.
Chaohuang Zeng, Nirmal Saxena, Edward J. McCluskey
ITC2
1998 Dependable adaptive computing systems-the ROAR project
abstract
Describes the ROAR (Reliability Obtained by Adaptive Reconfiguration) project. Adaptive computing system (ACS) environments provide opportunities for dependable computing solutions. We identify some of the issues in traditional dependable computing solutions. We show that adaptive and reconfigurable environments address these issues and also enable new innovative techniques that enhance the dependability of systems. Natural redundancy in applications and new synthesis algorithms are two other opportunities we use in enhancing the dependability.
Nirmal Saxena, Edward J. McCluskey
SMC1
1997 Parallel Signatur Analysis Design with Bounds on Aliasing
abstract
This paper presents parallel signature design techniques that guarantee the aliasing probability to be less than 2/L, where L is the test length. Using y signature samples, a parallel signature analysis design is proposed that guarantees the aliasing probability to be less than (y/L)/sup y/2/. Inaccuracies and incompleteness in previously published bounds on the aliasing probability are discussed. Simple bounds on the aliasing probability are derived for parallel signature designs using primitive polynomials.
Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Computers1
1996 Generation of Test Cases for Hardware Design Verification of a Super-Scalar Fetch Processor
abstract
We describe a method to generate test cases (or test programs) for hardware design verification. The proposed method uses non-functional errors defined over a software model of the specification to guide the generation of test programs. Application of the proposed method to the Hal. super-scalar Fetch Processor is also described. For this design, we present experimental results to demonstrate the effectiveness of the method in achieving a high design error coverage with relatively short test programs.
Irith Pomeranz, Nirmal Saxena, Richard Reeve, Paritosh Kulkarni, Yan A. Li
ITC2
1996 Counting Two-State Transition-Tour Sequences
abstract
This paper develops a closed-form formula, f(k), to count the number of transition-tour sequences of length k for bistable machines. It is shown that the function f(k) is related to Fibonacci numbers. Some applications of the results in this paper are in the areas of testable sequential machine designs, random testing of register data paths, and qualification tests for random pattern generators.
Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Computers1
1995 Floating Point Fault Tolerance with Backward Error Assertions
abstract
The paper introduces an assertion scheme based on the backward error analysis for error detection in algorithms that solve dense systems of linear equations, Ax=b. Unlike previous methods, this backward error assertion model is specifically designed to operate in an environment of floating point arithmetic subject to round-off errors, and it can be easily instrumented in a Watchdog processor environment. The complexity of verifying assertions is O(n/sup 2/), compared to the O(n/sup 3/) complexity of algorithms solving Ax=b. Unlike other proposed error detection methods, this assertion model does not require any encoding of the matrix A. Experimental results under various error models are presented to validate the effectiveness of this assertion scheme.>
Daniel Boley, Gene H. Golub, Samy Makar, Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Computers4
1995 Fault-Tolerant Features in the HaL Memory Management Unit
abstract
This paper describes fault-tolerant and error detection features in HaL's memory management unit (MMU). The proposed fault-tolerant features allow recovery from transient errors in the MMU. It is shown that these features were natural choices considering the architectural and implementation constraints in the MMU's design environment. Three concurrent error detection and correction methods employed in address translation and coherence tables in the MMU are described. Virtually-indexed and virtually-tagged cache architecture is exploited to provide an almost fault-secure hardware coherence mechanism in the MMU, with very small performance overhead (less than 0.01% in the instruction throughput). Low overhead linear polynomial codes have been chosen in these designs to minimize both the hardware and software instrumentation impact.>
Nirmal Saxena, David Chih-Wei Chang, Kevin Dawallu, Jaspal Kohli, Pat Helland
IEEE Trans. Computers1
1994 Linear Complexity Assertions for Sorting
abstract
Correctness of the execution of sorting programs can be checked by two assertions: the order assertion and the permutation assertion. The order assertion checks if the sorted data is in ascending or descending order. The permutation assertion checks if the output data produced by sorting is a permutation of the original input data. Permutation and order assertions are sufficient for the detection of errors in the execution of sorting programs; however, in terms of execution time these assertions cost the same as sorting programs. An assertion, called the order-sum assertion, that has lower execution cost than sorting programs is derived from permutation and order assertions. The reduction in cost is achieved at the expense of incomplete checking. Some metrics are derived to quantify the effectiveness of order-sum assertion under various error models. A natural connection between the effectiveness of the order-sum assertion and the partition theory of numbers is shown. Asymptotic formulae for partition functions are derived.>
Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Software Eng.1
1992 Simple Bounds on Serial Signature Analysis Aliasing for Random Testing
abstract
It is shown that the aliasing probability is bounded above by (1+ epsilon )/L approximately=1/L ( epsilon small for large L) for test lengths L less than the period, L/sub c/, of the signature polynomial; for test lengths L that are multiples of L/sub c/, the aliasing probability is bounded above by 1; for test lengths L greater than L/sub c/ and not a multiple of L/sub c/, the aliasing probability is bounded above by 2/(L/sub c/+1). These simple bounds avoid any exponential complexity associated with the exact computation of the aliasing probability. Simple bounds also apply to signature analysis based on any linear finite state machine (including linear cellular automaton). From these simple bounds it follows that the aliasing probability in a signature analysis design using beta intermediate signatures is bounded by ((1+ epsilon )/sup beta / beta /sup beta /)/L/sup beta /, for beta>
Nirmal Saxena, Piero Franco, Edward J. McCluskey
IEEE Trans. Computers1
1991 Refined Bounds on Signature Analysis Aliasing for Random Testing
abstract
in previous work a simple bound, ~+2 , on the aliasing probability in serial signature analysis for a random test pattern of length L was derived. This simple bound is sharpened here by almost a factor of two. For serial signature analysis, it is shown that the I+& 1 aliasing probability is bounded above by - = L (E L small for large L) for test lengths L less than the period, Lc, of the signature polynomial. The simple bounds derived are compared with exact as well as experimentally measured aliasing probability values. It is conjectured that L-l is the best monotonic bound on the aliasing probability for serial signature analysis.
Nirmal Saxena, Piero Franco, Edward J. McCluskey
ITC1
1990 Control-Flow Checking Using Watchdog Assists and Extended-Precision Checksums
abstract
A control-flow checking method using extended-precision checksums and watchdog assists is proposed. Control-flow checking based on extended-precision checksums is shown to have low error detection latency compared to previously proposed methods. Analytical measures are derived to demonstrate the effectiveness of using extended-precision checksums for control flow checking. It is shown that the error detection latency in the extended-precision-checksum-based control-flow checking remains relatively constant for both single and multiple sequence errors. In the case of signature-based methods, error detection latency increases linearly with the number of sequence errors. A watchdog assist architecture for control-flow checking in programs which addresses several architecture issues is proposed. This watchdog assist architecture can support control-flow checking for multiprocessor, multiprogramming, and cache-based environments. The Hewlett-Packard Precision Architecture is used as an example architecture to demonstrate the feasibility of watchdog assists.>
Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Computers1
1990 Analysis of Checksums, Extended-Precision Checksums, and Cyclic Redundancy Checks
abstract
The effectiveness of extended-precision checksums is thoroughly analyzed. It is demonstrated that the extended-precision checksums most effectively exploit natural redundancy occurring in program codes. Honeywell checksums and cyclic redundancy checks are compared to extended-precision checksums. Two's complement, unsigned, and one's complement arithmetic checksums are treated in a unified manner. Results are also extended to any general radix-p arithmetic checksum. Asymptotic and closed-form formulas of aliasing probabilities for the various error models are derived.>
Nirmal Saxena, Edward J. McCluskey
IEEE Trans. Computers1
1989 Arithmetic and galois checksums
abstract
An analysis is presented of error detection characteristics when galois checksums and arithmetic checksums are used simultaneously. By generalizing previous results, it is shown that galois checksums and arithmetic checksums exhibit orthogonal characteristics with respect to error detection. An analytic proof of orthogonality is presented for certain restricted cases. These orthogonal characteristics hold good for equally likely errors, restricted column errors, and restricted word errors. Double-length galois checksums are compared with combined arithmetic and galois checksums.>
Nirmal Saxena, Edward J. McCluskey
ICCAD1
1988 Simultaneous signature and syndrome compression
abstract
A commonly used organization for built-in self-test of VLSI (very large-scale integration) circuits uses complete or pseudorandom test input generators followed by output data reduction. Two compression techniques which have been used are polynomial division (signature) and ones counting (syndrome). The simultaneous use of both of these approaches in parallel is investigated. Analytic and enumerative results indicate that the number of error patterns which are missed by both methods together is nearly the theoretical minimum. The conclusion extends to signature compression combined with any other counter-based compression method such as the use of Walsh spectral coefficients. Some suggestions for CAD (computer-aided design) implementations of test design are given based on these results.>
John P. Robinson, Nirmal Saxena
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1988 Syndrome and transition count are uncorrelated
abstract
In the testing of logic circuits, two proposed data-compression methods use the number of ones (syndrome) and the number of sequence changes (transition count). An enumeration N(m, k, t) of the number of length-m binary sequences having syndrome value k and transition count t is developed. Examination of this result reveals that the parallel compression of these two methods has small overlap in error masking. An asymptotic expression for N(m, k, t) is developed.>
Nirmal Saxena, John P. Robinson
IEEE Trans. Inf. Theory1
1987 A Unified View of Test Compression Methods
abstract
A unified treatment of the various techniques to reduce the output data from a unit under test is given. The characteristics of time compression schemes with respect to errors detected are developed. The use of two or more of these methods together is considered. Methods to design efficient test compression structures for built-in-tests are proposed. The feasibility of the proposed approach is demonstrated by simulation results.
John P. Robinson, Nirmal Saxena
IEEE Trans. Computers2
1986 Accumulator Compression Testing
abstract
A new test data reduction technique called accumulator compression testng (ACT) is proposed. ACT is an extension of syndrome testing. It is shown that the enumeration of errors missed by ACT for a unit under test is equivalent to the number of restricted partitions of a number. Asymptotic results are obtained for independent and dependent error modes. Comparison is made between signature analysis (SA) and ACT. Theoretical results indicate that with ACT a better control over fault coverage can be obtained than with SA. Experimental results are supportive of this indication. Built-in self test for processor environments may be feasible with ACT. However, for general VLSI circuits the complexity of ACT may be a problem as an adder is necessary.
Nirmal Saxena, John P. Robinson
IEEE Trans. Computers1