Kewal K. Saluja

dblp:74/5143 · DBLP profile ↗
← Back
155ranked-venue papers
17as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 137 · 16 first-author · 1 since 2021Computer networks · 11Software engineering, systems software and programming languages · 6Security and privacy · 5Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
38 papers
Electronic design automation · 75% Distributed systems · 8% Hardware reliability and fault tolerance · 8%
Computer networks
4 papers
Internet of things and sensor networks · 34% Transport protocols and congestion control · 22% Wireless sensing and localization · 14%

Topics — the 30 heaviest of 87, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
hardware verification and test
0.5242011
Power and Thermal Constrained Test Scheduling Under Deep Submicron Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Post-silicon diagnosis of segments of failing speedpaths due to manufacturing variations · DAC 2010
Critical-Path-Aware X-Filling for Effective IR-Drop Reduction in At-Speed Scan Testing · DAC 2007
Internet of things and sensor networks
mobile sensor networks
0.222009
Modeling Detection Latency with Collaborative Mobile Sensing Architecture · IEEE Trans. Computers 2009
Analytic modeling of detection latency in mobile sensor networks · IPSN 2006
Wireless sensing and localization › radar signal processing
target detection
0.222009
Modeling Detection Latency with Collaborative Mobile Sensing Architecture · IEEE Trans. Computers 2009
Analytic modeling of detection latency in mobile sensor networks · IPSN 2006
Internet of things and sensor networks
wireless sensor network
0.122009
Modeling Detection Latency with Collaborative Mobile Sensing Architecture · IEEE Trans. Computers 2009
Fault Tolerance in Collaborative Sensor Networks for Target Detection · IEEE Trans. Computers 2004
Electronic design automation › hardware verification and test
test scheduling
0.122011
Power and Thermal Constrained Test Scheduling Under Deep Submicron Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Test Scheduling and Control for VLSI Built-In Self-Test · IEEE Trans. Computers 1988
Internet architecture and protocols
network coding
0.112011
Routing TCP Flows in Underwater Mesh Networks · IEEE J. Sel. Areas Commun. 2011
Transport protocols and congestion control › TCP performance
TCP performance over wireless
0.112011
Routing TCP Flows in Underwater Mesh Networks · IEEE J. Sel. Areas Commun. 2011
Transport protocols and congestion control
transport protocols
0.112011
Routing TCP Flows in Underwater Mesh Networks · IEEE J. Sel. Areas Commun. 2011
Routing and switching
wireless routing
0.112011
Routing TCP Flows in Underwater Mesh Networks · IEEE J. Sel. Areas Commun. 2011
Electronic design automation › hardware verification and test › test scheduling
power-aware test scheduling
0.112011
Power and Thermal Constrained Test Scheduling Under Deep Submicron Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation › hardware verification and test › test scheduling
thermal-aware test scheduling
0.112011
Power and Thermal Constrained Test Scheduling Under Deep Submicron Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation
hardware test
0.142005
A method for reducing the target fault list of crosstalk faults in synchronous sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
On diagnosing multiple stuck-at faults using multiple and singlefault simulation in combinational circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002
Testing reconfigured RAM's and scrambled address RAM's for pattern sensitive faults · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996
Electronic design automation › hardware verification and test › design validation
post-silicon validation
0.112010
Post-silicon diagnosis of segments of failing speedpaths due to manufacturing variations · DAC 2010
Electronic design automation
timing analysis
0.112010
Post-silicon diagnosis of segments of failing speedpaths due to manufacturing variations · DAC 2010
Internet of things and sensor networks
detection latency
0.112009
Modeling Detection Latency with Collaborative Mobile Sensing Architecture · IEEE Trans. Computers 2009
Distributed systems › fault tolerance
high availability
0.112008
Implementing high availability memory with a duplication cache · MICRO 2008
Hardware reliability and fault tolerance › memory fault tolerance
memory redundancy
0.112008
Implementing high availability memory with a duplication cache · MICRO 2008
Electronic design automation › hardware verification and test
test generation
0.152005
Combinational automatic test pattern generation for acyclic sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
A method of reducing aliasing in a built-in self-test environment · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Test application time reduction for sequential circuits with scan · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation › hardware verification and test › delay fault testing
at-speed testing
0.112007
Critical-Path-Aware X-Filling for Effective IR-Drop Reduction in At-Speed Scan Testing · DAC 2007
Electronic design automation › hardware verification and test › test data compression
x-filling
0.112007
Critical-Path-Aware X-Filling for Effective IR-Drop Reduction in At-Speed Scan Testing · DAC 2007
Electronic design automation › hardware verification and test › test generation
sequential circuit test generation
0.122005
Combinational automatic test pattern generation for acyclic sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
An Efficient Algorithm for Sequential Circuit Test Generation · IEEE Trans. Computers 1993
Performance modeling and evaluation
analytical modeling
0.112006
Analytic modeling of detection latency in mobile sensor networks · IPSN 2006
Hardware reliability and fault tolerance › error detection
error detection latency
0.112006
Analytic modeling of detection latency in mobile sensor networks · IPSN 2006
Electronic design automation › hardware test
crosstalk fault testing
0.112005
A method for reducing the target fault list of crosstalk faults in synchronous sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › hardware verification and test › VLSI testing
defect-based testing
0.112005
Optimizing program disturb fault tests using defect-based testing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Electronic design automation › hardware verification and test
delay fault testing
0.112005
A method for reducing the target fault list of crosstalk faults in synchronous sequential circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Hardware reliability and fault tolerance › memory reliability
memory fault modeling
0.112005
Optimizing program disturb fault tests using defect-based testing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Distributed systems
fault tolerance
0.122004
Fault Tolerance in Collaborative Sensor Networks for Target Detection · IEEE Trans. Computers 2004
Design and Analysis of a Gracefully Degrading Interleaved Memory System · IEEE Trans. Computers 1990
Network management and operations › network robustness
fault tolerance
0.012004
Fault Tolerance in Collaborative Sensor Networks for Target Detection · IEEE Trans. Computers 2004
Distributed systems
consensus
0.012004
Fault Tolerance in Collaborative Sensor Networks for Target Detection · IEEE Trans. Computers 2004

Methods — techniques the papers use, named apart from their topics

simulation · 0.3thermal simulation · 0.1superposition principle · 0.1network coding · 0.1path delay measurement · 0.1integer linear programming · 0.1analytic modeling · 0.1critical-path-aware x-filling · 0.1time-frame expansion · 0.1fault classification · 0.1electrical simulation · 0.1combinational ATPG · 0.1value fusion · 0.0hierarchical agreement · 0.0decision fusion · 0.0test algorithm design · 0.0hypergraph coloring · 0.0boolean matrix algorithm · 0.0
YearPublicationVenuePosition
2021 Fault Tolerant Lanczos Eigensolver via an Invariant Checking Method
Felix Loh, Kewal K. Saluja, Parameswaran Ramanathan
J. Electron. Test.2
2017 Exploiting path delay test generation to develop better TDF tests for small delay defects
abstract
Localized small delay defects, for example due to degraded transistor drive strength caused by a broken fin, are a growing concern in current FinFET and emerging gate all around (GAA) technologies. Such defects are currently targeted by timing-aware Transition Delay Fault (TDF) tests that aim to test the target nodes along the longest path. The resulting tests often require considerable test generation time, have high test data volume, and at times do not provide the desired coverage. In this paper, we show that Path Delay Fault (PDF) test generation can be exploited to not only generate the timing tests more efficiently, but the resulting TDF test sets are also more compact and perform better on commonly used delay test coverage metrics. This is because all TDF faults along a PDF targeted timing-critical path can be detected efficiently by generating a single PDF test. This efficiency is not explicitly exploited by node oriented TDF test generation even when the TDFs are targeted along the longest paths. We demonstrate the effectiveness of our methodology for a range of benchmark circuits by comparing the results from a commercial timing-aware ATPG (TA-ATPG) with our new approach that efficiently exploit PDF tests wherever possible. The proposed new approach results in approximately 12.5% reduction in pattern volume, 35% reduction in ATPG runtime and also a 5% improvement in delay test coverage (DTC) when compared to existing TA-ATPG approaches.
Ankush Srivastava, Adit D. Singh, Virendra Singh, Kewal K. Saluja
ITC4
2017 A Reliability-Aware Methodology to Isolate Timing-Critical Paths under Aging
Ankush Srivastava, Virendra Singh, Adit D. Singh, Kewal K. Saluja
J. Electron. Test.4
2016 Crypt-Delay: Encrypting IP Cores with Capabilities for Gate-level Logic and Delay Simulations
abstract
System-on-Chip is a promising model for design of complex integrated circuits. In this model, designers may easily incorporate licensed and/or purchased Intellectual Property (IP) modules from other vendors to significantly reduce the design cycle time. However, the need to fulfill the legal obligations in the associated license or purchase contracts usually imposes considerable burden on the design process. The burden is often in the form of design access constraints that have to be imposed on SoC design team and/or in the form of additional costs needed to acquire a less constraining contract. To alleviate these problems, this paper proposes a encryption approach that assures confidentiality of the gate-level circuit information for the module designer while allowing the SoC designers an ability to perform gate-level digital simulations with gate-level delays. Empirical evaluation on benchmark circuits shows that, in addition to not being able to identify the gate types, an attacker cannot decrypt the delays associated with most of the gates in the circuit.
Parameswaran Ramanathan, Kewal K. Saluja
ATS2
2016 Necessary and Sufficient Conditions for Thermal Schedulability of Periodic Real-Time Tasks Under Fluid Scheduling Model
abstract
With the growing need to address the thermal issues in modern processing platforms, various performance throttling schemes have been proposed in literature (DVFS, clock gating, and so on) to manage temperature. In real-time systems, such methods are often unacceptable, as they can result in potentially catastrophic deadline misses. As a result, real-time scheduling research has recently focused on developing algorithms that meet the compute deadline while satisfying power and thermal constraints. Basic bounds that can determine if a set of tasks can be scheduled or not were established in the 1970s based on computation utilization. Similar results for thermal bounds have not been forthcoming. In this article, we address the problem of thermal constraint schedulability of tasks and derive necessary and sufficient conditions for thermal feasibility of periodic tasksets on a unicore system. We prove that a GPS-inspired fluid scheduling scheme is thermally optimal when context switch/preemption overhead is ignored. Extension of sufficient conditions to a nonfluid model is still an open problem. We also extend some of the results to a multicore processing environment. We demonstrate the efficacy of our results through extensive simulations. We also evaluate the proposed concepts on a hardware testbed.
Parameswaran Ramanathan, Kewal K. Saluja
ACM Trans. Embed. Comput. Syst.3
2015 A Methodology for Identifying High Timing Variability Paths in Complex Designs
abstract
In some complex deep sub-micron designs, the variations in interconnect delay has a significant impact on the production yield of the product. In this paper, we develop a theoretical explanation for the unexpectedly higher process related timing variability shown by long interconnects that are driven by high drive strength gates. This gets even worse due to conventional gate delay variability and other random process effects. Our analysis is supported by actual silicon data and further validated by detailed Monte-Carlo (MC) simulations. Unfortunately, traditional scan based transition delay fault (TDF) timing tests can miss these variability induced delay faults on long interconnects which lies on the critical paths. We propose a methodology to identify high variability paths dominated by such long interconnects, with the aim of developing high quality delay timing tests. Specifically, we develop a heuristic based path selection algorithm to identify potentially slow paths that can contribute to test escapes in production. We further extend our approach to generate high quality delay timing tests for the target paths using the proposed "three pass" method.
Virendra Singh, Adit D. Singh, Kewal K. Saluja
ATS3
2014 Necessary and Sufficient Conditions for Thermal Schedulability of Periodic Real-Time Tasks
abstract
With growing need to address the thermal issues in modern processing platforms various performance throttling schemes have been proposed in literature (DVFS, clock gating etcetera). In real-time systems such methods are often unacceptable as they can result into potentially catastrophic deadline misses. As a result real-time scheduling research has been focused in developing algorithms which meet the compute deadline while satisfying power and thermal constraints. Basic bounds that can determine if a set of tasks can be scheduled or not were established in the 70's based on computation utilization of processing power and no new results have been forthcoming that deal with thermal effect based bounds. In this paper we address the problem of thermal constraint schedulability of tasks and derive necessary and sufficient conditions for thermal feasibility of periodic task sets for a unicore system. We then extend some of these results to multi-coreprocessing environment. We demonstrate the efficacy of our results through extensive simulations.
Parameswaran Ramanathan, Kewal K. Saluja
ECRTS3
2014 Optimal Test Scheduling Formulation under Power Constraints with Dynamic Voltage and Frequency Scaling
Spencer K. Millican, Kewal K. Saluja
J. Electron. Test.2
2013 Formulating Optimal Test Scheduling Problem with Dynamic Voltage and Frequency Scaling
abstract
Various techniques for modern high performance designs, such as clock gating and dynamic voltage frequency scaling (DVFS), have been adapted to address power issues. This is a consequence of technology scaling and it is important and desirable to address reliability needs as well as economic issues. From a testing point of view, introduction of power constraints during testing is needed for the desired product quality and to avoid yield loss. Unlike designers who have often benefited from the design for test hardware introduced for testing, test engineers have rarely taken advantage of the extra hardware introduced to meet design needs. In this paper, we make use of the DVFS technology and its associated hardware to improve test economics. We formulate the power constrained testing problem as an optimization problem that makes use of DVFS technology. We show that we can obtain superior test schedules for both session-based and session less testing methods relative to existing and traditional methods of obtaining test schedules.
Spencer K. Millican, Kewal K. Saluja
Asian Test Symposium2
2013 Diagnosing Resistive Open Faults Using Small Delay Fault Simulation
abstract
Modern high performance, high density integrated circuits use a very large number of metal layers, necessitating the need to deal with the problem of resistive open defects. Resistive opens often manifest as and are modeled as small delay faults. Furthermore, in deep sub-micron technologies, it is known that the additional delay of a line with resistive open fault is not only a function of the resistant of the faulty line but it is also dependent on the signal transition(s) on its adjacent lines. In this paper, we propose an efficient simulation method to simulate small delay faults and we use this simulator to diagnose resistive open faults. The fault simulator developed by us simulates all delay faults for one signal line simultaneously. This information is then used to deduce the candidate faulty lines in two steps. Experimental results for ISCAS'89 benchmark circuits show that by using the method proposed by us the faulty lines can be identified correctly in most cases.
Koji Yamazaki, Toshiyuki Tsutsumi, Hiroshi Takahashi, Yoshinobu Higami, Hironobu Yotsuyanagi, Masaki Hashizume, Kewal K. Saluja
Asian Test Symposium7
2013 DRMA: dynamically reconfigurable MPSoC architecture
abstract
Embedded systems are ubiquitous and are deployed in a large range of applications. Designing and fabricating Integrated Circuits (ICs) targeting such different range of applications is expensive. Designers seek flexible processors which efficiently execute a multitude of applications. FPGAs are considered affordable, but design cost, high reconfiguration delay and power consumption are all prohibitive. In this paper, we propose a novel ASIC based flexible MPSoC architecture, which can execute separate tasks in parallel, and it can be configured to execute single task with wide data widths or execute multiple tasks with varying data widths. The architecture presented, called Dynamically Reconfigurable MPSoC Architecture (DRMA), can be rapidly reconfigured through instructions. We present applications as case studies to showcase the flexibility and efficacy of DRMA. Results show for an additional area overhead of about 5%, the system is capable of working as four 32-bit processors, a single 128 bit processor or as a pipelined processing system.
Lawrance Zhang, Jude Angelo Ambrose, Jorgen Peddersen, Sri Parameswaran, Roshan G. Ragel, Swarnalatha Radhakrishnan, Kewal K. Saluja
ACM Great Lakes Symposium on VLSI7
2013 On thermal utilization of periodic task sets in uni-processor systems
abstract
In this paper we introduce a novel characterization of real-time tasks based on their temperature impact. This characterization is used to analyze the schedulability of periodic real-time tasks in thermally constrained systems. The proposed characterization is important because thermal constraints are becoming increasingly vital due to rapid and increasing rise in power densities of modern architectures. As part of this work, we introduce the concept of “Accumulated Thermal Impact” (ATI), which represents cumulative temperature increase due to execution of a given task. The concept of ATI is then used to determine thermal utilization of a periodic task set, which is somewhat analogous to traditional computation utilization, and perform schedulability analysis. We also propose a speed scaling scheme for minimizing thermal utilization/system temperature. Our results show that thermal utilization of a periodic task set is strongly correlated to its thermal feasibility. The proposed speed scaling scheme is also shown to perform significantly better than current schemes in terms of thermal utilization/temperature minimization.
Parameswaran Ramanathan, Kewal K. Saluja
RTCSA3
2012 Linear Programming Formulations for Thermal-Aware Test Scheduling of 3D-Stacked Integrated Circuits
abstract
With technology scaling towards smaller geometries, the power density of modern integrated circuits (ICs) can potentially result into high temperatures during test, a problem further compounded by stacking dies in 3D stacked structures (3DSICs). Scheduling tests in a way to minimize the total test time becomes a key issue when temperature constraints are involved, since a more compact schedule leads to a hotter device. Unfortunately, many previous attempts at temperature-bounded scheduling either use inferior temperature models leading to under compaction, or they can only be applied to traditional single-die designs. Simple thermal models based on steady state temperatures are inadequate to schedule tests in 3DSICs due to their limitations. This paper proposes two formulations for test scheduling under thermal constraints for 3DSICs using the superposition principle, which allows for accurate thermal modeling and superior test compaction. This paper then compares them to previous formulations which use steady-state models, and also discusses the inherent limitations of the steady-state model. Results of the algorithms proposed in this paper show the superiority of the schedules obtained for testing 3DSICs.
Spencer K. Millican, Kewal K. Saluja
Asian Test Symposium2
2012 Diagnosis for Bridging Faults on Clock Lines
abstract
This paper presents diagnosis methods for bridging faults between a clock line and a gate signal line. Scan-based simulation methods are applied while assuming that only scan-based flush tests are used. In view of the fact that initial states play an important role, we consider two possible scenarios: 1) all flip-flops are assumed to be reset table, and 2) flip-flops are not reset table. In order to handle unknown states due to the non-reset table flip-flops, we introduce heuristic techniques. The effectiveness of the proposed methods are evaluated by the experimental results for benchmark circuits.
Yoshinobu Higami, Hiroshi Takahashi, Shin-ya Kobayashi, Kewal K. Saluja
PRDC4
2011 Fault simulation and test generation for clock delay faults
abstract
In this paper, we investigate the effects of delay faults on clock lines under launch-on-capture test strategy. In this fault model we assume that scan-in and scan-out operations, being relatively slow, can perform correctly even in the presence of a fault. However, a flip-flop may fail to capture a value at correct timing during system clock operation, thus requiring the use of launch-on-capture test strategy to detect such a fault. In the paper, we first show simulation results providing a relation between the duration of the delay and difficulty of detecting such faults in the launch-on-capture test. Next, we propose test generation methods to detect such clock delay faults, and show some experimental results to establish the effectiveness of our methods.
Yoshinobu Higami, Hiroshi Takahashi, Shin-ya Kobayashi, Kewal K. Saluja
ASP-DAC4
2011 On Detecting Transition Faults in the Presence of Clock Delay Faults
abstract
Shrinking timing margins for modern high speed digital circuits require a careful reconsideration of faults and fault models. In this paper, we discuss detection of transition faults in the presence of small clock delay faults. We first show that in the presence of a delay fault on a clock line some transition faults may fail to be detected. We propose a test generation method for detecting such faults (simultaneous presence of two faults) which consist of a gate transition fault and a clock delay fault assuming launch-on-capture test environment. The proposed test generation method employs a standard stuck-at ATPG tool. In our test generation methodology, the conditions for detecting a clock delay fault are converted into those for detecting a stuck-at fault, by adding some modeling logic during the ATPG process. Experimental results for benchmark circuits show the effectiveness of the proposed methods.
Yoshinobu Higami, Hiroshi Takahashi, Shin-ya Kobayashi, Kewal K. Saluja
Asian Test Symposium4
2011 Temperature Dependent Test Scheduling for Multi-core System-on-Chip
abstract
Recent research has shown that some defects are detect resilient under normal or high temperature, therefore tests for those defects must be applied under lower temperature. On the other hand, some tests need to be applied under high temperature to improve the detection sensitivity. Thus temperature dependent testing which applies tests at different temperature ranges is needed. This paper discusses and gives a formulation of the temperature dependent test scheduling problem. In the proposed test scheduling scheme, each test is associated with a lower temperature bound and an upper temperature bound to define the temperature range within which the test must be applied. A list schedule based test scheduling algorithm is proposed to find the earliest starting time of each test. Cooling period is inserted when the core temperature is too high and heating sequence is applied when the core temperature is below the required specified temperature for the core. Simulation studies are performed for ITC'02 SoC benchmarks and test scheduling results are shown.
Chunhua Yao, Kewal K. Saluja, Parameswaran Ramanathan
Asian Test Symposium2
2011 Indirect detection of clock skew induced hold-time violations on functional paths using scan shift operations
abstract
Hold-time violations in a scan circuit may occur both in the scan chain and in its combinational logic part. If a hold-time violation occurs on the scan path from one scan cell to another, it is also likely to happen on short functional paths between the two cells, which have clock skew. This paper is intended to indirectly detect hold-time violations on short functional paths using scan shift operations. A greedy approach to scan chain ordering is presented to detect as many hold-time violations as possible using scan shift operations. Extensive experiments are conducted to detect hold-time violations on short functional paths of various lengths by scan shift operations. Experimental results show that many hold-time violations on short functional paths can be detected by choosing an appropriate order of the scan cells in the scan chain.
Tsuyoshi Iwagaki, Kewal K. Saluja
DDECS2
2011 Enhancement of Clock Delay Faults Testing
abstract
This paper addresses the problem of simultaneous presence of multiple faults consisting of clock delay and gate transitions faults. The conditions of detecting a target multiple fault are converted into those for detecting a single stuck-at fault by adding some logic during the ATPG process. Experimental results show the effectiveness of our method by achieving nearly 100% fault efficiency.
Yoshinobu Higami, Hiroshi Takahashi, Shin-ya Kobayashi, Kewal K. Saluja
ETS4
2011 Adaptive execution assistance for multiplexed fault-tolerant chip multiprocessors
abstract
Relentless scaling of CMOS fabrication technology has made contemporary integrated circuits increasingly susceptible to transient faults, wearout-related permanent faults, intermittent faults and process variations. Therefore, mechanisms to mitigate the effects of decreased reliability are expected to become essential components of future general-purpose microprocessors. In this paper, we introduce a new throughput-efficient architecture for multiplexed fault-tolerant chip multiprocessors (CMPs). Our proposal relies on the new technique of adaptive execution assistance, which dynamically varies instruction outcomes forwarded from the leading core to the trailing core based on measures of trailing core performance. We identify policies and design low overhead hardware mechanisms to achieve this. Our work also introduces a new priority-based thread-scheduling algorithm for multiplexed architectures that improves multiplexed fault tolerant CMP throughput by prioritizing stalled threads. Through simulation-based evaluation, we And that our proposal delivers 17.2% higher throughput than perfect dual modular redundant (DMR) execution and outperforms previous proposals for throughput-efficient CMP architectures.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
ICCD3
2011 Adaptive cost efficient deployment strategy for homogeneous wireless camera sensors
Kewal K. Saluja, Seapahn Megerian
Ad Hoc Networks2
2011 Calibrating On-chip Thermal Sensors in Integrated Circuits: A Design-for-Calibration Approach
Chunhua Yao, Kewal K. Saluja, Parameswaran Ramanathan
J. Electron. Test.2
2011 Routing TCP Flows in Underwater Mesh Networks
abstract
Due to the growing importance of coastline surveillance and protection, underwater communication is playing an increasingly important role in military networks. As compared to terrestrial networks, large propagation delays and low data rates are fundamental characteristics of underwater communication. Furthermore, due to some key differences in the factors causing fluctuations in the quality of the underwater channels, the corresponding communication links experience more prolonged data rate changes as compared to those in terrestrial networks. Large propagation delays and prolonged link data rate deteriorations severely degrade the end-to-end performance of Transmission Control Protocol (TCP) based applications. Since military applications often require the reliable data delivery provided by TCP, it is important to devise solutions to alleviate this problem. In this paper, we propose a new routing scheme called Linear Coded Digraph Routing (LCDR) to enhance the end-to-end throughput of TCP based packet flows in underwater mesh networks. LCDR is a fully distributed scheme designed to locally respond to changes in the link data rates. In LCDR, each ingress node forwards packets after network coding. Each intermediate node adaptively uses network coding before forwarding the packets to the outgoing links. Each terrestrial gateway decodes the network coded packets before forwarding them to terrestrial networks. Each node adapts its packet forwarding rate based on the available bandwidth on the outgoing links, such that the terrestrial gateway can successfully receive packets with higher probability without significantly affecting cross-traffic. The effectiveness of the proposed scheme is evaluated using simulation. The simulation results show that the proposed scheme uses the spare bandwidth on each link efficiently and it significantly improves end-to-end throughput of TCP flows.
Chin-Ya Huang, Parameswaran Ramanathan, Kewal K. Saluja
IEEE J. Sel. Areas Commun.3
2011 Power and Thermal Constrained Test Scheduling Under Deep Submicron Technologies
abstract
Conventional power constrained test scheduling methods do not guarantee a thermal-safe solution. In this paper, we propose a test scheduling algorithm that satisfies the resource, power, and thermal constraints. First, in contrast to existing schemes, the proposed algorithm exploits superposition principle to perform fast and accurate thermal simulation, which, in turn, allows the algorithm to search for solutions which introduce cooling periods between tests to reduce the overall test length. Second, we propose a test partition-based method to further improve the performance of the test scheduling. We apply our test scheduling algorithm to ITC'02 SoC benchmarks and the results show considerable improvement in the total test length over existing methods.
Chunhua Yao, Kewal K. Saluja, Parameswaran Ramanathan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Controlling Peak Power Consumption for Scan Based Multiple Weighted Random BIST
abstract
This paper presents a multiple weight set weighted random BIST scheme to perform power aware test which takes into account peak power constraint while reducing the overall energy consumption. Generally, exceeding the peak power budget during test provokes problems such as erroneous circuit function due to IR drop and electrical damage to the circuit. On the other hand, excessive power reduction during test can degrade fault screening capability of the test since the test environment becomes substantially less stringent than the environment in which the device is actually used. In this paper we address both these problem simultaneously. The main contribution of this paper is to provide an efficient scan based weighted random BIST scheme within peak power constraint and adequate energy consumption. In order to restrict the amount of test power consumption, fault clustering is performed. The size of the fault cluster is changed to adjust the amount of power consumption. We assume full scan environment, test per scan testing, and single stuck-at fault model. Genetic algorithm (GA) based optimization method is proposed to find effective weight sets.
Hiroshi Yokoyama, Hideo Tamamoto, Kewal K. Saluja
Asian Test Symposium3
2010 Post-silicon diagnosis of segments of failing speedpaths due to manufacturing variations
abstract
We study diagnosis of segments on speedpaths that fail the timing constraint at the post-silicon stage due to manufacturing variations. We propose a formal procedure that is applied after isolating the failing speedpaths which also incorporates post-silicon path-delay measurements for more accurate analysis. Our goal is to identify segments of the failing speedpaths that have a post-silicon delay larger than their estimated delays at the pre-silicon stage. We refer to such segments as "failing segments" and we rank them according to their degree of failure. Diagnosis of failing segments alleviates the problem of lack of observability inside a path. Moreover, root-cause analysis, and post-silicon tuning or repair, can be done more effectively by focusing on the failing segments. We propose an Integer Linear Programming formulation to breakdown a path into a set of non-failing segments, leaving the remaining to be likely-failing ones. Our algorithm yields a very high "diagnosis resolution" in identifying failing segments, and in ranking them.
Azadeh Davoodi, Kewal K. Saluja
DAC3
2010 Multiplexed redundant execution: A technique for efficient fault tolerance in chip multiprocessors
abstract
Continued CMOS scaling is expected to make future microprocessors susceptible to transient faults, hard faults, manufacturing defects and process variations causing fault tolerance to become important even for general purpose processors targeted at the commodity market. To mitigate the effect of decreased reliability, a number of fault-tolerant architectures have been proposed that exploit the natural coarse-grained redundancy available in chip multiprocessors (CMPs). These architectures execute a single application using two threads, typically as one leading thread and one trailing thread. Errors are detected by comparing the outputs produced by these two threads. These architectures schedule a single application on two cores or two thread contexts of a CMP. As a result, besides the additional energy consumption and performance overhead that is required to provide fault tolerance, such schemes also impose a throughput loss. Consequently a CMP which is capable of executing 2n threads in non-redundant mode can only execute half as many (n) threads in fault-tolerant mode. In this paper we propose multiplexed redundant execution (MRE), a low-overhead architectural technique that executes multiple trailing threads on a single processor core. MRE exploits the observation that it is possible to accelerate the execution of the trailing thread by providing execution assistance from the leading thread. Execution assistance combined with coarse-grained multithreading allows MRE to schedule multiple trailing threads concurrently on a single core with only a small performance penalty. Our results show that MRE increases the throughput of fault-tolerant CMP by 16% over an ideal dual modular redundant (DMR) architecture.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
DATE3
2010 Energy-efficient fault tolerance in chip multiprocessors using Critical Value Forwarding
abstract
Relentless CMOS scaling coupled with lower design tolerances is making ICs increasingly susceptible to wear-out related permanent faults and transient faults, necessitating on-chip fault tolerance in future chip microprocessors (CMPs). In this paper we introduce a new energy-efficient fault-tolerant CMP architecture known as Redundant Execution using Critical Value Forwarding (RECVF). RECVF is based on two observations: (i) forwarding critical instruction results from the leading to the trailing core enables the latter to execute faster, and (ii) this speedup can be exploited to reduce energy consumption by operating the trailing core at a lower voltage-frequency level. Our evaluation shows that RECVF consumes 37% less energy than conventional dual modular redundant (DMR) execution of a program. It consumes only 1.26 times the energy of a non-fault-tolerant baseline and has a performance overhead of just 1.2%.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
DSN3
2010 Modified T-Flip-Flop based scan cell for RAS
abstract
Testing using a random access scan (RAS) design-for-test approach is experiencing renewed interest because of the potential for lower test application time, low power dissipation, and low test data volume compared to standard serial scan. In this paper we propose a significant modification and enhancement to the T-Flip-Flop based cell design for Random Access Scan (RAS). Importantly, the new RAS cell can allow the overlap of the test response read out with the loading of the next test input patterns within the same memory addressing cycle, thereby masking out the need for a separate memory cycle to read the test response in many cases. This can greatly reduce test application time. Experimental results show that the Modified T-Flip-Flop based scan cell is able to mask about 33% to 76% of reads. Further, this new RAS cell also eliminates the need for clock gating and additionally achieves reduction in gate overhead as much as about 20% compared to the existing T-flip-flop based RAS cell design.
Raghavendra Adiga, Gandhi Arpit, Virendra Singh, Kewal K. Saluja, Adit D. Singh
ETS4
2010 SEU tolerant SRAM for FPGA applications
abstract
Modern integrated circuits require careful attention to the soft errors resulting into bit upsets, which are normally caused by alpha particle or neutron hits. These events, also referred to as single-event upsets (SEUs), will become more severe for future technologies. LUT-based FPGAs are heavily using SRAM and there is a growing concern on correct operations of such FPGAs. Although there have been researches on enhancing fault tolerance of such FPGAs, they are based on TMR (triple modular redundancy) and simply too costly for normal application. In this paper we propose a novel 10T SEU tolerant SRAM cell and discuss its modifications for storage of configuration bits in FPGA so that reasonable protection against soft errors can be achieved with small area increase.
Sudipta Sarkar, Anubhav Adak, Virendra Singh, Kewal K. Saluja, Masahiro Fujita 0004
FPT4
2010 Energy-efficient redundant execution for chip multiprocessors
abstract
Relentless CMOS scaling coupled with lower design tolerances is making ICs increasingly susceptible to wear-out related permanent faults and transient faults, necessitating on-chip fault tolerance in future chip microprocessors (CMPs). In this paper, we describe a power-efficient architecture for redundant execution on chip multiprocessors (CMPs) which when coupled with our per-core dynamic voltage and frequency scaling (DVFS) algorithm significantly reduces the energy overhead of redundant execution without sacrificing performance. Our evaluation shows that this architecture has a performance overhead of only 0.3% and consumes only 1.48 times the energy of a non-fault-tolerant baseline.
Pramod Subramanyan, Virendra Singh, Kewal K. Saluja, Erik Larsson
ACM Great Lakes Symposium on VLSI3
2010 Test application time minimization for RAS using basis optimization of column decoder
abstract
Random Access Scan, which addresses individual flip-flops in a design using a memory array like row and column decoder architecture, has recently attracted widespread attention, due to its potential for lower test application time, test data volume and test power dissipation when compared to traditional Serial Scan. This is because typically only a very limited number of random "care" bits in a test response need be modified to create the next test vector. Unlike traditional scan, most flip-flops need not be updated. Test application efficiency can be further improved by organizing the access by word instead of by bit. In this paper we present a new decoder structure that takes advantage of basis vectors and linear algebra to further significantly optimize test application in RAS by performing the write operations on multiple bits consecutively. Simulations performed on benchmark circuits show an average of 2-3 times speed up in test write time compared to conventional RAS.
A. Abhishek, Amanulla Khan, Virendra Singh, Kewal K. Saluja, Adit D. Singh
ISCAS4
2010 Detection of inter-port bridging faults in dual-port memories
abstract
This paper presents an approach and test sequence to detect inter-port bridging faults in dual-port memories. Unlike other approaches we model a fault as a four-way bridging fault which is more reflective of a real defect. In our test approach, we consider word- and bit- line bridges, both in structure and in functional modes. Further, faults for all scenarios namely, read-read, write-write, and read-write ports are considered. We also propose the use of additional logic, three wide-OR gates, to detect certain inter-port faults that may remain undetected otherwise. Our approach achieves 100% coverage of all inter-port bridging faults.
Ho-Yong Choi, Kewal K. Saluja
ISCAS2
2010 On techniques for handling soft errors in digital circuits
abstract
Dealing with soft errors due to particle strikes is the next major challenge in implementing digital systems. This study thoroughly investigates the effect of device size on circuit soft error rate and identifies methods to reduce soft error rate in combinational circuits. In particular, we propose three novel methods that upsize only selected gates and /or transistor networks. In order to obtain the most appropriate technique for soft error rate reduction in small technology node circuits, we conduct experiments and compare the results for several upsizing techniques including all gates, selected gates and transistor networks based on their fault sensitivities, and parallel networks with soft error rate saturation consideration. Consequently, it is discovered that some upsizing scenarios perform large improvement whereas others do not or even increase the soft error rate. The use of fault sensitivity analysis approach with parallel transistor network upsizing based on the contribution of each sensitive gate can reasonably reduce overall circuit sensitivity. Experimental results show an average reduction in soft error rate about 20% with a very small area overhead of 2% for benchmark circuits using our technique.
Warin Sootkaneung, Kewal K. Saluja
ITC2
2010 Connected Barrier Coverage on a Narrow Band: Analysis and Deployment
abstract
Barrier coverage indicates the capability of a deployed wireless sensor network to detect intruders crossing the sensing field, and it has been widely studied in recent years. Most of the existing works are asymptotic and focusing on the critical conditions (sensor density, sensing radius, etc.) to achieve barrier coverage. However these results are not very useful in practice since the sensing field generally has finite region. Also, the critical conditions may not be adequate for making deployment decisions if sensor cost and deployment cost are taken into consideration. In this paper we analyze the probability of achieving connected barrier coverage on a finite narrow band while sensors with given sensing/communicating radius are randomly deployed with given density. Moreover, we apply our analytical result and propose a cost efficient deployment strategy that uses minimal number of sensors to achieve connected barrier coverage within at most k iterations. Both the correctness of the analysis and the performance of the proposed deployment strategy are evaluated via simulations.
Kewal K. Saluja, Parameswaran Ramanathan
SECON2
2010 Modeling latency - lifetime trade-off for target detection in mobile sensor networks
abstract
Two important measures of performance for the surveillance applications of the mobile sensor networks are detection latency and system lifetime. Previous work on modeling detection delay has assumed that sensor measurements are delivered to the fusion center with zero delay. Such approaches can require excessive energy, resulting into reduced lifetime. This article argues that a trade-off between detection latency and system lifetime can be made by employing an energy aware transmission scheme. The article formulates the trade-off as an optimization problem, and presents an analytic method to model both detection latency and system lifetime. The model is substantiated by using simulation.
Parameswaran Ramanathan, Kewal K. Saluja
ACM Trans. Sens. Networks3
2009 Partition Based SoC Test Scheduling with Thermal and Power Constraints under Deep Submicron Technologies
abstract
For core-based system-on-chip (SoC) testing, conventional power-constrained test scheduling methods do not guarantee a thermal-safe solution. Also, most of the test scheduling schemes make poor assumptions about power consumption. In deep submicron era, leakage power and wake-up power consumption can not be neglected. In this paper, we propose a partition based thermal-aware test scheduling algorithm with more realistic assumptions of recent SoCs. In our test scheduling algorithm, each test is partitioned and the earliest starting time of each partition is searched. To reduce the execution time of thermal simulation, we also exploit superposition principle to compute the power and thermal profile rapidly and accurately. We apply our test scheduling algorithm to ITC'02 SoC benchmarks and the results show improvements in the total test time over scheduling schemes without partitioning.
Chunhua Yao, Kewal K. Saluja, Parameswaran Ramanathan
Asian Test Symposium2
2009 DX-compactor: distributed X-compaction for SoCs
abstract
The emergence of System-on-Chip (SoC) devices has led to a complex on-chip interconnect structure that consumes significant area. Distributed compaction is a test response compaction scheme for an SoC that aims at reducing the area occupied for the purpose of testing the chip. This technique involves the design of compactors for individual cores on the chip. These are interconnected suitably to achieve the required functionality while reducing the area overhead by decreasing the length and the number of interconnects that are routed from scan chain outputs to the output pins. The distributed compaction technique matches the performance of an X-Compactor in terms of error detection and X-masking.
Reshma C. Jumani, Niraj Bharatkumar Jain, Virendra Singh, Kewal K. Saluja
ACM Great Lakes Symposium on VLSI4
2009 Power and thermal constrained test scheduling
abstract
We propose a test scheduling algorithm that ensures the resource compatibility and satisfies both power and thermal constraints. The proposed algorithm can start a test at an arbitrary time and it has the capability of delaying a test to let a core cool down to find a valid schedule even when traditional scheduling schemes cannot find a solution. To reduce the execution time of thermal simulation, we exploit superposition principle to compute the thermal profile rapidly and accurately. We apply our scheduling algorithm to ITC'02 SoC benchmarks and the results show a remarkable improvement in the total test length over other methods, while meeting the thermal and power constraints.
Chunhua Yao, Kewal K. Saluja, Parameswaran Ramanathan
ITC2
2009 Blindly Calibrating Mobile Sensors Using Piecewise Linear Functions
abstract
Calibrating nonlinear mobile sensors in-field is a challenging task due to the unavailability of controlled signal field and pre-calibrated sensor devices. In this paper, we propose a Density Guided blind Calibration (DGC) scheme for nonlinear mobile sensors by approximating the nonlinear calibration functions using piecewise linear functions. The DGC scheme exploits the fact that sensors moving in the same region collect similar fraction of true values in any given interval over time. The proposed scheme tackles the nonlinear calibration problem through an optimization formulation which is very easy to solve. The effectiveness of the proposed scheme is verified through simulations and an experiment with MICA2 light sensors.
Parameswaran Ramanathan, Kewal K. Saluja
SECON3
2009 False Path Aware Timing Yield Estimation under Variability
abstract
Effects of fluctuations in circuit timing due to process and environmental variations are becoming increasingly important as we move into sub-45 nm technology. Since the delay of each gate is dependent on its input vectors, the timing yield, the probability that the circuit meets the given timing constraint, varies with different primary input patterns. Traditional timing yield estimation approaches assumed worst case delay models for each gate over all its input vectors, which results in much pessimism. To overcome the aforementioned problems, this paper proposes a Monte Carlo based approach which can obtain a much tighter lower bound on the circuit timing yield compared to the existing timing yield estimation techniques. Specifically, our approach builds multiple input-vector-dependent variation-aware delay models for each logic gate, and considers the impact of false paths, both static and dynamic false paths, which are carefully selected from the likely timing-critical paths under variability. We demonstrate gradual improvement in the estimated timing yield in the simulation results, and show that the timing yield computed using traditional worst-case delay models is highly pessimistic.
Azadeh Davoodi, Kewal K. Saluja, Abhishek A. Sinkar
VTS3
2009 Low-Area Wrapper Cell Design for Hierarchical SoC Testing
Kyuchull Kim, Kewal K. Saluja
J. Electron. Test.2
2009 Modeling Detection Latency with Collaborative Mobile Sensing Architecture
abstract
Detection latency, which is defined as the time from the target arrival to the time of the first detection, is an important metric for the performance of sensor networks carrying out target detection, especially when the target is malicious or hostile. It characterizes the efficiency of detecting the presence of a target in a region of interest. Traditionally, stationary sensor networks are used to perform such sensing tasks. Consequently, nearly all research literature for the target detection problem has focused on stationary sensor networks. This paper addresses the problem of detecting the presence/absence of a target using a mobile sensor network. An analytic method is proposed to model the detection latency based on a collaborative sensing architecture. Detection latency for different node mobility models is presented. The accuracy of the analytic model is verified by simulations. This paper also compares the performance of mobile and stationary sensor networks. The comparison shows that if the target is present at the worst possible location in a given deployment, then detection latency of mobile sensor networks is considerably shorter as compared to that of stationary networks with the same number of nodes.
Tai-Lin Chin, Parameswaran Ramanathan, Kewal K. Saluja
IEEE Trans. Computers3
2008 Increasing Defect Coverage by Generating Test Vectors for Stuck-Open Faults
abstract
Defects in the modern LSIs manufactured by the deep-submicron technologies are known to cause complex faulty phenomena. Testing by targeting only stuck-at or bridging faults is no longer sufficient. Yet, increasing defect coverage is even more important. A stuck-open fault model considers transistor level defects, many of which are not covered by a stuck-at fault model. Further, test vectors for stuck-open faults also have the ability to detect the defects modeled by delay faults. This paper presents test generation methods for stuck-open faults using stuck-at test vectors and stuck-at test generation tools. The resultant test vectors achieve high coverage of stuck open faults while maintaining the original stuck-at fault coverage, thus offering the benefit of potential better defect coverage. We consider two types of test application mechanisms, namely launch on capture test and enhanced scan test. The effectiveness of the proposed methods is established by experimental results for benchmark circuits.
Yoshinobu Higami, Kewal K. Saluja, Hiroshi Takahashi, Shin-ya Kobayashi, Yuzo Takamatsu
ATS2
2008 An accurate flip-flop selection technique for reducing logic SER
abstract
The combination of continued technology scaling and increased on-chip transistor densities has made vulnerability to radiation induced soft errors a significant design concern. In particular, the effects of these errors on logic nodes are predicted to play an increasingly large role in determining the overall failure rate of future VLSI chips. While a myriad of techniques have been proposed to mitigate the effects of soft errors, system designers must ensure that the application of these solutions does not come at the expense of other design goals. This work presents a heuristic to selectively apply temporal redundancy to flip-flops within a pipelined logic unit, achieving significant reductions in failures associated with soft errors with minimal overhead.
Eric L. Hill, Mikko H. Lipasti, Kewal K. Saluja
DSN3
2008 A Capture-Safe Test Generation Scheme for At-Speed Scan Testing
abstract
Capture-safety, defined as the avoidance of any timing error due to unduly high launch switching activity in capture mode during at-speed scan testing, is critical for avoiding test- induced yield loss. Although point techniques are available for reducing capture IR-drop, there is a lack of complete capture-safe test generation flows. The paper addresses this problem by proposing a novel and practical capture-safe test generation scheme, featuring (1) reliable capture-safety checking and (2) effective capture-safety improvement by combining X-bit identification & X-filling with low launch- switching-activity test generation. This scheme is compatible with existing ATPG flows, and achieves capture-safety with no changes in the circuit-under-test or the clocking scheme.
Xiaoqing Wen, Kohei Miyase, Seiji Kajihara, Hiroshi Furukawa, Yuta Yamato, Atsushi Takashima, Kenji Noda, Hideaki Ito, Kazumi Hatayama, Takashi Aikyo, Kewal K. Saluja
ETS11
2008 Moments Based Blind Calibration in Mobile Sensor Networks
abstract
In-field calibration of sensor devices is known to be a challenging problem because there is often no access to a controlled signal field and/or a pre-calibrated device to measure the existing signal field. In this paper, we describe a blind calibration scheme that is tailored for sensor networks with mobile nodes. The scheme proposed in this paper exploits the fact that sensor devices are moving in the same region and hence the signal statistics they observe over time are almost the same. Analysis and simulation results are included to demonstrate the effectiveness of the proposed scheme.
Parameswaran Ramanathan, Kewal K. Saluja
ICC3
2008 Implementing high availability memory with a duplication cache
abstract
High availability systems typically rely on redundant components and functionality to achieve fault detection, isolation and fail over. In the future, increases in error rates will make high availability important even in the commodity and volume market. Systems will be built out of chip multiprocessors (CMPs) with multiple identical components that can be configured to provide redundancy for high availability. However, the 100% overhead of making all components redundant is going to be unacceptable for the commodity market, especially when all applications might not require high availability. In particular, duplicating the entire memory like the current high availability systems (e.g. NonStop and Stratus) do is particularly problematic given the fact that system costs are going to be dominated by the cost of memory. In this paper, we propose a novel technique called a duplication cache to reduce the overhead of memory duplication in CMP-based high availability systems. A duplication cache is a reserved area of main memory that holds copies of pages belonging to the current write working set (set of actively modified pages) of running processes. All other pages are marked as read-only and are kept only as a single, shared copy. The size of the duplication cache can be configured dynamically at runtime and allows system designers to trade off the cost of memory duplication with minor performance overhead. We extensively analyze the effectiveness of our duplication cache technique and show that for a range of benchmarks memory duplication can be reduced by 60-90% with performance degradation ranging from 1-12%. On average, a duplication cache can reduce memory duplication by 60% for a performance overhead of 4% and by 90% for a performance overhead of 5%.
Nidhi Aggarwal, James E. Smith 0001, Kewal K. Saluja, Norman P. Jouppi, Parthasarathy Ranganathan
MICRO3
2008 Calibrating Nonlinear Mobile Sensors
abstract
In-field calibration of sensor devices is known to be a challenging problem because there is often no access to a controlled signal field and/or a pre-calibrated device to provide the ground truth. Nonlinear characteristics of sensor devices make the calibration problem even harder. In this paper, we describe two blind calibration schemes for nonlinear mobile sensor nodes: nullspace based calibration (NBC) and moments based calibration (MBC). Simulation results are included to demonstrate the effectiveness of the proposed schemes. MBC scheme is also used to calibrate light sensors on MICA2 motes in a light field generated by a light bulb. Results show that significant error reduction can be achieved when nonlinearity is considered.
Parameswaran Ramanathan, Kewal K. Saluja
SECON3
2008 Low Capture Switching Activity Test Generation for Reducing IR-Drop in At-Speed Scan Testing
Xiaoqing Wen, Kohei Miyase, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita
J. Electron. Test.6
2007 Test Generation for Transistor Shorts using Stuck-at Fault Simulator and Test Generator
abstract
Test generation methods for transistor shorts using logic test environment are proposed. The fault models used are strong shorts and weak shorts, introduced in our earlier work. Our methodology consists of fault simulation, test generation and test compaction using gate-level tools to detect transistor faults but without resorting to use of transistor-level tools.
Yoshinobu Higami, Kewal K. Saluja, Hiroshi Takahashi, Shin-ya Kobayashi, Yuzo Takamatsu
ATS2
2007 Critical-Path-Aware X-Filling for Effective IR-Drop Reduction in At-Speed Scan Testing
abstract
IR-drop-induced malfunction is mostly caused by timing violations on activated critical paths during the capture cycle of at-speed scan testing. A critical-path-aware X-filling method is proposed for reducing IR-drop, especially on gates that are close to activated critical paths, thus effectively preventing test-induced yield loss.
Xiaoqing Wen, Kohei Miyase, Seiji Kajihara, Yuji Ohsumi, Kewal K. Saluja
DAC6
2007 Diagnosing At-Speed Scan BIST Circuits Using a Low Speed and Low Memory Tester
abstract
Numerous solutions have been proposed to reduce test data volume and test application time during manufacturing testing of digital devices. However, time to market challenge also requires a very efficient debug phase. Error identification in the test responses can become impractically slow in the debug phase due to large debug data, slow tester speed, and limited memory of the tester. In this paper, we investigate the problems and solutions related to using a relatively slow and limited memory tester to observe the at-speed behavior of fast circuits. Our method can identify all errors in at-speed scan BIST environment without any aliasing and using only little extra overhead by way of a multiplexer and masking circuit for diagnosis. Our solution takes into account the relatively slower speed of the tester and the reload time of the expected data to the tester memory due to limited tester memory while reducing the test/debug cost. Experimental results show that the test application time by our method can be reduced by a factor of 10 with very little hardware overhead to achieve such advantage.
Yoshiyuki Nakamura, Thomas Clouqueur, Kewal K. Saluja, Hideo Fujiwara
IEEE Trans. Very Large Scale Integr. Syst.3
2006 Compaction of pass/fail-based diagnostic test vectors for combinational and sequential circuits
abstract
Substantial attention is being paid to the fault diagnosis problem in recent test literature. Yet, the compaction of test vectors for fault diagnosis is little explored. The compaction of diagnostic test vectors must take care of all fault pairs that need to be distinguished by a given test vector set. Clearly, the number of fault pairs is much larger than the number of faults thus making this problem very difficult and challenging. The key contributions of this paper are: 1) to use techniques for reducing the size of fault pairs to be considered at a time, 2) to use novel variants of the fault distinguishing table method for combinational circuits and reverse order restoration method for sequential circuits, and 3) to introduce heuristics to manage the space complexity of considering all fault pairs for large circuits. Finally, the experimental results for ISCAS benchmark circuits are presented to demonstrate the effectiveness of the proposed methods
Yoshinobu Higami, Kewal K. Saluja, Hiroshi Takahashi, Shin-ya Kobayashi, Yuzo Takamatsu
ASP-DAC2
2006 Diagnosis of Transistor Shorts in Logic Test Environment
abstract
For deep-sub micron technology based LSIs, conventional stuck-at fault model is no longer sufficient for fault test and diagnosis. This paper presents a method of fault diagnosis for transistor shorts in combinational and full-scan circuits under logic test environment. Description of a short requires a very large number of physical parameters, and hence it is difficult, if not impossible, to describe precisely the behavior of transistor shorts. Therefore, two types of transistor short models were defined and algorithms to address the diagnostic problem were developed. The novelty of the algorithms is that they use conventional stuck-at fault simulation methodologies to diagnose transistor level shorts. Experiments were conducted on benchmark circuits to demonstrate the effectiveness of the method
Yoshinobu Higami, Kewal K. Saluja, Hiroshi Takahashi, Sin-ya Kobayashi, Yuzo Takamatsu
ATS2
2006 Diagnosing At-Speed Scan BIST Circuits Using a Low Speed and Low Memory Tester
abstract
Numerous solutions have been proposed to reduce test data volume and test application time during manufacturing testing of digital devices. However, time to market challenge also requires a very efficient debug phase. Error identification in the test responses can become impractically slow in the debug phase due to large debug data, slow tester speed and limited memory of the tester. In this paper, the authors investigate how a relatively slow and limited memory tester can observe the at-speed behavior of fast circuits. Our method can identify all errors in at-speed scan BIST environment without any aliasing and negligible extra hardware while taking into account the relatively slower speed of the tester and the re-load time of the expected data to the tester memory due to limited tester memory. Experimental results show that the test application time by our method can be reduced by a factor of 10 with very little hardware overhead to achieve such advantage
Yoshiyuki Nakamura, Thomas Clouqueur, Kewal K. Saluja, Hideo Fujiwara
ATS3
2006 Optimal Sensor Distribution for Maximum Exposure in A Region with Obstacles
abstract
Sensor networks have been envisioned to enhance the ability of human beings in observing the environment and understanding the world. A potential application of a sensor network is to detect the presence or absence of a target in a region of interest. Many heuristics have been proposed in literature for placing sensors to achieve better coverage in the monitored region. However, none of them guarantee an optimal sensor deployment especially when there are obstacles in the region. Unlike the prior work, this paper focuses on the problem of determining the optimal sensor distribution in a region with or without obstacles. The detection performance is characterized using a metric called ldquoexposurerdquo, which is defined as the least probability of detecting a target over all possible target locations subject to a fixed false alarm probability. A linear programming based approach is proposed to find the optimal sensor distribution by maximizing the exposure in a given region with or without obstacles. The optimal sensor distribution can also be used as weights of sensor measurements taken at different locations for decision-making.
Tai-Lin Chin, Parameswaran Ramanathan, Kewal K. Saluja
GLOBECOM3
2006 Highly-Guided X-Filling Method for Effective Low-Capture-Power Scan Test Generation
abstract
X-filling is preferred for low-capture-power scan test generation, since it reduces IR-drop-induced yield loss without the need of any circuit modification. However, the effectiveness of previous X-filling methods suffers from lack of guidance in selecting targets and values for X-filling. This paper addresses this problem with a highly-guided X-filling method based on two novel concepts: (1) X-score for X-filling target selection and (2) probabilistic weighted capture transition count for Y-filling value selection. Experimental results show the superiority of the new X-filling method for capture power reduction.
Xiaoqing Wen, Kohei Miyase, Yuta Yamato, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja
ICCD7
2006 Analytic modeling of detection latency in mobile sensor networks
abstract
An envisioned usage of sensor networks is in surveillance systems for detecting a target or monitoring a physical phenomenon in a region. Traditionally, stationary sensor networks are deployed to carry out the sensing operations. In many applications, if the monitored region is relatively large compared to the sensing range of a node, a large number of nodes are required in the region to achieve high coverage. Using mobile nodes in such situations can be an attractive alternative. Mobility of sensor nodes has been studied in sensor networks for many purposes such as power saving, data collection, and packet delivery. However, nearly all research literature for the target detection problem has focused on stationary sensor networks. This paper investigates the problem of detecting the presence/absence of a target using mobile sensor networks. It presents an analytic method to evaluate the detection latency based on a collaborative sensing approach using nodes with uncoordinated mobility. We verify the analytic model through simulations. The analytic method provides a simple way of analyzing the tradeoff between number of nodes and detection latency in a mobile sensor network. The analysis is also used to compare the performance of mobile and stationary sensor networks with respect to these measures. Results show that if the target is present at the worst possible location in a given deployment, then detection latency of mobile sensor networks is considerably less as compared to that of stationary networks with the same number of nodes.
Tai-Lin Chin, Parameswaran Ramanathan, Kewal K. Saluja
IPSN3
2006 A New ATPG Method for Efficient Capture Power Reduction During Scan Testing
abstract
High power dissipation can occur when the response to a test vector is captured by flip-flops in scan testing, resulting in excessive JR drop, which may cause significant capture-induced yield loss in the DSM era. This paper addresses this serious problem with a novel test generation method, featuring a unique algorithm that deterministically generates test cubes not only for fault detection but also for capture power reduction. Compared with previous methods that passively conduct X-filling for unspecified bits in test cubes generated only for fault detection, the new method achieves more capture power reduction with less test set inflation. Experimental results show its effectiveness.
Xiaoqing Wen, Seiji Kajihara, Kohei Miyase, Kewal K. Saluja, Laung-Terng Wang, Khader S. Abdel-Hafez, Kozo Kinoshita
VTS5
2006 Instruction-Based Self-Testing of Delay Faults in Pipelined Processors
abstract
Aggressive processor design methodology using high-speed clock and deep submicrometer technology is necessitating the use of at-speed delay fault testing. Although nearly all modern processors use pipelined architecture, no method has been proposed in literature to model these for the purpose of test generation. This paper proposes a graph theoretic model of pipelined processors and develops a systematic approach to path delay fault testing of such processor cores using the processor instruction set. The proposed methodology generates test vectors under the extracted architectural constraints. These test vectors can be applied in functional mode of operation, hence, self-test becomes possible. Self-test in a functional mode can also be used for online periodic testing. Our approach uses a graph model for architectural constraint extraction and path classification. Test vectors are generated using constrained automatic test pattern generation (ATPG) under the extracted constraints. Finally, a test program consisting of an instruction sequence is generated for the application of generated test vectors. We applied our method to two example processors, namely a 16-bit 5-stage VPRO pipelined processor and a 32-bit pipelined DLX processor, to demonstrate the effectiveness of our methodology
Virendra Singh, Michiko Inoue, Kewal K. Saluja, Hideo Fujiwara
IEEE Trans. Very Large Scale Integr. Syst.3
2005 State-reuse Test Generation for Progressive Random Access Scan: Solution to Test Power, Application Time and Data Size
abstract
Three issues that are dominating test research today are test application time, test data volume and test power. Researchers have focused on these issues mostly considering the popular serial scan architecture for its relatively low hardware overhead while ignoring the fact that exponential drop in hardware cost offers opportunities for implementing a test architecture that previously may have been un-acceptable. This paper takes such a paradigm shift into account and studies the simultaneous solution of all three problems of serial scan by making use of progressive random access scan test architecture. This architecture only increases the hardware cost marginally while providing marked improvements for the three issues. This paper explains the test architecture and then develops a test generation methodology which reduces the test application time by nearly 75%, test data volume by 50% for the benchmark circuits. Above all, the architecture is inherently so efficient that it reduces the test power by nearly 99% or more of the test power consumption compared to serial scan.
Dong Hyun Baik, Kewal K. Saluja
Asian Test Symposium2
2005 A Class of Linear Space Compactors for Enhanced Diagnostic
abstract
Testing of VLSI circuits is challenged by the increasing volume of test data that adds constraints on tester memory and impacts test application time substantially. Space compactors are commonly used to reduce the test volume by one or two orders of magnitude. However, such level of compaction reduces the quality of the diagnostic of faults because it is difficult to identify the locations of errors in the compacted response. In this paper, we introduce a design of space compactors that can be used in pass/fail mode as well as in diagnostic mode with enhanced performance by trading off compaction ratio for diagnostic ability. We analyze the properties of the compactors and evaluate their performance through simulations
Thomas Clouqueur, Hideo Fujiwara, Kewal K. Saluja
Asian Test Symposium3
2005 Testing Superscalar Processors in Functional Mode
abstract
This paper presents a methodology for testing a superscalar processor using functional mode of operation for the performance oriented delay faults. The functional mode test issues for superscalar are discussed. A graph based model is developed and used to develop for the generation of test programs.
Virendra Singh, Michiko Inoue, Kewal K. Saluja, Hideo Fujiwara
FPL3
2005 Progressive random access scan: a simultaneous solution to test power, test data volume and test time
abstract
Traditional testing research for testing VLSI circuits has been confined to the use of serial scan test architecture whose origin lies in keeping the hardware overhead low. However, there has been a paradigm shift in the cost factor - the transistor cost has been dropping exponentially whereas the test cost is starting to increase. We believe that adding marginally more hardware is acceptable provided the test cost can be reduced considerably. This paper takes such a view of testing and rejuvenates the random access scan as a design for testability method that simultaneously addresses three limitations of the traditional serial scan namely, test data volume, test application time, and test power. The novelty of the progressive random access scan approach proposed in this paper lies in developing the test architecture and formulating the test application time and test data volume reduction problems. We provide a traveling salesman formulation of these problems in our test architecture setting. Experimental results show the practicality of our approach as the hardware cost components, consisting of routing and transistor count, increase only marginally compared to the serial scan approach whereas there is a dramatic decrease in test power consumption (nearly a 1000 fold decrease in average test power) as well as the test data volume and the test times are halved.
Dong Hyun Baik, Kewal K. Saluja
ITC2
2005 Design and analysis of multiple weight linear compactors of responses containing unknown values
abstract
Occurrence of unknown values in scan chains in response to test vectors is a common phenomenon. This paper presents a method for designing matrices for linear test output compactors by using rows of multiple weights. Compared to previously proposed compactors, the method reduces the masking caused by unknowns by an order of magnitude provided that the unknowns are non-uniformally distributed among the scan chains. Also, using multiple rather than single weight compactors increases the compaction ratio and reduces the hardware overhead. The effectiveness of multiple weight compactors is demonstrated through analysis, simulations and experiments with test response from an industrial design.
Thomas Clouqueur, Kamran Zarrineh, Kewal K. Saluja, Hideo Fujiwara
ITC3
2005 Low-capture-power test generation for scan-based at-speed testing
abstract
Scan-based at-speed testing is a key technology to guarantee timing-related test quality in the deep submicron era. However, its applicability is being severely challenged since significant yield loss may occur from circuit malfunction due to excessive IR drop caused by high power dissipation when a test response is captured. This paper addresses this critical problem with a novel low-capture-power X-filling method of assigning 0's and 1's to unspecified (X) bits in a test cube obtained during ATPG. This method reduces the circuit switching activity in capture mode and can be easily incorporated into any test generation flow to achieve capture power reduction without any area, timing, or fault coverage impact. Test vectors generated with this practical method greatly improve the applicability of scan-based at-speed testing by reducing the risk of test yield loss.
Xiaoqing Wen, Yoshiyuki Yamashita, Shohei Morishima, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita
ITC6
2005 Exposure for collaborative detection using mobile sensor networks
abstract
Sensor networks possess the inherent potential to detect the presence of a target in a monitored region. Although a stationary sensor network is often adequate to meet application requirements, it is not suited to many situations, for example, a huge number of nodes are required to monitor a large region. In such situations, mobile sensor networks can be used to resolve the communication and sensing coverage problems. This paper addresses the problem of detecting a target using mobile sensor networks. One of the fundamental issues in target detection problems is exposure, which measures how the region is covered by the sensor network. While traditional studies focus on stationary sensor networks, this paper formally defines and evaluates exposure in mobile sensor networks with the presence of obstacles and noise. To conform with practical situations, detection is conducted without presuming the target's activities and moving directions. As there is no fixed layout of node positions, a time expansion technique is developed to evaluate exposure. Since determining exposure can be computationally expensive, algorithms to calculate the upper and lower bounds on exposure are developed. Simulation results are also presented to illustrate the effectiveness of the algorithms
Tai-Lin Chin, Parameswaran Ramanathan, Kewal K. Saluja, Kuang-Ching Wang
MASS3
2005 On Low-Capture-Power Test Generation for Scan Testing
abstract
Research on low-power scan testing has been focused on the shift mode, with little or no consideration given to the capture mode power. However, high switching activity when capturing a test response can cause excessive IR drop, resulting in significant yield loss. This paper addresses this problem with a novel low-capture-power X-filling method by assigning 0's and 1's to unspecified (X) bits in a test cube to reduce the switching activity in capture mode. This method can be easily incorporated into any test generation flow, where test cubes are obtained during ATPG or by X-bit identification. Experimental results show the effectiveness of this method in reducing capture power dissipation without any impact on area, timing, and fault coverage.
Xiaoqing Wen, Yoshiyuki Yamashita, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita
VTS5
2005 Fault Diagnosis of Physical Defects Using Unknown Behavior Model
Xiaoqing Wen, Hideo Tamamoto, Kewal K. Saluja, Kozo Kinoshita
J. Comput. Sci. Technol.3
2005 Combinational automatic test pattern generation for acyclic sequential circuits
abstract
It is known that the complexity of automatic test pattern generation (ATPG) for acyclic sequential circuits is similar to that of combinational ATPG. The general problem, however, requires time-frame expansion and multiple-fault detection and hence does not allow the use of available combinational ATPG programs. The first contribution of this work is a combinational single-fault ATPG method for the most general class of acyclic sequential circuits. Without inserting any real hardware, we create a functionally equivalent "balanced" ATPG model of the circuit in which all reconverging paths have the same sequential depth. Some primary inputs and gates are duplicated in this model, which is converted into a combinational circuit by shorting all flip-flops. A test vector obtained by a combinational ATPG program for a fault in this combinational circuit is transformed into a test sequence to detect a corresponding fault in the original sequential circuit. A combinational ATPG program finds tests for all but a small set of faults that must be explicitly detected as multiple-faults. Those are modeled for ATPG using the second contribution of this work, which is a generalized method to model any given multiple stuck-at fault as a single stuck-at fault. The procedure requires insertion of at most n+3 modeling gates for a fault of multiplicity n. We show that the modeled circuit is functionally equivalent to the original circuit and the targeted multiple fault is equivalent to the modeled single stuck-at fault. Benchmark results show at least an order of magnitude saving in the ATPG CPU time by the new combinational method over sequential ATPG.
Yong Chang Kim, Vishwani D. Agrawal, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 Optimizing program disturb fault tests using defect-based testing
abstract
Nonvolatile memories (NVMs) are susceptible to a special type of faults known as program disturb faults. These faults are described using logical fault models and often functional tests are used to detect different faults that occur under such models. The use of functional fault models and tests results in the simplification of the testing process, although such tests can be very long. In this paper, we present a defect-based model that can be used to model different disturb faults in NVM. The relationship between defect location and fault manifestation is first established using electrical simulation. Next, the use of stress tests and margin read schemes and how they are used to detect disturb faults is discussed. Using electrical simulation results, we show that defect-based testing can be used to optimize the cost of program disturb tests of NVM.
Mohammad AlFailakawi, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 A method for reducing the target fault list of crosstalk faults in synchronous sequential circuits
abstract
We describe a method of identifying a set of target crosstalk faults which may need to be tested in synchronous sequential circuits. Our method classifies the pairs of aggressor and victim lines, using topological and timing information, to deduce a set of target crosstalk faults. In this process, our method also identifies the false crosstalk faults that need not (and/or cannot) be tested in synchronous sequential circuits. Experimental results for ISCAS'89 and ITC'99 benchmark circuits show that the proposed method is CPU time efficient in obtaining the reduced lists of the target crosstalk faults. Also, the lists of the target crosstalk faults obtained by our method are substantially smaller than the sets of all possible combinations of faults.
Hiroshi Takahashi, Keith J. Keller, Kim T. Le, Kewal K. Saluja, Yuzo Takamatsu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2004 Enhanced 3-valued logic/fault simulation for full scan circuits using implicit logic values
abstract
When test vectors for a full scan logic circuit include unspecified values, conventional 3-valued fault simulation may not compute the exact fault coverage for single stuck-at faults. This paper first addresses the incompleteness of logic/fault simulation based on the conventional 3-valued logic. Then we propose an enhanced method of logic/fault simulation to compute more accurate fault coverage using implicit logic values. The proposed method employs indirect implications. We also propose a new learning criterion to identify indirect implications that are not identified by earlier static learning procedure. Since some indirect implications derived from a fault-free circuit become invalid in the presence of a fault, we use a sufficient condition for an indirect implication to remain valid for the faulty circuit, and give an efficient procedure for more accurate fault simulation. Experimental results demonstrate that the proposed method reduces the number of unknown values at the circuit outputs in logic simulation, and hence it discovers several detected faults that are not declared as detected by the conventional fault simulation.
Seiji Kajihara, Kewal K. Saluja, Sudhakar M. Reddy
ETS2
2004 A yield improvement methodology using pre- and post-silicon statistical clock scheduling
abstract
In deep sub-micron technologies, process variations can cause significant path delay and clock skew uncertainties thereby lead to timing failure and yield loss. In this paper, we propose a comprehensive clock scheduling methodology that improves timing and yield through both pre-silicon clock scheduling and post-silicon clock tuning. First, an optimal clock scheduling algorithm has been developed to allocate the slack for each path according to its timing uncertainty. To balance the skew that can be caused by process variations, programmable delay elements are inserted at the clock inputs of a small set of flip-flops on the timing critical paths. A delay-fault testing scheme combined with linear programming is used to identify and eliminate timing violations in the manufactured chips. Experimental results show that our methodology achieves substantial yield improvement over a traditional clock scheduling algorithm in many of the ISCAS89 benchmark circuits, and obtain an average yield improvement of 13.6%.
Jeng-Liang Tsai, Dong Hyun Baik, Charlie Chung-Ping Chen, Kewal K. Saluja
ICCAD4
2004 On per-test fault diagnosis using the X-fault model
abstract
This work proposes a new per-test fault diagnosis method based on the X-fault model. The X-fault model represents all possible behaviors of a physical defect or defects in a gate and/or on its fanout branches by using different X symbols on the fanout branches. A novel technique is proposed for analyzing the relation between observed and simulated responses to extract diagnostic information and to score the results of diagnosis. Experimental results show the effectiveness of our method.
Xiaoqing Wen, Tokiharu Miyoshi, Seiji Kajihara, Laung-Terng Wang, Kewal K. Saluja, Kozo Kinoshita
ICCAD5
2004 Testing of Hard Faults in Simultaneous Multithreaded Processors
Eric F. Weglarz, Kewal K. Saluja
IOLTS2
2004 Testing Micropipelined Asynchronous Circuits
abstract
Despite advances in the design of asynchronous circuits, little progress has been made in their testing or design for testability. This work proposes a new strategy for testing micropipelines, by treating the asynchronous elements, such as the C-element, as atomic state elements for testing purposes. By treating asynchronous elements as finite state machines, tests for these can be generated that verify their correct operation and also detect nearly all testable faults in the circuit. Design for testability methods for a micropipeline are also presented that reduce the amount of hardware added to the design while increasing its overall testability compared to other micropipeline testing methods.
Matthew L. King, Kewal K. Saluja
ITC2
2004 Fault Tolerance in Collaborative Sensor Networks for Target Detection
abstract
Collaboration in sensor networks must be fault-tolerant due to the harsh environmental conditions in which such networks can be deployed. We focus on finding algorithms for collaborative target detection that are efficient in terms of communication cost, precision, accuracy, and number of faulty sensors tolerable in the network. Two algorithms, namely, value fusion and decision fusion, are identified first. When comparing their performance and communication overhead, decision fusion is found to become superior to value fusion as the ratio of faulty sensors to fault free sensors increases. As robust data fusion requires agreement among nodes in the network, an analysis of fully distributed and hierarchical agreement is also presented. The impact of hierarchical agreement on communication cost and system failure probability is evaluated and a method for determining the number of tolerable faults is identified.
Thomas Clouqueur, Kewal K. Saluja, Parameswaran Ramanathan
IEEE Trans. Computers2
2003 Stress Test for Disturb Faults in Non-Volatile Memories
abstract
Non-volatile memories are susceptible to special type of faults known as program disturb faults. Testing for such faults requires the application of stress tests which have long application time to distinguish faulty cells from non-faulty cells. In this paper we present a new sensing scheme that can be used with stress tests to allow for efficient detection of faulty cells based on the notion of margin reads. We demonstrate the efficiency of the margin-read approach for distinguishing between faulty and fault-free cells using electrical simulations.
Mohammad AlFailakawi, Kewal K. Saluja
Asian Test Symposium2
2003 Outstanding Challenges in Testing Nanotechnology Based Integrated Circuits
abstract
Summary form only given. This presentation raises questions in all three dominant domains in the area of testing: traditional issues, power issues, and test application time issues. Traditional issues such as ATPG systems, fault simulators, and stuck-at faults are discussed, as well as new faults like capacitive and inductive crosstalks. The requirements in power delivery to the devices during normal operation as well as during test are expected to be largely in the area of test scheduling. Finally, the variable cost, the cost of testing each device, is likely to be dominated by the test application time. Hence, we need to develop novel solutions to the test application problem. This may mean re-looking at the most often used DFT and BIST methods. Also, we need to develop new and superior test generators and fault simulators that can handle faults from the new fault models for nanotechnologies and possibly deal with multiple faults.
Kewal K. Saluja
Asian Test Symposium1
2003 Software-Based Delay Fault Testing of Processor Cores
abstract
This paper presents a software-based self-testing methodology for delay fault testing. Delay faults affect the circuit functionality only when it can be activated in functional mode. A systematic approach or the generation of test vectors, which are applicable in functional mode, is presented. A graph theoretic model (represented by IE-Graph) is developed in order to model the datapath. A finite state machine model is used for the controller. These models are used for constraint extraction so that the generated test can be applied in functional mode.
Virendra Singh, Michiko Inoue, Kewal K. Saluja, Hideo Fujiwara
Asian Test Symposium3
2003 Fault Diagnosis for Physical Defects of Unknown Behaviors
abstract
This paper proposes an X-fault model for fault diagnosis of physical defects with unknown behaviors by using X symbols. An efficient X-fault simulation method and an efficient X-fault diagnostic reasoning method are presented. Based on these, an X-fault diagnosis method is described to improve the failure analysis for a wide range of physical defects in complex IC circuits.
Xiaoqing Wen, Hideo Tamamoto, Kewal K. Saluja, Kozo Kinoshita
Asian Test Symposium3
2003 Event-Centric Simulation of Crosstalk Pulse Faults in Sequential Circuits
abstract
The essence of existing methods to simulate crosstalk pulse faults in sequential circuits is the use of logic waveform on each line in the circuit. Explicitly keeping track of timing information is costly in terms of memory usage and computational effort. We propose and develop a novel approach to the simulation of crosstalk pulse faults due to coupling between aggressor lines and victim (flip-flop) clock lines. Our algorithm extends existing ideas fundamental to logic event-driven simulation to crosstalk faults excitation and fault grouping. In addition, our simulator treats issues related to timing in a more precise manner. Experimental results on ISCAS '89 benchmark circuits show extraordinary improvement on all fronts, including CPU time, fault coverage, and memory usage.
Marong Phadoongsidhi, Kewal K. Saluja
ICCD2
2003 Sensor Deployment Strategy for Detection of Targets Traversing a Region
Thomas Clouqueur, Veradej Phipatanasuphorn, Parameswaran Ramanathan, Kewal K. Saluja
Mob. Networks Appl.4
2002 Reduction of Target Fault List for Crosstalk-Induced Delay Faults by using Layout Constraints
abstract
We propose a method of identifying a set of crosstalk induced delay faults which may need to be tested in synchronous sequential circuits. During the fault list generation 1) we take into account all clocking effects, and 2) infer layout information front the logic level description. With regard to layout constraints we introduce two methods, namely the distance based layout constraint and the cone based layout constraint. The lists of the target faults obtained by the proposed methods are substantially smaller than the sets of all possible combinations of faults.
Keith J. Keller, Hiroshi Takahashi, Kim T. Le, Kewal K. Saluja, Yuzo Takamatsu
Asian Test Symposium4
2002 A Concurrent Fault Simulation for Crosstalk Faults in Sequential Circuits
abstract
Existing principles for crosstalk fault simulation require the storage of waveform representation at each node in the circuit throughout a time frame. At the end of each time frame a pair of waveforms, one belonging to an aggressor node, and one depicting a victim node, is inspected. If the fault is captured, it will be simulated until it is either detected or the test vectors are exhausted. This fault detection method can require a prohibitive amount of computation time for a large sequential circuit with high number of possible fault pairs to be tested. With our simulation technique, introduced in this paper, these operations can be processed concurrently for many faults. The fault list dynamically adjusts itself during the simulation to accommodate fault injection and fault dropping. Experimental results on ISCAS'89 benchmark circuits show that a substantial improvement in CPU time, over a conventional method, is achieved with a trade-off in the amount of memory consumed.
Marong Phadoongsidhi, Kim T. Le, Kewal K. Saluja
Asian Test Symposium3
2002 An Alternative Method of Generating Tests for Path Delay Faults Using N -Detection Test Sets
abstract
In order to generate tests for path delay faults we propose an alternative method that does not generate a test for each path delay fault directly. The proposed method generates an n-propagation test-pair set by using an N/sub i/-detection test set for single stuck-at faults. The n-propagation test-pair set is a set of vector pairs which contains n distinct vector pairs for every transition fault at a checkpoint (primary inputs and fanout branches in a circuit are called check points). We do not target the path delay faults for test generation, instead, the n-propagation test-pair set is generated for the transition (both rising and falling) faults of check points in the circuit, and simulated to determine their effectiveness for singly testable path delay faults and robust path delay faults. Results of experiments on the ISCAS'85 benchmark circuits show that the n-propagation test-pair sets obtained by our method are very effective in testing path delay faults.
Hiroshi Takahashi, Kewal K. Saluja, Yuzo Takamatsu
PRDC2
2002 On diagnosing multiple stuck-at faults using multiple and singlefault simulation in combinational circuits
abstract
Diagnosing multiple stuck-at faults in combinational circuits using singleand multiple-fault simulation is proposed. The proposed method adds (removes) faults from a set of suspected faults depending on the result of multiple-fault simulation at a primary output agreeing (disagreeing) with the observed value. However, the faults that are added or removed from the set of suspected faults are determined using single-fault simulation. Diagnosis is carried out by repeated addition and removal of faults. The effectiveness of the diagnosis method is evaluated by experiments conducted on benchmark circuits and it is found to be substantially superior compared to the previous known solutions. The method proposed in this paper can be used as a powerful tool at the preprocessing stage of diagnosis in an electron-beam tester environment.
Hiroshi Takahashi, Kwame Osei Boateng, Kewal K. Saluja, Yuzo Takamatsu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2001 Simulation-Based Diagnosis for Crosstalk Faults in Sequential Circuits
abstract
Describes two methods of diagnosing crosstalk-induced pulse faults in sequential circuits using crosstalk fault simulation. These methods compare with observed responses and simulated values at primary outputs to identify a set of suspected faults that are consistent with the observed responses. In these methods, if the simulated values agree with the observed responses, then the simulated fault is added to a set of suspected faults, otherwise the simulated fault is removed from the set of suspected faults. The diagnosis methods repeat the above process for each time frame to identify the suspected faults. The first method is a basic method which determines the suspected fault list by using the knowledge about the first and last failures of the test sequence. The second method uses state information and focuses on reducing the CPU time for diagnosing the faults. The CPU time is reduced by using stored state information to calculate the primary output values at the present time frame. Experimental results for ISCAS'89 benchmark circuits show that the number of suspected faults obtained by our methods is sufficiently small, and the second method is substantially faster than the first method.
Hiroshi Takahashi, Marong Phadoongsidhi, Yoshinobu Higami, Kewal K. Saluja, Yuzo Takamatsu
Asian Test Symposium4
2001 On reducing the target fault list of crosstalk-induced delay faults in synchronous sequential circuits
abstract
This paper describes a method of identifying a set of crosstalk-induced delay faults which may need to be tested in synchronous sequential circuits. In this process, the false crosstalk-induced delay faults that need not (and/or can not) be tested in synchronous sequential circuits are also identify. Our method classifies the pairs of aggressor and victim lines, using topological information and timing information, to deduce a set of faults that need to be tested in a sequential circuit. Experimental results for ISCAS'89 benchmark circuits show that the lists of the target faults obtained by the proposed method are sufficiently smaller than the sets of all possible combinations of faults.
Keith J. Keller, Hiroshi Takahashi, Kewal K. Saluja, Yuzo Takamatsu
ITC3
2001 Combinational test generation for various classes of acyclic sequential circuits
abstract
It is known that a class of acyclic sequential circuits called balanced circuits can be tested by combinational ATPG. The first contribution of this paper is a modified and efficient combinational single fault ATPG method for any general (not necessarily balanced) acyclic circuit. Without inserting real hardware, we create a "balanced" ATPG model of the circuit in which all reconverging paths have the same sequential depth. Some primary inputs are duplicated and each combinational A TPG vector for this model circuit is transformed into a test sequence. Although no time-frame expansion is used, a small set of faults still map onto multiple faults. Those are identified and dealt with again by the single fault combinational A TPG. The results show nearly an order of magnitude or greater saving in the A TPG CPU time over sequential ATPG. The second contribution consists of new partial-scan algorithms to obtain three subclasses of acyclic circuits, namely, internally balanced, balanced, and strongly balanced, which have been described in the literature. Results on ISCAS '89 circuits show that such structures require extra scan overhead, sometimes almost approaching that of full-scan, and their advantages in ATPG are marginal considering the present contribution.
Yong Chang Kim, Vishwani D. Agrawal, Kewal K. Saluja
ITC3
2001 Testable Sequential Circuit Design: A Partition and Resynthesis Approach
abstract
In this work, we present a divide and conquer approach to improve the testability of large sequential circuits while reducing the area overhead required. Specifically, we partition the circuits into more manageable size circuits for efficient synthesis with testability constraints. We resynthesize each circuit partition and restitch the partitions together to achieve our objectives. Experimental results are presented to demonstrate the effectiveness of our approach.
Richard M. Chou, Kewal K. Saluja
VTS2
2001 Flash Memory Disturbances: Modeling and Test
abstract
Nonvolatile Memories (NVMs) can undergo different types of disturbances. These disturbances are particular to the technology and the cell structure of the memory element. In this paper we develop a coupling fault model that appropriately models disturbances in flash memories that use floating gate transistor as their core memory element. We describe the behavior of faulty cells under different fault models and how their characteristics change under each model. We demonstrate the inappropriateness of conventional march algorithms for testing flash memories and present a procedure to derive pseudo-algorithms that can be used in testing flash memories. In addition we present an efficient test that detects these disturbances under different fault models developed in this paper.
Mohammad AlFailakawi, Kewal K. Saluja
VTS2
2001 Fault Models and Test Procedures for Flash Memory Disturbances
Mohammad AlFailakawi, Kewal K. Saluja, Alex S. Yap
J. Electron. Test.2
2000 Fault models and test generation for IDDQ testing: embedded tutorial
abstract
Abstract| This paper surveys recent research related to IDDQ testing, particularly focuses on fault models and test generation methods.(1) The paper pro videsa taxonomy of fault models that hav e been studied in literature, and classi es these models into a small set of faults.(2) The paper describes ecient test generation methods and fault simulation methods.T est compaction methods, including reduction of the total number of test v ectors and selection of IDDQ measurement v ectors, are also described.
Yoshinobu Higami, Yuzo Takamatsu, Kewal K. Saluja, Kozo Kinoshita
ASP-DAC3
2000 Fault Tolerance through Re-Execution in Multiscalar Architecture
abstract
Multi-threading and multiscaling are two fundamental microarchitecture approaches that are expected to stay on the existing performance gain curve. Both of these approaches assume that integrated circuits with over billion transistors will become available in the near future. Such large integrated circuits imply reduced design tolerances and hence increased failure probability. Conventional hardware redundancy techniques for desired reliability in computation may severely limit the performance of such high performance processors. Hence we need to study novel methods to exploit the inherent redundancy of the microarchitectures, without unduly affecting the performance, to provide correct program execution and/or detect failures (permanent or transient) that can occur in the hardware. This paper proposes a time redundancy technique suitable for multiscalar architectures. In the multiscalar architecture, there are usually several processing units to exploit the instruction level parallelism that exists in a given program. The technique in this paper uses a majority of the processing units for executing the program as in the traditional multiscalar paradigm while using the remainder of the processing units for re-executing the committed instructions. By comparing the results from the two program executions, errors caused by permanent or transient faults in the processing units can be detected. Simulation results presented in this paper demonstrate that this can be achieved with about 5-15% performance degradation.
Faisal Rashid, Kewal K. Saluja, Parameswaran Ramanathan
DSN2
2000 Algorithms to Select IDDQ Measurement Vectors for Bridging Faults in Sequential Circuits
Yoshinobu Higami, Yuzo Takamatsu, Kewal K. Saluja, Kozo Kinoshita
J. Electron. Test.3
1999 Fault Simulation Techniques to Reduce IDDQ Measurement Vectors for Sequential Circuits
abstract
This paper presents fault simulation techniques for selecting a small number of IDDQ measurement vectors from a given test sequence while maintaining the original fault coverage. The proposed method covers a class of bridging faults and uses parallel fault simulation wherever possible. Experimental results are presented to demonstrate the effectiveness of the proposed method.
Yoshinobu Higami, Yuzo Takamatsu, Kewal K. Saluja, Kozo Kinoshita
Asian Test Symposium3
1999 A Correlation Matrix Method of Clock Partitioning for Sequential Circuit Testability
abstract
We propose a method of partitioning the set of all flip-flops in a circuit for multiple clock testing. In the multiple clock testing, flip-flops are partitioned into different groups and each group of flip-flops has an independent clock control. In our method, we use a test generator assuming an independent clock control for each flip-flop. We than determine correlation between clock activity for all pairs of flip-flops. This information is then used to an optimal or near optimal partition of flip-flops in the circuit. Through experiments, we demonstrate that our partitioning method increases fault coverage and reduces test length with almost no hardware overhead or performance penalty.
Yong Chang Kim, Kewal K. Saluja, Vishwani D. Agrawal
Great Lakes Symposium on VLSI2
1998 Observation Time Reduction for IDDQ Testing of Briding Faults in Sequential Circuits
abstract
One of the major unsolved and ignored but significant problems is the reduction of the long testing time for IDDQ testing of CMOS circuits. Since IDDQ must be observed after dynamic current disappears, testing time is much longer than logic testing. This paper presents a method to reduce the observation time for IDDQ testing. The proposed method is a static method which focuses on selection of vectors to be observed instead of removing vectors. Experimental results are presented to demonstrate the effectiveness of the proposed method.
Yoshinobu Higami, Kewal K. Saluja, Kozo Kinoshita
Asian Test Symposium2
1998 Design for Diagnosability of CMOS Circuits
abstract
This paper presents a new approach to improving the diagnosability of a CMOS circuit by dividing it into independent partitions and using a separate power supply for each partition. This technique makes it possible to implement multiple I/sub DDQ/ measurement points. As a result, the diagnosability of the circuit can be improved. The problem of partitioning a circuit is addressed and optimum and heuristic solutions are proposed. The effectiveness of our approach is demonstrated through experimental results.
Xiaoqing Wen, Tooru Honzawa, Hideo Tamamoto, Kewal K. Saluja, Kozo Kinoshita
Asian Test Symposium4
1998 A Heuristic Measure to Maximize Detected Faults per Test
Kim T. Le, Kewal K. Saluja
J. Electron. Test.2
1998 Sequential test generators: past, present and future
Yong Chang Kim, Kewal K. Saluja
Integr.2
1998 A Novel Approach to Random Pattern Testing of Sequential Circuits
abstract
Random pattern testing methods are known to result in poor fault coverage for most sequential circuits unless costly circuit modifications are made. In this paper, we propose a novel approach to improve the random pattern testability of sequential circuits. We introduce the concept of holding signals at primary inputs and scan flipflops of a partially scanned sequential circuit for a certain length of time, instead of applying a new random vector at each clock cycle. When a random vector is held at the primary inputs of the circuit under test or at the scan flip-flops, the system clock is applied and the primary outputs of the circuit are observed. Information obtained from a testability analysis or test generator is used to determine the number of clock cycles for which each random vector is to be held constant. The method is low cost and the results of our experiment on the benchmark circuits show that it is very effective in providing fault coverage close to the maximum obtainable fault coverage using random patterns with full scan.
Lama Nachman, Kewal K. Saluja, Shambhu J. Upadhyaya, Robert Reuse
IEEE Trans. Computers2
1997 Guest Editorial
Kwang-Ting Cheng, Kewal K. Saluja, Hans-Joachim Wunderlich
J. Electron. Test.2
1997 Scheduling tests for VLSI systems under power constraints
abstract
This paper considers the problem of testing VLSI integrated circuits in minimum time without exceeding their power ratings during test. We use a resource graph formulation for the test problem. The solution requires finding a power-constrained schedule of tests. Two formulations of this problem are given as follows: (1) scheduling equal length tests with power constraints and (2) scheduling unequal length tests with power constraints. Optimum solutions are obtained for both formulations. Algorithms consist of four basic steps. First, a test compatibility graph is constructed from the resource graph. Second, the test compatibility graph is used to identify a complete set of time compatible tests with power dissipation information associated with each test. Third, from the set of compatible tests, lists of power compatible tests are extracted. Finally, a minimum cover table approach is used to find an optimum schedule of power compatible tests.
Richard M. Chou, Kewal K. Saluja, Vishwani D. Agrawal
IEEE Trans. Very Large Scale Integr. Syst.2
1996 A new method towards achieving global optimality in technology mapping
abstract
This paper presents a new method for covering a Boolean network by library cells. In this method, matches are classified according to their properties. Some matches are selected unconditionally into a cover and the remaining nodes are divided into independent portions. Then, a match compatibility graph (MCG) is constructed for each portion and an optimum cover is found for it using the MCG. Thus our method finds an efficient and closer to optimum cover for the complete network.
Xiaoqing Wen, Kewal K. Saluja
ICCAD2
1996 Testing reconfigured RAM's and scrambled address RAM's for pattern sensitive faults
abstract
State-of-the-art RAM chips are designed with spare rows and columns for reconfiguration purposes. After a RAM chip is reconfigured, physically adjacent spare cells may no longer have consecutive logical addresses. Another aspect that results in distinct logical and physical neighborhoods in RAM's is the scrambling of address lines, necessitated by the need to minimize the overall silicon area and the critical path lengths. Test algorithms used to test reconfigured RAM's and scrambled address RAM's for the detection of physical neighborhood pattern sensitive faults have to consider that the physical and logical neighborhoods are different and that the address mapping of the reconfigured RAM is no longer available. In this paper, we present a test algorithm that detects static five-cell physical neighborhood pattern sensitive faults in reconfigured RAMs and RAMs with scrambled address lines. This algorithm is based on the widely used MSCAN and Marching test algorithms, and requires only O(N[log/sub 2/N]) reads and writes to test an N-bit RAM array.
Manoj Franklin, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1996 Incorporating performance and testability constraints during binding in high-level synthesis
abstract
Module and register binding during high-level synthesis is one of the most important steps in generating an RTL design from a behavioral description. The binding phase determines the structure of the final design, and hence issues related to area, performance and testability of the RTL design have to be addressed in this step. In this paper, we present algorithms for module and register binding which generate RTL designs having high performance and/or high testability. The binding problem is decomposed into a sequence of subproblems, each of which is modeled as a minimum-cost network flow problem. The relative impact of the possible bindings is expressed in terms of the costs associated with the edges of the network. The model is simple and can be solved quickly to obtain a low cost flow solution. Putting together the solutions to the subproblems gives low cost bindings. We also propose cost functions that can be used with varying emphasis on delay and testing. The results demonstrate the effectiveness of our algorithm; the final designs produced by the algorithms require a smaller clock cycle or are easier to test as compared to designs generated without the performance or testability constraints.
Ashutosh Mujumdar, Rajiv Jain, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1995 Test application time reduction for scan based sequential circuits
abstract
This paper addresses the issue of reducing test application time in sequential circuits with partial scan using a single clock configuration without freezing the state of the non-scan flip-flops. Experimental results show that this technique significantly reduces test application time. Further, we study the effect of ordering the scan flip-flops on the test vector length and also present a non-atomic two-clock scan method which can be easily incorporated in conventional test generation environment.
Kewal K. Saluja, Rajiv Jain
Great Lakes Symposium on VLSI2
1995 An optimized testable architecture for finite state machines
abstract
This paper presents a testable architecture for FSM synthesis. The transfer, synchronizing and distinguishing sequences are obtained simultaneously by adding extra edges, if necessary, and their associated inputs and outputs to the original FSM. The algorithm that achieves this minimizes the number of extra edges that make a machine testable. The testable machine has the following properties: (1) transfer sequences of length at most [log/sub 2/n] where n is the number of the states in the machine, to carry the machine from state S/sub i/ to state S/sub j/ for all i and j, (2) a synchronizing sequence of length at most [log/sub 2/n] which sets the machine to a specific state S1, and (3) a distinguishing sequence of length at most [log/sub 2/n]. The states can be observed at the output. Several synthesis benchmark circuits were investigated for area by using the architecture.
Ting-Yu Kuo, Chun-Yeh Liu, Kewal K. Saluja
VTS3
1995 Test application time reduction for sequential circuits with scan
abstract
Scan designs alleviate the test generation problem for sequential circuits. However, scan operations substantially increase the total number of test clocks during test application stage. Classical methods used to solve this problem perform test compaction and obtain fewer test vectors. In this paper we show that such a strategy does not always reduce the test clocks or test application time. Our approach is to associate a scan strategy function with each test vector during test generation for circuits with full or partial scan. The paper presents two algorithms to generate test sequences that reduce the number of test clocks required to apply the test sequences. The algorithms are based on: (1) heuristics that determine the need for scan operations; and (2) controlling sequential test generation process by choosing an appropriate target fault. In this paper we define and investigate different scan strategies for full and partial scan designs. We propose approximate measures that can be used for selection of a target fault during sequential test generation. These concepts are integrated into the algorithms Test Application time Reduction for Full scan (TARF) and Test Application time Reduction for Partial scan (TARP). The algorithms are implemented, and their efficiencies are demonstrated by using them for a set of ISCAS sequential benchmark circuits. The experiments show that, in full scan designs, TARF generated vectors require 36% fewer test clocks compared to the vectors from COMPACTEST that produces near optimal test sets. Similarly for partial scan designs, TARP achieves over 30% cumulative test clock reduction compared to the results from FASTEST which produced generally fewer vectors than other ATPG systems.>
Soo Young Lee, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 Sequential test generation with reduced test clocks for partial scan designs
abstract
Partial scan design technique is often preferred to full scan because the use of smaller number of scan flip-flops leads to less performance degradation and less overhead. However, the number of clocks required to apply a test vector is proportional to the number of flip-flops in the scan path whenever scan is performed. This tends to increase the test application considerably. In this paper we presents an algorithm to generate a test with fewer test clocks for partial scan designs by using sequential test generation and scan strategies. The objective is to find a test that requires less test clocks while achieving high fault coverage. The algorithm, Test Application time Reduction for Partial scan design (TARP), is implemented and tested on a set of ISCAS sequential benchmark circuits. The algorithm produces a test with substantial reduction in the number of test clocks, compared to a test in which each test vector is associated with a scan operation.>
Soo Young Lee, Kewal K. Saluja
VTS2
1994 Incorporating testability considerations in high-level synthesis
Ashutosh Mujumdar, Rajiv Jain, Kewal K. Saluja
J. Electron. Test.3
1994 On-chip testing of random access memories
Kewal K. Saluja
J. Electron. Test.1
1994 Hypergraph Coloring and Reconfigured RAM Testing
abstract
RAM decoders are designed with a view to minimize the overall silicon area and critical path lengths. This can result in designs in which-physically adjacent rows (and columns) are not logically adjacent. Even if physically adjacent rows (and columns) are logically adjacent, there are other issues that preclude the possibility of identical physical and logical addresses. State-of-the-art memory chips are designed with spare rows and spare columns for reconfiguration purposes. After a memory chip is reconfigured, physically adjacent cells may no longer have consecutive logical addresses. Test algorithms used at later stages for the detection of physical neighborhood pattern sensitive faults have to consider the fact that the address mapping of the memory chip is no longer available. We present test algorithms to detect 5-cell and 9-cell physical neighborhood pattern sensitive faults and arbitrary 3-coupling faults, even if the logical and physical addresses are different and the physical-to-logical address mapping is not available. These algorithms have test lengths of O(N[log/sub 3/N]/sup 4/) and O(N[log/sub 3/N]/sup 2/), respectively, for N-bit RAMs, and are especially suited for testing reconfigured DRAMs. They also detect other conventional faults such as stuck-at faults and decoder faults. These test algorithms are based on efficiently identifying all triplets of objects among a group of n objects. We formulated this triplet identification problem as a hypergraph coloring problem, and developed an efficient 3-coloring algorithm that colors the n vertices of a complete uniform hypergraph of rank 3 such that each edge of the hypergraph is trichromatically colored in at most [log/sub 3/n]/sup 2/ coloring steps.>
Manoj Franklin, Kewal K. Saluja
IEEE Trans. Computers2
1993 Efficient Test Vectors for ISCAS Sequential Benchmark Circuits
Soo Young Lee, Kewal K. Saluja
ISCAS2
1993 CCSTG: an efficient test pattern generator for sequential circuits
abstract
A simple method which combines the efficiencies of an event driven implication method and speed of a compiled code implication is proposed for use in a sequential test pattern generator. This method, in conjunction with several other concepts and heuristics, is used to implement a sequential test pattern generator CCSTG based on the sequential test generation algorithm used in FASTEST. It is shown that the performance of a sequential test pattern generator improves substantially when methods proposed in this paper are incorporated in a test pattern generator. The authors verified this assertion by comparing the performances of test pattern generators, with and without these features, for ISCAS-89 benchmark sequential circuits.>
Kyuchull Kim, Kewal K. Saluja
VTS2
1993 An Efficient Algorithm for Sequential Circuit Test Generation
abstract
This paper presents an efficient sequential circuit automatic test generation algorithm. The algorithm is based on PODEM and uses a nine-valued logic model. Among the novel features of the algorithm are use of Initial Timeframe Algorithm and correct implementation of a solution to the Previous State Information Problem. The Initial Timeframe Algorithm, one of the most important aspects of the test generator, determines the number of timeframes required to excite the fault for which a test is to be derived and the number of timeframes required to observe the excited fault. Correct determination of the number of timeframes in which the fault should be excited (activated) and observed saves the test generator from performing unnecessary search in the input space. Test generation is unidirectional, i.e., it is done strictly in forward time, and flip-flops in the initial timeframe are never assigned a state that needs to be justified later. The algorithm saves both the good and the faulty machine states after finding a test to aid in subsequent test generation. The Previous State Information Problem, which has often been ignored by existing test generators, is presented and discussed in the paper. Experimental results are presented to demonstrate the effectiveness of the algorithm.>
Todd P. Kelsey, Kewal K. Saluja, Soo Young Lee
IEEE Trans. Computers2
1993 An efficient algorithm for bipartite PLA folding
abstract
Programmable logic arrays (PLAs) provide a flexible and efficient way of synthesizing arbitrary combinational functions as well as sequential logic circuits. They are used in both LSI and VLSI technologies. The disadvantage of using PLAs is that most PLAs are very sparse. The high sparsity of the PLA results in a significant waste of silicon area. PLA folding is a technique which reclaims unused area in the original PLA. This paper proposes a column bipartite folding algorithm based on matrix representation. Heuristics are used to reduce the search space and to speed up the search processes. The algorithm has been implemented in C programming language on a SUN-4 workstation. The program was used to study several large PLAs of varying sizes. The experimental results show that in most cases the proposed algorithm finds optimal solution in a reasonable CPU time.>
Chun-Yeh Liu, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1992 An algorithm to reduce test application time in full scan designs
abstract
An algorithm for generating a test with fewer test clocks for full scan designs by using combinational and sequential test generation algorithms adaptively is presented. Heuristics combining tests measures and scan strategies are introduced. The algorithm, 'Test Application time Reduction for Full scan designs' (TARF), is implemented and tested on a set of ISCAS sequential benchmark circuits. The results show that TARF achieves the same test coverage as combinational test generators but with fewer test clocks.>
Soo Young Lee, Kewal K. Saluja
ICCAD2
1992 On fault deletion problem in concurrent fault simulation for synchronous sequential circuits
abstract
A method for improving the performance of concurrent fault simulators for combinational and synchronous sequential circuits is proposed. The paper identifies two causes of inefficiencies and a simple and uniform method to eliminate them. A simulator, FASTS, based on the method proposed in the paper is implemented and it is shown that FASTS outperforms the existing concurrent simulation methods proposed in literature.>
Kyuchull Kim, Kewal K. Saluja
VTS2
1992 Zero cost testing of check bits in RAMs with on-chip ECC
abstract
The authors address the problem of testing the check bits in RAMs with on-chip ECC. A solution is proposed in which the check bits are tested in parallel with the testing of the information bits. The solution entails finding parity-check matrices such that all the check bits are tested while the information bits are being tested, without any increase in the length of the test sequence. The resulting parity-check matrix is such that there is no loss in error-correction capabilities and with minimal penalty in the worst-case delay of the error-correcting logic.>
Parameswaran Ramanathan, Kewal K. Saluja, Michael J. Franklin
VTS2
1991 An Algorithm to Test Rams for Physical Neighborhood Pattern Sensitive Faults
abstract
State-of-the-art memory chips are designed with spare rows and columns for reconfiguration purposes. After a memory chip is reconfigured, physically adjacent cells may no longer have consecutive logical addresses. Test algorithms used at later stages for the detection of physical neighborhood pattern sensitive faults have to consider the fact that the address mapping of the memory chip is no longer available. Furthermore, RAM decoders are designed with a view to minimize the overall silicon area and critical path lengths. This can also result in designs in which physically adjacent rows (and columns) are not logically adjacent. In this paper, we present new test algorithms to detect 5-cell and 9-cell physical neighborhood pattern sensitive faults in dynamic RAMs, even if the logical and physical addresses are different and the physical-to-logical address mapping is not available. These algorithms have test lengths of O(Nr10g3M4) for N-bit RAMs, and also detect other faults such as stuck-at and coupling faults. The algorithms depend on the development of an efficient 3-coloring algorithm that michromatically colors all the triplets among a group off n objects in at most r1og34 coloring steps.
Manoj Franklin, Kewal K. Saluja
ITC2
1991 A method of reducing aliasing in a built-in self-test environment
abstract
A method of reducing aliasing in built-in self-test of VLSI circuits is proposed. The method is based on the use of transition count testing. A new formulation of the problem is given in terms of finding a test generator as opposed to solving the problem at the data compaction end. An algorithm is proposed which can be used to find a counter-based test pattern generator. This test generator tests a circuit exhaustively or pseudo-exhaustively so that the aliasing is reduced substantially provided the data compactor used is a transition counter. Experimental results are presented to substantiate these claims.>
Keiho Akiyama, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1990 Design and Analysis of a Gracefully Degrading Interleaved Memory System
abstract
The organization of interleaved memories in such a way that faults in the memory system degrade the performance in a graceful manner is studied. Attention is restricted to an interleaved memory system that starts out with 2/sup q/ memory banks and uses a low-order interleaving scheme. The motivation and design objectives of the memory system are described. A new reconfiguration scheme and the design of the hardware needed to implement it are presented. The reconfiguration scheme is evaluated using trace-driven simulation for a number of benchmarks. The ideas presented can easily be extended to other interleaved memory schemes.>
Kifung C. Cheung, Gurindar S. Sohi, Kewal K. Saluja, Dhiraj K. Pradhan
IEEE Trans. Computers3
1989 Fast test generation for sequential circuits
abstract
An efficient sequential circuit test generation algorithm is presented. The algorithm is based on PODEM and uses a nine-valued logic model. Among the novel features of the algorithm are use of an initial time-frame algorithm and correct implementation of a solution to the previous state information problem. The initial time-frame algorithm determines the number of time-frames required to excite the fault under test and the number of time-frames required to observe the excited fault. This step saves the test generator from doing unnecessary search in the input space. Test generation is done strictly in forward time. The algorithm saves good machine circuit state after test generation to aid in future test generation. Faulty machine state is set to unknown whenever test generation for a fault is begun. This solves the previous state information problem, which has often been ignored by existing test generators.>
Todd P. Kelsey, Kewal K. Saluja
ICCAD2
1989 Design of a BIST RAM with Row/Column Pattern Sensitive Fault Detection Capability
abstract
A novel fault model is developed for random-access memories for a class of pattern-sensitive faults called row/column weight-sensitive faults. A test procedure is developed to detect faults from the defined fault model. This test sequence also tests the memory array for the 5-cell-neighborhood static pattern-sensitive faults and other faults, such as stuck-at-faults and coupling faults. A built-in self-test (BIST) version of the algorithm has been implemented by completing the logic design and layout in 2- mu m CMOS technology. The silicon area overhead for a 4M RAM is as little as 0.8%. The number of extra pins can be as low as one if clock is available on-chip. The number of extra pins can be as low as one if clock is available on-chip. The delay introduced to the normal paths, as estimated by simulation tools, is small and can be reduced even further.>
Manoj Franklin, Kewal K. Saluja, Kozo Kinoshita
ITC2
1988 Test Scheduling and Control for VLSI Built-In Self-Test
abstract
The test scheduling problem for equal length and unequal length tests for VLSI circuits using built-in self-test (BIST) has been modeled. A hierarchical model for VLSI circuit testing is introduced. The test resource sharing model from C. Kime and K. Saluja (1982) is employed to exploit the potential parallelism. Based on this model, very efficient suboptimum algorithms are proposed for defining test schedules for both the equal length test and unequal length test cases. For the unequal length test case, three different scheduling disciplines are defined, and scheduling algorithms are given for two of the three cases. Data on algorithm performance are presented. The issue of the control of the test schedule is also addressed, and a number of structures are proposed for implementation of control.>
Gary L. Craig, Charles R. Kime, Kewal K. Saluja
IEEE Trans. Computers3
1988 A Data Compression Technique for Built-In Self-Test
abstract
A data compression technique called self-testable and error-propagating space compression is proposed and analyzed. Faults in a realization of Exclusive-OR and Exclusive-NOR gates are analyzed, and the use of these gates in the design of self-testing and error propagating space compressors is discussed. It is argued that the proposed data-compression technique reduce the hardware complexity in built-in self-test (BIST) logic designs using external tester environments.>
Sudhakar M. Reddy, Kewal K. Saluja, Mark G. Karpovsky
IEEE Trans. Computers2
1988 An experimental study to determine task size for rollback recovery systems
abstract
The effects of using a recovery cache to save the variables of a program are studied. A novel optimization model for rollback is formulated to include the effects of a recovery cache in rollback systems. The parameters of the model proposed are the maximum recovery time, the cache size, and the save and load time associated with the task size. The results are also discussed of an experimental study conducted to estimate the parameters of the programs that are critical for arriving at a suitable task size or cache size to minimize the cost of recovery.>
Shambhu J. Upadhyaya, Kewal K. Saluja
IEEE Trans. Computers2
1988 A concurrent testing technique for digital circuits
abstract
A method is presented for testing digital circuits during normal operation. The resources used to perform online testing are those which are inserted to alleviate the offline testing problem. The offline testing resources are modified so that during system operation they can also observe the normal inputs and outputs of a combinational circuit under test. The normal inputs to the circuit under test are with test vectors in its test set. When a normal input matches a test vector, the circuit output for such an input is typically compressed into a developing signature. When all of the test vectors in the test set have appeared as normal inputs, the signature is read and verified. With this method, the length of time required for all of the test vectors to appear, possibly in some order, among the normal inputs to the circuit under test is of considerable importance.>
Kewal K. Saluja, Rajiv Sharma, Charles R. Kime
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1988 A new approach to the design of built-in self-testing PLAs for high fault coverage
abstract
Four critical requirements are identified for the built-in self-testing of programmable logic arrays (BIST PLAs): the test set to test the PLA as well as the output response must be independent of the function of the PLA; the test pattern generator (TPG) and the response evaluator circuits must be simple to keep the extra logic overhead to a minimum; the fault coverage of the PLA must be within acceptable limits; and the speed of the test application must be high. A design that meets all of these goals is proposed. The approach is based on counting crosspoints, as opposed to the conventional parity technique. The TPG and RE circuits are simple and consist of shift registers and counters. The design requires a reorganization of the columns of the PLA on the basis of the number of crosspoints. This design provides extremely high fault coverage: the coverage for multiple faults is higher than that of any BIST design known to the authors, and the single-fault coverage is 100%. The design is simple and can easily be incorporated into existing computer-aided design systems.>
Shambhu J. Upadhyaya, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1987 BIST-PLA: A Built-in Self-Test Design of Large Programmable Logic Arrays
abstract
A new method for designing a Built-In Self-Test Programmable Logic Array (BIST-PLA) is presented. In the proposed design, the Test Pattern Generator and the Response Evaluator circuits are very simple. The design requires a re-arrangement of the AND (OR) planes on the basis of number of crosspoints in the product (output) lines in the PLA.
Chun-Yeh Liu, Kewal K. Saluja, Shambhu J. Upadhyaya
DAC2
1987 Organization and Analysis of a Gracefully-Degrading Interleaved Memory System
abstract
A hardware mechanism has been proposed to reconfigure an interleaved memory system. The reconfiguration scheme is such that, at any instant all fault-free memory banks in the memory system are utilized in interleaved manner. A performance metric is defined which takes into account the bandwidth and the page-fault rate in an interleaved memory system. The reconfiguration scheme proposed in this paper is analyzed for a number of distinct programs using the performance metric defined in the paper. It is shown that the system performance degrades slowly, as the number of faulty banks increase, in a memory system using the proposed reconfiguration scheme.
Kifung C. Cheung, Gurindar S. Sohi, Kewal K. Saluja, Dhiraj K. Pradhan
ISCA3
1986 A Novel Approach for Testing Memories Using a Built-In Self Testing Technique
Kim T. Le, Kewal K. Saluja
ITC2
1986 Built-In Testing of Memory Using an On-Chip Compact Testing Scheme
abstract
In this paper we study the problem of testing RAM. A new fault model, which encompasses the existing fault models, is proposed. We then propose a scheme of testing faults from the new fault model using built-in testing techniques. We introduce concept of p-hard and determine the complexity of the extra hardware required for built-in self-testing on our hardness scale. A novel approach using microcoded ROM for implementation of built-in testing is also proposed and its complexity is determined.
Kozo Kinoshita, Kewal K. Saluja
IEEE Trans. Computers2
1986 An Alternative to Scan Design Methods for Sequential Machines
abstract
The problem of testing sequential machines using a checking experiment is investigated. An algorithm is given to augment sequential machines by adding extra input(s) to make them testable. We also present a circuit modification method, similar to scan methods, such that the augmented machine can be tested by the checking experiment. A justification of our method for a VLSI environment is given by determining the overheads.
Kewal K. Saluja, Ramaswami Dandapani
IEEE Trans. Computers1
1986 Testable Design of Single-Output Sequential Machines Using Checking Experiments
abstract
The problem of testing sequential machines using checking experiments is investigated. A method of modifying sequential machines by adding a controllable input is presented. A procedure is given to construct checking experiments for the modified machine and it is shown that only one output observation is sufficient to determine whether the machine is fault free.
Kewal K. Saluja, Ramaswami Dandapani
IEEE Trans. Computers1
1986 A Wachtdog Processor Based General Rollback Technique with Multiple Retries
abstract
A common assumption in the existing rollback techniques is that transients, the cause of most failures, subside very quickly, implying that a single story retry of the program from the previous rollback point is sufficient. The authors discuss a general rollback strategy withn(n≥2) retries which takes into consideration multiple transient failures as well as transients of long duration. Ways of deriving practical values ofnfor a given program are also discussed. Furthermore, the authors propose the use of a watchdog processor as an error detection tool to initiate recovery action through rollback, since the watchdog processor offers low error latency. They also discuss the merging of the watchdog processor with rollback recovery technique for enhancing the overall system reliability.
Shambhu J. Upadhyaya, Kewal K. Saluja
IEEE Trans. Software Eng.2
1985 A Testable Design of Programmable Logic Arrays with Universal Control and Minimal Overhead
Hideo Fujiwara, Kewal K. Saluja, Kozo Kinoshita
ITC2
1985 Test Pattern Generation for API Faults in RAM
abstract
In this correspondence we consider the problem of test pattern generation for random-access memory to detect pattern-sensitive faults. A test algorithm is presented which contains a near-optimal WRITE sequence and is an improvement over existing algorithms. The algorithm is well suited for built-in testing applications.
Kewal K. Saluja, Kozo Kinoshita
IEEE Trans. Computers1
1984 Built-in Testing of Memory Using On-chip Compact Testing Scheme
Kozo Kinoshita, Kewal K. Saluja
ITC2
1984 Testable design of large random access memories
Kewal K. Saluja, Kim T. Le
Integr.1
1983 Testing Computer Hardware through Data Compression in Space and Time
Kewal K. Saluja, Mark G. Karpovsky
ITC1
1983 A Simplified Algorithm for Testing Microprocessors
Kewal K. Saluja, Stephen Y. H. Su
ITC1
1983 An Easily Testable Design of Programmable Logic Arrays for Multiple Faults
abstract
In this paper, the problem of fault detection for multiple faults in programmable logic arrays (PLA's) is discussed. An easily testable design of PLA's has been proposed which has the following properties: 1) for a PLA with n inputs, m product terms, there exists a test set such that the test patterns do not depend on the function realized by the PLA; 2) the number of tests to detect multiple stuck type and cross point faults is m(2n + 1) + 4n + 4; 3) the number of additional pins for the testable design is 3; 4) the design philosophy is compatible with the built-in-testing approaches.
Kewal K. Saluja, Kozo Kinoshita, Hideo Fujiwara
IEEE Trans. Computers1
1982 An enhancement of lssd to reduce test pattern generation effort and increase fault coverage
abstract
In this paper we propose designs of latches which can be used in Level Sensitive Scan Design (LSSD). These new designs can use the existing software support for design rule checks but result into a reduction of effort in test pattern generation and provide a better fault coverage. The system performance is not degraded with the use of latches proposed in this paper.
Kewal K. Saluja
DAC1
1980 Fault diagnosis in loop-connected systems
Kewal K. Saluja, Brian D. O. Anderson
Inf. Sci.1
1980 Synchronous Sequential Machines: A Modular and Testable Design
abstract
It will be shown that a single-input n-definite machine realized by a universal modular tree, in which each module consists of AND-EXCLUSIVE-OR-DELAY (AND-EOR-DELAY) as a basic element, can be tested for single stuck-type-faults by tests of length 2n + 3 only. This is a marked improvement over the previous results for trees consisting of AND-OR-DELAYS, which are known to have test lengths of exponential growth.
Kewal K. Saluja
IEEE Trans. Computers1
1979 Minimization of Reed-Muller Canonic Expansion
abstract
It is shown that Swamy's [1] approach to generate generalized Reed–Muller canonic expansion is in error. A different algorithm is presented which uses a single Boolean matrix and successive modifications in function vector to generate all the solutions sequentially.
Kewal K. Saluja, E. H. Ong
IEEE Trans. Computers1
1975 Fault Detecting Test Sets for Reed-Muller Canonic Networks
abstract
Fault detecting test sets to detect multiple stuck-at-faults (s-a-faults) in certain networks, realizing Reed-Muller(RM) canonic expressions called RM canonic (RMC) networks, are given. It is shown that to detect t faults, t ≥ 1, in a network realizing an arbitrary n-variable logic function only tests need be applied ([x] is the integer part of x) and that these tests are independent of the function being realized.
Kewal K. Saluja, Sudhakar M. Reddy
IEEE Trans. Computers1
1974 On Minimally Testable Logic Networks
abstract
A new technique to modify any logic network to facilitate diagnosis is given. By providing extra controllable inputs (at most six) and observable outputs it is shown that any number of stuck-at-faults in a logic network can be detected by applying only three tests. This number is believed to be minimal for networks using current technologies. Example of logic module that can be used to realize any logic function such that only two tests detect stuck-at-faults is also given.
Kewal K. Saluja, Sudhakar M. Reddy
IEEE Trans. Computers1
1974 Easily Testable Two-Dimensional Cellular Logic Arrays
abstract
An algorithm to synthesize two-dimensional AND-EOR arrays is given. The design criterion chosen is to minimize the number of columns in the two-dimensional cellular arrays. It is also shown that And-Eor arrays, synthesized using the algorithm presented, can be modified such that 2n + 5 test vectors will detect any single fault in the array realizing an n-variable function with only one observable output. Furthermore, this test set is shown to be independent of the function being realized by the cellular array under test.
Kewal K. Saluja, Sudhakar M. Reddy
IEEE Trans. Computers1