Nicola Nicolici

dblp:20/5950 · DBLP profile ↗
← Back
109ranked-venue papers
11as first author
5since 2021 · last 2025
0000-0001-6345-5908ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 109 · 11 first-author · 5 since 2021Software engineering, systems software and programming languages · 15 · 3 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Karatsuba Matrix Multiplication and Its Efficient Custom Hardware Implementations
abstract
While the Karatsuba algorithm reduces the complexity of large integer multiplication, the extra additions required minimize its benefits for smaller integers of more commonly-used bitwidths. In this work, we propose the extension of the scalar Karatsuba multiplication algorithm to matrix multiplication, showing how this maintains the reduction in multiplication complexity of the original Karatsuba algorithm while reducing the complexity of the extra additions. Furthermore, we propose new matrix multiplication hardware architectures for efficiently exploiting this extension of the Karatsuba algorithm in custom hardware. We show that the proposed algorithm and hardware architectures can provide real area or execution time improvements for integer matrix multiplication compared to scalar Karatsuba or conventional matrix multiplication algorithms, while also supporting implementation through proven systolic array and conventional multiplier architectures at the core. We provide a complexity analysis of the algorithm and architectures and evaluate the proposed designs both in isolation and in an end-to-end accelerator system compared to baseline designs and prior state-of-the-art works implemented on the same type of compute platform, demonstrating their ability to increase the performance-per-area of matrix multiplication hardware.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Computers2
2025 An Embedded Architecture for DDR5 DFE Calibration Based on Channel Stimulus Inversion
abstract
The increase in performance promised by the recent generation of double data rate (DDR) memory, DDR5, is conditioned by addressing its signal integrity challenges. The DDR5 standard specifies a 4-tap decision feedback equalizer (DFE) at the memory receiver to deal with these challenges. Although adaptive equalization is a mature field, known methods for DFE calibration are limited by the DDR5 interface complexity and the equalization requirements mandated by its specification. In this article, we propose a novel approach based on linear inversion of channel stimulus that leverages specific architectural details of DDR5 and can tune memory devices deterministically at runtime. In addition to using few hardware resources relative to a modern memory controller, by operating at very low latency, this new approach facilitates periodic equalization when the DFE is offline, thus avoiding DFE error propagation during training inherent to adaptive techniques.
Mitchell Cooke, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2025 Strassen Multisystolic Array Hardware Architectures
abstract
While Strassen’s matrix multiplication algorithm reduces the complexity of naive matrix multiplication, general-purpose hardware is not suitable for achieving the algorithm’s promised theoretical speedups. This leaves the question of whether it could be better exploited in custom hardware architectures designed specifically for executing the algorithm. However, there is limited prior work on this and it is not immediately clear how to derive such architectures or whether they can ultimately lead to real improvements. We bridge this gap, presenting and evaluating new systolic array architectures that efficiently translate the theoretical complexity reductions of Strassen’s algorithm directly into hardware resource savings. Furthermore, the architectures are multisystolic array designs that can multiply smaller matrices with higher utilization than single-systolic array designs. The proposed designs implemented on FPGA reduce DSP requirements by a factor of$1.14^{r}$for r implemented Strassen recursion levels, and otherwise require overall similar soft logic resources when instantiated to support matrix sizes down to$32\times 32$and$24\times 24$at one to two levels of Strassen recursion, respectively. We evaluate the proposed designs in both isolation and an end-to-end machine learning accelerator compared with baseline designs and prior works, achieving state-of-the-art performance.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2024 Fast Inner-Product Algorithms and Architectures for Deep Neural Network Accelerators
abstract
We introduce a new algorithm called the Free-pipeline Fast Inner Product (FFIP) and its hardware architecture that improve an under-explored fast inner-product algorithm (FIP) proposed by Winograd in 1968. Unlike the unrelated Winograd minimal filtering algorithms for convolutional layers, FIP is applicable to all machine learning (ML) model layers that can mainly decompose to matrix multiplication, including fully-connected, convolutional, recurrent, and attention/transformer layers. We implement FIP for the first time in an ML accelerator then present our FFIP algorithm and generalized architecture which inherently improve FIP's clock frequency and, as a consequence, throughput for a similar hardware cost. Finally, we contribute ML-specific optimizations for the FIP and FFIP algorithms and architectures. We show that FFIP can be seamlessly incorporated into traditional fixed-point systolic array ML accelerators to achieve the same throughput with half the number of multiply-accumulate (MAC) units, or it can double the maximum systolic array size that can fit onto devices with a fixed hardware budget. Our FFIP implementation for non-sparse ML models with 8 to 16-bit fixed-point inputs achieves higher throughput and compute efficiency than the best-in-class prior solutions on the same type of compute platform.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Computers2
2024 Thresholding Decision-Directed Descent (T3D): A Tuning Solution for DDR5 DRAM DFEs
abstract
Emerging memory technologies, such as DDR5, offer increased data rates and storage capacities, at the expense of signal integrity challenges. To address these challenges, the DDR5 standard incorporates a four-tap decision feedback equalizer (DFE). As elaborated in this article, known methods for DFE tuning are limited due to interface complexity and distinct equalization requirements for DDR5. We propose a decision-directed DFE tuning method called thresholding decision-directed descent (T3D). By leveraging DDR5 architectural features, our novel method tracks the eye envelope as it opens, which facilitates rapid convergence compared to the state of the art. To validate the performance of T3D, silicon measurements are presented alongside a virtual testbench methodology. By demonstrating the high correlation between silicon and simulation results, the virtual testbench can be beneficial for the design, validation, and prototyping of future DFE tuning methods.
Mitchell Cooke, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Incremental Fault Analysis: Relaxing the Fault Model of Differential Fault Attacks
abstract
This article presents a new fault analysis technique against cryptographic devices called the incremental fault analysis (IFA), which can be adapted into fault attacks using more traditional differential fault analysis (DFA) techniques in order to increase their feasibility under more practical fault injection conditions. Many previous attack methods require precise fault injection techniques such as clock glitching. By contrast, IFA is compatible with a more practical overclocking fault injection technique in which a cryptosystem is stressed at a constant level throughout the entire encryption, and this constant stress level is then increased between consecutive encryptions. It is observed that as new faults occur incrementally between increased stress levels, they often become superimposed upon faults first appearing at lower stress levels. IFA exploits these incremental fault differentials to deduce the cipher key more rapidly. Attacks were tested using practical fault injection methods on the advanced encryption standard (AES) both with and without IFA applied. Using IFA, allowed cipher keys to be retrieved with a success rate of 100% from 10 times less faulty ciphertexts and 6.4 times less computational time, requiring 16, 86, and 43 ciphertexts on average for AES-128, AES-192, and AES-256, respectively.
Trevor E. Pogue, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Bit-Flip Detection-Driven Selection of Trace Signals
abstract
Since integrating memory blocks on-chip became affordable, embedded logic analysis has been used extensively for post-silicon validation and debugging. Deciding at design time which signals to be traceable at the post-silicon phase, has been posed as an algorithmic problem a decade ago. The primary focus of the subsequent approaches on this topic was to restore as much data as possible within a software simulator in order to facilitate the analysis of functional bugs, assuming there are no electrically induced design errors, e.g., bit-flips. In this paper, we show that analyzing post-silicon traces can also aid with the identification bit-flips. We present a new trace signals selection algorithm that is driven by the detection of bit-flips.
Amin Vali, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 An automated SAT-based method for the design of on-chip bit-flip detectors
abstract
Hardware invariants are known to facilitate bit-flip detection during post-silicon validation. In this paper, we present a fully automated SAT-based methodology for fast generation of hardware invariants by using the built-in pruning mechanisms within SAT solvers, namely learned clauses. These candidates are evaluated for their potential to detect bit-flips using a new incremental SAT-based approach. In addition to speeding-up the simulation-based approaches for invariant generation and evaluation, when compared to the known art, our results show improvements in both the number of flip-flops that can be covered for bit-flip detection, as well as for the on-chip area for the bit-flip detection unit.
Pouya Taatizadeh, Nicola Nicolici
ICCAD2
2017 A generic embedded sequence generator for constrained-random validation with weighted distributions
abstract
Post-silicon validation is concerned for discovering design errors that escape to the silicon prototypes. Recent research efforts have shown how to reuse the constraints from pre-silicon verification to support post-silicon constrained-random validation. The objective is to subject the prototype to a large volume of random, yet functionally-compliant stimuli. In this paper, we present a new method that facilitates on-chip stimuli generation compliant to constraints with weighted distributions.
Xiaobing Shi, Nicola Nicolici
IOLTS2
2017 Emulation Infrastructure for the Evaluation of Hardware Assertions for Post-Silicon Validation
abstract
The objective of post-silicon validation is to identify design errors that remain undetected after pre-silicon verification and, therefore, manifest themselves in the silicon prototypes. These errors are often associated with the subtle interactions between the electrical states of the systems and commonly manifest in the logic domain as bit-flips in flip-flops. They occur under unique operating conditions, which are often not-easily repeatable. In order to shorten the long detection latencies from an error's occurrence until its observation (i.e., system crash), embedded assertion checkers can be employed. Nonetheless, relying on simulation-based experiments for selecting and assessing the practical effectiveness of a subset of assertion checkers (to be implemented in the physical device) suffers from the slow simulation speed. To address this concern, in this paper, we present a systematic methodology to automatically design emulation-based experiments that can aid the selection and assessment of the embedded assertion checkers. Our results indicate improvements of up to 10% on average for the coverage of flip-flops that are affected by bit-flips when compared with results obtained by simulation-based experiments.
Pouya Taatizadeh, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Bit-flip detection-driven selection of trace signals
abstract
Analyzing the traces that are collected on-chip is an effective method employed during post-silicon validation. Deciding at design time which signals to trace at the post-silicon phase, has been posed as an algorithmic problem several years ago. The primary focus of the subsequent approaches on this topic was to restore data in order to facilitate functional debugging. In this paper we show that analyzing post-silicon traces can also aid with the identification of electrically-induced design errors, e.g., bit-flips. We present a new trace signals selection algorithm that is driven by the detection of bit-flips.
Amin Vali, Nicola Nicolici
ETS2
2016 Generating Cyclic-Random Sequences in a Constrained Space for In-System Validation
abstract
The constrained-random methodology is widely used during the pre-silicon verification of very-large scale integrated circuits. Recently, research efforts have been made to support the application of constrained-random patterns during the post-silicon validation stage. In this paper, we present a new method, including both software algorithms and on-chip hardware structures, for in-system constrained-random generation of stimuli sequences that are uniformly distributed. More specifically, we facilitate in-system application of constrained-random sequences that are cyclic-random, i.e., all the valid values from the user-constrained space are generated only once before the entire sample space is exhausted. While software simulation environments commonly support this feature, e.g., randc in SystemVerilog, to the best of our knowledge this is the first time it is shown how such feature can be ported to hardware environments.
Xiaobing Shi, Nicola Nicolici
IEEE Trans. Computers2
2016 On-Chip Cube-Based Constrained-Random Stimuli Generation for Post-Silicon Validation
abstract
Post-silicon validation is critical for exposing subtle design errors that have escaped to the silicon prototypes. Its effectiveness is conditioned by in-system application of a large volume of functionally-compliant stimuli. In this paper, we present a methodology to design constrained-random stimuli generators, which are placed on-chip and are configurable at design-time to generate in-system functionally-compliant stimuli subject to user-programmable constraints provided at validation time. Central to our method is a cube-based representation of constraints. These cubes are used as masks that force pseudo-random sequences to map onto functionally-compliant stimuli. To reduce the on-chip storage requirements, masks are compressed at design-time and expanded on-the-fly at validation time using decompression circuitry. Experimental results evaluate the impact of our method on the requirements for on-chip logic and memory resources.
Xiaobing Shi, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Automated Selection of Assertions for Bit-Flip Detection During Post-Silicon Validation
abstract
Post-silicon validation deals with detection and diagnosis of errors that, due to existing limitations in pre-silicon verification, escape to the silicon prototypes and need to be fixed before committing to high-volume manufacturing. Electrical errors, such as those caused by cross-talk or power droops, are particularly difficult to catch during the pre-silicon phase because of the insufficient accuracy of device models, which is often traded-off against simulation time. This challenge is further aggravated by the rising number of voltage domains, especially if subtle errors are excited in unique electrical states. In fact these electrically-induced subtle errors most commonly manifest in the logic domain as bit-flips and, to the best of our knowledge, there are no systematic methods for designing embedded hardware monitors for generic logic blocks that can detect bit-flips with low detection latency. Moreover, unlike pre-silicon verification and manufacturing test that benefit from well-defined and universally accepted coverage metrics, there is no generic metric from which confidence can be implied at the end of post-silicon validation. Toward these goals, we present a method that relies on design invariants (assertions) that are ranked based on their potential to detect bit-flips. We also introduce two metrics bit-flip coverage estimate and flip-flop coverage estimate that can be used to assess the quality of the selected assertions, and, in general, the effectiveness of the post-silicon validation process.
Pouya Taatizadeh, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 A methodology for automated design of embedded bit-flips detectors in post-silicon validation
Pouya Taatizadeh, Nicola Nicolici
DATE2
2015 On-Chip Generation of Uniformly Distributed Constrained-Random Stimuli for Post-Silicon Validation
abstract
Post-silicon validation is becoming widely adopted because it runs significantly faster than pre-silicon verification and hence it helps uncover subtle design errors that escape to silicon prototypes. However, it is hindered by limited controllability and observability, which makes it challenging to reuse pre-silicon content. In order to enable the reuse of stimuli constraints from pre-silicon verification environments, we present a method that facilitates the on-chip generation of uniformly distributed constrained-random stimuli. More specifically, our method, which relies on new pre-processing steps and on-chip hardware features, can generate in real-time pseudo cyclic-random stimuli with no repetition until the space of the compliant stimuli is exhausted.
Xiaobing Shi, Nicola Nicolici
ICCAD2
2015 SAT Solving using FPGA-based Heterogeneous Computing
abstract
We present a heterogeneous computing solution to the Boolean satisfiability (SAT) problem. Our field-programmable gate array (FPGA)-based implementation for accelerating the common case computation within a SAT solver utilizes most of the FPGA resources and it seamlessly integrates with our software host. Algorithms and data structures were redesigned to maximize the strengths of customized computing and generalizable optimizations are proposed to maximize throughput, minimize communication latencies, and compact hardware memory. We are significantly faster than state-of-the-art SAT solvers in software and hardware.
Jason Thong, Nicola Nicolici
ICCAD2
2015 Emulation-based selection and assessment of assertion checkers for post-silicon validation
abstract
The objective of post-silicon validation is to detect design errors on early silicon prototypes. Electrically-induced errors commonly manifest as bit-flips in the logic domain and they occur under unique operating conditions, which are often not-easily-repeatable. In order to shorten the long detection latencies from an error's manifestation until its observation (i.e. system crash), embedded assertion checkers can be employed. Nonetheless, relying on simulation-based experiments for selecting and assessing the usefulness of a subset of assertion checkers (to be committed to silicon) suffers from limitations associated with the slow simulation speed. To address this concern, in this paper we present a systematic method to automatically design emulation-based experiments that can aid the selection and assessment of the embedded assertion checkers. Our results indicate improvements of up to 10% on average for the coverage of flip-flops that are affected by bit-flips when compared to results obtained from simulation-based experiments.
Pouya Taatizadeh, Nicola Nicolici
ICCD2
2014 On Supporting Sequential Constraints for On-Chip Generation of Post-silicon Validation Stimuli
abstract
Post-silicon validation plays a critical role in exposing design errors in early silicon prototypes. Its effectiveness is conditioned by in-system application of functionally-compliant stimuli for extensive periods of time. This is achieved by expanding on-the-fly randomized functional sequences, which are subjected to user-programmable constraints. In this paper we present a method to extend the existing work for on-chip generation of functionally-compliant randomized sequences with support for sequential constraints.
Xiaobing Shi, Nicola Nicolici
ATS2
2014 On-chip constrained random stimuli generation for post-silicon validation using compact masks
abstract
During post-silicon validation a large number of constrained random stimuli are applied to expose the subtle design errors that have escaped to the silicon prototypes. In this paper we present a new method to design constrained random stimuli generators, which are programmable and can be placed on-chip to generate extensive random, yet functionally-compliant, sequences for real-time/in-system validation. The basic idea is to translate the constraints for constrained-random variables into binary cubes, whose specified values are used as masks to correct random sequences. To reduce the volume of data needed to be placed on-chip, the cubes are efficiently encoded and expanded in real-time. Experimental results confirm the effectiveness of this new method when compared against the prior work on the topic.
Xiaobing Shi, Nicola Nicolici
ITC2
2014 A Novel Algorithmic Approach to Aid Post-Silicon Delay Measurement and Clock Tuning
abstract
The number of speedpaths in modern high-performance designs is in the range of millions and, due to unmodelled electrical effects, they are difficult to be measured accurately before the first silicon samples are available. As a consequence, clock tuning elements are employed to aid the post-silicon clock tuning. However, as the number of these elements continues to grow, it becomes increasingly difficult to determine their configurations in a compute effective manner. In this paper we describe a novel exact algorithm for post-silicon clock tuning, which employs smart pruning techniques that exploit the characteristics of the clock tuning buffers.
Zahra Lak, Nicola Nicolici
IEEE Trans. Computers2
2014 A Multiple-FPGA parallel computing architecture for real-time simulation of soft-object deformation
abstract
Hardware-based parallel computing is proposed for acceleration of finite-element (FE) analysis of linear elastic deformation models. An implementation of the Preconditioned Conjugate Gradient algorithm on N Field Programmable Gate Array (FPGA) devices solves the large linear system of equations arising from the FE discretization. The system employs a large number of customized fixed-point computing units with a high-throughput memory architecture. An implementation of this scalable architecture on four Altera EP3SE110 FPGA devices yields a peak performance of 604 Giga Operations per second. This enables haptic simulation of a 3-dimensional deformable object of 21000 elements at an update rate of 400Hz.
Behzad Mahdavikhah, Ramin Mafi, Shahin Sirouspour, Nicola Nicolici
ACM Trans. Embed. Comput. Syst.4
2013 Hardware-efficient on-chip generation of time-extensive constrained-random sequences for in-system validation
abstract
Linear Feedback Shift Registers (LFSRs) have been extensively used for compressed manufacturing test. They have been recently employed as a foundation for porting constrained-random stimuli from a pre-silicon verification environment to in-system validation. This work advances this concept by improving both the hardware efficiency and the duration of in-system validation experiments.
Adam B. Kinsman, Ho Fai Ko, Nicola Nicolici
DAC3
2013 FPGA acceleration of enhanced boolean constraint propagation for SAT solvers
abstract
We propose a hardware architecture to accelerate boolean constraint propagation (BCP). Although satisfiability (SAT) solvers in software use varying search and learning strategies, BCP is a fundamental component and by far consumes the most CPU time. Our field-programmable gate array (FPGA) design uses on-chip SRAM to facilitate the acceleration of BCP. We discuss many insights to our innovative hardware memory layout, which is very compact and enables extremely fast BCP. It also supports multithreading to minimize the idle time in hardware and to fully utilize the multicore processor host. Additionally, many industrial SAT instances encode logic gates as constraints. We compact these to simultaneously reduce the hardware memory usage as well as speed up the computation (enhanced BCP). We implemented our enhanced BCP core and integrated it with a simple software SAT solver which communicates over PCI Express. Hardware performance counters show that a single processing engine is up to 4x faster than a state-of-the-art software SAT solver.
Jason Thong, Nicola Nicolici
ICCAD2
2013 NoC-Based FPGA Acceleration for Monte Carlo Simulations with Applications to SPECT Imaging
abstract
As the number of transistors that are integrated onto a silicon die continues to increase, the compute power is becoming a commodity. This has enabled a whole host of new applications that rely on high-throughput computations. Recently, the need for faster and cost-effective applications in form-factor constrained environments has driven an interest in on-chip acceleration of algorithms based on Monte Carlo simulations. Though Field Programmable Gate Arrays (FPGAs), with hundreds of on-chip arithmetic units, show significant promise for accelerating these embarrassingly parallel simulations, a challenge exists in sharing access to simulation data among many concurrent experiments. This paper presents a compute architecture for accelerating Monte Carlo simulations based on the Network-on-Chip (NOC) paradigm for on-chip communication. We demonstrate through the complete implementation of a Monte Carlo-based image reconstruction algorithm for Single-Photon Emission Computed Tomography (SPECT) imaging that this complex problem can be accelerated by two orders of magnitude on even a modestly sized FPGA over a 2 GHz Intel Core 2 Duo Processor. The architecture and the methodology that we present in this paper is modular and hence it is scalable to problem instances of different sizes, with application to other domains that rely on Monte Carlo simulations.
Phillip Kinsman, Nicola Nicolici
IEEE Trans. Computers2
2012 Automated data analysis techniques for a modern silicon debug environment
abstract
With the growing size of modern designs and more strict time-to-market constraints, design errors unavoidably escape pre-silicon verification and reside in silicon prototypes. As a result, silicon debug has become a necessary step in the digital integrated circuit design flow. Although embedded hardware blocks, such as scan chains and trace buffers, provide a means to acquire data of internal signals in real time for debugging, there is a relative shortage in methodologies to efficiently analyze this vast data to identify root-causes. This paper presents an automated software solution that attempts to fill-in the gap. The presented techniques automate the configuration process for trace-buffer based hardware in order to acquire helpful information for debugging the failure, and detect suspects of the failure in both the spatial and temporal domain.
Yu-Shen Yang, Andreas G. Veneris, Nicola Nicolici
ASP-DAC3
2012 In-system constrained-random stimuli generation for post-silicon validation
abstract
When generating the verification stimuli in a pre-silicon environment, the primary objectives are to reduce the simulation time and the pattern count for achieving the target coverage goals. In a hardware environment, because an increase in the number of stimuli is inherently compensated by the advantage of real-time execution, the objective augments to considering hardware complexity when designing in-system stimuli generators that must operate according to user-programmable constraints. In this paper we introduce a structured methodology for porting in-system the constrained-random stimuli generation aspect from a pre-silicon verification environment.
Adam B. Kinsman, Ho Fai Ko, Nicola Nicolici
ITC3
2012 Mapping Trigger Conditions onto Trigger Units during Post-silicon Validation and Debugging
abstract
On-chip trigger units are employed for detecting events of interest during post-silicon validation and debugging. Their implementation constrains the trigger conditions that can be programmed at runtime. It is often the case that some trigger events of interest, which were not accounted for during design time, cannot be detected due to the constraints imposed by the hardware implementation of the trigger units. To address this issue, we present architectural features that can be included into the trigger units and discuss the algorithmic approach for automatically mapping trigger conditions onto the trigger units.
Ho Fai Ko, Nicola Nicolici
IEEE Trans. Computers2
2012 On Using On-Chip Clock Tuning Elements to Address Delay Degradation Due to Circuit Aging
abstract
Lifetime performance of digital integrated circuits degrades as a consequence of circuit aging. In the past few years, there has been extensive research to reduce the impact of aging by different design techniques, or to predict the degradation and adapt the circuit accordingly. In this paper, we explore a novel perspective to this problem by exploiting the presence of clock tuning elements in high-performance designs. By combining on-chip sensors to predict setup or hold-time violations with the clock tuning elements, we provide an effective self-tuning mechanism for each circuit sample. The proposed method can operate in-system to prolong the circuit's maximum performance in its unique operating environment.
Zahra Lak, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2012 Automating Data Analysis and Acquisition Setup in a Silicon Debug Environment
abstract
With the growing size of modern designs and more strict time-to-market constraints, design errors can unavoidably escape pre-silicon verification and reside in silicon prototypes. Due to those errors and faults in the fabrication process, silicon debug has become a necessary step in the digital integrated circuit design flow. Embedded hardware blocks, such as scan chains and trace buffers, provide a means to acquire data of internal signals in real time for debugging. However, the amount of the data is limited compared to pre-silicon debugging. This paper presents an automated software solution to analyze this sparse data to detect suspects of the failure in both the spatial and temporal domain. It also introduces a technique to automate the configuration process for trace-buffer-based hardware in order to acquire helpful information for debugging the failure. The technique takes the hardware constraints into account and identifies alternatives for signals not part of the traceable set so that their values can be restored by implications. The experiments demonstrate the effectiveness of the proposed software solution in terms of run-time and resolution.
Yu-Shen Yang, Andreas G. Veneris, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Dynamic binary translation to a reconfigurable target for on-the-fly acceleration
abstract
Dynamic binary translation has been extensively used in porting applications from one platform to another at runtime. At the same time, reconfigurable computing has been commonly employed for hardware acceleration for compute-intensive applications. In this paper we attempt to understand the technical challenges of applying dynamic binary translation to a reconfigurable computing environment, when the translation focus is on-the-fly acceleration.
Phillip Kinsman, Nicola Nicolici
DAC2
2011 In-system and on-the-fly clock tuning mechanism to combat lifetime performance degradation
abstract
Addressing lifetime performance degradation caused by circuit ageing has been a topic of active research for the past few years. In this paper we present a different perspective to this problem, by leveraging the presence of clock tuning elements that are commonly available in high-performance designs. By combining clock tuning elements with on-chip sensors for predicting setup/hold-time violations, we introduce a new clock tuning mechanism that operates on-the-fly and it maintains the maximum achievable performance in-system for each circuit sample affected by ageing.
Zahra Lak, Nicola Nicolici
ICCAD2
2011 On Using Lossy Compression for Repeatable Experiments during Silicon Debug
abstract
The amount of data that is observed during at-speed silicon debug is limited by the capacity of the on-chip trace buffers. To increase the debug observation window, we propose a low-cost debug architecture for at-speed silicon debug based on lossy compression. The proposed architecture enables a new debug methodology that accelerates the identification of the erroneous samples that occur intermittently over a long observation window by avoiding debug experiments that capture only error-free data. The proposed solution is applicable to both automatic test equipment-based debug and in-field debug on application boards, as long as the debug experiments are repeatable and the reference data at the probe signals are deterministically computed using a fast behavioral model of the circuit under debug.
Ehab Anis Daoud, Nicola Nicolici
IEEE Trans. Computers2
2011 Trade-Offs in Test Data Compression and Deterministic X-Masking of Responses
abstract
While a large body of work on output unknown (X) tolerance exists, to the best of the authors' knowledge, no study is provided in the literature which explores the trade-off between X density and compression of circuit stimuli without reducing fault coverage. To this end, we introduce an architectural and algorithmic framework through which we explore this trade-off, the findings of which we discuss in the experimental results section.
Adam B. Kinsman, Nicola Nicolici
IEEE Trans. Computers2
2011 Computational Vector-Magnitude-Based Range Determination for Scientific Abstract Data Types
abstract
As interest mounts in using hardware accelerators to speed up numerical scientific calculations, automation tool support is required to aid designers in mapping applications to custom hardware. One key step in designing this custom hardware is bit-width allocation where the known-art faces challenges when dealing with applications from the scientific computing domain, thus motivating the use of computational methods based on Satisfiability-Modulo Theory. Many real-life applications are, however, specified in terms of vectors and matrices which are of sufficient size to make expansion into scalar equations infeasible. The proposed vector-magnitude method and its extension via block vectors enable computational methods to be leveraged in tackling calculations of practically relevant complexity. Application to case studies confirms that through a more compact computational instance, search efficiency is improved leading to tighter bounds and thus smaller bit-widths.
Adam B. Kinsman, Nicola Nicolici
IEEE Trans. Computers2
2011 Automated Range and Precision Bit-Width Allocation for Iterative Computations
abstract
As scientific computing becomes more widespread in environments where form-factor considerations necessitate hardware acceleration, the problem of selecting numerical data representations (bit-width allocation), key to accelerator design, is faced with shortcomings in the existing techniques. To address this problem for scientific computing dataflows, we propose a methodology for determining custom hybrid fixed/floating-point data representations for iterative computations.
Adam B. Kinsman, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2011 An Optimal and Practical Approach to Single Constant Multiplication
abstract
We propose an exact solution to the single constant multiplication (SCM) problem. Existing optimal algorithms are limited to constants of up to 19 bits. Our algorithm requires less than 10 s on average to find a solution for a 32 bit constant. Optimality is guaranteed via an exhaustive search. We analyze two common SCM frameworks and the corresponding search strategies that each framework facilitates. Combining the strengths of both frameworks, we obtain highly aggressive pruning. The various strategies used in our algorithm and their underlying intuition are discussed extensively in this paper.
Jason Thong, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2011 Embedded Debug Architecture for Bypassing Blocking Bugs During Post-Silicon Validation
abstract
Once a bug is found during post-silicon validation, before committing to a silicon respin of the design it is expected that any other bugs, which have escaped pre-silicon verification, to be also identified. This will minimize the number of respins, which in turn will reduce the implementation costs. However, this is hindered by the presence of blocking bugs in one erroneous module that inhibit the search for bugs in other parts of the chip that process data received from this erroneous module. To address this problem, in this paper we propose a novel embedded debug architecture for bypassing the blocking bugs when dealing with deterministic debug experiments.
Ehab Anis Daoud, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2011 A VLSI Architecture and the FPGA Prototype for MPEG-2 Audio/Video Decoding
abstract
This paper details our experience of developing an MPEG-2 audio/video decoder which operates at main level/main profile, 720 × 480 4:2:0 at 29.97 frames per second, with audio at 16 bits, 48 000 samples per second. The design has been developed with a focus on energy-efficiency.
Adam B. Kinsman, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Design-for-Debug Architecture for Distributed Embedded Logic Analysis
abstract
In multi-core designs, distributed embedded logic analyzers with multiple trigger units and trace buffers with real-time offload capability through high-speed trace ports can be placed on-chip. This brings new challenges on how to connect the debug units together in such way that the limited storage space in the trace buffers can be used efficiently. This problem is further aggravated when shadow registers are used to capture data for some signals in the design. In this paper, we propose a new architecture that can dynamically allocate the trace buffers at runtime based on the needs for debug data acquisition coming from multiple data sources and user-programmable priorities. Experimental results show that using the proposed architecture, real-time observability can be improved using only a small amount of on-chip logic hardware, while avoiding excessive storage on-chip.
Ho Fai Ko, Adam B. Kinsman, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.3
2010 Embedded memory binding in FPGAs
abstract
Embedded memory blocks have been integrated infield-programmable gate-arrays (FPGAs) for over a decade. Their count, as well as their capacity and the number of configurations, has increased over time. This growth poses unique challenges to binding the large number of embedded memory blocks to the data vectors that exist in the applications mapped onto FPGAs. In this paper we discuss how this challenge can be addressed algorithmically.
Kaveh Elizeh, Nicola Nicolici
DAC2
2010 Robust design methods for hardware accelerators for iterative algorithms in scientific computing
abstract
The ubiquity of embedded systems of increasing complexity in domains like scientific computing requires computation on models whose complexity has grown beyond what is economical to manage purely in software to requiring hardware acceleration - a key part of which is selecting numerical data representations (bit-width allocation). To address the shortcomings of existing techniques when applied to scientific computing dataflows, we propose a methodology for determining custom hybrid fixed/floating-point data representations for iterative scientific computing applications.
Adam B. Kinsman, Nicola Nicolici
DAC2
2010 Post-silicon validation opportunities, challenges and recent advances
abstract
Post-silicon validation is used to detect and fix bugs in integrated circuits and systems after manufacture. Due to sheer design complexity, it is nearly impossible to detect and fix all bugs before manufacture. Post-silicon validation is a major challenge for future systems. Today, it is largely viewed as an art with very few systematic solutions. As a result, post-silicon validation is an emerging research topic with several exciting opportunities for major innovations in electronic design automation. In this paper, we provide an overview of the post-silicon validation problem and how it differs from traditional pre-silicon verification and manufacturing testing. We also discuss major postsilicon validation challenges and recent advances.
Subhasish Mitra, Sanjit A. Seshia, Nicola Nicolici
DAC3
2010 A novel optimal single constant multiplication algorithm
abstract
Existing optimal single constant multiplication (SCM) algorithms are limited to 19 bit constants. We propose an exact SCM algorithm. For 32 bit constants, the average run time is under 10 seconds. Optimality is ensured via an exhaustive search. The novelty of our algorithm is in how aggressive pruning is achieved by combining two SCM frameworks.
Jason Thong, Nicola Nicolici
DAC2
2010 Combining scan and trace buffers for enhancing real-time observability in post-silicon debugging
abstract
Scan is a known design-for-test technique in manufacturing test that has been successfully applied also to aid post-silicon debugging on testers. However, to achieve real-time observability in-field, embedded trace buffers are needed. In this paper, we discuss how in the presence of enhanced scan chains, trace buffers can be utilized efficiently for real-time debug data acquisition in-field.
Ho Fai Ko, Nicola Nicolici
ETS2
2010 Haptic rendering of deformable objects using a multiple FPGA parallel computing architecture
abstract
High-fidelity simulations of haptic interaction with deformable objects is computationally challenging. In this paper, hardwarebased parallel computing is proposed for finite-element (FE) analysis of soft-object deformation models. A distributed implementation of the Preconditioned Conjugate Gradient (PCG) algorithms on N Field Programmable Gate Array (FPGA) devices can solve the large system of equations arising from FE models at high update rates required for stable haptic interaction. Massive parallelization of the computations is achieved by customizing the hardware architecture to the problem at hand and concurrently employing a large number of adaptive fixed-point computing units. An implementation of this scalable hardware accelerator on four Altera EP3SE110 FPGA devices is capable of performing 230.4 Giga Operations per second in Sparse Matrix by Vector (SpMxV) multiplication. This architecture has successfully enabled real-time simulation of haptic interaction with a 3-dimensional FE model of 6000 nodes at an update rate of 200 Hz. Both static and dynamic linear elastic models have been successfully simulated.
Behzad Mahdavikhah, Ramin Mafi, Shahin Sirouspour, Nicola Nicolici
FPGA4
2010 Combined optimal and heuristic approaches for multiple constant multiplication
abstract
We propose new optimal and heuristic approaches for solving the multiple constant multiplication (MCM) problem. Bounded depth first search (BDFS), our proposed optimal algorithm, is benchmarked on problem sizes that are impractical for the existing optimal method. We focus on MCM problems with few constants but on large bit widths. In this scenario, we outperform the existing heuristics in minimizing the number of adders. In addition, subject to a given quality of solution, our run time is faster. We reuse our heuristics for pruning within BDFS.
Jason Thong, Nicola Nicolici
ICCD2
2010 Automated trace signals selection using the RTL descriptions
abstract
Pre-silicon verification has been traditionally used for eliminating design bugs before tape-out. However, due to the increasing design complexity and the limited accuracy in circuit modelling, the number of the design errors that escape to silicon continues to grow. This is aggravated by the interactions between multiple clock and power domains in the modern system-on-a-chip devices. As a result, structured methods for post-silicon debugging, which aim to detect and localize the bug escapes in silicon, have gained increasing attention in recent years. However, the existing approaches to aid post-silicon debugging primarily rely on the analysis performed using gate-level circuit descriptions. Since design entry is commonly done at the register transfer-level (RTL), the RTL information can be leveraged for the design of the on-chip debug hardware. In particular, in this paper we investigate how to automatically decide which signals to trace in real-time using the RTL information.
Ho Fai Ko, Nicola Nicolici
ITC2
2010 Bit-Width Allocation for Hardware Accelerators for Scientific Computing Using SAT-Modulo Theory
abstract
This paper investigates the application of computational methods via Satisfiability Modulo Theory (SMT) to the bit-width allocation problem for finite precision implementation of numerical calculations, specifically in the context of scientific computing where division frequently occurs. In contrast to the large body of work targeted at the precision aspect of the problem, this paper addresses the range problem where employing SMT leads to more accurate bounds estimation than those provided by other analytical methods, in turn yielding smaller bit-widths, and hence a reduction in hardware cost and/or increased parallelism, while maintaining robustness as is necessary for scientific applications.
Adam B. Kinsman, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 Time-Multiplexed Compressed Test of SOC Designs
abstract
In this paper we observe that the necessary amount of compressed test data transferred from the tester to the embedded cores in a system-on-a-chip (SOC) varies significantly during the testing process. This motivates a novel approach to compressed system-on-a-chip testing based on time-multiplexing the tester channels. It is shown how the introduction of a few control channels will enable the sharing of data channels, on which compressed seeds are passed to every embedded core. Through the use of modular and scalable hardware for on-chip test control and test data decompression, we define a new algorithmic framework for test data compression that is applicable to system-on-a-chip devices comprising intellectual property-protected blocks.
Adam B. Kinsman, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Finite Precision bit-width allocation using SAT-Modulo Theory
abstract
This paper explores the use of SAT-Modulo Theory in determination of bit-widths for finite precision implementation of numerical calculations, specifically in the context of scientific computing where division frequently occurs. Employing SAT-Modulo Theory leads to more accurate bounds estimation than those provided by other analytical methods, in turn yielding smaller bit-widths.
Adam B. Kinsman, Nicola Nicolici
DATE2
2009 Automated data analysis solutions to silicon debug
abstract
Since pre-silicon functional verification is insufficient to detect all design errors, re-spins are often needed due to malfunctions that escape into the silicon. This paper presents an automated software solution to analyze the data collected during silicon debug. The proposed methodology analyzes the test sequences to detect suspects in both the spatial and the temporal domain. A set of software debug techniques are proposed to analyze the acquired data from the hardware testing and provide suggestions for the setup of the test environment in the next debug session. A comprehensive set of experiments demonstrate its effectiveness in terms of run-time and resolution.
Yu-Shen Yang, Nicola Nicolici, Andreas G. Veneris
DATE2
2009 Resource-Efficient Programmable Trigger Units for Post-Silicon Validation
abstract
The decisions on when to acquire debug data during post-silicon validation are determined by trigger events that are programmed into on-chip trigger units. In this paper, we investigate how to design trigger units that are both resource-efficient and runtime programmable. To achieve these two goals, we introduce new architectural features, as well as an algorithm for automatically mapping trigger events onto trigger units.
Ho Fai Ko, Nicola Nicolici
ETS2
2009 Computational bit-width allocation for operations in vector calculus
abstract
Automated bit-width allocation is a key step required for the design of hardware accelerators. The use of computational methods based on SAT-Modulo Theory to the problem of finite-precision bit-width allocation has recently been shown to overcome challenges faced by the known-art, particularly in the scientific computing domain. However, many such real-life applications are specified in terms of vectors and matrices and they are rendered infeasible by expansion into scalar equations. This paper proposes a framework to include operations from vector calculus and thus it enables tackling applications of practically relevant complexity.
Adam B. Kinsman, Nicola Nicolici
ICCD2
2009 Real-Time Lossless Compression for Silicon Debug
abstract
Silicon debug is becoming a key step in the implementation flow for the purpose of identifying and fixing design errors that have escaped pre-silicon verification. To address the lack of observability for the internal circuit nodes during silicon debug, embedded logic analysis enables real-time data acquisition from a limited number of internal signals. In this paper, we propose a novel architecture for embedded logic analysis that enables real-time lossless compression of debug data. To quantify the gain from using lossless compression in embedded logic analysis, we present a new compression-ratio metric that captures the trade-off between the area and the increase in the observation window. The proposed architecture is particularly suitable for in-field debugging on application boards, which have asynchronous events that inhibit the deterministic replay of debug experiments.
Ehab Anis Daoud, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 Algorithms for State Restoration and Trace-Signal Selection for Data Acquisition in Silicon Debug
abstract
To locate and correct design errors that escape pre-silicon verification, silicon debug has become a necessary step in the implementation flow of digital integrated circuits. Embedded logic analysis, which employs on-chip storage units to acquire data in real time from the internal signals of the circuit-under-debug, has emerged as a powerful technique for improving observability during in-system debug. However, as the amount of data that can be acquired is limited by the on-chip storage capacity, the decision on which signals to sample is essential when it is not knownaprioriwhere the bugs will occur. In this paper, we present accelerated algorithms for restoring circuit state elements from the traces collected during a debug session, by exploiting bitwise parallelism. We also introduce new metrics that guide the automated selection of trace signals, which can enhance the real-time observability during in-system debug.
Ho Fai Ko, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 Time-Efficient Single Constant Multiplication Based on Overlapping Digit Patterns
abstract
Common subexpression elimination (CSE) algorithms try to minimize the number of adders (or subtracters) required to implement constant multiplication by searching and substituting common patterns in the CSE representation of a constant. CSE algorithms, in general, cannot find certain patterns due to inherent restrictions in the CSE representation. We propose overlapping digit patterns (ODPs) to remove some of these restrictions. We integrate ODPs into H(k), the best existing heuristic algorithm for single constant multiplication (SCM). H(k) is not applicable to the multiple constant multiplication (MCM) problem, so we cannot consider this problem. Generally, H(k) finds solutions very close to optimal, so there is a strict limitation on any further improvement which applies to any new heuristic. Instead, by integrating ODPs within H(k), we can on average significantly improve the run time of the algorithm (typically by one order of magnitude) while still reducing the number of adders.
Jason Thong, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Power-Aware Testing and Test Strategies for Low Power Devices
Dimitris Gizopoulos, Kaushik Roy 0001, Patrick Girard 0001, Nicola Nicolici, Xiaoqing Wen
DATE4
2008 Automated Trace Signals Identification and State Restoration for Improving Observability in Post-Silicon Validation
abstract
Embedded logic analysis has emerged as a powerful technique for identifying functional bugs during post-silicon validation, as it enables at-speed acquisition of data from the circuit nodes in real-time. Nonetheless, the amount of data that is observed is limited by the capacity of the on-chip trace buffers. This paper introduces an automated method for improving the utilization of the on-chip storage, by identifying a small set of trace signals from which a large number of states can be restored using a compute-efficient algorithm. This enlarged set of data can then be used to aid the search of functional bugs in the fabricated circuit.
Ho Fai Ko, Nicola Nicolici
DATE2
2008 On Automated Trigger Event Generation in Post-Silicon Validation
abstract
When searching for functional bugs in silicon, debug data is acquired after a trigger event occurs. A trigger event can be configured at run-time using a set of control registers that uniquely identify the event that initiates data acquisition. Nonetheless the values loaded in these programmable registers interact only with a set of pre-defined trigger signals that are selected at design-time. If the state conditions required for triggering cannot be expressed directly in terms of the pre-defined trigger signals, the common practice is that the designer manually searches for an equivalent trigger event that can be programmed on-chip. In this paper we investigate if trigger events can be automatically generated from a set of state conditions.
Ho Fai Ko, Nicola Nicolici
DATE2
2008 On Bypassing Blocking Bugs during Post-Silicon Validation
abstract
Design errors (or bugs) inadvertently escape the pre- silicon verification process. Before committing to a re-spin, it is expected that the escaped bugs have been identified during post-silicon validation. This is however hindered by the presence of blocking bugs in one erroneous module that inhibit the search for bugs in other parts of the chip that process data received from the erroneous module. In this paper we discuss how to design a novel embedded debug module that can bypass blocking bugs and aid the designer in validating the first silicon.
Ehab Anis Daoud, Nicola Nicolici
ETS2
2008 Hardware-based parallel computing for real-time haptic rendering of deformable objects
abstract
In this work, a new hardware-based parallel implementation of the iterative conjugate gradient (CG) algorithm for solving such systems of equations is proposed. Fixed point computations are employed to optimize hardware resource usage and to increase parallelism. The proposed implementation adaptively adjusts to variations in the dynamic range of data operands in order to enhance computation accuracy and avoid divergence due to overflow and quantization errors.
Ramin Mafi, Shahin Sirouspour, Brian Moody, Behzad Mahdavikhah, Kaveh Elizeh, Adam B. Kinsman, Nicola Nicolici, Mahyar Fotoohi, D. Madill
IROS7
2008 Distributed Embedded Logic Analysis for Post-Silicon Validation of SOCs
abstract
Post-silicon validation is used to identify design errors in silicon. Its main limitation is real-time observability of the circuit's internal nodes. In this paper, we introduce a novel design-for-debug architecture which automatically allocates distributed trace buffers to handle debug data acquisition requests from multiple sources located in different cores. Using resource-efficient and intelligent control placed on-chip, we show how real-time observability can be improved, thus helping bridge the gap between pre-silicon verification and post-silicon validation for SOC designs.
Ho Fai Ko, Adam B. Kinsman, Nicola Nicolici
ITC3
2008 Scan Division Algorithm for Shift and Capture Power Reduction for At-Speed Test Using Skewed-Load Test Application Strategy
Ho Fai Ko, Nicola Nicolici
J. Electron. Test.2
2008 Guest Editorial
Nicola Nicolici, Patrick Girard 0001
J. Electron. Test.1
2008 Automated Scan Chain Division for Reducing Shift and Capture Power During Broadside At-Speed Test
abstract
Scan chain division has been successfully used to control shift power by enabling mutually exclusive flip-flops at different times during the scan cycle. However, to control capture power without losing transition fault coverage during at-speed scan test, the existing automatic test pattern generation (ATPG) flows need to be modified. In this paper, we present a novel scan chain division algorithm that analyzes the signal dependencies and creates the circuit partitions such that both shift and capture power can be reduced when using the existing ATPG flows. This novel algorithm has been designed for the broadside test application strategy, and a technique for employing partial scan when dividing the scan chains is also proposed.
Ho Fai Ko, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Interactive presentation: Low cost debug architecture using lossy compression for silicon debug
abstract
The size of on-chip trace buffers used for at-speed silicon debug limits the observation window in any debug session. Whenever the debug experiment can be repeated, we propose a novel architecture for at-speed silicon debug that enables a methodology where the designer can iteratively zoom only in the intervals containing erroneous samples. When compared to increasing the size of the trace buffer, the proposed architecture has a small impact on silicon area, while significantly reducing the number of debug sessions
Ehab Anis Daoud, Nicola Nicolici
DATE2
2007 Embedded Tutorial on Low Power Test
abstract
Excessive power during test affects the reliability of digital integrated circuits, test throughput and manufacturing yield. Numerous low power test methods have been investigated over the past decade and new power-aware automatic test pattern generation, design-for-test and test planning techniques have emerged. This embedded tutorial introduces the topic of low power test and it overviews the basic techniques and some recent advancements in this field.
Nicola Nicolici, Xiaoqing Wen
ETS1
2007 On using lossless compression of debug data in embedded logic analysis
abstract
The capacity of on-chip trace buffers employed for embedded logic analysis limits the observation window of a debug experiment. To increase the debug observation window, we propose a novel architecture for embedded logic analysis based on lossless compression. The proposed architecture is particularly useful for in-field debugging of custom circuits that have sources of nondeterministic behavior such as asynchronous interfaces. In order to measure the tradeoff between the area overhead and the increase in the observation window, we also introduce a new compression ratio metric. We use this metric to quantify the performance gain of three lossless compression algorithms suitable for embedded logic analysis.
Ehab Anis Daoud, Nicola Nicolici
ITC2
2007 Test Wrapper Design and Optimization Under Power Constraints for Embedded Cores With Multiple Clock Domains
abstract
Even though many embedded cores contain several clock domains, most published methods for wrapper design have been limited to single-frequency cores. Cumbersome and invasive design techniques, such as insertion of test points, are needed to make these methods applicable to current-generation embedded cores. This paper presents a new method for designing test wrappers for embedded cores with multiple clock domains. The proposed 1500-compliant wrapper prevents clock skew and allows scan chains in different clock domains to shift test data at distinct clock frequencies, which enables a better control of power dissipation during test. We present an integer linear programming (ILP) model that can be used to minimize the core testing time under power constraints for small problem instances, and which can be combined with LP-relaxation to obtain lower bounds on the testing time for larger instances. We also present an efficient heuristic method that is applicable to large problem instances, and which yields the same (optimal) testing time as ILP for small problem instances.
Qiang Xu 0001, Nicola Nicolici, Krishnendu Chakrabarty
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 RTL Scan Design for Skewed-Load At-speed Test under Power Constraints
abstract
This paper discusses an automated method to build scan chains at the register-transfer level (RTL) for power-constrained at-speed testing. By analyzing a circuit at the RTL, where design complexity is lower than at the gate netlist level, one can divide a circuit into multiple partitions, which can be tested independently in order to reduce test power. Despite activating one partition at a time, we show how through conscious construction of scan chains, high transition fault coverage can be achieved, while reducing test time of the circuit when employing third party test generation tools. Furthermore, as shown in experimental results, by constructing scan chains for the partitioned circuit at the RTL, area and performance penalty of the design-for-test hardware may be reduced.
Ho Fai Ko, Nicola Nicolici
ICCD2
2006 DFT Infrastructure for Broadside Two-Pattern Test of Core-Based SOCs
abstract
Existing approaches for modular manufacturing test of core-based system-on-a-chip (SOC) devices do not provide any explicit mechanism for delivering two-pattern tests in the broadside mode, which is necessary to achieve reliable coverage of delay and stuck-open faults. Although wrapper input cells can be enhanced with two memory elements to address this problem, this incur a large test area overhead. This paper proposes a novel architecture for broadside two-pattern test of core-based SOCs without any loss in fault coverage and without increasing the size of the wrapper input cells. The proposed solution combines the dedicated bus-based test access mechanism and functional interconnects for test data transfer in order to provide full controllability of the wrapper input cells in the two consecutive clock cycles required by two-pattern testing. New algorithms for test access mechanism design and test scheduling are proposed and design trade-offs between test area and testing time are discussed using experimental results.
Qiang Xu 0001, Nicola Nicolici
IEEE Trans. Computers2
2006 Multifrequency TAM design for hierarchical SOCs
abstract
The emergence of megacores in hierarchical system-on-a-chip (SOC) presents new challenges to electronic test automation. This paper describes a new framework for designing test access mechanisms (TAMs) for modular testing of hierarchical SOCs. We first explore the concept that TAMs on the same level of design hierarchy employ multiple frequencies for test data transportation. Then we extend this concept to hierarchical SOCs and, by introducing frequency converters at the inputs and outputs of the megacores, the proposed solution not only removes the constraint that the system level TAM width must be wider than the internal TAM width of the megacores, but also facilitates rapid exploration of the tradeoffs between the test application time and the required test area. Experimental results for the ITC'02 SOC Test Benchmarks show that the proposed TAM design algorithms increase the size of the solution space that is explored, which, in turn, will lower the test application time when compared to the existing solutions.
Qiang Xu 0001, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 Diagnosis of Logic Circuits Using Compressed Deterministic Data and On-Chip Response Comparison
abstract
While manufacturing test helps to isolate faulty devices from the good ones, diagnosis is enabling a faster transition from the yield learning to the volume production phase of a new process technology. Given the escalating design complexity, new methods such as embedded deterministic test have been proposed in recent years to deal with the cost of manufacturing test. This paper discusses diagnosis of logic blocks by leveraging the existing embedded deterministic test hardware. The proposed method is based on new techniques for on-chip decompression and comparison of incompletely specified test patterns and test responses. Using experimental data, the tradeoffs between the number of tester channels, on-chip area, and scan time are discussed.
Adam B. Kinsman, Scott Ollivierre, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Register-transfer level functional scan for hierarchical designs
abstract
This paper discusses the potential benefits of inserting scan chains (SCs) in hierarchical designs at the register-transfer level (RTL) of design abstraction. Using new algorithms for functional scan chain design, it is shown how tight timing constraints for design-for-test (DFT) planning at RTL can improve the performance of a circuit, when compared to its gate level counterpart, without any loss in testability.
Ho Fai Ko, Qiang Xu 0001, Nicola Nicolici
ASP-DAC3
2005 Multi-frequency wrapper design and optimization for embedded cores under average power constraints
abstract
This paper presents a new method for designing test wrappers for embedded cores with multiple clock domains. By exploiting the use of multiple shift frequencies, the proposed method improves upon a recent wrapper design method that requires a common shift frequency for the scan elements in the different clock domains. We present an integer linear programming (ILP) model that can be used to minimize the testing time for small problem instances. We also present an efficient heuristic method that is applicable to large problem instances, and which yields the same (optimal) testing time as ILP for small problem instances. Compared to recent work on wrapper design using a single shift frequency, we obtain lower testing times and the reduction in testing time is especially significant under power constraints.
Qiang Xu 0001, Nicola Nicolici, Krishnendu Chakrabarty
DAC2
2005 Time-multiplexed test data decompression architecture for core-based SOCs with improved utilization of tester channels
abstract
In this paper we first observe that the required amount of compressed test data transferred from the tester to the embedded cores in a system-on-a-chip (SOC) varies significantly during the testing process. This motivates a new approach to test data compression, mainly applicable to core-based SOCs, based on time-multiplexing the tester channels through which the on-chip decompressors will receive compressed test data only when necessary. The distinguishing advantage of this approach is that it is suitable for concurrent testing of multiple cores in an SOC and not only will the volume of test data be reduced but also the tester channel utilization will be increased.
Adam B. Kinsman, Nicola Nicolici
ETS2
2005 On concurrent test of wrapped cores and unwrapped logic blocks in SOCs
abstract
System-on-a-chip designs may contain user defined logic or embedded cores that cannot be wrapped for test purposes due to area constraints or timing violations. This paper discusses how these unwrapped logic blocks can be tested rapidly through the TestRail architecture using only the test control mechanism and the test instructions available through IEEE standard for embedded core test. A new test scheduling algorithm, which facilitates concurrent test of both unwrapped logic blocks and wrapped cores, is proposed and experiments show that it outperforms a previous approach when the available number of tester channels and/or the number of unwrapped logic blocks are small.
Qiang Xu 0001, Nicola Nicolici
ITC2
2005 Modular SOC testing with reduced wrapper count
abstract
Motivated by the increasing design for test (DFT) area overhead and potential performance degradation caused by wrapping all the embedded cores for modular system-on-a-chip (SOC) testing, this paper proposes a solution for reducing the number of wrapper boundary register (WBR) cells. By utilizing the functional interconnect topology and the WBRs of the surrounding cores to transfer test stimuli and responses, the WBRs of some cores can be removed without affecting the testability of the SOC. We denote the cores without WBRs as light-wrapped cores and present a new modular SOC test architecture for concurrently testing both the wrapped and the light-wrapped logic cores. Since the WBRs of cores that transfer test stimuli and test responses for light-wrapped cores become shared resources during test, conflicts arise during test scheduling that will negatively impact the test application time. As a consequence, to alleviate this problem, we present a novel test access mechanism (TAM) design algorithm for the proposed SOC test architecture. We conduct experiments on several SOC benchmark circuits and demonstrate that, with an acceptable increase in test application time, the number of WBRs can be significantly decreased. This will ultimately lessen the necessary DFT area for modular SOC testing and reduce the propagation delays between cores.
Qiang Xu 0001, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 Synchronization overhead in SOC compressed test
abstract
Test data compression is an enabling technology for low-cost test. Compression schemes however, require communication between the system under test and the automated test equipment. This communication, referred to in this paper as synchronization overhead, may hinder the effective deployment of this new test technology for core-based systems-on-chip. This paper analyzes the sources of synchronization overhead and discusses the different tradeoffs, such as area overhead, test time and automatic test equipment extensions. A novel scalable and programmable on-chip distribution architecture is proposed, which addresses the synchronization overhead problem and facilitates the use of low cost testers for manufacturing test. The design of the proposed architecture is introduced in a generic framework, and the implementation issues (including the test controller and test set preparation) have been considered for a particular case.
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Wrapper design for multifrequency IP cores
abstract
This paper addresses the testability problems raised by intellectual property cores with multiple clock domains. The proposed solution is based on a novel core wrapper architecture and a new wrapper design algorithm. It is shown how multifrequency at-speed test response capture can be achieved via the design of capture windows without any structural modifications to the logic within the embedded core. The new features in the core wrapper architecture, which introduce limited hardware overhead, can also synchronize the external tester channels with the core's internal scan chains in the shift mode. Thus, the wrapper implementation space can be explored in order to efficiently utilize the available tester bandwidth while meeting the constraints on the maximum internal shift frequency that guarantees low testing time within the given power ratings. Using experimental data, the benefits of the proposed solution are demonstrated by analyzing the tradeoffs between the number of tester channels, testing time, area overhead, and power dissipation.
Qiang Xu 0001, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Modular and rapid testing of SOCs with unwrapped logic blocks
abstract
Extensive research has been carried out for test planning of core-based system-on-a-chip devices. Most of the prior work assumes that all of the embedded cores are wrapped for test purpose. However, some designs may contain user-defined logic or cores that cannot be wrapped due to area constraints or timing violations. This paper discusses how these unwrapped logic blocks can be tested rapidly by adapting the TestRail architecture, which uses only the test control mechanism and the test instructions available through the IEEE 1500 standard for embedded core test. A new test scheduling algorithm, which facilitates a concurrent test of both unwrapped logic blocks and IEEE 1500-wrapped cores, is proposed, and experiments show that it outperforms a previous approach when the available number of tester channels and/or the number of unwrapped logic blocks are small.
Qiang Xu 0001, Nicola Nicolici
IEEE Trans. Very Large Scale Integr. Syst.2
2004 Functional Scan Chain Design at RTL for Skewed-Load Delay Fault Testing
abstract
This paper introduces a new method to construct functional scan chains at the register-transfer level aimed at increasing the delay fault coverage when using the skewed-load test application strategy. It is shown how by consciously creating scan paths prior to logic synthesis, both the transition delay fault coverage and circuit speed can be improved.
Ho Fai Ko, Nicola Nicolici
Asian Test Symposium2
2004 Multi-Frequency Test Access Mechanism Design for Modular SOC Testing
abstract
This paper investigates the applicability of multi-frequency test access mechanism (TAM) design for reducing the system-on-a-chip (SOC) test application time. Based on the bandwidth matching concept the proposed algorithms explore a larger solution space, which, as shown by experimental data, can lead to improved test application time.
Qiang Xu 0001, Nicola Nicolici
Asian Test Symposium2
2004 Wrapper Design for Testing IP Cores with Multiple Clock Domains
abstract
This paper addresses the testability problems raised by embedded cores with multiple clock domains. The proposed solution, based on a novel core wrapper architecture, shows how multi-frequency at-speed test response capture can be achieved using low-speed testers synchronized with high-speed on-chip generated clocks. Using experimental data, the trade-offs between the number of tester channels, testing time, area overhead and power dissipation are discussed.
Qiang Xu 0001, Nicola Nicolici
DATE2
2004 Functional Illinois Scan Design at RTL
abstract
This paper shows that by creating functional scan chains at the register-transfer level (RTL), not only the timing of the circuit can be improved, but also the test data compression provided from the Illinois scan architecture is similar or even better than the gate level counterpart. It was found that DFT infrastructure built using only the control and data flow information available at the RTL can lead to similar improvements in test data compression (for a fault coverage target over 99%), regardless of the final implementation of the logic network or the manufacturing test set.
Ho Fai Ko, Nicola Nicolici
ICCD2
2004 Compressed Embedded Diagnosis of Logic Cores
abstract
This paper introduces a new method for deterministic diagnosis of logic cores. The proposed method is based on on-chip decompression and comparison of incompletely specified test patterns and test responses. Using experimental data, the trade-offs between the number of tester channels, on-chip area and scan time are discussed.
Scott Ollivierre, Adam B. Kinsman, Nicola Nicolici
ICCD3
2004 Time/Area Tradeoffs in Testing Hierarchical SOCs With Hard Mega-Cores
abstract
Motivated by the presence of mega-cores in hierarchical systems-on-a-chip, This work describes a new framework for the design space exploration of multi-level test access mechanisms. Test resources are placed next to the mega-core wrappers, which removes the constraint that upper-level TAM width must be wider than the internal TAM width of the mega-core. The proposed solution can rapidly analyze the tradeoffs between test application time and area overhead and it facilitates test data reuse for hard mega-cores.
Qiang Xu 0001, Nicola Nicolici
ITC2
2004 Testability Trade-Offs for BIST Data Paths
Nicola Nicolici, Bashir M. Al-Hashimi
J. Electron. Test.1
2004 Scan architecture with mutually exclusive scan segment activation for shift- and capture-power reduction
abstract
Power dissipation during scan testing is becoming an important concern as design sizes and gate densities increase. While several approaches have been recently proposed for reducing power dissipation during the shift cycle (minimum transition don't care fill, special scan cells and scan chain partitioning), very little work has been carried out towards reducing the peak power during test response capture and the few existing approaches for reducing capture power rely on complex ATPG algorithms. This paper proposes a scan architecture with mutually exclusive scan segment activation which overcomes the shortcomings of previous approaches. The proposed architecture achieves both shift and capture power reduction with no impact on the performance of the design, and with minimal impact on area and testing time (typically 2-3%). An algorithmic procedure for assigning flip-flips to scan segments enables reuse of test patterns generated by standard ATPG tools. An implementation of the proposed method had been integrated into an automated design flow using commercial synthesis and simulation tools which was used on a wide range of benchmark designs. Reductions up to 57% in average power, and up to 44% and 34% in peak power dissipation during shift and capture cycles, respectively, were obtained when using two scan segments. Increasing the number of scan segments to six leads to reductions of 96% and 80% in average power and respectively maximum number of simultaneous transitions.
Paul M. Rosinger, Bashir M. Al-Hashimi, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2003 Test Data Compression: The System Integrator's Perspective
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici
DATE3
2003 Delay Fault Testing of Core-Based Systems-on-a-Chi
Qiang Xu 0001, Nicola Nicolici
DATE2
2003 Hardware/Software Co-testing of Embedded Memories in Complex SOCs
abstract
A novel approach for testing embedded memories in complex systems-on-a-chip (SOCs) is presented. The proposed solution aims to balance the usage of the existing on-chip resources and dedicated design for test (DFT) hardware such that the functional power constraints are not exceeded during test while trading-off the testing time against DFT area and performance overhead. The suitability of software-centric and hardware-centric approaches for embedded memory testing is examined and to combine the advantages of both directions, a new built-in self-test (BIST)-based method, called hardware/software co-testing, is introduced. The proposed solution is programmable, scalable and guarantees low routine overhead.
Bai Hong Fang, Qiang Xu 0001, Nicola Nicolici
ICCAD3
2003 On Reducing Wrapper Boundary Register Cells in Modular SOC Testing
Qiang Xu 0001, Nicola Nicolici
ITC2
2003 Variable-length input Huffman coding for system-on-a-chip test
abstract
This paper presents a new compression method for embedded core-based system-on-a-chip test. In addition to the new compression method, this paper analyzes the three test data compression environment (TDCE) parameters: compression ratio, area overhead, and test application time, and explains the impact of the factors which influence these three parameters. The proposed method is based on a new variable-length input Huffman coding scheme, which proves to be the key element that determines all the factors that influence the TDCE parameters. Extensive experimental comparisons show that, when compared with three previous approaches, which reduce some test data compression environment's parameters at the expense of the others, the proposed method is capable of improving on all the three TDCE parameters simultaneously.
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2003 Addressing useless test data in core-based system-on-a-chip test
abstract
This paper analyzes the test memory requirements for core-based systems-on-a-chips and identifies useless test data as one of the contributors to the total amount of test data. The useless test data comprises the padding bits necessary to compensate for the difference between the lengths of different chains in multiple scan chain designs. Although useless test data does not represent any relevant test information, it is often unavoidable, and leads to the tradeoff between the test bus width and the volume of test data in multiple scan chain-based cores. Ultimately, this tradeoff influences the test access mechanism design algorithms leading to solutions that have either short test time or low volume of test data. Therefore, in this paper, a novel test methodology is proposed which, by dividing the wrapper scan chains (WSCs) into two or more partitions, and by exploiting automated test equipment memory management features, reduces the amount of useless test data. Extensive experimental results using ISCAS'89 and ITC'02 benchmark circuits are provided to analyze the implications of the number of WSCs in the partition, and the number of partitions on the proposed methodology.
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2002 Improving Compression Ratio, Area Overhead, and Test Application Time for System-on-a-Chip Test Data Compression/Decompression
abstract
Proposes a new test data compression/decompression method for systems-on-a-chip. The method is based on analyzing the factors that influence test parameters: compression ratio, area overhead and test application time. To improve compression ratio, the new method is based on a variable-length input Huffman coding (VIHC), which fully exploits the type and length of the patterns, as well as a novel mapping and reordering algorithm proposed in a pre-processing step. The new VIHC algorithm is combined with a novel parallel on-chip decoder that simultaneously leads to low test application time and low area overhead. It is shown that, unlike three previous approaches which reduce some test parameters at the expense of the others, the proposed method is capable of improving all the three parameters simultaneously. An experimental comparison on benchmark circuits validates the proposed method.
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici
DATE3
2002 Low Power Mixed-Mode BIST Based on Mask Pattern Generation Using Dual LFSR Re-Seeding
abstract
Low power design techniques have been employed for more than two decades, however an emerging problem is satisfying the test power constraints for avoiding destructive test and improving the yield. Our research addresses this problem by proposing a new method which maintains the benefits of mixed-mode built-in self-test (BIST) (low test application time and high fault coverage), and reduces the excessive power dissipation associated with scan-based test. This is achieved by employing dual linear feedback shift register (LFSR) re-seeding and generating mask patterns to reduce the switching activity. Theoretical analysis and experimental results show that the proposed method consistently reduces the switching activity by 25% when compared to the traditional approaches, at the expense of a limited increase in storage requirements.
Paul M. Rosinger, Bashir M. Al-Hashimi, Nicola Nicolici
ICCD3
2002 Integrated Test Data Decompression and Core Wrapper Design for Low-Cost System-on-a-Chip Testing
abstract
This paper discusses an integrated solution for reducing the volume of test data for deterministic system-on-a-chip testing. The proposed solution is based on a new test data decompression architecture which exploits the features of a core wrapper design algorithm targeting the elimination of useless test data. The compressed test data can be transferred from the automatic test equipment to the on-chip decompression architecture using only one test pin, thus providing an efficient reduced pin count test methodology for multiple scan chains-based embedded cores. In addition to reducing the volume of test data, the proposed solution decreases the control overhead, test application time and power dissipation during scan. Further, it also requires lower on-chip area when compared to the testing scenarios which employ decompression architectures for every scan chain and it eliminates the synchronization overhead between the automatic test equipment and the system-on-a-chip. Moreover, the proposed solution is scalable and programmable and, since it can be considered as an add-on to a test access mechanism of a given width, it provides seamless integration with any design flow. Thus, the proposed integrated solution is an efficient low-cost test methodology for systems-on-a-chip.
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici
ITC3
2002 Useless Memory Allocation in System-on-a-Chip Test: Problems and Solutions
abstract
Unlike the existing research direction that focuses on useful test data reduction, this paper analyzes the useless test data memory requirements for system-on-a-chip test. The proposed solution to minimize the useless test memory is based on a new test methodology which combines a novel core wrapper design algorithm with a new test vector deployment procedure stored in the automatic test equipment (ATE). To reduce memory requirements, the proposed core wrapper design finds the minimum number of wrapper scan chain partitions such that the useless memory allocation is minimized in each partition, which facilitates efficient usage of ATE capabilities. Further the new test vector deployment procedure provides a seamless integration with the ATE. When compared to the previously proposed core wrapper design algorithms, the proposed test methodology reduces the memory requirements up to 45%, without any penalties in test area overhead.
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici
VTS3
2002 Multiple Scan Chains for Power Minimization during Test Application in Sequential Circuits
abstract
The paper presents a novel technique for power minimization during test application in sequential circuits using multiple scan chains. The technique is based on a new design for test architecture and a novel test application strategy which reduces spurious transitions in the circuit under test. To facilitate the reduction of spurious transitions, the proposed design for test architecture is based on classifying scan latches into compatible, incompatible and independent scan latches. Based on their classification, the scan latches are partitioned into multiple scan chains and a single extra test vector associated with each scan chain is computed. A new test application strategy which applies the extra test vector to primary inputs while shifting out test responses for each scan chain, minimizes power dissipation by eliminating the spurious transitions which occur in the combinational part of the circuit. The newly introduced multiple scan chain-based technique does not introduce performance degradation and minimizes clock tree power dissipation with minimal impact on both test area and test data overhead. Unlike previous approaches which are test set dependent, and hence are not able to handle large circuits due to the complexity of the design space, the paper shows that with low test area and test data overhead substantial savings in power dissipation during test application are achieved in very low computational time for both small and large test sets. For example, in the case of the benchmark circuit s15850, it takes <6009 in computational time and <1 percent in test area and test data overhead to achieve over 80 percent savings in power dissipation.
Nicola Nicolici, Bashir M. Al-Hashimi
IEEE Trans. Computers1
2002 Power profile manipulation: a new approach for reducing test application time under power constraints
abstract
This paper proposes a power profile manipulation approach which merges two distinct research directions in low power testing: minimization of test power dissipation and test application time reduction under power constraints. It is shown how complementary techniques can be easily combined through this approach to significantly increase test concurrency under power constraints. This is achieved in two steps: in the first step power dissipation is considered a design objective and, consequently, it is minimized; results are further exploited in the second step, when power becomes a design constraint under which the test application time is reduced. A distinctive feature of the proposed power profile manipulation approach is that it can be included in, and consequently improve, any existing power constrained test scheduling algorithm. Extensive experimental results using benchmark circuits, considering test-per-clock, as well as test-per-scan schemes, show that by integrating the proposed power profile manipulation approach into any existing power constrained test scheduling algorithm, savings up to 41 % in test application time are achieved.
Paul M. Rosinger, Bashir M. Al-Hashimi, Nicola Nicolici
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2001 Testability trade-offs for BIST RTL data paths: the case for three dimensional design space
abstract
Power dissipation during test application is an emerging problem due to yield and reliability concerns. This paper focuses on BIST for RTL data paths and discusses testability trade-offs in terms of test application time, BIST area overhead and power dissipation.
Nicola Nicolici, Bashir M. Al-Hashimi
DATE1
2001 Tackling test trade-offs for BIST RTL data paths: BIST area overhead, test application time and power dissipation
abstract
Power dissipation during test application is an emerging problem due to yield and reliability concerns. This paper focuses on BIST for RTL data paths and discusses testability trade-offs in terms of test application time, BIST area overhead and power dissipation. Using a complex validation flow and experimental data for over 30,000 testable data paths, it is shown how test application time decreases asymptotically when increasing power constraints. Further, it is experimentally demonstrated why power conscious test synthesis and test scheduling algorithms are required due to large variations in useless power dissipation as test application time decreases. Finally, while previous research has outlined that test application time decreases as BIST area overhead increases, this paper shows that in order to reach high quality solutions in terms of test application time and BIST area overhead under given power constraints, a three dimensional design space needs to be explored.
Nicola Nicolici, Bashir M. Al-Hashimi
ITC1
2000 Scan Latch Partitioning into Multiple Scan Chains for Power Minimization in Full Scan Sequential Circuits
abstract
Power dissipated during test application is substantially higher than power dissipated during functional operation which can decrease the reliability and lead to yield loss. This paper presents a new technique for power minimization during test application in full scan sequential circuits. The technique is based on classifying scan latches into compatible, incompatible and independent scan latches. Based on their classification, scan latches are partitioned into multiple scan chains. A new test application strategy which applies an extra test vector to primary inputs while shifting out test responses for each scan chain, minimizes power dissipation by eliminating the spurious transitions which occur in the combinational part of the circuit. Unlike previous approaches which are test vector and scan latch order dependent and hence are not able to handle large circuits due to the complexity of the design space, this paper shows that with low test area and test data overhead substantial savings in power dissipation during test application are achieved in a very low computational time. For example, in the case of benchmark circuit s15850 it takes <3600s in computational time and <1% in test area and test data overhead to achieve 80% savings in power dissipation.
Nicola Nicolici, Bashir M. Al-Hashimi
DATE1
2000 Power conscious test synthesis and scheduling for BIST RTL data paths
abstract
Previous research has outlined that power dissipated during test application is substantially higher than during functional operation, which leads to loss of yield and decreases reliability. This paper shows for the first time how power is minimized in BIST RTL data paths by using power conscious test synthesis and test scheduling. According to the necessity for achieving the required test efficiency, power dissipation is classified into necessary and useless power dissipation. According to the occurrence during the testing process, power dissipation is classified into test application and shifting power dissipation. The effect of test synthesis and scheduling on power dissipation is analyzed and power minimization is achieved in two steps. Firstly, during the testable design space exploration only power conscious test synthesis moves are accepted leading to minimization of useless power dissipation. Secondly, module selection during power conscious test scheduling satisfies power constraints while reducing test application time. Experimental results using generic power models show savings up to 28% in test application power dissipation and up to 29% in shifting power dissipation.
Nicola Nicolici, Bashir M. Al-Hashimi
ITC1
2000 BIST hardware synthesis for RTL data paths based on testcompatibility classes
abstract
A new built-in self-test (BIST) methodology for register transfer level (RTL) data paths is presented. The proposed BIST methodology takes advantage of the structural information of the RTL data path and reduces the test application time by grouping same-type modules into test compatibility classes (TCCs). During testing, compatible modules share a small number of test pattern generators at the same test time leading to significant reductions in BIST area overhead, performance degradation and test application time. Module output responses from each TCC are checked by comparators leading to substantial reduction in fault-escape probability. Only a single signature analysis register is required to compress the responses of each TCC which leads to high reductions in volume of output data and overall test application time (the sum of test application time and shifting time required to shift out test responses). This paper shows how the proposed TCC grouping methodology is a general case of the traditional BIST embedding methodology for RTL data paths with both uniform and variable bit width. A new BIST hardware synthesis algorithm employs efficient tabu search-based testable design space exploration which combines the accuracy of incremental test scheduling algorithms and the exploration speed of test scheduling algorithms based on fixed test resource allocation. To illustrate TCC grouping methodology efficiency, various benchmark and complex hypothetical data paths have been evaluated and significant improvements over the BIST embedding methodology are achieved.
Nicola Nicolici, Bashir M. Al-Hashimi, Andrew D. Brown, Alan Christopher Williams
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1999 Efficient BIST Hardware Insertion with Low Test Application Time for Synthesized Data Paths
abstract
In this paper new and efficient BIST methodology and BIST hardware insertion algorithms are presented for RTL data paths obtained from high level synthesis. The methodology is based on concurrent testing of modules with identical physical information by sharing the test pattern generators in a partial intrusion BIST environment. Furthermore, to reduce the number of signature analysis registers and test application time the same type modules are grouped in test compatibility classes and n-input k-bit comparators are used to check the results. The test application time is computed using an incremental test scheduling approach. An existing test scheduling algorithm is modified to obtain an efficient trade-off between the algorithm complexity and testable design space exploration. A cost function based on both test application time and area overhead is defined and a tabu search-based heuristic capable of exploring the solution space in a very rapid time is presented. To reduce the computational time testable design space exploration is carried our in two phases: test application time reduction phase and BIST area reduction phase. Experimental results are included confirming the efficiency of the proposed methodology.
Nicola Nicolici, Bashir M. Al-Hashimi
DATE1
1998 Correction to the Proof of Theorem 2 in "Parallel Signature Analysis Design with Bounds on Aliasing"
Nicola Nicolici, Bashir M. Al-Hashimi
IEEE Trans. Computers1