VLDB 2026 Research / reviewers in the wild / expert
Zainalabedin Navabi
dblp:29/2510 · also Zain Navabi
· DBLP profile ↗
109ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 100 · 4 first-author · 11 since 2021Software engineering, systems software and programming languages · 21 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Domain-Specific Design Abstraction with Emphasis on Neural NetworksabstractThe increasing complexity in digital system design has introduced significant challenges, particularly in maintaining productivity relative to the rapid advancement of technology. This paper addresses the issue by presenting a new level of abstraction, termed Data Transaction Level (DTL), which operates at a higher level than the traditional Register Transfer Level (RTL). The proposed DTL design methodology draws inspiration from RTL but is tailored for the growing complexity of modern systems, particularly in neural network implementations. As neural networks are becoming essential in various applications, this paper applies the DTL methodology to hardware implementations of these networks. A SystemC-based hardware description language is introduced to facilitate the design process. The results section evaluates the effectiveness of the proposed methodology, highlighting increased productivity using the new abstraction and description language. Maryam Rajabalipanah, Zahra Hojati, Zahra Jahanpeima, Zainalabedin Navabi |
DDECS | 4 |
| 2025 | European Test Symposium Teams: an Anniversary SnapshotabstractThe IEEE European Test Symposium (ETS) has been facilitating progress in electronic systems testing since its launch in 1996. On the occasion of its 30th anniversary, this collaborative paper gathers sections by 21 ETS teams to outline their influential ideas and milestones. Each team’s section highlights historical perspective, current research, frameworks and projects as well as forward-looking research agendas in the area of electronic-based circuits and systems testing, reliability, safety, security and validation. This anniversary summary documents how research of various ETS teams, exemplifying the test community, has been evolving and transitioning from concepts to practical standards and Electronic Design Automation (EDA) tools and flows. This legacy is a strong base to drive the next generation of advances in electronic systems testing. Maksim Jenihhin, Jaan Raik, Artur Jutman, Natalia Cherezova, Raimund Ubar, Liviu Miclea, Szilárd Enyedi, Iulia Stefan, Ovidiu Stan, Cosmina Corches, Zebo Peng, Petru Eles, Rolf Drechsler, S. Eggersglüß, Görschwin Fey, Andreas Glowatz, Daniel Tille, Georges Gielen, Anthony Coyette, Wim Dobbelaere, Ronny Vanhooren, Po-Yao Chuang, Erik Jan Marinissen, Giorgio Di Natale, M. Barragan, Paolo Maistri, S. Mir, Vatajelu I. Vatajelu, Paolo Bernardi 0002, Stefano Di Carlo, Paolo Prinetto, Matteo Sonza Reorda, Massimo Violante, Haralampos-G. D. Stratigopoulos, M. K. Michael, Stelios Neophytou, Stavros Hadjitheophanous, Kyriakos Christou, M. Skitsas, Alberto Bosio, Bastien Deveautour, Patrick Girard 0001, Marcello Traiola, Arnaud Virazel, Fernando Santos 0001, Angeliki Kritikakou, Gioele Casagranda, Marzio Vallero, Flavio Vella, Paolo Rech, Letícia Maria Veiras Bolzani, Milos Krstic, Marko S. Andjelkovic, Fabian Vargas 0001, Grigor Tshagharyan, Gurgen Harutunyan, Valery A. Vardanian, Samvel K. Shoukourian, Yervant Zorian, Jennifer Dworak, Kundan Nepal, Theodore W. Manikas, Mottaqiallah Taouil, Moritz Fieback, Anteneh Gebregiorgis, Rajendra Bishnoi, Said Hamdioui, Abhijit Chatterjee, Anurup Saha, Suhasini Komarraju, K. Ma, Chandramouli N. Amarnath, Mehdi Baradaran Tahoori, Mahta Mayahinia, Maryam Rajabalipanah, Katayoon Basharkhah, N. Nosrati, Zahra Jahanpeima, Zainalabedin Navabi, Hans-Joachim Wunderlich, Sybille Hellebrand |
ETS | 79 |
| 2025 | An Integrated Framework for Aging Analysis Based on an Age-Aware Cell Library
Amirmahdi Joudi, Negin Safari, Fatemeh Mohammadzadeh, Katayoon Basharkhah, Fatemeh Sheikhshoaei, Zainalabedin Navabi |
ETS | 6 |
| 2024 | Pico-Programmable Neurons to Reduce Computations for Deep Neural Network AcceleratorsabstractDeep neural networks (DNNs) have shown impressive success in various fields. As a response to the ever-growing precision demand of DNN applications, more complex computational models are created. The growing computational volume has become a challenge for the power and performance efficiency of DNN accelerators. This article presents a new neural architecture to prevent ineffective and redundant computations by using neurons with memory that have decision-making power. In addition, another local memory is used to keep calculation history for removing redundancy by computational reuse. Sparse computing, as another feature, is supported to remove computations of not only zero weights but also zero bits of each weight. The results on conventional datasets such as IMAGENET show a computational reduction of more than$18 \times -150 \times $. This scalable architecture enables 124 GOPS by using 197-mW power. Alireza Nahvy, Zainalabedin Navabi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | A Low-Cost Combinational Approximate MultiplierabstractThis document provides instructions for the design of a purely combinational, multiplexer-based, low-cost multiplier, for machine learning multiplications. The main idea of the algorithm is to remove the most significant bits. Starting from the most significant 1 and after finishing the main multiplication, add back the number of ignored right-hand bits. The main intention of this design is area and power reduction. Since it’s an all-combinational circuit its critical path’s delay is one clock cycle of the main module that it’s going to be used in. Zahra Hojati, Zainalabedin Navabi |
DDECS | 2 |
| 2023 | A Low-cost Residue-based Scheme for Error-resiliency of RNN AcceleratorsabstractAcceleration and power reduction requirements are usually the main constraints for the design of Artificial Neural Network (ANN) accelerators. However, in the case of safety-critical applications like autonomous driving, reliability takes precedence over other requirements. Although ANN algorithms provide a degree of inherent resiliency, the hardware part is still vulnerable to faults and may cause catastrophic failures. This paper proposes using residue codes for detecting soft errors in Recurrent Neural Networks (RNNs), and in particular, Long Short-Term Memory (LSTM) networks. We attach Concurrent Error Detection (CED) hardware units to an entire LSTM structure or its substructures. Depending on the granularity of the components to which they are applied, CEDs are referred to as coarse-grain or fine-grain CEDs. The simulation results show that in fault detection rate and misprediction coverage rate, fine-grain CEDs have a better performance than coarse-grain. Specifically, fine-grain residue-based CEDs provide up to 97% fault detection for extremely large (10-2) bit error rates. Moreover, they reduce the misprediction rate by 84% compared to unprotected LSTM. Nooshin Nosrati, Zainalabedin Navabi |
DDECS | 2 |
| 2023 | Learning Electrical Behavior of Core Interconnects for System-Level Crosstalk PredictionabstractEfficient distribution of tasks in an SOC between various components of an embedded system affect rate of data exchange between cores and obviously the number and fanout of interconnecting cores. Data rate and interconnect fanouts depend on post-layout wire characteristics that, in the worst-case situation, must be evaluated for abovementioned system level decisions. In this work we are making provisions for avoiding this large gap between high-level decision making and low-level physical properties. IP-core interconnects can be fully characterized by post layout information of the IP-core, load properties, and the number of destination cores they are driving. This information can be back-annotated into abstract system-level interconnect models to be used by core integrators for design space exploration (DSE). Fanout and/or frequency of operation of an IP-core can be decided by this DSE environment. In this work, we propose a machine-learning based methodology that uses signoff parasitic information and the actual wire data to generate the dataset and train a model. The model was evaluated in fast high-level SystemC environment for two RISC-V based processors in two SoCs. The models were 26 times faster than the low-level simulations with a crosstalk fault coverage of 1.5% error. Katayoon Basharkhah, Raheleh Sadat Mirhashemi, Nooshin Nosrati, Mohammad-Javad Zare, Zainalabedin Navabi |
ETS | 5 |
| 2022 | Concurrent Error Detection for LSTM AcceleratorsabstractThe widespread usage of Long Short-Term Memory (LSTM) accelerators in time-series related applications necessitates using a protection mechanism against faults caused by wear-out and environmental effects. This paper proposes a Concurrent Error Detection (CED) scheme combining low overhead duplication and residue codes to detect faults in multiply and add stages of LSTM accelerators. For the multiply stage, the CED consists of a multiplier for every LSTM multiplier with a temporal selection of data. For the add stage, the CED adders are shared among the LSTM adders, thus spatial selection is performed. The experimental results show that the proposed method yields good detection probability with a lower area and power overhead in comparison with the traditional duplication techniques that indiscriminately duplicate all hardware structures all the time. Nooshin Nosrati, Seyedeh Maryam Ghasemi, Mahboobe Sadeghipourrudsari, Zainalabedin Navabi |
ETS | 4 |
| 2022 | MLC: A Machine Learning Based Checker For Soft Error Detection In Embedded ProcessorsabstractWith deep submicron scaling, the occurrence of soft errors has become a major reliability challenge for electronic systems. This work proposes a Machine Learning-based Checker (MLC) to protect hard-core processors against radiation-induced soft errors. MLC is an independent hardware unit that implements an ML algorithm to detect soft errors in a processor. The work presented here selects input features from key processor signals for creating a dataset for training. The dataset trains an ML model offline for learning the correct behavior of the processor and detecting soft errors at run-time. The inference of this trained ML is implemented in the MLC hardware that runs along with the processor. Several ML models have been considered for the inference phase, and XGBoost implementation has shown to be the best in terms of hardware overhead and accuracy. The proposed scheme is applied to a RISC-V-like processor, called SAYAC, as a case study. Nooshin Nosrati, Maksim Jenihhin, Zainalabedin Navabi |
IOLTS | 3 |
| 2021 | Online Testing of a Row-Stationary Convolution AcceleratorabstractConvolution accelerators used in safety-critical systems require robustness and resilience to runtime faults. In this paper, we design a processing element for the convolution accelerators that support a row-stationary dataflow. The proposed processing element is equipped with an embedded built-in mechanism for online testing of its computational parts without degrading performance by taking advantage of CNN data sparsity. Furthermore, the proposed architecture uses a low-cost error-detecting code to verify PE interconnections and local memory. The result obtained using post-synthesis simulations reveal, on average, 34% and 30% overheads on the power and area, respectively. This is in exchange for equipping the architecture with components that detect and locate faulty processing elements at runtime. Mohammad Rasoul Roshanshah, Katayoon Basharkhah, Zainalabedin Navabi |
ETS | 3 |
| 2021 | Integrating an Interconnect BIST with Crosstalk Avoidance HardwareabstractAn online interconnect BIST for detecting and avoiding interconnect faults is proposed in this paper. This BIST structure is built upon a new shielding method, the required hardware of which is reutilized in transmitting redundant test data across the interconnect for testing it. Together, BIST and the shielding parts, provide an integrated and novel interconnect testing and crosstalk fault-avoidance platform that share the same hardware. Our proposed integrated technique is modeled in SPICE to evaluate its functionality in terms of avoiding crosstalk faults, and online detection of those that cannot be avoided. We show that faults due to crosstalk are reduced by 49%, and those that still remain are 100% detected. Mahsa Akhsham, Zainalabedin Navabi |
IOLTS | 2 |
| 2020 | DiBA: n-Dimensional Bitslice Architecture for LSTM ImplementationabstractA hardware architecture for the implementation of LSTM neural networks that can be sized to the specific size of the problem is proposed here. Implementation of an LSTM application requires iteration of multiplications, additions, and the activation functions that operate on the stream of data inputs. To handle the iterations, the concept of bitslicing is done to cascade enough slices for an optimum performance depending on the problem size. In order to avoid a large linear array of MAC slices, which would require large adders, these slices are arranged into an n-dimensional structure. Such a structure forces the adder units to become slices of their own, which also operate concurrent with the rest of the hardware in a pipeline fashion. This paper presents this bitslice architecture that can become a fabric for a programmable general-purpose LSTM implementation. The paper also shows an FPGA implementation that uses an on-chip FPGA RAM for the LSTM required memory. The work is compared with other works not considering multidimensional structures, as well as one that considers multi-dimensional cascading. In both cases we show that our structure is faster and uses smaller adder structures. Mahboobe Sadeghipourrudsari, Mohamad Ali Saber, Zainalabedin Navabi |
DDECS | 3 |
| 2020 | Reconfiguration of Embedded Accelerators by Microprogramming for Intensive Loop ComputationsabstractThe work presented in this paper is on reconfigurable accelerators for the implementation of iterative computations and loops that form the core computations of applications like those in digital signal processing and machine learning. The accelerators become computation engines of an embedded system that can be reconfigured by an embedded processor for handling various kernels of embedded applications. This paper presents our MicroProgramed Configurable Accelerator (iMPAC) architecture and compares implementing a kernel (here a matrix multiplication) on this architecture with a) a program running of an embedded processor and b) with a hardwired controller accelerator. Our prototyping on an FPGA shows very little penalty in terms of energy consumption and required clock cycles when compared with the latter, and significant improvement of both energy and timing when compared with the former. At the same time, we have the programming flexibility of the former. Saba Yousefzadeh, Katayoon Basharkhah, Nooshin Nosrati, Maryam Rajabalipanah, Seyedeh Maryam Ghasemi, Zainalabedin Navabi |
DDECS | 6 |
| 2020 | Built-In Predictors for Dynamic Crosstalk AvoidanceabstractIn very deep sub-micrometer technology nodes, signal integrity of interconnects has been drastically jeopardized by crosstalk noise. To make a communication link reliable against crosstalk faults, different detection, correction, and avoidance methods have been proposed at the cost of redundant spatial and information overheads. In this paper, we propose a crosstalk prediction hardware based on an abstract model deduced from low-level interconnect evaluation for new technologies. This predictor monitors the data pattern to be sent through a communication bus and predicts those likely subjected to the crosstalk fault. Thus, to prevent crosstalk faults, the predictor dynamically engages an avoidance or detection mechanism. Top-level buses are the target of the proposed method. Simulation results reveal that the proposed communication channel is more efficient in terms of crosstalk alleviation as well as area and performance overhead compared to the state of the art reliability methods. Rezgar Sadeghi, Zainalabedin Navabi |
ETS | 2 |
| 2020 | ESL, Back-annotating Crosstalk Fault Models into High-level Communication LinksabstractAt the system-level, cores are put together using interconnects that we refer to as high-level communication links. This paper presents an abstract interconnect model for cores connecting to each other to estimate, and thus model, crosstalk noise resulting from the physical properties of interconnects. Such models consider the effects of adjacent wires on each other in the form of weighted transitions. Transition weights are extracted by DC analysis of interconnect SPICE models. These weights form our raw-models, which are then specialized by AC analysis of RLC interconnect models in a mixed-signal simulation environment. The latter analyses establish weight thresholds for glitch faults. Our simulations show that if we were to use only DC-based models for crosstalk faults, we would be over / under-estimating faults as compared with models that are specialized by AC simulation runs. For higher data rates, Specialized models perform an order of magnitude better than DC-based models for crosstalk fault detection. Katayoon Basharkhah, Rezgar Sadeghi, Nooshin Nosrati, Zainalabedin Navabi |
VTS | 4 |
| 2020 | LUT Input Reordering to Reduce Aging Impact on FPGA LUTsabstractIn this article, we propose a fine-grained FPGA aging mitigation method. Our method focuses on Look Up Tables (LUTs) on which Boolean functions are mapped. Based on our observations, for any configuration, even if it is carefully selected, a number of LUT transistors experience severe stress rates. Therefore, an algorithm is presented to select several alternative configurations for each LUT. Alternative configurations are obtained by LUT input reordering. These alternative configurations are rotationally loaded into the FPGA. Experimental results shows that our method achieves 263 and 14.1 percent Mean Time To Failure (MTTF) improvement for Hot Carrier Injection (HCI) and Bias Temperature Instability (BTI), respectively. Additionally, due to changing only local routings, our method imposes up to 1 percent performance overhead to the systems. Rezgar Sadeghi, Zainalabedin Navabi |
IEEE Trans. Computers | 3 |
| 2020 | Selecting Representative Critical Paths for Sensor Placement Provides Early FPGA Aging InformationabstractThis article proposes a methodology for critical path selection and delay sensor insertion for aging monitoring in field-programmable gate arrays (FPGAs). The aging information can be used to reconfigure FPGAs to achieve better performance or wear leveling. To filter out the paths aging less aggressively, our methodology decides based on both physical-level parameters [e.g., path delay, process variation, temperature, static stress (duty cycle), and dynamic stress (switching activity) as well as aging-relevant parameters (e.g., fan-out, and endpoint physical location)]. After selection of an optimal set of paths, age sensors, introduced in our earlier work, are placed in a distributed fashion throughout the FPGA area to provide aging information of its various parts as they are being used. Accurate aging models for FPGAs are required to determine the contribution of the above-mentioned parameters in aging. In this article, ASICs' aging models are adapted for FPGAs through measurement-based fitting approach. The experimental results using various benchmarks reveal that our algorithm selects paths with a minimum error. Zainalabedin Navabi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Test Adapted Shielding by a Multipurpose Crosstalk Avoidance SchemeabstractCrosstalk noise due to the electromagnetic coupling between wires adversely affects VLSI circuit performance. This makes interconnect testing an important issue in reliability evaluation that causes extra area and hardware overhead. In this paper, we present a novel test methodology that we refer to as Test Adapted Shielding (TAS) in order to improve testing challenges and also to optimize crosstalk noise. The test structure of TAS is designed to improve circuit performance by insertion of modified shield lines. The developed hardware for test data is smart enough to avoid aggregation over victim lines in crosstalk. Besides, the TAS test methodology is optimized in power consumption, complexity, fault occurrence, and fault detection. The proposed technique is implemented in HSPICE and MATLAB and the extracted results are used to evaluate its functionality. We show that crosstalk noise is decreased by about 49% and test coverage is 100% for negative and positive glitches. Mahsa Akhsham, Atefesadat Seyedolhosseini, Zainalabedin Navabi |
ETS | 3 |
| 2019 | Back-annotation of Interconnect Physical Properties for System-Level Crosstalk ModelingabstractAs digital design moves into higher abstraction levels, chip-level communications become more complex, thus harder to consider low-level interconnection signal effects. Abstract interconnect models are required in order to be able to bring signal integrity issues such as crosstalk into the hands of the high-level system designer. Such abstraction can be based on, and back-annotated from, the existing crosstalk fault models, of which MDSI is a candidate. While MDSI, by considering RLC interconnect effects and not just RC, is an improvement over some other proposed models, its simplified adjacent line effects makes it inadequate for the newer technologies. For the purpose of back-annotating physical properties for system-level crosstalk modeling, we have developed a new model based on MDSI that we refer to as Weighted-MDSI. In this modeling, the effect of lines causing a cross-talk noise on a victim line is weighted by their distances to the victim. This paper extracts parameters of our presented Weighted-MDSI from HSPICE simulation runs. The parameters extracted as such are programmed into SystemC communication channels for high-level reliability evaluations and other system-level decision makings. Furthermore, SystemC-AMS models are considered as an alternative for adjusting Weighted-MDSI parameters, and for verification purposes. Rezgar Sadeghi, Nooshin Nosrati, Katayoon Basharkhah, Zainalabedin Navabi |
ETS | 4 |
| 2018 | Performance and Energy Enhancement through an Online Single/Multi Level Mode Switching Cache ArchitectureabstractSTT-RAM cells can be considered as an alternative or a hybrid addition to today's SRAM-based cache memories. This is mostly because of their scalability and low leakage power. Moreover, their data storing mechanism (storing the value as resistance) makes them very suitable and applicable for multivalue cache architectures. This feature results in system performance enhancement without any area overhead. On the other hand, the required two-step read/write procedure in multilevel cells results in a non-uniform time access and energy and power overhead on the system. In this paper, we propose a new architecture to dynamically swap data between soft (fast read access) and hard (slow read access) bits in ML cell. Moreover, by reconfiguring cache block size, the proposed architecture can switch between ML and SL modes at runtime. In other words, the swapping method places the hot part of each cache block into soft-bits and the less accessed part into the hard-bits. The SL/ML switching method benefits from the low latency and energy of SL mode and the high storing capacity of ML mode at the same time. Although experimental results show that our proposed method slightly increases the miss rate compared with the conventional ML caches, the performance and energy are improved by 4.9% and 6.5%, respectively. Also, the storage overhead of our method is about 1% that is negligible. Ramin Rezaeizadeh Rookerd, Somayeh Sadeghi Kohan, Zainalabedin Navabi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | Near-Optimal Node Selection Procedure for Aging Monitor PlacementabstractTransistor and interconnect wearout is accelerated with transistor scaling resulting in timing variations and consequently reliability challenges in digital circuits. With the emergence of new issues like Electro-migration these problems are getting more crucial. Age monitoring methods can be used to predict and deal with the aging problem. Selecting appropriate locations for placement of aging monitors is an important issue. In this work we propose a procedure for selection of appropriate internal nodes that expose smaller overheads to the circuit, using correlation between nodes and the shareability amongst them. To select internal nodes, we first prune some nodes based on some attributes and thus provide a near-optimal solution that can effectively get a number of internal nodes and consider the effects of electro-migration as well. We have applied our proposed scheme to severalprocessors and ITC benchmarks and have looked at its effectiveness for these circuits. Somayeh Sadeghi Kohan, Arash Vafaei, Zainalabedin Navabi |
IOLTS | 3 |
| 2018 | Scalable Symbolic Simulation-Based Automatic Correction of Modern Processors
Fatemeh Refan, Bijan Alizadeh, Zainalabedin Navabi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Automatic Correction of Dynamic Power Management Architecture in Modern ProcessorsabstractThe increasing demand for lower power forces designers to use sophisticated power management strategies such as multivoltage and power gating which are often accompanied with many design bugs. Correcting such bugs can be a time-consuming process that requires considerable manual efforts. In this paper, we propose a scalable automated method for correcting dynamic power management architectures by an incremental SAT-based mechanism. First, an initial counterexample (CEX) is generated by checking the equivalency between the specification model of the processor and its buggy implementation model. Then, we find two candidate solutions instead of one to satisfy this CEX. If two solutions are not equivalent, we generate new CEX in an iterative process which effectively converges into the final solution. The proposed method enables designers to correct multiple bugs such as missing isolation cells between two power domains, disordering in the sequence of control signals, error in the data restoring or saving, and powering off in always-on domains which are not addressed by existing methods. We have shown the effectiveness of our method on modern processors supporting complex power management mechanisms. The results confirm that our proposed method, respectively, reduces symbolic simulation steps and runtime by 2.33× and 47.93× compared to the state-of-the-art methods. Reza Sharafinejad, Bijan Alizadeh, Zainalabedin Navabi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | TruncApp: A truncation-based approximate divider for energy efficient DSP applicationsabstractIn this paper, we present a high speed yet energy efficient approximate divider where the division operation is performed by multiplying the dividend by the inverse of the divisor. In this structure, truncated value of the dividend is multiplied exactly (approximately) by the approximate inverse value of divisor. To assess the efficacy of the proposed divider, its design parameters are extracted and compared to those of a number of prior art dividers in a 45nm CMOS technology. Results reveal that this structure provides 66% and 52% improvements in the area and energy consumption, respectively, compared to the most advanced prior art approximate divider. In addition, delay and energy consumption of the division operation are reduced about 94.4% and 99.93%, respectively, compared to those of an exact SRT radix-4 divider. Finally, the efficacy of the proposed divider in image processing application is studied. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Zainalabedin Navabi |
DATE | 5 |
| 2017 | Online Profiling for cluster-specific variable rate refreshing in high-density DRAM systemsabstractMulti-rate refresh techniques are among the methods that use non-uniformity in retention time of DRAM cells to reduce the DRAM refresh overheads. Unfortunately, retention time of some DRAM cells may change unpredictably over time due to variable retention time (VRT). In this paper, we propose an Online Profiler that divides DRAM cells into clusters and proactively tests and measures retention time of each cluster over time. The Online Profiler decides on increasing refresh period of a cluster based on a measured retention time, where this retention time has passed all tests of different data sets. Also, for ensuring maximum data integrity, the Online Profiler reads the entire memory periodically for correction of possible errors. We show that our proposed mechanism, that uses cluster-specific variable rate refreshing, can provide reliable operation while reducing refresh overhead of the performance by 6%, 13%, and 23%, and Energy-Delay Product (EDP) by 7%, 13%, and 27% for 32GB, 64GB, and 128GB DRAM modules, respectively. Rasool Sharifi, Zainalabedin Navabi |
ETS | 2 |
| 2017 | SENSIBle: A Highly Scalable SENsor DeSIgn for Path-Based Age Monitoring in FPGAsabstractThis paper proposes a highly scalable sensor design for late transition detection in FPGA based platforms. Transition delays occur because of aging mechanisms such as Biased Temperature Instability(BTI) and Hot Carrier Injection (HCI). We propose a sensor clock (SCLK) that is a function of minimum slack time of a set of paths selected for age monitoring. There will be one such clock for many sensors as are needed in an entire FPGA. Our proposed sensor architecture makes it possible for a single SCLK to be shared by all sensors. Additionally, the proposed sensor occupies one slice (basic FPGA logic block), which leads to low area, power, and performance overhead. Using Artix-7-based board, experimental results demonstrate that the proposed aging sensor detects aging earlier than existing sensors and provides less power and performance overheads. Zana Ghaderi, Zainalabedin Navabi, Elaheh Bozorgzadeh, Nader Bagherzadeh |
IEEE Trans. Computers | 3 |
| 2017 | Bridging Presilicon and Postsilicon Debugging by Instruction-Based Trace Signal Selection in Modern ProcessorsabstractAlthough using presilicon information in postsilicon debugging phase seems interesting, space and time limitations of existing formal verification tools restrict the possibility of this idea. In this paper, the effective usage of presilicon information to enhance postsilicon trace signal selection in modern processors is discussed. Furthermore, a novel architecture for dynamic per-cycle selection of signals based on the present instruction is implemented and synthesized. In presilicon phase, first, a set of controlling signals and their corresponding rules are extracted manually. Based on these rules, a set of data from model is extracted using an automatic formal method, which determines which signals should be traced at postsilicon according to the values of controlling signals. This mechanism alone results in an average of 79% and 54% bits to be pruned from the traceable signals for Leon3 and multithreaded DLX processors and 86% and 75% improvement when used in conjunction with traditional methods, respectively. Fatemeh Refan, Bijan Alizadeh, Zainalabedin Navabi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Path selection and sensor insertion flow for age monitoring in FPGAs
Zana Ghaderi, Elaheh Bozorgzadeh, Zainalabedin Navabi |
DATE | 4 |
| 2016 | Prolonging Lifetime of Non-volatile Last Level Caches with Cluster MappingabstractRecently, work has been done on using nonvolatile cells, such as Spin Transfer Torque RAM (STT-RAM) or Magnetic RAM (M-RAM), to construct last level caches (LLC). These structures mitigate the leakage power and density problem found in traditional SRAM cells. However, the low endurance of nonvolatile caches decreases the lifetime of the LLC. Therefore, an effective wear-leveling technique is required to tackle this issue. In this paper, we propose the inter-set algorithm that distributes the write traffic to all portions of the cache. Our method is based on cluster mapping that dynamically replaces two clusters during the operation of system. Since the inter-set algorithm is based on data movement, a large amount of data must transfer in each replacement. For an efficient data movement with a minimum effect on performance, we develop the novel scheduling technique that utilizes the idle time of the LLC in the computation phase of the processors. Our approach effectively improves the lifetime of LLC with negligible performance and area overhead. Using these methods in a quad core system with 2MB LLC, we can improve the lifetime of non-volatile LLC by 30% on average. Morteza Soltani, Zainalabedin Navabi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2016 | Optimistic clock adjustment for preventing Better-than-worst-case violationsabstractThis paper proposes a technique for reducing power consumption and increasing pipeline performance using the state of art “Better than worst case” design. This method uses a violation predication mechanism that places the optimum number of Transition Detectors and a novel Time Borrowing technique to prevent potential timing errors. In this work the tradeoff between power consumption (Transition Detectors dynamic and leakage power) and accuracy of prediction is examined and the optimum number of transistors is placed on the circuit based on the available power budget. Seyedeh Hanieh Hashemi, Reza Namazian, Zainalabedin Navabi |
VLSI-SoC | 3 |
| 2016 | Stochastic testing of processing cores in a many-core architecture
Arezoo Kamran, Zainalabedin Navabi |
Integr. | 2 |
| 2015 | Power-aware online testing of manycore systems in the dark silicon era
M. H. Haghbayan, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Zainalabedin Navabi, Hannu Tenhunen |
DATE | 6 |
| 2015 | Signature oriented model pruning to facilitate multi-threaded processors debuggingabstractIn this paper, we propose a signature based pruning technique to facilitate the debugging of multi-threaded processors. To accomplish this, a pipelined implementation of the multi-threaded processor model is checked for correspondence against the specification model based on flushing proof. Then, a two-stage signature oriented pruning method is proposed to avoid the space explosion problem caused by inserting debugging facilities in the model. The results show an average improvement of 47%, and 71% in the size of decision formula and CPU time for the DLX processor, respectively. Fatemeh Refan, Bijan Alizadeh, Zainalabedin Navabi |
VTS | 3 |
| 2015 | A near-threshold 7T SRAM cell with high write and read margins and low write time for sub-20 nm FinFET technologies
Mohammad Ansari, Hassan Afzali-Kusha, Behzad Ebrahimi, Zainalabedin Navabi, Ali Afzali-Kusha, Massoud Pedram |
Integr. | 4 |
| 2015 | Automatic High-Level Data-Flow Synthesis and Optimization of Polynomial Datapaths Using Functional DecompositionabstractThis paper concentrates on high-level data-flow optimization and synthesis techniques for datapath intensive designs such as those in Digital Signal Processing (DSP), computer graphics and embedded systems applications, which are modeled as polynomial computations over Z2n1x Z2n2x . . . x Z2ndto Z2m. Our main contribution in this paper is proposing an optimization method based on functional decomposition of multivariate polynomial in the form of f(x) = g(x) o h(x) + f0= g(h(x)) + f0to obtain good building blocks, and vanishing polynomials over Z2mto add/delete redundancy to/from given polynomial functions to extract further common sub-expressions. Experimental results for combinational implementation of the designs have shown an average saving of 38.85 and 18.85 percent in the number of gates and critical path delay, respectively, compared with the state-of-the-art techniques. Regarding the comparison with our previous works, the area and delay are improved by 10.87 and 11.22 percent, respectively. Furthermore, experimental results of sequential implementations have shown an average saving of 39.26 and 34.70 percent in the area and the latency, respectively, compared with the state-of-the-art techniques. Samaneh Ghandali, Bijan Alizadeh, Zainalabedin Navabi |
IEEE Trans. Computers | 4 |
| 2014 | Automatic correction of certain design errors using mutation techniqueabstractIn this paper, we introduce a new technique that makes use of satisfiability (SAT) based debugging techniques along with a mutation-based technique to correct certain design errors in digital designs automatically. The experimental results demonstrate that our proposed method enables us to locate and correct multiple bugs by targeting gate replacements and wire exchanging within reasonable run-time and memory usage for several designs. Payman Behnam, Bijan Alizadeh, Zainalabedin Navabi |
ETS | 3 |
| 2014 | Homogeneous many-core processor system test distribution and execution mechanismabstractIn response to reliability challenges of new systems being built, we are proposing a scalable Self-Test architecture for many-core processor systems. This BIST architecture periodically distributes test stimuli among identical processing cores in a many-core processor system, suspends normal operation of individual processing cores, applies test, detects faulty cores, and removes them from the system if any are found faulty. Test is continuously performed without any perceptible down-time to the end-user, realizing a many-core processor system with self-healing capability. Arezoo Kamran, Zainalabedin Navabi |
ETS | 2 |
| 2014 | Improving polynomial datapath debugging with HEDsabstractIn this paper, we introduce a formal and scalable debugging approach to derive a reduced ordered set of design error candidates in polynomial datapath designs. To make our debugging method scalable for large designs, we utilize a Modular Horner Expansion Diagram (M-HED), which has been shown to be a scalable high level decision model. In our method, we extract data dependency graphs from the polynomial datapath designs using static slicing. Then we combine backward and forward path tracing to extract a reduced set of error candidates. In order to increase the accuracy of the method in the presence of multiple design errors, we rank the error candidates in decreasing order of their probability of being an error using a proposed priority criterion. In order to evaluate the effectiveness of our method, we have applied it to several large designs. The experimental results show that the proposed method enables us to locate even multiple errors with high accuracy in a short run time. Somayeh Sadeghi Kohan, Payman Behnam, Bijan Alizadeh, Zainalabedin Navabi |
ETS | 5 |
| 2014 | An off-line MDSI interconnect BIST incorporated in BS 1149.1abstractThis paper presents an off-line interconnect test methodology that implements the MDSI (Maximal Dominant Signal Integrity) crosstalk fault model. The test methodology consists of MDSI test pattern generators and response analyzers that are incorporated into the IEEE BS 1149.1 Standard on the two sides of an interconnect. This work is the first in implementing MDSI hardware structure. Our method is compared with hardware structures implementing MA interconnect tests. Marzieh Mohammadi, Somayeh Sadeghi Kohan, Nasser Masoumi, Zainalabedin Navabi |
ETS | 4 |
| 2014 | High-level design space exploration of locally linear neuro-fuzzy models for embedded systems
Mohammadreza Baharani, Hamid Noori, Mohammad Aliasgari, Zainalabedin Navabi |
Fuzzy Sets Syst. | 4 |
| 2013 | Online periodic test mechanism for homogeneous many-core processorsabstractA possible solution to the reliability challenges of new fabrication technologies is self-test and self-reconfiguration with no or limited external control. This necessitates the inclusion of mechanisms for detection, recovery and reconfiguration, providing the chip with self-healing capability. While appropriate mechanisms for system recovery and reconfiguration exist, detection remains to be a challenge in realization of self-healing and graceful degradable many-core processors. In this paper we propose a scalable test architecture to distribute test stimuli among homogeneous processing cores in a many-core processor. We have incorporated test infrastructure in the architecture that periodically suspends normal operation of processing cores and applies test to them. This procedure is performed in an online fashion without any downtime visible to the end user. Our proposed fault detection mechanism can be accompanied with appropriate recovery and reconfiguration techniques in order to realize a many-core processor with self-healing capability. Arezoo Kamran, Zainalabedin Navabi |
VLSI-SoC | 2 |
| 2013 | Embedded tutorials: Embedded tutorial 1: Cell-aware test-from gates to transistorsabstractDevices manufactured in 20 nm and smaller geometry technologies will potentially be very large by today's standards, they will also have new characteristics implied by things like process variability and adoption of FinFET transistors. The industry has cumulatively adopted more and more sophisticated fault models that use timing as well as layout information. There is a growing body of experimental data showing it is still insufficient. The next area of focus will be the quality of test. Cell-aware test is one of the most promising approaches developed over the last five years aimed at improving the quality of test while maintaining the efficiency of gate-level approach. This approach combines two levels of abstraction to provide trade-offs between accuracy and efficiency. The first step creates the cell-aware test library models. It starts with standard cell libraries and performs layout extraction. Realistic defects (bridges and opens) are injected into the SPICE netlist, and analog fault simulation is performed to determine the conditions under which the defects are detected. Those conditions are aggregated to create a compact and efficient representation of the libraries for ATPG done at the gate-level. Generation of library views for cell-aware test is performed only once for a given standard cell library. The final cell-aware ATPG generates the high quality test patterns based on the cell-aware library views. This guarantees that the investment in gate-level ATPG infrastructure could be efficiently utilized. The technology has been used on a number of high-volume industrial designs. The experimental data show a significant increase of defect coverage and the corresponding improvement of defect rate. Janusz Rajski, Miodrag Potkonjak, Adit D. Singh, Abhijit Chatterjee, Zainalabedin Navabi, Matthew R. Guthaus, Sezer Gören 0001 |
VLSI-SoC | 5 |
| 2012 | A Probabilistic and Constraint Based Approach for Low Power Test GenerationabstractInserting scan chain into the circuit alongside the combinational automatic test pattern generation (ATPG) is the most commonly used test technique for digital circuits. Since power dissipation in test mode is generally much higher than in the functional mode, some considerations should be made during ATPG. This paper presents a probabilistic and constraint based approach for scan-based low power test generation. This ATPG exploits signal probability analysis to estimate fault detection probability as well as signal transition activities to guide test generation process. The effectiveness of the proposed approach has been evaluated by applying it to ISCAS85 and ISCAS89 benchmarks with three different power constraints, including propagation, capture, and shift. Hossein Sabaghian Bidgoli, Majid Namaki-Shoushtari, Zainalabedin Navabi |
Asian Test Symposium | 3 |
| 2012 | Power constraint testing for multi-clock domain SoCs using concurrent hybrid BISTabstractThis paper presents a novel approach for selecting optimal pseudo random and deterministic test patterns and minimizing test time for multi-clock domain SoCs based on a hybrid BIST architecture for each core. For test scheduling, a concurrent method considering peak power upper bound is used. A test scheduling graph is presented for modeling concurrent hybrid BIST test scheduling. Furthermore, a heuristic is proposed for selecting cores to be tested concurrently and the order of applying sequence of test patterns to each core. Experimental results show that the proposed heuristics for both selecting groups of cores to be tested concurrently during the SoC test process, and determining the amount of deterministic and pseudo random test patterns for each core, give us an optimized method for multi clock domain SoC testing compared with the existing methods. M. H. Haghbayan, Saeed Safari, Zainalabedin Navabi |
DDECS | 3 |
| 2012 | Effective RT-level software-based self-testing of embedded processor coresabstractEmbedded processors are often used in systems that are both safety-critical and long-living. Therefore testing of these devices is critical not only after production, but also in the field. Due to the limited accessibility of embedded processors, testing of such systems is a major challenge in terms of ultra-deep submicron issues. Scan based testing methods cannot be applied to embedded processor cores which cannot be modified to meet the design requirements for scan insertion. From the other side, generating test vectors for these high gate count devices is a major task. This paper presents a high level, component-oriented software-based self-testing method which achieves a high stuck-at-fault coverage for an embedded processor core. The method requires no DFT or changing the processor architecture. The proposed method is high level in the sense that it is based on the knowledge of the Instruction Set Architecture (ISA) and Register Transfer Level (RTL) description of the processor. The method is well suited for meeting the challenges of testing SoCs which contain embedded processor cores. Our methodology is superior in terms of test quality in such a way that significantly increases fault coverage and reduces test time. The proposed method outperforms all existing method in terms of fault coverage. Parisa Kabiri, Zainalabedin Navabi |
DDECS | 2 |
| 2012 | Soft Error Analysis on Communication Channels in On-Chip Communication NetworksabstractIncrease in the number of transistors has resulted in more vulnerability of digital systems to soft-errors. For a long time, designers have used the worst case scenarios in their designs to alleviate such disturbances, but due to the overall increase in cost in relation with system performance, manufacturing based on worst case scenarios for transient faults is no longer a viable option. Meanwhile, demand for higher performance systems, multi-core processors, and Networks on Chip (NoC) communication structures are affecting design methodologies, and must be considered while dealing with soft errors. Demand for more flexible and high performance systems have made such communication infrastructures so complex that a perfect error free communication is hard to come by. This paper proposes a method to evaluate the reliability of communication infrastructures in a Multi-Processor System on Chip (MPSoC). In this method, we evaluate vulnerability of communication systems based on data flow vulnerability in the network. The experimental results show estimated values for silence data corruption (SDC) of MPSoC's communication infrastructures under various routing algorithms and traffic loads. Mohammadreza Najafi, Saeed Safari, Zainalabedin Navabi |
DSD | 3 |
| 2012 | BS 1149.1 extensions for an online interconnect fault detection and recoveryabstractLoss of signal integrity in today's deep sub-micron designs puts communication links at a higher risk of permanent or more frequent intermittent faults. This results in performance and reliability reduction. This paper presents an online interconnect BIST method that applies to a hybrid serial/parallel communication scheme. The proposed BIST method is implemented by a simple extension to the boundary scan standard, which facilitates online testing methodology with negligible hardware overhead. The online hardware works in the idle state of the Boundary Scan TAP controller. The proposed method includes fault detection and diagnosis phases. Moreover, for error handling, it uses the same test hardware added to the communication interface. It effectively reuses the existing boundary scan structure to act as a signature generator, an error detector and locater for testing interconnects, and an error handling mechanism. Our method can detect about 90% of the interconnect faults after six block transfers. Somayeh Sadeghi Kohan, Majid Namaki-Shoushtari, Fatemeh Javaheri, Zainalabedin Navabi |
ITC | 4 |
| 2012 | Polynomial datapath synthesis and optimization based on vanishing polynomial over Z2m and algebraic techniquesabstractThe growing market for Digital Signal Processing (DSP), Computer graphics and embedded systems applications that can be modeled as polynomial computations in their datapath designs, requires improvements in high-level synthesis and optimization techniques for such systems. This paper concentrates on how to find common sub-expressions between s given polynomial functions over Z2n1× Z2n2× ... × Z2ndto Z2min order to optimize the area and delay as much as possible. Our main contributions in this paper is proposing an optimization method based on adding/deleting vanishing polynomials over Z2m, i.e., those polynomials that are equivalent to zero over Z2m, to/from given polynomial functions in the hope of achieving further common sub-expressions. After applying our optimization techniques, experimental comparisons with the state-of-the-art techniques show an average improvement in the area by 36.80% with an average delay decrease of 2.41%. Regarding the comparison with our previous works, the area and delay are improved by 21.4% and 8.7% respectively. Samaneh Ghandali, Bijan Alizadeh, Zainalabedin Navabi |
MEMOCODE | 3 |
| 2011 | Mapping Transaction Level Faults to Stuck-At Faults in Communication HardwareabstractAdvances in semiconductor technology, the increasing complexity of digital systems and demand for faster time to market, have raised the level of design from transistor to Electronic System Level (ESL). However, digital system testing remains at lower levels of abstraction. To cover the gap between system-level design and test, this paper presents a method of testing communication links at the ESL. For this purpose, system-level communication links are formally represented by Timed Automata (TA) to perform the fault simulation process automatically. This is facilitated by a set of high-level fault models that includes faults for data and control parts of a communication link. We show how the proposed high-level fault models map into faults at the gate level in communication hardware. The proposed test strategy not only applies communication links, but also can be used for testing processing elements using an appropriate fault model. Fatemeh Javaheri, Majid Namaki-Shoushtari, Parastoo Kamranfar, Zainalabedin Navabi |
Asian Test Symposium | 4 |
| 2011 | Online Test Macro Scheduling and Assignment in MPSoC DesignabstractDue to unreliability of the cores in embedded systems in deep sub-micron technologies, a method for testing cores in the field is needed. In this paper an online method for testing cores of embedded designs is presented. The proposed task scheduling method runs the test routine ASAP periodically considering the real time constraints. A software test routine based on a proposed method will be generated and a task scheduling process including the test task (for each core) and other existing applications of the embedded system will be presented. A software based checksum is issued for online test result analysis that shortens the memory usage of the test process. Experimental results show that this method improves the test application time (TAT) and fault coverage (in proportion to TAT) as compared with the existing methods. Behnam Khodabandeloo, Seyyed Alireza Hoseini, Sajjad Taheri, M. H. Haghbayan, Mahmood Reza Babaei, Zainalabedin Navabi |
Asian Test Symposium | 6 |
| 2011 | Adaptation of Standard RT Level BIST Architectures for System Level Communication TestingabstractTest and testability are essential concerns for design in any abstraction level, and are even more challenging for high level designs. Because of complexity of today's designs, design at ESL (electronic system level) using transaction level modeling (TLM) has become a focal point of today's system level designers. However, there are no standard test methods or conventions proposed for this level of abstraction. Built-In Self-Test is a conventional DFT method, well defined in gate level and RT level. In this work by inspiration from the standard RTL BIST architectures, and finding similarities in TLM-2 designs and the RTL designs being tested by standard RTL BISTs, a number of TLM-2 BIST architectures are proposed. The overhead of inserting these BISTs in the original design is calculated. Nastaran Nemati, Zainalabedin Navabi |
Asian Test Symposium | 2 |
| 2010 | Test Pattern Selection and Compaction for Sequential Circuits in an HDL EnvironmentabstractIn this paper we are revisiting the issue of sequential circuit test generation, and use a selective random pattern test generation method implemented in an HDL environment. The method uses a statistical expectation graph and states of the sequential circuit for selecting the appropriate test vectors to achieve better fault coverage and a more compact test set. To further reduce the size of the generated test set, a static compaction method, which is also implemented in an HDL environment, is used after the test generation process. The experimental results show that selecting good test patterns among random test patterns, not only can be implemented dynamically in an HDL design environment, but also results in a better fault coverage and shorter test pattern length in comparison with some traditional deterministic methods. In addition, it will be shown that static test set compaction methods can considerably reduce the test length of test patterns for sequential designs obtained by our proposed method. M. H. Haghbayan, Sara Karamati, Fatemeh Javaheri, Zainalabedin Navabi |
Asian Test Symposium | 4 |
| 2010 | A reconfigurable online BIST for combinational hardware using digital neural networksabstractOnline testing, one of the most challenging issues in design for test domain, is intended for inspection of digital systems behavior during their working period. This paper presents a novel approach for simultaneous online testing of several combinational circuits using a reconfigurable neural network implemented along the original hardware. Automatic generation of the neural network to model the behavior of each design is proposed as well as required techniques to obtain optimum configuration for its hardware realization. Advantages and shortcomings of this approach in terms of area overhead, fault latency and reliability are discussed as well. S. Behdad Hosseini, Ali Shahabi, Hasan Sohofi, Zainalabedin Navabi |
ETS | 4 |
| 2010 | A partitioning approach to improve reconfigurable neuron-inspired online BISTabstractTwo of the most challenging issues in online testing are deriving a general tester scheme for various circuits and reducing the area overhead. This paper presents a novel reconfigurable online tester using artificial neural networks to test combinational hardware. Our proposed BIST architecture has the capability of testing a number of arbitrary sub-modules of a big design simultaneously by time-multiplexing between them. Output partitioning method is proposed as a powerful technique to reduce neural network training time and the tester area overhead. Our experimental results show that after proper partitioning the average area overhead is reduced by 16% in data-path and 33% in memory area. Also average fault detection latency has been improved by 14%. Ali Shahabi, S. Behdad Hosseini, Hasan Sohofi, Zainalabedin Navabi |
IOLTS | 4 |
| 2010 | Using context based methods for test data compressionabstractThis paper proposes a new test data compression method based on a context based binary arithmetic coding. The proposed method is suitable for fully specified test vector, benefiting from elimination of the relaxation step used in former methods. This method is evaluated using ISCAS89 full scan circuits. Sara Karamati, Zainalabedin Navabi |
ITC | 2 |
| 2010 | EDXY - A low cost congestion-aware routing algorithm for network-on-chips
Pejman Lotfi-Kamran, Amir-Mohammad Rahmani, Masoud Daneshtalab, Ali Afzali-Kusha, Zainalabedin Navabi |
J. Syst. Archit. | 5 |
| 2010 | Real-time embedded emotional controller
Mohammad Reza Jamali, Masoud Dehyadegari, Arash Arami, Caro Lucas, Zainalabedin Navabi |
Neural Comput. Appl. | 5 |
| 2009 | Emotion on FPGA: Model driven approach
Mohammad Reza Jamali, Arash Arami, Masoud Dehyadegari, Caro Lucas, Zainalabedin Navabi |
Expert Syst. Appl. | 5 |
| 2008 | BARP-A Dynamic Routing Protocol for Balanced Distribution of Traffic in NoCsabstractA novel routing algorithm, named balanced adaptive routing protocol (BARP), is proposed for NoCs to provide adaptive routing and ensure deadlock-free and livelock-free routing at the same time. By evenly distributing input packets of a router among all its shortest path output ports, a novel adaptive routing protocol for avoiding congestion condition emerges. It is observed that BARP can achieve better performance compared to static XY routing, odd- even routing and dynamic XY routing. Pejman Lotfi-Kamran, Masoud Daneshtalab, Caro Lucas, Zainalabedin Navabi |
DATE | 4 |
| 2008 | A Novel GA-Based High-Level Synthesis Technique to Enhance RT-Level Concurrent TestingabstractThis paper presents an efficient high-level synthesis (HLS) approach to improve RT-level concurrent testing. The proposed method used for both fault detection and fault location. At first the available resources are used in their dead intervals to test active resources for fault detection, and then some changes are applied to the RT-level controller to locate the faults. The fault detection step is based on a genetic algorithm (GA) search technique. This genetic algorithm is applied to the design after high level synthesis process to explore the test map. The proposed method has been evaluated based on dependability enhancement and area/latency overhead imposed to different benchmarks after applying our algorithm. The dependability has been considered in terms of fault coverage. The experimental result shows that applying our algorithm, the associated area overhead and performance penalty are negligible while the online fault coverage improvement is considerable. Naghmeh Karimi, Soheil Aminzadeh, Saeed Safari, Zainalabedin Navabi |
IOLTS | 4 |
| 2008 | Reliability in Application Specific Mesh-Based NoC ArchitecturesabstractNetworks on chips (NoCs) provide a mechanism for handling complex communications in the next generation of integrated circuits. At the same time, lower yield in nano-technology, makes self repair communication channels a necessity in design of digital systems. This paper proposes a reliable NoC architecture based on specific application mapped onto an NoC. This architecture is capable of recovering from permanent switch failures via replacing them by neighboring switches. This method has hardware and power consumption overhead, but significantly improves reliability and has a very little effect on the performance of the system. We suggest a reliability analysis method based on the combinatorial reliability models and use it to evaluate our proposed fault-tolerant NoC architecture. Fatemeh Refan, Homa Alemzadeh, Saeed Safari, Paolo Prinetto, Zainalabedin Navabi |
IOLTS | 5 |
| 2008 | NoC Reconfiguration for Utilizing the Largest Fault-free Connected Sub-structureabstractThis paper proposes an offline test strategy for finding the largest fault-free connected sub-structure of a mesh-based NoC. Faulty switch ports are found by flooding the NoC with test packets. Then, NoC routers are reconfigured according to the degraded NoC structure to route incoming packets. Armin Alaghi, Mahshid Sedghi, Naghmeh Karimi, Zainalabedin Navabi |
ITC | 4 |
| 2008 | "Plug & Test" at System Level via Testable TLM PrimitivesabstractWith the evolution of Electronic System Level (ESL) design methodologies, we are experiencing an extensive use of Transaction-Level Modeling (TLM). TLM is a high-level approach to modeling digital systems where details of the communication among modules are separated from the those of the implementation of functional units. This paper represents a first step toward the automatic insertion of testing capabilities at the transaction level by definition of testable TLM primitives. The use of testable TLM primitives should help designers to easily get testable transaction level descriptions implementing what we call a "Plug & Test" design methodology. The proposed approach is intended to work both with hardware and software implementations. In particular, in this paper we will focus on the design of a testable FIFO communication channel to show how designers are given the freedom of trading-off complexity, testability levels, and cost. Homa Alemzadeh, Stefano Di Carlo, Fatemeh Refan, Paolo Prinetto, Zainalabedin Navabi |
ITC | 5 |
| 2008 | A Selective Trigger Scan Architecture for VLSI TestingabstractTime, power, and data volume are among some of the most challenging issues for testing system-on-chip (SoC) and have not been fully resolved, even if a scan-based technique is employed. A novel architecture, referred to the selective trigger scan architecture, is introduced in this paper to address these issues. This architecture reduces switching activity in the circuit-under-test (CUT) and increases the clock frequency of the scanning process. An auxiliary chain is utilized in this architecture to avoid the large number of transitions to the CUT during the scan-in process, as well as enabling retention of the currently applied test vectors and applying only necessary changes to them. The auxiliary chain shifts in the difference between consecutive test vectors and only the required transitions (referred to as trigger data) are applied to the CUT. Power requirements are substantially reduced; moreover, DFT penalties are reduced because no additional multiplexer is utilized along the scan path. Data reformatting is applied in order to make the proposed architecture amenable to data compression, thus permitting a further reduction in test time. It also permits delay fault testing. Using ISCAS 85 and 89 benchmark circuits, the effectiveness of this architecture for improving SoC test measures (such as power, time, and data volume) is experimentally evaluated and confirmed. Mohammad Hosseinabady, Shervin Sharifi, Fabrizio Lombardi, Zainalabedin Navabi |
IEEE Trans. Computers | 4 |
| 2007 | Optimized Assignment Coverage Computation in Formal Verification of Digital SystemsabstractModel checking thoroughly verifies the design correctness with respect to a specification. When the verification process succeeds, we can only postulate the correctness of the design relative to the given specification. How far can we affirm the verified design implements all the behavior of the desired system? With this regard we need to estimate the completeness of the properties by using some coverage metrics. In this paper, we have proposed a new metric called assignment coverage and an optimized method to overcome the intensive computations required for the multiple transformations among the abstract layers in the verification tool. The proposed coverage computation method provides adequate information to complete the set of properties. Finally, we have applied the proposed metric to some verification benchmark to reveal the effectiveness of this metric in finding undetected coverage holes. Majid Nabi, Hamid Shojaei, Siamak Mohammadi, Zainalabedin Navabi |
ATS | 4 |
| 2007 | An HDL-Based Platform for High Level NoC Switch TestingabstractThis paper presents a non-scan method of NoC switch testing. The method requires addition of test-mode hardware for NoC switches and processing elements which is much less than what is required for most scan methods. Associated with our proposed test-mode of an NoC, we have developed a test environment based on high-level switch faults. The test environment applies test packets to the NoC-under-test in its test-mode and generates an NoC fault dictionary to be used for error detection of an NoC running in the test-mode. Proposed fault models and test strategy will be discussed in this paper. Mahshid Sedghi, Armin Alaghi, Elnaz Koopahi, Zainalabedin Navabi |
ATS | 4 |
| 2007 | Using the inter- and intra-switch regularity in NoC switch testingabstractThis paper proposes an efficient test methodology to test switches in a network-on-chip (NoC) architecture. A switch in a NoC consists of a number of ports and a router. Using the intra-switch regularity among ports of a switch and inter-switch regularity among routers of switches, the proposed method decreases the test application time and test data volume of NoC testing. Using a test source to generate test vectors and scan-based testing, this methodology broadcasts test vectors through the minimum spanning tree of the NoC and concurrently tests its switches. In addition, a possible fault is detected by comparing test results using inter- or intra- switch comparisons. The logic and memory parts of a switch are tested by appropriate memory and logic testing methods. Experimental results show less test application time and test power consumption, as compared with other methods in the literature Mohammad Hosseinabady, Atefe Dalirsani, Zainalabedin Navabi |
DATE | 3 |
| 2007 | On-Chip Verification of NoCs Using Assertion ProcessorsabstractNoC verification has become an increasingly difficult task due to the growing complexity of these systems. In this paper, we propose a methodology based on assertions for on-chip verification of NoCs. We have a local assertion processor (LAP) in each core to manage the outputs of assertions inside it. To route assertions' outputs toward this processor, we offer a boundary scan chain mechanism. Moreover, after detecting error, each LAP dispatches a packet called error packet to a global assertion processor (GAP) which receives error packets from all cores and performs necessary actions regarding to errors and their severities. Finally, in order to evaluate our method, we apply it on an NoC structure and show the experimental results. Mohammad Reza Kakoee, Mohammad Hossein Neishaburi, Masoud Daneshtalab, Saeed Safari, Zainalabedin Navabi |
DSD | 5 |
| 2007 | APDL: A Processor Description Language For Design Space Exploration of Embedded Processors
Nima Honarmand, Hasan Sohofi, Maghsoud Abbaspour, Zainalabedin Navabi |
FDL | 4 |
| 2007 | A Configurable Transaction Level Model of a Generic Interconnection Part of Embedded Systems Used in an ESL Design Library
Parisa Razaghi, Shahrzad Mirkhani, Zainalabedin Navabi |
FDL | 3 |
| 2007 | RT level reliability enhancement by constructing dynamic TMRSabstractThis paper presents a novel and efficient approach for reliability enhancement at the RT level. The reliability enhancement is performed by utilizing the available resources of a design in their dead intervals. Such resources are used for constructing dynamic TMR structures that can change per clock cycle. In this method all resources participate in constructing TMR structures at least once per a system input to output flow.To evaluate the proposed fault tolerance technique we consider dependability, and area/latency overhead imposed on a circuit by applying our method. In order to evaluate dependability, faults are injected into our test circuits before and after applying our algorithm and fault coverage is measured. Experimental results show that after applying our method, fault coverage is significantly reduced indicating that the reliability of designs is improved. Naghmeh Karimi, Shahrzad Mirkhani, Zainalabedin Navabi, Fabrizio Lombardi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | An Analytical Model for Reliability Evaluation of NoC ArchitecturesabstractThis paper proposes an analytical model to assess Reliability Factor of an NoC based System-on-Chip design. Reliability Factor is the probability that faults in the NoC infrastructure can be recovered without any effect on system functionality. The proposed method classifies switch faults of an NoC according to their impact on system functionality. Based on this classification, the contribution of each transient fault lowering the reliability of the NoC is calculated. This model can be used to decide which fault tolerant techniques cause more improvement on system reliability. Atefe Dalirsani, Mohammad Hosseinabady, Zainalabedin Navabi |
IOLTS | 3 |
| 2007 | Analysis of System-Failure Rate Caused by Soft-Errors using a UML-Based Systematic Methodology in an SoCabstractThis paper proposes an analytical method to assess the soft-error rate (SER) in the early stages of a System-on-Chip (SoC) platform-based design methodology. The proposed method gets an executable UML (Unified Modeling Language) model of the SoC and the raw softerror rate of different parts of the platform as its inputs. Soft-errors on the design are modeled by disturbances on the value of attributes in the classes of the UML model and disturbances on opcodes of software cores. The Dynamic behavior of each core is used to determine the propagation probability of each variable disturbance to the core outputs. Furthermore, the SER and the execution time of each core in the SoC and a Failure Modes and Effects Analysis (FMEA) that determines the severity of each failure mode in the SoC are used to compute the System-Failure Rate (SFR) of the SoC. Mohammad Hosseinabady, Mohammad Hossein Neishaburi, Zainalabedin Navabi, Alfredo Benso, Stefano Di Carlo, Paolo Prinetto, Giorgio Di Natale |
IOLTS | 3 |
| 2007 | A New Approach for Design and Verification of Transaction Level ModelsabstractTransaction level modeling allows exploring several SoC design architectures leading to better performance and easier verification of the final product. In this paper, we present an approach for design and verification of transaction level models. Verification is integrated as part of the design-flow. In the proposed method, we first model the design in UML. Then, we translate it into the reactive objects language, Rebeca (Marjan Sirjani et al., 2004), which is an actor-based language with formal foundation. A model in Rebeca is a set of concurrently executed reactive objects (called rebecs) interacted by asynchronous message passing. After mapping UML to Rebeca, Rebeca code will be translated into Promela which is a language for formal verification. Checking the correctness of the design is performed on-the-fly with the LTL properties using the SPIN model checker. Finally, we translate the verified design to SystemC and map the properties to a set of assertions that can be re-used to validate the design at lower levels through simulation. Mohammad Reza Kakoee, Hamid Shojaei, Hassan Ghasemzadeh 0001, Marjan Sirjani, Zainalabedin Navabi |
ISCAS | 5 |
| 2007 | Programmable Routing Tables for Degradable Torus-Based Networks on ChipsabstractThe decreasing manufacturing yield of integrated circuits, as a result of rising complexity and decreased feature size, and the emergence of NoC-based design techniques, has necessitated the search for network reconfiguration techniques for reusing NoCs with faulty communication hardware. In this paper, we propose a method to cope with the problem of faulty communication links in NoCs with torus topology. The method is based on the use of programmable routing tables in network switches. We investigate this technique for a conventional routing mechanism and our optimized routing mechanism. The conventional mechanism uses one entry in its routing table for every destination address while our proposed routing mechanism uses a fixed number of entries per table and routes based on the address value comparison of the current switch and the destination switch. Ali Shahabi, Nima Honarmand, Zainalabedin Navabi |
ISCAS | 3 |
| 2007 | High Level Synthesis of Degradable ASICs Using Virtual BindingabstractAs the complexity of the integrated circuits increases, they become more susceptible to manufacturing faults, decreasing the total process yield. Thus, it would be desirable to develop techniques for reusing faulty dies, even with a degraded performance. In this paper, a new method for high level synthesis of degradable ASICs is presented. Our technique introduces the concept of Virtual Binding. In this approach, the operations are bound to virtual components that are linked with actual non-faulty components using a set of configuration multiplexers and flip-flops embedded in the data-path. Using virtual components simplifies the synthesis algorithm and decreases the size of generated control unit. Virtual-to-physical mapping of the components will be established by programming the configuration flip-flops after diagnosing the faulty components. The experimental results show that the area and delay overhead of the resulting circuits have acceptable values compared to the original, non-degradable circuits. Nima Honarmand, Ali Shahabi, Hasan Sohofi, Maghsoud Abbaspour, Zainalabedin Navabi |
VTS | 5 |
| 2007 | A UML Based System Level Failure Rate Assessment Technique for SoC DesignsabstractThis paper proposes an analytical method to assess soft-error rate (SER) in the early stages of a system-on-chip (SoC) platform-based design methodology. The proposed method uses an executable UML model of the SoC for its input. Soft-errors on the design are modeled by disturbances on the value of attributes in the classes of the UML model and disturbances on opcodes of software cores. SER and execution time of each core in the SoC and a failure modes and effects analysis (FMEA) that determines the severity of each failure mode in the SoC are used to compute the system-failure rate (SFR) of the SoC Mohammad Hosseinabady, Mohammad Hossein Neishaburi, Pejman Lotfi-Kamran, Zainalabedin Navabi |
VTS | 4 |
| 2007 | Low test application time resource binding for behavioral synthesisabstractRecent advances in process technology have led to a rapid increase in the density of integrated circuits (ICs). Increased density and the need to test for new types of defects in nanometer technologies have resulted in a tremendous increase in test application time (TAT). This article presents a test synthesis method to reduce test application time for testing the datapath of a design. The test application time is reduced by applying a test-time-aware resource sharing algorithm on a scheduled control data flow graph (CDFG) of a design. Mohammad Hosseinabady, Pejman Lotfi-Kamran, Zainalabedin Navabi |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | NoC Hot Spot minimization Using AntNet Dynamic Routing AlgorithmabstractIn this paper, a routing model for minimizing hot spots in the network on chip (NOC) is presented. The model makes use of AntNet routing algorithm which is based on Ant colony. Using this algorithm, which we call AntNet routing algorithm, heavy packet traffics are distributed on the chip minimizing the occurrence of hot spots. To evaluate the efficiency of the scheme, the proposed algorithm was compared to the XY, Odd- Even, and DyAD routing models. The simulation results show that in realistic (Transpose) traffic as well as in heavy packet traffic, the proposed model has less average delay and peak power compared to the other routing models. In addition, the maximum temperature in the proposed algorithm is less than those of the other routing algorithms. Masoud Daneshtalab, Ashkan Sobhani, Ali Afzali-Kusha, Omid Fatemi, Zainalabedin Navabi |
ASAP | 5 |
| 2006 | An Optimum ORA BIST for Multiple Fault FPGA Look-Up Table TestingabstractThis paper presents BIST architecture for FPGA look-up table testing using a minimum number of logic elements for its ORA. The propagation of faults in the TPGs and CUTs is formulated so that the ORA can detect multiple faults by monitoring a single signal. At the cost of using more cells for the ORA, the granularity of error detection can be reduced to as low as one fault per five LUTs. The increase in the ORA overhead, and thus the untested FPGA areas, can be compensated by more configurations. We will show that 100% test coverage and a maximum granularity can be achieved simultaneously by a reasonable number of FPGA configurations Armin Alaghi, Mahnaz Sadoughi Yarandi, Zainalabedin Navabi |
ATS | 3 |
| 2006 | ESTA: An Efficient Method for Reliability Enhancement of RT-Level DesignsabstractThis paper proposes a novel and efficient method for RT level online testing. Our method makes every RT-level resource online-testable, and guarantees high single stuck-at fault detection (i.e., high reliability) with low area/latency overhead. This method uses available resources in their dead intervals (the intervals during which a resource is not being used) to test active resources. The area and/or latency overhead are due to concurrent operation of active and inactive resources. This method is evaluated by fault simulating several benchmark designs before and after applying the proposed algorithm. Experimental results show that after applying our method, online fault coverage is significantly improved Naghmeh Karimi, Shahrzad Mirkhani, Zainalabedin Navabi |
ATS | 3 |
| 2006 | A concurrent testing method for NoC switchesabstractThis paper proposes reuse of on-chip networks for testing switches in network on chips (NoCs). The proposed algorithm broadcasts test vectors of switches through the on-chip networks and detects faults by comparing output responses of switches with each other. This algorithm alleviates the need for: (1) external comparison of the output response of the circuit-under-test with the response of a fault free circuit stored on a tester (2) on-chip signature analysis (3) a dedicated test-bus to reach test vectors and collect their responses. Experimental results on a few test benches compare the proposed algorithm with traditional system on chip (SoC) test methods Mohammad Hosseinabady, Abbas BanaiyanMofrad, Mahdi Nazm Bojnordi, Zainalabedin Navabi |
DATE | 4 |
| 2006 | DCim++: a C++ library for object oriented hardware design and distributed simulationabstractDCim++ is a C++ library developed for object oriented hardware design, modeling and distributed simulation. DCim++ enables C++ to be used as an OO HDL, which supports concurrency in description, inheritance in design and distributedness in simulation. Design simulation results are obtained by running C++ programs on a network of workstations. The message passing interface (MPI) library has been used in the implementation of DCim++ as the basis of communications required for distributed simulation. In our simulation scheme, we have not considered any central management unit in order to defy performance degradation, instead only a coarse-grain synchronizer is used to keep the distributed components synchronized. This paper explores the structure of the DCim++ library and its mechanisms. The process a designer has to go through in order to design a system using DCim++ and conduct its distributed simulation leaving communication complications to DCim++, has also been presented. Finally, the results of our uniprocessor and distributed simulations for ISCAS benchmark circuits show high degrees of performance gains. Hadi Esmaeilzadeh, A. Moghimi, Eiman Ebrahimi, Caro Lucas, Zainalabedin Navabi, A. M. Fakhraie |
ISCAS | 5 |
| 2006 | Low-power and low-latency cluster topology for local traffic NoCsabstractIn this paper, we introduce a topology for network on chips that is named cluster-mesh (CM) topology. This architecture reduces dynamic and static power consumption in NoCs and can reduce latency of communications in low traffic or local traffic applications. With cluster-mesh topology, area reduction in routers is about 44% and in links we can save more than 50% in area too. The dynamic power in this architecture is reduced more than 20%. The idea of clustering may be applied to some other topologies such as Torus and Octagon Mohsen Saneei, Ali Afzali-Kusha, Zainalabedin Navabi |
ISCAS | 3 |
| 2006 | A Test Approach for Look-Up Table Based FPGAs
Ehsan Atoofian, Zainalabedin Navabi |
J. Comput. Sci. Technol. | 2 |
| 2005 | TED+: a data structure for microprocessor verificationabstractFormal verification of microprocessors requires a mechanism for efficient representation and manipulation of both arithmetic and random Boolean functions. Recently, a new canonical and graph-based representation called TED has been introduced for verification of digital systems. Although TED can be used effectively to represent arithmetic expressions at the word-level, it is not memory efficient in representing bit-level logic expressions. In this paper, we present modifications to TED to improve its ability for bit-level logic representation while maintaining its robustness in arithmetic word-level representation. It will be shown that for random Boolean expressions, the modified TED performs the same as BDD representation. Pejman Lotfi-Kamran, Mohammad Hosseinabady, Hamid Shojaei, Mehran Massoumi, Zainalabedin Navabi |
ASP-DAC | 5 |
| 2005 | ISC: Reconfigurable Scan-Cell Architecture for Low Power TestingabstractViolation of power constraints in the test mode may cause permanent failure in a circuit. Thus, Low power testing is essential for low power circuits. This paper proposes a reconfigurable scan-cell architecture that eliminates the propagation of unnecessary transitions during shift-in and shift-out. The proposed reconfigurable scanpath rearranges its latches to mask its outputs when a test-vector/test-result shifts in/out to/from. The rearrangement is performed without any need to extra latches or buffers. In fact, the native latches of a basic scan-path are reconfigured to keep the outputs of the scan-path (inputs of the combinational cloud) intact in the shifting phase. A few primitive gates are required for the rearrangement of the latches which means that the architecture has a low area overhead. The rearrangement implies that even and odd bits of test-vectors/test-results are interleaved in the shifting, and then, we called this reconfigurable architecture Interleaved Scan-Cell (ISC). The proposed scan-cell supports all required operations such as scan-in, scan-out, test-vector application, and test-result collection. The reconfigurable interleaved scan-path is inserted in a number of ISCAS benchmark circuits and the total area overhead and test power consumptions are presented. The results and comparisons show that using interleaved scancell architecture reduces test power dissipation while it has a low area overhead. Also, it is shown that the proposed scan-cell architecture adds a negligible delay to the propagation time of the scan-path registers and thus does not alter the clock frequency. Hadi Esmaeilzadeh, Saeed Shamshiri, Pooya Saeedi, Zainalabedin Navabi |
Asian Test Symposium | 4 |
| 2005 | Enhancing Fault Simulation Performance by Dynamic Fault ClusteringabstractFault simulation algorithms used for large designs propagate a list of faults instead of a single fault in each simulation. Concurrent (Ulrich and Baker, 1974) and deductive (Armstrong, 1972) fault simulation algorithms are two examples of this kind of algorithm. In this paper, we utilize an optimization concept, which can be added to fault list propagating algorithms. In this concept, faults can be grouped into several disjoint fault sets. All faults in a group affect every line of the circuit in a similar way. Fault clustering is performed dynamically, based on a particular test vector, during the fault simulation process. This method causes less memory fragmentation, since there are a limited number of fault groups in each simulation time. On the other hand, it reduces faulty circuit calculation in fault simulation process compared with the traditional fault simulation methods. In addition, the generality of this concept makes it useful for behavioral fault simulation methods as well as traditional gate-level ones. We have implemented this method in the VHDL environment and tested it on ISCAS'85 benchmarks. Experimental results show that in large circuits the performance is at least doubled by this technique Shahrzad Mirkhani, Zainalabedin Navabi |
Asian Test Symposium | 2 |
| 2005 | Sign bit reduction encoding for low power applicationsabstractThis paper proposes a low power technique, called SBR (Sign Bit Reduction) which may reduce the switching activity in multipliers as well as data buses. Utilizing the multipliers based on this scheme, the dynamic power consumption of some digital systems such as digital filters based on CMOS logic system can be reduced considerably compared to those based on 2's complement implementation. To verify the efficacy of the SBR, a 16-bit multiplier was implemented by this scheme. The results for voice data show an average of 29% to 35% switching reduction compared to the 2's complement implementation. For 16-bit random data, this scheme decreases the switching of 16-bit multipliers by an average of 21%. Finally, the application of the technique to a 16-bit data bus leads up to 14.5% switching reduction on average. Mohsen Saneei, Ali Afzali-Kusha, Zainalabedin Navabi |
DAC | 3 |
| 2005 | Simultaneous Reduction of Dynamic and Static Power in Scan StructuresabstractPower dissipation during test is a major challenge in testing integrated circuits. Dynamic power has been the dominant part of power dissipation in CMOS circuits, however, in future technologies the static portion of power dissipation will outreach the dynamic portion. This paper proposes an efficient technique to reduce both dynamic and static power dissipation in scan structures. Scan cell outputs which are not on the critical path(s) are multiplexed to fixed values during scan mode. These constant values and primary inputs are selected such that the transitions occurring on nonmultiplexed scan cells are suppressed and the leakage current during scan mode is decreased. A method for finding these vectors is also proposed. The effectiveness of this technique is proved by experiments performed on ISCAS89 benchmark circuits. Shervin Sharifi, Javid Jaffari, Mohammad Hosseinabady, Ali Afzali-Kusha, Zainalabedin Navabi |
DATE | 5 |
| 2005 | Combination of Assertion and HSAT Methods For Automated Test Vectors Generation
Mostafa Naderi, Zainalabedin Navabi |
FDL | 2 |
| 2005 | Instruction-level test methodology for CPU core self-testingabstractTIS is an instruction-level methodology for processor core self-testing that enhances instruction set of a CPU with test instructions. Since the functionality of test instructions is the same as the NOP instruction, NOP instructions can be replaced with test instructions. Online testing can be accomplished without any performance penalty. TIS tests different parts of the processor and detects stuck-at faults. This method can be employed in offline and online testing of single-cycle, multicycle and pipelined processors. But, TIS is more appropriate for online testing of pipelined architectures in which NOP instructions are frequently executed because of data, control and structural hazards. Running test instructions instead of these NOP instructions, TIS utilizes the time that is otherwise wasted by NOPs. In this article, two different implementations of TIS are presented. One implementation employs a dedicated hardware modules for test vector generation, while the other is a software-based approach that reads test vectors from memory. These two approaches are implemented on a pipelined processor core and their area overheads are compared. To demonstrate the appropriateness of the TIS test technique, several programs are executed and fault coverage results are presented. Saeed Shamshiri, Hadi Esmaeilzadeh, Zainalabedin Navabi |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2004 | Test Instruction Set (TIS) for High Level Self-Testing of CPU CoresabstractTIS (test instruction set) is an instruction level technique for CPU core self-testing. This method is based on enhancing a CPU instruction set with test instructions. TIS replaces the NOP instruction that is available in most processors with test instructions so that online testing can be done with no performance penalty. This method can be applied to both offline and online (concurrent) testing of all types of processors (single-cycle, multi-cycle and pipelined). TIS is appropriate for pipelined architectures in which one or many NOP instructions (or stalls) are inserted between instructions that are data or control dependent. We have implemented this test method on a pipelined CPU core and several test programs for this pipelined CPU are used to illustrate the method. Also fault coverage results are presented to demonstrate the effectiveness of the TIS test technique. Saeed Shamshiri, Hadi Esmaeilzadeh, Zainalabedin Navabi |
Asian Test Symposium | 3 |
| 2004 | Using RT Level Component Descriptions for Single Stuck-at Hierarchical Fault Simulation
Zainalabedin Navabi, Shahrzad Mirkhani, Meisam Lavasani, Fabrizio Lombardi |
J. Electron. Test. | 1 |
| 2004 | A Low-Cost At-Speed BIST Architecture for Embedded Processor and SRAM Cores
Mohammad H. Tehranipour, Sied Mehdi Fakhraie, Zainalabedin Navabi, M. R. Movahedin |
J. Electron. Test. | 3 |
| 2003 | A BIST Architecture for FPGA Look-Up Table Testing Reduces ReconfigurationsabstractThis paper describes a test architecture for minimum number of test configurations for test of FPGA (field programmable gate array) LUTs (look up tables). Our test architecture includes a TPG (test pattern generator) that is tested while it is generating test data for LE (logic elements) that form our CUT (circuit under test). This scheme eliminates the need for switching LEs between CUT, TPG and ORA (output response analyzer) and having to perform many reconfigurations of the FPGA. An external ORA locates faults of the FPGA under test. In addition to the LUTs, we are also presenting a scheme for testing other parts of the LEs. Compared with other methods, our method uses the least number of reconfigurations of an FPGA for its LUT testing. Ehsan Atoofian, Zainalabedin Navabi |
Asian Test Symposium | 2 |
| 2003 | The VPI-Based Combinational IP Core Module-Based Mixed Level Serial Fault Simulation and Test Generation MethodologyabstractIn this paper we are presenting a test methodology for performing module-bused mixed level fault simulation and test generation on System-on-Chip (SOC) combinational Intellectual Property (IP) cores for which both a pre-synthesis behavioral description and a post-synthesis netlist is available but in an analyzer output intermediate format not readable by core integraters. We use the Verilog Procedural Interface (VPI) to access and perform serial fault simulation on a pre-compiled core available as a mixed behavioral structural level design. We also use VPI to prepare a testbench environment for performing random pattern test generation. The simulation time results of applying this VPI-based test methodology on ISCAS85 Verilog benchmarks are also presented and compared to the flat (non-mixed level) version the proposed VPI-based environment. Pedram A. Riahi, Zainalabedin Navabi, Fabrizio Lombardi |
Asian Test Symposium | 2 |
| 2003 | A Low Power BIST Architecture for FPGA Look-Up Table Testing
Ehsan Atoofian, Zainalabedin Navabi |
VLSI-SOC | 2 |
| 2003 | Processor Testing Using an ADL Description and Genetic Algorithms
Elham Safi, Reihaneh Saberi, Zohreh Karimi, Zainalabedin Navabi |
VLSI-SOC | 4 |
| 2003 | Selective Trigger Scan Architecture for Reducing Power, Time and Data Volume in SoC Testing
Shervin Sharifi, Mohammad Hosseinabady, Zainalabedin Navabi |
VLSI-SOC | 3 |
| 2002 | Hierarchical Fault Simulation Using Behavioral and Gate Level Hardware ModelsabstractThis paper presents a fault simulation environment that takes advantage of available models at the behavioral and gate levels of abstraction. The simulation takes place in VHDL and for fault simulation, special VHDL models are written that are capable of propagating circuit faults. Behavioral VHDL models propagate fault effects that appear on their input ports; in addition to this, gate level VHDL models are capable of injecting faults on their output lines. The fault simulation environment assumes the existence of the gate level and behavioral models for every component, and uses the appropriate model depending on whether a fault belongs to it or another component. A wrapper simulation model that encloses both models of a component switches automatically between the models. The wrapper takes care of feedback in the sequential circuits by always selecting the gate level of a component for propagating its own faults. This environment fits well with the hardware description language settings in which pre-synthesis behavioral models, post-synthesis gate-level models and a mixed simulation environment are available. The paper shows a mathematical analysis illustrating the performance improvement of this method over the traditional gate-level fault simulation. Shahrzad Mirkhani, Meisam Lavasani, Zainalabedin Navabi |
Asian Test Symposium | 3 |
| 2001 | Fault Simulation for VHDL Based Test Bench and BIST EvaluationabstractA VHDL based Fault simulation procedure for test bench and test hardware evaluation has been developed. This work is aimed to utilize features of VHDL for more efficient fault simulation. Information about fault detection can be obtained in this environment using fault simulation method and guidelines presented in this report. This environment consists of automated steps, which will lead to fault simulation. Information such as fault coverage, efficiency of test patterns and capability of test hardware to detect faults, can be extracted. Using this environment, one can evaluate test benches and order test vectors or configure BIST (Built-In Self Test) architectures. Hamed Farshbaf, Mina Zolfy, Shahrzad Mirkhani, Zainalabedin Navabi |
Asian Test Symposium | 4 |
| 2001 | Adaptation of an event-driven simulation environment to sequentially propagated concurrent fault simulationabstractSummary form only given. A new fault simulation method is presented here. The method relies on the simulation cycle timing of event-driven simulators (delta delays in VHDL). This timing is used for propagation of faulty values in faulty sections of a circuit. This method is based on concurrent fault simulation and is implemented in VHDL. VHDL gate models that are capable of propagating faults in fault queues perform this fault simulation. The gate models process their fault queues and propagate them in delta time units. In these models, gates with faulty input values are expanded in delta time to evaluate faulty output values and propagate them to other sections of the circuit. Using ISCAS benchmarks, a performance improvement of up to 500X over serial fault simulation has been obtained. This work is useful for fault simulation of post-synthesis VHDL outputs. Mina Zolfy, Shahrzad Mirkhani, Zainalabedin Navabi |
DATE | 3 |
| 1992 | A high-level language for design and modeling of hardware
Zainalabedin Navabi |
J. Syst. Softw. | 1 |
| 1986 | Compiling an RT level hardware description language into layout of NMOS cells
Zainalabedin Navabi, Kia Doroudi |
Microprocessing and Microprogramming | 1 |
| 1985 | Generating gate level two phase dynamic MOS logic from AHPL
Zainalabedin Navabi |
Microprocessing and Microprogramming | 1 |
| 1984 | Hardware Compilation from an RTL to a Storage Logic Array TargetabstractThis paper treats the automatic translation of register transfer level (RTL) descriptions of digital systems to VLSI realization. The target technology is the storage logic array or SLA. The approach is aimed at applications where the emphasis is on reducing engineering effort and design turnaround time rather than maximizing chip area utilization. The paper develops a mapping between the register transfer language, AHPL, and the SLA. It is shown that each primitive explicitly appearing in an AHPL description can be mapped into an area of real estate in an SLA realization. A detailed development of some of the algorithms is presented. The entire process has been successfully implemented and applied to a set of examples. This is accomplished by developing a final stage for an already existing three-stage multi-application compiler for AHPL. Layout and routing are shown to be a single optimization process if the hardware target is an SLA. Fredrick J. Hill, Zainalabedin Navabi, Chen H. Chiang, Duan-Ping Chen, Manzer Masud |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1981 | Structure Specification with a Procedural Hardware Description LanguageabstractDescribes the extension and formalization of the hardware description language AHPI to form AHPL III. This language provides for nesting AHPL descriptions within descriptions. It incorporates a general index extension mechanism which permits the efficient representation of sets of duplicate descriptions of any complexity. Three types of structures, procedural structures, functional registers, and combinational logic units are permitted. Procedural structures may be primitive or nonprimitive. All but primitive procedural structures share a common syntax. Nesting, declaration, and invocation rules for these distinct structures are specified in a semantics table. Fredrick J. Hill, R. E. Swanson, Manzer Masud, Zainalabedin Navabi |
IEEE Trans. Computers | 4 |
| 1979 | Efficient simulation of AHPL
Zainalabedin Navabi, Fredrick J. Hill |
DAC | 1 |