Seyed Ghassem Miremadi

dblp:80/4361 · DBLP profile ↗
← Back
80ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0003-4347-4380ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 53 · 2 first-authorSecurity and privacy · 24Software engineering, systems software and programming languages · 19Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Memory systems · 45% Hardware reliability and fault tolerance · 32% Storage systems · 7%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.732019
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
An Efficient Protection Technique for Last Level STT-RAM Caches in Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2017
Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM Caches · IEEE Trans. Computers 2016
Memory systems › cache management
cache replacement
0.522019
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
RAW-Tag: Replicating in Altered Cache Ways for Correcting Multiple-Bit Errors in Tag Array · IEEE Trans. Dependable Secur. Comput. 2019
Hardware reliability and fault tolerance
soft errors
0.522019
RAW-Tag: Replicating in Altered Cache Ways for Correcting Multiple-Bit Errors in Tag Array · IEEE Trans. Dependable Secur. Comput. 2019
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Hardware reliability and fault tolerance › memory reliability
cache reliability
0.412019
RAW-Tag: Replicating in Altered Cache Ways for Correcting Multiple-Bit Errors in Tag Array · IEEE Trans. Dependable Secur. Comput. 2019
Hardware reliability and fault tolerance › error correction
multi-bit error correction
0.412019
RAW-Tag: Replicating in Altered Cache Ways for Correcting Multiple-Bit Errors in Tag Array · IEEE Trans. Dependable Secur. Comput. 2019
Storage systems › storage reliability › durability
retention failure
0.412019
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Memory systems › non-volatile memory › magnetic random access memory › STT-MRAM
STT-MRAM cache
0.412019
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Hardware reliability and fault tolerance
error correction
0.312017
An Efficient Protection Technique for Last Level STT-RAM Caches in Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2017
Electronic design automation › physical design › routing › message routing
network-on-chip routing
0.312017
LAXY: A Location-Based Aging-Resilient Xy-Yx Routing Algorithm for Network on Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Interconnection networks and networks-on-chip › routing algorithms › interconnection routing
oblivious routing
0.312017
LAXY: A Location-Based Aging-Resilient Xy-Yx Routing Algorithm for Network on Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Memory systems › cache
STT-RAM cache
0.312017
An Efficient Protection Technique for Last Level STT-RAM Caches in Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems
non-volatile memory
0.212016
Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM Caches · IEEE Trans. Computers 2016
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.212016
Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM Caches · IEEE Trans. Computers 2016
Hardware reliability and fault tolerance
aging
0.112017
LAXY: A Location-Based Aging-Resilient Xy-Yx Routing Algorithm for Network on Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Processor architecture and microarchitecture
chip multiprocessor
0.112017
An Efficient Protection Technique for Last Level STT-RAM Caches in Multi-Core Processors · IEEE Trans. Parallel Distributed Syst. 2017
Hardware reliability and fault tolerance › aging › transistor aging
negative bias temperature instability
0.112017
LAXY: A Location-Based Aging-Resilient Xy-Yx Routing Algorithm for Network on Chip · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Hardware reliability and fault tolerance › error correction
error-correcting codes
0.112016
Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM Caches · IEEE Trans. Computers 2016
Memory systems › non-volatile memory › write reliability
write endurance
0.112016
Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM Caches · IEEE Trans. Computers 2016
Electronic design automation › hardware verification and test › functional verification
emulation
0.012003
Switch-level emulation · DAC 2003
Reconfigurable computing and FPGAs
FPGA-based emulation
0.012003
Switch-level emulation · DAC 2003
Electronic design automation
hardware verification and test
0.012003
Switch-level emulation · DAC 2003

Methods — techniques the papers use, named apart from their topics

trace-driven simulation · 0.4simulation · 0.4replication · 0.4parity check · 0.4error-correcting codes · 0.3asymmetry-aware protection · 0.3bit relocation · 0.2SECDED · 0.2Floating-ECC · 0.2switch-level modeling · 0.0
YearPublicationVenuePosition
2019 TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches
abstract
As technology process node scales down, on-chip SRAM caches lose their efficiency because of their low scalability, high leakage power, and increasing rate of soft errors. Among emerging memory technologies,$Spin$-$Transfer\; Torque\; Magnetic\; RAM$(STT-MRAM) is known as the most promising replacement for SRAM-based cache memories. The main advantages of STT-MRAM are its non-volatility, near-zero leakage power, higher density, soft-error immunity, and higher scalability. Despite these advantages, high error rate in STT-MRAM cells due to$retention\; failure$,$write\; failure$, and$read\; disturbance$threatens the reliability of cache memories built upon STT-MRAM technology. The error rate is significantly increased in higher temperature, which further affects the reliability of STT-MRAM-based cache memories. The major source of heat generation and temperature increase in STT-MRAM cache memories is write operations, which are managed by cache$replacement\; policy$. To the best of our knowledge, none of previous studies have attempted to mitigate heat generation and high temperature of STT-MRAM cache memories using replacement policy. In this paper, we first analyze the cache behavior in conventional$Least$-$Recently\; Used$(LRU) replacement policy and demonstrate that the majority of consecutive write operations (more than 66 percent) are committed to adjacent cache blocks. These adjacent write operations cause accumulated heat and increased temperature, which significantly increase the cache error rate. To eliminate heat accumulation and the adjacency of consecutive writes, we propose a cache replacement policy, named$Thermal$-$Aware\; Least$-$Recently\; Written$(TA-LRW), to smoothly distribute the generated heat by conducting consecutive write operations in distant cache blocks. TA-LRW guarantees the distance of at least three blocks for each two consecutive write operations in an 8-way associative cache. This distant write scheme reduces the temperature-induced error rate by 94.8 percent, on average, compared with the conventional LRU policy, which results in 6.9x reduction in cache error rate. The implementation cost and complexity of TA-LRW is as low as$First$-$In,\; First$-$Out$(FIFO) policy while providing a near-LRU performance, having the advantages of both replacement policies. The significantly reduced error rate is achieved by imposing only 2.3 percent performance overhead compared with the LRU policy.
Elham Cheshmikhani, Hamed Farbeh, Seyed Ghassem Miremadi, Hossein Asadi 0001
IEEE Trans. Computers3
2019 RAW-Tag: Replicating in Altered Cache Ways for Correcting Multiple-Bit Errors in Tag Array
abstract
Tag array in on-chip caches is one of the most vulnerable components to radiation-induced soft errors. Protecting the tag array in some processors is limited to error detection using the parity check, since the overheads of error correcting codes are not affordable in this component. State-of-the-art tag protection schemes combine the parity check with replication to provide error correction capability. Classifying these replication-based schemes into partial-replication and full-replication, the former offers a low overhead protection in which a large fraction of detectable errors remain uncorrectable, whereas the latter imposes a significant overhead to correct all of the errors. This paper proposes a low overhead full-replication scheme, so called Replicating in Altered Ways of Tag (RAW-Tag), to correct all detectable errors. RAW-Tag manipulates the cache replacement algorithm and keeps track of the incoming/evicting cache lines to not only provide a replica for all tags, but also eliminate the simultaneous susceptibility of both a tag and its replica to a single Multiple-Bit Upset (MBU). The simulation results show that RAW-Tag imposes no performance overhead and increases the energy consumption of L1 and L2 caches by only 6.6 and 0.3 percent, respectively, as compared with the baseline.
Hamed Farbeh, Fereshte Mozafari, Masoume Zabihi, Seyed Ghassem Miremadi
IEEE Trans. Dependable Secur. Comput.4
2017 WIPE: Wearout Informed Pattern Elimination to Improve the Endurance of NVM-based Caches
abstract
With the recent development in Non-Volatile Memory (NVM) technologies, several studies have suggested using them as an alternative to SRAMs in on-chip caches. However, limited endurance of NVMs is a major challenge when employed in the caches. This paper proposes a data manipulation technique, so-called Wearout Informed Pattern Elimination (WIPE), to improve the endurance of NVM-based caches by reducing the activity of frequent data patterns. Simulation results show that WIPE improves the endurance by up to 93% with negligible overheads.
Sina Asadi, Amir Mahdi Hosseini Monazzah, Hamed Farbeh, Seyed Ghassem Miremadi
ASP-DAC4
2017 QuARK: Quality-configurable approximate STT-MRAM cache by fine-grained tuning of reliability-energy knobs
abstract
Emerging STT-MRAM memories are promising alternatives for SRAM memories to tackle their low density and high static power consumption, but impose high energy consumption for reliable read/write operations. However, absolute data integrity is not required for many approximate computing applications, allowing energy savings with minimal quality loss. This paper proposes QuARK, a hardware/software approach for trading reliability of STT-MRAM caches for energy savings in the on-chip memory hierarchy of multi- and many-core systems running approximate applications. In contrast to SRAM-based cache-way-level actuators, QuARK utilizes fine-grained cache-line-level actuation knobs with different levels of reliability for individual read and write accesses which are unique to STT-MRAM and suitable for systems running multiple applications with mixed accuracy sensitivity, thus avoiding interapplication actuation interference. Our experimental results with a set of recognition, mining and synthesis (RMS) benchmarks demonstrate up to 40% energy savings over a fully-protected STT-MRAM cache, with negligible loss in the quality of the generated outputs.
Amir Mahdi Hosseini Monazzah, Majid Namaki-Shoushtari, Seyed Ghassem Miremadi, Amir-Mohammad Rahmani, Nikil Dutt
ISLPED3
2017 ANMR: Aging-aware adaptive N-modular redundancy for homogeneous multicore embedded processors
Farshad Baharvand, Seyed Ghassem Miremadi
J. Parallel Distributed Comput.2
2017 LAXY: A Location-Based Aging-Resilient Xy-Yx Routing Algorithm for Network on Chip
abstract
Network on chip (NoC) is a scalable interconnection architecture for ever increasing communication demand between processing cores. However, in nanoscale technology size, NoC lifetime is limited due to aging processes of negative bias temperature instability, hot carrier injection, and electromigration. Usually, because of unbalanced utilization of NoC resources, some parts of the network experience more thermal stress and duty cycle in comparison with other parts, which may accelerate chip failure. To slow down the aging rate of NoC, this paper proposes an oblivious routing algorithm called location-based aging-resilient Xy-Yx (LAXY) to distribute packet flow over entire network. LAXY is based on the fact that dimension-ordered routing algorithms imposes the highest traffic load on the central nodes in mesh topologies. To balance the traffic over the network, certain routers at the east and the west of NoC, with dimension-order XY routing, statically are configured as YX. Various configurations have been explored for LAXY and the simulations show a specific configuration, called Fishtail, increases mean time to failure of the routers and interconnects by about 42% and 56%, respectively. Moreover, by balancing the load over the network, LAXY improves overall packet latency by about 7% in average, with negligible area overhead.
Nezam Rohbani, Zahra Shirmohammadi, Maryam Zare, Seyed Ghassem Miremadi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2017 An Efficient Protection Technique for Last Level STT-RAM Caches in Multi-Core Processors
abstract
Due to serious problems of SRAM-based caches in nano-scale technologies, researchers seek for new alternatives. Among the existing options, STT-RAM seems to be the most promising alternative. With high density and negligible leakage power, STT-RAMs open a new doorto respond to future demands of multi-core systems, i.e., large on-chip caches. However, several problems in STT-RAMs should be overcome to make it applicable in on-chip caches. High probability of write error due to stochastic switching is a major problem in STT-RAMs. Conventional Error-Correcting Codes (ECCs) impose significant area and energy consumption overheads to protect STT-RAM caches. These overheads in multi-core processors with large last-level caches are not affordable. In this paper, we propose Asymmetry-Aware Protection Technique (A2PT) to efficiently protect the STT-RAM caches. A2PT benefits from error rate asymmetry of STT-RAM write operations to provide the required level of cache protection with significantly lower overheads. Compared with the conventional ECC configuration, the evaluation results show that A2PT reduces the area and energy consumption overheads by about 42 and 50 percent, respectively, while providing the same level of protection. Moreover, A2PT decreases the number of bit switching in write operations by 28 percent, which leads to about 25 percent saving in write energy consumption.
Zahra Azad, Hamed Farbeh, Amir Mahdi Hosseini Monazzah, Seyed Ghassem Miremadi
IEEE Trans. Parallel Distributed Syst.4
2017 A Low Area Overhead NBTI/PBTI Sensor for SRAM Memories
abstract
Bias temperature instability (BTI) is known as one serious reliability concern in nanoscale technologies. BTI gradually increases the absolute value of threshold voltage (Vth) of MOS transistors. The main consequence of Vth shift of the SRAM cell transistors is the static noise margin (SNM) degradation. The SNM degradation of SRAM cells results in bit-flip occurrences due to transient faults and should be monitored accurately. This paper proposes a sensor called write current-based BTI sensor (WCBS) to assess the BTI-aging state of SRAM cells. The WCBS measures BTI-induced SNM degradation of SRAM cells by monitoring the maximum write current shifts due to BTI. The observations show that the maximum current consumption during write operation is an effective identifier to measure Vth and SNM shifts. The granularity of BTI assessment of one cell up to a row of memory can be achieved by writing special bit patterns on the memory block during the test. We evaluated the sensor through SPICE-level simulations in 32-nm technology size. The precision of WCBS is about ±1.25 mV (±3.2% error). One sensor is enough for the entire SRAM memory block with negligible area/power overhead; less than 1%. The effects of process variation and temperature changes on WCBS are investigated in detail.
Nezam Rohbani, Seyed Ghassem Miremadi
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Bias Temperature Instability Mitigation via Adaptive Cache Size Management
abstract
Bias temperature instability (BTI) is one of the major CMOS reliability issues in nanoscales. The main impact of BTI on SRAM memory cells is the degradation of the static noise margin (SNM), which leads to a higher susceptibility to failures. A variety of techniques for mitigating the impact of BTI on caches have been proposed at architecture level. However, their considerable overheads limit the application of such techniques. Recent studies showed that the utilization of the cache capacity widely varies from one workload to another and even within a workload. When cache utilization is low, for the majority of the cells, the same value is stored for a very long period, which significantly degrades SNM due to BTI. In this paper, we propose a technique to dynamically adjust the cache size according to the running workload cache requirement by monitoring the cache miss rate. The unused cache capacity is power gated to increase the energy efficiency and mitigate aging of the entire cache. The experimental results show that the proposed technique reduces hold and read SNM degradation by up to 48.1% and 33.3%, respectively, at the cost of 2.0% performance penalty.
Nezam Rohbani, Mojtaba Ebrahimi, Seyed Ghassem Miremadi, Mehdi Baradaran Tahoori
IEEE Trans. Very Large Scale Integr. Syst.3
2016 ACM: Accurate crosstalk modeling to predict channel delay in Network-on-Chips
abstract
The severity of timing delay in the communication channels of Network on Chip (NoC) depends on the transition patterns appearing on the wires. An analytical model can estimate the timing delay in NoC channels in the presence of crosstalk faults. However, recently proposed analytical model does not have enough accuracy and is based on 3-wire delay model. In this paper, an Accurate Crosstalk Model (ACM) based on 5-wire delay model is proposed to estimate the delay of communication channels in the presence of crosstalk faults. ACM is more accurate due to considering more wires in the delay model and also considering the overlaps between locations of transition patterns.
Zeinab Mahdavi, Zahra Shirmohammadi, Seyed Ghassem Miremadi
IOLTS3
2016 Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM Caches
abstract
Spin-Transfer Torque RAM (STT-RAM) is a promising alternative to SRAM for implementing on-chip L2 and L3 caches. One of the most critical challenges in STT-RAM is reliability due to limited write endurance, which results in insufficient lifetime, as well as various types of errors. Previous studies have focused on either presenting various cache architectures/management techniques to improve the lifetime of STT-RAM caches or utilizing different Error Correcting Codes (ECCs) to protect against the permanent and transient errors. However, there is no quantitative analysis in the literature to determine the impact of ECCs on the lifetime of the STT-RAM caches. This paper formulates this impact and demonstrates that ECCs shorten the lifetime of STT-RAM cache lines by more than 50 percent due to ECCs high write activity. Then, we propose the Floating-ECC architecture for increasing the lifetime of the STT-RAM caches. The main idea is to evenly distribute the ECC write activity over all bits of cache lines by periodically relocating the ECC bits inside the cache lines. The simulation results for the most conventional ECC scheme, i.e., interleaved Single Error Correction-Double Error Detection (SEC-DED), show that Floating-ECC increases the lifetime of L2 and L3 caches by more than 318 percent and 254 percent, respectively.
Hamed Farbeh, Hyeonggyu Kim, Seyed Ghassem Miremadi, Soontae Kim
IEEE Trans. Computers3
2016 A Cache-Assisted Scratchpad Memory for Multiple-Bit-Error Correction
abstract
Scratchpad memory (SPM) is widely used in modern embedded processors to overcome the limitations of cache memory. The high vulnerability of SPM to soft errors, however, limits its usage in safety-critical applications. This paper proposes an efficient fault-tolerant scheme, called cache-assisted duplicated SPM (CADS), to protect SPM against soft errors. The main aim of CADS is to utilize cache memory to provide a replica for SPM lines. Using cache memory, CADS is able to guarantee a full duplication of all SPM lines. We also further enhance the proposed scheme by presenting buffered CADS (BCADS) that significantly improves the CADS energy efficiency. BCADS is compared with two well-known duplication schemes as well as single-error correction scheme. The comparison results reveal that: 1) BCADS imposes a 13.6% less energy-delay product (EDP) overhead than the duplication schemes and it does not require to modify the SPM manager and target application and 2) in comparison with the conventional single-error correction double-error detection (SEC-DED) scheme, BCADS provides a significantly higher error correction capability by correcting up to 4-b burst errors using a low-cost 4-b interleaved parity code. Moreover, the area overhead for error correction and the performance overhead of BCADS are negligible (less than 1%), whereas the area and performance overheads are 21.9% and 6.1% for SEC-DED, respectively. Furthermore, BCADS imposes about a 10.7% lower EDP overhead compared with the SEC-DED scheme.
Hamed Farbeh, Nooshin Sadat Mirzadeh, Nahid Farhady Ghalaty, Seyed Ghassem Miremadi, Mahdi Fazeli, Hossein Asadi 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Addressing NoC Reliability Through an Efficient Fibonacci-Based Crosstalk Avoidance Codec Design
Zahra Shirmohammadi, Seyed Ghassem Miremadi
ICA3PP (3)2
2015 A fault-tolerant and energy-aware mechanism for cluster-based routing algorithm of WSNs
abstract
Wireless Sensor Networks (WSNs) are prone to faults due to battery depletion of nodes. A node failure can disturb routing as it plays a key role in transferring sensed data to the end users. This paper presents a Fault-Tolerant and Energy-Aware Mechanism (FTEAM), which prolongs the lifetime of WSNs. This mechanism can be applied to cluster-based WSN protocols. The main idea behind the FTEAM is to identify overlapped nodes and configure the most powerful ones to the sleep mode to save their energy for the purpose of replacing a failed Cluster Head (CH) with them. FTEAM not only provides fault tolerant sensor nodes, but also tackles the problem of emerging dead area in the network. Our experimental results and simulations show that FTEAM outperforms conventional protocols in terms of network lifetime and energy consumption. In addition, an analytical evaluation using the Markov model is performed to determine the reliability of the FTEAM.
Maryam Hezaveh, Zahra Shirmohammadi, Nezam Rohbani, Seyed Ghassem Miremadi
IM4
2015 ARMOR: Adaptive Reliability Management by On-the-Fly Redundancy in Multicore Embedded Processors
abstract
Multicore processors are expected to play a key role in the future of critical embedded systems such as automotive and avionics. This is primarily due to their advantages offered to the embedded systems such as increase in processing capability and reduction in power, size, and cost. However, reliability is one of the most compelling factors for critical applications. Improvement of the reliability may adversely affect the parameters of an embedded system such as power and resource utilization. This paper proposes an adaptive yet proactive method by which, in response to the requests from part of the application software, the reliability of the system is enhanced using physical redundancy. Having considered the resource and power limits in an embedded system, the method assigns some cores to perform the targeted critical tasks in a resilient form. It takes advantage of generated slacks by on-the-fly managing frequency of the cores without violating hard deadlines of the real-time system. Analytical results show that this method is up to 12 times more efficient than standby sparing in a multicore processor considering combination of energy consumption and resource utilization factors.
Farshad Baharvand, Seyed Ghassem Miremadi
PRDC2
2015 In-Scratchpad Memory Replication: Protecting Scratchpad Memories in Multicore Embedded Systems against Soft Errors
abstract
Scratchpad memories (SPMs) are widely employed in multicore embedded processors. Reliability is one of the major constraints in the embedded processor design, which is threatened with the increasing susceptibility of memory cells to multiple-bit upsets (MBUs) due to continuous technology down-scaling. This article proposes a low-cost and efficient data replication mechanism, called In-Scratchpad Memory Replication (ISMR), to correct MBUs in SPMs of multicore embedded processors. The main feature of ISMR is a smart controller, called Replication Management Unit (RMU), which is responsible for dynamically analyzing the activity of the SPM blocks at runtime and efficiently replicating the vulnerable SPM blocks into currently inactive SPM blocks. RMU exploits a 2-bit tag for each SPM block, where the value of each tag is determined by RMU according to the SPM access pattern. Accordingly, the proposed mechanism guarantees the replication of all vulnerable SPM blocks to provide error correction without decreasing the SPM utilization. To detect errors in SPM blocks, ISMR uses a 2-bit interleaved-parity code. As compared with the previous E-RAID 1 mechanism, the simulation results illustrate that for an 8-core embedded processor, the ISMR mechanism experiences 81% less energy consumption overhead and 48% less performance overhead.
Leila Delshadtehrani, Hamed Farbeh, Seyed Ghassem Miremadi
ACM Trans. Design Autom. Electr. Syst.3
2014 PSP-Cache: A low-cost fault-tolerant cache memory architecture
abstract
Cache memories constitute a large fraction of processor chip area and are highly vulnerable to soft errors caused by energetic particles. To protect these memories, most of the modern processors employ Error Detection Codes (EDCs) or Error Correction Codes (ECCs). EDCs/ECCs impose significant overheads in terms of area and energy; these overheads increase as a function of interleaving EDCs/ECCs to detect/correct multiple errors. This paper proposes a new cache architecture to minimize the area and energy overheads of EDCs/ECCs in set-associative L1-caches. Simulation results for a 4-way set-associative cache show that the proposed architecture reduces both the area and static power overheads of parity code by about 75% and the dynamic energy overhead by about 73% in comparison to conventional cache architecture. These reduction figures are about 68% and about 66%, respectively, for SEC-DED code. The above reductions are achieved without affecting the error coverage.
Hamed Farbeh, Seyed Ghassem Miremadi
DATE2
2014 EA-EO: Endurance Aware Erasure Code for SSD-Based Storage Systems
abstract
One of the main issues in Solid State Drive (SSD)-based storage systems is endurance which is directly affected by the number of Program/Erase (P/E) cycles. The increment of P/E cycles increases the bit error rate threatening the reliability of SSDs. Erasure codes are used to leverage the reliability of storage systems but they also affect the number of P/E cycles based on their code pattern. A lower dependency between data and parities in the code pattern may lead to smaller number of P/E cycles providing better endurance. This paper introduces an Endurance Aware EVENODD (EA-EO), which minimizes the dependency between data and parities in the coding pattern. A simulation environment is used to compare the write-cycles of EA-EO with EVENODD in terms of different request size. The results show that the endurance improvement of EA-EO code could be as high as 44%. Furthermore, performance analysis of these codes in terms of parity construction and failure recovery shows that the number of XOR-operations is reduced in EA-EO compared to EVENODD.
Saeideh Alinezhad Chamazcoti, Seyed Ghassem Miremadi
PRDC2
2014 Developing Inherently Resilient Software Against Soft-Errors Based on Algorithm Level Inherent Features
Bahman Arasteh, Seyed Ghassem Miremadi, Amir Masoud Rahmani
J. Electron. Test.2
2013 FTSPM: A Fault-Tolerant ScratchPad Memory
abstract
ScratchPad Memory (SPM) is an important part of most modern embedded processors. The use of embedded processors in safety-critical applications implies including fault tolerance in the design of SPM. This paper proposes a method, called FTSPM, which integrates a multi-priority mapping algorithm with a hybrid SPM structure. The proposed structure divides SPM into three parts: 1) a part is equipped with Non-Volatile Memory (NVM) which is immune against soft errors, 2) a part is equipped with Error-Correcting Code, and 3) a part is equipped with parity. The proposed mapping algorithm is responsible to distribute the program blocks among the above three parts with regards to their vulnerability level. The simulation results demonstrate that the FTSPM reduces the SPM vulnerability by about 7x in comparison to a pure SRAM-based SPM. In addition, the dynamic energy consumption of the proposed method is 77% and 47% less than that of a pure NVM-based SPM and a pure SRAM-based SPM, respectively.
Amir Mahdi Hosseini Monazzah, Hamed Farbeh, Seyed Ghassem Miremadi, Mahdi Fazeli, Hossein Asadi 0001
DSN3
2013 A non-intrusive portable fault injection framework to assess reliability of FPGA-based designs
abstract
This paper proposes a full-featured fault injection framework to assess reliability of FPGA-based designs. The framework provides non-intrusiveness, portability, flexibility and performance in reliability evaluation of FPGA-based designs against adverse effects of SEUs. It works in a non-intrusive manner, allowing the reliability of ready-to-be-released designs to be assessed independently, without any intrusion into their place and route characteristics. We have studied implications of framework's intrusiveness into design under test by comparing proposed non-intrusive framework with previous intrusive methods; up to 5% deviation in the number of effective faults is observed in intrusive methods. Providing portability, the framework can be applied for a wide variety of FPGAs. Allowing the user to define desired parameters for different fault injection strategies confirms framework's flexibility. Finally, the framework performs the process of injecting faults, evaluating design and removing faults in about 17ms, on average.
Elyas Abolhassani Ghazaani, Zana Ghaderi, Seyed Ghassem Miremadi
FPT3
2013 Low-Cost Scan-Chain-Based Technique to Recover Multiple Errors in TMR Systems
abstract
In this paper, we present a scan-chain-based multiple error recovery technique for triple modular redundancy (TMR) systems (SMERTMR). The proposed technique reuses scan-chain flip-flops fabricated for testability purposes to detect and correct faulty modules in the presence of single or multiple transient faults. In the proposed technique, the manifested errors are detected at the modules' outputs, while the latent faults are detected by comparing the internal states of the TMR modules. Upon detection of any mismatch, the faulty modules are located and the state of a fault-free module is copied into the faulty modules. In case of detecting a permanent fault, the system is degraded to a master/checker configuration by disregarding the faulty module. FPGA-based fault injection experiments reveal that SMERTMR has the error detection and recovery coverage of 100% and 99.7% in the presence of single and two faulty modules, respectively, while imposing negligible area and performance overheads on the traditional TMR systems.
Mojtaba Ebrahimi, Seyed Ghassem Miremadi, Hossein Asadi 0001, Mahdi Fazeli
IEEE Trans. Very Large Scale Integr. Syst.2
2012 SCFIT: A FPGA-based fault injection technique for SEU fault model
abstract
In this paper, we have proposed a fast and easy-to-develop FPGA-based fault injection technique. This technique uses the Altera FPGAs debugging facilities in order to inject SEU fault model in both flip-flops and memory units. Since this method uses the FPGAs built-in facilities, it imposes a negligible performance and area overhead on the system. The experimental results on Leon2 processor shows that the proposed technique is on average four orders of magnitude faster than a simulation-based fault injection.
Abbas Mohammadi 0001, Mojtaba Ebrahimi, Alireza Ejlali, Seyed Ghassem Miremadi
DATE4
2012 Using Genetic Algorithm to Identify Soft-Error Derating Blocks of an Application Program
abstract
Soft-errors are increasingly considered as a major cause for computer system failures. Software techniques are used as cost-effective and flexible techniques to tolerate soft-errors but the introduced overhead is not acceptable in some safety-critical real-time systems. The identification of the program blocks and protecting only vulnerable blocks against soft-errors reduces the performance overhead. In this paper, we present a genetic algorithm to identify the vulnerable program blocks as well as the derating program blocks against soft-errors. Then, only vulnerable blocks are protected by some software-based soft-error tolerance techniques to achieve a lower performance and space overhead. This genetic algorithm is implemented by the C++ programming languages as an automatic tool. To evaluate the algorithm, errors are injected using the Simple scalar toolset. The experimental results indicate that the effectiveness of this method is higher than the previous methods.
Bahman Arasteh, Amir Masoud Rahmani, Ali Mansoor, Seyed Ghassem Miremadi
DSD4
2011 ScTMR: A scan chain-based error recovery technique for TMR systems in safety-critical applications
abstract
We propose a roll-forward error recovery technique based on multiple scan chains for TMR systems, called Scan chained TMR (ScTMR). ScTMR reuses the scan chain flip-flops employed for testability purposes to restore the correct state of a TMR system in the presence of transient or permanent errors. In the proposed ScTMR technique, we present a voter circuitry to locate the faulty module and a controller circuitry to restore the system to the fault-free state. As a case study, we have implemented the proposed ScTMR technique on an embedded processor, suited for safety-critical applications. Exhaustive fault injection experiments reveal that the proposed architecture has the error detection and recovery coverage of 100% with respect to Single Event Upset (SEU) while imposing a negligible area and performance overhead as compared to traditional TMR-based techniques.
Mojtaba Ebrahimi, Seyed Ghassem Miremadi, Hossein Asadi 0001
DATE2
2011 Soft error rate estimation of digital circuits in the presence of Multiple Event Transients (METs)
abstract
In this paper, we present a very fast and accurate technique to estimate the soft error rate of digital circuits in the presence of Multiple Event Transients (METs). In the proposed technique, called Multiple Event Probability Propagation (MEPP), a four-value logic and probability set are used to accurately propagate the effects of multiple erroneous values (transients) due to METs to the outputs and obtain soft error rate. MEPP considers a unified treatment of all three masking mechanisms i.e., logical, electrical, and timing, while propagating the transient glitches. Experimental results through comparisons with statistical fault injection confirm accuracy (only 2.5% difference) and speed-up (10,000X faster) of MEPP.
Mahdi Fazeli, Seyed Nematollah Ahmadian, Seyed Ghassem Miremadi, Hossein Asadi 0001, Mehdi Baradaran Tahoori
DATE3
2011 Numeral-Based Crosstalk Avoidance Coding to Reliable NoC Design
abstract
This paper proposes a Numeral-Based Crosstalk Avoidance Coding (NB-CAC) to protect communication channels of Network-on-Chips (NoCs) against crosstalk faults. The NB-CAC scheme produces code words without bit patterns '101' and '010' to eliminate harmful transition patterns from NoC channels. This is done by the use of a new numeral system proposed in the paper. Using the proposed numeral system, the NB-CAC scheme 1) can be utilized in NoC channels with any arbitrary width, and 2) can be implemented with low area, power, and timing overheads. VHDL and SPICE simulations have been carried out for a wide range of channel widths to evaluate delay, area, and power consumption of the NB-CAC codecs. Results of simulations reveal that the NB-CAC scheme completely removes crosstalk faults from NoC channel. In addition, the NB-CAC scheme provides reductions of 17.3% in area and 31.9% in power-delay product with respect to Fibonacci-based coding which has been recently proposed in literature.
Mansour Shafaei, Ahmad Patooghy, Seyed Ghassem Miremadi
DSD3
2011 Low Cost Concurrent Error Detection for On-Chip Memory Based Embedded Processors
abstract
This paper proposes an efficient concurrent error detection method using control flow checking for embedded processors. The proposed method is based on the co-operation of two hardware modules: 1) an on-chip hardware component to detect branch instructions and generate signatures for the running program, and 2) an external watchdog processor to compare runtime signatures and branch addresses with the information extracted offline. The proposed method is implemented on an embedded processor core and is evaluated by a simulation based statistical fault injection approach where faults are injected into cache and main memory. Experimental results show that the proposed method detects more than 96.7% of all errors with only 2.6% overhead in area and less than 1% increase in power consumption. Furthermore, this technique imposes almost no performance degradation.
Faramarz Khosravi, Hamed Farbeh, Mahdi Fazeli, Seyed Ghassem Miremadi
EUC4
2011 Software-based control flow error detection and correction using branch triplication
abstract
Ever Increasing use of commercial off-the-shelf (COTS) processors to reduce cost and time to market in embedded systems has brought significant challenges in error detection and recovery methods employing in such systems. This paper presents a software based control flow error detection and correction technique, so called branch TMR (BTMR), suitable for use in COTS-based embedded systems. In BTMR method, each branch instruction is triplicated and a software interrupt routine is invoked to check the correctness of the branch instruction. During the execution of a program, when a branch instruction is executed, it is compared with the second redundant branch in the interrupt routine. If a mismatch is detected, the third redundant branch instruction is considered as the error-free branch instruction; otherwise, no error has occurred. The main advantage of BTMR over previously proposed control flow checking (CFC) methods is its ability to correct CFEs as well as protection of indirect branch instructions. The BTMR method is evaluated on LEON2 embedded processor. The results show that, error correction coverage is about 96%, while memory overhead and performance overhead of BTMR is about 28% and 10%, respectively.
Nahid Farhady Ghalaty, Mahdi Fazeli, Hossein Izadi Rad, Seyed Ghassem Miremadi
IOLTS4
2010 A Fast Analytical Approach to Multi-cycle Soft Error Rate Estimation of Sequential Circuits
abstract
In this paper, we propose a very fast analytical approach to measure the overall circuit Soft Error Rate (SER) and to identify the most vulnerable gates and flip-flops. In the proposed approach, we first compute the error propagation probability from an error site to primary outputs as well as system bistables. Then, we perform a multi-cycle error propagation analysis in the sequential circuit. The results show that the proposed approach is four to five orders of magnitude faster than the Monte Carlo (MC) simulation-based fault injection approach with 92% accuracy. This makes the proposed approach applicable to industrial-scale circuits.
Mahdi Fazeli, Seyed Ghassem Miremadi, Hossein Asadi 0001, Mehdi Baradaran Tahoori
DSD2
2010 An Efficient Method to Reliable Data Transmission in Network-on-Chips
abstract
Data transmission in Network-on-Chips (NoCs) is a serious problem due to cross talk faults happening in adjacent communication links. This paper proposes an efficient flow-control method to enhance the reliability of packet transmission in Network-on-Chips. The method investigates the opposite direction transitions appearing between flits of a packet to reorder the flits in the packet. Flits are reordered in a fixed-size window to reduce: 1) the probability of cross talk occurrence, and 2) the total power consumed for packet delivery. The proposed flow-control method is evaluated by a VHDL-based simulator under different window sizes and various channel widths. Simulation results enable NoC designers to make a trade-off between window size, reliability and power consumption of packet delivery. This method is also compared with other cross talk tolerant methods in terms of reliability and power consumption. Comparison results confirm that the method is a cost efficient solution to overcome the cross talk problem.
Ahmad Patooghy, Hamed Tabkhi, Seyed Ghassem Miremadi
DSD3
2010 A fast and accurate multi-cycle soft error rate estimation approach to resilient embedded systems design
abstract
In this paper, we propose a very fast and accurate analytical approach to estimate the overall SER and to identify the most vulnerable gates, flip-flops, and paths of a circuit. Using such information, designers can selectively protect the vulnerable parts resulting in lower power and area overheads that are the most important factors in embedded systems. Unlike previous approaches, the proposed approach firstly does not rely on fault injection or fault simulation; secondly it measures the SER for multi cycles of circuit operation; thirdly, the proposed approach accurately computes all three masking factors, namely, logical, electrical, and timing masking; fourthly, the effects of error propagation in re-convergent fanouts are considered in the proposed approach. SERs estimated by the proposed approach for some ISCAS89 circuit benchmarks are compared with that estimated by the Monte Carlo (MC) simulation based fault injection approach. The results show that the proposed approach is about four orders of magnitude faster than the MC fault injection approach while having an accuracy of about 97%. This level of fastness and accuracy makes the proposed approach a viable solution to measure the SER of very large size circuits used in industry.
Mahdi Fazeli, Seyed Ghassem Miremadi, Hossein Asadi 0001, Seyed Nematollah Ahmadian
DSN2
2010 Crosstalk modeling to predict channel delay in Network-on-Chips
abstract
Communication channels in Network-on-Chips (NoCs) are highly susceptible to crosstalk faults due to the use of nano-scale VLSI technologies in the fabrication of NoCs. Crosstalk faults cause variable timing delay in NoC channels based on the patterns of transitions appearing on the channels. This paper proposes an analytical model to estimate the timing delay of an NoC channel in the presence of crosstalk faults. The model calculates expected number of 4C, 3C, 2C, and 1C transition patterns to predict delay of a K-bit communication channel. The model is applicable for both non-protected channels and channels which are protected by crosstalk mitigation methods. Spice simulations are done in a wide range of working conditions to validate the proposed model. Delays extracted from the simulations are compared with those obtained from the model. Comparisons show that the proposed model accurately estimates the delay of NoC channels. In addition, the proposed model accelerates the evaluation phase of any crosstalk mitigation method by at least three orders of magnitude.
Ahmad Patooghy, Seyed Ghassem Miremadi, Mansour Shafaei
ICCD2
2010 Investigating the Effects of Schedulability Conditions on the Power Efficiency of Task Scheduling in an Embedded System
abstract
Power consumption, performance and reliability are the most important parameters in modern safety-critical distributed real-time embedded systems. This paper evaluates and compares different schedulability conditions in fault-tolerant Rate-Monotonic (RM) and Earliest-Deadline-First (EDF) algorithms, with respect to their power efficiency. The primary-backup scheme is used to implement fault tolerance in the algorithms. To evaluate the algorithms, a software tool is developed that can simulate an embedded system consisting of n processors and m periodic tasks. The results show that depending on the different schedulability conditions, the EDF algorithm implemented with the Best-Fit policy is on average 9.6% more power efficient than other algorithms when n=1500 and m=1000. between the two selected schedulability conditions in the RM algorithm, the Utilization Oriented (UO) condition implemented with the Best-Fit policy is on average 5% more power efficient than the other schedulability condition.
Mohsen Bashiri, Seyed Ghassem Miremadi
ISORC2
2010 FiRot: An Efficient Crosstalk Mitigation Method for Network-on-Chips
abstract
This paper proposes an efficient cross talk mitigation method for Network-on-Chips (NoCs). The proposed method investigates flits in each packet to minimize the number of harmful transition patterns appearing on the communication channels of NoC. To do this, the content of every flit is rotated with respect to the previously flit sent through the channel. Rotation is done to find a rotated version of the flit which minimizes the number of harmful transition patterns. A tag field is added into the rotated flit to enable the receiving side to recover the original flit. Maximum number of rotations is bounded by a fixed value to minimize the timing and power overheads of the proposed method. Evaluation of the proposed method is done in both analytical and simulation manners. VHDL-based simulations are carried out for several channel widths and several tag widths. Simulation results confirm that the proposed method effectively overcomes the cross talk problem while its timing and power overheads are negligible. Results of analytical evaluation are also in agreement with the simulation results.
Ahmad Patooghy, Mansour Shafaei, Seyed Ghassem Miremadi, Hajar Falahati, Somayyeh Taheri
PRDC3
2010 Classification of Activated Faults in the FlexRay-Based Networks
Yasser Sedaghat, Seyed Ghassem Miremadi
J. Electron. Test.2
2010 A low-overhead and reliable switch architecture for Network-on-Chips
Ahmad Patooghy, Seyed Ghassem Miremadi, Mahdi Fazeli
Integr.2
2010 Performability/Energy Tradeoff in Error-Control Schemes for On-Chip Networks
abstract
High reliability against noise, high performance, and low energy consumption are key objectives in the design of on-chip networks. Recently some researchers have considered the impact of various error-control schemes on these objectives and on the tradeoff between them. In all these works performance and reliability are measured separately. However, we will argue in this paper that the use of error-control schemes in on-chip networks results indegradable systems, hence, performance and reliability must be measured jointly using a unified measure, i.e.,performability. Based on the traditional concept of performability, we provide a definition for the ¿Interconnect Performability¿. Analytical models are developed for interconnect performability and expected energy consumption. A detailed comparative analysis of the error-control schemes using the performability analytical models and SPICE simulations is provided taking into consideration voltage swing variations (used to reduce interconnect energy consumption) and variations in wire length. Furthermore, the impact of noise power and time constraint on the effectiveness of error-control schemes are analyzed.
Alireza Ejlali, Bashir M. Al-Hashimi, Paul M. Rosinger, Seyed Ghassem Miremadi, Luca Benini
IEEE Trans. Very Large Scale Integr. Syst.4
2009 Fault Tolerant and Low Energy Write-Back Heterogeneous Set Associative Cache for DSM Technologies
abstract
This paper presents a fault tolerant and energy efficient write-back set-associative cache, which has a heterogeneous structure. The cache architecture is based on partitioning the ways of each set into two different parts. In each set, one cache way uses SEC-DED code and maintains dirty blocks while the other ways employ parity bit and keep clean blocks. To evaluate the set-associative cache, SIMPLESCALAR tool and CACTI analytical model are used. The experimental results show that as the feature size decreases and the associativity increases, the energy saving of the proposed cache increases. The experimental results express that for an 8-way set-associative cache in 32 nm, about 7% area and 2%-17% energy consumption are saved. These figures are achieved by keeping the reliability in the same level of the conventional SEC-DED protected cache.
Mehrtash Manoochehri, Alireza Ejlali, Seyed Ghassem Miremadi
ARES3
2009 A High Speed and Low Cost Error Correction Technique for the Carry Select Adder
abstract
In this paper, a high speed and low cost error correction technique is proposed for the Carry Select Adder (CSA) which can correct both transient and permanent errors and is applicable on all partitioning types of the basic CSA circuit. The proposed error correction technique is compatible with all existing error detection techniques which are proposed for the CSA adder. The synthesized results show that applying this novel error correction technique to a CSA with error detection technique results in up to 18.4%, 3.1% and 14.9%, increase in power consumption, delay and area respectively.
Alireza Namazi, Seyed Ghassem Miremadi, Alireza Ejlali
ARES2
2009 A Micro-FT-UART for Safety-Critical SoC-Based Applications
abstract
This paper presents the design of a fault-tolerant universal asynchronous receiver transmitter (UART) called micro-FT-UART for safety-critical SoC-based applications. This UART exploits advantages of three fault-tolerant techniques to tolerate soft errors. The three techniques are triple modular redundancy (TMR), Hamming code and a new technique called correction by parity storing (CPS). An VHDL model of a micro-UART is simulated by the ModelSim v.6.0 and synthesized by the Synopsys Design Compiler v.X-2005.09-SP2. About 1000 single-bit errors and 1000 multiple-bit errors are injected into different parts of the micro-UART to find out the error sensitivity of each specific part. Considering tradeoff between reliability and power consumption, an optimum fault-tolerant technique is assigned to each part to design the micro-FT-UART. This UART corrects all single-bit errors and on average 24% of multiple-bit errors with about 81% power consumption overhead and 152% area overhead.
Mohammad-Hamed Razmkhah, Seyed Ghassem Miremadi, Alireza Ejlali
ARES2
2009 An energy efficient circuit level technique to protect register file from MBUs and SETs in embedded processors
abstract
This paper presents a circuit level soft error-tolerant-technique, called RRC (robust register caching), for the register file of embedded processors. The basic idea behind the RRC is to effectively cache the most vulnerable registers in a small highly robust register cache built by circuit level SEU and SET protected memory cells. To decide which cache entry should be replaced, the average number of read operations during a register ACE time is used as a criterion to judge. In fact, the victim cache entry is one which has the maximum read count. To minimize the power overhead of the RRC, the clock gating technique is efficiently exploited for the main register file resulting in significantly low power consumption. The RRC is able to protect the register file not only against single bit upsets (SBUs) but also against multiple bit upsets (MBUs) and single event transients (SETs). The RRC is experimentally evaluated using the LEON processor. The experimental results show that, if the cache size is selected properly, the architectural vulnerability factor (AVF) of the register file becomes about 1% while it imposes low power, area and performance overheads to the processor.
Mahdi Fazeli, Alireza Namazi, Seyed Ghassem Miremadi
DSN3
2009 Categorizing and Analysis of Activated Faults in the FlexRay Communication Controller Registers
abstract
FlexRay communication protocol is expected becoming the de-facto standard for distributed safety-critical systems. In this paper, transient single bit-flip faults were injected into the FlexRay communication controller to categorize and analyze the activated faults. In this protocol, an activated fault results in one or more error types which are boundary violation, conflict, content, freeze, synchronization, and syntax. To study the activated faults, a FlexRay bus network, composed of four nodes, was modeled by verilog HDL; and a total of 135,600 transient faults were injected in only one node, where 9,342 (6.9%) of the faults were activated. The results show that the synchronization error is the widespread error with the occurrence ratio of about 70.1%. The Boundary violation and the syntax errors have the occurrence ratios of 32.4% and 24.6%, respectively. The results also show that the Freeze error which more frequent resulted system failures has the occurrence ratio of about 17.3%.
Yasser Sedaghat, Seyed Ghassem Miremadi
ETS2
2009 A low-cost fault-tolerant technique for Carry Look-Ahead adder
abstract
This paper proposes a low-cost fault-tolerant Carry Look-Ahead (CLA) adder which consumes much less power and area overheads in comparison with other fault-tolerant CLA adders. Analytical and experimental results show that this adder corrects all single-bit and multiple-bit transient faults. The Power-Delay Product (PDP) and area overheads of this technique are decreased at least 82% and 71%, respectively, as compared to adders which use traditional TMR, parity prediction, and duplication techniques.
Alireza Namazi, Yasser Sedaghat, Seyed Ghassem Miremadi, Alireza Ejlali
IOLTS3
2009 XYX: A Power & Performance Efficient Fault-Tolerant Routing Algorithm for Network on Chip
abstract
Reliability is one of the main concerns in the design of network on chips due to the use of deep-sub micron technologies in fabrication of such products. This paper proposes a fault-tolerant routing algorithm called XYX which is based on sending redundant packets through the paths with lower traffic loads. The XYX routing algorithm makes a redundant copy of each packet at the source node and exploits two different routing algorithms to route the original and the redundant packets. Since two copies of each packet reach the destination node, the erroneous packet is detected and replaced with the correct one. Due to the use of paths with lower traffic rates for sending redundant packets and minimizing the number of sent redundant packets, the XYX routing algorithm provides lower performance and power overheads as compared to flood-based routing algorithms. Experimental results show that the XYX routing algorithm imposes negligible performance and power consumption overheads while providing almost the same reliability in comparison with flood-based routing algorithms.
Ahmad Patooghy, Seyed Ghassem Miremadi
PDP2
2008 Fault Effects in FlexRay-Based Networks with Hybrid Topology
abstract
This paper investigates fault effects and error propagation in a FlexRay-based network with hybrid topology that includes a bus subnetwork and a star subnetwork. The investigation is based on about 43500 bit-flip fault injection inside different parts of the FlexRay communication controller. To do this, a FlexRay communication controller is modeled by Verilog HDL at the behavioral level. Then, this controller is exploited to setup a FlexRay-based network composed of eight nodes (four nodes in the bus subnetwork and four nodes in the star subnetwork). The faults are injected in a node of the bus subnetwork and a node of the star subnetwork of the hybrid network Then, the faults resulting in the three kinds of errors, namely, content errors, syntax errors and boundary violation errors are characterized. The results of fault injection show that boundary violation errors and content errors are negligibly propagated to the star subnetwork and syntax errors propagation is almost equal in the both bus and star subnetworks. Totally, the percentage of errors propagation in the bus subnetwork is more than the star subnetwork.
Mehdi Dehbashi, Vahid Lari, Seyed Ghassem Miremadi, Mohammad Shokrollah-Shirazi
ARES3
2008 FEDC: Control Flow Error Detection and Correction for Embedded Systems without Program Interruption
abstract
This paper proposes a new technique called CFEDC to detect and correct control flow errors (CFEs) without program interruption. The proposed technique is based on the modification of application software and minor changes in the underlying hardware. To demonstrate the effectiveness of CFEDC, it has been implemented on an OpenRISC 1200 as a case study. Analytical results for three workload programs show that this technique detects all CFEs and corrects on average about 81.6% of CFEs. These figures are achieved with zero error detection /correction latency. According to the experimental results, the overheads are generally low as compared to other techniques; the performance overhead and the memory overhead are on average 8.5% and 9.1%, respectively. The area overhead is about 4% and the power dissipation increases by the amount of 1.5% on average.
Navid Farazmand, Mahdi Fazeli, Seyed Ghassem Miremadi
ARES3
2008 Analyzing fault effects in the 32-bit OpenRISC 1200 microprocessor
abstract
This paper presents an analysis of the effects and propagation of faults in the open-core 32-bit OpenRISC 1200 microprocessor. The analysis is based on a total of 13,000 transient faults injected into 65 parts of the CPU module in the OpenRISC 1200 core described at the RTL model. A comparison of the effects of faults on the various parts of the CPU including the pipeline's registers, the CPU component such as the register file, the control unit, and the ALU, and the data and address buses is done. It is shown that about 30%, 40% and 27% of injected faults terminated in address, data, and control errors respectively. About 28% of all injected faults resulted in failures.
Nima Mehdizadeh, Mohammad Shokrollah-Shirazi, Seyed Ghassem Miremadi
ARES3
2008 Control-Flow Checking Using Branch Instructions
abstract
This paper presents a hardware control-flow checking scheme for RISC processor-based systems. This scheme combines two error detection mechanisms to provide high coverage. The first mechanism uses parity bits to detect faults occurring in the opcodes and in the target addresses of branch instructions which lead to erroneous branches. The second mechanism uses signature monitoring to detect errors occurring in the sequential instructions. The scheme is implemented using a watchdog processor for an VHDL model of the LEON2 processor. About 31800 simulation faults were injected into the LEON2 processor. The results show that the error detection coverage is about 99.5% with average detection latency of 7 cycles. The performance loss of presented scheme is about 8.4%.
Mostafa Jafari-Nodoushan, Seyed Ghassem Miremadi, Alireza Ejlali
EUC (1)2
2008 A Low Power Error Detection Technique for Floating-Point Units in Embedded Applications
abstract
Reliability and low power consumption are two major design objectives in today's embedded systems. Since floating-point units (FPU) are required for some embedded applications (e.g., multimedia applications), careful considerations should be given to the reliability and power consumptions of FPUs used in embedded systems. When using existing fault handling mechanisms for FPUs, it has been observed that the division operation imposes a considerable hardware overhead as compared to the addition, subtraction, and multiplication operations. Although the division operation is less frequently used, in reliable applications it is a must that all the components operate properly. In this paper, we present a low power error detection mechanism for the division operation in FPUs. In this technique the FPU multiplier circuitry is modified so that it can be used to detect the errors that may happen in the divider circuitry. The experimental results show that while the proposed technique can detect almost all the errors in the division circuitry, its power-delay-product (PDP) is about 23% lower than that of the traditional error detection techniques.
Seyed Mohammad Hossein Shekarian, Alireza Ejlali, Seyed Ghassem Miremadi
EUC (1)3
2008 Investigation and Reduction of Fault Sensitivity in the FlexRay Communication Controller Registers
Yasser Sedaghat, Seyed Ghassem Miremadi
SAFECOMP2
2008 Error Detection Enhancement in PowerPC Architecture-based Embedded Processors
Mahdi Fazeli, Reza Farivar 0003, Seyed Ghassem Miremadi
J. Electron. Test.3
2007 The Effect of Routing-Update Time on Network's Performability
abstract
By occurring failures in computer networks, routing protocols are triggered to update routing and forwarding tables. Because of invalid tables during update-time, transient loop may occur and packet-drop rate and end- to-end delay increase which means that the quality of service decreases. This paper studies the effect of routing- tables update-time on networks' performability, i.e. the ability of network to deliver services at predefined level. A sample network is studied and the simulation results show that faster updates of routing table, improve network's performability in the presence of failures. Since it may not worth or even be practical to accelerate all routers in the network, this paper suggests finding bottleneck routers and accelerating them in order to improve the performability of the network. The simulation results show that by speeding-up the bottleneck routers of the network, instead of all routers, the desired performability could be achieved.
Mostafa Shaad Zolpirani, Mohammad-Mahdi Bidmeshki, Seyed Ghassem Miremadi
AICCSA3
2007 Joint consideration of fault-tolerance, energy-efficiency and performance in on-chip networks
Alireza Ejlali, Bashir M. Al-Hashimi, Paul M. Rosinger, Seyed Ghassem Miremadi
DATE4
2007 Feedback Redundancy: A Power Efficient SEU-Tolerant Latch Design for Deep Sub-Micron Technologies
abstract
The continuous decrease in CMOS technology feature size increases the susceptibility of such circuits to single event upsets (SEU) caused by the impact of particle strikes on system flip flops. This paper presents a novel SEU-tolerant latch where redundant feedback lines are used to mask the effects of SEUs. The power dissipation, area, reliability, and propagation delay of the presented SEU-tolerant latch are analyzed by SPICE simulations. The results show that this latch consumes about 50% less power and occupies 42% less area than a TMR-latch. However, the reliability and the propagation delay of the proposed latch are still the same as the TMR-latch. the reliability of the proposed latch is also compared with other SEU-tolerant latches.
Mahdi Fazeli, Ahmad Patooghy, Seyed Ghassem Miremadi, Alireza Ejlali
DSN3
2007 Fault-Tolerant Earliest-Deadline-First Scheduling Algorithm
abstract
The general approach to fault tolerance in uniprocessor systems is to maintain enough time redundancy in the schedule so that any task instance can be re-executed in presence of faults during the execution. In this paper a scheme is presented to add enough and efficient time redundancy to the earliest-deadline-first (EDF) scheduling policy for periodic real-time tasks. This scheme can be used to tolerate transient faults during the execution of tasks. We describe a recovery scheme which can be used to re-execute tasks in the event of transient faults and discuss conditions that must be met by any such recovery scheme. For performance evaluation of this idea a tool is developed.
Hakem Beitollahi, Seyed Ghassem Miremadi, Geert Deconinck
IPDPS2
2007 Fast SEU Detection and Correction in LUT Configuration Bits of SRAM-based FPGAs
abstract
FPGAs are an appealing solution for the space-based remote sensing applications. However, in a low-earth orbit, configuration bits of SRAM-based FPGAs are susceptible to single-event upsets (SEUs). In this paper, a new protected CLB and FPGA architecture are proposed which utilize error detection and correction codes to correct SEUs occurred in LUTs of the FPGA. The fault detection and correction is achieved using online or offline fast detection and correction cycles. In the latter, detection and correction is performed in predefined error-correction intervals. In both of them error detections and corrections of k-input LUTs are performed with a latency of 2kclock cycle without any required reconfiguration and significant area overhead. The power and area analysis of the proposed techniques show that these methods are more efficient than the traditional schemes such as duplication with comparison and TMR circuit design in the FPGAs.
Hamid R. Zarandi, Seyed Ghassem Miremadi, Costas Argyrides, Dhiraj K. Pradhan
IPDPS2
2007 CLB-based Detection and Correction of Bit-flip faults in SRAM-based FPGAs
abstract
This paper presents a bit-flip tolerance in SRAM-based FPGAs which suffers from high energy particles, alpha and neutrons in the atmosphere. For each of protections, the applicability, efficiency and implementation issues are discussed. Moreover, the area, the power and the protection capability of the methods are mentioned and compared with previous work. Based on the results of experiments and their analysis, one method is selected as best one. The selected method is much better than previous work e.g., duplication with comparison, triple modular redundancy which impose two and three area and power overheads, respectively.
Hamid R. Zarandi, Seyed Ghassem Miremadi, Costas Argyrides, Dhiraj K. Pradhan
ISCAS2
2007 Soft Error Mitigation in Switch Modules of SRAM-based FPGAs
abstract
In this paper, we propose two techniques to mitigate soft error effects on the switch modules of SRAM-based FPGAs: 1) The first technique tolerates SEU-caused open errors based on a new programming method for SRAM-bits of switch modules, and 2) The second technique mitigates SEU-cause short errors in the switch modules based on a mixed programmable and hard-wired switch module structure in the FPGAs. The effects of these two techniques on the delay, area and power consumption for 20 MCNC benchmark circuits are achieved using a minor modification in VPR and T-VPack FPGA CAD tools. The experimental results show that the first technique increase reliability of connections of switch module up to 30% while the second technique decreases the susceptibility of switch modules to SEUs about 50% compared to the traditional ones
Hamid R. Zarandi, Seyed Ghassem Miremadi, Dhiraj K. Pradhan, Jimson Mathew
ISCAS2
2007 CAD-Directed SEU Susceptibility Reduction in FPGA Circuits Designs
abstract
This paper presents a SEU-mitigative placement and route of circuits in the FPGAs which is based on the popular placement and route tool. The tool is modified so that during placement and routing, decisions are taken with awareness of SEU-mitigation and no redundancies during the placement and routing are used but the algorithms are based on the SEU avoidance. We have investigated the effect of this tool on several MCNC benchmarks and the results of the placement and routing have been compared to the traditional one. The evaluations of results show that placement and routing can decrease the SEU rate of circuits implemented on FPGAs about 22%. However, it increases critical path delay and power consumptions of the circuits.
Hamid R. Zarandi, Seyed Ghassem Miremadi, Dhiraj K. Pradhan, Jimson Mathew
ISCAS2
2007 Assessment of Message Missing Failures in FlexRay-Based Networks
abstract
This paper assesses message missing failures in a FlexRay-based network. The assessment is based on about 35680 bit-flip fault injections inside different parts of the FlexRay communication controller; the parts are: controller host interface, protocol operation control, coding and decoding unit, media access control and clock synchronization process. To do this, a FlexRay communication controller is modeled by Verilog HDL at the behavioral level. This HDL model of the controller is exploited to setup a FlexRay-based network composed of four nodes. The results of fault injection show that about 35% of faults led to the message missing failures. The controller host interface and the clock synchronization process of the FlexRay were the most sensitive parts to the message missing failures. The coding and decoding unit of the FlexRay was the least sensitive part to these failures.
Vahid Lari, Mehdi Dehbashi, Seyed Ghassem Miremadi, Navid Farazmand
PRDC3
2007 A Low-Power and SEU-Tolerant Switch Architecture for Network on Chips
abstract
High reliability, high performance, low power consumption are the main objectives in the design of NoCs. These three design objectives are mostly conflicting and should be considered simultaneously in order to have an optimal design. This paper proposes a method based on duplicating the virtual channels of each NoC node as well as parity codes to prevent SEUs from producing erroneous data. The method is compared with two widely used SEU-tolerant methods i.e., the switch to switch and the end to end flow control methods, in terms of reliability, power consumption and performance. A flit level VHDL-based simulator and Synopsys power compiler tool have been used to extract experimental results. The simulation results show the same reliability for all three methods, while the proposed method shows the lowest power consumption and the highest performance almost in all traffic generation rates and all packet error rates.
Ahmad Patooghy, Mahdi Fazeli, Seyed Ghassem Miremadi
PRDC3
2006 A Solution to Single Point of Failure Using Voter Replication and Disagreement Detection
abstract
This paper suggests a method, called distributed voting, to overcome the problem of the single point of failure in a TMR system used in robotics and industrial control applications. It uses time redundancy and is based on TMR with disagreement detector feature. This method masks faults occurring in the voter where the TMR system can continue its function properly. The method has been evaluated by injecting faults into Vertex2Pro and Vertex4 Xilinx FPGAs An analytical evolution is also performed. The results of both evaluation approaches show that the proposed method can improve the reliability and the mean time to failure (MTTF) of a TMR system by at least a factor of (2-RV(t)) where RV(t) is the reliability of the voter
Ahmad Patooghy, Seyed Ghassem Miremadi, Abbas Javadtalab, Mahdi Fazeli, Navid Farazmand
DASC2
2006 Combined time and information redundancy for SEU-tolerance in energy-efficient real-time systems
abstract
Recently, the tradeoff between energy consumption and fault-tolerance in real-time systems has been highlighted. These works have focused on dynamic voltage scaling (DVS) to reduce dynamic energy dissipation and on-time redundancy to achieve transient-fault tolerance. While the time redundancy technique exploits the available slack-time to increase the fault-tolerance by performing recovery executions, DVS exploits slack-time to save energy. Therefore, we believe there is a resource conflict between the time-redundancy technique and DVS. The first aim of this paper is to propose the use of information redundancy to solve this problem. We demonstrate through analytical and experimental studies that it is possible to achieve both higher transient fault-tolerance [tolerance to single event upsets (SEUs)] and less energy using a combination of information and time redundancy when compared with using time redundancy alone. The second aim of this paper is to analyze the interplay of transient-fault tolerance (SEU-tolerance) and adaptive body biasing (ABB) used to reduce static leakage energy, which has not been addressed in previous studies. We show that the same technique (i.e., the combination of time and information redundancy) is applicable to ABB-enabled systems and provides more advantages than time redundancy alone.
Alireza Ejlali, Bashir M. Al-Hashimi, Marcus T. Schmitz, Paul M. Rosinger, Seyed Ghassem Miremadi
IEEE Trans. Very Large Scale Integr. Syst.5
2005 Energy efficient SEU-tolerance in DVS-enabled real-time systems through information redundancy
abstract
Concerns about the reliability of real-time embedded systems that employ dynamic voltage scaling has recently been highlighted [1,2,3], focusing on transient-fault-tolerance techniques based on time-redundancy. In this paper we analyze the usage of information redundancy in DVS-enabled systems with the aim of improving both the system tolerance to transient faults as well as the energy consumption. We demonstrate through a case study that it is possible to achieve both higher fault-tolerance and less energy using a combination of information and time redundancy when compared with using time redundancy alone. This even holds despite the impact of the information redundancy hardware overhead and its associated switching activities
Alireza Ejlali, Marcus T. Schmitz, Bashir M. Al-Hashimi, Seyed Ghassem Miremadi, Paul M. Rosinger
ISLPED4
2005 A Hardware Approach to Concurrent Error Detection Capability Enhancement in COTS Processors
abstract
To enhance the error detection capability in COTS (commercial off-the-shelf)-based design of safety-critical systems, a new hardware-based control flow checking (CFC) technique is presented. This technique, control flow checking by execution tracing (CFCET), employs the internal execution tracing features available in COTS processors and an external watchdog processor (WDP) to monitor the addresses of taken branches in a program. This is done without any modification of application programs, therefore, the program overhead is zero. The external hardware overhead is about 3.5% using an Altera Flex 10K30 FPGA. For different workload programs, the execution time overhead and the error detection coverage of the technique vary between 33.3 and 140.8% and between 79.7 and 84.6% respectively. The errors are detected with about zero latency.
Amir Rajabzadeh, Seyed Ghassem Miremadi
PRDC2
2005 Contribution of Controller Area Networks Controllers to Masquerade Failures
abstract
This paper scrutinizes faults in a CAN controller that may result in masquerade failures, and suggests an even parity mechanism to detect them with minimum hardware overhead. To do this, a CAN controller is modeled by VHDL at behavioral level and is exploited to setup a CAN-based network composed of two nodes. A total of 5,500 faults are injected into essential parts of one of the controllers. The results show that about 3.44% of faults terminate in masquerade failures. The results, also, show that register bank in the CAN controller are the most sensitive portions in which 92.10% of faults occurring in the register bank result in masquerade failures. The even parity mechanism detects about 96.31% of all masquerade failures.
Hassan Salmani, Seyed Ghassem Miremadi
PRDC2
2004 Experimental Evaluation of Master/Checker Architecture Using Power Supply- and Software-Based Fault Injection
Amir Rajabzadeh, Seyed Ghassem Miremadi, Mirzad Mohandespour
IOLTS2
2004 Fault Detection Enhancement in Cache Memories Using a High Performance Placement Algorithm
Hamid R. Zarandi, Seyed Ghassem Miremadi, Hamid Sarbazi-Azad
IOLTS2
2004 Evaluation of Fault-Tolerant Designs Implemented on SRAM-Based FPGAs
abstract
The technology of SRAM-based devices is sensible to single event upsets (SEUs) that may be induced mainly by high energy heavy ions and neutrons. We present a framework for the evaluation of fault-tolerant designs implemented on SRAM-based FPGAs using emulated SEUs. The SEU injection process is performed by inserting emulated SEUs in the device using its configuration bitstream file. An Altera FPGA, i.e. the Flex10K200, and the ITC'99 benchmark circuits are used to experimentally evaluate the method. The results show that between 32 to 45 percent of SEUs injected to the device propagate to the output terminals of the device.
Hossein Asadi 0001, Seyed Ghassem Miremadi, Hamid R. Zarandi, Alireza Ejlali
PRDC2
2004 Error Detection Enhancement in COTS Superscalar Processors with Event Monitoring Features
abstract
Increasing use of commercial off-the-shelf (COTS) superscalar processors in industrial, embedded, and real-time systems necessitates the development of error detection mechanisms for such systems. This shows an error detection scheme called committed instructions counting (CIC) to increase error detection in such systems. The scheme uses internal performance monitoring features and an external watchdog processor (WDP). The performance monitoring features enable counting the number of committed instructions in a program. The scheme is experimentally evaluated on a 32-bit Pentium/spl reg/ processor using software implemented fault injection (SWIFI). A total of 8181 errors were injected into the Pentium/spl reg/ processor. The results show that the error detection coverage varies between to 90.92% and 98.41%, for different workloads.
Amir Rajabzadeh, Mirzad Mohandespour, Seyed Ghassem Miremadi
PRDC3
2004 A Highly Fault Detectable Cache Architecture for Dependable Computing
Hamid R. Zarandi, Seyed Ghassem Miremadi
SAFECOMP2
2004 Error Detection Enhancement in COTS Superscalar Processors with Performance Monitoring Features
Amir Rajabzadeh, Seyed Ghassem Miremadi, Mirzad Mohandespour
J. Electron. Test.2
2003 Switch-level emulation
abstract
This paper presents a method for the fast emulation of switch-level circuits using FPGAs. In this method, logic gates are used to model switch-level circuits without any abstraction. In contrast to the abstraction methods for which transistors are grouped together to form gates, in this method, gates are grouped together to form the switch models of transistors. Unlike the abstraction methods, the method presented in this paper can emulate many important features of switch-level models, such as bi-directional signal propagation and variations in driving strength. In order to attain a better utilization of FPGA resources a mixed-mode emulation approach has been used. In this approach parts of the circuit are emulated at the switch-level while the rest of the circuit is emulated at the gate-level.
Alireza Ejlali, Seyed Ghassem Miremadi
DAC2
2003 A Hybrid Fault Injection Approach Based on Simulation and Emulation Co-operation
abstract
This paper presents a new fault injection approach, which is based on a co-operation between a simulator and an emulator. This hybrid approach utilizes the advantages of both simulation-based fault injection as well as physical fault injection to provide a good controllability, observability and also a high speed in the fault injection experiments. To do this, parts of a circuit are simulated while the rest parts of the circuit are emulated. A fault injection tool called FITSEC (Fault Injection Tool based on Simulation and Emulation Cooperation) is developed, which supports the entire process of a system design. This is based on both Verilog and VHDL languages and can be used to inject faults at different levels of abstraction. The experimental results show that this approach can significantly reduce the time needed for executing fault injection campaigns.
Alireza Ejlali, Seyed Ghassem Miremadi, Hamid R. Zarandi, Hossein Asadi 0001, Siavash Bayat Sarmadi
DSN2
2003 Switch Level Fault Emulation
Seyed Ghassem Miremadi, Alireza Ejlali
FPL1
2003 Fault injection into SRAM-based FPGAs for the analysis of SEU effects
abstract
SRAM-based FPGAs are currently utilized in applications such as industrial and space applications where high availability and reliability and low cost are important constraints. The technology of such devices is sensible to Single Event Upsets (SEUs) that may be originated mainly from heavy ion radiation. This paper presents a fault injection method that is based on emulated SEU on the configuration bitstream file of commercial SRAM-based FPGA devices to study the error propagation in these devices. To demonstrate the method, an Altera FPGA, i.e. the Flex10K200, and the ITC'99 benchmark circuits are used. A fault injection tool is developed to inject emulated SEU faults into the circuits. The results show that between 33 to 45 percent of the SEUs injected to the FPGA device have propagated to the output terminals of the device.
Hossein Asadi 0001, Seyed Ghassem Miremadi, Hamid R. Zarandi, Alireza Ejlali
FPT2
2003 Fault Injection into Verilog Models for Dependability Evaluation of Digital Systems
abstract
This paper presents transient and permanent fault injection into Verilog models of digital systems during the design phase by a developed simulation-based fault injection tool called INJECT. With this fault injection tool, it is possible to inject crucial fault models in all abstraction levels (such as swith-level) supported by Verilog HDL. Several fault models for injecting into Verilog models are specified and described. Analyzing the results obtained from the fault injections, using INJECT enables system designers to inform from dependable parameters, such as fault latency, propagation and coverage. As a case study, a 32-bit processor, namely DP32, has been evaluated and effects of faults on some important observation points have been presented. In this study, recovered errors are distinguished from those that affected the system behavior. The errors that lead to wrong results are separated from those that do not affect the correct results.
Hamid R. Zarandi, Seyed Ghassem Miremadi, Alireza Ejlali
ISPDC2
2002 Fast Prototyping with Co-operation of Simulation and Emulation
Siavash Bayat Sarmadi, Seyed Ghassem Miremadi, Hossein Asadi 0001, Alireza Ejlali
FPL2
2002 Speedup analysis in simulation-emulation co-operation
abstract
This paper presents an analytical approach to estimate the speedup in a simulation-emulation cooperation environment. The speedup of this approach as compared with the speedup of a pure simulation is analyzed. Also, an analysis of the speedup is given when different types of application instructions are utilized. The analysis is based on using both Verilog and VHDL. The results show that when only the simulation part of the simulation-emulation co-operation is used, the speedup is higher, than when the pure simulation is used. The total speedup is also depended on the type of application instructions and the communication cycle time between the simulator and the emulator.
Seyed Ghassem Miremadi, Siavash Bayat Sarmadi, Hossein Asadi 0001
FPT1