EDBT 2026 Demo / reviewers in the wild / expert
Hamed Farbeh
dblp:57/10661
· DBLP profile ↗
27ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-4204-9131ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReSAFT: An efficient stuck-at fault-tolerant scheme for ReRAM-based process-in-memory accelerators
Aniseh Dorostkar, Hamed Farbeh, Hamid R. Zarandi |
Future Gener. Comput. Syst. | 2 |
| 2025 | An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable CacheabstractAssociative cache memory plays a decisive role in enhancing the performance and energy consumption of modern processors. Meanwhile, by occupying more than half of the processor chip area, cache memory is susceptible to transient and permanent faults, threatening the system's dependability. As the onlyhardware-managedmemory module in the system, the tag array of the caches is the most critical and active component contributing a large fraction of energy consumption and error occurrence.Tag partitioningis a widespread approach for both tag energy consumption reduction and reliability enhancement. This approach splits the tag comparison operation into two steps, and only the tags whoseklower order bits are matched with that of the input address in the first step are activated for comparing their remaining higher order bits in the second step. The key decision parameter for tag partitioning is properly adjusting the tag-splitting point (k) to achieve the maximum reduction in the number of reads. This parameter has been intuitively, randomly, or experimentally selected in the existing studies without any justification. Even for an appropriate selection of this parameter via extensive experiments, its sensitivity to various cache configuration parameters makes it ad-hoc and not extendable to other scenarios. In this paper, we analytically illustrate that selecting an inappropriately large or small value for the tag-splitting point significantly downgrades the efficiency of tag partitioning and then formulate this parameter to determine its optimum value. As a function of cache configuration parameters, the proposed formulation is proven to be convex and differentiable for determining the optimum splitting point, besides its ability to accurately report the degree of the tag partitioning efficiency for any splitting point and configuration parameters. To approve the correctness and accuracy of the proposed formulation, we experimentally investigate the tag partitioning efficiency and optimum splitting point for a wide range of cache configurations and demonstrate a very close matching between the two. The proposed formulation is a guarantee for the designers and researchers to instantaneously determine the optimum tag-splitting point and calculate the read reduction of tag partitioning. Elham Cheshmikhani, Hamed Farbeh |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | An Empirical Fault Vulnerability Exploration of ReRAM-Based Process-in-Memory CNN AcceleratorsabstractResistive random-access memory (ReRAM)-basedprocessing-in-memory(PIM) accelerator is a promising platform for processing massively memory intensive matrix-vector multiplications of neural networks in parallel domain, due to its capability of analog computation, ultra-high density, near-zero leakage current, and nonvolatility. Despite many advantages, ReRAM-based accelerators are highly error-prone due to limitations of technology fabrication that lead to process variations and defects. These limitations degrade the accuracy of deep convolutional neural networks (CNNs) (Deep CNNs) running on PIM accelerators. While these CNNs accelerators are widely deployed in safety-critical systems, their vulnerability to fault is not well explored. In this article, we have developed a fault-injection framework to investigate the vulnerability of large-scale CNNs at both software- and hardware-level of inference phases. Faulty ReRAM devices are another reliability challenges due to significant degradation of classification accuracy when CNN parameters are mapped to the accelerators. To investigate this challenge, we map the CNN learning parameter to the ReRAM crossbar and inject faults into crossbar arrays. The proposed framework analyzes the impact ofstuck-at high(SaH) andstuck-at low(SaL) fault models on different layers and locations of CNN learning parameters. By performing extensive fault injections, we illustrate that the vulnerability behavior of ReRAM-based PIM accelerator for CNNs is greatly impressible to the types and depth of layers, the location of the learning parameter in every layer, and the value and types of faults. Our observations show that different models have different vulnerabilities to faults. Specifically, we show that SaL further reduces classification accuracy than SaH. Aniseh Dorostkar, Hamed Farbeh, Hamid R. Zarandi |
IEEE Trans. Reliab. | 2 |
| 2023 | An adaptive data coding scheme for energy consumption reduction in SDN-based Internet of Things
Shahab Salehi, Hamed Farbeh, Alireza Rokhsari |
Comput. Networks | 2 |
| 2022 | A Novel Neuromorphic Processors Realization of Spiking Deep Reinforcement Learning for Portfolio ManagementabstractThe process of constantly reallocating budgets into financial assets, aiming to increase the anticipated return of assets and minimizing the risk, is known as portfolio management. Processing speed and energy consumption of portfolio management have become crucial as the complexity of their real-world applications increasingly involves high-dimensional observation and action spaces and environment uncertainty, which their limited onboard resources cannot offset. Emerging neuromorphic chips inspired by the human brain increase processing speed by up to 500 times and reduce power consumption by several orders of magnitude. This paper proposes a spiking deep reinforcement learning (SDRL) algorithm that can predict financial markets based on unpredictable environments and achieve the defined portfolio management goal of profitability and risk reduction. This algorithm is optimized for Intel's Loihi neuromorphic processor and provides 186x and 516x energy consumption reduction compared to a high-end processor and GPU, respectively. In addition, a 1.3x and 2.0x speed-up is observed over the high-end processors and GPUs, respectively. The evaluations are performed on cryptocurrency market benchmark between 2016 and 2021. Seyyed Amirhossein Saeidi, Forouzan Fallah, Soroush Barmaki, Hamed Farbeh |
DATE | 4 |
| 2022 | GraphA: An efficient ReRAM-based architecture to accelerate large scale graph processing
Seyed Ali Ghasemi, Belal Jahannia, Hamed Farbeh |
J. Syst. Archit. | 3 |
| 2022 | 3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag ComparisonabstractRecent development in memory technologies has introduced Spin-Transfer Torque Magnetic RAM (STT-MRAM) as the most promising replacement for SRAMs in on-chip cache memories. Besides its lower leakage power, higher density, immunity to radiation-induced particles, and non-volatility, an unintentional bit flip during read operation, referred to as read disturbance error, is a severe reliability challenge in STT-MRAM caches. One major source of read disturbance error in STT-MRAM caches is simultaneous accesses to all tags for parallel comparison operation in a cache set, which has not been addressed in previous work. This paper first demonstrates that high read accesses to tag arrays extremely increase the read disturbance rate and then proposes a low-cost scheme, so-called Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison (3RSeT), to reduce the error rate by eliminating a significant portion of tag reads. 3RSeT proactively disables the tags that have no chance for hit, using low significant bits of the tags on each access request. Our evaluations using gem5 full-system cycle-accurate simulator show that 3RSeT reduces the read disturbance rate in the tag array by 71.8%, which results in 3.6x improvement in Mean Time To Failure. In addition, the energy consumption is reduced by 62.1% without compromising performance and with less than 0.4% area overhead. Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001 |
IEEE Trans. Computers | 2 |
| 2022 | Data block manipulation for error rate reduction in STT-MRAM based main memory
Nooshin Mahdavi, Farhad Razaghian, Hamed Farbeh |
J. Supercomput. | 3 |
| 2022 | LETHOR: a thermal-aware proactive routing algorithm for 3D NoCs with less entrance to hot regions
Maede Safari, Zahra Shirmohammadi, Nezam Rohbani, Hamed Farbeh |
J. Supercomput. | 4 |
| 2021 | ECC-United Cache: Maximizing Efficiency of Error Detection/Correction Codes in Associative Cache MemoriesabstractError Detection/Correction Codes (EDCs/ECCs) are the most conventional approaches to protect on-chip caches against radiation-induced soft errors. The overhead of EDCs/ECCs is a major concern and is of decisive importance when a higher protection capability is required to tolerate multiple adjacent bit errors (burst errors). This article proposes the ECC-United Cache (EUC) architecture to improve the efficiency of EDCs/ECCs in set-associative L1 caches. EUC architecture extends the data protection granularity from a single word to multiple words by exploiting the parallel cache lines access, which is inherently available in the cache. As compared with the conventional architecture, EUC can be configured to provide: 1) the same protection capability with a significantly lower overhead, 2) a significantly higher protection capability with the same number of check bits, or 3) a trade-off between the former two features. Simulation results show that, when configured to minimize the overhead, EUC reduces the number of check bits by 69 and 75 percent in data-cache and instruction-cache, respectively. When configured to maximize the protection capability, EUC provides fourfold higher burst error detection/correction capability. Moreover, EUC is orthogonal to previous protection schemes and they can be redesigned based on the EUC architecture to further improve their efficiency. Hamed Farbeh, Leila Delshadtehrani, Hyeonggyu Kim, Soontae Kim |
IEEE Trans. Computers | 1 |
| 2021 | TAMER: an adaptive task allocation method for aging reduction in multi-core embedded real-time systems
Faezeh Sadat Saadatmand, Nezam Rohbani, Farshad Baharvand, Hamed Farbeh |
J. Supercomput. | 4 |
| 2020 | A System-Level Framework for Analytical and Empirical Reliability Exploration of STT-MRAM CachesabstractSpin-transfer torque magnetic RAM (STT-MRAM) is known as the most promising replacement for static random access memory (SRAM) technology in large last-level cache memories (LLC). Despite its high density, nonvolatility, near-zero leakage power, and immunity to radiation as the major advantages, STT-MRAM-based cache memory suffers from high error rates mainly due to retention failure (RF), read disturbance, and write failure. Existing studies are limited to estimate the rate of only one or two of these error types for STT-MRAM cache. However, the overall vulnerability of STT-MRAM caches, whose estimation is a must to design cost-efficient reliable caches, has not been studied previously. In this paper, we propose a system-level framework for reliability exploration and characterization of errors' behavior in STT-MRAM caches. To this end, we formulate the cache vulnerability considering the intercorrelation of the error types including RF, read disturbance, and write failure as well as the dependency of error rates to workloads' behavior and process variations (PVs). Our analysis reveals that STT-MRAM cache vulnerability is highly workload-dependent and varies by orders of magnitude in different cache access patterns. Our analytical study also shows that this vulnerability divergence significantly increases by PVs in STT-MRAM cells. To take the effects of system workloads and PVs into account, we implement the error types in gem5 full-system simulator. The experimental results using a comprehensive set of multiprogrammed workloads from SPEC CPU2006 benchmark suite on a quad-core processor show that the total error rate in a shared STT-MRAM LLC varies by 32.0× for different workloads. A further 6.5× vulnerability variation is observed when considering PVs in the STT-MRAM cells. In addition, the contribution of each error type in total LLC vulnerability highly varies in different cache access patterns and moreover, error rates are differently affected by PVs. The proposed analytical and empirical studies can significantly help system architects for efficient utilization of error mitigation techniques and designing highly reliable and low-cost STT-MRAM LLCs. Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001 |
IEEE Trans. Reliab. | 2 |
| 2019 | ROBIN: incremental oblique interleaved ECC for reliability improvement in STT-MRAM cachesabstractSpin-Transfer Torque Magnetic RAM (STT-MRAM) is a promising alternative for SRAMs in on-chip cache memories. Besides all its advantages, high error rate in STT-MRAM is a major limiting factor for on-chip cache memories. In this paper, we first present a comprehensive analysis that reveals that the conventional Error-Correcting Codes (ECCs) lose their efficiency due to data-dependent error patterns, and then propose an efficient ECC configuration, so-called ROBIN, to improve the correction capability. The evaluations show that the inefficiency of conventional ECC increases the cache error rate by an average of 151.7% while ROBIN reduces this value by more than 28.6x. Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001 |
ASP-DAC | 2 |
| 2019 | Enhancing Reliability of STT-MRAM Caches by Eliminating Read Disturbance AccumulationabstractSpin-Transfer Torque Magnetic RAM (STT-MRAM) as one of the most promising replacements for SRAMs in on-chip cache memories benefits from higher density and scalability, near-zero leakage power, and non-volatility, but its reliability is threatened by high read disturbance error rate. Error-Correcting Codes (ECCs) are conventionally suggested to overcome the read disturbance errors in STT-MRAM caches. By employing aggressive ECCs and checking out a cache block on every read access, a high level of cache reliability is achieved. However, to minimize the cache access time in modern processors, all blocks in the target cache set are simultaneously read in parallel for tags comparison operation and only the requested block is sent out, if any, after checking its ECC. These extra cache block reads without checking their ECCs until requesting the blocks by the processor cause the accumulation of read disturbance error, which significantly degrades the cache reliability. In this paper, we first introduce and formulate the read disturbance accumulation phenomenon and reveal that this accumulation due to conventional parallel accesses of cache blocks significantly increases the cache error rate. Then, we propose a simple yet effective scheme, so-called Read Error Accumulation Preventer cache (REAP-cache) to completely eliminate the accumulation of read disturbances without compromising the cache performance. Our evaluations show that the proposed REAP-cache extends the cache Mean Time To Failure (MTTF) by 171x, while increases the cache area by less than 1% and energy consumption by only 2.7%. Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001 |
DATE | 2 |
| 2019 | TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM CachesabstractAs technology process node scales down, on-chip SRAM caches lose their efficiency because of their low scalability, high leakage power, and increasing rate of soft errors. Among emerging memory technologies,$Spin$-$Transfer\; Torque\; Magnetic\; RAM$(STT-MRAM) is known as the most promising replacement for SRAM-based cache memories. The main advantages of STT-MRAM are its non-volatility, near-zero leakage power, higher density, soft-error immunity, and higher scalability. Despite these advantages, high error rate in STT-MRAM cells due to$retention\; failure$,$write\; failure$, and$read\; disturbance$threatens the reliability of cache memories built upon STT-MRAM technology. The error rate is significantly increased in higher temperature, which further affects the reliability of STT-MRAM-based cache memories. The major source of heat generation and temperature increase in STT-MRAM cache memories is write operations, which are managed by cache$replacement\; policy$. To the best of our knowledge, none of previous studies have attempted to mitigate heat generation and high temperature of STT-MRAM cache memories using replacement policy. In this paper, we first analyze the cache behavior in conventional$Least$-$Recently\; Used$(LRU) replacement policy and demonstrate that the majority of consecutive write operations (more than 66 percent) are committed to adjacent cache blocks. These adjacent write operations cause accumulated heat and increased temperature, which significantly increase the cache error rate. To eliminate heat accumulation and the adjacency of consecutive writes, we propose a cache replacement policy, named$Thermal$-$Aware\; Least$-$Recently\; Written$(TA-LRW), to smoothly distribute the generated heat by conducting consecutive write operations in distant cache blocks. TA-LRW guarantees the distance of at least three blocks for each two consecutive write operations in an 8-way associative cache. This distant write scheme reduces the temperature-induced error rate by 94.8 percent, on average, compared with the conventional LRU policy, which results in 6.9x reduction in cache error rate. The implementation cost and complexity of TA-LRW is as low as$First$-$In,\; First$-$Out$(FIFO) policy while providing a near-LRU performance, having the advantages of both replacement policies. The significantly reduced error rate is achieved by imposing only 2.3 percent performance overhead compared with the LRU policy. Elham Cheshmikhani, Hamed Farbeh, Seyed Ghassem Miremadi, Hossein Asadi 0001 |
IEEE Trans. Computers | 2 |
| 2019 | RAW-Tag: Replicating in Altered Cache Ways for Correcting Multiple-Bit Errors in Tag ArrayabstractTag array in on-chip caches is one of the most vulnerable components to radiation-induced soft errors. Protecting the tag array in some processors is limited to error detection using the parity check, since the overheads of error correcting codes are not affordable in this component. State-of-the-art tag protection schemes combine the parity check with replication to provide error correction capability. Classifying these replication-based schemes into partial-replication and full-replication, the former offers a low overhead protection in which a large fraction of detectable errors remain uncorrectable, whereas the latter imposes a significant overhead to correct all of the errors. This paper proposes a low overhead full-replication scheme, so called Replicating in Altered Ways of Tag (RAW-Tag), to correct all detectable errors. RAW-Tag manipulates the cache replacement algorithm and keeps track of the incoming/evicting cache lines to not only provide a replica for all tags, but also eliminate the simultaneous susceptibility of both a tag and its replica to a single Multiple-Bit Upset (MBU). The simulation results show that RAW-Tag imposes no performance overhead and increases the energy consumption of L1 and L2 caches by only 6.6 and 0.3 percent, respectively, as compared with the baseline. Hamed Farbeh, Fereshte Mozafari, Masoume Zabihi, Seyed Ghassem Miremadi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2019 | Sleepy-LRU: extending the lifetime of non-volatile caches by reducing activity of age bits
Seyedeh Golsana Ghaemi, Iman Ahmadpour, Mehdi Ardebili, Hamed Farbeh |
J. Supercomput. | 4 |
| 2018 | ORIENT: Organized interleaved ECCs for new STT-MRAM cachesabstractSpin-Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising alternative to SRAM in cache memories. However, STT-MRAMs face with high probability of write errors due to its stochastic switching behavior. To correct the write errors, Error-Correcting Codes (ECCs) used in SRAM caches are conventionally employed. A cache line consists of several codewords and the data bits are selected in such a way that the maximum correction capability is provided based on the error patterns in SRAMs. However, the different write error patterns in STT-MRAM caches leads to inefficiency of conventional ECC configurations. In this paper, first we investigate the efficiency of ECC configurations and demonstrate that the vulnerability of codewords in a cache line varies by up to 17x. This variation means that, while some words are overprotected, some others are highly probable to experience uncorrectable errors. Then, we propose an ECC bit selection scheme, so-called ORIENT, to reduce the vulnerability variation of codewords to 1.4x. The simulation results show that conventional ECC configuration increases the write error rate by up to about 64.4% compared with the optimum ECC bit selection, whereas this value for ORIENT is only 4.5%. Zahra Azad, Hamed Farbeh, Amir Mahdi Hosseini Monazzah |
DATE | 2 |
| 2018 | SMARTag: Error Correction in Cache Tag Array by Exploiting Address LocalityabstractSoft errors in on-chip caches are the major cause of processors failure. Partitioning the cache into data and tag arrays, recent reports show that the vulnerability of the latter is as high as or even higher than that of the former. Although Error-Correcting Codes (ECCs) are widely used to protect the data array, their overheads are not affordable in the tag array and its protection is conventionally limited to parity code. In this paper, we propose Similarity-Managed Robust Tag (SMARTag) technique to provide the error correction capability in parity-protected tags. SMARTag exploits the inherent similarity between the upper parts of the tags in a cache set to share these parts between addresses and ECCs. Using SMARTag, the cache access time is intact since the ECC part is bypassed in normal cache operation and no extra memory is required since ECCs are stored in available tag space. The simulation results show that SMARTag is capable of correcting more than 98% of errors in the tag array, on average, and its energy consumption, area, and performance overhead is less than 0.2%. Seyedeh Golsana Ghaemi, Iman Ahmadpour, Mehdi Ardebili, Hamed Farbeh |
DATE | 4 |
| 2017 | WIPE: Wearout Informed Pattern Elimination to Improve the Endurance of NVM-based CachesabstractWith the recent development in Non-Volatile Memory (NVM) technologies, several studies have suggested using them as an alternative to SRAMs in on-chip caches. However, limited endurance of NVMs is a major challenge when employed in the caches. This paper proposes a data manipulation technique, so-called Wearout Informed Pattern Elimination (WIPE), to improve the endurance of NVM-based caches by reducing the activity of frequent data patterns. Simulation results show that WIPE improves the endurance by up to 93% with negligible overheads. Sina Asadi, Amir Mahdi Hosseini Monazzah, Hamed Farbeh, Seyed Ghassem Miremadi |
ASP-DAC | 3 |
| 2017 | An Efficient Protection Technique for Last Level STT-RAM Caches in Multi-Core ProcessorsabstractDue to serious problems of SRAM-based caches in nano-scale technologies, researchers seek for new alternatives. Among the existing options, STT-RAM seems to be the most promising alternative. With high density and negligible leakage power, STT-RAMs open a new doorto respond to future demands of multi-core systems, i.e., large on-chip caches. However, several problems in STT-RAMs should be overcome to make it applicable in on-chip caches. High probability of write error due to stochastic switching is a major problem in STT-RAMs. Conventional Error-Correcting Codes (ECCs) impose significant area and energy consumption overheads to protect STT-RAM caches. These overheads in multi-core processors with large last-level caches are not affordable. In this paper, we propose Asymmetry-Aware Protection Technique (A2PT) to efficiently protect the STT-RAM caches. A2PT benefits from error rate asymmetry of STT-RAM write operations to provide the required level of cache protection with significantly lower overheads. Compared with the conventional ECC configuration, the evaluation results show that A2PT reduces the area and energy consumption overheads by about 42 and 50 percent, respectively, while providing the same level of protection. Moreover, A2PT decreases the number of bit switching in write operations by 28 percent, which leads to about 25 percent saving in write energy consumption. Zahra Azad, Hamed Farbeh, Amir Mahdi Hosseini Monazzah, Seyed Ghassem Miremadi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Floating-ECC: Dynamic Repositioning of Error Correcting Code Bits for Extending the Lifetime of STT-RAM CachesabstractSpin-Transfer Torque RAM (STT-RAM) is a promising alternative to SRAM for implementing on-chip L2 and L3 caches. One of the most critical challenges in STT-RAM is reliability due to limited write endurance, which results in insufficient lifetime, as well as various types of errors. Previous studies have focused on either presenting various cache architectures/management techniques to improve the lifetime of STT-RAM caches or utilizing different Error Correcting Codes (ECCs) to protect against the permanent and transient errors. However, there is no quantitative analysis in the literature to determine the impact of ECCs on the lifetime of the STT-RAM caches. This paper formulates this impact and demonstrates that ECCs shorten the lifetime of STT-RAM cache lines by more than 50 percent due to ECCs high write activity. Then, we propose the Floating-ECC architecture for increasing the lifetime of the STT-RAM caches. The main idea is to evenly distribute the ECC write activity over all bits of cache lines by periodically relocating the ECC bits inside the cache lines. The simulation results for the most conventional ECC scheme, i.e., interleaved Single Error Correction-Double Error Detection (SEC-DED), show that Floating-ECC increases the lifetime of L2 and L3 caches by more than 318 percent and 254 percent, respectively. Hamed Farbeh, Hyeonggyu Kim, Seyed Ghassem Miremadi, Soontae Kim |
IEEE Trans. Computers | 1 |
| 2016 | A Cache-Assisted Scratchpad Memory for Multiple-Bit-Error CorrectionabstractScratchpad memory (SPM) is widely used in modern embedded processors to overcome the limitations of cache memory. The high vulnerability of SPM to soft errors, however, limits its usage in safety-critical applications. This paper proposes an efficient fault-tolerant scheme, called cache-assisted duplicated SPM (CADS), to protect SPM against soft errors. The main aim of CADS is to utilize cache memory to provide a replica for SPM lines. Using cache memory, CADS is able to guarantee a full duplication of all SPM lines. We also further enhance the proposed scheme by presenting buffered CADS (BCADS) that significantly improves the CADS energy efficiency. BCADS is compared with two well-known duplication schemes as well as single-error correction scheme. The comparison results reveal that: 1) BCADS imposes a 13.6% less energy-delay product (EDP) overhead than the duplication schemes and it does not require to modify the SPM manager and target application and 2) in comparison with the conventional single-error correction double-error detection (SEC-DED) scheme, BCADS provides a significantly higher error correction capability by correcting up to 4-b burst errors using a low-cost 4-b interleaved parity code. Moreover, the area overhead for error correction and the performance overhead of BCADS are negligible (less than 1%), whereas the area and performance overheads are 21.9% and 6.1% for SEC-DED, respectively. Furthermore, BCADS imposes about a 10.7% lower EDP overhead compared with the SEC-DED scheme. Hamed Farbeh, Nooshin Sadat Mirzadeh, Nahid Farhady Ghalaty, Seyed Ghassem Miremadi, Mahdi Fazeli, Hossein Asadi 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | In-Scratchpad Memory Replication: Protecting Scratchpad Memories in Multicore Embedded Systems against Soft ErrorsabstractScratchpad memories (SPMs) are widely employed in multicore embedded processors. Reliability is one of the major constraints in the embedded processor design, which is threatened with the increasing susceptibility of memory cells to multiple-bit upsets (MBUs) due to continuous technology down-scaling. This article proposes a low-cost and efficient data replication mechanism, called In-Scratchpad Memory Replication (ISMR), to correct MBUs in SPMs of multicore embedded processors. The main feature of ISMR is a smart controller, called Replication Management Unit (RMU), which is responsible for dynamically analyzing the activity of the SPM blocks at runtime and efficiently replicating the vulnerable SPM blocks into currently inactive SPM blocks. RMU exploits a 2-bit tag for each SPM block, where the value of each tag is determined by RMU according to the SPM access pattern. Accordingly, the proposed mechanism guarantees the replication of all vulnerable SPM blocks to provide error correction without decreasing the SPM utilization. To detect errors in SPM blocks, ISMR uses a 2-bit interleaved-parity code. As compared with the previous E-RAID 1 mechanism, the simulation results illustrate that for an 8-core embedded processor, the ISMR mechanism experiences 81% less energy consumption overhead and 48% less performance overhead. Leila Delshadtehrani, Hamed Farbeh, Seyed Ghassem Miremadi |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | PSP-Cache: A low-cost fault-tolerant cache memory architectureabstractCache memories constitute a large fraction of processor chip area and are highly vulnerable to soft errors caused by energetic particles. To protect these memories, most of the modern processors employ Error Detection Codes (EDCs) or Error Correction Codes (ECCs). EDCs/ECCs impose significant overheads in terms of area and energy; these overheads increase as a function of interleaving EDCs/ECCs to detect/correct multiple errors. This paper proposes a new cache architecture to minimize the area and energy overheads of EDCs/ECCs in set-associative L1-caches. Simulation results for a 4-way set-associative cache show that the proposed architecture reduces both the area and static power overheads of parity code by about 75% and the dynamic energy overhead by about 73% in comparison to conventional cache architecture. These reduction figures are about 68% and about 66%, respectively, for SEC-DED code. The above reductions are achieved without affecting the error coverage. Hamed Farbeh, Seyed Ghassem Miremadi |
DATE | 1 |
| 2013 | FTSPM: A Fault-Tolerant ScratchPad MemoryabstractScratchPad Memory (SPM) is an important part of most modern embedded processors. The use of embedded processors in safety-critical applications implies including fault tolerance in the design of SPM. This paper proposes a method, called FTSPM, which integrates a multi-priority mapping algorithm with a hybrid SPM structure. The proposed structure divides SPM into three parts: 1) a part is equipped with Non-Volatile Memory (NVM) which is immune against soft errors, 2) a part is equipped with Error-Correcting Code, and 3) a part is equipped with parity. The proposed mapping algorithm is responsible to distribute the program blocks among the above three parts with regards to their vulnerability level. The simulation results demonstrate that the FTSPM reduces the SPM vulnerability by about 7x in comparison to a pure SRAM-based SPM. In addition, the dynamic energy consumption of the proposed method is 77% and 47% less than that of a pure NVM-based SPM and a pure SRAM-based SPM, respectively. Amir Mahdi Hosseini Monazzah, Hamed Farbeh, Seyed Ghassem Miremadi, Mahdi Fazeli, Hossein Asadi 0001 |
DSN | 2 |
| 2011 | Low Cost Concurrent Error Detection for On-Chip Memory Based Embedded ProcessorsabstractThis paper proposes an efficient concurrent error detection method using control flow checking for embedded processors. The proposed method is based on the co-operation of two hardware modules: 1) an on-chip hardware component to detect branch instructions and generate signatures for the running program, and 2) an external watchdog processor to compare runtime signatures and branch addresses with the information extracted offline. The proposed method is implemented on an embedded processor core and is evaluated by a simulation based statistical fault injection approach where faults are injected into cache and main memory. Experimental results show that the proposed method detects more than 96.7% of all errors with only 2.6% overhead in area and less than 1% increase in power consumption. Furthermore, this technique imposes almost no performance degradation. Faramarz Khosravi, Hamed Farbeh, Mahdi Fazeli, Seyed Ghassem Miremadi |
EUC | 2 |