Elham Cheshmikhani

dblp:173/7261 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
4since 2021 · last 2025
0000-0003-3737-683XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 66% Storage systems · 25% Hardware reliability and fault tolerance · 6%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
1.832025
An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable Cache · IEEE Trans. Dependable Secur. Comput. 2025
3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison · IEEE Trans. Computers 2022
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Storage systems › storage reliability › durability
retention failure
1.022022
CoPA: Cold Page Awakening to Overcome Retention Failures in STT-MRAM Based I/O Buffers · IEEE Trans. Parallel Distributed Syst. 2022
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Memory systems › non-volatile memory › magnetic random access memory › STT-MRAM
STT-MRAM cache
1.022022
3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison · IEEE Trans. Computers 2022
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Memory systems › cache design
cache tag storage
0.912025
An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable Cache · IEEE Trans. Dependable Secur. Comput. 2025
Memory systems
non-volatile memory
0.612022
CoPA: Cold Page Awakening to Overcome Retention Failures in STT-MRAM Based I/O Buffers · IEEE Trans. Parallel Distributed Syst. 2022
Storage systems › flash and SSD › flash memory › flash storage
read disturbance error
0.612022
3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison · IEEE Trans. Computers 2022
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.612022
CoPA: Cold Page Awakening to Overcome Retention Failures in STT-MRAM Based I/O Buffers · IEEE Trans. Parallel Distributed Syst. 2022
Hardware reliability and fault tolerance
soft errors
0.532025
An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable Cache · IEEE Trans. Dependable Secur. Comput. 2025
3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison · IEEE Trans. Computers 2022
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Memory systems › cache management
cache replacement
0.412019
TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches · IEEE Trans. Computers 2019
Energy-efficient computing › power management › memory power management
cache energy reduction
0.312025
An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable Cache · IEEE Trans. Dependable Secur. Comput. 2025
Memory systems
DRAM
0.212022
CoPA: Cold Page Awakening to Overcome Retention Failures in STT-MRAM Based I/O Buffers · IEEE Trans. Parallel Distributed Syst. 2022
Memory systems › DRAM
rowhammer
0.212022
3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison · IEEE Trans. Computers 2022

Methods — techniques the papers use, named apart from their topics

convex optimization · 0.9analytical modeling · 0.9distant refreshing · 0.6cycle-accurate simulation · 0.6trace-driven simulation · 0.4
YearPublicationVenuePosition
2025 An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable Cache
abstract
Associative cache memory plays a decisive role in enhancing the performance and energy consumption of modern processors. Meanwhile, by occupying more than half of the processor chip area, cache memory is susceptible to transient and permanent faults, threatening the system's dependability. As the onlyhardware-managedmemory module in the system, the tag array of the caches is the most critical and active component contributing a large fraction of energy consumption and error occurrence.Tag partitioningis a widespread approach for both tag energy consumption reduction and reliability enhancement. This approach splits the tag comparison operation into two steps, and only the tags whoseklower order bits are matched with that of the input address in the first step are activated for comparing their remaining higher order bits in the second step. The key decision parameter for tag partitioning is properly adjusting the tag-splitting point (k) to achieve the maximum reduction in the number of reads. This parameter has been intuitively, randomly, or experimentally selected in the existing studies without any justification. Even for an appropriate selection of this parameter via extensive experiments, its sensitivity to various cache configuration parameters makes it ad-hoc and not extendable to other scenarios. In this paper, we analytically illustrate that selecting an inappropriately large or small value for the tag-splitting point significantly downgrades the efficiency of tag partitioning and then formulate this parameter to determine its optimum value. As a function of cache configuration parameters, the proposed formulation is proven to be convex and differentiable for determining the optimum splitting point, besides its ability to accurately report the degree of the tag partitioning efficiency for any splitting point and configuration parameters. To approve the correctness and accuracy of the proposed formulation, we experimentally investigate the tag partitioning efficiency and optimum splitting point for a wide range of cache configurations and demonstrate a very close matching between the two. The proposed formulation is a guarantee for the designers and researchers to instantaneously determine the optimum tag-splitting point and calculate the read reduction of tag partitioning.
Elham Cheshmikhani, Hamed Farbeh
IEEE Trans. Dependable Secur. Comput.1
2025 A Reliability-Aware Replacement Policy for STT-MRAM Caches in Server-Class Processors
abstract
Spin-transfer torque magnetic RAM(STT-MRAM) has several advantages over conventional SRAM technology in on-chip caches, such as low leakage, soft error tolerance, and high density and scalability. These advantages make it the most promising nonvolatile memory for SRAM replacement. However, STT-MRAM cache memory faces two main reliability challenges in emergingnanoscaletechnology nodes, i.e.,retention failureandread disturbance. Because of the lower data access rate in thelast-level caches(LLCs) compared to the higher cache levels, as well as the higher contribution of read accesses, these two reliability challenges have become severe in the LLCs of server-class processors. The existing approaches to overcome these challenges impose significant area and performance overhead or adversely affect the other failure types. In this article, we first investigate the parameters that affect the reliability of STT-MRAM-based LLCs due to the retention failure and read disturbance. Our investigation shows that 1) the duration ofdead dirty blocksis the main contributor to the retention failure rate of STT-MRAM LLCs while 2) the high number ofriskyreads, i.e., those that can affect the cache reliability, in the dirty blocks is the main contributor to the read disturbance. Based on these observations, we propose a simple yet effective cache replacement policy, calledRetentionfailureandread disturbancereduction, to decrease the length of dirty intervals and the number of reads, which results in a significant reduction in retention failure and read disturbance rate. Our evaluations demonstrate that the proposed replacement policy cuts down the probability of retention failure and the number of risky reads per dirty block by 56% and 61%, respectively. The area overhead of this scheme is negligible (0.2%) with no adverse effect on the energy consumption.
Abdollah Mohammadi, Elham Cheshmikhani, Hossein Asadi 0001
IEEE Trans. Reliab.2
2022 3RSeT: Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison
abstract
Recent development in memory technologies has introduced Spin-Transfer Torque Magnetic RAM (STT-MRAM) as the most promising replacement for SRAMs in on-chip cache memories. Besides its lower leakage power, higher density, immunity to radiation-induced particles, and non-volatility, an unintentional bit flip during read operation, referred to as read disturbance error, is a severe reliability challenge in STT-MRAM caches. One major source of read disturbance error in STT-MRAM caches is simultaneous accesses to all tags for parallel comparison operation in a cache set, which has not been addressed in previous work. This paper first demonstrates that high read accesses to tag arrays extremely increase the read disturbance rate and then proposes a low-cost scheme, so-called Read Disturbance Rate Reduction in STT-MRAM Caches by Selective Tag Comparison (3RSeT), to reduce the error rate by eliminating a significant portion of tag reads. 3RSeT proactively disables the tags that have no chance for hit, using low significant bits of the tags on each access request. Our evaluations using gem5 full-system cycle-accurate simulator show that 3RSeT reduces the read disturbance rate in the tag array by 71.8%, which results in 3.6x improvement in Mean Time To Failure. In addition, the energy consumption is reduced by 62.1% without compromising performance and with less than 0.4% area overhead.
Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001
IEEE Trans. Computers1
2022 CoPA: Cold Page Awakening to Overcome Retention Failures in STT-MRAM Based I/O Buffers
abstract
Performance and reliability are two prominent factors in the design of data storage systems. To achieve higher performance, recently storage system designers use$Dynamic$$RAM$(DRAM)-based buffers. The volatility of DRAM brings up the possibility of data loss and data inconsistency. Thus, a part of the main storage is conventionally used as the journal area to be able of recovering unflushed data pages in the case of power failure. Moreover, periodically flushing buffered data pages to the main storage is a common mechanism to preserve a high level of reliability. This scheme, however, leads to a considerable increase in storage write traffic, which adversely affects the performance. To address this shortcoming, recent studies offer a small$Non-Volatile$$Memory$(NVM) as the$Persistent$$Journal$$Area$(PJA) along with DRAM as an efficient approach to overcome DRAM vulnerability against power failure while effectively reducing storage write traffic. This approach, named$NVM-Backed$$Buffer$(NVB-Buffer), features from advantages of NVMs and addresses DRAM shortcomings. In this article, we employ the most promising technologies for PJA among the emerging technologies, which is$Spin-Transfer$$Torque$$Magnetic$$Random$$Access$$Memory$(STT-MRAM) to meet the requirements of efficient PJA by providing high endurance, non-volatility, and DRAM-like latency. Despite these advantages, STT-MRAM faces major reliability challenges, i.e.,Retention Failure,Read Disturbance, andWrite Failure, which havenotbeen addressed in previously suggested NVB-Buffers. In this article, we first demonstrate that the retention failure is the dominant source of errors in NVB-Buffers as it suffers from long and unpredictable page idle intervals (i.e., the time interval between two consecutive accesses to a PJA page). Then, we propose a novel NVB-Buffer management scheme, named,$\underline{Co}ld$$\underline{P}age$$\underline{A}wakening$(CoPA), which predictably reduces the idle time of PJA pages. To this aim, CoPA employs$Distant$$Refreshing$to periodically overwrite the vulnerable PJA page contents by opportunistically using their replica in DRAM-based buffer. We compare CoPA with the state-of-the-art schemes over several well-known storage workloads based on physical journaling. Our evaluations show that CoPA significantly reduces the maximum page idle time, which leads to three orders of magnitude lower failure rate with negligible performance degradation (1.1%) and memory overhead (1.2%).
Mostafa Hadizadeh, Elham Cheshmikhani, Maysam Rahmanpour, Onur Mutlu, Hossein Asadi 0001
IEEE Trans. Parallel Distributed Syst.2
2020 STAIR: High Reliable STT-MRAM Aware Multi-Level I/O Cache Architecture by Adaptive ECC Allocation
abstract
Hybrid Multi-Level Cache Architectures (HCAs) are promising solutions for the growing need of high-performance and cost-efficient data storage systems. HCAs employ a high endurable memory as the first-level cache and a Solid-State Drive (SSD) as the second-level cache. Spin-Transfer Torque Magnetic RAM (STT-MRAM) is one of the most promising candidates for the first-level cache of HCAs because of its high endurance and DRAM-comparable performance along with non-volatility. However, STT-MRAM faces with three major reliability challenges named Read Disturbance, Write Failure, and Retention Failure. To provide a reliable HCA, the reliability challenges of STT-MRAM should be carefully addressed. To this end, this paper first makes a careful distinction between clean and dirty pages to classify and prioritize their different vulnerabilities. Then, we investigate the distribution of more vulnerable pages in the first-level cache of HCAs over 17 storage workloads. Our observations show that the protection overhead can be significantly reduced by adjusting the protection level of data pages based on their vulnerability. To this aim, we propose a STT-MRAM Aware Multi-Level I/O Cache Architecture (STAIR) to improve HCA reliability by dynamically generating extra strong Error- Correction Codes (ECCs) for the dirty data pages. STAIR adaptively allocates under-utilized parts of the first-level cache to store these extra ECCs. Our evaluations show that STAIR decreases the data loss probability by five orders of magnitude, on average, with negligible performance overhead (0.12% hit ratio reduction in the worst case) and 1.56% memory overhead for the cache controller.
Mostafa Hadizadeh, Elham Cheshmikhani, Hossein Asadi 0001
DATE2
2020 A System-Level Framework for Analytical and Empirical Reliability Exploration of STT-MRAM Caches
abstract
Spin-transfer torque magnetic RAM (STT-MRAM) is known as the most promising replacement for static random access memory (SRAM) technology in large last-level cache memories (LLC). Despite its high density, nonvolatility, near-zero leakage power, and immunity to radiation as the major advantages, STT-MRAM-based cache memory suffers from high error rates mainly due to retention failure (RF), read disturbance, and write failure. Existing studies are limited to estimate the rate of only one or two of these error types for STT-MRAM cache. However, the overall vulnerability of STT-MRAM caches, whose estimation is a must to design cost-efficient reliable caches, has not been studied previously. In this paper, we propose a system-level framework for reliability exploration and characterization of errors' behavior in STT-MRAM caches. To this end, we formulate the cache vulnerability considering the intercorrelation of the error types including RF, read disturbance, and write failure as well as the dependency of error rates to workloads' behavior and process variations (PVs). Our analysis reveals that STT-MRAM cache vulnerability is highly workload-dependent and varies by orders of magnitude in different cache access patterns. Our analytical study also shows that this vulnerability divergence significantly increases by PVs in STT-MRAM cells. To take the effects of system workloads and PVs into account, we implement the error types in gem5 full-system simulator. The experimental results using a comprehensive set of multiprogrammed workloads from SPEC CPU2006 benchmark suite on a quad-core processor show that the total error rate in a shared STT-MRAM LLC varies by 32.0× for different workloads. A further 6.5× vulnerability variation is observed when considering PVs in the STT-MRAM cells. In addition, the contribution of each error type in total LLC vulnerability highly varies in different cache access patterns and moreover, error rates are differently affected by PVs. The proposed analytical and empirical studies can significantly help system architects for efficient utilization of error mitigation techniques and designing highly reliable and low-cost STT-MRAM LLCs.
Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001
IEEE Trans. Reliab.1
2019 ROBIN: incremental oblique interleaved ECC for reliability improvement in STT-MRAM caches
abstract
Spin-Transfer Torque Magnetic RAM (STT-MRAM) is a promising alternative for SRAMs in on-chip cache memories. Besides all its advantages, high error rate in STT-MRAM is a major limiting factor for on-chip cache memories. In this paper, we first present a comprehensive analysis that reveals that the conventional Error-Correcting Codes (ECCs) lose their efficiency due to data-dependent error patterns, and then propose an efficient ECC configuration, so-called ROBIN, to improve the correction capability. The evaluations show that the inefficiency of conventional ECC increases the cache error rate by an average of 151.7% while ROBIN reduces this value by more than 28.6x.
Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001
ASP-DAC1
2019 Enhancing Reliability of STT-MRAM Caches by Eliminating Read Disturbance Accumulation
abstract
Spin-Transfer Torque Magnetic RAM (STT-MRAM) as one of the most promising replacements for SRAMs in on-chip cache memories benefits from higher density and scalability, near-zero leakage power, and non-volatility, but its reliability is threatened by high read disturbance error rate. Error-Correcting Codes (ECCs) are conventionally suggested to overcome the read disturbance errors in STT-MRAM caches. By employing aggressive ECCs and checking out a cache block on every read access, a high level of cache reliability is achieved. However, to minimize the cache access time in modern processors, all blocks in the target cache set are simultaneously read in parallel for tags comparison operation and only the requested block is sent out, if any, after checking its ECC. These extra cache block reads without checking their ECCs until requesting the blocks by the processor cause the accumulation of read disturbance error, which significantly degrades the cache reliability. In this paper, we first introduce and formulate the read disturbance accumulation phenomenon and reveal that this accumulation due to conventional parallel accesses of cache blocks significantly increases the cache error rate. Then, we propose a simple yet effective scheme, so-called Read Error Accumulation Preventer cache (REAP-cache) to completely eliminate the accumulation of read disturbances without compromising the cache performance. Our evaluations show that the proposed REAP-cache extends the cache Mean Time To Failure (MTTF) by 171x, while increases the cache area by less than 1% and energy consumption by only 2.7%.
Elham Cheshmikhani, Hamed Farbeh, Hossein Asadi 0001
DATE1
2019 TA-LRW: A Replacement Policy for Error Rate Reduction in STT-MRAM Caches
abstract
As technology process node scales down, on-chip SRAM caches lose their efficiency because of their low scalability, high leakage power, and increasing rate of soft errors. Among emerging memory technologies,$Spin$-$Transfer\; Torque\; Magnetic\; RAM$(STT-MRAM) is known as the most promising replacement for SRAM-based cache memories. The main advantages of STT-MRAM are its non-volatility, near-zero leakage power, higher density, soft-error immunity, and higher scalability. Despite these advantages, high error rate in STT-MRAM cells due to$retention\; failure$,$write\; failure$, and$read\; disturbance$threatens the reliability of cache memories built upon STT-MRAM technology. The error rate is significantly increased in higher temperature, which further affects the reliability of STT-MRAM-based cache memories. The major source of heat generation and temperature increase in STT-MRAM cache memories is write operations, which are managed by cache$replacement\; policy$. To the best of our knowledge, none of previous studies have attempted to mitigate heat generation and high temperature of STT-MRAM cache memories using replacement policy. In this paper, we first analyze the cache behavior in conventional$Least$-$Recently\; Used$(LRU) replacement policy and demonstrate that the majority of consecutive write operations (more than 66 percent) are committed to adjacent cache blocks. These adjacent write operations cause accumulated heat and increased temperature, which significantly increase the cache error rate. To eliminate heat accumulation and the adjacency of consecutive writes, we propose a cache replacement policy, named$Thermal$-$Aware\; Least$-$Recently\; Written$(TA-LRW), to smoothly distribute the generated heat by conducting consecutive write operations in distant cache blocks. TA-LRW guarantees the distance of at least three blocks for each two consecutive write operations in an 8-way associative cache. This distant write scheme reduces the temperature-induced error rate by 94.8 percent, on average, compared with the conventional LRU policy, which results in 6.9x reduction in cache error rate. The implementation cost and complexity of TA-LRW is as low as$First$-$In,\; First$-$Out$(FIFO) policy while providing a near-LRU performance, having the advantages of both replacement policies. The significantly reduced error rate is achieved by imposing only 2.3 percent performance overhead compared with the LRU policy.
Elham Cheshmikhani, Hamed Farbeh, Seyed Ghassem Miremadi, Hossein Asadi 0001
IEEE Trans. Computers1
2016 Accelerating Dynamic Fault Tree Analysis Based on Stochastic Logic Utilizing GPGPUs
abstract
This paper demonstrates on speeding up an accurate analysis of fault trees using stochastic logic through GPGPUs. Actually, probability models of dynamic gates and new accurate models for different combinations of cold spare gate e.g., two cold spare gates with a share spare and a cold spare gate with more than one spare inputs are developed in this paper. Experimental results show that on average, the proposed analysis method is 235 times faster than CPU simulation time. Moreover, proposing new stochastic models results accuracy and simplicity as additional advantages of the proposed method.
Elham Cheshmikhani, Hamid R. Zarandi
PDP1