Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Majid Jalili 0001

dblp:123/4090-1 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
1since 2021 · last 2022
0000-0001-5166-4788ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 7 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 84% Hardware reliability and fault tolerance · 16%
Network and information security
1 paper
Hardware security and side channels · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
0.932018
Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance Levels · ACM Trans. Archit. Code Optim. 2018
Endurance-Aware Security Enhancement in Non-Volatile Memories Using Compression and Selective Encryption · IEEE Trans. Computers 2017
BLESS: a simple and efficient scheme for prolonging PCM lifetime · DAC 2016
Memory systems › non-volatile memory
phase change memory
0.622018
Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance Levels · ACM Trans. Archit. Code Optim. 2018
BLESS: a simple and efficient scheme for prolonging PCM lifetime · DAC 2016
Memory systems
cache
0.612022
Reducing Load Latency with Cache Level Prediction · HPCA 2022
Memory systems
memory access latency
0.612022
Reducing Load Latency with Cache Level Prediction · HPCA 2022
Hardware reliability and fault tolerance
memory reliability
0.312018
Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance Levels · ACM Trans. Archit. Code Optim. 2018
Memory systems › non-volatile memory › multi-level cell
multi-level cell memory
0.312018
Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance Levels · ACM Trans. Archit. Code Optim. 2018
Memory systems › non-volatile memory › phase change memory
multi-level cell PCM
0.312018
Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance Levels · ACM Trans. Archit. Code Optim. 2018
Hardware reliability and fault tolerance › soft errors
soft error mitigation
0.312018
Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance Levels · ACM Trans. Archit. Code Optim. 2018
Memory systems › cache
prefetching
0.212022
Reducing Load Latency with Cache Level Prediction · HPCA 2022
Hardware security and side channels › physical attacks
probing attack
0.112017
Endurance-Aware Security Enhancement in Non-Volatile Memories Using Compression and Selective Encryption · IEEE Trans. Computers 2017
Memory systems › non-volatile memory › write reliability
write endurance
0.112016
BLESS: a simple and efficient scheme for prolonging PCM lifetime · DAC 2016

Methods — techniques the papers use, named apart from their topics

wear leveling · 0.6simulation · 0.6selective encryption · 0.6compression · 0.6relaxed write and read · 0.3error correction metadata · 0.3FPC compression · 0.3multi-level cell error recovery · 0.2byte-level shifting · 0.2
YearPublicationVenuePosition
2022 Reducing Load Latency with Cache Level Prediction
abstract
High load latency that results from deep cache hierarchies and relatively slow main memory is an important limiter of single-thread performance. Data prefetch helps reduce this latency by fetching data up the hierarchy before it is requested by load instructions. However, data prefetching has shown to be imperfect in many situations. We propose cache-level prediction to complement prefetchers. Our method predicts which memory hierarchy level a load will access allowing the memory loads to start earlier, and thereby saves many cycles. The predictor provides high prediction accuracy at the cost of just one cycle added latency to L1 misses. Level prediction reduces the memory access latency by 20% on average, and provides speedup of 10.3% over a conventional baseline, and 6.1% over a boosted baseline on generic, graph, and HPC applications.
Majid Jalili 0001, Mattan Erez
HPCA1
2018 Improving MLC PCM Performance through Relaxed Write and Read for Intermediate Resistance Levels
abstract
Phase Change Memory (PCM) is one of the most promising candidates to be used at the main memory level of the memory hierarchy due to poor scalability, considerable leakage power, and high cost/bit of DRAM. PCM is a new resistive memory that is capable of storing data based on resistance values. The wide resistance range of PCM allows for storing multiple bits per cell (MLC) rather than a single bit per cell (SLC). Unfortunately, higher density of MLC PCM comes at the expense of longer read/write latency, higher soft error rate, higher energy consumption, and earlier wearout compared to the SLC PCM. Some studies suggest removing the most error-prone level to mitigate soft error and write latency of MLC PCM, hence introducing a less dense memory called Tri-Level memory. Another scheme, called M-Metric, proposes a new read metric to address the soft error problem in MLC PCM. In order to deal with the limited lifetime of PCM, some extra storage per memory line is required to correct permanent hard errors (stuck-at faults). Since the extra storage is used only when permanent faults occur, it has a low utilization for a long time before hard errors start to occur. In this article, we utilize the extra storage to improve the read/write latency in a 2-bit MLC PCM using a relaxation scheme for reading and writing the cells for intermediate resistance levels. More specifically, we combine the most time-consuming levels (intermediate resistance levels) to reduce the number of resistance levels (making a Tri-Level PCM) and therefore improve write latency. We then store some error correction metadata in the extra storage section to successfully retrieve the exact data values in the read operation. We also modify the Tri-Level PCM cell to increase its read latency when the M-Metric scheme is used. Evaluation results show that the proposed scheme improves read latency by 57.2%, write latency by 56.1%, and overall system performance (IPC) by 26.9% over the baseline. It is noteworthy that combining the proposed scheme and FPC compression method improves read latency by 75.2%, write latency by 67%, and overall system performance (IPC) by 37.4%.
Saeed Rashidi, Majid Jalili 0001, Hamid Sarbazi-Azad
ACM Trans. Archit. Code Optim.2
2018 Express Read in MLC Phase Change Memories
abstract
In the era of big data, the capability of computer systems must be enhanced to support 2.5 quintillion byte/day data delivery. Among the components of a computer system, main memory has a great impact on overall system performance. DRAM technology has been used over the past four decades to build main memories. However, the scalability of DRAM technology has faced serious challenges. To keep pace with the ever-increasing demand for larger main memory, some new alternative technologies have been introduced. Phase change memory (PCM) is considered as one of such technologies for substituting DRAM. PCM offers some noteworthy properties such as low static power consumption, nonvolatility, and capability of storing more than one bit per cell (multilevel cell, or MLC). However, the short lifetime and long access latency of PCM (specifically MLC PCM) require feasible and efficient solutions. In this article, based on the observation that applications access a significant number of read-friendly data blocks, we propose Express Read to prevent the MLC PCM read circuit to spend unnecessary time sensing the cells of a memory block. A read-friendly data block (RFDB) is composed of only “11” and “00” bit pairs, and thus upon sensing the most significant bit of a cell, the read operation can be early terminated to reduce the MLC read time and power consumption. Moreover, we increase the number of RFDBs using two simple techniques to better exploit the benefits of Express Read. Results obtained from full-system simulation near 6% performance improvement and 21% energy gain, on average, over the baseline system.
Majid Jalili 0001, Hamid Sarbazi-Azad
ACM Trans. Design Autom. Electr. Syst.1
2017 Data Block Partitioning for Recovering Stuck-at Faults in PCMs
abstract
Main burdens to the DRAM scalability are leakage and charge storage restrictions. Phase Change Memory (PCM) is being known as a promising candidate for the replacement of DRAM among competitive non-volatile memories. However, this memory suffers from low cell reliability due to limited write endurance. This problem can lead to some memory cells permanently stuck at either '0' or '1'. Therefore, a robust error recovery scheme is needed to overcome this problem and recover from hard errors. State-of-the-art solutions apply error correction and recovery techniques at inter- line or intra-line level. Precisely, they can improve PCM endurance either by remapping failed lines to spares (in inter-line level schemes) or by using data-block partitioning and bit- inversion scheme (in intra-line level schemes). Although techniques of the latter type are effective, proper partitioning of data blocks and spreading out faults across different groups are required. In this paper, we propose and evaluate a novel intra-line level scheme that statically partition a data-block into some groups and efficiently recover multi-bit stuck-at faults per partition. This method benefits from the advantage of a simple shifting mechanism in order to increase the chance of storing data in presence of failed cells. Evaluation results for multi- threaded workloads show enhancement in the number of recoverable failures and improvement of lifetime over existing techniques.
Marjan Asadinia, Majid Jalili 0001, Hamid Sarbazi-Azad
NAS2
2017 Endurance-Aware Security Enhancement in Non-Volatile Memories Using Compression and Selective Encryption
abstract
Emerging non-volatile memories (NVMs) are notable candidates for replacing traditional DRAMs. Although NVMs are scalable, dissipate lower power, and do not require refreshes, they face new challenges including shorter lifetime and security issues. Efforts toward securing the NVMs against probe attacks pose a serious downside in terms of lifetime. Cryptography algorithms increase the information density of data blocks and consequently handicap the existing lifetime enhancement solutions like Flip-N-Write. In this paper, based on the insight that compression can relax the constraints of lifetime-security trade-off, we propose CryptoComp, an architecture that, taking the advantage of block size reduction after compression, aims to enhance the memory system lifetime and security. Our idea is to limit the avalanche effect caused by encryption algorithms in a lower space through compression and selective encryption. This way, for highly compressible data blocks, we follow a fully-encryption approach while for poorly compressible data blocks, we rely on a non-deterministic fine-grain selective-encryption mechanism. Additionally, a simple and block-oriented wear-leveling scheme is presented to fairly distribute the bit flips on memory cells. Our experimental results show 3.59$\times$and 3.66$\times$lifetime improvements over two state-of-the-art schemes, DEUCE and i-NVMM, while imposing a negligible performance degradation of 2.1 percent, on average.
Majid Jalili 0001, Hamid Sarbazi-Azad
IEEE Trans. Computers1
2016 BLESS: a simple and efficient scheme for prolonging PCM lifetime
abstract
Limited endurance problem and low cell reliability are main challenges of phase change memory (PCM) as an alternative to DRAM. To further prolong the lifetime of a PCM device, there exist a number of techniques that can be grouped in two categories: 1) reducing the write rate to PCM cells, and 2) handling cell failures when faults occur. Our experiments confirm that during write operations, an extensive non-uniformity in bit ips is exhibited. To reduce this non-uniformity, we present byte-level shifting scheme (BLESS) which reduces write pressure over hot cells of blocks. Additionally, this shifting mechanism can be used for error recovery purpose by using the MLC capability of PCM and manipulating the data block to recover faulty cells. Evaluation results for multi-threaded workloads reveal 14-25% improvement in lifetime over existing state-of-the-art schemes.
Marjan Asadinia, Majid Jalili 0001, Hamid Sarbazi-Azad
DAC2
2016 Captopril: Reducing the pressure of bit flips on hot locations in non-volatile main memories
Majid Jalili 0001, Hamid Sarbazi-Azad
DATE1
2016 Tolerating more hard errors in MLC PCMs using compression
abstract
Modern computer systems require fast, large and reliable memories to handle information explosion. With this goal in mind, not only deployment of main memories with new technologies are necessary, but also adopting innovative solutions for addressing newfound challenges must be considered as a priority. Recently, phase change memory (PCM) appeared as a preferred candidate for substituting DRAM. PCM with non-volatility, low static power consumption and storing multiple level cells (MLC) capability has opened a new way to the future of memories. Although PCM presents considerable potentials, its short lifetime is a critical concern. Worse still, adopting multiple bits per cell capability degrades the lifetime of PCM more quickly. Hence, convincing the industry for massive production of PCM needs effective and pragmatic solutions. In this paper, we extend the lifetime of a MLC PCM main memory by relying on two strategies: postponing the occurrence of stuck-at failures, and effectively tolerating hard errors. Our scheme postpones the occurrence of hard errors by converting blocks from MLC to SLC mode using byte-level compression. Employing some pointers, the proposed method attempts to cover hard errors by fitting a compressed block either in its physical line location or elsewhere in the page. Full-system and statistical evaluation of the proposed solution shows 48% lifetime improvement of the memory system along with 9% IPC improvement, compared to a state-of-the-art design.
Majid Jalili 0001, Hamid Sarbazi-Azad
ICCD1
2016 Efficient processor allocation in a reconfigurable CMP architecture for dark silicon era
abstract
The continuance of Moore's law and failure of Dennard scaling force future chip multiprocessors (CMPs) to have considerable dark regions. How to use up available dark resources is an important concern for computer architects. In harmony with these changes, we must revise processor allocation schemes that severely affect the performance of a parallel on-chip system. A suitable allocation algorithm should reduce runtime and increase the power efficiency with proper thermal distribution to avoid hotspots. With this motivation, this paper proposes a power-efficient and high performance general purpose infrastructure for which a Dark Silicon Aware Processor Allocation (DSAPA) scheme is proposed which targets future many-core systems. To obtain high performance, we suggest a tunable-clustered mesh with the capability of sharing NoC resources in each cluster. We also employ a buffer-level power gating technique is used to improve power efficiency. Evaluation results reveal that the maximum achieved performance and power consumption improvements are 38.7% and 29.4% for multi-threaded workloads over the equivalent conventional design.
Fatemeh Aghaaliakbari, Mohaddeseh Hoveida, Mohammad Arjomand, Majid Jalili 0001, Hamid Sarbazi-Azad
ICCD4
2014 A compression-based morphable PCM architecture for improving resistance drift tolerance
abstract
Due to the growing demand for large memories, using emerging technologies such as Phase Change Memories (PCM) are inevitable. PCM with appropriate scalability, power consumption and multiple bits per cell storage capability is a probable candidate for substituting DRAM. Although storing multiple bits per cell seems to be a rational response to large memory demands, there is a significant problem to achieve this goal. Resistance drift problem is an important reliability concern that is coupled to a multi-level cell PCM (MLC PCM) memory system. In this paper, we propose a memory system architecture that, by exploiting the benefits of compression, converts resistance drift prone blocks to drift resilient blocks in order to protect the memory system from resistance drift. Evaluations on a full-system simulator, consisting of a quad-core ALPHA CMP and a banked PCM memory, show that our proposed approach provides up to 9.8× reduction in bit error rate, an average of 8% reduction in energy consumption, and 21% IPC improvement.
Majid Jalili 0001, Hamid Sarbazi-Azad
ASAP1
2014 A Reliable 3D MLC PCM Architecture with Resistance Drift Predictor
abstract
In this paper, we study the problem of resistance drift in an MLC Phase Change Memory (PCM) and propose a solution to circumvent its thermally-affected accelerated rate in 3D CMPs. Our scheme is based on the observation that instead of alleviating the problem of resistance drift by using large margins or error correction codes, the PCM read circuit can be reconfigured for tolerating most of the resistance drift errors in a dynamic manner. Through detailed characterization of memory access patterns for 22 applications, we propose an efficient mechanism to facilitate such reliable read scheme via tolerating (a) early-cycle resistance drifts by using narrow margins so that considerably saving energy of writes and improving cell endurance, and (b) late-cycle resistance drifts by accurately estimating resistance thresholds that separate states for sensing. Evaluations on a true 3D architecture, consisting of a 4-core CMP and a banked 2-bit PCM memory, show that our proposal provides 106× lower error rate compared to the state-of-the-art designs of PCMs.
Majid Jalili 0001, Mohammad Arjomand, Hamid Sarbazi-Azad
DSN1