Muhammad Imran 0010

dblp:78/5250-10 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0002-6246-6143ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2024 HYDRA: A Hybrid Resistance Drift Resilient Architecture for Phase Change Memory-Based Neural Network Accelerators
abstract
In-memory Computing (IMC) using Phase Change Memory (PCM) has proven to be effective for efficient processing of Deep Neural Networks (DNNs). However, with the use of multi-level cell PCM (MLC-PCM) in NVMs-based accelerators, errors due to resistance drift in MLC-PCM can severely degrade the DNNs accuracy. In this paper, an analysis of the impact of resistance drift errors on accuracy of MLC-PCM based DNN accelerator shows that the drift errors alone can significantly impact the accuracy. This paper proposes Hydra, which is a hybrid resistance drift resilient architecture for MLC-PCM based DNN accelerators which use IMC for efficient computations. Hydra utilizes Tri-level cell PCM, which has a negligible resistance drift error rate, to store the critical bits of DNNs parameters and MLC-PCM (4-level cell), which has a higher error rate (but offers more storage density), for the non-critical bits. Experimental results on various DNN architectures, configurations and datasets show that, with the presence of resistance drift errors in PCM, Hydra can maintain the baseline accuracy of DNNs for up to 1 year (resistance drift is time-dependent), whereas conventional drift tolerance techniques lead to a significant accuracy drop in just a few seconds.
Thai-Hoang Nguyen, Muhammad Imran 0010, Jaehyuk Choi 0001, Joon-Sung Yang
IEEE Trans. Computers2
2023 CRAFT: Criticality-Aware Fault-Tolerance Enhancement Techniques for Emerging Memories-Based Deep Neural Networks
abstract
Deep neural networks (DNNs) have emerged as the most effective programming paradigm for computer vision and natural language processing applications. With the rapid development of DNNs, efficient hardware architectures for deploying DNN-based applications on edge devices have been extensively studied. Emerging nonvolatile memories (NVMs), with their better scalability, nonvolatility, and good read performance, are found to be promising candidates for deploying DNNs. However, despite the promise, emerging NVMs often suffer from reliability issues, such as stuck-at faults, which decrease the chip yield/memory lifetime and severely impact the accuracy of DNNs. A stuck-at cell can be read but not reprogrammed, thus, stuck-at faults in NVMs may or may not result in errors depending on the data to be stored. By reducing the number of errors caused by stuck-at faults, the reliability of a DNN-based system can be enhanced. This article proposes CRAFT, i.e., criticality-aware fault-tolerance enhancement techniques to enhance the reliability of NVM-based DNNs in the presence of stuck-at faults. A data block remapping technique is used to reduce the impact of stuck-at faults on DNNs accuracy. Additionally, by performing bit-level criticality analysis on various DNNs, the critical-bit positions in network parameters that can significantly impact the accuracy are identified. Based on this analysis, we propose an encoding method which effectively swaps the critical bit positions with that of noncritical bits when more errors (due to stuck-at faults) are present in the critical bits. Experiments of CRAFT architecture with various DNN models indicate that the robustness of a DNN against stuck-at faults can be enhanced by up to$10^{5}$times on the CIFAR-10 dataset and up to 29 times on ImageNet dataset with only a minimal amount of storage overhead, i.e., 1.17%. Being orthogonal, CRAFT can be integrated with existing fault-tolerance schemes to further enhance the robustness of DNNs against stuck-at faults in NVMs.
Thai-Hoang Nguyen, Muhammad Imran 0010, Jaehyuk Choi 0001, Joon-Sung Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 DynaPAT: A Dynamic Pattern-Aware Encoding Technique for Robust MLC PCM-Based Deep Neural Networks
abstract
As the effectiveness of Deep Neural Networks (DNNs) is rising over time, so is the need for highly scalable and efficient hardware architectures to capitalize this effectiveness in many practical applications. Emerging non-volatile Phase Change Memory (PCM) technology has been found to be a promising candidate for future memory systems due to its better scalability, non-volatility and low leakage/dynamic power consumption, compared to conventional charged-based memories. Additionally, with its cell's wide resistance span, PCM also has the Flash-like Multi-Level Cell (MLC) capability, which has enhanced storage density, providing an opportunity for the deployment of data-intensive applications such as DNNs on resource-constrained edge devices. However, the practical deployment of MLC PCM is hampered by certain reliability challenges, among which, the resistance drift is considered to be a critical concern. In a DNN application, the presence of resistance drift in MLC PCM can cause a severe impact to DNN's accuracy if no drift-error-tolerance technique is utilized. This paper proposes DynaPAT, a low-cost and effective pattern-aware encoding technique to enhance the drift-error-tolerance of MLC PCM-based Deep Neural Networks. DynaPAT has been constructed on the insight into DNN's vulnerability against different data pattern switching. Based on this insight, DynaPAT efficiently maps the most-frequent data pattern in DNN's parameters to the least-drift-prone level of the MLC PCM, thus significantly enhancing the robustness of the system against drift errors. Various experiments on different DNN models and configurations demonstrate the effectiveness of DynaPAT. The experimental results indicate that DynaPAT can achieve up to 500× enhancement in the drift-errors-tolerance capability over the baseline MLC PCM based DNN while requiring only a negligible hardware overhead (below 1% storage overhead). Being orthogonal, DynaPAT can be integrated with existing drift-tolerance schemes for even higher gains in reliability.
Thai-Hoang Nguyen, Muhammad Imran 0010, Joon-Sung Yang
ICCAD2
2022 CEnT: An Efficient Architecture to Eliminate Intra-Array Write Disturbance in PCM
abstract
Phase Change Memory (PCM), with its better scaling potential compared to DRAM, is seen as a promising candidate to replace or complement DRAM. The heat generated from a RESET programming pulse to a PCM cell can disturb the neighboring cells which are not being programmed. Write disturbance (WD) poses a critical reliability challenge in high-density PCM memory with scaling below 20nm process technology node. Increasing the intra-cell space can eliminate the WD, however, it reduces the storage density which counteracts the benefits of scalability in PCM. At architectural level, a verify and correct (VnC) technique can be used to address this problem. However, this leads to an increased number of write operations, thus degrading performance, energy efficiency and memory lifetime. Due to its dependence on the type of programming operation and the state of the neighboring cell, WD is a data-dependent problem. Exploiting this property, encoding techniques have been proposed to reduce the frequency of WD-vulnerable data patterns. These techniques, however, do not eliminate the WD in an array and ultimately rely on the VnC method to ensure reliable memory operation. This article introduces a novel architecture, based on encoding and multi-level programming characteristics of PCM, to eliminate the intra-array WD in PCM. By eliminating WD and hence the need for a VnC operation, the proposed architecture improves performance, energy efficiency and memory lifetime. Our evaluation of the proposed architecture shows an average reduction of 57 percent in the number of writes (to service one write request) over the existing state-of-the-art intra-array WD-mitigation technique. Depending on the PCM write bandwidth, the proposed architecture can reduce the write service time by up to 27 percent, on average, compared to the existing best-performing technique. This leads to an average improvement of 15 percent in IPC. Additionally, by eliminating the overhead of a verify operation, the write energy efficiency is also improved by 8 percent over the previous art. Finally, with an average reduction of 26 percent in bit flips, the proposed method also improves the memory lifetime. The proposed method is also proven to be effective when considering WD both within the word-lines and across the bit-lines.
Muhammad Imran 0010, Nur A. Touba, Joon-Sung Yang
IEEE Trans. Computers1
2022 ADAPT: A Write Disturbance-Aware Programming Technique for Scaled Phase Change Memory
abstract
Phase change memory (PCM) is an emerging, resistance-based, nonvolatile memory. With a promising scaling potential, PCM can replace the existing charge-based memory technologies. A highly scaled PCM is prone to write disturbance (WD) because of a high-current RESET programming pulse. Exploiting the data-dependent nature of WD, encoding techniques have been proposed to reduce the frequency of WD-vulnerable data patterns. These techniques work along with a verify and correct (VnC) method to ensure memory reliability. However, the effectiveness of these techniques varies depending on the data patterns. Unlike the conventional methods, this article introduces a WD-aware programming technique to mitigate WD in PCM. The proposed method encodes the data based on the number of WD-vulnerable cells and the bit flips. By reducing the number of WD-vulnerable cells as well as the bit flips, the proposed method is more effective than the existing encoding techniques, in mitigating WD as well as improving the memory lifetime. Evaluation using various realistic workloads shows that the proposed method can reduce the average word-line WD errors by 62%, compared to the existing state of the art. With a reduced number of WD errors, the frequency of a VnC operation is also reduced. This leads to a reduction of 44% in the number of extra writes and 20% in the average write time. With reduction in the number of writes and the write time, instructions-per-cycle is improved by 9% and the write energy by 11% over the existing art. By reducing the number of bit flips compared to the previous state of the art, the proposed method improves the PCM memory lifetime by 13% to 33%, considering the asymmetry of SET and RESET operations in impacting the cell endurance.
Muhammad Imran 0010, Joon-Sung Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Low-Cost and Effective Fault-Tolerance Enhancement Techniques for Emerging Memories-Based Deep Neural Networks
abstract
Deep Neural Networks (DNNs) have been found to outperform conventional programming approaches in several applications such as computer vision and natural language processing. Efficient hardware architectures for deploying DNNs on edge devices have been actively studied. Emerging memory technologies with their better scalability, non-volatility, and good read performance are ideal candidates for DNNs which are trained once and deployed over many devices. Emerging memories have also been used in DNNs accelerators for efficient computations of dot-product. However, due to immature manufacturing and limited cell endurance, emerging resistive memories often result in reliability issues like stuck-at faults, which reduce the chip yield and pose a challenge to the accuracy of DNNs. Depending on the state, stuck-at faults may or may not cause error. Fault-tolerance of DNNs can be enhanced by reducing the impact of errors resulting from the stuck-at faults. In this work, we introduce simple and light-weight Intra-block Address remapping and weight encoding techniques to improve the fault-tolerance for DNNs. The proposed schemes effectively work at the network deployment time while preserving the network organization and the original values of the parameters. Experimental results on state-of-the-art DNN models indicate that, with a small storage overhead of just 0.98%, the proposed techniques achieve up to 300× stuck-at faults tolerance capability on Cifar10 dataset and 125× on Imagenet datatset, compared to the baseline DNNs without any fault-tolerance method. By integrating with the existing schemes, the proposed schemes can further enhance the fault resilience of DNNs.
Thai-Hoang Nguyen, Muhammad Imran 0010, Jaehyuk Choi 0001, Joon-Sung Yang
DAC2
2021 Reliability Enhanced Heterogeneous Phase Change Memory Architecture for Performance and Energy Efficiency
abstract
Next-generation memories have been actively researched to replace the existing memories like DRAM and flash in deep sub-micron process technology. Unlike the conventional charge-based memories, next-generation memories utilize the resistive properties of different materials to store and read a data. Among the next-generation memories, Phase Change Memory (PCM) is seen as a good choice for future memory systems, given its good read performance, process compatibility and scaling potential. To enhance the storage density, multi-level cell (MLC) operation is seemed promising which can store more than one bit in each PCM cell. However, MLC operation significantly degrades the reliability of PCM, thus requiring a strong Error Correction Code (ECC) to guarantee correct memory operation. The use of heavyweight ECC comes at cost of significant degradations in storage density, performance and energy efficiency. In this article, we propose a heterogeneous PCM architecture which uses both multi-level cell and single-level cell (SLC) together for a single word line. With highly-reliable SLC cells, the overall array reliability is enhanced. To improve the reliability further, a dynamic self-encoding/decoding scheme is performed before the data is written to the PCM cells. The dynamic scheme automatically determines the locations of MLC and SLC cells and sets the corresponding resistance levels to be programmed. Since the proposed encoding/decoding scheme does not require any additional stages or storages for encoding and decoding, the overhead is negligible. The improved reliability allows to use lighter ECC scheme which in turn helps to improve performance and energy efficiency of the MLC PCM. The experimental results show that the reliability is improved by approximately$10^6$times compared to the conventional 4LC and more than$10^3$times compared to the existing encoding methods. The performance improvement is 21.5 percent over the conventional 4LC and is more than 4.1 percent higher than the prior encoding techniques. The proposed method is 30.3 percent more energy efficient than the conventional 4LC and this is similar or higher than other energy efficiency improvement methods.
Muhammad Imran 0010, Joon-Sung Yang
IEEE Trans. Computers2
2020 Effective Write Disturbance Mitigation Encoding Scheme for High-density PCM
abstract
Write Disturbance (WD) is a crucial reliability concern in a high-density PCM with below 20nm scaling. WD occurs because of the inter-cell heat transfer during a RESET operation. Being dependent on the type of programming pulse and the state of the vulnerable cell, WD is significantly impacted by the data patterns. Existing encoding techniques to mitigate WD reduce the percentage of a single WD-vulnerable pattern in the data. However, it is observed that reducing the frequency of a single bit pattern may not be effective to mitigate WD for certain data patterns. This work proposes a significantly more effective encoding method which minimizes the number of vulnerable cells instead of a single bit pattern. The proposed method mitigates WD both within a word-line and across the bit-lines. In addition to WD-mitigation, the proposed method encodes the data to minimize the bit flips, thus improving the memory lifetime compared to the conventional WD-mitigation techniques. Our evaluation using SPEC CPU2006 benchmarks shows that the proposed method can reduce the aggregate (word- line+bit-line) WD errors by 42% compared to the existing state- of-the-art (SD-PCM). Compared to the state-of-the-art SD-PCM method, the proposed method improves the average write time, instructions-per-cycle (IPC) and write energy by 12%, 12% and 9%, respectively, by reducing the frequency of verify and correct operations to address WD errors. With reduction in bit flips, memory lifetime is also improved by 18% to 37% compared to SD-PCM, given an asymmetric cost of the bit flips. By integrating with the orthogonal techniques of SD-PCM, the proposed method can further enhance the performance and energy efficiency.
Muhammad Imran 0010, Joon-Sung Yang
DATE1
2020 Pattern-Aware Encoding for MLC PCM Storage Density, Energy Efficiency, and Performance Enhancement
abstract
With the scaling limitations and increasing leakage power of the existing charge-based memories, next-generation memory technologies to overcome the issues are in development. Among the various emerging memories, phase change memory (PCM) is considered as a promising candidate due to its scalability potential and negligible leakage power. For enhanced storage density, the multilevel cell (MLC) operation has been proposed for PCM. This, however, comes at cost of poor reliability, write energy increase and performance degradation. Unlike DRAM, the MLC PCM has a much higher soft error rate due to the resistance drift phenomenon. Error correction code (ECC) schemes can be utilized to improve the MLC PCM reliability, however, this would lead to a lower storage density and an increase in write energy and latency. The iterative programming required for the MLC PCM also degrades its energy efficiency and performance. This paper introduces a simple yet effective encoding scheme to mitigate the problems of the MLC PCM. By using a simple XOR-based encoding, the proposed architecture minimizes the most drift-prone state in the data. The method divides the original data into several encoding blocks and analyzes initial pattern frequencies for each 2-bit pattern. Based on the initial pattern frequencies, the inputs for the XOR encoding are selected that result in minimal frequency of the drift-prone state. This considerably enhances the MLC PCM reliability, leading to a high storage density with a reduced ECC overhead. The energy efficiency and performance are also improved due to reduction in iterative current pulses and ECC overhead. The simulation results show a reduction of about$10^{5}$X in soft error rate. The improvements in energy efficiency and performance over the conventional 4-level cell (4LC) PCM are 11.5% and 31.9%, respectively.
Muhammad Imran 0010, Joon-Sung Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Virtual-Tile-Based Flip-Flop Alignment Methodology for Clock Network Power Optimization
abstract
Clock network plays the most significant role in power consumption in IC design. Since a clock network normally has a high switching ratio, power optimization of the clock network is one of the best solutions to minimize dynamic power and total power in modern IC designs. The clock network is synthesized based on an initial flip-flop placement. The number of clock buffers and their sizes are decided by the initial placement. Moreover, clock wires, which are the major sources of clock power consumption, are also constructed based on the flip-flop placement. As a result, the flip-flop placement determines the quality of the clock network. In this article, we propose a new clock network optimization method to reduce the dynamic power consumption of clock network. The method first creates virtual tiles over the entire design area and selects the most effective columns to align flip-flops in lines. Once the effective columns are determined, flip-flops are relocated based on the virtual tiles in the columns considering the minimum moving distance. By aligning flip-flops, it is possible to significantly reduce both wire capacitance and wire length. Since it does not change the clock structure, unlike the conventional clock network optimization techniques which use multibit flip-flop or register bank, there is no degradation in timing or other constraints. Experimental results show that the proposed method reduces the wire capacitance, wire length, and via count up to 23.2%, 10.2%, and 16.4%, respectively, in five industrial intellectual property (IP) designs. The reduction in clock network power is 14.1% on average.
Muhammad Imran 0010, David Z. Pan, Joon-Sung Yang
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Flipcy: Efficient Pattern Redistribution for Enhancing MLC PCM Reliability and Storage Density
abstract
Phase change memory (PCM) is a scalable, non-volatile emerging memory. The storage density of the PCM can be enhanced by using the multi-level cell (MLC) operation. However, the MLC PCM suffers from low reliability due to resistance drift. The rate of resistance drift is proportional to the initial resistance of the cell with intermediate storage levels being particularly vulnerable. Using heavy Error Correction Codes (ECC) results in poor effective storage density (data bits per cell) of the MLC PCM. The state-of-the-art Tri-level cell technique improves reliability by using only three out of four storage levels, thus eliminating the ECC overhead. However, its storage density is much less than the ideal MLC PCM. Moreover, its storage density is fixed even if practically the MLC PCM reliability is improved. This paper introduces a more flexible pattern redistribution technique, Flipcy, to improve the MLC PCM reliability and effective storage density. The proposed method proportions the data-patterns according to the rate of resistance drift for different storage levels. A simple flip or a complement operation is used to reduce the percentage of the most error-prone pattern. The simulation results show up to 107X reduction in the error rate and 31% improvement in performance compared to the conventional MLC PCM. With a reduced overhead of the auxiliary bits and ECC parity bits, the proposed method can achieve about 25% improvement in effective storage density over the Tri-level cell approach for a similar level of reliability while incurring about 11% degradation in performance compared to the Tri-level cell approach. This performance degradation can be reduced when the proposed method is accompanied with orthogonal techniques to improve MLC PCM reliability and efficient scrubbing methods.
Muhammad Imran 0010, Jung Min You, Joon-Sung Yang
ICCAD1
2018 Heterogeneous PCM array architecture for reliability, performance and lifetime enhancement
abstract
Conventional DRAM and flash memory are reaching their scaling limits thus motivating research in various emerging memory technologies as a potential replacement. Among these, phase change memory (PCM) has received considerable attention owing to its high scalability and multi-level cell (MLC) operation for high storage density. However, due to the resistance drift over time, the soft error rate in MLC PCM is high. Additionally, the iterative programming in MLC negatively impacts performance and cell endurance. The conventional methods to overcome the drift problem incur large overheads, impact memory lifetime and are inadequate in terms of acceptable soft error rate (SER). In this paper, we propose a new PCM memory architecture with heterogeneous PCM arrays to increase reliability, performance and lifetime. The basic storage unit in the proposed architecture consists of two single-level cells (SLCs) and one four-level cell (4LC). Using the reduced number of 4LCs compared to conventional homogeneous 4LC PCM arrays, the drift-induced error rate is considerably reduced. By alternating each cell operation between SLC and 4LC over time, the overall lifetime can also be significantly enhanced. The proposed architecture achieves up to 105times lower soft error rate with considerably less ECC overhead. With simple ECC scheme, about 22% performance improvement is achieved and additionally, the overall lifetime is also enhanced by about 57%.
Muhammad Imran 0010, Jung Min You, Joon-Sung Yang
DATE2