Fateme S. Hosseini

dblp:202/9522 · also Fateme Sadat Hosseini · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
3since 2021 · last 2021
0000-0002-9091-3908ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2021 A Self-Test Framework for Detecting Fault-induced Accuracy Drop in Neural Network Accelerators
abstract
Hardware accelerators built with SRAM or emerging memory devices are essential to the accommodation of the ever-increasing Deep Neural Network (DNN) workloads on resource-constrained devices. After deployment, however, the performance of these accelerators is threatened by the faults in their on-chip and off-chip memories where millions of DNN weights are held. Different types of faults may exist depending on the underlying memory technology, degrading inference accuracy. To tackle this challenge, this paper proposes an online self-test framework that monitors the accuracy of the accelerator with a small set of test images selected from the test dataset. Upon detecting a noticeable level of accuracy drop, the framework uses additional test images to identify the corresponding fault type and predict the severeness of faults by analyzing the change in the ranking of the test images. Experimental results show that our method can quickly detect the fault status of a DNN accelerator and provide accurate fault type and fault severeness information, allowing for subsequent recovery and self-healing process.
Fanruo Meng, Fateme S. Hosseini, Chengmo Yang
ASP-DAC2
2021 A Compile-Time Framework for Tolerating Read Disturbance in STT-RAM
abstract
Spin-transfer torque magnetic random access memory (STT-RAM) is one of the most promising candidates for next-generation on-chip memories. While STT-RAM offers high density, negligible leakage power, and fast access speed, it also suffers from read-disturbance errors, that is, read operations might accidentally change the value of the accessed memory location. Although these errors could be mitigated by applying restore-after-read operations, the energy overhead would be significant. To reduce such overhead, this article presents an application-level and architecture-independent framework, which selectively inserts restore operations under the guidance of a compiler. This work first introduces a new concept of disturbance chain and then analyzes the vulnerability of each load instruction on a chain to read disturbance errors. This work further proposes a number of compile-time code optimizations to reduce the number of vulnerable loads and hence the associated restore overhead. The proposed compiler optimizations are implemented in LLVM. Experiments in Gem5 show up to 98.6% reduction in the number of restore operations and 48% savings of the energy overhead while maintaining 99.8% coverage of read disturbance errors.
Fateme S. Hosseini, Chengmo Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Tolerating Defects in Low-Power Neural Network Accelerators Via Retraining-Free Weight Approximation
abstract
Hardware accelerators are essential to the accommodation of ever-increasing Deep Neural Network (DNN) workloads on the resource-constrained embedded devices. While accelerators facilitate fast and energy-efficient DNN operations, their accuracy is threatened by faults in their on-chip and off-chip memories, where millions of DNN weights are held. The use of emerging Non-Volatile Memories (NVM) further exposes DNN accelerators to a non-negligible rate of permanent defects due to immature fabrication, limited endurance, and aging. To tolerate defects in NVM-based DNN accelerators, previous work either requires extra redundancy in hardware or performs defect-aware retraining, imposing significant overhead. In comparison, this paper proposes a set of algorithms that exploit the flexibility in setting the fault-free bits in weight memory to effectively approximate weight values, so as to mitigate defect-induced accuracy drop. These algorithms can be applied as a one-step solution when loading the weights to embedded devices. They only require trivial hardware support and impose negligible run-time overhead. Experiments on popular DNN models show that the proposed techniques successfully boost inference accuracy even in the face of elevated defect rates in the weight memory.
Fateme S. Hosseini, Fanruo Meng, Chengmo Yang, Wujie Wen, Rosario Cammarota
ACM Trans. Embed. Comput. Syst.1
2019 Compiler-Directed and Architecture-Independent Mitigation of Read Disturbance Errors in STT-RAM
abstract
High density, negligible leakage power, and fast read speed have made Spin-Transfer Torque Random Access Memory (STT-RAM) one of the most promising candidates for next generation on-chip memories. However, STT-RAM suffers from read-disturbance errors, that is, read operations might accidentally change the value of the accessed memory location. Although these errors could be mitigated by applying a restore-after-read operation, the energy overhead would be significant. This paper presents an architecture-independent framework to mitigate read disturbance errors while reducing the energy overhead, by selectively inserting restore operations under the guidance of the compiler. For that purpose, the vulnerability of load operations to read disturbance errors is evaluated using a specifically designed fault model; a code transformation technique is developed to reduce the number of vulnerable loads; and, an algorithm is proposed to selectively insert restore operations. The evaluation results show that the proposed technique can effectively reduce up to 97% of restore operations and 66% of the energy overhead while maintaining 99.8% coverage of read disturbance errors.
Fateme S. Hosseini, Chengmo Yang
DATE1
2017 Leveraging Compiler Optimizations to Reduce Runtime Fault Recovery Overhead
abstract
Smaller feature size, lower supply voltage, and faster clock rates have made modern computer systems more susceptible to faults. Although previous fault tolerance techniques usually target a relatively low fault rate and consider error recovery less critical, with the advent of higher fault rates, recovery overhead is no longer negligible. In this paper, we propose a scheme that leverages and revises a set of compiler optimizations to design, for each application hotspot, a smart recovery plan that identifies the minimal set of instructions to be re-executed in different fault scenarios. Such fault scenario and recovery plan information is efficiently delivered to the processor for runtime fault recovery. The proposed optimizations are implemented in LLVM and GEM5. The results show that the proposed scheme can significantly reduce runtime recovery overhead by 72%.
Fateme S. Hosseini, Pouya Fotouhi, Chengmo Yang, Guang R. Gao
DAC1
2017 Segment and Conflict Aware Page Allocation and Migration in DRAM-PCM Hybrid Main Memory
abstract
Phase change memory (PCM), given its nonvolatility, potential high density, and low standby power, is a promising candidate to be used as main memory in next generation computer systems. However, to hide its shortcomings of limited endurance and slow write performance, state-of-the-art solutions tend to construct a dynamic RAM (DRAM)-PCM hybrid memory and place write-intensive pages in DRAM. While existing optimizations to this hybrid architecture focus on tuning DRAM configurations to reduce the number of write operations to PCM, this paper explores the interactions between DRAM and PCM to improve both the performance and the endurance of a DRAM-PCM hybrid main memory. Specifically, it exploits the flexibility of mapping virtual pages to physical pages, and develops a proactive strategy to allocate pages taking both program segments and DRAM conflict misses into consideration, thus distributing those heavily written pages across different DRAM sets. Meanwhile, a lifetime-aware DRAM replacement algorithm and a conflict-aware page remapping strategy are proposed to further reduce DRAM misses and PCM writes. Experiments confirm that the proposed techniques are able to improve average memory hit time and reduce maximum PCM write counts thus enhancing both performance and lifetime of a DRAM-PCM hybrid main memory.
Hoda Aghaei Khouzani, Fateme S. Hosseini, Chengmo Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2