Atieh Lotfi

dblp:68/11076 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorSecurity and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 23% Parallel and multicore computing · 23% Hardware reliability and fault tolerance · 18%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
high-level synthesis
0.312017
RxRE: Throughput Optimization for High-Level Synthesis using Resource-Aware Regularity Extraction (Abstract Only) · FPGA 2017
Parallel and multicore computing › parallel scheduling
resource-aware scheduling
0.312017
RxRE: Throughput Optimization for High-Level Synthesis using Resource-Aware Regularity Extraction (Abstract Only) · FPGA 2017
Hardware reliability and fault tolerance › aging and degradation
aging and lifetime reliability
0.212015
Aging-Aware Compilation for GP-GPUs · ACM Trans. Archit. Code Optim. 2015
Storage systems › flash and SSD › flash memory management
wear leveling
0.212015
Aging-Aware Compilation for GP-GPUs · ACM Trans. Archit. Code Optim. 2015
Compilers and program optimization
static compilation
0.112015
Aging-Aware Compilation for GP-GPUs · ACM Trans. Archit. Code Optim. 2015

Methods — techniques the papers use, named apart from their topics

workload allocation · 0.4online NBTI monitoring · 0.4resource-aware regularity extraction · 0.3
YearPublicationVenuePosition
2022 Error Model (EM) - A New Way of Doing Fault Simulation
abstract
This paper introduces the concept of error model (EM) that replaces a traditional fault-model-based simulation. EM is a temporal simulation of fault symptoms in an application processor. This paper shows that the application resiliency metrics, such as fault coverage, derived through EM are more comprehensive and accurate than those derived through empirical models like single stuck-at faults. In addition, EM can be used on high-level simulation models (behavioral RTL, emulation or in some cases in-silicon). EM approach gives greater than three orders of performance improvement over gate netlist models using stuck-at fault simulation. This paper shows that the coverage metrics, for a billion-logic-gate GPU design, obtained through in-silicon EM closely match the corresponding coverage metrics estimated from a low-level netlist with single stuck-at fault simulation.
Nirmal Saxena, Atieh Lotfi
ITC2
2019 Accelerating Local Binary Pattern Networks with Software-Programmable FPGAs
abstract
Fueled by the success of mobile devices, the computational demands on these platforms have been rising faster than the computational and storage capacities or energy availability to perform tasks ranging from recognizing speech, images to automated reasoning and cognition. While the success of convolutional neural networks (CNNs) have contributed to such a vision, these algorithms stay out of the reach of limited computing and storage capabilities of mobile platforms. It is clear to most researchers that such a transition can only be achieved by using dedicated hardware accelerators on these platforms. However, CNNs with arithmetic-intensive operations remain particularly unsuitable for such acceleration both computationally as well as for the high memory bandwidth needs of highly parallel processing required. In this paper, we implement and optimize an alternative genre of networks, local binary pattern network (LBPNet) which eliminates arithmetic operations by combinatorial operations thus substantially boosting the efficiency of hardware implementation. LBPNet is built upon a radically different view of the arithmetic operations sought by conventional neural networks to overcome limitations posed by compression and quantization methods used for hardware implementation of CNNs. This paper explores in depth the design and implementation of both an architecture and critical optimizations of LBPNet for realization in accelerator hardware and provides a comparison of results with the state-of-art CNN on multiple datasets.
Jeng-Hau Lin, Atieh Lotfi, Vahideh Akhlaghi, Zhuowen Tu, Rajesh K. Gupta 0001
DATE2
2019 Resiliency of automotive object detection networks on GPU architectures
abstract
Safety is the most important aspect of an autonomous driving platform. Deep neural networks (DNNs) play an increasingly critical role in localization, perception, and control in these systems. The object detection and classification inference are of particular importance to construct a precise picture of a vehicle's surrounding objects. Graphics Processing Units (GPU) are well-suited to accelerate such DNN-based inference applications since they leverage data and thread-level parallelism in GPU architectures. Understanding the vulnerability of such DNNs to random hardware faults (including transient and permanent faults) in GPU-based systems is essential to meet the safety requirements of auto safety standards such as the ISO 26262, as well as to influence the design of hardware and software-based safety features in current and future generations of GPU architectures and GPU-based automotive platforms. In this paper, we assess the vulnerability of object detection and classification DNNs to permanent and transient faults using fault injection experiments and accelerated neutron beam testing respectively. We also evaluate the effectiveness of chip-level safety mechanisms in GPU architectures, such as ECC and parity, in detecting these random hardware faults. Our studies demonstrate that such object detection networks tend to be vulnerable to random hardware faults, which cause incorrect or mispredicted object detection outcomes. The neutron beam experiments show that existing chip-level protections successfully mitigate all silent data corruption events caused by transient faults. For permanent faults, while ECC and parity are effective in some cases, our results suggest the need for exploring other complementary detection methods, such as periodic online and offline diagnostic testing.
Atieh Lotfi, Saurabh Hukerikar, Keshav Balasubramanian, Paul Racunas, Nirmal Saxena, Richard Bramley, Yanxiang Huang
ITC1
2018 Low Overhead Tag Error Mitigation for GPU Architectures
abstract
Cache structures on modern GPUs or CPUs occupy a large area and are frequently accessed. This increases their vulnerability to transient errors. With some area and energy overhead, these structures are often protected by ECC or parity checking. However, in deference to the energy efficiency and scalability challenges in high-performance computing, it is crucial to minimize any unnecessary overhead while maintaining the desired reliability. This paper evaluates the reliability of unprotected tag SRAM structures in modern GPUs, and studies the use of a low-overhead tag error mitigation mechanism. The proposed mechanism exploits Galois-based hash functions for set-index calculation to mitigate some pathological address strides that cause false hit events. Extensive analysis on a modern GPU indicates that the hash-based mechanism yields 10x reduction in false hit probability (with 2% improvement in hit rate) for write-through data caches when compared to a baseline cache indexing scheme.
Atieh Lotfi, Nirmal Saxena, Richard Bramley, Paul Racunas, Philip P. Shirvani
DSN1
2017 RxRE: Throughput Optimization for High-Level Synthesis using Resource-Aware Regularity Extraction (Abstract Only)
Atieh Lotfi, Rajesh K. Gupta 0001
FPGA1
2017 ReHLS: Resource-Aware Program Transformation Workflow for High-Level Synthesis
abstract
Despite considerable improvements in existing HLS tools, they still require designer interventions to provide efficient synthesis results. This manual design space exploration and code rewriting and optimization takes significant time and negates the HLS design productivity gains. To overcome this challenge, this paper uses compiler frontend as an independent preprocessing step to explore the design space and adds an automated sourceto- source transformation step before HLS. In particular, it shows how inherent regularity in applications can be used to construct a workflow that analyzes the program, explores the design space for resource optimization opportunity, and transforms the program accordingly. When the transformed program is synthesized using the HLS tool, it uses less hardware resources with similar latency comparing to the original design. The synthesis results on a modern Xilinx Virtex-7 FPGA for a diverse set of applications show that our automated transformation can reduce the design area by an average of 15.4% with less than 1% performance overhead compared to the state-of-the-art Xilinx HLS tool solutions. This automated tool reduces the design time and especially can be useful for non-expert FPGA designers.
Atieh Lotfi, Rajesh K. Gupta 0001
ICCD1
2016 Grater: An approximation workflow for exploiting data-level parallelism in FPGA acceleration
Atieh Lotfi, Abbas Rahimi, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, Rajesh K. Gupta 0001
DATE1
2015 Aging-Aware Compilation for GP-GPUs
abstract
General-purpose graphic processing units (GP-GPUs) offer high computational throughput using thousands of integrated processing elements (PEs). These PEs are stressed during workload execution, and negative bias temperature instability (NBTI) adversely affects their reliability by introducing new delay-induced faults. However, the effect of these delay variations is not uniformly spread across the PEs: some are affected more—hence less reliable—than others. This variation causes significant reduction in the lifetime of GP-GPU parts. In this article, we address the problem of “wear leveling” across processing units to mitigate lifetime uncertainty in GP-GPUs. We propose innovations in the static compiled code that can improve healing in PEs and stream cores (SCs) based on their degradation status. PE healing is a fine-grained very long instruction word (VLIW) slot assignment scheme that balances the stress of instructions across the PEs within an SC. SC healing is a coarse-grained workload allocation scheme that distributes workload across SCs in GP-GPUs. Both schemes share a common property: they adaptively shift workload from less reliable units to more reliable units, either spatially or temporally. These software schemes are based on online calibration with NBTI monitoring that equalizes the expected lifetime of PEs and SCs by regenerating adaptive compiled codes to respond to the specific health state of the GP-GPUs. We evaluate the effectiveness of the proposed schemes for various OpenCL kernels from the AMD APP SDK on Evergreen and Southern Island GPU architectures. The aging-aware healthy kernels generated by the PE (or SC) healing scheme reduce NBTI-induced voltage threshold shift by 30% (77% in the case of SCs), with no (moderate) performance penalty compared to the naive kernels.
Atieh Lotfi, Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001
ACM Trans. Archit. Code Optim.1
2013 A fine-grain distortion and complexity aware parameter tuning model for the H.264/AVC encoder
Mehdi Semsarzadeh, Atieh Lotfi, Mahmoud Reza Hashemi, Shervin Shirmohammadi
Signal Process. Image Commun.2
2012 Architectural vulnerability aware checkpoint placement in a multicore processor
abstract
As the system complexity increases, the failure probability increases substantially. Therefore, the system requires techniques for supporting fault tolerance. Checkpointing technique is widely used to reduce the execution time of long-running programs in presence of failures and enhancing the reliability of such systems. Several methods were studied thus far in order to determine the checkpointing interval which optimizes system performance. The crucial parameter in all of these solutions is system failure model which is primarily assumed as exponential or Weibull distributions. But, these models are not perfectly accurate since they fail to model the effect of soft errors. In this paper, we introduce a more realistic failure model based on the processors AVF. In addition, we propose three checkpoint placement methods with constant and variable intervals that determine suitable checkpoint places for the proposed failure model. Our experimental results show that our method, which is implementable on any multicore system, can find the suitable points in which checkpoints should be taken.
Atieh Lotfi, Arash Bayat, Saeed Safari
IOLTS1