Navid Khoshavi

dblp:45/10032 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-4010-1354ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 INSPIRE: In-Sensor Compressed Weight Retrieval for Enhancing ViT Efficiency at Edge
abstract
Deploying Vision Transformer (ViT) models on edge devices poses significant challenges due to the high bandwidth, energy demands, and latency associated with transmitting large weight parameter sets to the sensing unit, along with limited on-chip memory resources, which are often insufficient for storing these parameters. To address these constraints, we present a software-hardware co-design framework that incorporates a novel in-sensor Compressed Weight Retrieval mechanism within an intelligent vision sensor. This framework offers two key contributions. First, we propose an innovative hardware-friendly weight compression algorithm that substantially reduces bandwidth and power consumption by optimizing on-chip memory usage for storing weight parameters. Second, we leverage the exceptional efficiency of Silicon Photonic (SiPh) devices and design a novel in-sensor accelerator called INSPIRE for the first time to perform in-sensor retrieval of the compressed weights and parallel fine-grained convolution operations next to the pixel array, enabling low-power adaptable ViT inference on resource-constrained edge platforms. Our extensive simulation results show that INSPIRE can remarkably reduce the memory footprint of ViT results with favorable accuracy. Besides, INSPIRE significantly reduces the bandwidth and power requirements associated with storing weight parameters in on-chip memory. INSPIRE achieves up to 245.4 Kilo FPS/W and reduces the data transfer energy by a factor of ∼11× on average compared with 4-bit quantized ViTs.
Deniz Najafi, Mohaiminul Al Nahian, Navid Khoshavi, Abdullah Al Arafat, Mamshad Nayeem Rizve, Mahdi Nikdast, Adnan Siraj Rakin, Shaahin Angizi
DATE4
2025 Event-Driven Spatiotemporal Processing-In-Sensor with Phase Change Memory-based Optical Acceleration
Mehrdad Morsali, Deniz Najafi, Amin Shafiee, Sepehr Tabrizchi, Pietro Mercati, Mohsen Imani, Arman Roohi, Navid Khoshavi, Mahdi Nikdast, Shaahin Angizi
ACM Great Lakes Symposium on VLSI8
2025 Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics
abstract
Vision Transformers (ViTs) have emerged as a powerful architecture for computer vision tasks due to their ability to model long-range dependencies and global contextual relationships. However, their substantial compute and memory demands hinder efficient deployment in scenarios with strict energy and bandwidth limitations. In this work, we propose Opto-ViT, the first near-sensor, region-aware ViT accelerator leveraging silicon photonics (SiPh) for real-time and energy-efficient vision processing. Opto-ViT features a hybrid electronic-photonic architecture, where the optical core handles compute-intensive matrix multiplications using Vertical-Cavity Surface-Emitting Lasers (VCSELs) and Microring Resonators (MRs), while nonlinear functions and normalization are executed electronically. To reduce redundant computation and patch processing, we introduce a lightweight Mask Generation Network (MGNet) that identifies regions of interest in the current frame and prunes irrelevant patches before ViT encoding. We further co-optimize the ViT backbone using quantization-aware training and matrix decomposition tailored for photonic constraints. Experiments across device fabrication, circuit and architecture co-design, to classification, detection, and video tasks demonstrate that Opto-ViT achieves 100.4 KFPS/W with up to 84% energy savings with less than 1.6% accuracy loss, while enabling scalable and efficient ViT deployment at the edge.
Mehrdad Morsali, Chengwei Zhou, Deniz Najafi, Sreetama Sarkar, Pietro Mercati, Navid Khoshavi, Peter A. Beerel, Mahdi Nikdast, Gourav Datta, Shaahin Angizi
ICCAD6
2022 Path and Floor Detection in Outdoor Environments for Fall Prevention of the Visually Impaired Population
abstract
According to the world report in 2019 about vision health presented by the World Health Organization (WHO), there are 2.2 billion people in the world with some kind of visual impairment. Furthermore, in a recent National Health Interview Survey report, 25.5 million adult Americans 18 and older reported experiencing vision loss. Of these 25.5 million American adults, 15.3 million women and 10.1 million men report experiencing significant vision loss. As a result, this population is constantly at risk of a fall and its consequences. This paper presents the first module of our fall prevention system for visually impaired people. The module corresponds to a system of floor and path recognition in outdoor environments. At the core of the system is a Mask R-CNN trained with around 40.000 outdoor images. The proposed approach has reached a performance in a combination of training and testing of 92%, using a combined set of indoor and outdoor environment images.
Juan M. Calderón, Navid Khoshavi, Luis Gabriel Jaimes
CCNC3
2022 MeF-RAM: A New Non-Volatile Cache Memory Based on Magneto-Electric FET
abstract
Magneto-Electric FET ( MEFET ) is a recently developed post-CMOS FET, which offers intriguing characteristics for high-speed and low-power design in both logic and memory applications. In this article, we present MeF-RAM , a non-volatile cache memory design based on 2-Transistor-1-MEFET ( 2T1M ) memory bit-cell with separate read and write paths. We show that with proper co-design across MEFET device, memory cell circuit, and array architecture, MeF-RAM is a promising candidate for fast non-volatile memory ( NVM ). To evaluate its cache performance in the memory system, we, for the first time, build a device-to-architecture cross-layer evaluation framework to quantitatively analyze and benchmark the MeF-RAM design with other memory technologies, including both volatile memory (i.e., SRAM, eDRAM) and other popular non-volatile emerging memory (i.e., ReRAM, STT-MRAM, and SOT-MRAM). The experiment results for the PARSEC benchmark suite indicate that, as an L2 cache memory, MeF-RAM reduces Energy Area Latency ( EAT ) product on average by ~98% and ~70% compared with typical 6T-SRAM and 2T1R SOT-MRAM counterparts, respectively.
Shaahin Angizi, Navid Khoshavi, Andrew Marshall, Peter Dowben, Deliang Fan
ACM Trans. Design Autom. Electr. Syst.2
2021 Entropy-Based Modeling for Estimating Adversarial Bit-flip Attack Impact on Binarized Neural Network
abstract
Over past years, the high demand to efficiently process deep learning (DL) models has driven the market of the chip design companies. However, the new Deep Chip architectures, a common term to refer to DL hardware accelerator, have slightly paid attention to the security requirements in quantized neural networks (QNNs), while the black/white -box adversarial attacks can jeopardize the integrity of the inference accelerator. Therefore in this paper, a comprehensive study of the resiliency of QNN topologies to black-box attacks is examined. Herein, different attack scenarios are performed on an FPGA-processor co-design, and the collected results are extensively analyzed to give an estimation of the impact's degree of different types of attacks on the QNN topology. To be specific, we evaluated the sensitivity of the QNN accelerator to a range number of bit-flip attacks (BFAs) that might occur in the operational lifetime of the device. The BFAs are injected at uniformly distributed times either across the entire QNN or per individual layer during the image classification. The acquired results are utilized to build the entropy-based model that can be leveraged to construct resilient QNN architectures to bit-flip attacks.
Navid Khoshavi, Saman Sargolzaei, Yu Bi, Arman Roohi
ASP-DAC1
2020 SHIELDeNN: Online Accelerated Framework for Fault-Tolerant Deep Neural Network Architectures
abstract
We propose SHIELDeNN, an end-to-end inference accelerator frame-work that synergizes the mitigation approach and computational resources to realize a low-overhead error-resilient Neural Network (NN) overlay. We develop a rigorous fault assessment paradigm to delineate a ground-truth fault-skeleton map for revealing the most vulnerable parameters in NN. The error-susceptible parameters and resource constraints are given to a function to find superior design. The error-resiliency magnitude offered by SHIELDeNN can be adjusted based on the given boundaries. SHIELDeNN methodology improves the error-resiliency magnitude of cnvW1A1 by 17.19% and 96.15% for 100 MBUs that target weight and activation layers, respectively.
Navid Khoshavi, Arman Roohi, Connor Broyles, Saman Sargolzaei, Yu Bi, David Z. Pan
DAC1
2020 Fiji-FIN: A Fault Injection Framework on Quantized Neural Network Inference Accelerator
abstract
In recent years, the big data booming has boosted the development of highly accurate prediction models driven from machine learning (ML) and deep learning (DL) algorithms. These models can be orchestrated on the customized hardware in the safety-critical missions to accelerate the inference process in ML/DL -powered IoT. However, the radiation-induced transient faults and black/white -box attacks can potentially impact the individual parameters in ML/DL models which may result in generating noisy data/labels or compromising the pre-trained model. In this paper, we propose Fiji-FIN1, a suitable framework for evaluating the resiliency of IoT devices during the ML/DL model execution with respect to the major security challenges such as bit perturbation attacks and soft errors. Fiji-FIN is capable of injecting both single bit/event flip/upset and multi-bit flip/upset faults on the architectural ML/DL accelerator embedded in ML/DL -powered IoT. Fiji-FIN is significantly more accurate compared to the existing software-level fault injections paradigms on ML/DL -driven IoT devices.
Navid Khoshavi, Connor Broyles, Yu Bi, Arman Roohi
ICMLA1
2020 A survey on attack vectors in stack cache memory
Navid Khoshavi, Mohammad Maghsoudloo, Yu Bi, William Francois, Luis Gabriel Jaimes, Arman Sargolzaei
Integr.1
2020 Elastic HDFS: interconnected distributed architecture for availability-scalability enhancement of large-scale cloud storages
Mohammad Maghsoudloo, Navid Khoshavi
J. Supercomput.2
2019 Self-Organized Sub-bank SHE-MRAM-based LLC: An energy-efficient and variation-immune read and write architecture
Soheil Salehi, Navid Khoshavi, Ramtin Zand, Ronald F. DeMara
Integr.2
2017 Contemporary CMOS aging mitigation techniques: Survey, taxonomy, and methods
Navid Khoshavi, Rizwan A. Ashraf, Ronald F. DeMara, Saman Kiamehr, Fabian Oboril, Mehdi Baradaran Tahoori
Integr.1
2017 Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache
abstract
For the sake of higher cell density while achieving near-zero standby power, recent research progress in Magnetic Tunneling Junction (MTJ) devices has leveraged Multi-Level Cell (MLC) configurations of Spin-Transfer Torque Random Access Memory (STT-RAM). However, in orderto mitigate the write disturbance in an MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. Furthermore, as the result of MTJ feature size scaling, the soft bit can be expected to become disturbed by the read sensing current, thus requiring an immediate restore operation to ensure the data reliability. In this paper, we design and analyze a novel Adaptive Restore Scheme for Write Disturbance (ARS-WD) and Read Disturbance (ARS-RD), respectively. ARS-WD alleviates restoration overhead by intentionally overwriting soft bit lines which are less likely to be read. ARS-RD, on the other hand, aggregates the potential writes and restore the soft bit line at the time of its eviction from higher level cache. Both of these two schemes are based on a lightweight forecasting approach for the future read behavior of the cache block. Our experimental results show substantial reduction in soft bit line restore operations, delivering 17.9 percent decrease in overall energy consumption and 9.4 percent increase in IPC, while incurring negligible capacity overhead. Moreover, ARS promotes advantages of MLC to provide a preferable L2 design alternative in terms of energy, area and latency product compared to SLC STT-RAM alternatives.
Xunchao Chen, Navid Khoshavi, Ronald F. DeMara, Jun Wang 0001, Dan Huang 0001, Wujie Wen, Yiran Chen 0001
IEEE Trans. Computers2
2016 AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cache
abstract
Spin-Transfer Torque Random Access Memory (STT-RAM) has been identified as an advantageous candidate for on-chip memory technology due to its high density and ultra low leakage power. Recent research progress in Magnetic Tunneling Junction (MTJ) devices has developed Multi-Level Cell (MLC) STT-RAM to further enhance cell density. To avoid the write disturbance in MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. However, frequent restores are not only unnecessary, but also introduce a significant energy consumption overhead. In this paper, we propose an Adaptive Overwrite Scheme (AOS) which alleviates restoration overhead by intentionally overwriting selected soft bits based on RRD (Read Reuse Distance). Our experimental results show 54.6% reduction in soft bit restoration, delivering 10.8% decrease in overall energy consumption. Moreover, AOS promotes MLC to be a preferable L2 design alternative in terms of energy, area and latency product.
Xunchao Chen, Navid Khoshavi, Jian Zhou 0004, Dan Huang 0001, Ronald F. DeMara, Jun Wang 0001, Wujie Wen, Yiran Chen 0001
DAC2
2015 Reactive rejuvenation of CMOS logic paths using self-activating voltage domains
abstract
Although the trend of technology scaling is sought to realize higher performance computer systems, it also results in Integrated Circuits (ICs) suffering from increasing Process, Voltage, and Temperature (PVT) variations and adverse aging effects. In most cases, these reliability threats manifest themselves as timing errors on critical speed-paths of the circuit, if a large design guardband is not reserved. In this work, we propose the Reactive Rejuvenation (RR) architectural approach consisting of detection and recovery phases to mitigate circuit from BTI-induced aging. The BTI impact on the critical and near critical paths performance is continuously examined through a lightweight logic circuit which asserts an error signal in the case of any timing violation in those paths. By utilizing timing violation occurrence in the system, the timing-sensitive portion of the circuit is recovered from BTI through switching computations to redundant aging-critical voltage domain. The proposed technique achieves aging mitigation and reduced energy consumption as compared to a baseline circuit. Thus, significant voltage guardbands to meet the desired timing specification are avoided.
Rizwan A. Ashraf, Ahmad Alzahrani 0001, Navid Khoshavi, Ramtin Zand, Soheil Salehi, Arman Roohi, Mingjie Lin, Ronald F. DeMara
ISCAS3
2011 Soft Error Detection Technique in Multi-threaded Architectures Using Control-Flow Monitoring
abstract
This paper presents a software-based error detection technique through monitoring flow of the programs in multithreaded architectures. This technique is based on the analysis of two key ideas: 1) Modifying the structure of traditional control-flow graphs used by control-flow checking methods so that they can be applied on multi-core and multi-threaded architectures. These achievements in designing control-flow error detectors lead to increase their applicability in current architectures. 2) Adjusting the locations of additional checking assertions in a given program in order to increase the ability of detecting possible control-flow errors along with significant reduction in overheads. The experimental results, through taking into account both detection coverage and overheads, demonstrate that on average about 94% of the control-flow errors can be detected by the proposed technique, more efficient compared to previous works.
Mohammad Maghsoudloo, Hamid R. Zarandi, Saadat Pour-Mozafari, Navid Khoshavi
DSD4
2011 Control-flow error recovery using commodity multi-core architecture features
abstract
This paper presents a software-based technique to recover control-flow errors using inherent redundancy in commodity multi-core processors. The proposed recovery technique is composed of two phases of control-flow error detection and control-flow error recovery. Previous research shows that modern superscalar microprocessors already contain significant amounts of redundancy. CFEs can be tolerated by leveraging existing microprocessors redundancy. Therefore, the cost of adding extra redundancy for fault tolerance is eliminated.
Navid Khoshavi, Hamid R. Zarandi, Mohammad Maghsoudloo
IOLTS1
2010 Two Efficient Software Techniques to Detect and Correct Control-Flow Errors
abstract
This paper proposes two efficient software techniques, Control-flow and Data Errors Correction using Data-flow Graph Consideration (CDCC) and Miniaturized Check-Pointing (MCP), to detect and correct control-flow errors. These techniques have been implemented based on addition of redundant codes in a given program. The creativity applied in the methods for online detection and correction of the control-flow errors is using data-flow graph alongside of using control-flow graph. These techniques can detect most of the control-flow errors in the program firstly, and next can correct them, automatically. Therefore, both errors in the control-flow and program data which is caused by control-flow errors can be corrected, efficiently. In order to evaluate the proposed techniques, a post compiler is used, so that the techniques can be applied to every 80×86 binaries, transparently. Three benchmarks quick sort, matrix multiplication and linked list are used, and a total of 5000 transient faults are injected on several executable points in each program. The experimental results demonstrate that at least 93% and 89% of the control-flow errors can be detected and corrected without any data error generation by the CDCC and MCP, respectively. Moreover, the strength of these techniques is significant reduction in the performance and memory overheads in compare to traditional methods, for as much as remarkable correction abilities.
Hamid R. Zarandi, Mohammad Maghsoudloo, Navid Khoshavi
PRDC3