VLDB 2026 Research / reviewers in the wild / expert
Ali Nezhadi
dblp:243/8732 · also Ali Nezhadi Khelejani
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-9394-2627ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trust, but Verify: Reliable Compute-in-Memory via Double-Reference Sensing and Selective Recompute
Ali Nezhadi, Odysseas Chatzopoulos, Dimitris Gizopoulos, Mehdi Baradaran Tahoori |
IOLTS | 1 |
| 2026 | Algorithm-Technology Co-Optimization for Reliable NVM-CAM Systems
Ali Nezhadi, Sina Bakhtavari Mamaghani, Mehdi Baradaran Tahoori |
VTS | 1 |
| 2026 | Temporal Reference Scouting Logic for PVT Reliable Logic Computation-in-Memory
Shanmukha Mangadahalli Siddaramu, Ali Nezhadi, Mahta Mayahinia, Sule Ozev, Mehdi Baradaran Tahoori |
VTS | 2 |
| 2025 | Sisyphus: Cross-Layer Efficiency Across NVM Technologies in Compute-in-Memory ArchitecturesabstractCompute-in-Memory (CiM) employing Non-Volatile Memory (NVM) technology is an emerging paradigm that promises higher power efficiency for important data-intensive computations. The performance, power, and resilience properties of emerging NVM technologies determine the efficiency of architectures built around processors and computational memories, and affect design decisions. Thus, fast exploration of the broad design space is necessary to assist decision-making. We present Sisyphus, the first cross-layer framework built to facilitate computer architecture research when such an exploration is required. Sisyphus incorporates detailed technology information for various CiM circuit designs based on STT-MRAM, ReRAM, and PCM technologies and integrates them in fast microarchitecture level system models in gem5 to evaluate performance, power, and resilience (through fault injection) across a large space of design options. Sisyphus’ holistic modeling enables the comprehensive evaluation of all efficiency aspects during the execution of actual workloads on the CPU-CiM architecture. This allows for comparisons to a baseline CPU-only system. In our experimental evaluation, we demonstrate how Sisyphus can derive conclusions regarding the prevalence of one NVM type over another, depending on the prioritized optimization aspect(s). Ali Nezhadi, Odysseas Chatzopoulos, Mahta Mayahinia, George Papadimitriou 0001, Mehdi Baradaran Tahoori, Dimitris Gizopoulos |
ITC | 1 |
| 2024 | SHERLOCK: Scheduling Efficient and Reliable Bulk Bitwise Operations in NVMsabstractBulk bitwise operations are commonplace in application domains such as databases, web search, cryptography, and image processing. The ever-growing volume of data and processing demands of these domains often result in high energy consumption and latency in conventional system architectures, mainly due to data movement between the processing and memory subsystems. Non-volatile memories (NVMs), such as RRAM, PCM and STT-MRAM, facilitate conducting bulk-bitwise logic operations in-memory (CIM). Efficient mapping of complex applications to these CIM-capable NVMs is non-trivial and can even lead to slowdowns. This paper presents Sherlock, a novel mapping and scheduling method for efficient execution of bulk bitwise operations in NVMs. Sherlock collaboratively optimizes for performance and energy consumption and outperforms the state-of-the-art by 10× and 4.6×, respectively. Hamid Farzaneh, João Paulo C. de Lima, Ali Nezhadi, Asif Ali Khan, Mahta Mayahinia, Mehdi Baradaran Tahoori, Jerónimo Castrillón |
DAC | 3 |
| 2024 | Cross-Layer Reliability Evaluation of In-Memory Similarity ComputationabstractThe Memory Wall represents a significant performance and energy bottleneck in conventional computer architecture, caused by the frequent and costly data transfers between memory and processor cores. Computation in Memory (CiM) offers a promising solution, particularly benefiting similarity computation—a key component in various computer science applications—through the use of emerging non-volatile memory (NVM) technologies. However, intrinsic non-idealities of NVM technologies coupled with noise-sensitivity of analog circuitry can impair the reliability and correct functionality of the applications. This paper performs a cross-layer technology to application reliability and performance analysis of NVM-based similarity computation in memory. We consider various NVM technologies, different CiM-based similarity computation modules, and different real-world applications that are accurately simulated in a full-system environment. The results demonstrate significant improvements in both latency (up to 12x) and energy efficiency (up to 6.7x) of NVM-CiM compared to traditional architectures. Importantly, the study finds that latency and energy are largely unaffected by the specific NVM technology used, though reliability varies significantly—up to 28% in terms of architectural vulnerability factor (AVF). Ali Nezhadi, Mahta Mayahinia, Mehdi Baradaran Tahoori |
ITC | 1 |
| 2024 | Hardware and Software Co-Design for Optimized Decoding Schemes and Application Mapping in NVM Compute-in-Memory ArchitecturesabstractThe computation-in nonvolatile memory (NVM-CiM) approach addresses the growing computational demands and the memory-wall problem faced by traditional processor-centric architectures. Computation-in-memory (CiM) capitalizes on the parallel nature of memory arrays enabling effective computation through multirow memristor reading and sensing. In this context, the conventional design of memory decoders needs to be accordingly modified for efficient multirow activation and parallel data processing. This article presents the design and optimization of address decoders for NVM-CiM system architectures, employing a cross-layer co-optimization approach that integrates circuit and architecture design with application requirements. Our methodology starts at the circuit level, examining various decoder designs, including cascaded, hierarchical, latched, and hybrid models. An in-depth application-level characterization follows, utilizing an extended NVM-CiM-capable gem5 simulator to assess the impact of these decoders on the mapping of CiM-friendly applications and the resulting system performance, particularly in facilitating rapid and efficient activation of multirow memory configurations. This holistic analysis allows us to identify the bottlenecks and requirements from the application side and adjust the design of the decoder accordingly. Our analysis reveals that Hybrid Decoders significantly decrease latency and power consumption compared to other decoder designs within NVM-CiM systems. This highlights the crucial role of the decoder’s row selection flexibility, reducing additional system-level data movement even at the expense of its performance, can substantially improve the overall efficiency of NVM-CiM systems. Shanmukha Mangadahalli Siddaramu, Ali Nezhadi, Mahta Mayahinia, Seyedeh Maryam Ghasemi, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | semiMul: Floating-Point Free Implementations for Efficient and Accurate Neural Network TrainingabstractMultiply–accumulate operation (MAC) is a fundamental component of machine learning tasks, where multiplication (either integer or float multiplication) compared to addition is costly in terms of hardware implementation or power consumption. In this paper, we approximate floating-point multiplication by converting it to integer addition while preserving the test accuracy of shallow and deep neural networks. We mathematically show and prove that our proposed method can be utilized with any floating-point format (e.g., FP8, FP16, FP32, etc.). It is also highly compatible with conventional hardware architectures and can be employed in CPU, GPU, or ASIC accelerators for neural network tasks with minimum hardware cost. Moreover, the proposed method can be utilized in embedded processors without a floating-point unit to perform neural network tasks. We evaluated our method on various datasets such as MNIST, FashionMNIST, SVHN, Cifar-10, and Cifar-100, with both FP16 and FP32 arithmetics. The proposed method preserves the test accuracy and, in some cases, overcomes the overfitting problem and improves the test accuracy. Ali Nezhadi, Shaahin Angizi, Arman Roohi |
ICMLA | 1 |