MohammadHosein Gholamrezaei

dblp:308/5975 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-5549-2901ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Characterizing Digital DRAM PIM through Modeling and Benchmarking
abstract
The disparity between processor speed and memory bandwidth has become a growing performance bottleneck, particularly for memory-intensive workloads. Processing-in-Memory (PIM) mitigates this bottleneck by integrating computation directly within DRAM. However, the effectiveness of PIM varies significantly across workloads, architectures, and DRAM technologies, yet it is often assessed using tightly coupled simulators and benchmarks that lack portability and generality. This article extends PIMbench and PIMeval —a generalizable benchmark suite and an extensible PIM simulator to support a broader range of workloads, PIM architectures, and DRAM technologies. This evaluation incorporates roofline analysis and a breakdown of intra-memory execution stages to identify PIM-specific bottlenecks and performance scaling limits. The evaluation spans three classes of digital PIM architectures: subarray-level bit-serial, subarray-level bit-parallel, and bank-level bit-parallel. It further demonstrates how internal DRAM parameters such as subarray count and GDL width impact PIM performance. The code is publicly available at: https://github.com/UVA-LavaLab/PIMeval-PIMbench .
Farzana Siddique, Deyuan Guo, Hugo Abbot, Kyle Durrer, MohammadHosein Gholamrezaei, Morteza Baradaran, Ethan Ermovick, Alif Ahmed, Zhenxing Fan, Beenish Gul, Ashish Venkat, Kevin Skadron
ACM Trans. Archit. Code Optim.5
2026 Optimization and Benchmarking of Monolithically Stackable Gain Cell Memory for Last-Level Cache
abstract
The Last Level Cache (LLC) is the processor’s critical bridge between on-chip and off-chip memory levels - optimized for high density, high bandwidth, and low operation energy. To date, high-density (HD) SRAM has been the conventional device of choice; however, with the slowing of transistor scaling, as reflected in the industry’s almost identical HD SRAM cell size from 5 nm to 3 nm, alternative solutions such as 3D stacking with advanced packaging (i.e., hybrid bonding) are pursued (as demonstrated in AMD’s V-cache). Escalating data demands necessitate ultra-large on-chip caches to decrease costly off-chip memory movement, pushing the exploration of device technology towards monolithic 3D (M3D) integration, where transistors can be stacked in the back-end-of-line (BEOL) at the interconnect level. M3D integration requires fabrication techniques compatible with a low thermal budget (seconds) when used in a gain-cell configuration. This paper examines device, circuit, and system-level tradeoffs made when optimizing BEOL-compatible AOS-based 2-transistor gain cells (2T-GC) for LLC. A cache early-exploration tool, NS-Cache, is developed to model caches in advanced 7 & 3 nm nodes and is integrated with the Gem5 simulator to systematically benchmark the impact of the newfound density/performance when compared to HD-SRAM, MRAM, and 1T1C eDRAM alternatives for LLC.
Faaiq G. Waqar, Jungyoun Kwak, Omkar Phadke, Minji Shon, MohammadHosein Gholamrezaei, Kevin Skadron, Shimeng Yu
IEEE Trans. Computers6
2023 A Generalized Residue Number System Design Approach for Ultralow-Power Arithmetic Circuits Based on Deterministic Bit-Streams
abstract
The peak power consumption has become an important concern in the hardware design process of some of today’s applications, such as energy harvesting (EH) and bio-implantable (BI) electronic devices. The limited peak harvested power in EH devices and heating concerns in BI devices are the main reasons for power control’s importance in these devices. This article proposes a generalized design approach for ultralow-power arithmetic circuits. The proposed circuits are based on residue number system (RNS) combined with deterministic bit-streams. The resulting circuits can be used in systems with a restricted power budget. We suggest several approaches to design generic hardware-efficient adders, multipliers, multiply-accumulate (MAC) unit, forward converters (FCs), and reverse converters (RCs). Using the proposed approach, designing these components for any moduli of the RNS can be performed through simple bit-width adjustments in the circuits. The synthesis results show that the proposed adder achieves, on average, 69% and 2% lower area compared to the bit-serial and a state-of-the-art RNS adder, respectively. Furthermore, the proposed multiplier outperforms the bit-serial, interleaved, and a state-of-the-art design for multiplying RNS numbers by, on average, 57%, 60%, and 77% in terms of power consumption, respectively. The efficiency of our approach is shown via two essential applications, digital signal processing, and machine learning. We implement an FFT engine using the proposed method. Compared to prior RNS implementations, our design achieves 47% lower power consumption. We also implement a CNN accelerator’s processing element (PE) with the proposed computation elements. Our design provides considerable speedup and lower power consumption compared to a state-of-the-art ultralower-power design.
Kamyar Givaki, Ahmad Khonsari, MohammadHosein Gholamrezaei, Saeid Gorgin 0001, M. Hassan Najafi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 HyperDbg: Reinventing Hardware-Assisted Debugging
abstract
Software analysis, debugging, and reverse engineering have a crucial impact in today's software industry. Efficient and stealthy debuggers are especially relevant for malware analysis. However, existing debugging platforms fail to address a transparent, effective, and high-performance low-level debugger due to their detectable fingerprints, complexity, and implementation restrictions.
Mohammad Sina Karvandi, MohammadHosein Gholamrezaei, Saleh Khalaj Monfared, Soroush Meghdadi Zanjani, Behrooz Abbassi, Reza Mortazavi, Saeid Gorgin 0001, Dara Rahmati, Michael Schwarz 0001
CCS2
2022 An Efficient FPGA Implementation of k-Nearest Neighbors via Online Arithmetic
abstract
k-NN, as one of the well-employed classification algorithms, severely suffers from a computationally intensive nature. This paper exploits the parallelism and digit level pipelining opportunities via FPGA devices and Online arithmetic to offer an efficient k-NN FPGA implementation. All the required operations for computing distances and sorting are applied to serially coming data. Moreover, we dynamically terminate the unnecessary computations once they are detected. To the best of our knowledge, the proposed k-NN implementation is the first one that used FPGA and Online arithmetic effectively. It provides up to 34% speedup compared to the best state-of-the-art design.
Saeid Gorgin 0001, MohammadHosein Gholamrezaei, Danial Javaheri, Jeong-A Lee
FCCM2
2022 An Energy-Efficient K-means Clustering FPGA Accelerator via Most-Significant Digit First Arithmetic
abstract
K-means clustering is the most well-known unsupervised learning method that partitions the input dataset into$K$clusters based on the similarity between the data samples. In this paper, to achieve an energy-efficient implementation without sacrificing performance, we take advantage of massive parallelism and digit-level pipelining via FPGA and the most-significant digit first arithmetic. Having the result of the most-significant digits in advance provides the possibility of early termination for unnecessary computations and fetching just the required most-significant part of data points from memory. This early termination technique significantly increases performance and decreases energy consumption. Our experimental results from various datasets and comparisons with the state-of-the-art FPGA accelerators indicate that our proposed design has effectively reduced energy consumption without any performance loss.
Saeid Gorgin 0001, MohammadHosein Gholamrezaei, Danial Javaheri, Jeong-A Lee
FPT2