Mohammad Arman Soleimani

dblp:327/2113 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0007-7719-2556ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 WISEDRAM: A Reliable Bitwise In-DRAM Accelerator
abstract
Processing-in-Memory (PIM) aims to address the costly data movement between processing elements and memory subsystem, by computing simple operations inside DRAM in parallel. The large capacity, wide activation size during cell access, and the maturity of DRAM technology, make this technology a great choice for PIM techniques. Nonetheless, vulnerability to process variation and noises, internal leakage of the cells, and high latency in cell access, limit the utilization of processing in DRAMs for real-world applications. This work proposes a fast PIM technique, called WISEDRAM, which leverages one row of special cells, called X-cells, to enable in-DRAM bulk-bitwise operations. Unlike previous approaches, WISEDRAM retains the conventional DRAM cell access procedure, thereby ensuring the reliability of cell access for reads and writes at a level equivalent to that of conventional DRAMs. Compared with the state-of-the-art, WISEDRAM exhibits $22 \%$ reduction in average bitwise computation latency and a $71 \%$ improvement in XOR operation execution speed, while imposing an area overhead of $1.6 \%$.
Mohammad Arman Soleimani, Nezam Rohbani, Adrián Cristal, Osman S. Unsal, Hamid Sarbazi-Azad
DAC1
2025 BIMAX: A Bitwise In-Memory Accelerator Using 6T-SRAM Structure
abstract
In-memory computing (IMC) paradigm reduces costly and inefficient data transfer between memory modules and processing cores by implementing simple and parallel operations inside the memory subsystem. SRAM, the fastest memory structure in the memory hierarchy, is an appropriate platform to implement IMC. However, the main challenges of implementing IMC in SRAM are the limited operations and unreliable accuracy due to environmental noise and process variations. This work proposes a low-latency, energy-efficient, and noise-robust IMC technique, called Bitwise In-Memory Accelerator using 6T-SRAM Structure (BIMAX). BIMAX performs parallel bitwise operations (i.e., (N)AND, (N)OR, NOT, X(N)OR) as well as row-copy with the capability of writing the computation result back to a target memory row. BIMAX functionality is based on an imbalanced differential sense amplifier (SA) that reads and writes data from and into multiple 6T-SRAM cells. The simulations show BIMAX performs these operations with 52.7% lower energy dissipation compared to the state-of-the-art IMC technique, with 5.7% average higher performance rate. Furthermore, BIMAX is about 5.4× more robust against environmental noises compared to the state-of-the-art.
Nezam Rohbani, Mohammad Arman Soleimani, Behzad Salami 0001, Osman S. Unsal, Adrián Cristal, Hamid Sarbazi-Azad
DATE2
2025 DIST: Distributed Learning-Based Energy-Efficient and Reliable Task Scheduling and Resource Allocation in Fog Computing
abstract
This paper presents DIST, a novel distributed reinforcement learning-based (DRL) framework for energyefficient and reliable task scheduling and resource allocation in fog computing, low-latency computing solutions driven by the rapid deployment of IoT devices, and time-sensitive applications. DIST is built based on a novel distributed Q-learning to enable fog nodes to learn an optimal strategy to balance energy consumption, task execution time, and system reliability. The main novelty includes a cooperative Dynamic Voltage and Frequency Scaling-enabled task scheduling policy that dynamically adjusts node energy level to ensure power consumption reduction without sacrificing deadline adherence or reliability. The results demonstrate that DIST reduces energy consumption by up to 52.26%, realizes 38% higher success rates, and reduces task wait times by up to 46.77%, compared with state-of-the-art algorithms.
Elyas Oustad, Abolfazl Younesi, Mohsen Ansari, Sepideh Safari, Mohammad Arman Soleimani, Jörg Henkel, Alireza Ejlali
IEEE Trans. Serv. Comput.5
2024 A Built-In Integrated Rowhammer, Rowpress, and Leakage Detection Sensor for DRAM
abstract
The increasing density of DRAM chips has led to heightened susceptibility of memory cells to bit-flips. Rowhammer and Rowpress attacks on DRAMs have obtained significant attention due to their effectiveness in compromising system security and data integrity. A substantial portion of the previously proposed techniques for detecting rowhammer attacks focus on counting the number of row activations. However, these methods suffer from several limitations, including high overheads, lack of consideration for environmental conditions during DRAM operation, and limited tunability to detect rowpress attacks. To address these challenges, this work introduces a novel low-overhead built-in DRAM sensor designed specifically for detecting and locating rowhammer and rowpress attacks. The core idea behind the sensor is to utilize an additional non-operational DRAM sensor cell for each row. This auxiliary cell experiences the same potential rowhammer attacks as the other cells within that row. The sensor operates by controlling an extra precharged match-line based on the charge stored in the sensor cells. This sensor not only detects rowhammer attacks but also other environmental conditions that may impact DRAM cells retention time, such as temperature elevation or changes in operating voltage. Simulation results demonstrate that the proposed sensor detects all of the rowhammer attacks in the presence of process variation with a variance of up to 7.5% in circuit parameters. Furthermore, the area overhead introduced by the sensor remains about 1%, making it a promising solution for enhancing DRAM security with the minimum memory structure change.
Nezam Rohbani, Rouzbeh Pirayadi, Mohammad Arman Soleimani, Adrián Cristal, Osman S. Unsal, Hamid Sarbazi-Azad
ICCAD3
2023 CoolDRAM: An Energy-Efficient and Robust DRAM
abstract
DRAM is the most mature and widely-utilized memory structure as main memory in computing systems. However, energy dissipation and latency of DRAM are two of the most serious limiting factors of this technology. All DRAM main operations are initiated by a Precharge phase, which is time-consuming and power-hungry. This work proposes a novel DRAM cell access scheme that entirely eliminates Precharge phase from DRAM read, write, and refresh operations, with a very slight modification in commodity DRAM structure. The proposed DRAM design, called CoolDRAM, operates using a single extra cell row as reference cells. CoolDRAM reduces energy dissipation by about 34% on average, with a negligible area overhead of about 0.4%. The robustness of CoolDRAM against process variation and environmental noises is 61× and 1.78 × of the state-of-the-art, respectively, while maintaining the same power consumption and latency.
Nezam Rohbani, Mohammad Arman Soleimani, Hamid Sarbazi-Azad
ISLPED2
2022 PIPF-DRAM: processing in precharge-free DRAM
abstract
To alleviate costly data communication among processing cores and memory modules, parallel processing-in-memory (PIM) is a promising approach which exploits the huge available internal memory bandwidth. High capacity, wide row size, and maturity of DRAM technology, make DRAM an alluring structure for PIM. However, dense layout, high process variation, and noise vulnerability of DRAMs make it very challenging to apply PIM for DRAMs in practice. This work proposes a PIM structure which eliminates these DRAM limitations, exploiting a precharge-free DRAM (PF-DRAM) structure. The proposed PIM structure, called PIPF-DRAM, performs parallel bitwise operations only by modifying control signal sequences in PF-DRAM, with almost zero structural and circuit modifications. Comparing the state-of-the-art PIM techniques, PIPF-DRAM is 4.2× more robust to process variation, 4.1% faster in average cycle time of operations, and consumes 66.1% less energy.
Nezam Rohbani, Mohammad Arman Soleimani, Hamid Sarbazi-Azad
DAC2