EDBT 2026 Demo / reviewers in the wild / expert
Nezam Rohbani
dblp:164/8041
· DBLP profile ↗
19ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-1935-7830ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 10 first-author · 13 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoolDawn: A Thermal-Aware Technique to Enhance Lifetime of Neural Network AcceleratorsabstractEffective thermal management has become more essential as Neural Network (NN) accelerators continue to grow in power and complexity. High operating temperatures can degrade performance, accelerate hardware aging, increase power consumption, and raise failure rates. Simultaneously, these accelerators face up to 60% underutilization due to mismatches between network layers and the accelerator architecture. This work presents a technique that leverages idle resources to alleviate thermal hotspots in NN accelerators through strategic workload redistribution, all while preserving performance. Experimental results demonstrate that the proposed technique reduces the time that the NN accelerator spends in hotspot temperature by 47.2% compared to the state-of-the-art, leading to an improvement in the Mean Time to Failure (MTTF) by approximately 20.2%, which extends the lifespan of the NN accelerator by 25.1%. An improvement in MTTF allows for a reduction in the guardband for operating voltage and frequency, potentially leading to better performance and increased power efficiency. Paria Darbani, Nezam Rohbani, AmirReza Imani, Pejman Lotfi-Kamran |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | WISEDRAM: A Reliable Bitwise In-DRAM AcceleratorabstractProcessing-in-Memory (PIM) aims to address the costly data movement between processing elements and memory subsystem, by computing simple operations inside DRAM in parallel. The large capacity, wide activation size during cell access, and the maturity of DRAM technology, make this technology a great choice for PIM techniques. Nonetheless, vulnerability to process variation and noises, internal leakage of the cells, and high latency in cell access, limit the utilization of processing in DRAMs for real-world applications. This work proposes a fast PIM technique, called WISEDRAM, which leverages one row of special cells, called X-cells, to enable in-DRAM bulk-bitwise operations. Unlike previous approaches, WISEDRAM retains the conventional DRAM cell access procedure, thereby ensuring the reliability of cell access for reads and writes at a level equivalent to that of conventional DRAMs. Compared with the state-of-the-art, WISEDRAM exhibits $22 \%$ reduction in average bitwise computation latency and a $71 \%$ improvement in XOR operation execution speed, while imposing an area overhead of $1.6 \%$. Mohammad Arman Soleimani, Nezam Rohbani, Adrián Cristal, Osman S. Unsal, Hamid Sarbazi-Azad |
DAC | 2 |
| 2025 | BIMAX: A Bitwise In-Memory Accelerator Using 6T-SRAM StructureabstractIn-memory computing (IMC) paradigm reduces costly and inefficient data transfer between memory modules and processing cores by implementing simple and parallel operations inside the memory subsystem. SRAM, the fastest memory structure in the memory hierarchy, is an appropriate platform to implement IMC. However, the main challenges of implementing IMC in SRAM are the limited operations and unreliable accuracy due to environmental noise and process variations. This work proposes a low-latency, energy-efficient, and noise-robust IMC technique, called Bitwise In-Memory Accelerator using 6T-SRAM Structure (BIMAX). BIMAX performs parallel bitwise operations (i.e., (N)AND, (N)OR, NOT, X(N)OR) as well as row-copy with the capability of writing the computation result back to a target memory row. BIMAX functionality is based on an imbalanced differential sense amplifier (SA) that reads and writes data from and into multiple 6T-SRAM cells. The simulations show BIMAX performs these operations with 52.7% lower energy dissipation compared to the state-of-the-art IMC technique, with 5.7% average higher performance rate. Furthermore, BIMAX is about 5.4× more robust against environmental noises compared to the state-of-the-art. Nezam Rohbani, Mohammad Arman Soleimani, Behzad Salami 0001, Osman S. Unsal, Adrián Cristal, Hamid Sarbazi-Azad |
DATE | 1 |
| 2024 | A Built-In Integrated Rowhammer, Rowpress, and Leakage Detection Sensor for DRAMabstractThe increasing density of DRAM chips has led to heightened susceptibility of memory cells to bit-flips. Rowhammer and Rowpress attacks on DRAMs have obtained significant attention due to their effectiveness in compromising system security and data integrity. A substantial portion of the previously proposed techniques for detecting rowhammer attacks focus on counting the number of row activations. However, these methods suffer from several limitations, including high overheads, lack of consideration for environmental conditions during DRAM operation, and limited tunability to detect rowpress attacks. To address these challenges, this work introduces a novel low-overhead built-in DRAM sensor designed specifically for detecting and locating rowhammer and rowpress attacks. The core idea behind the sensor is to utilize an additional non-operational DRAM sensor cell for each row. This auxiliary cell experiences the same potential rowhammer attacks as the other cells within that row. The sensor operates by controlling an extra precharged match-line based on the charge stored in the sensor cells. This sensor not only detects rowhammer attacks but also other environmental conditions that may impact DRAM cells retention time, such as temperature elevation or changes in operating voltage. Simulation results demonstrate that the proposed sensor detects all of the rowhammer attacks in the presence of process variation with a variance of up to 7.5% in circuit parameters. Furthermore, the area overhead introduced by the sensor remains about 1%, making it a promising solution for enhancing DRAM security with the minimum memory structure change. Nezam Rohbani, Rouzbeh Pirayadi, Mohammad Arman Soleimani, Adrián Cristal, Osman S. Unsal, Hamid Sarbazi-Azad |
ICCAD | 1 |
| 2024 | An Efficient FPGA Architecture with Turn-Restricted Switch BoxesabstractAbstract. Field-Programmable Gate Arrays (FPGAs) employ a large number of SRAM cells to provide a flexible routing architecture which have a significant impact on the FPGA’s area and power consumption. This flexible routing allows for a rather easy realization of the desired functionality, but our evaluations show that the full routing flexibility is not required in many occasions. In this work, we focus on what is actually needed and introduce a new switch-box realization what we call Turn-Restricted Switch-Boxes which supports only a subset of possible turns. The proposed method increases the utilization rate of FPGA switch-boxes by eliminating the unemployed resources. Experimental evaluations confirm that the area and average power consumption can be reduced by 12.8% and 14.1%, on average, respectively and the FPGA routing susceptibility to SEU and MBU can be improved by 18.2%, on average, by imposing negligible performance. 1 Fatemeh Serajeh-hassani, Mohammad Sadrosadati, Nezam Rohbani, Sebastian Pointner, Robert Wille, Hamid Sarbazi-Azad |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | CoolDRAM: An Energy-Efficient and Robust DRAMabstractDRAM is the most mature and widely-utilized memory structure as main memory in computing systems. However, energy dissipation and latency of DRAM are two of the most serious limiting factors of this technology. All DRAM main operations are initiated by a Precharge phase, which is time-consuming and power-hungry. This work proposes a novel DRAM cell access scheme that entirely eliminates Precharge phase from DRAM read, write, and refresh operations, with a very slight modification in commodity DRAM structure. The proposed DRAM design, called CoolDRAM, operates using a single extra cell row as reference cells. CoolDRAM reduces energy dissipation by about 34% on average, with a negligible area overhead of about 0.4%. The robustness of CoolDRAM against process variation and environmental noises is 61× and 1.78 × of the state-of-the-art, respectively, while maintaining the same power consumption and latency. Nezam Rohbani, Mohammad Arman Soleimani, Hamid Sarbazi-Azad |
ISLPED | 1 |
| 2023 | Power-Efficient and Aging-Aware Primary/Backup Technique for Heterogeneous Embedded SystemsabstractOne of the essential requirements of embedded systems is a guaranteed level of reliability. In this regard, fault-tolerance techniques are broadly applied to these systems to enhance reliability. However, fault-tolerance techniques may increase power consumption due to their inherent redundancy. For this purpose, power management techniques are applied, along with fault-tolerance techniques, which generally prolong the system lifespan by decreasing the temperature and leading to an aging rate reduction. Yet, some power management techniques, such as Dynamic voltage and frequency scaling (DVFS), increase the transient fault rate and timing error. For this reason, heterogeneous multicore platforms have received much attention due to their ability to make a trade-off between power consumption and performance. Still, it is more complicated to map and schedule tasks in a heterogeneous multicore system. In this paper, for the first time, we propose a power management method for a heterogeneous multicore system that reduces power consumption and tolerates both transient and permanent faults through primary/backup technique while considering core-level power constraint, real-time constraint, and aging effect. Experimental evaluations demonstrate the efficiency of our proposed method in terms of reducing power consumption compared to the state-of-the-art schemes, together with guaranteeing reliability and considering the aging effect. Mohsen Ansari, Sepideh Safari, Nezam Rohbani, Alireza Ejlali, Bashir M. Al-Hashimi |
IEEE Trans. Sustain. Comput. | 3 |
| 2022 | PIPF-DRAM: processing in precharge-free DRAMabstractTo alleviate costly data communication among processing cores and memory modules, parallel processing-in-memory (PIM) is a promising approach which exploits the huge available internal memory bandwidth. High capacity, wide row size, and maturity of DRAM technology, make DRAM an alluring structure for PIM. However, dense layout, high process variation, and noise vulnerability of DRAMs make it very challenging to apply PIM for DRAMs in practice. This work proposes a PIM structure which eliminates these DRAM limitations, exploiting a precharge-free DRAM (PF-DRAM) structure. The proposed PIM structure, called PIPF-DRAM, performs parallel bitwise operations only by modifying control signal sequences in PF-DRAM, with almost zero structural and circuit modifications. Comparing the state-of-the-art PIM techniques, PIPF-DRAM is 4.2× more robust to process variation, 4.1% faster in average cycle time of operations, and consumes 66.1% less energy. Nezam Rohbani, Mohammad Arman Soleimani, Hamid Sarbazi-Azad |
DAC | 1 |
| 2022 | LETHOR: a thermal-aware proactive routing algorithm for 3D NoCs with less entrance to hot regions
Maede Safari, Zahra Shirmohammadi, Nezam Rohbani, Hamed Farbeh |
J. Supercomput. | 3 |
| 2022 | RASHT: A Partially Reconfigurable Architecture for Efficient Implementation of CNNsabstractConvolutional neural networks (CNNs) are widely used in machine learning (ML) applications such as image processing. CNN requires heavy computations to provide significant accuracy for many ML tasks. Therefore, the efficient implementations of CNNs to improve performance using limited resources without accuracy reduction is a challenge for ML systems. One of the architectures for the efficient execution of CNNs is the array-based accelerator, that consists of an array of similar processing elements (PEs). The array accelerators are popular as high-performance architecture using the features of parallel computing and data reuse. These accelerators are optimized for a set of CNN layers, not for individual layers. Using the same accelerator dimension size to compute all CNN layers with varying shapes and sizes leads to the resource underutilization problem. We propose a flexible and scalable architecture for array-based accelerator that increases resource utilization by resizing PEs to better match the different shapes of CNN layers. The low-cost partial reconfiguration improves resource utilization and performance, resulting in a 23.2% reduction in computational times of GoogLeNet compared to the state-of-the-art accelerators. The proposed architecture decreases the on-chip memory access rate by 26.5% with no accuracy loss. Paria Darbani, Nezam Rohbani, Hakem Beitollahi, Pejman Lotfi-Kamran |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | PF-DRAM: A Precharge-Free DRAM StructureabstractAlthough DRAM capacity and bandwidth have increased sharply by the advances in technology and standards, its latency and energy per access have remained almost constant in recent generations. The main portion of DRAM power/energy is dissipated by Read, Write, and Refresh operations, all initiated by a Precharge phase. Precharge phase not only imposes a large amount of energy consumption, but also increases the delay of closing a row in a memory block to open another one. By reduction of row-hit rate in recent workloads, especially in multi-core systems, precharge rate increases which exacerbates DRAM power dissipation and access latency. This work proposes a novel DRAM structure, called Precharge-Free DRAM (PF-DRAM), that eliminates the Precharge phase of DRAM. PF-DRAM uses the charge on bitlines from the previous Activation phase, as the starting point for the next Activation. The difference between PF-DRAM and conventional DRAM structure is limited to precharge and equalizer circuitry and simple modifications in sense amplifier, which are all limited to subarray level. PF-DRAM is compatible with the mainstream JEDEC memory standards like DDRx and HBM, with minimum modifications in memory controller. Furthermore, almost all of the previously proposed power/energy reduction techniques in DRAM are still applicable to PF-DRAM for further improvement. Our experimental results on a 8GB memory system running SPEC CPU2017 and PARSEC2.1 workloads show an average of 35.3% memory power consumption reduction (up to 54.2%) achieved by the system using PF-DRAM with respect to the system using conventional DRAM. Moreover, the overall performance is improved by 8.6%, in average (up to 24.3%). According to our analysis, all such improvements are achieved for less than 9% area overhead. Nezam Rohbani, Sina Darabi, Hamid Sarbazi-Azad |
ISCA | 1 |
| 2021 | SRAM Gauge: SRAM Health Monitoring via Cells RaceabstractBy shrinking transistors’ dimensions and, consequently, reducing the operating voltage in nano-scale CMOS technologies, the stability of SRAM cells has become a major reliability concern. SRAM cells’ robustness against undesirable bit-flips is commonly measured by Static Noise Margin (SNM). Degradation in SNM is mainly because of the gradual variations in transistors’ parameters due to aging. This work proposes a built-in SRAM health sensor capable of monitoring the SNM of individual SRAM cells in a memory block. The sensor is composed of extra non-operational sensor cells with different predefined SNMs. These sensor cells are put in a race with operational SRAM cells to determine their strength. The precision, sensing range, and robustness of the proposed sensor against process variation are adjustable at the cost of small area overhead. In our simulation setup, with the area overhead of 0.29%, the sensor monitors a wide range of SNMs from 275mV to 325mV, with a precision of 5mV. Nezam Rohbani, Masoumeh Ebrahimi |
ISLPED | 1 |
| 2021 | TAMER: an adaptive task allocation method for aging reduction in multi-core embedded real-time systems
Faezeh Sadat Saadatmand, Nezam Rohbani, Farshad Baharvand, Hamed Farbeh |
J. Supercomput. | 2 |
| 2019 | NVDL-Cache: Narrow-Width Value Aware Variable Delay Low-Power Data CacheabstractCache memories dissipate a large portion of processors' power budget. On the other hand, due to unbalanced stress condition on their SRAMs, aging of cache memories is one of the most challenging reliability issues in modern processors. Therefore, power management and aging mitigation of these memories are mandatory in recent nanoscale technologies to achieve reliable and stable functionality of processor. Regarding the fact that the rate of Narrow-Width Values (NWVs) stored in data-caches memory is more than 80%, this work proposes an NWV-aware power consumption reduction and aging mitigation technique. In the proposed data-cache memory, the operating voltage of memory blocks which store the most significant bits of cache words is adjustable according to the rate of NWVs to reduce cache power consumption and aging rate. Our simulations show the proposed technique decreases the overall power dissipation of a 32KB cache memory by 44.20%, with 0.55% and 2.4% performance and area overheads, respectively. This technique prolongs the lifetime of the cache by up to 1.96x. Nezam Rohbani, Tapas K. Maiti, Dondee Navarro, Mitiko Miura-Mattausch, Hans Jürgen Mattausch, Hirotaka Takatsuka |
ICCD | 1 |
| 2019 | Power Reduction and BTI Mitigation of Data-Cache Memory Based on the Storage Management of Narrow-Width ValuesabstractPower dissipation of on-chip cache memories contributes a large portion of a processor's power consumption. Therefore, power management of cache memories is crucial in modern processors. On the other hand, bias temperature instability (BTI) is one of the most serious reliability concerns in SRAM-based on-chip memories. The effect of BTI on SRAM-cell transistors is manifested as their threshold-voltage shift over stress duration, which decreases the robustness of these structures. This paper presents a power consumption-reduction technique for data-cache memories, based on a storage management considering narrow-width values (NWVs), which additionally mitigates the BTI rate on the most aging susceptible SRAM cells of data-cache memory, as well. In the proposed technique, the most significant bits of data-cache memory words are stored in SRAM blocks that operate with lower VDD, to achieve the decrease of the related power consumption and aging rate. This is shown to reduce leakage-current and dynamic current of data-cache memory by 37.6% and 22.1%, respectively, besides improving its lifetime by up to 2.25x, all with negligible performance and area overheads. Nezam Rohbani, Hiroaki Gau, Sara Mohammadinejad, Tapas K. Maiti, Dondee Navarro, Mitiko Miura-Mattausch, Hans Jürgen Mattausch, Hirotaka Takatsuka |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | LAXY: A Location-Based Aging-Resilient Xy-Yx Routing Algorithm for Network on ChipabstractNetwork on chip (NoC) is a scalable interconnection architecture for ever increasing communication demand between processing cores. However, in nanoscale technology size, NoC lifetime is limited due to aging processes of negative bias temperature instability, hot carrier injection, and electromigration. Usually, because of unbalanced utilization of NoC resources, some parts of the network experience more thermal stress and duty cycle in comparison with other parts, which may accelerate chip failure. To slow down the aging rate of NoC, this paper proposes an oblivious routing algorithm called location-based aging-resilient Xy-Yx (LAXY) to distribute packet flow over entire network. LAXY is based on the fact that dimension-ordered routing algorithms imposes the highest traffic load on the central nodes in mesh topologies. To balance the traffic over the network, certain routers at the east and the west of NoC, with dimension-order XY routing, statically are configured as YX. Various configurations have been explored for LAXY and the simulations show a specific configuration, called Fishtail, increases mean time to failure of the routers and interconnects by about 42% and 56%, respectively. Moreover, by balancing the load over the network, LAXY improves overall packet latency by about 7% in average, with negligible area overhead. Nezam Rohbani, Zahra Shirmohammadi, Maryam Zare, Seyed Ghassem Miremadi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | A Low Area Overhead NBTI/PBTI Sensor for SRAM MemoriesabstractBias temperature instability (BTI) is known as one serious reliability concern in nanoscale technologies. BTI gradually increases the absolute value of threshold voltage (Vth) of MOS transistors. The main consequence of Vth shift of the SRAM cell transistors is the static noise margin (SNM) degradation. The SNM degradation of SRAM cells results in bit-flip occurrences due to transient faults and should be monitored accurately. This paper proposes a sensor called write current-based BTI sensor (WCBS) to assess the BTI-aging state of SRAM cells. The WCBS measures BTI-induced SNM degradation of SRAM cells by monitoring the maximum write current shifts due to BTI. The observations show that the maximum current consumption during write operation is an effective identifier to measure Vth and SNM shifts. The granularity of BTI assessment of one cell up to a row of memory can be achieved by writing special bit patterns on the memory block during the test. We evaluated the sensor through SPICE-level simulations in 32-nm technology size. The precision of WCBS is about ±1.25 mV (±3.2% error). One sensor is enough for the entire SRAM memory block with negligible area/power overhead; less than 1%. The effects of process variation and temperature changes on WCBS are investigated in detail. Nezam Rohbani, Seyed Ghassem Miremadi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Bias Temperature Instability Mitigation via Adaptive Cache Size ManagementabstractBias temperature instability (BTI) is one of the major CMOS reliability issues in nanoscales. The main impact of BTI on SRAM memory cells is the degradation of the static noise margin (SNM), which leads to a higher susceptibility to failures. A variety of techniques for mitigating the impact of BTI on caches have been proposed at architecture level. However, their considerable overheads limit the application of such techniques. Recent studies showed that the utilization of the cache capacity widely varies from one workload to another and even within a workload. When cache utilization is low, for the majority of the cells, the same value is stored for a very long period, which significantly degrades SNM due to BTI. In this paper, we propose a technique to dynamically adjust the cache size according to the running workload cache requirement by monitoring the cache miss rate. The unused cache capacity is power gated to increase the energy efficiency and mitigate aging of the entire cache. The experimental results show that the proposed technique reduces hold and read SNM degradation by up to 48.1% and 33.3%, respectively, at the cost of 2.0% performance penalty. Nezam Rohbani, Mojtaba Ebrahimi, Seyed Ghassem Miremadi, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | A fault-tolerant and energy-aware mechanism for cluster-based routing algorithm of WSNsabstractWireless Sensor Networks (WSNs) are prone to faults due to battery depletion of nodes. A node failure can disturb routing as it plays a key role in transferring sensed data to the end users. This paper presents a Fault-Tolerant and Energy-Aware Mechanism (FTEAM), which prolongs the lifetime of WSNs. This mechanism can be applied to cluster-based WSN protocols. The main idea behind the FTEAM is to identify overlapped nodes and configure the most powerful ones to the sleep mode to save their energy for the purpose of replacing a failed Cluster Head (CH) with them. FTEAM not only provides fault tolerant sensor nodes, but also tackles the problem of emerging dead area in the network. Our experimental results and simulations show that FTEAM outperforms conventional protocols in terms of network lifetime and energy consumption. In addition, an analytical evaluation using the Markov model is performed to determine the reliability of the FTEAM. Maryam Hezaveh, Zahra Shirmohammadi, Nezam Rohbani, Seyed Ghassem Miremadi |
IM | 3 |