EDBT 2026 Demo / reviewers in the wild / expert
Paria Darbani
dblp:323/5433
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-1173-0936ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoolDawn: A Thermal-Aware Technique to Enhance Lifetime of Neural Network AcceleratorsabstractEffective thermal management has become more essential as Neural Network (NN) accelerators continue to grow in power and complexity. High operating temperatures can degrade performance, accelerate hardware aging, increase power consumption, and raise failure rates. Simultaneously, these accelerators face up to 60% underutilization due to mismatches between network layers and the accelerator architecture. This work presents a technique that leverages idle resources to alleviate thermal hotspots in NN accelerators through strategic workload redistribution, all while preserving performance. Experimental results demonstrate that the proposed technique reduces the time that the NN accelerator spends in hotspot temperature by 47.2% compared to the state-of-the-art, leading to an improvement in the Mean Time to Failure (MTTF) by approximately 20.2%, which extends the lifespan of the NN accelerator by 25.1%. An improvement in MTTF allows for a reduction in the guardband for operating voltage and frequency, potentially leading to better performance and increased power efficiency. Paria Darbani, Nezam Rohbani, AmirReza Imani, Pejman Lotfi-Kamran |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | RADI: A High-Performance Reconfigurable Array-Based Accelerator for DNN ImplementationsabstractDeep neural networks (DNNs) have gained significant attention due to the rapid growth of learning-based applications. However, the computational demands of DNNs limit their performance in many of these applications. As a result, extensive research has focused on hardware implementations of these networks as accelerators. Array-based accelerators are an efficient architecture type that employs an array of processing elements (PEs) for parallel computations. However, array-based accelerators cannot reach their potential performance due to having fixed dimensions to execute different layers of DNNs. This article proposes a reconfigurable architecture to address this limitation by adaptively selecting the size of PEs to better align with the dimensions of the active DNN layers. Simulations demonstrate significant improvements for various DNN models compared to state-of-the-art architectures. Experimental results show that the proposed architecture achieves, on average, 43% higher speed, 32% more resource utilization, and a 38% reduction in on-chip memory access rate compared to the baseline architecture when executing GoogLeNet model layers. These enhancements are achieved with only a 1.6% area overhead, making the proposed architecture a cost-effective design. Furthermore, by incorporating multithreading into the simulator's source code, we significantly accelerate simulations compared to the basic version. Mobina Ranjbar Malidareh, Mojtaba Valinataj, Paria Darbani |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2026 | Corrigendum: RADI: A High-Performance Reconfigurable Array-Based Accelerator for DNN ImplementationsabstractThis is a corrigendum for the article "RADI: A High-Performance Reconfigurable Array-Based Accelerator for DNN Implementations" published in ACM Trans. Des. Autom. Electron. Syst. 31, 5, Article 116 (April 2026), 27 pages. Mobina Ranjbar Malidareh, Mojtaba Valinataj, Paria Darbani |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | RASHT: A Partially Reconfigurable Architecture for Efficient Implementation of CNNsabstractConvolutional neural networks (CNNs) are widely used in machine learning (ML) applications such as image processing. CNN requires heavy computations to provide significant accuracy for many ML tasks. Therefore, the efficient implementations of CNNs to improve performance using limited resources without accuracy reduction is a challenge for ML systems. One of the architectures for the efficient execution of CNNs is the array-based accelerator, that consists of an array of similar processing elements (PEs). The array accelerators are popular as high-performance architecture using the features of parallel computing and data reuse. These accelerators are optimized for a set of CNN layers, not for individual layers. Using the same accelerator dimension size to compute all CNN layers with varying shapes and sizes leads to the resource underutilization problem. We propose a flexible and scalable architecture for array-based accelerator that increases resource utilization by resizing PEs to better match the different shapes of CNN layers. The low-cost partial reconfiguration improves resource utilization and performance, resulting in a 23.2% reduction in computational times of GoogLeNet compared to the state-of-the-art accelerators. The proposed architecture decreases the on-chip memory access rate by 26.5% with no accuracy loss. Paria Darbani, Nezam Rohbani, Hakem Beitollahi, Pejman Lotfi-Kamran |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |