EDBT 2026 Demo / reviewers in the wild / expert
Dara Rahmati
dblp:56/6575
· DBLP profile ↗
22ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-0104-4016ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 5 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Modular Addition for FPGA-Based Cryptographic OperationsabstractModular adders are essential components in finite field arithmetic, serving as key components in public-key cryptographic algorithms like Elliptic Curve Cryptography (ECC) and Post-Quantum Cryptography (PQC). Naïve implementation of modular adders, due to two cascaded adders with large operand bit widths struggle to meet high-frequency requirements. On the other hand, parallel implementations, while faster, demand excessive resources and power, making them impractical for many applications. This paper introduces a novel modular addition algorithm leveraging a novel operand representation based on the two-valued digit encoding (Twit). In this approach, each operand is represented as an n-bit unsigned number augmented by a Twit value {0,±δ}. The algorithm efficiently computes modular addition by speculating and dynamically adjusting the twit value in the result, achieving both computational and resource efficiency. The proposed design has been implemented on a Xilinx 7-series FPGA, demonstrating superior performance in achieving high operating frequencies (i.e., 8% to 36% depending on operand bit widths) while significantly reducing resource utilization (i.e., >36%). In addition to extensive analytical and synthesis-based evaluations, we further demonstrate the benefits of the proposed adder within application-level cryptographic datapaths (ECC and PQC). Saeid Gorgin 0001, Amirhossein Sadr, Dara Rahmati, Ali Jahanian 0001, Jungrae Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | Efficient neural network acceleration using redundant residue number systems
Soudabeh Mousavi, Dara Rahmati, Amirhossein Sadr, Saeid Gorgin 0001, Jungrae Kim |
J. Supercomput. | 2 |
| 2026 | FPGA-accelerated real-time DCGANs via Xilinx DPUs and Vitis AI
Amirhossein Sadr, Aida Pakniyat, Dara Rahmati, Saeid Gorgin 0001 |
J. Supercomput. | 3 |
| 2025 | A Generic Modulo-(2n±δ) Addition Algorithm via Two-Valued Digit EncodingabstractModular adders are essential arithmetic components in Residue Number System (RNS)-based applications, including digital signal processing, cryptography, and machine learning. These applications consistently push the boundaries of dynamic range (DR) and operating frequency, making the design of efficient generic modular adders a critical and evolving challenge. This paper presents a novel algorithm for modulo$-(2^{n}\pm \delta)$addition, where$\delta$is an integer within the range$0\leq\delta\leq 2^{n-1}-1$. The proposed approach leverages a two-valued digit (twit) for encoding the value of$\pm\delta$and uses a faithful representation of operands. In this representation, each operand is encoded as an n-bit unsigned number augmented by a twit value$\{0,\pm\delta\}$. The algorithm efficiently performs modular addition by speculating and adjusting the twit value in the addition result. When the result exceeds the modulus, it subtracts$\mathrm{z}^{n}\pm\delta$by ignoring the carry-out and adjusting the speculated twit value. This adjustment is achieved through an XOR operation between the carry-out and the speculated twit value, simplifying the modular reduction process. The proposed design has been synthesized for practical n$(4\leq n\leq 16$using a FreePDK 45 nm process. The results demonstrate superior performance across key metrics such as delay, area, and power consumption compared to previous designs, highlighting the efficacy and scalability of the approach. Saeid Gorgin 0001, Amirhossein Sadr, Dara Rahmati, Jungrae Kim |
ARITH | 3 |
| 2022 | HyperDbg: Reinventing Hardware-Assisted DebuggingabstractSoftware analysis, debugging, and reverse engineering have a crucial impact in today's software industry. Efficient and stealthy debuggers are especially relevant for malware analysis. However, existing debugging platforms fail to address a transparent, effective, and high-performance low-level debugger due to their detectable fingerprints, complexity, and implementation restrictions. Mohammad Sina Karvandi, MohammadHosein Gholamrezaei, Saleh Khalaj Monfared, Soroush Meghdadi Zanjani, Behrooz Abbassi, Reza Mortazavi, Saeid Gorgin 0001, Dara Rahmati, Michael Schwarz 0001 |
CCS | 9 |
| 2022 | Hardware Efficient FIR Filter Architectures Using Accurate Unary Stochastic ComputingabstractFinite Impulse Response (FIR) filters are commonly used due to lower sensitivity to noise than their recursive counterparts. Computations of FIR filters require numerous multiply-and-accumulate (MAC) operations. Therefore, hard-ware implementation of high-order adaptive FIR filters results in a considerable area and power consumption. This paper proposes a hardware-efficient FIR engine based on the integration of deterministic approaches to Stochastic Computing (SC) with Residue Number Systems (RNS). The design inherits the intrinsic simplicity and low hardware requirements of SC circuits. As a contribution of our work, in contrast to other SC-based methods that impose errors on computations, our proposed method offers exact results like the binary implementations of FIR filters. Furthermore, the design decreases the required clock cycles, which can be translated to higher throughput in comparison with its SC predecessor (for example, 4× for 8-bit computations) at the cost of acceptable hardware overhead. Kamyar Givaki, Ahmad Khonsari, M. Hossein Gholamrezaei, Dara Rahmati, Saeid Gorgin 0001 |
ICCD | 4 |
| 2022 | A multi-application approach for synthesizing custom network-on-chips
Somayeh Kashi, Ahmad Patooghy, Dara Rahmati, Mahdi Fazeli |
J. Supercomput. | 3 |
| 2021 | A TSX-Based KASLR Break: Bypassing UMIP and Descriptor-Table Exiting
Mohammad Sina Karvandi, Saleh Khalaj Monfared, Sina Kiarostami, Dara Rahmati, Saeid Gorgin 0001 |
CRiSIS | 4 |
| 2021 | An energy efficient synthesis flow for application specific SoC design
Somayeh Kashi, Ahmad Patooghy, Dara Rahmati, Mahdi Fazeli |
Integr. | 3 |
| 2020 | TaxoNN: A Light-Weight Accelerator for Deep Neural Network TrainingabstractEmerging intelligent embedded devices rely on Deep Neural Networks (DNNs) to be able to interact with the real-world environment. This interaction comes with the ability to retrain DNNs, since environmental conditions change continuously in time. Stochastic Gradient Descent (SGD) is a widely used algorithm to train DNNs by optimizing the parameters over the training data iteratively. In this work, first we present a novel approach to add the training ability to a baseline DNN accelerator (inference only) by splitting the SGD algorithm into simple computational elements. Then, based on this heuristic approach we propose TaxoNN, a light-weight accelerator for DNN training. TaxoNN can easily tune the DNN weights by reusing the hardware resources used in the inference process using a time-multiplexing approach and low-bitwidth units. Our experimental results show that TaxoNN delivers, on average, 0.97% higher misclassification rate compared to a full-precision implementation. Moreover, TaxoNN provides 2.1× power saving and 1.65× area reduction over the state-of-the-art DNN training accelerator. Reza Hojabr, Kamyar Givaki, Kossar Pourahmadi, Parsa Nooralinejad, Ahmad Khonsari, Dara Rahmati, M. Hassan Najafi |
ISCAS | 6 |
| 2020 | On the Resilience of Deep Learning for Reduced-voltage FPGAsabstractDeep Neural Networks (DNNs) are inherently computation-intensive and also power-hungry. Hardware accelerators such as Field Programmable Gate Arrays (FPGAs) are a promising solution that can satisfy these requirements for both embedded and High-Performance Computing (HPC) systems. In FPGAs, as well as CPUs and GPUs, aggressive voltage scaling below the nominal level is an effective technique for power dissipation minimization. Unfortunately, bit-flip faults start to appear as the voltage is scaled down closer to the transistor threshold due to timing issues, thus creating a resilience issue.This paper experimentally evaluates the resilience of the training phase of DNNs in the presence of voltage underscaling related faults of FPGAs, especially in on-chip memories. Toward this goal, we have experimentally evaluated the resilience of LeNet-5 and also a specially designed network for CIFAR-10 dataset with different activation functions of Rectified Linear Unit (Relu) and Hyperbolic Tangent (Tanh). We have found that modern FPGAs are robust enough in extremely low-voltage levels and that low-voltage related faults can be automatically masked within the training iterations, so there is no need for costly software-or hardware-oriented fault mitigation techniques like ECC. Approximately 10% more training iterations are needed to fill the gap in the accuracy. This observation is the result of the relatively low rate of undervolting faults, i.e., <0.1%, measured on real FPGA fabrics. We have also increased the fault rate significantly for the LeNet-5 network by randomly generated fault injection campaigns and observed that the training accuracy starts to degrade. When the fault rate increases, the network with Tanh activation function outperforms the one with Relu in terms of accuracy, e.g., when the fault rate is 30% the accuracy difference is 4.92%. Kamyar Givaki, Behzad Salami 0001, Reza Hojabr, S. M. Reza Tayaranian, Ahmad Khonsari, Dara Rahmati, Saeid Gorgin 0001, Adrián Cristal, Osman S. Unsal |
PDP | 6 |
| 2019 | Using Residue Number Systems to Accelerate Deterministic Bit-stream MultiplicationabstractInaccuracy of computations is an important challenge with Stochastic Computing (SC). Deterministic approaches are proposed to produce completely accurate results with SC circuits. Current deterministic methods need a large number of clock cycles to produce exact result. This directly translates to a very high energy consumption. We propose a method based on the Residue Number Systems (RNS) to mitigate the high processing time of the deterministic methods. Compared to the state-of-the-art deterministic methods of SC, our approach delivers 760x and 170x improvement in terms of processing time and energy consumption. Kamyar Givaki, Reza Hojabr, M. Hassan Najafi, Ahmad Khonsari, M. Hossein Gholamrezayi, Saeid Gorgin 0001, Dara Rahmati |
ASAP | 7 |
| 2019 | Multi-Agent non-Overlapping Pathfinding with Monte-Carlo Tree SearchabstractIn this work, we propose a novel implementation of Monte-Carlo Tree Search (MCTS) algorithm to solve a multiagent pathfinding (MAPF) problem. We employ an optimization of MCTS with low time-complexity and acceptable reliability to approach the MAPF problems with no time constraint. To examine the efficiency and performance of the proposed approach, the NumberLink problem as a MAPF is investigated. We show that the addressed problem could be characterized as multi-agent pathfinding problem with no overlapping paths for the agents. Furthermore, we define this problem to be a simplified and special case of Multi-commodity flow problem (MCFP). Our MCTS solution utilizes a modified search-tree structure to efficiently solve the problem based on a 2-dimensional search space which performs in quadratic time complexity (O(m4) where input size is m2) and linear memory complexity (O(m2)). To evaluate our algorithm, we investigate the efficiency of the proposed solution for the well-known Flow Free puzzle. Our implementation solves a large 40 × 40 Numberlink puzzle in 21 minutes. To the best of our knowledge, there is no other efficient solution for this puzzle where the size of the problem is considerably large. Sina Kiarostami, Mohammad Reza Daneshvaramoli, Saleh Khalaj Monfared, Dara Rahmati, Saeid Gorgin 0001 |
CoG | 4 |
| 2019 | SkippyNN: An Embedded Stochastic-Computing Accelerator for Convolutional Neural NetworksabstractEmploying convolutional neural networks (CNNs) in embedded devices seeks novel low-cost and energy efficient CNN accelerators. Stochastic computing (SC) is a promising low-cost alternative to conventional binary implementations of CNNs. Despite the low-cost advantage, SC-based arithmetic units suffer from prohibitive execution time due to processing long bit-streams. In particular, multiplication as the main operation in convolution computation, is an extremely time-consuming operation which hampers employing SC methods in designing embedded CNNs. Reza Hojabr, Kamyar Givaki, S. M. Reza Tayaranian, Parsa Esfahanian, Ahmad Khonsari, Dara Rahmati, M. Hassan Najafi |
DAC | 6 |
| 2018 | Accurate Performance Bounds Calculation for Dynamic Voltage-Freq Islands in Best Effort NoCsabstractDynamic voltage and frequency scaling (DVFS) is a technique used to meet the power budget limitations in multi-core embedded systems. DVFS is applied to Networks-on-Chip (NoC) as a major contributor of the dissipated power on-chip. We propose an analytic model to accurately calculate the performance bounds on the best effort NoC with multiple voltage-frequency islands. The model supports dual-port buffer or handshaking mechanisms. We also suggest a method to set the frequencies of the islands to decrease the dissipated power. We examine our method and models and show their effectiveness. Dara Rahmati, Sobhan Masoudi, Ahmad Khonsari, Reza Sabbaghi-Nadooshan |
ICCD | 1 |
| 2018 | Classified Round Robin: A Simple Prioritized Arbitration to Equip Best Effort NoCs With Effective Hard QoSabstractAdvances in semiconductor technology enable integrating tens of cores on a single chip. Providing quality-of-service (QoS) for communication flows in complex embedded applications is critical. In this paper, we present a new approach for designing guaranteed service (GS) networks-on-chip by introducing a new arbitration algorithm and differentiating high and low priority traffic flows in best-effort (BE) networks. An analytical model is provided to compute accurate performance bound parameters in the network with the new arbitration. When the flows have the same priorities in a switch, the new algorithm acts exactly the same as the basic round robin arbitration. It works as a superset of the basic algorithm, when the flows have different priorities. The proposed method helps designers to easily equip traditional BE networks with effective hard QoS, changing it to a GS network. This is done without the need to get involved in the designing complexity of traditional GS networks and still benefit from the superior properties of BE networks. We show substantial improvement in performance bounds for high priority flows (more than 40% in delay and 80% in bandwidth, on average) compared to the known approaches. Dara Rahmati, Hamid Sarbazi-Azad |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Computing Accurate Performance Bounds for Best Effort Networks-on-ChipabstractReal-time (RT) communication support is a critical requirement for many complex embedded applications which are currently targeted to Network-on-chip (NoC) platforms. In this paper, we present novel methods to efficiently calculate worst case bandwidth and latency bounds for RT traffic streams on wormhole-switched NoCs with arbitrary topology. The proposed methods apply to best-effort NoC architectures, with no extra hardware dedicated to RT traffic support. By applying our methods to several realistic NoC designs, we show substantial improvements (more than 30 percent in bandwidth and 50 percent in latency, on average) in bound tightness with respect to existing approaches. Dara Rahmati, Srinivasan Murali, Luca Benini, Federico Angiolini, Giovanni De Micheli, Hamid Sarbazi-Azad |
IEEE Trans. Computers | 1 |
| 2013 | Designing best effort networks-on-chip to meet hard latency constraintsabstractMany classes of applications require Quality of Service (QoS) guarantees from the system interconnect. In Networks-on-Chip (NoC) QoS guarantees usually translate into bandwidth and latency constraints for the traffic flows and require hardware support in the NoC fabric and its interfaces. In this article we present a novel NoC synthesis framework to automatically build networks that meet hard latency constraints of end-to-end traffic streams without requiring specialized hardware for the network components. The hard latency constraints are met by carefully designing the NoC topology and selecting the appropriate routes for flow using lean best-effort network components. We perform experiments on several System on Chip (SoC) benchmarks. We compared against a topology synthesis method with no support for real-time constraints and we show that the proposed method can produce topologies that can meet significantly tighter worst case latency constraints (on average 44%). We also show that the tightest worst case latency can be provided with little overhead on power consumption (on average 8.5%). Ciprian Seiculescu, Dara Rahmati, Srinivasan Murali, Hamid Sarbazi-Azad, Luca Benini, Giovanni De Micheli |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2009 | A method for calculating hard QoS guarantees for Networks-on-ChipabstractMany Networks-on-Chip (NoC) applications exhibit one or \nmore critical traffic flows that require hard Quality of Service \n(QoS). Guaranteeing bandwidth and latency for such real time \nflows is crucial. In this paper, we present novel methods to \nefficiently calculate worst-case bandwidth and latency bounds \nand thereby provide hard QoS guarantees. Importantly, the \nproposed methods apply even to best-effort NoC architectures, \nwith no extra hardware dedicated to QoS support. By applying \nour methods to several realistic NoC designs, we show \nsubstantial improvements (on average, more than 30% in \nbandwidth and 50% in latency) in bound tightness with respect \nto existing approaches.1 Dara Rahmati, Srinivasan Murali, Luca Benini, Federico Angiolini, Giovanni De Micheli, Hamid Sarbazi-Azad |
ICCAD | 1 |
| 2008 | A Markovian Performance Model for Networks-on-ChipabstractNetwork-on-chip (NoC) has been proposed as a solution for addressing the design challenges of future high-performance nanoscale architectures. Thus, it is of crucial importance for a designer to have access to last methods for evaluating the performance of on-chip networks. To this end, we present a Markovian model for evaluating the latency and energy consumption of on-chip networks. We compute the average delay due to path contention, virtual channel and crossbar switch arbitration using a queuing-based approach, which can capture the blocking phenomena of wormhole switching quite accurately. The model is then used to estimate the power consumption of all routers in NoCs. The performance results from the analytical models are validated with those obtained from a synthesizable VHDL-based cycle accurate simulator. Comparison with simulation results indicate that the proposed analytical model is quite accurate and can be used as an efficient design tool by SoC designers. Abbas Eslami Kiasari, Dara Rahmati, Hamid Sarbazi-Azad, Shaahin Hessabi |
PDP | 2 |
| 2007 | Effect of number of faults on NoC power and performanceabstractAccording to international technology roadmap for semiconductors (ITRS), before the end of this decade, we will be entering the era of a billion transistors on a single chip. The major threat toward the achievement of billion transistor on a chip is poor scalability of current interconnect infrastructure. With the advent of "network on chip (NoC)" various characters and methodologies of traditional networks were hardly considered on-chip. Failure, power and area are the major concepts that should be considered when migrating from traditional interconnection networks to NoCs. In this paper we study the effects of faulty links and nodes on power and performance of mesh based NoC, Also several routing algorithms have been implemented and simulated using a cycle accurate VHDL model of NoC. Mahdiar Hosein Ghadiry, Mahdieh Nadi Senejani, Mohammad T. Manzuri Shalmani, Dara Rahmati |
ICPADS | 4 |
| 2006 | A performance and power analysis of WK-Recursive and Mesh Networks for Network-on-ChipsabstractNetwork-on-chip (NoC) has been proposed as an attractive alternative to traditional dedicated wires to achieve high performance and modularity. Power efficiency is one of the most important concerns in NoC architecture design. The choice of network topology is important in designing a low-power and high-performance NoC. In this paper, we propose the use of the WK-recursive networks to be used as the underlying topology in NoC. We have implemented VHDL hardware model of mesh and WK-recursive topologies and measured the latency results using simulation with these implementation. We also propose a novel approach in high level power modeling based on latency for these topologies and show that the power consumption of WK-recursive topology is less than that of the equivalent mesh on a chip. Dara Rahmati, Abbas Eslami Kiasari, Shaahin Hessabi, Hamid Sarbazi-Azad |
ICCD | 1 |