EDBT 2026 Demo / reviewers in the wild / expert
Reza Hojabr
dblp:172/4942
· DBLP profile ↗
13ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MLB-MAC: Multi-Level Binary MAC Array for Energy Efficient ML AcceleratorsabstractQuantization is a critical compression technique for optimizing deep neural networks (DNNs) on resource-constrained embedded devices. Efficient hardware utilization hinges on effective number representation. Integer representation, a widely adopted method, uses scaling factors and offsets to enhance network accuracy and simplify fractional bit selection through uniform quantization. On the other hand, non-uniform quantization is well-suited for DNNs with parameters following a normal distribution, helping reduce data width requirements. This paper introduces a novel non-uniform representation called MLB (Multi-Level Binary), which encompasses and extends integer representation. We propose an architecture that objectively compares these representations, demonstrating that MLB is a superset of integer representation in terms of accuracy. Our comprehensive analysis spans various data widths, parallel factorization, and DNN models. Our findings indicate that MLB outperforms integer representation in energy efficiency for lower bit widths (2–5 bits), whereas integer representation is more advantageous for higher bit widths (4–8 bits). Specifically, our work shows an average energy improvement of 1.1 to 2.3× and area saving up to 1.7× compared to integer multiply-accumulate (MAC) units, while preserving network accuracy. This research provides insights into the optimal choice of number representation based on bit-width requirements, highlighting the potential of MLB in enhancing the performance and efficiency of DNNs on embedded devices. Ali Ansarmohammadi, Reza Hojabr, Marzie Mastalizade, Najmeh Nazari, Mostafa E. Salehi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | mu-grind: A Framework for Dynamically Instrumenting HLS-Generated RTLabstractHigh-level synthesis compilers (HLS) enable the rapid creation of accelerator circuits. Unfortunately, compiler generated RTL (H-RTL) is inconsistent in terms of quality, hard to comprehend, and tends to be brittle [28, 41]. This paper develops a framework to help HLS compiler architects inspect and profile H-RTL. Prior state-of-the-art tools [23, 57] have predominantly focused on tracing. Tracing requires massive amount of on-chip buffering, limits the H-RTL design size, and only support post-mortem analysis at the end of the execution. Parmida Vahdatniya, Amirali Sharifian, Reza Hojabr, Arrvindh Shriraman |
PACT | 3 |
| 2022 | X-cache: a modular architecture for domain-specific cachesabstractWith Dennard scaling ending, architects are turning to domain-specific accelerators (DSAs). State-of-the-art DSAs work with sparse data [37] and indirectly-indexed data structures [18, 30]. They introduce non-affine and dynamic memory accesses [7, 35], and require domain-specific caches. Unfortunately, cache controllers are notorious for being difficult to architect; domain-specialization compounds the problem. DSA caches need to support custom tags, data-structure walks, multiple refills, and preloading. Prior DSAs include ad-hoc cache structures, and do not implement the cache controller. We propose X-Cache, a reusable caching idiom for DSAs. We will be open-sourcing a toolchain for both generating the RTL and programming X-Cache. There are three key ideas: i) DSA-specific Tags (Meta-tag): The designer can use any combination of fields from the DSA-metadata as the tag. Meta-tags eliminate the overhead of walking and translating metadata to global addresses. This saves energy, and improves load-to-use latency. ii) DSA-programmable walkers (X-Actions): We find that a common set of microcode actions can be used to implement the DSA-specific walking, data block, and tag management. We develop a programmable microcode engine that can efficiently realize the data orchestration. iii) DSA-portable controller (X-Routines): We use a portable abstraction, coroutines, to let the designer express walking and orchestration. Coroutines capture the block-level parallelism, remain lightweight, and minimize controller occupancy. We create caches for four different DSA families: Sparse GEMM [35, 37], GraphPulse [30], DASX [22], and Widx [18]. X-Cache outperforms address-based caches by 1.7 × and remains competitive with hardwired DSAs (even 50% improvement in one case). We demonstrate that meta-tags save 26--79% energy compared to address-tags. In X-Cache, meta-tags consume 1.5--6.5% of data RAM energy and the programmable microcode adds a further 7%. Ali Sedaghati, Milad Hakimi, Reza Hojabr, Arrvindh Shriraman |
ISCA | 3 |
| 2021 | X-Layer: Building Composable Pipelined Dataflows for Low-Rank ConvolutionsabstractPrior research in hardware accelerators has largely focused on spatial convolutions (CONV). However, state-of-the-art DNNs employ low-rank convolutions (LR-CONV). LR-CONVs such as depthwise and pointwise convolutions exhibit lower arithmetic intensity and lower data re-use. LR-CONV s result in low hardware utilization and high latency. However, they provide opportunities for inter-layer data reuse. We propose X-Layer, which systematically explores the design space of cross-layer dataflows. We develop novel fine-grain cross-layer dataflows for LR-CONVs that support partial loop dimension completion. X-Layer decouples the nested loops in a pipeline and combines them to create a common outer dataflow and several inner dataflows. X-layer discovers additional opportunities for optimizing LR-CONVs: i) it overlaps adjacent layers at fine-granularity with partially completed channels and filters. This minimizes the intermediate storage required. ii) it enables each pipelined layer to independently choose optimal outer and inner dataflows by supporting streaming activation transformations. We explore a large design space of cross-layer dataflows and evaluate them for depth-separable, inverted residual, and CONV layers across six different DNNs. We also find that coarse-grain dataflows are sensitive to on-chip memory (≥ 1.5 MB) and performance drops steeply if enough on-chip SRAM is not provided. X-Layer dataflows find optimal performance across a wide range of on-chip memory (≥ 32KB). Compared to the existing coarse-grain and medium-grain dataflows, X-Layer improves the performance by 7.8× and 16.6×, while requiring 8.3× and 2× less SRAM. Naveen Vedula, Reza Hojabr, Ahmad Khonsari, Arrvindh Shriraman |
PACT | 2 |
| 2021 | SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsabstractGeneralized Sparse Matrix-Matrix Multiplication (Sparse GEMM) is widely used across multiple domains, but the computation’s regularity is dependent on the input sparsity pattern. The majority of sparse GEMM accelerators are based on the inner product method and propose new storage formats [5], [28], [31] to regularize computation. We find that these storage formats are more suited for denser matrices. Accelerators [26], [34] adopting the outer product algorithm are more suitable for highly sparse inputs $(\lt 1$% density), since they support CSC/CSR storage formats. The current state-of-the-art, SpArch [34], condenses inputs to improve output reuse, but then spoils input reuse. The condensing effectiveness varies across inputs leading to high variance in DRAM utilization and speedup across inputs. SpArch also requires a complex memory hierarchy (e.g., prefetch caches) to re-capture input reuse.We propose Spaghetti, an open-source Chisel generator for creating FPGA-optimized outer product accelerators. The key novelty in Spaghetti is a new pattern-aware software scheduler that analyzes the sparsity pattern and schedules row-col pairs of the inputs onto the fixed microarchitecture. Spaghetti takes advantage of our observation that the rows in the input matrix lead to mutually independent rows in the final output. Thus the scheduler can partition the input into tiles that maximize reuse and eliminate re-fetching the partial matrices from the DRAM. The microarchitecture template we create has the following key benefits: i) we can statically schedule the inputs in a streaming fashion and maximize DRAM utilization, ii) we can parallelize the merge phase and generate multiple rows of the output in parallel maximally using the output DRAM bandwidth, iii) we can adapt to the varying logic resources and bandwidth across various FPGA devices and attain maximal roofline performance (only limited by memory bandwidth). We auto-generate sparse GEMM accelerators on Amazon AWS FPGAs and demonstrate that we can achieve performance improvement over CPUs and GPUs between 1.1 – 34.5 x. Compared to SpArch [34], our design improves performance by an average of $2.6 \times$, and reduces DRAM accesses by an average of $4 \times$. Reza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari, Arrvindh Shriraman |
HPCA | 1 |
| 2021 | Real-Time Hamilton-Jacobi Reachability Analysis of Autonomous System With An FPGAabstractHamilton-Jacobi (HJ) reachability analysis is a powerful technique used to verify the safety of autonomous systems. HJ reachability is ideal for analysing nonlinear systems with disturbances and flexible set representations. A drawback to this approach is that it suffers from the curse of dimensionality, which prevents real-time deployment on safety-critical systems. In this paper, we show that a customized hardware design on an Field Programmable Gate Array (FPGA) could accelerate 4D grid-based HJ reachability analysis up to 14 times compared to an optimized implementation and 103 times compared to state-of-the-art MATLAB toolboxes on a 16-thread CPU. Because of this, we are able to achieve guaranteed real- time collision avoidance in dynamic environments that abruptly change with a 4D car model by re-solving the HJ partial differential equation (PDE) at a frequency of 4Hz on an FPGA. Our design can overcome the complex data access pattern while taking advantage of the parallel nature of the computations for solving the HJ PDE. The low latency of our computation is consistent, which is crucial for safety-critical systems. The methodology presented here is without loss of generality: it can potentially be applied to different systems dynamics, and more- over, leveraged for higher dimensional systems. We validate our approach in real world collision avoidance experiments with a robot car in a changing environment. We also provide the code of our hardware design and an AWS AFI image. Minh Bui, Michael Lu, Reza Hojabr, Mo Chen 0001, Arrvindh Shriraman |
IROS | 3 |
| 2020 | TaxoNN: A Light-Weight Accelerator for Deep Neural Network TrainingabstractEmerging intelligent embedded devices rely on Deep Neural Networks (DNNs) to be able to interact with the real-world environment. This interaction comes with the ability to retrain DNNs, since environmental conditions change continuously in time. Stochastic Gradient Descent (SGD) is a widely used algorithm to train DNNs by optimizing the parameters over the training data iteratively. In this work, first we present a novel approach to add the training ability to a baseline DNN accelerator (inference only) by splitting the SGD algorithm into simple computational elements. Then, based on this heuristic approach we propose TaxoNN, a light-weight accelerator for DNN training. TaxoNN can easily tune the DNN weights by reusing the hardware resources used in the inference process using a time-multiplexing approach and low-bitwidth units. Our experimental results show that TaxoNN delivers, on average, 0.97% higher misclassification rate compared to a full-precision implementation. Moreover, TaxoNN provides 2.1× power saving and 1.65× area reduction over the state-of-the-art DNN training accelerator. Reza Hojabr, Kamyar Givaki, Kossar Pourahmadi, Parsa Nooralinejad, Ahmad Khonsari, Dara Rahmati, M. Hassan Najafi |
ISCAS | 1 |
| 2020 | On the Resilience of Deep Learning for Reduced-voltage FPGAsabstractDeep Neural Networks (DNNs) are inherently computation-intensive and also power-hungry. Hardware accelerators such as Field Programmable Gate Arrays (FPGAs) are a promising solution that can satisfy these requirements for both embedded and High-Performance Computing (HPC) systems. In FPGAs, as well as CPUs and GPUs, aggressive voltage scaling below the nominal level is an effective technique for power dissipation minimization. Unfortunately, bit-flip faults start to appear as the voltage is scaled down closer to the transistor threshold due to timing issues, thus creating a resilience issue.This paper experimentally evaluates the resilience of the training phase of DNNs in the presence of voltage underscaling related faults of FPGAs, especially in on-chip memories. Toward this goal, we have experimentally evaluated the resilience of LeNet-5 and also a specially designed network for CIFAR-10 dataset with different activation functions of Rectified Linear Unit (Relu) and Hyperbolic Tangent (Tanh). We have found that modern FPGAs are robust enough in extremely low-voltage levels and that low-voltage related faults can be automatically masked within the training iterations, so there is no need for costly software-or hardware-oriented fault mitigation techniques like ECC. Approximately 10% more training iterations are needed to fill the gap in the accuracy. This observation is the result of the relatively low rate of undervolting faults, i.e., <0.1%, measured on real FPGA fabrics. We have also increased the fault rate significantly for the LeNet-5 network by randomly generated fault injection campaigns and observed that the training accuracy starts to degrade. When the fault rate increases, the network with Tanh activation function outperforms the one with Relu in terms of accuracy, e.g., when the fault rate is 30% the accuracy difference is 4.92%. Kamyar Givaki, Behzad Salami 0001, Reza Hojabr, S. M. Reza Tayaranian, Ahmad Khonsari, Dara Rahmati, Saeid Gorgin 0001, Adrián Cristal, Osman S. Unsal |
PDP | 3 |
| 2019 | Using Residue Number Systems to Accelerate Deterministic Bit-stream MultiplicationabstractInaccuracy of computations is an important challenge with Stochastic Computing (SC). Deterministic approaches are proposed to produce completely accurate results with SC circuits. Current deterministic methods need a large number of clock cycles to produce exact result. This directly translates to a very high energy consumption. We propose a method based on the Residue Number Systems (RNS) to mitigate the high processing time of the deterministic methods. Compared to the state-of-the-art deterministic methods of SC, our approach delivers 760x and 170x improvement in terms of processing time and energy consumption. Kamyar Givaki, Reza Hojabr, M. Hassan Najafi, Ahmad Khonsari, M. Hossein Gholamrezayi, Saeid Gorgin 0001, Dara Rahmati |
ASAP | 2 |
| 2019 | SkippyNN: An Embedded Stochastic-Computing Accelerator for Convolutional Neural NetworksabstractEmploying convolutional neural networks (CNNs) in embedded devices seeks novel low-cost and energy efficient CNN accelerators. Stochastic computing (SC) is a promising low-cost alternative to conventional binary implementations of CNNs. Despite the low-cost advantage, SC-based arithmetic units suffer from prohibitive execution time due to processing long bit-streams. In particular, multiplication as the main operation in convolution computation, is an extremely time-consuming operation which hampers employing SC methods in designing embedded CNNs. Reza Hojabr, Kamyar Givaki, S. M. Reza Tayaranian, Parsa Esfahanian, Ahmad Khonsari, Dara Rahmati, M. Hassan Najafi |
DAC | 1 |
| 2019 | μIR -An intermediate representation for transforming and optimizing the microarchitecture of application acceleratorsabstractCreating high quality application-specific accelerators requires us to make iterative changes to both algorithm behavior and microarchitecture, and this is a tedious and error-prone process. High-Level Synthesis (HLS) tools [5, 10] generate RTL for application accelerators from annotated software. Unfortunately, the generated RTL is challenging to change and optimize. The primary limitation of HLS is that the functionality and microarchitecture are conflated together in a single language (such as C++). Making changes to the accelerator design may require code restructuring, and microarchitecture optimizations are tied with program correctness. Amirali Sharifian, Reza Hojabr, Navid Rahimi, Sihao Liu, Apala Guha, Tony Nowatzki, Arrvindh Shriraman |
MICRO | 2 |
| 2017 | Customizing Clos Network-on-Chip for Neural NetworksabstractLarge-scale neural network accelerators are often implemented as a many-core chip and rely on a network-on-chip to manage the huge amount of inter-neuron traffic. The baseline and different variations of the well-known mesh and tree topologies are the most popular topologies in prior many-core implementations of neural networks. However, the grid-like mesh and hierarchical tree topologies suffer from high diameter and low bisection bandwidth, respectively. In this paper, we present ClosNN, a customized Clos topology for Neural Networks. The inherent capability of Clos to support multicast and broadcast traffic in a simple and efficient way, as well as its adaptable bisection bandwidth, is the major motivation behind proposing a customized version of this topology as the communication infrastructure of large-scale neural network implementations. We compare ClosNN with some state-of-the-art NoC topologies adopted in recent neural network hardware accelerators and show that it offers lower average message hop count and higher throughput, which directly translates to faster neural information processing. Reza Hojabr, Mehdi Modarressi, Masoud Daneshtalab, Ali Yasoubi, Ahmad Khonsari |
IEEE Trans. Computers | 1 |
| 2015 | CuPAN - High Throughput On-chip Interconnection for Neural Networks
Ali Yasoubi, Reza Hojabr, Hengameh Takshi, Mehdi Modarressi, Masoud Daneshtalab |
ICONIP (3) | 2 |