Bahareh Khabbazan

dblp:251/4291 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0001-6726-2804ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Toward Efficient LUT-Based PIM: A Scalable and Low-Power Approach for Modern Workloads
abstract
Data movement in memory-intensive workloads is several orders of magnitude more energy-consuming than computation. Frequent memory-processor transfers make this a critical bottleneck. Processing-using-memory (PuM) alleviates this by enabling in-memory computation and reducing data movement. In this paper, we proposeLama, a LUT-based PuM architecture that efficiently executes SIMD operations by supporting independent column accesses within each mat of a DRAM subarray. Lama minimizes the number of energy-intensive memory activation commands, which are the primary source of overhead in most PuM architectures. Besides, Lama supports up to 8-bit operand precision without decomposing computations, unlike prior PuM solutions, while incurring only a 2.47% area overhead. Our evaluation shows that Lama achieves an average performance improvement of 8.5× over state-of-the-art PuM architectures and a 3.8× improvement over CPU, along with energy efficiency gains of 6.9$\times /8\times$, respectively, for bulk 8-bit multiplication. We introduceLamaAccel, an HBM-based PuM accelerator that leverages Lama to accelerate LLM inference. By transforming MAC computations into energy-efficient addition and counting operations, LamaAccel achieves up to 9.3$\times /19.2\times$energy reduction and 4.8$\times /9.8\times$speedup over a TPU/GPU, as well as up to 5.8× lower energy consumption and 2.1× higher performance compared to a state-of-the-art PuM baseline.
Bahareh Khabbazan, Marc Riera, Antonio González 0001
IEEE Trans. Parallel Distributed Syst.1
2025 An energy-efficient near-data processing accelerator for DNNs to optimize memory accesses
abstract
The constant growth of DNNs makes them challenging to implement and run efficiently on traditional computecentric architectures. Some accelerators have attempted to add more compute units and on-chip buffers to solve the memory wall problem without much success, and sometimes even worsening the issue since more compute units also require higher memory bandwidth. Prior works have proposed the design of memorycentric architectures based on the Near-Data Processing (NDP) paradigm. NDP seeks to break the memory wall by moving the computations closer to the memory hierarchy, reducing the data movements and their cost as much as possible. The 3D-stacked memory is especially appealing for DNN accelerators due to its high-density/low-energy storage and near-memory computation capabilities to perform the DNN operations massively in parallel. However, memory accesses remain as the main bottleneck for running modern DNNs efficiently. To improve the efficiency of DNN inference we present QeiHaN, a hardware accelerator that implements a 3D-stacked memory-centric weight storage scheme to take advantage of a logarithmic quantization of activations. In particular, since activations of FC and CONV layers of modern DNNs are commonly represented as powers of two with negative exponents, QeiHaN performs an implicit in-memory bit-shifting of the DNN weights to reduce memory activity. Only the meaningful bits of the weights required for the bit-shift operation are accessed. Overall, QeiHaN reduces memory accesses by 25% compared to a standard memory organization. We evaluate QeiHaN on a popular set of DNNs. On average, QeiHaN provides4.3¿speedup and3.5¿energy savings over a Neurocube-like accelerator.
Bahareh Khabbazan, Mohammad Sabri, Marc Riera, Antonio González 0001
J. Syst. Archit.1
2023 QeiHaN: An Energy-Efficient DNN Accelerator that Leverages Log Quantization in NDP Architectures
abstract
The constant growth of DNNs makes them challenging to implement and run efficiently on traditional computecentric architectures. Some works have attempted to enhance accelerators by adding more compute units and on-chip buffers, but they often worsen the memory issue due to increased bandwidth demands. Memory-centric designs based on Near-Data Processing (NDP) have been proposed to mitigate this problem by moving computations closer to the memory hierarchy. Leveraging 3D-stacked memory for its storage density and near-memory processing capabilities, this paper introduces QeiHaN, a hardware accelerator that optimizes DNN inference efficiency. QeiHaN employs a 3D-stacked memory-centric weight storage scheme combined with a logarithmic quantization of activations, resulting in reduced memory accesses by 25%. Evaluation demonstrates significant speedup and energy savings compared to a Neurocube-like accelerator across various DNNs.
Bahareh Khabbazan, Marc Riera, Antonio González 0001
PACT1
2023 DNA-TEQ: An Adaptive Exponential Quantization of Tensors for DNN Inference
abstract
Quantization is commonly used in Deep Neural Networks (DNNs) to reduce the storage and computational complexity by decreasing the arithmetical precision of activations and weights, a.k.a. tensors. Efficient hardware architectures employ linear quantization to enable the deployment of recent DNNs onto embedded systems and mobile devices. However, linear uniform quantization cannot usually reduce the numerical precision to less than 8 bits without sacrificing high performance in terms of model accuracy. The performance loss is due to the fact that tensors do not follow uniform distributions. In this paper, we show that a significant amount of tensors fit into an exponential distribution. Then, we propose DNA-TEQ to exponentially quantize DNN tensors with an adaptive scheme that achieves the best trade-off between numerical precision and accuracy loss. The experimental results show that DNA-TEQ provides a much lower quantization bit-width compared to previous proposals, resulting in an average compression ratio of 40 % over the linear INT8 baseline, with negligible accuracy loss and without retraining the DNNs. Besides, DNA-TEQ leads the way in performing dot-product operations in the exponential domain. On average for a set of widely used DNNs, DNA-TEQ provides 1.5x speedup and 2.5x energy savings over a baseline DNN accelerator based on 3D-stacked memory.
Bahareh Khabbazan, Marc Riera, Antonio González 0001
HiPC1
2021 Area and Power-Efficient Variable-Sized DCT Architecture for HEVC Using Muxed-MCM Problem
abstract
This paper presents an area and power-efficient variable-size DCT architecture for HEVC application. We develop a reconfigurable and scalable shift-and-add unit (SAU) embedded in our 1D-DCT architecture by leveraging Muxed-MCM problem with the aim of increasing the hardware reusability in the arithmetic units, while reducing the hardware cost. The key idea behind the proposed architecture is the fact that in most of the times (≈90%) the lower point DCTs are performed when the higher point SAUs remain unused. Accordingly, we focus on merging the SAUs of lower point DCTs into the higher point DCTs to compute multiple lower point DCTs in parallel as well as processing any combination of transform sizes. The experimental results show that the proposed folded and fully-parallel 2D-DCT architectures achieve the best hardware cost by 45% and 30% reduction in gate count, respectively, amongst the existing architectures. Moreover, power saving of 55% and 32% can be achieved for the proposed folded and fully-parallel architectures, respectively, where they can process 60 fps of 4K and 30 fps of 8K UHD video sequences in 300 MHz operating frequency.
Ahmad Shabani, Mohammad Sabri Abrebekoh, Bahareh Khabbazan, Somayeh Timarchi
IEEE Trans. Circuits Syst. I Regul. Pap.3
2019 Design and Implementation of a Low-Power, Embedded CNN Accelerator on a Low-end FPGA
abstract
In this paper, an optimized hardware for Convolutional Neural Networks with the purpose of implementation on embedded vision systems is presented. This design method is meant to be implemented with minimum resource consumption on a low-end hardware platform. We propose an architecture on a Z-turn evaluation board featuring a Xilinx Zynq-7000 system on chip (SoC). All computations in this architecture are optimized as 8-bit. Also, the accelerator has a frequency of 160 MHz and power consumption of 1.77 watts which leads to a performance of 40.96GOP/s, using only 134 computing units and 601 KB of internal memory. So we can claim that the acceptable speed and low power and low area consumption of this architecture make it an ideal choice for portable and embedded CNN applications.
Bahareh Khabbazan, Sattar Mirzakuchaki
DSD1