Foroozan Karimzadeh

dblp:169/3604 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0001-8849-7784ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CommVQ: Commutative Vector Quantization for KV Cache Compression
abstract
Large Language Models (LLMs) are increasingly used in applications requiring long context lengths, but the key-value (KV) cache often becomes a memory bottleneck on GPUs as context grows. To address this, we propose Commutative Vector Quantization (CommVQ) to significantly reduce memory usage for long-context LLM inference. We first introduce additive quantization with a lightweight encoder and codebook to compress the KV cache, which can be decoded via simple matrix multiplication. To further reduce computational costs during decoding, we design the codebook to be commutative with Rotary Position Embedding (RoPE) and train it using an Expectation-Maximization (EM) algorithm. This enables efficient integration of decoding into the self-attention mechanism. Our approach achieves high accuracy with additive quantization and low overhead via the RoPE-commutative codebook. Experiments on long-context benchmarks and GSM8K show that our method reduces FP16 KV cache size by 87.5% with 2-bit quantization, while outperforming state-of-the-art KV cache quantization methods. Notably, it enables 1-bit KV cache quantization with minimal accuracy loss, allowing a LLaMA-3.1 8B model to run with a 128K context length on a single RTX 4090 GPU. The source code is available at: https://github.com/UMass-Embodied-AGI/CommVQ.
Yang Zhang 0001, Muhammad Yusuf Hassan, Talha Chafekar, Tianle Cai, Zhile Ren, Pengsheng Guo, Foroozan Karimzadeh, Colorado Reed, Chuang Gan 0001
ICML8
2023 Memory-Based Computing for Energy-Efficient AI: Grand Challenges
abstract
The remarkable progress in artificial intelligence (AI) has ushered in a new era characterized by models with billions of parameters, enabling extraordinary capabilities across diverse domains. However, these achievements come at a significant cost in terms of memory and energy consumption. The growing demand for computational resources raises grand challenges for the sustainable development of energy-efficient AI systems. This paper delves into the paradigm of memory-based computing as a promising avenue to address these challenges. By capitalizing on the inherent characteristics of memory and its efficient utilization, memory-based computing offers a novel approach to enhance AI performance while reducing the associated energy costs. Our paper systematically analyzes the multifaceted aspects of this paradigm, highlighting its potential benefits and outlining the challenges it poses. Through an exploration of various methodologies, architectures, and algorithms, we elucidate the intricate interplay between memory utilization, computational efficiency, and AI model complexity. Furthermore, we review the evolving area of hardware and software solutions for memory-based computing, underscoring their implications for achieving energy-efficient AI systems. As AI continues its rapid evolution, identifying the key challenges and insights presented in this paper serve as a foundational guide for researchers striving to navigate the complex field of memory-based computing and its pivotal role in shaping the future of energy-efficient AI.
Foroozan Karimzadeh, Mohsen Imani, Bahar Asgari, Ningyuan Cao, Yingyan (Celine) Lin, Yan Fang 0002
VLSI-SoC1
2022 Towards Energy Efficient DNN accelerator via Sparsified Gradual Knowledge Distillation
abstract
Artificial intelligence (AI) is becoming increasingly popular in many applications. However, the computation cost of deep neural network (DNN) , which is a powerful form of AI, calls for efficient DNN compression technique to make energy efficient networks. In this paper, we proposed SKG, a method to jointly sparsify and quantize DNN models to ultra-low bit-precision using Knowledge Distillation and gradual quantization (SKG). We demonstrated that our method can preserve the accuracy more than 20% for uniform quantization with 2 bit-width compared to the baseline methods on ImageNet and ResNet-18. In addition, our method can achieve up to 2.7x lower energy consumption using compute-in-memory (CIM) architecture compared to a traditional 65nm CMOS architecture for both pruned and unpruned network during inference and eventually enabling using DNN models on resource constrained edge devices.
Foroozan Karimzadeh, Arijit Raychowdhury
VLSI-SoC1
2022 Towards CIM-friendly and Energy-Efficient DNN Accelerator via Bit-level Sparsity
abstract
The rising popularity of deep neural network (DNN) algorithms calls for energy-efficient accelerators to enable DNNs run on edge devices. In this paper, we presented BitS-Net, a bit-level sparsity method that quantize the network to desirable numbers with more zeros in their bit representation. We demonstrated that BitS-Net can preserve the accuracy (67.73 %) with accuracy drop ¡1% compared to the original network. Moreover, it achieved up to 5x energy efficiency for ResNet-18 models on the ImageNet dataset compared to the baseline methods.
Foroozan Karimzadeh, Arijit Raychowdhury
VLSI-SoC1
2022 BitS-Net: Bit-Sparse Deep Neural Network for Energy-Efficient RRAM-Based Compute-In-Memory
abstract
The rising popularity of intelligent mobile devices and the computational cost of deep learning-based models call for efficient and accurate on-device inference schemes. We propose a novel model compression scheme that allows inference to be carried out using bit-level sparsity, which can be efficiently implemented using in-memory computing macros. In this paper, we introduce a method called BitS-Net to leverage the benefits of bit-sparsity (where the number of zeros are more than number of ones in binary representation of weight/activation values) when applied to compute-in-memory (CIM) with resistive RAM (RRAM) to develop energy efficient DNN accelerators operating in the inference mode. We demonstrate that BitS-Net improves the energy efficiency by up to 5x for ResNet models on the ImageNet dataset.
Foroozan Karimzadeh, Jong-Hyeok Yoon, Arijit Raychowdhury
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 A Hardware-Friendly Approach Towards Sparse Neural Networks Based on LFSR-Generated Pseudo-Random Sequences
abstract
The increase in the number of edge devices has led to the emergence of edge computing where the computations are performed on the device. In recent years, deep neural networks (DNNs) have become the state-of-the-art method in a broad range of applications, from image recognition, to cognitive tasks to control. However, neural network models are typically large and computationally expensive and therefore not deployable on power and memory constrained edge devices. Sparsification techniques have been proposed to reduce the memory foot-print of neural network models. However, they typically lead to substantial hardware and memory overhead. In this article, we propose a hardware-aware pruning method using linear feedback shift register (LFSRs) to generate the locations of non-zero weights in real-time during inference. We call this LFSR-generated pseudorandom sequence based sparsity (LGPS) technique. We explore two different architectures for our hardware-friendly LGPS technique, based on (1) row/column indexing with LFSRs and (2) column-wise indexing with nested LFSRs, respectively. Using the proposed method, we present a total saving of energy and area up to 37.47% and 49.93% respectively and speed up of 1.53× w.r.t the baseline pruning method, for the VGG-16 network on down-sampled ImageNet.
Foroozan Karimzadeh, Ningyuan Cao, Brian Crafton, Justin K. Romberg, Arijit Raychowdhury
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 Hardware-Aware Pruning of DNNs using LFSR-Generated Pseudo-Random Indices
abstract
Deep neural networks (DNNs) have been emerged as the state-of-the-art algorithms in broad range of applications. To reduce the memory foot-print of DNNs, in particular for embedded applications, sparsification techniques have been proposed. Unfortunately, these techniques come with a large hardware overhead. In this paper, we present a hardware-aware pruning method where the locations of non-zero weights are derived in real-time from a Linear Feedback Shift Registers (LFSRs). Using the proposed method, we demonstrate a total saving of energy and area up to 63.96% and 64.23% for VGG-16 network on down-sampled ImageNet, respectively for iso-compression-rate and iso-accuracy.
Foroozan Karimzadeh, Ningyuan Cao, Brian Crafton, Justin K. Romberg, Arijit Raychowdhury
ISCAS1
2020 Memory and Energy Efficient Method Toward Sparse Neural Network Using LFSR Indexing
abstract
Deep Neural Networks (DNNs) require enormous computational power and storage memory. This impose a critical challenge to their efficient deployment on resource-constrained computing platforms such as edge devices. In this paper, we present a novel pruning algorithm and its hardware implementation to reduce the required memory-footprint and power usage of DNNs to enable them to be deployable on edge and mobile devices. we demonstrated a hardware-friendly pruning method where the locations of non-zero weights are derived from a Linear Feedback Shift Registers (LFSRs) in real-time. The results show a total power and area savings up to 49.97 % and 50.20 % for VGG-16 network on down-sampled ImageNet, respectively.
Foroozan Karimzadeh, Arijit Raychowdhury
VLSI-SOC1