EDBT 2026 Demo / reviewers in the wild / expert
Utkarsh Saxena
dblp:230/8317
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0007-0042-2413ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRAGE:MRAM-Based Near ADC-Less Compute-In-Memory Macro for Deep Learning AccelerationabstractNon-volatile memory (NVM) based Compute-in-Memory (CiM) architectures have emerged as a promising compute primitive for accelerating deep neural networks (DNNs) by performing in-situ matrix–vector multiplications (MVMs). Among various NVMs, STT-MRAM (Spin Transfer Torque based Magnetoresistive Random Access Memory) shows potential due to its high endurance, low energy consumption and high density. However, existing STT-MRAM CiM designs typically rely on multi-bit analog-to-digital converters (ADCs) at the peripherals to digitize accumulated bit-line currents. While enabling high-precision computation, ADCs add substantial energy, latency, and area overheads. To alleviate such problems, we propose a system-technology co-design approach to a Near ADC-Less CiM design with ternary partial-sums called MIRAGE. The accuracy is maintained by considering hardware level partial sum quantization in the training loop. Specifically, we develop an STT-MRAM based CiM macro which features differential bitcells and an adaptive threshold sensing that is amenable to the requirements posed by ternary partial-sum quantization. We do a thorough energy, area, latency, and sense margin analysis along with robust benchmarking against conventional 1T-1MTJ (1 transistor-1 Magnetoresistive Tunnel Junction) based MRAM CiM. The proposed CiM macro occupies ∼ 20% less area, consumes 1.8× less MVM energy and shows 5× better latency with improved distinguishability compared to 1T-1MTJ CiM macro while achieving better accuracy. Mainakh Mukherjee, Ayan B. Pranta, Utkarsh Saxena, Anushka Mukherjee, K. Gaurav Kumar, Kaushik Roy 0001 |
DATE | 3 |
| 2025 | HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning WorkloadsabstractAnalog Compute-in-Memory (CiM) accelerators are increasingly recognized for their efficiency in accelerating Deep Neural Networks (DNNs). However, their dependence on Analog-to-Digital Converters (ADCs) for accumulating partial sums from crossbars leads to substantial power and area overhead. Moreover, the high area overhead of ADCs constrains throughput due to the limited number of ADCs that can be integrated per crossbar. To mitigate this issue, extreme low-precision quantization (binary or ternary) for partial sums can be adopted, eliminating the need for ADCs. While this strategy effectively reduces ADC costs, it introduces the challenge of managing numerous floating-point scale factors, which are trainable parameters like DNN weights. These scale factors must be multiplied with the binary or ternary outputs at the crossbar columns to maintain system accuracy, offsetting the benefits of CiM and partial sum quantization. To that effect, we propose an algorithm-hardware co-design approach. Initially, DNNs are trained with two-stage quantization-aware training. Subsequently, we introduce HCiM, an ADC-Less Hybrid Analog-Digital CiM accelerator. HCiM uses analog CiM crossbars for performing Matrix-Vector Multiplication operations, coupled with a digital CiM array for processing scale factors. Compared to an analog CiM baseline architecture using 7 and 2-bit ADCs, HCiM provides energy reductions up to 28× and 11×, respectively, with a minimal drop in accuracy. Shubham Negi, Utkarsh Saxena, Kaushik Roy 0001 |
ASP-DAC | 2 |
| 2025 | ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank ResidualsabstractPost-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors to 4-bit without significantly degrading generalizability is challenging, due to the high quantization error caused by extreme outliers in activations. To tackle this problem, we propose ResQ, a PTQ method that pushes further the state-of-the-art. By means of principal component analysis (PCA), it identifies a low-rank subspace (in practice 1/8 of the hidden dimension) in which activation variances are highest, and keep the coefficients within this subspace in high precision, e.g. 8-bit, while quantizing the rest to 4-bit. Within each subspace, invariant random rotation is applied to further suppress outliers. We show that this is a provably optimal mixed precision quantization scheme that minimizes error. With the Llama and Qwen2.5 families of models, we demonstrate that ResQ outperforms recent uniform and mixed precision PTQ methods on a variety of benchmarks, achieving up to 33% lower perplexity on Wikitext than the next best method SpinQuant, and upto 3X speedup over 16-bit baseline. Anonymous code repository available at https://anonymous.4open.science/r/project-resq-2142. Utkarsh Saxena, Sayeh Sharify, Kaushik Roy 0001 |
ICML | 1 |
| 2023 | McQueen: Mixed Precision Quantization of Early Exit Networks
Utkarsh Saxena, Kaushik Roy 0001 |
BMVC | 1 |
| 2023 | Partial-Sum Quantization for Near ADC-Less Compute-In-Memory AcceleratorsabstractResistive Crossbar (Xbar) Array based Compute-in-Memory (CiM) accelerators form an attractive hardware substrate for acceleration of Deep Neural Networks (DNNs) on edge devices. They perform highly efficient Matrix Vector Multiplication (MVM) operation, employing the power of analog compute. However, efficiency gains with CiM accelerators are limited due to the overhead posed by peripheral circuits, primarily the Analog-to-Digital Converters (ADCs). In this work, we improve efficiency of CiM accelerators by developing ADC-Less and near ADC-Less CiM accelerators which either eliminate or minimize the ADC overhead. More specifically, we leverage partial-sum quantization to reduce ADC precision to binary (1-bit) or ternary (1.5-bit) values. Xbars with binary partial sums require a sense amplifier for analog-to-digital conversion leading to ADC-Less design. Xbars with ternary partial-sums require two comparators for the conversion process leading to a near ADC-Less design. We develop a CiM hardware aware DNN quantization methodology to mitigate accuracy degradation with partial-sum quantization. We show the effectiveness of our training methodology by achieving high accuracies and minimal accuracy degradation on CIFAR-10 and Imagenet datasets. Consequently, we achieve 14x, 178x and 131x improvements over baseline (8-bit ADC) in Energy, Latency and Compute Efficiency (TOPS/mm2), respectively on Resnet-20 (CIFAR-10) with ADC-Less design and 11x, 55x and 36x improvements over baseline (8-bit ADC) in Energy, Latency and TOPS/mm2, respectively on Resnet-18 (ImageNet) with Near ADC-Less design. Utkarsh Saxena, Kaushik Roy 0001 |
ISLPED | 1 |
| 2022 | Towards ADC-Less Compute-In-Memory Accelerators for Energy Efficient Deep LearningabstractCompute-in-Memory (CiM) hardware has shown great potential in accelerating Deep Neural Networks (DNNs). However, most CiM accelerators for matrix vector multiplication rely on costly analog to digital converters (ADCs) which becomes a bottleneck in achieving high energy efficiency. In this work, we propose a hardware-software co-design approach to reduce the aforementioned ADC costs through partial-sum quantization. Specifically, we replace ADCs with 1-bit sense amplifiers and develop a quantization aware training methodology to compensate for the loss in representation ability. We show that the proposed ADC-less DNN model achieves 1.1x-9.6x reduction in energy consumption while maintaining accuracy within 1% of the DNN model without partial-sum quantization. Utkarsh Saxena, Indranil Chakraborty, Kaushik Roy 0001 |
DATE | 1 |
| 2022 | Compute-in-Memory Technologies and Architectures for Deep Learning WorkloadsabstractThe use of deep learning (DL) to real-world applications, such as computer vision, speech recognition, and robotics, has become ubiquitous. This can be largely attributed to a virtuous cycle between algorithms, data and computing, and storage capacity, which has driven rapid advances in all these dimensions. The ever-increasing demand for computation and memory from DL workloads presents challenges across the entire spectrum of computing platforms, from edge devices to the cloud. Hence, there is a need to explore new hardware paradigms that go well beyond the current mainstays such as graphical processing units (GPUs), tensor processing units (TPUs), and neural processing units (NPUs). A key bottleneck of current platforms is the so-called memory wall, which arises from the need to move large amounts of data between memory and compute units, expending considerable time and energy. One promising solution to this challenge is to move some computations either closer to memory, within the memory subsystem, or even within individual memory arrays. This approach, which is broadly referred to as compute-in-memory (CiM), has the potential to break the memory wall, and thereby greatly improve speed and power consumption. In this article, we provide an overview of CiM techniques used at different levels of the memory hierarchy and based on different memory technologies, including static random access memories (SRAMs), nonvolatile memories (NVMs), and DRAMs. We also discuss architectural approaches to designing CiM-based DL accelerators. Finally, we discuss the challenges associated with adopting CiM in future DL accelerators. Mustafa Fayez Ali, Sourjya Roy, Utkarsh Saxena, Tanvi Sharma, Anand Raghunathan, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Revisiting Stochastic Computing in the Era of Nanoscale Nonvolatile TechnologiesabstractIn this era of nanoscale technologies, the inherent characteristics of some nonvolatile devices, such as resistive random access memory (ReRAM), phase-change material (PCM), and spintronics, can emulate stochastic functionalities. Traditionally, these devices have been engineered to suppress the stochastic switching behavior as it poses reliability concerns for memory storage and logic applications. However, leveraging stochasticity in such devices led to a renewed interest in hardware-software codesign of stochastic algorithms since the CMOS-based implementations of stochastic algorithms involve cumbersome circuitry to generate “stochastic bits.” In this article, we consider two classes of problems: deep neural networks (DNNs) and combinatorial optimization. The rapidly growing demands of artificial intelligence (AI) have sparked an interest in energy-efficient implementations of large DNNs, with binary representations of synaptic weights and neuronal activities. Stochasticity plays an important role in leveraging the benefits of these binary representations, leading to model compression and optimization during training. In combinatorial optimization, such as graph coloring or traveling salesman problems, stochastic algorithms, such as the Ising computing model, have been shown to be effective. These problems require exhaustive computational procedures, and the Ising model uses a natural annealing agent to achieve near-optimal solutions in a reasonable timescale, without getting stuck in “local minima.” In this article, we present a broad review of stochastic computing utilizing the stochastic switching characteristics of devices based on nanoscale nonvolatile technologies. We show how to codesign of the devices and algorithms that can enable optimal solutions for both combinatorial problems and binary neural networks for local learning and inference. Directly mapping the nonvolatile device characteristics to the stochastic algorithms without the need for storing the bits in a separate memory leads to efficient use of hardware. Amogh Agrawal, Indranil Chakraborty, Deboleena Roy, Utkarsh Saxena, Saima Sharmin, Minsuk Koo, Yong Shim, Gopalakrishnan Srinivasan, Chamika M. Liyanagedera, Abhronil Sengupta, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |