Shubham Negi

dblp:204/6363 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
abstract
Modern machine learning accelerators are designed to efficiently execute deep neural networks, but emerging models increasingly rely on compound operations that introduce significant off-chip memory traffic challenges. As model sizes continue to grow, computation must be distributed across spatial clusters, which requires frequent and complex collective communication. Existing dataflow optimization frameworks and performance models lack the explicit modeling of these collective communication costs, limiting their applicability. To address this, we propose COMET, a framework that introduces a novel representation to explicitly model collective communication across spatial clusters, alongside latency and energy cost models for operation-level dependencies. By enabling collective-aware modeling, COMET allows for broader mapping exploration, achieving up to a $1.42 \times$ speedup for GEMM-Softmax and $3.46 \times$ for GEMM-LayerNorm and $1.82 \times$ for self-attention compared to unfused baselines.
Shubham Negi, Manik Singhal, Aayush Ankit, Sudeep Bhoja, Kaushik Roy 0001
ISPASS1
2025 HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads
abstract
Analog Compute-in-Memory (CiM) accelerators are increasingly recognized for their efficiency in accelerating Deep Neural Networks (DNNs). However, their dependence on Analog-to-Digital Converters (ADCs) for accumulating partial sums from crossbars leads to substantial power and area overhead. Moreover, the high area overhead of ADCs constrains throughput due to the limited number of ADCs that can be integrated per crossbar. To mitigate this issue, extreme low-precision quantization (binary or ternary) for partial sums can be adopted, eliminating the need for ADCs. While this strategy effectively reduces ADC costs, it introduces the challenge of managing numerous floating-point scale factors, which are trainable parameters like DNN weights. These scale factors must be multiplied with the binary or ternary outputs at the crossbar columns to maintain system accuracy, offsetting the benefits of CiM and partial sum quantization. To that effect, we propose an algorithm-hardware co-design approach. Initially, DNNs are trained with two-stage quantization-aware training. Subsequently, we introduce HCiM, an ADC-Less Hybrid Analog-Digital CiM accelerator. HCiM uses analog CiM crossbars for performing Matrix-Vector Multiplication operations, coupled with a digital CiM array for processing scale factors. Compared to an analog CiM baseline architecture using 7 and 2-bit ADCs, HCiM provides energy reductions up to 28× and 11×, respectively, with a minimal drop in accuracy.
Shubham Negi, Utkarsh Saxena, Kaushik Roy 0001
ASP-DAC1
2024 Best of Both Worlds: Hybrid SNN-ANN Architecture for Event-based Optical Flow Estimation
abstract
In the field of robotics, event-based cameras are emerging as a promising low-power alternative to traditional frame-based cameras for capturing high-speed motion and high dynamic range scenes. This is due to their sparse and asynchronous event outputs. Spiking Neural Networks (SNNs) with their asynchronous event-driven compute, show great potential for extracting the spatio-temporal features from these event streams. In contrast, the standard Analog Neural Networks (ANNs1) fail to process event data effectively. However, training SNNs is difficult due to additional trainable parameters (thresholds and leaks), vanishing spikes at deeper layers, and a non-differentiable binary activation function. Furthermore, an additional data structure, "membrane potential", responsible for keeping track of temporal information, must be fetched and updated at every timestep in SNNs. To overcome these challenges, we propose a novel SNN-ANN hybrid architecture that combines the strengths of both. Specifically, we leverage the asynchronous compute capabilities of SNN layers to effectively extract the input temporal information. Concurrently, the ANN layers facilitate training and efficient hardware deployment on traditional machine learning hardware such as GPUs. We provide extensive experimental analysis for assigning each layer to be spiking or analog, leading to a network configuration optimized for performance and ease of training. We evaluate our hybrid architecture for optical flow estimation on DSEC-flow and Multi-Vehicle Stereo Event-Camera (MVSEC) datasets. On the DSEC-flow dataset, the hybrid SNN-ANN architecture achieves a 40% reduction in average endpoint error (AEE) with 22% lower energy consumption compared to Full-SNN, and 48% lower AEE compared to Full-ANN, while maintaining comparable energy usage.
Shubham Negi, Adarsh Kosta, Kaushik Roy 0001
IROS1
2022 NAX: neural architecture and memristive xbar based accelerator co-design
abstract
Neural Architecture Search (NAS) has provided the ability to design efficient deep neural network (DNN) catered towards different hardwares like GPUs, CPUs etc. However, integrating NAS with Memristive Crossbar Array (MCA) based In-Memory Computing (IMC) accelerator remains an open problem. The hardware efficiency (energy, latency and area) as well as application accuracy (considering device and circuit non-idealities) of DNNs mapped to such hardware are co-dependent on network parameters such as kernel size, depth etc. and hardware architecture parameters such as crossbar size and the precision of analog-to-digital converters. Co-optimization of both network and hardware parameters presents a challenging search space comprising of different kernel sizes mapped to varying crossbar sizes. To that effect, we propose NAX - an efficient neural architecture search engine that co-designs neural network and IMC based hardware architecture. NAX explores the aforementioned search space to determine kernel and corresponding crossbar sizes for each DNN layer to achieve optimal tradeoffs between hardware efficiency and application accuracy. For CIFAR-10 and Tiny ImageNet, our models achieve 0.9% and 18.57% higher accuracy at 30% and -10.47% lower EDAP (energy-delay-area product), compared to baseline ResNet-20 and ResNet-18 models, respectively.
Shubham Negi, Indranil Chakraborty, Aayush Ankit, Kaushik Roy 0001
DAC1
2021 Modeling and Analysis of High-Performance Triple Hole Block Layer Organic LED Based Light Sensor for Detection of Ovarian Cancer
abstract
In this paper a novel triple hole block layer (HBL) structure of the OLED is proposed that depicts an enhanced luminescence of 25285 cd/m2with an improvement of 47% over multilayered OLED architecture. It also owes 74% improvement in luminous power efficiency. An in-depth numerical analysis based on Poisson and drift diffusion equation is undertaken and validated against the internal device analysis. The analysis results highlight an enhanced recombination rate within the proposed device. High electron injection and efficient hole blocking contributes to improved recombination rate. Triple HBL OLED is therefore used for diagnosis of ovarian cancer. The device illustrated good response towards varying wavelengths generating a maximum photo current value of 93 mA. A healthy person can be differentiated from an oncological cancer patient based on fluorescence produced by their urine. The fluorescence values for healthy person and oncological cancer patient are in the range of 420 and 440 nm, correspondingly. The cathode current produced by OLED corresponding to these two wavelengths are 5 and 1 mA respectively. Hence, the proposed device can successfully diagnose the ovarian cancer patient. Further, the methodology proposed for diagnosis of ovarian cancer can help in developing a portable, flexible low cost biomedical sensor.
Shubham Negi, Poornima Mittal
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 Methodology for Realizing VMM with Binary RRAM Arrays: Experimental Demonstration of Binarized-ADALINE using OxRAM Crossbar
abstract
In this paper, we present an efficient hardware mapping methodology for realizing vector matrix multiplication (VMM) on resistive memory (RRAM) arrays. Using the proposed VMM computation technique, we experimentally demonstrate a binarized-ADALINE (Adaptive Linear) classifier on an OxRAM crossbar. An 8×8 OxRAM crossbar with Ni/3-nm HfO2/7 nm Al-doped-TiO2/TiN device stack is used. Weight training for the binarized-ADALINE classifier is performed ex-situ on UCI cancer dataset. Post weight generation the OxRAM array is carefully programmed to binary weight-states using the proposed weight mapping technique on a custom-built testbench. Our VMM powered binarized-ADALINE network achieves a classification accuracy of 78% in simulation and 67% in experiments. Experimental accuracy was found to drop mainly due to crossbar inherent sneak-path issues and RRAM device programming variability.
Sandeep Kaur Kingra, Vivek Parmar, Shubham Negi, Sufyan Khan, Boris Hudec, Tuo-Hung Hou, Manan Suri
ISCAS3