EDBT 2026 Demo / reviewers in the wild / expert
Ashish Reddy Bommana
dblp:338/7537
· DBLP profile ↗
11ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0001-6321-6544ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 first-author · 10 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noise-Agnostic One-Shot Training and Retraining for Robust DNN Inferencing on Analog Compute-in-Memory SystemsabstractAnalog Compute-in-Memory (ACiM) architectures are a promising alternatives to traditional von Neumann-based systems for accelerating deep neural networks (DNNs), as they alleviate the memory bottleneck by performing in-situ matrixvector multiplications. However, the analog nature of computation in ACiM makes DNNs highly susceptible to noise and process variations. To mitigate the effects of analog noise, existing approaches rely on variation-aware or noise-aware training, retraining, or fine-tuning. These methods, however, are not scalable, as they require chip-specific retraining and typically involve separate training runs for different levels of noise tolerance. Moreover, they overlook the inherent fault tolerance of analog-to-digital converters (ADCs). To address these limitations, we propose a one-shot training and retraining strategy for robust DNN inferencing on ACiM platforms. Our method is guided by a detailed analysis of error propagation through ADCs, revealing that robustness can be enhanced by strategically reshaping the weight distribution to better align with ADC resilience characteristics. Simulation results and experimental results with fabricated chips show that the proposed method improves inferencing accuracy by $\mathbf{7 0 \%} \boldsymbol{-} \mathbf{9 0 \%}$ for ResNet-18 and DenseNet-121 under $\mathbf{7 0 \%}$ noise injection on CIFAR-10 and SVHN, and by $\mathbf{5 0 \%}$-80% for VGG-16 under $\mathbf{5 0 \%}$ noise. These gains are achieved with only a $5 \%$ energy overhead due to the modified weight distribution. Ashish Reddy Bommana, Ben Feinberg, T. Patrick Xiao, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty |
ASP-DAC | 1 |
| 2026 | WARP: Workload-Aware Reference Prediction for Reliable Multi-Bit FeFET Readout under Charge-Trapping DegradationabstractFerroelectric FET (FeFET)-based arrays are promising candidates for energy-efficient, high-density non-volatile memory in data-intensive applications. However, charge-trappinginduced degradation and process variations pose significant reliability challenges. These effects lead to reduced memory window and degraded read accuracy over time. We propose a workload-aware degradation modeling and readout framework for FeFET arrays. First, we select a small set of representative workloads to efficiently capture degradation trends across a large workload space. We apply a two-step method to reduce read error: (a) adjust intermediate state currents to widen the separation between states; (b) select optimum reference thresholds based on the shifted distributions. Next, we perform detailed tradeoff analysis involving degradation improvement, on-chip area, and the overhead of a memory-mapped CPU polling system for in-field workload tracking. This is the first work to propose adaptive reference prediction for FeFETs based on runtime workload characteristics. Our framework improves read reliability with minimum hardware overhead and enables scalable in-field monitoring for future FeFET-based systems. Dhruv Thapar, Ashish Reddy Bommana, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
ASP-DAC | 2 |
| 2026 | Encoded Repair Configuration Chains for Die-to-Die Interconnect Test and Repair Language
Ashish Reddy Bommana, Anshuman Chandra, Moiz Khan |
ETS | 1 |
| 2026 | MoD-CiM: A Mixture-of-Defenses Framework Against Power-Hammering Attacks in Multi-Tenant Compute-in-Memory
Ashish Reddy Bommana, Soyed Tuhin Ahmed, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 1 |
| 2026 | COMET-3D: Compute-in-Memory-Based Transformer Accelerator With Optimized Pipeline and 3D Heterogeneous IntegrationabstractTransformers have become the backbone of large language decoder and encoder models, but their compute- and memory-intensive nature makes them inefficient on traditional von Neumann architectures. Compute-in-memory (CIM) architectures offer a promising path forward by enablingin situmatrix operations and reducing memory access overhead. However, existing CIM-based accelerators suffer from: (1) unrealistic assumption of single-cycle activation of entire crossbar; (2) inefficient pipelines that are not optimized for operational unit (OU)-based execution; and (3) high analog-to-digital converter (ADC) cost. To address these limitations, we propose a latency-optimized pipeline tailored specifically for OU-based CIM execution and introduce COMET-3D—a 3D heterogeneous architecture that integrates SRAM and ReRAM-based CIM arrays with a logic die containing digital MAC units and softmax modules. The proposed architecture and dataflow maximize hardware utilization and enable efficient acceleration of the multi-head self-attention (MHSA) layer. Experimental results across BERT, GPT2, and LLAMA demonstrate that COMET-3D outperforms baseline architectures with similar compute resources by up to 34× in energy-delay product (EDP) for LLAMA, with gains of 7.4× for GPT2 and 4.3× for BERT-Large. Ashish Reddy Bommana, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | PHANTOM: Power Hammering Attack and Countermeasure on Multi-Tenant ReRAM Compute-in-Memory AcceleratorsabstractThe increasing demand for efficient and low-power deep neural network (DNN) inference has advanced the adoption of ReRAM-based compute-in-memory (CiM) accelerators, which perform computations directly within memory to reduce energy consumption and enhance throughput. However, such architectures are vulnerable to security threats, especially in a multi-tenant environment where multiple users share the same physical resources. This paper introduces a new attack model for multi-tenant ReRAM-based CiM, power hammering, that exploits the temperature sensitivity of ReRAM cells, inducing local temperature increases that lead to conductance drift and ultimately result in erroneous inference outcomes. This serves as a denial-of-service (DoS) attack, where malicious co-tenants degrade inferencing accuracy and system reliability for legitimate users in a shared environment, ultimately undermining trust and causing potential losses to the service provider. Additionally, we propose a novel strategy to counter this security vulnerability. In this technique, we focus on selectively protecting important weights with error compensation hardware. These important weights are treated as faults, and their computation is offloaded to compensation hardware. Simulation results confirm the effectiveness of the proposed method in ensuring accurate classification results even under adversarial conditions, thereby enabling secure multi-tenant inference on ReRAM-based CiM accelerators. Ashish Reddy Bommana, Rajendra Bishnoi, Naghmeh Karimi, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | DEAR: Dependable 3D Architecture for Robust DNN TrainingabstractReRAM-based compute-in-memory (CiM) architectures present an attractive design choice for accelerating deep neural network (DNN) training. However, these architectures are susceptible to stuck-at faults (SAFs) in ReRAM cells, which arise from manufacturing defects and cell wearout over time, particularly due to the continuous weight updates during DNN training. These faults significantly degrade accuracy and compromise dependability. To address this issue, we propose DEAR: dependable 3D architecture for robust DNN training. DEAR introduces a novel online compensation method that employs a digital compensation unit to correct SAF-induced errors dynamically during both forward and backward propagation. This approach mitigates errors induced by SAFs during both the forward and backward phases of DNN training. Additionally, DEAR leverages an HBM-based 3D memory structure to store fault-related error information efficiently. Experimental results show that DEAR limits inferencing accuracy loss to under 2% even when up to 10% of cells are faulty with uniformly distributed faults, and under 2% for up to 5% faulty cells in clustered distributions. This high fault tolerance is achieved with an area overhead of 11.5% and energy overhead of less than 6% for VGG networks and less than 12% for ResNet networks. Ashish Reddy Bommana, Farshad Firouzi, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
DATE | 1 |
| 2025 | Taming Sparse Giants: Deploying Mixture-of-Experts on 3D Heterogeneous Compute-in-Memory SystemsabstractThe deployment of large Mixture-of-Experts (MoE) models on 3D heterogeneous integrated (3D-HI) Compute-in-Memory (CiM) architectures presents unique challenges, requiring joint optimization of area, energy, latency, and perplexity (PPL). We first introduce a detailed 3D-stacked CiM architecture model, incorporating both SRAM and ReRAM tiers with thermal and device-level considerations. Building on this foundation, we present OPTIMEX, a multi-objective optimization framework that efficiently maps MoE expert projections onto heterogeneous tiers. Our evaluation demonstrates substantial benefits: up to$\text{6 0. 9 \%}$area and$\text{5 4. 7 \%}$energy reduction versus an all-SRAM baseline, while lowering PPL by as much as 98.4% compared to all-ReRAM configurations. Furthermore, OPTIMEX outperforms common heuristics, delivering improvements of up to 73.1% in area, 67.0% in energy, 96.1% in PPL, and 74.9% in latency. Together, these contributions highlight a path toward scalable, energy-efficient, and reliable MoE deployment on advanced CiM platforms. Ashish Reddy Bommana, Farshad Firouzi, Krishnendu Chakrabarty |
ICCD | 2 |
| 2025 | Descriptive Language For 3D IC Die-to-Die Interconnect Repair For IEEE P3405 StandardabstractFor any unambiguous transfer of information it is critical to provide a comprehensive description of components, operations, procedures, methodology etc. defined by a technical standard. Many standards provide a descriptive language to capture and transfer information. In this paper we present a descriptive language model for Die-to-Die (D2D) interconnect test and repair, which is under development in the framework of IEEE P3405 standard. We show its capability to describe interconnect, repair scheme, and its application to existing interconnect standards like UCIe, AIB and HBM. We report the impact of various design parameters in a repair IP on the algorithm’s runtime, which extracts the repair algorithm and generates valid repair solutions from the given descriptive language in JSON format. Ashish Reddy Bommana, Anshuman Chandra, Moiz Khan |
ITC-Asia | 1 |
| 2024 | SEC-CiM: Selective Error Compensation for ReRAM-based Compute-in-Memory*abstractReRAM-based Compute-in-Memory (CiM) architectures offer an attractive design choice for accelerating Convolutional Neural Network (CNN) inferencing in edge computing environments. However, these architectures are susceptible to stuck-at-faults (SAFs) in ReRAM cells stemming from manufacturing defects and cell wearout over time, significantly degrading CNN inferencing accuracy. To address this challenge, we propose a technique called Selective Error Compensation for CiM (SEC-CiM). This technique strategically mitigates errors by leveraging the insight that compensating for errors in a limited number of selected columns in a crossbar is sufficient to maintain CNN inferencing accuracy. With this strategy, SEC-CiM achieves significantly lower overhead compared to previous work. Notably, it effectively addresses errors resulting from stuck-at intermediate levels, a critical aspect that was previously overlooked. We develop a theoretical framework to determine the minimum number of columns requiring error compensation. Simulation results demonstrate that SEC-CiM limits the drop in inferencing accuracy to 2% for the ResNet18 and VGG16 models, even when up to 30% of the ReRAM cells in the crossbar are faulty. Similarly, for the Densenet121 CNN, comparable accuracy results are obtained when up to 15% of the ReRAM cells are faulty. We achieve this high level of fault tolerance with moderate area and power consumption overhead of 12.2% and 10.2%, respectively. Ashish Reddy Bommana, Farshad Firouzi, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ITC | 1 |
| 2023 | Design of Synthesis-time Vectorized Arithmetic Hardware for Tapered Floating-point Addition and SubtractionabstractEnergy efficiency has become the new performance criterion in this era of pervasive embedded computing; thus, accelerator-rich multi-processor system-on-chips are commonly used in embedded computing hardware. Once computationally intensive machine learning applications gained much traction, they are now deployed in many application domains due to abundant and cheaply available computational capacity. In addition, there is a growing trend toward developing hardware accelerators for machine learning applications for embedded edge devices where performance and energy efficiency are critical. Although these hardware accelerators frequently use floating-point operations for accuracy, reduced-width floating-point formats are also used to reduce hardware complexity; thus, power consumption while maintaining accuracy. Vectorization concepts can also be used to improve performance, energy efficiency, and memory bandwidth. We propose the design of a vectorized floating-point adder/subtractor that supports arbitrary length floating-point formats with varying exponent and mantissa widths in this article. In comparison to existing designs in the literature, the proposed design is 2.57× area- and 1.56× power-efficient, and it supports true vectorization with no restrictions on exponent and mantissa widths. Ashish Reddy Bommana, Susheel Ujwal Siddamshetty, Pudi Dhilleswararao, Arvind Thumatti K. R., Srinivas Boppu, M. Sabarimalai Manikandan, Linga Reddy Cenkeramaddi |
ACM Trans. Design Autom. Electr. Syst. | 1 |