Deboleena Roy

dblp:210/1009 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0001-9973-7237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Meta's Second Generation AI Chip: Model-Chip Co-Design and Productionization Experiences
abstract
The rapid growth of AI workloads at Meta has motivated our inhouse development of AI chips, aiming to significantly reduce the total cost of ownership and mitigate risks posed by unpredictable GPU supplies.At ISCA'23, we presented Meta's first-generation AI chip, MTIA 1.This paper describes its successor, MTIA 2i, now deployed at scale and serving billions of users.MTIA 2i significantly improves upon MTIA 1, reducing total cost of ownership by 44% compared to GPUs while delivering competitive performance per watt.A key differentiator is its memory hierarchy: instead of costly HBM, it uses large SRAM alongside LPDDR.Although there has been a proliferation of publications on AI chips, they often focus on architectural design and overlook three critical aspects:(1) co-designing and optimizing ML models to work effectively with the AI chip; (2) demonstrating sufficient flexibility to support a wide range of models; and (3) during the productionization process, addressing challenges unanticipated or decisions deferred at design time, such as dealing with memory errors, safe overclocking, reducing provisioned power, and implementing real-time firmware updates to mitigate silicon design defects.A key contribution of this paper is sharing our experience with these aspects, based on our journey of productionizing MTIA 2i at scale.
Joel Coburn, Chunqiang Tang, Sameer Abu Asal, Neeraj Agrawal, Raviteja Chinta, Harish Dattatraya Dixit, Brian Dodds, Saritha Dwarakapuram, Amin Firoozshahian, Cao Gao, Kaustubh Gondkar, Tyler Graf, Junhan Hu, Sterling Hughes, Adam Hutchin, Bhasker Jakka, Guoqiang Jerry Chen, Indu Kalyanaraman, Ashwin Kamath, Pankaj Kansal, Erum Kazi, Roman Levenstein, Mahesh Maddury, Alex Mastro, Siji Medaiyese, Pritesh Modi, Jack Montgomery, Nadathur Satish, Amit Nagpal, Ashwin Narasimha, Maxim Naumov, Eleanor Ozer, Jongsoo Park, Poorvaja Ramani, Harikrishna Reddy, David Reiss, Deboleena Roy, Sathish Sekar, Pavan Shetty, Aravind Sukumaran-Rajam, Eran Tal, Mike Tsai, Shreya Varshini, Richard Wareing, Olívia Wu, Xiaolong Xie, Hangchen Yu, Tanmay Zargar, Zitong Zeng, Feixiong Zhang, Ajit Mathews, Jiyuan Zhang 0008, Emmanuel Menage, Truls Edvard Stokke, Mohammed Sourouri
ISCA38
2022 On Noise Stability and Robustness of Adversarially Trained Networks on NVM Crossbars
abstract
Applications based on deep neural networks (DNNs) have grown exponentially in the past decade. To match their increasing computational needs, several nonvolatile memory (NVM) crossbar-based accelerators have been proposed. Recently, researchers have shown that apart from improved energy efficiency and performance, such approximate hardware also possess intrinsic robustness for defense against adversarial attacks. Prior works have focused on quantifying this intrinsic robustness for vanilla networks, that is DNNs trained on unperturbed inputs. However, adversarial training of DNNs, i.e., training with adversarially perturbed images, is the benchmark technique for robustness, and sole reliance on intrinsic robustness of the hardware may not be sufficient. In this work, we explore the design of robust DNNs through the amalgamation of adversarial training and the intrinsic robustness offered by NVM crossbar-based analog hardware. First, we study the noise stability of such networks on unperturbed inputs and observe that internal activations of adversarially trained networks have lower signal-to-noise ratio (SNR), and are sensitive to noise compared to vanilla networks. As a result, they suffer significantly higher performance degradation due to the approximate computations on analog hardware; on an average$2\times $accuracy drop. Noise stability analyses clearly show the instability of adversarially trained DNNs. On the other hand, for adversarial images generated using Square Black Box attacks, ResNet-10/20 adversarially trained on CIFAR-10/100 display a robustness improvement of 20%–30% under high$\epsilon _{\mathrm{ attack}}$(degree of input perturbation). For adversarial images generated using projected-gradient-descent (PGD) White-Box attacks, the adversarially trained DNNs present a 5%–10% gain in robust accuracy due to the underlying NVM crossbar when$\epsilon _{\mathrm{ attack}}$is greater than theepsilonof the adversarial training ($\epsilon _{\mathrm{ train}}$). Our results indicate that implementing adversarially trained networks on analog hardware requires careful calibration between hardware nonidealities and$\epsilon _{\mathrm{ train}}$to achieve optimum robustness and performance.
Chun Tao, Deboleena Roy, Indranil Chakraborty, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2021 On the Intrinsic Robustness of NVM Crossbars Against Adversarial Attacks
abstract
The increasing computational demand of Deep Learning has propelled research in special-purpose inference accelerators based on emerging non-volatile memory (NVM) technologies. Such NVM crossbars promise fast and energy-efficient in-situ Matrix Vector Multiplication (MVM) thus alleviating the long-standing von Neuman bottleneck in today’s digital hardware. However, the analog nature of computing in these crossbars is inherently approximate and results in deviations from ideal output values, which reduces the overall performance of Deep Neural Networks (DNNs) under normal circumstances. In this paper, we study the impact of these non-idealities under adversarial circumstances. We show that the non-ideal behavior of analog computing lowers the effectiveness of adversarial attacks, in both Black-Box and White-Box attack scenarios. In a non-adaptive attack, where the attacker is unaware of the analog hardware, we observe that analog computing offers a varying degree of intrinsic robustness, with a peak adversarial accuracy improvement of 35.34%, 22.69%, and 9.90% for white box PGD (ϵ=1/255, iter =30) for CIFAR-10, CIFAR-100, and ImageNet respectively. We also demonstrate “Hardware-in-Loop” adaptive attacks that circumvent this robustness by utilizing the knowledge of the NVM model.
Deboleena Roy, Indranil Chakraborty, Timur Ibrayev, Kaushik Roy 0001
DAC1
2020 Tree-CNN: A hierarchical Deep Convolutional Neural Network for incremental learning
Deboleena Roy, Priyadarshini Panda, Kaushik Roy 0001
Neural Networks1
2020 Revisiting Stochastic Computing in the Era of Nanoscale Nonvolatile Technologies
abstract
In this era of nanoscale technologies, the inherent characteristics of some nonvolatile devices, such as resistive random access memory (ReRAM), phase-change material (PCM), and spintronics, can emulate stochastic functionalities. Traditionally, these devices have been engineered to suppress the stochastic switching behavior as it poses reliability concerns for memory storage and logic applications. However, leveraging stochasticity in such devices led to a renewed interest in hardware-software codesign of stochastic algorithms since the CMOS-based implementations of stochastic algorithms involve cumbersome circuitry to generate “stochastic bits.” In this article, we consider two classes of problems: deep neural networks (DNNs) and combinatorial optimization. The rapidly growing demands of artificial intelligence (AI) have sparked an interest in energy-efficient implementations of large DNNs, with binary representations of synaptic weights and neuronal activities. Stochasticity plays an important role in leveraging the benefits of these binary representations, leading to model compression and optimization during training. In combinatorial optimization, such as graph coloring or traveling salesman problems, stochastic algorithms, such as the Ising computing model, have been shown to be effective. These problems require exhaustive computational procedures, and the Ising model uses a natural annealing agent to achieve near-optimal solutions in a reasonable timescale, without getting stuck in “local minima.” In this article, we present a broad review of stochastic computing utilizing the stochastic switching characteristics of devices based on nanoscale nonvolatile technologies. We show how to codesign of the devices and algorithms that can enable optimal solutions for both combinatorial problems and binary neural networks for local learning and inference. Directly mapping the nonvolatile device characteristics to the stochastic algorithms without the need for storing the bits in a separate memory leads to efficient use of hardware.
Amogh Agrawal, Indranil Chakraborty, Deboleena Roy, Utkarsh Saxena, Saima Sharmin, Minsuk Koo, Yong Shim, Gopalakrishnan Srinivasan, Chamika M. Liyanagedera, Abhronil Sengupta, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2019 On Robustness of Spin-Orbit-Torque Based Stochastic Sigmoid Neurons for Spiking Neural Networks
abstract
Nano-scale neuro-mimetic devices have recently gained wide research interest in the quest to enable brain-like energy-efficiency with cognitive computing abilities. Traditionally, neuromorphic devices have exploited deterministic nano-scale devices for emulating the intrinsic neuronal and synaptic behavior. However, of particular interest are stochastic neuromorphic devices owing to - 1) availability of nano-scale devices that are inherently stochastic based on intrinsic device physics 2) various neuroscience experiments have demonstrated that cortical neurons are stochastic in nature. In this paper, we focus on spin orbit torque based Magnetic Tunnel Junction (SOT-MTJ) that exhibit stochastic sigmoid behavior with respect to the switching process. We first discuss the modeling framework that was used to study the effect of dimensional variations in SOT-MTJs and the resulting changes in the stochastic sigmoid behavior. Our model is based on the well-known stochastic-Landau-Lifshitz-Gilbert-Slonczewski equation under mono-domain approximation. Subsequently, we abstract the sigmoid characteristic of the device into a behavioral model and study the effect of variations in sigmoid characteristics on a deep binary network. Our results show that the variations in the sigmoidal neuron behavior results in a minimal loss in accuracy (for CIFAR 10 dataset). Additionally, the degradation in accuracy monotonically increases with increase in induced variations. This highlights the robustness of stochastic neural networks based on SOT-MTJs in presence of dimensional variations.
Akhilesh Jaiswal 0001, Amogh Agrawal, Indranil Chakraborty, Deboleena Roy, Kaushik Roy 0001
IJCNN4
2019 Neural Networks at the Edge
abstract
As neural networks gain importance with several successful applications of them, this paper raises the question of how they can be applied in the context of coalition operations. A key challenge in military coalition operations is that of energy and severe bandwidth constraints. We address this challenge by exploring the use of Deep Neural Networks (DNNs) and splitting them across multiple edge nodes. Further, we explore the idea of using spiking neural networks that can lower the energy consumption significantly. Preliminary results show that both these approaches can have significant impact on coalition operations.
Deboleena Roy, Gopalakrishnan Srinivasan, Priyadarshini Panda, Richard Tomsett, Nirmit Desai, Raghu K. Ganti, Kaushik Roy 0001
SMARTCOMP1