Anupam Golder

dblp:241/4284 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-0725-1593ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 POSTER: SMAUG-SCA: Machine Learning Based Power Side-Channel Attack on SMAUG-T
abstract
status: Published
Madhumitha Ramaswamy, Rishav Saha, Suparna Kundu, Anupam Golder, Angshuman Karmakar, Debayan Das
AsiaCCS4
2026 Leveraging ASIC AI Chips for Homomorphic Encryption
abstract
Homomorphic Encryption (HE) provides strong data privacy for cloud services but at the cost of prohibitive computational overhead. While GPUs have emerged as a practical platform for accelerating HE, there remains an order-ofmagnitude energy-efficiency gap compared to specialized (but expensive) HE ASICs. This paper explores an alternate direction: leveraging existing AI accelerators, like Google's TPUs with coarse-grained compute and memory architectures, to offer a path toward ASIC-level energy efficiency for HE. However, this architectural paradigm creates a fundamental mismatch with SoTA HE algorithms designed for GPUs. These algorithms rely heavily on: (1) high-precision (32-bit) integer arithmetic to now run on a TPU's low-throughput vector unit, leaving its high-throughput low-precision (8-bit) matrix engine (MXU) idle, and (2) fine-grained data permutations that are inefficient on the TPU's coarse-grained memory subsystem. Consequently, porting GPU-optimized HE libraries to TPUs results in severe resource under-utilization and performance degradation. To tackle above challenges, we introduce CROSS, a compiler framework that systematically transforms HE workloads to align with the TPU's architecture. CROSS makes two key contributions: (1) Basis-Aligned Transformation (BAT), a novel technique that converts high-precision modular arithmetic into dense, low-precision (INT8) matrix multiplications, unlocking and improving the utilization of TPU's MXU for HE, and (2) Memory-Aligned Transformation (MAT), which eliminates costly runtime data reordering by embedding reordering into compute kernels through offline parameter transformation. Our evaluation on a real single-host Google TPU v6e refreshes the SoTA Number Theoretic Transform (NTT) throughput record with up-to$1.43 \times$throughput improvement over WarpDrive on a NVIDIA A100. Furthermore, CROSS achieves$451 \times, 7.81 \times, 1.83 \times, 1.31 \times, 1.86 \times$, and$1.15 \times$higher throughput per watt than OpenFHE, WarpDrive, FIDESlib, FAB, HEAP, and Cheddar, respectively, establishing AI ASIC as the SotA efficient platform for HE operators. Code: https://github.com/EfficientPPML/CROSS.
Jianming Tong, Jingtian Dang, Leo de Castro, Anirudh Itagi, Anupam Golder, Asra Ali, Jeremy Kun, Jevin Jiang, Arvind 0001, G. Edward Suh, Tushar Krishna
HPCA6
2023 Power Side-Channel Vulnerability Assessment of Lightweight Cryptographic Scheme, XOODYAK
abstract
This work presents a power side-channel analysis (SCA) of a lightweight cryptography (LWC) algorithm, XOODYAK, implemented on an FPGA. First, we perform generic leakage detection tests for two phases of authenticated encryption with associated data (AEAD) mode, namely INITIALIZE, and ABSORB. Second, we develop novel hypothetical attack models for correlation power analysis (CPA) and demonstrate a success rate (SR) of 92%/82% and minimum-traces-to-disclosure (MTD)=13K/38K on the INITIALIZE/ABSORB phases, respectively. Third, we evaluate ABSORB against Profiled SCA using convolutional neural network (CNN), and achieve SR=96%/64% and MTD=2K/16K on the test set for the same/different keys used for training, respectively. Finally, we suggest low-overhead countermeasures to protect against these SCA attacks.
Anupam Golder, Debayan Das, Santosh Ghosh, Avinash L. Varna, Majid Sabbagh, Sayak Ray, Rana Elnaggar, Joseph Friel, Daniel Dinu, Jason M. Fung
DAC1
2023 RAGA: Resource-Aware Tree-Splitting for High Performance Knuth-Yao-based Discrete Gaussian Sampling on FPGAs
abstract
In this work, we present a high-performance architecture for Discrete Gaussian (DG) sampling used in lattice-based cryptography (LBC) through lookup table (LUT) optimization for the combinational datapath as well as FPGA architecture aware pipelining to reduce resource utilization and decrease latency. The proposed tree-splitting technique translates the discrete distribution generating (DDG) tree for Knuth-Yao-based non-uniform sampling into a LUT-based logic with pipelined architecture. This allows a reduction in the area-time product (ATP) compared to the most efficient current state-of-the-art (SotA) designs for DG sampling with standard deviation, (σ = 3.19/6.15543/8.5) by 96%/14%/56% on Virtex-7/Virtex-6/Artix-7 FPGAs, while running at 709/580/440 MHz, respectively. The simplicity of the translation technique also allows us to automate the design flow, resulting in quick design-space exploration for different DG distributions.
Zachary J. Ellis, Anupam Golder, Addison J. Elliott, Arijit Raychowdhury
ACM Great Lakes Symposium on VLSI2
2022 Exploration into the Explainability of Neural Network Models for Power Side-Channel Analysis
abstract
In this work, we present a comprehensive analysis of explainability of Neural Network (NN) models in the context of power Side-Channel Analysis (SCA), to gain insight into which features or Points of Interest (PoI) contribute the most to the classification decision. Although many existing works claim state-of-the-art accuracy in recovering secret key from cryptographic implementations, it remains to be seen whether the models actually learn representations from the leakage points. In this work, we evaluated the reasoning behind the success of a NN model, by validating the relevance scores of features derived from the network to the ones identified by traditional statistical PoI selection methods. Thus, utilizing the explainability techniques as a standard validation technique for NN models is justified.
Anupam Golder, Ashwin Bhat, Arijit Raychowdhury
ACM Great Lakes Symposium on VLSI1
2022 EM-X-DL: Efficient Cross-device Deep Learning Side-channel Attack With Noisy EM Signatures
abstract
This work presents a Cross-device Deep-Learning based Electromagnetic (EM-X-DL) side-channel analysis (SCA) on AES-128, in the presence of a significantly lower signal-to-noise ratio (SNR) compared to previous works. Using a novel algorithm to intelligently select multiple training devices and proper choice of hyperparameters, the proposed 256-class deep neural network (DNN) can be trained efficiently utilizing pre-processing techniques like PCA, LDA, and FFT on measurements from the target encryption engine running on an 8-bit Atmel microcontroller. In this way, EM-X-DL achieves >90% single-trace attack accuracy. Finally, an efficient end-to-end SCA leakage detection and attack framework using EM-X-DL demonstrates high confidence of an attacker with <20 averaged EM traces.
Josef Danial, Debayan Das, Anupam Golder, Santosh Ghosh, Arijit Raychowdhury, Shreyas Sen
ACM J. Emerg. Technol. Comput. Syst.3
2019 X-DeepSCA: Cross-Device Deep Learning Side Channel Attack
abstract
This article, for the first time, demonstrates Cross-device Deep Learning Side-Channel Attack (X-DeepSCA), achieving an accuracy of > 99.9%, even in presence of significantly higher inter-device variations compared to the inter-key variations. Augmenting traces captured from multiple devices for training and with proper choice of hyper-parameters, the proposed 256-class Deep Neural Network (DNN) learns accurately from the power side-channel leakage of an AES-128 target encryption engine, and an N-trace (N ≤ 10) X-DeepSCA attack breaks different target devices within seconds compared to a few minutes for a correlational power analysis (CPA) attack, thereby increasing the threat surface for embedded devices significantly. Even for low SNR scenarios, the proposed X-DeepSCA attack achieves ~ 10× lower minimum traces to disclosure (MTD) compared to a traditional CPA.
Debayan Das, Anupam Golder, Josef Danial, Santosh Ghosh, Arijit Raychowdhury, Shreyas Sen
DAC2
2019 Practical Approaches Toward Deep-Learning-Based Cross-Device Power Side-Channel Attack
abstract
Power side-channel analysis (SCA) has been of immense interest to most embedded designers to evaluate the physical security of the system. This work presents profiling-based cross-device power SCA attacks using deep-learning techniques on 8-bit AVR microcontroller devices running AES-128. First, we show the practical issues that arise in these profiling-based cross-device attacks due to significant device-to-device variations. Second, we show that utilizing principal component analysis (PCA)-based preprocessing and multidevice training, a multilayer perceptron (MLP)-based 256-class classifier can achieve an average accuracy of 99.43% in recovering the first keybyte from all the 30 devices in our data set, even in the presence of significant interdevice variations. Results show that the designed MLP with PCA-based preprocessing outperforms a convolutional neural network (CNN) with four-device training by ~20% in terms of the average test accuracy of cross-device attack for the aligned traces captured using the ChipWhisperer hardware. Finally, to extend the practicality of these cross-device attacks, another preprocessing step, namely, dynamic time warping (DTW) has been utilized to remove any misalignment among the traces, before performing PCA. DTW along with PCA followed by the 256-class MLP classifier provides ≥10.97% higher accuracy than the CNN-based approach for cross-device attack even in the presence of up to 50 time-sample misalignments between the traces.
Anupam Golder, Debayan Das, Josef Danial, Santosh Ghosh, Shreyas Sen, Arijit Raychowdhury
IEEE Trans. Very Large Scale Integr. Syst.1