Naina Gupta 0001

dblp:83/7953-1 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0003-3056-9241ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-author · 6 since 2021Security and privacy · 2Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 AI Attacks AI: Recovering Neural Network Architecture from NVDLA Using AI-Assisted Side Channel Attack
abstract
During the last decade, there has been a stunning progress in the domain of Artificial Intelligence (AI) aided by highly trained Machine Learning (ML) models. Such models are valuable Intellectual Property (IP) and, therefore, have been subjected to various model recovery attacks. In this work, we study the vulnerabilities of commercial, open-source accelerator NVDLA and present the first successful model recovery attack. For this purpose, we used power and timing information from the side-channel leakage of convolutional neural networks (CNN) models to train CNN-based attack models. Utilizing these attack models, we demonstrate that even with a highly pipelined architecture, multiple parallel execution in the accelerator along with Linux OS running tasks in the background, recovery of number of layers, kernel sizes, output neurons and distinguishing different layers, is possible with very high accuracy. This is also the first work to show the impact of differences in hyperparameters on the power traces. Our solution is fully automated, AI-based, and portable to other hardware neural networks, thus presenting a greater threat toward IP protection. Using LeNet as the target victim model, we demonstrate an accuracy of more than 95% in recovering various parameters. This study presents a serious practical threat, in the form of side-channel attack, toward complex commercial architectures. Furthermore, we show that AI-guided attack significantly boosts the attacker capability.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
ACM Trans. Embed. Comput. Syst.1
2024 A Configurable CRYSTALS-Kyber Hardware Implementation with Side-Channel Protection
abstract
In this work, we present a configurable and side channel resistant implementation of the post-quantum key-exchange algorithm CRYSTALS-Kyber . The implemented design can be configured for different performance and area requirements leading to different trade-offs for different applications. A low area implementation can be achieved in 5,269 LUTs and 2,422 FFs, whereas a high performance implementation required 7,151 LUTs and 3,730 FFs. Due to a deeply pipelined architecture, a high operating speed of more than 250 MHz could be achieved on 28nm Xilinx FPGAs. The side channel resistance is implemented using a carefully chosen set of novel and known techniques such as Fault Detection Hashes, Instruction Randomization, FSM Protection and so on. resulting in a low overhead of less than 5% while being highly configurable. To the best of our knowledge, this work presents the first side-channel and fault attack protected configurable accelerator for CRYSTALS-Kyber . Using TVLA (test vector leakage assessment), we validate the implemented protection techniques and demonstrate that the design does not leak information even after 200 K traces. Furthermore, one of the configuration choices results in the smallest hardware implementation of CRYSTALS-Kyber known in the literature.
Arpan Jati, Naina Gupta 0001, Anupam Chattopadhyay, Somitra Kumar Sanadhya
ACM Trans. Embed. Comput. Syst.2
2023 CRYSTALS-Dilithium on RISC-V Processor: Lightweight Secure Boot Using Post-Quantum Digital Signature
abstract
With the ongoing efforts for transitioning towards post-quantum security, NIST has recently selected the digital signature algorithm CRYSTALS-Dilithium for standardization. In this work, we demonstrate the first Dilithium based hardware accelerated secure boot architecture developed around Ariane, an open-source RISC- V core. By utilizing a compact design with novel verification engine, a secure boot flow is implemented with only 3.48ms runtime overhead compared to normal boot, while requiring 10.4K LUTs and 5.7K FFs on an FPGA. Compared to the state-of-the-art we achieve a reduction of 3.42× and 7.88 × for LUTs and FFs respectively. Also, the design when realized in 65nm ASIC requires only 125 kGE and 6.3 mW power at 100 MHz. Further, as secure boot is one of the critical processes and the security of the whole system depends on it, we implemented hardware fault countermeasures and evaluated their effectiveness in preventing secure boot bypass.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
ICCAD1
2023 Lightweight Hardware Accelerator for Post-Quantum Digital Signature CRYSTALS-Dilithium
abstract
The looming threat of an adversary with quantum computing capability led to a worldwide research effort towards identifying and standardizing novel post-quantum cryptographic primitives. Post-standardization, all existing security protocols will need to support efficient implementation of these primitives. In this work, we contribute to these efforts by reporting the smallest implementation of CRYSTALS-Dilithium, one of the chosen post-quantum digital signature scheme for NIST standardization process. By invoking multiple optimizations to leverage parallelism, pre-computation and memory access sharing, we obtain an implementation that could be fit into one of the smallest Zynq FPGA. On Zynq Ultrascale+, our design achieves an improvement of about 36.7%/35.4%/42.3% in Area$\times $Time (LUTs$\times \text{s}$) trade-off for KeyGen/Sign/Verify respectively over state-of-the-art implementation. We also evaluate our design as a co-processor on three different hardware platforms and compare the results with software implementation, thus presenting a detailed evaluation of CRYSTALS-Dilithium targeted for embedded applications. Further, on ASIC using TSMC 65nm technology, our design requires 0.227mm2area and can operate at a frequency of 1.176 GHz. As a result, it only requires$53.7\mu \text{s}/96.9\mu \text{s}/57.7\mu \text{s}$for KeyGen/Sign/Verify operation for the best-case scenario.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay, Gautam Jha
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 In Quest for Fast and Secure SoC
abstract
Over the past few years, edge devices has gained a lot of attention. It is mainly due to significant improvements in technology, available processing power and efficiency. This evolution has resulted in edge devices becoming intelligent, smarter and more responsive. Such a growth has also resulted in many security challenges especially with the rise in side-channel attack possibilities. As these devices collect a lot of data and the decision making process is data driven; security of such devices becomes a necessity for safety-critical and time-critical applications.It is a well known fact that security in any system comes with a cost. As many IoT devices are constrained either due to available resources or the time sensitiveness of the decision they are required to make; therefore, in this work, we focus on individual System on Chip (SoC) components to integrate security measures while maintaining a balance between performance and resource utilization.
Naina Gupta 0001, Anupam Chattopadhyay
VLSI-SoC1
2021 MemEnc: A Lightweight, Low-Power, and Transparent Memory Encryption Engine for IoT
abstract
Recent advancement in technologies has led to the widespread adoption and deployment of Internet-of-Things devices. Because of the ubiquitous nature of these devices, they process large amounts of personal and sensitive data. These data are typically stored on DRAM chips, and hence becomes an easy target for attackers. Memory encryption is a commonly adopted solution to provide confidentiality. However, realizing a lightweight, low-latency, low-power solution for resource-constrained devices is a challenge. To address this, we designed MemEnc, a purely hardware-based solution that performs encryption on-the-fly and handles memory requests transparently without any OS intervention. MemEnc runs at a maximum frequency of 401 MHz and requires about 23.2 kGE (gate equivalents) on 65-nm ASIC and consumes only 1.9 mW of power at 250 MHz. Using comprehensive benchmarking, we also analyze the applicability of the proposed solution on real-world workloads. Our experiments show that certain real-time applications can run with full memory encryption and still meet system requirements. Moreover, we integrated our memory encryption engine with ARM TrustZone and present comparative results for the different case studies with Intel SGX. We show that static and dynamic efficiency-security tradeoff is necessary for all scenarios.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
IEEE Internet Things J.1
2021 PQC Acceleration Using GPUs: FrodoKEM, NewHope, and Kyber
abstract
In this article, we present the first GPU implementation for FrodoKEM-976, NewHope-1024, and Kyber-1024. These algorithms belong to three different classes of post-quantum algorithms: Learning with errors (LWE), Ring-LWE, and Module-LWE. We show the practical applicability of the algorithms in different scenarios using two different implementation approaches. Moreover, we achieve highly efficient realization of computationally expensive operations such as NTT (Number Theoretic Transform), matrix multiplication, and Keccak. Since, these are the most common operations in lattice-based cryptographic algorithms, the techniques presented in this article will likely benefit other similar algorithms. Using a NVIDIA QUADRO GV100 graphics card, we undertook a detailed experimental study. For NewHope and Kyber we were able to perform approximately 504K and 473K key exchanges per second, demonstrating a speedup of almost 53.1× and 51.05× compared to the reference C implementation. Compared to the optimized AVX2 versions we obtain speedups of 25.7× and 14.6×, respectively. Further, implementation of FrodoKEM resulted in a speedup of 50.6×, 44.2×, and 36.9× for KeyGen, Encaps and Decaps operations. Compared to its AVX2 counterpart, we achieved a speedup of about 7.3×, 4.7× and 4.9×, respectively. We also show that using multiple streams resulted in further speedup of about 28-38 percent.
Naina Gupta 0001, Arpan Jati, Amit Kumar Chauhan, Anupam Chattopadhyay
IEEE Trans. Parallel Distributed Syst.1
2020 Post-Quantum Secure Boot
abstract
A secure boot protocol is fundamental to ensuring the integrity of the trusted computing base of a secure system. The use of digital signature algorithms (DSAs) based on traditional asymmetric cryptography, particularly for secure boot, leaves such systems vulnerable to the threat of quantum computers. This paper presents the first post-quantum secure boot solution, implemented fully as hardware for reasons of security and performance. In particular, this work uses the eXtended Merkle Signature Scheme (XMSS), a hash-based scheme that has been specified as an IETF RFC. The solution has been integrated into a secure SoC platform around RISC-V cores and evaluated on an FPGA and is shown to be orders of magnitude faster compared to corresponding hardware/software implementations and to compare competitively with a fully hardware elliptic curve DSA based solution.
Vinay B. Y. Kumar, Naina Gupta 0001, Anupam Chattopadhyay, Michael Kasper, Christoph Krauß, Ruben Niederhagen
DATE2
2020 Threshold Implementations of <tt>GIFT</tt>: A Trade-Off Analysis
abstract
Threshold Implementation (TI) is one of the most widely used countermeasure for side channel attacks. Over the years several TI techniques have been proposed for randomizing cipher execution using different variations of secret-sharing and implementation techniques. For instance, sharing without decomposition (4-shares) is the most straightforward implementation of the threshold countermeasure. However, its usage is limited due to its high area requirements. On the other hand, sharing using decomposition (3-shares) countermeasure for cubic non-linear functions significantly reduces area and complexity in comparison to 4-shares. Nowadays, security of ciphers using a side channel countermeasure is of utmost importance. This is due to the wide range of security critical applications from smart cards, battery operated IoT devices, to accelerated crypto-processors. Such applications have different requirements (higher speed, energy efficiency, low latency, small area etc.) and hence need different implementation techniques. Although, many TI strategies and implementation techniques are known for different ciphers, there is no single study comparing these on a single cipher. Such a study would allow a fair comparison of the various methodologies. In this work, we present an in-depth analysis of the various ways in which TI can be implemented for a lightweight cipher. We chose GIFT for our analysis as it is currently one of the most energy-efficient lightweight ciphers. The experimental results show that different implementation techniques have distinct applications. For example, the 4-shares technique is good for applications demanding high throughput whereas 3-shares is suitable for constrained environments with less area and moderate throughput requirements. The techniques presented in the paper are also applicable to other blockciphers. For security evaluation, we performed TVLA (test vector leakage assessment) on all the design strategies. Experiments using up to 50 million traces show that the designs are protected against first-order attacks.
Arpan Jati, Naina Gupta 0001, Anupam Chattopadhyay, Somitra Kumar Sanadhya, Donghoon Chang
IEEE Trans. Inf. Forensics Secur.2
2019 XMSS and Embedded Systems
Wen Wang 0007, Bernhard Jungk, Julian Wälde, Shuwen Deng, Naina Gupta 0001, Jakub Szefer, Ruben Niederhagen
SAC5