Arpan Jati

dblp:157/0154 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-1082-4049ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Security and privacy · 3 · 1 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AI Attacks AI: Recovering Neural Network Architecture from NVDLA Using AI-Assisted Side Channel Attack
abstract
During the last decade, there has been a stunning progress in the domain of Artificial Intelligence (AI) aided by highly trained Machine Learning (ML) models. Such models are valuable Intellectual Property (IP) and, therefore, have been subjected to various model recovery attacks. In this work, we study the vulnerabilities of commercial, open-source accelerator NVDLA and present the first successful model recovery attack. For this purpose, we used power and timing information from the side-channel leakage of convolutional neural networks (CNN) models to train CNN-based attack models. Utilizing these attack models, we demonstrate that even with a highly pipelined architecture, multiple parallel execution in the accelerator along with Linux OS running tasks in the background, recovery of number of layers, kernel sizes, output neurons and distinguishing different layers, is possible with very high accuracy. This is also the first work to show the impact of differences in hyperparameters on the power traces. Our solution is fully automated, AI-based, and portable to other hardware neural networks, thus presenting a greater threat toward IP protection. Using LeNet as the target victim model, we demonstrate an accuracy of more than 95% in recovering various parameters. This study presents a serious practical threat, in the form of side-channel attack, toward complex commercial architectures. Furthermore, we show that AI-guided attack significantly boosts the attacker capability.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
ACM Trans. Embed. Comput. Syst.2
2024 Formal Verification of Secure Boot Process
abstract
Formal verification is widely used for checking the correctness of a system with respect to the succinct properties at early design phase. In recent times, these methods are also adopted to validate the security of a system by rightly abstracting the security-oriented properties. In this work, we focus on a recently reported attack on the secure boot process of Zynq-7000 platform. We study the vulnerability with formal analysis of its boot loader code. The challenge is to model the system and identify the right set of properties so that, the vulnerability is exposed quickly. It is shown that the UPPAAL model checking analysis can be used, which helps to find the vulnerability instantaneously with the help of four properties and a virtual memory usage peak of about 50MB to test each property. To the best of our knowledge, we perform the first analysis of UPPAAL on a concrete software implementation and perform a case study on a real security vulnerability reported in a CVE [1].
Sriram Vasudevan, Prasanna Ravi, Arpan Jati, Shivam Bhasin, Anupam Chattopadhyay
DATE3
2024 A Configurable CRYSTALS-Kyber Hardware Implementation with Side-Channel Protection
abstract
In this work, we present a configurable and side channel resistant implementation of the post-quantum key-exchange algorithm CRYSTALS-Kyber . The implemented design can be configured for different performance and area requirements leading to different trade-offs for different applications. A low area implementation can be achieved in 5,269 LUTs and 2,422 FFs, whereas a high performance implementation required 7,151 LUTs and 3,730 FFs. Due to a deeply pipelined architecture, a high operating speed of more than 250 MHz could be achieved on 28nm Xilinx FPGAs. The side channel resistance is implemented using a carefully chosen set of novel and known techniques such as Fault Detection Hashes, Instruction Randomization, FSM Protection and so on. resulting in a low overhead of less than 5% while being highly configurable. To the best of our knowledge, this work presents the first side-channel and fault attack protected configurable accelerator for CRYSTALS-Kyber . Using TVLA (test vector leakage assessment), we validate the implemented protection techniques and demonstrate that the design does not leak information even after 200 K traces. Furthermore, one of the configuration choices results in the smallest hardware implementation of CRYSTALS-Kyber known in the literature.
Arpan Jati, Naina Gupta 0001, Anupam Chattopadhyay, Somitra Kumar Sanadhya
ACM Trans. Embed. Comput. Syst.1
2023 CRYSTALS-Dilithium on RISC-V Processor: Lightweight Secure Boot Using Post-Quantum Digital Signature
abstract
With the ongoing efforts for transitioning towards post-quantum security, NIST has recently selected the digital signature algorithm CRYSTALS-Dilithium for standardization. In this work, we demonstrate the first Dilithium based hardware accelerated secure boot architecture developed around Ariane, an open-source RISC- V core. By utilizing a compact design with novel verification engine, a secure boot flow is implemented with only 3.48ms runtime overhead compared to normal boot, while requiring 10.4K LUTs and 5.7K FFs on an FPGA. Compared to the state-of-the-art we achieve a reduction of 3.42× and 7.88 × for LUTs and FFs respectively. Also, the design when realized in 65nm ASIC requires only 125 kGE and 6.3 mW power at 100 MHz. Further, as secure boot is one of the critical processes and the security of the whole system depends on it, we implemented hardware fault countermeasures and evaluated their effectiveness in preventing secure boot bypass.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
ICCAD2
2023 Lightweight Hardware Accelerator for Post-Quantum Digital Signature CRYSTALS-Dilithium
abstract
The looming threat of an adversary with quantum computing capability led to a worldwide research effort towards identifying and standardizing novel post-quantum cryptographic primitives. Post-standardization, all existing security protocols will need to support efficient implementation of these primitives. In this work, we contribute to these efforts by reporting the smallest implementation of CRYSTALS-Dilithium, one of the chosen post-quantum digital signature scheme for NIST standardization process. By invoking multiple optimizations to leverage parallelism, pre-computation and memory access sharing, we obtain an implementation that could be fit into one of the smallest Zynq FPGA. On Zynq Ultrascale+, our design achieves an improvement of about 36.7%/35.4%/42.3% in Area$\times $Time (LUTs$\times \text{s}$) trade-off for KeyGen/Sign/Verify respectively over state-of-the-art implementation. We also evaluate our design as a co-processor on three different hardware platforms and compare the results with software implementation, thus presenting a detailed evaluation of CRYSTALS-Dilithium targeted for embedded applications. Further, on ASIC using TSMC 65nm technology, our design requires 0.227mm2area and can operate at a frequency of 1.176 GHz. As a result, it only requires$53.7\mu \text{s}/96.9\mu \text{s}/57.7\mu \text{s}$for KeyGen/Sign/Verify operation for the best-case scenario.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay, Gautam Jha
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 MemEnc: A Lightweight, Low-Power, and Transparent Memory Encryption Engine for IoT
abstract
Recent advancement in technologies has led to the widespread adoption and deployment of Internet-of-Things devices. Because of the ubiquitous nature of these devices, they process large amounts of personal and sensitive data. These data are typically stored on DRAM chips, and hence becomes an easy target for attackers. Memory encryption is a commonly adopted solution to provide confidentiality. However, realizing a lightweight, low-latency, low-power solution for resource-constrained devices is a challenge. To address this, we designed MemEnc, a purely hardware-based solution that performs encryption on-the-fly and handles memory requests transparently without any OS intervention. MemEnc runs at a maximum frequency of 401 MHz and requires about 23.2 kGE (gate equivalents) on 65-nm ASIC and consumes only 1.9 mW of power at 250 MHz. Using comprehensive benchmarking, we also analyze the applicability of the proposed solution on real-world workloads. Our experiments show that certain real-time applications can run with full memory encryption and still meet system requirements. Moreover, we integrated our memory encryption engine with ARM TrustZone and present comparative results for the different case studies with Intel SGX. We show that static and dynamic efficiency-security tradeoff is necessary for all scenarios.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
IEEE Internet Things J.2
2021 PQC Acceleration Using GPUs: FrodoKEM, NewHope, and Kyber
abstract
In this article, we present the first GPU implementation for FrodoKEM-976, NewHope-1024, and Kyber-1024. These algorithms belong to three different classes of post-quantum algorithms: Learning with errors (LWE), Ring-LWE, and Module-LWE. We show the practical applicability of the algorithms in different scenarios using two different implementation approaches. Moreover, we achieve highly efficient realization of computationally expensive operations such as NTT (Number Theoretic Transform), matrix multiplication, and Keccak. Since, these are the most common operations in lattice-based cryptographic algorithms, the techniques presented in this article will likely benefit other similar algorithms. Using a NVIDIA QUADRO GV100 graphics card, we undertook a detailed experimental study. For NewHope and Kyber we were able to perform approximately 504K and 473K key exchanges per second, demonstrating a speedup of almost 53.1× and 51.05× compared to the reference C implementation. Compared to the optimized AVX2 versions we obtain speedups of 25.7× and 14.6×, respectively. Further, implementation of FrodoKEM resulted in a speedup of 50.6×, 44.2×, and 36.9× for KeyGen, Encaps and Decaps operations. Compared to its AVX2 counterpart, we achieved a speedup of about 7.3×, 4.7× and 4.9×, respectively. We also show that using multiple streams resulted in further speedup of about 28-38 percent.
Naina Gupta 0001, Arpan Jati, Amit Kumar Chauhan, Anupam Chattopadhyay
IEEE Trans. Parallel Distributed Syst.2
2020 Threshold Implementations of <tt>GIFT</tt>: A Trade-Off Analysis
abstract
Threshold Implementation (TI) is one of the most widely used countermeasure for side channel attacks. Over the years several TI techniques have been proposed for randomizing cipher execution using different variations of secret-sharing and implementation techniques. For instance, sharing without decomposition (4-shares) is the most straightforward implementation of the threshold countermeasure. However, its usage is limited due to its high area requirements. On the other hand, sharing using decomposition (3-shares) countermeasure for cubic non-linear functions significantly reduces area and complexity in comparison to 4-shares. Nowadays, security of ciphers using a side channel countermeasure is of utmost importance. This is due to the wide range of security critical applications from smart cards, battery operated IoT devices, to accelerated crypto-processors. Such applications have different requirements (higher speed, energy efficiency, low latency, small area etc.) and hence need different implementation techniques. Although, many TI strategies and implementation techniques are known for different ciphers, there is no single study comparing these on a single cipher. Such a study would allow a fair comparison of the various methodologies. In this work, we present an in-depth analysis of the various ways in which TI can be implemented for a lightweight cipher. We chose GIFT for our analysis as it is currently one of the most energy-efficient lightweight ciphers. The experimental results show that different implementation techniques have distinct applications. For example, the 4-shares technique is good for applications demanding high throughput whereas 3-shares is suitable for constrained environments with less area and moderate throughput requirements. The techniques presented in the paper are also applicable to other blockciphers. For security evaluation, we performed TVLA (test vector leakage assessment) on all the design strategies. Experiments using up to 50 million traces show that the designs are protected against first-order attacks.
Arpan Jati, Naina Gupta 0001, Anupam Chattopadhyay, Somitra Kumar Sanadhya, Donghoon Chang
IEEE Trans. Inf. Forensics Secur.1
2016 SPF: A New Family of Efficient Format-Preserving Encryption Algorithms
Donghoon Chang, Mohona Ghosh, Kishan Chand Gupta, Arpan Jati, Abhishek Kumar 0002, Dukjae Moon, Indranil Ghosh Ray, Somitra Kumar Sanadhya
Inscrypt4
2014 Rig: A Simple, Secure and Flexible Design for Password Hashing
Donghoon Chang, Arpan Jati, Sweta Mishra, Somitra Kumar Sanadhya
Inscrypt2