Debapriya Basu Roy

dblp:116/4686 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0003-4664-5237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 6 first-author · 7 since 2021Security and privacy · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Hardware Acceleration for Zero-Knowledge Proof: Recent Advances and Challenges
Pengzhou He, Debapriya Basu Roy, Jiafeng Xie
VTS2
2025 Three Eyed Raven: An On-Chip Side Channel Analysis Framework for Run-Time Evaluation
abstract
Side-channel attacks exploit the physical leakages from hardware components, such as power consumption, to break secure cryptographic algorithms and retrieve their secret key. Evaluating implementations of cryptographic algorithms against such analysis is crucial but traditional frameworks require expensive external devices like oscilloscopes, making the process expensive and time-consuming. Recent advancements in on-chip sensors offer a cost-effective, fully on-chip SCA framework, eliminating the need for external devices. In this paper, we propose Raven, an on-chip SCA framework with hardware implementations of Test Vector Leakage Assessment (TVLA), Correlation Power Analysis (CPA), and Deep Learningbased Leakage Assessment (DL-LA), for run-time evaluation of cryptographic implementations. RAVEN leverages on-chip sensors to efficiently assess side-channel security, without requiring any external measurement devices or any customized evaluation platform. Our proposed hardware implementations of TVLA, CPA, and DL-LA are lightweight and the entire architecture including the sensors can fit within the lightweight and low-cost AMD-Xilinx PYNQ FPGA platform. The proposed framework is verified on an FPGA implementation of AES-128 and the corresponding result of TVLA, CPA, and DL-LA closely matches with these algorithm's software implementation while requiring significantly less time and storage.
M. Dhilipkumar, Priyanka Bagade, Debapriya Basu Roy
DATE3
2024 Multiplierless Design of High-Speed Very Large Constant Multiplications
abstract
In cryptographic algorithms, the constants to be multiplied by a variable can be very large due to security requirements. Thus, the hardware complexity of such algorithms heavily depends on the design architecture handling large constants. In this paper, we introduce an electronic design automation tool, called LEIGER, which can automatically generate the realizations of very large constant multiplications for low-complexity and high-speed applications, targeting the ASIC design platform. LEIGER can utilize the shift-adds architecture and use 3-input operations, i.e., carry-save adders (CSAs), where the number of CSAs is reduced using a prominent optimization algorithm. It can also generate constant multiplications under a hybrid design architecture, where 2-and 3-input operations are used at different stages. Moreover, it can describe constant multiplications under a design architecture using compressor trees. As a case study, high-speed Montgomery multiplication, which is a fundamental operation in cryptographic algorithms, is designed with its constant multiplication block realized under the proposed architectures. Experimental results indicate that LEIGER enables a designer to explore the trade-off between area and delay of the very large constant and Montgomery multiplications and leads to designs with area-delay product, latency, and energy consumption values significantly better than those obtained by a recently proposed algorithm.
Levent Aksoy, Debapriya Basu Roy, Malik Imran, Samuel Nascimento Pagliarini
ASPDAC2
2024 Automatic Generation of Modular Multipliers Upon Pseudo-Mersenne Primes Using DSP Blocks on FPGAs
abstract
Modular multiplication is crucial and often forms the critical path in the execution of complex cryptographic algorithms such as Elliptic Curve Cryptography (ECC) and Homomorphic Encryption. Execution in ECC primarily requires multiplication and addition in$GF(p)$. On the other hand, homomorphic encryption, a cryptographic technique that allows arithmetic operations over encrypted data, requires addition and multiplication in ring$R_{q}=Z_{q}[x]/(x^{n}+1)$, which in turn requires coefficient multiplication over the integer$q$. Efforts to enhance the efficiency of modular multipliers have been extensively pursued due to the pivotal role they play in the critical path of various cryptographic operations. In recent times, FPGA boards containing DSP blocks which are capable of efficiently performing various mathematical operations stood out to be a desirable platform to design modular multipliers. But designing multipliers over FPG As stands tricky as the operand sizes supported by the DSPs are asymmetric in nature. In this paper, we propose an automated multiplier design technique that facilitates the design of modular multipliers upon pseudo-Mersenne primes. Our proposed automated multiplier design tool generates efficient hardware architectures of modular multipliers with pseudo-Mersenne primes that can be directly used in RNS (residue number system) and cryptographic algorithms like ECC in GF(p). Using the tool, we generated modular multiplier designs for various primes supporting several elliptic curves and for several RNS systems supporting a large prime. Our tool-generated hardware implementations' resource requirements regarding critical path delay and resource utilization are significantly more efficient than those of existing implementations.
Shree Harish S, Debapriya Basu Roy
DSD2
2024 Design of a Lightweight Fast Fourier Transformation for FALCON using Hardware-Software Co-Design
abstract
Lattice-based post-quantum cryptographic algorithm FALCON needs to execute the time-critical Fast-Fourier Transformation (FFT). Existing works in the literature have explored hardware for FFT of FALCON using Cooley-Tukey. In this work, we have designed an efficient hardware-software co-design of FFT for FALCON using Winograd’s FFT method. Winograd’s FFT is a widely adopted technique for FFT and reduces the multiplication counts for higher-radix FFT than the Cooley-Tukey, with a penalty of some extra addition/subtraction. Our Winograd radix-8 framework for FFT outperforms the traditional Cooley-Tukey method. Moreover,our proposed architecture is flexible in adopting different instruction sets and can also be configured for any type of FFT method with specific instruction sets.
Suraj Mandal, Debapriya Basu Roy
ACM Great Lakes Symposium on VLSI2
2024 Winograd for NTT: A Case Study on Higher-Radix and Low-Latency Implementation of NTT for Post Quantum Cryptography on FPGA
abstract
Number Theoretic Transform (NTT) plays an important role in efficiently implementing lattice-based cryptographic algorithms like CRYSTALS-Kyber, Dilithium, and FALCON. Existing implementations of NTT for these algorithms are mostly based on radix-2 or radix-4 realization of Cooley-Tukey and Gentleman-Sande architectures. In this work, we explore an alternative method of performing NTT known as Winograd’s NTT that requires fewer number of modular multipliers than the conventional Coole-Tukey/Gentleman-Sande for higher radix NTT. We have proposed three different low-latency implementations of Winograd’s NTT, applicable to CRYSTALS-Dilithium, FALCON, and CRYSTALS-Kyber, respectively. Our first implementation of Winograd NTT focuses on radix-16 NTT multiplication unit for polynomials of length 256 and can be directly used for CRYSTALS-Dilithium. The NTT of CRYSTALS-Dilithium is also benefited from our proposed K-RED modular multiplication. Our radix-16-based Winograd outperforms existing Cooley-Tukey/Gentleman-Sande based NTT multipliers of CRYSTALS-Dilithium. Our second implementation of NTT is based on radix-8 Winograd structure with a novel modular multiplication method that targets polynomials of length 512 and can be directly applied for FALCON. For CRYSTALS-Kyber, we have designed a radix-16 Winograd Butterfly Unit (BFU) that can be configured as two parallel radix-8 Winograd BFUs during mixed-radix computation. To the best of our knowledge, this is the first work that applied the Winograd technique for NTT multiplication for post-quantum secure lattice-based cryptographic algorithms.
Suraj Mandal, Debapriya Basu Roy
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 FlexiPair: An Automated Programmable Framework for Pairing Cryptosystems
abstract
Pairing cryptosystems are extremely powerful mathematical tools for developing cryptographic protocols that can provide end-to-end security for applications like Internet-of-Things (IoT), cloud services and cyber-physical systems (CPS). However, these applications require the implementations to be light-weight but still real-time, with the additional feature of being flexible. The flexibility can come from different choices of underlying algorithms along with suitable parameter choices. A software implementation offers better flexibility but lacks in timing performance, whereas custom hardware delivers better performance but has poor flexibility. Furthermore, the designs over small characteristic curves are now insecure against recent attacks. Existing designs do not address the drawback of less flexibility and huge resource consumption collectively. In this article, we present a micro-program controlled hardware design which has the least resource consumption among the similar existing designs on FPGA that offer such programmability and flexibility. This redundant number arithmetic-based architecture consumes only 2506 slices on Xilinx Virtex-7 FPGA. It can be migrated to other device families or updated for different algorithms without data-path or control-path modification. To enhance the flexibility, we developed a custom assembly-like finite state machine (FSM) description, called Prism, and necessary tool to generate the micro-program states. To illustrate the functionality of Prism, we present designs for Tate and Optimal-Ate pairing with the micro-program states generated using this tool.
Arnab Bag, Debapriya Basu Roy, Sikhar Patranabis, Debdeep Mukhopadhyay
IEEE Trans. Computers2
2020 Fault Template Attacks on Block Ciphers Exploiting Fault Propagation
Sayandeep Saha, Arnab Bag, Debapriya Basu Roy, Sikhar Patranabis, Debdeep Mukhopadhyay
EUROCRYPT (1)3
2020 Efficient Hardware/Software Co-Design for Post-Quantum Crypto Algorithm SIKE on ARM and RISC-V based Microcontrollers
abstract
Post-quantum cryptography has emerged as a very attractive research topic due to the recent advancements in the development of quantum computers. Among the different available post-quantum public-key algorithms, Supersingular Isogeny Key-Encapsulation (SIKE) has posed a unique design challenge due to its resource intensive arithmetic but is characterized by small key sizes. Existing implementations of SIKE either focus on dedicated accelerators on FPGA platforms or on assembly optimized software implementations on ARM. A full FPGA implementation, though offering low latency and high performance, suffers from the disadvantage of having a large area footprint and a low flexibility. On the other hand, a pure software implementation has lower performance compared to FPGA implementations. In this paper, we propose hardware/software co-design methodologies for SIKE and integrate a redundant number based finite field accelerator into two microcontroller platforms based on ARM and RISC-V. The result shows that our implementation on ARM Cortex-A9 enhanced with a field accelerator offers significant speedup in terms of clock cycles when compared to standalone software implementations on ARM32 and ARM64. Moreover, to show how the communication overhead between processor and accelerator can be mitigated, we integrated the finite field accelerator directly into the core of a RISC-V processor. To the best of our knowledge, this is the first design that applies hardware/software co-design methodologies to implement SIKE on ARM and RISC-V platforms. Our proposed design requires 65500 K clock cycles to execute SIKEp434 on an ARM Cortex-A9 processor. On RISC-V, our proposed design requires only 36900 K clock cycles.
Debapriya Basu Roy, Tim Fritzmann, Georg Sigl
ICCAD1
2020 A Minimalistic Perspective on Koblitz Curve Scalar Multiplication for FPGA Platforms
abstract
Koblitz Curves offer excellent optimization opportunities for characteristic-2 Elliptic Curve Cryptosystems (ECC). However, porting such choices onto a lightweight and cost-effective FPGA platform is a major challenge. The underlying characteristic-2 algebra is not aligned with the on-chip components like DSP multipliers, which if not utilized leads to large LUT counts of the designs. In this work, we develop several techniques to propose a minimal instruction set, centered around ADDN (Add and Branch if less than zero) based OISC (One-Instruction-Set-Computing) coupled with a lightweight comba characteristic-2 finite field multiplier to ensure the aggressive utilization of the FPGA resources. The paper develops an ADDN based scalar multiplier for Koblitz curves which has a minimal footprint on the LUTs, leveraging methods for utilizing the underlying DSP blocks in FPGAs for performing GF(2) operations, in conjunction with proposed reduced variants of the ADDN instruction to perform the computations. The architecture uses FPGA resources like BRAMs, DSPs effectively to have a minimal requirement for FPGA LUTs, thus leaving room for other peripheral designs to be hosted in a single FPGA. The same design can be configured to realize scalar multiplications by two popular strategies, namely due to Solinas and Montgomery, to demonstrate the modular nature of the design. The minimalistic design style leads to a resource-constrained architecture performing scalar multiplication in less than 1000 slices, needing less than 0.5 ms on a cost-effective Artix-7 FPGA. The design has been compared with reported literature to highlight the area efficiency of the design, however ensuring a competitive AT (area-time) product.
Siddhartha Chowdhury, Debapriya Basu Roy, Debdeep Mukhopadhyay
VLSI-SOC2
2020 Neural Network-based Inherently Fault-tolerant Hardware Cryptographic Primitives without Explicit Redundancy Checks
abstract
Fault injection-based cryptanalysis is one of the most powerful practical threats to modern cryptographic primitives. Popular countermeasures to such fault-based attacks generally use some form of redundant computation to detect and react/correct the injected faults. However, such countermeasures are shown to be vulnerable to selective fault injections. In this article, we aim to develop a cryptographic primitive that is fault tolerant by its construction and does not require to compute the same value multiple times. We utilize the effectiveness of Neural Networks (NNs), which show “some degree” of robustness by functioning correctly even after the occurrence of faults in any of its parameters. We also propose a novel strategy that enhances the fault tolerance of the implementation to “high degree” (close to 100%) by incorporating selective constraints in the NN parameters during the training phase. We evaluated the performance of revised NN considering both software and FPGA implementations for standard cryptographic primitives like 8×8 AES SBox and 4×4 PRESENT SBox. The results show that the fault tolerance of such implementations can be significantly increased with the proposed methodology. Such NN-based cryptographic primitives will provide inherent resistance against fault injections without requiring any redundancy countermeasures.
Manaar Alam, Arnab Bag, Debapriya Basu Roy, Dirmanto Jap, Jakub Breier, Shivam Bhasin, Debdeep Mukhopadhyay
ACM J. Emerg. Technol. Comput. Syst.3
2020 A Framework to Counter Statistical Ineffective Fault Analysis of Block Ciphers Using Domain Transformation and Error Correction
abstract
Right from its introduction, fault attacks (FA) have been established to be one of the most practical threats to both public key and symmetric key based cryptosystems. Statistical Ineffective Fault Analysis (SIFA) is a recently proposed class of fault attacks introduced at CHES 2018. The fascinating feature of this attack is that it exploits the correct ciphertexts obtained during a fault injection campaign, instead of the faulty ciphertexts. SIFA has been shown to bypass almost all of the existing fault attack countermeasures even when they are combined with masking schemes for side-channel resistance. The goal of this work is to propose a countermeasure framework for SIFA. It has been observed that a randomized domain transformation of the intermediate computation combined with bit-level error correction can prevent SIFA attacks. The domain transformation (Transform) can be realized by standard masking schemes. In fact, we prove that if biased faults are injected at the state register of a block cipher at a certain target round, then masking is sufficient for SIFA protection, until all the shares for a specific bit are corrupted. However, masking alone cannot prevent SIFA if the faults are injected at certain specific locations inside the S-Boxes. To address this issue, we incorporate a bit-level error-correction mechanism (Encode). An instantiation of this Transform-and-Encode (TaE) framework, called AntiSIFA, has been proposed and realized for the block cipher PRESENT as a proof-of-concept. Practical evaluation of the countermeasure implementation in both hardware and software ensures our theoretical claims regarding SIFA security, as well as protection against Side-Channel-Attacks (SCA).
Sayandeep Saha, Dirmanto Jap, Debapriya Basu Roy, Avik Chakraborty, Shivam Bhasin, Debdeep Mukhopadhyay
IEEE Trans. Inf. Forensics Secur.3
2019 Count Your Toggles: a New Leakage Model for Pre-Silicon Power Analysis of Crypto Designs
Rajat Sadhukhan, Paulson Mathew, Debapriya Basu Roy, Debdeep Mukhopadhyay
J. Electron. Test.3
2019 CC Meets FIPS: A Hybrid Test Methodology for First Order Side Channel Analysis
abstract
Common Criteria (CC) and FIPS 140-3 are two popular side channel testing methodologies. Test Vector Leakage Assessment Methodology (TVLA), a potential candidate for FIPS, can detect the presence of side-channel information in leakage measurements. However, TVLA results cannot be used to quantify side-channel vulnerability and it is an open problem to derive its relationship with side channel attack success rate (SR), i.e., a common metric for CC. In this paper, we extend the TVLA testing beyond its current scope. Precisely, we derive a concrete relationship between TVLA and signal to noise ratio (SNR). The linking of the two metrics allows direct computation of success rate (SR) from TVLA for given choice of intermediate variable and leakage model and thus unify these popular side channel detection and evaluation metrics. An end-to-end methodology is proposed, which can be easily automated, to derive attack SR starting from TVLA testing. The methodology works under both univariate and multivariate setting and is capable of quantifying any first order leakage. Detailed experiments have been provided using both simulated traces and real traces on SAKURA-GW platform. Additionally, the proposed methodology is benchmarked against previously published attacks on DPA contest v4.0 traces, followed by extension to jitter based countermeasure. The result shows that the proposed methodology provides a quick estimate of SR without performing actual attacks, thus bridging the gap between CC and FIPS.
Debapriya Basu Roy, Shivam Bhasin, Sylvain Guilley, Annelie Heuser, Sikhar Patranabis, Debdeep Mukhopadhyay
IEEE Trans. Computers1
2019 Combining PUF with RLUTs: A Two-party Pay-per-device IP Licensing Scheme on FPGAs
abstract
With the popularity of modern FPGAs, the business of FPGA specific intellectual properties (IP) is expanding rapidly. This also brings in the concern of IP protection. FPGA vendors are making serious efforts toward IP protection, leading to standardization schemes like IEEE P1735. However, efficient techniques to prevent unauthorized overuse of IP still remain an open question. In this article, we propose a two-party IP protection scheme combining the re-configurable look-up table primitive of modern FPGAs with physically unclonable functions (PUF). The proposed scheme works with the assumption that the FPGA vendor provides the assurance of confidentiality and integrity of the developed IP. The proposed scheme is considerably lightweight compared to existing schemes, prevents overuse, and does not involve FPGA vendors or trusted third parties for IP licensing. The validation of the proposed scheme is done on MCNC’91 benchmark and third-party IPs like AES and lightweight MIPS processors.
Debapriya Basu Roy, Shivam Bhasin, Ivica Nikolic, Debdeep Mukhopadhyay
ACM Trans. Embed. Comput. Syst.1
2019 High-Speed Implementation of ECC Scalar Multiplication in GF(p) for Generic Montgomery Curves
abstract
Elliptic curve-based cryptography (ECC) has become the automatic choice for public key cryptography due to its lightweightness compared to Rivest-Shamir-Adleman (RSA). The most important operation in ECC is elliptic curve scalar multiplication, and its efficient implementation has gathered significant attention in the research community. Fast implementation of ECC scalar multiplication is often desired for speed-critical applications such as runtime authentication in automated cars, web server certification, and so on. Such fast architectures are achieved by implementing ECC scalar multiplication in fields with pseudo-Mersenne prime or Solinas prime. In this paper, we aim to implement a fast implementation of ECC scalar multiplication for any generic Montgomery curve in Galois Field in p [GF(p)] without having the constraint of using any specialized modulus. We will show that the proposed ECC scalar multiplication architecture is as fast as scalar multiplication in special curves like Curve25519, albeit with little area overhead. The proposed architecture can be modified to support ECC scalar multiplication in both Montgomery and short Weierstrass curves.
Debapriya Basu Roy, Debdeep Mukhopadhyay
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Revisiting FPGA Implementation of Montgomery Multiplier in Redundant Number System for Efficient ECC Application in GF(p)
abstract
The fast implementations of ECC in GF(p) are generally implemented using specialized prime field, and henceforth, they are dependent on the structure of the prime. But, these implementations cannot be ported to generic curves which do not support such prime structures. Such generic curves are often used in various crypto-applications like pairing and post-quantum secure supersingular isogeny based key exchange. In those cases, modular multiplication is executed through Montgomery multiplier which is slower compared to modular multiplication using specialized primes. This work aims to reduce the speed gap between Montgomery multiplication and modular multiplication in specialized prime field by presenting an efficient implementation of Montgomery multiplier on FPGA using the redundant number system.
Debdeep Mukhopadhyay, Debapriya Basu Roy
FPL2
2017 Side Channel Evaluation of PUF-Based Pseudorandom Permutation
abstract
PUF-PRFs are Pseudorandom Functions (PRFs) constructed using Physically Unclonable Functions (PUFs) as a hardware building block to provide the random input-output mapping. Since PUF-PRFs inherit all the principal properties of PUFs such as memory-leakage resilience, unclonablity, tampering-resistance, pseudo-randomness, and provable security, PUF-PRFs hold great promise as an extremely useful cryptographic hardware primitive. In this paper, we evaluate the security of PUF-PRFs against Side Channel Attacks. Two different attacks based on analysis of power side channel are developed, and demonstrated through the experiments on Xilinx FPGAs. In addition, we reduce the complexity of Correlation Power Analysis (CPA) to recover n-bit secret, from O(2 2n) to O(3n 2n). Based on our experimental results, we conclude that the security of PUF-PRFs, when subjected to side channel attacks, depends on not only the security of the used PUFs, but also the PUF-PRF architecture.
Durga Prasad Sahoo, Phuong Ha Nguyen, Debapriya Basu Roy, Debdeep Mukhopadhyay, Rajat Subhra Chakraborty
DSD3
2016 Shuffling across rounds: A lightweight strategy to counter side-channel attacks
abstract
Side-channel attacks are a potent threat to the security of devices implementing cryptographic algorithms. Designing lightweight countermeasures against side-channel analysis that can run on resource constrained devices is a major challenge. One such lightweight countermeasure is shuffling, in which the designer randomly permutes the order of execution of potentially vulnerable operations. State of the art shuffling countermeasures advocate shuffling a set of independent operations in a single round of a cryptographic algorithm, but are often found to be insufficient as standalone countermeasures. In this paper, we propose a two-round version of the shuffling countermeasure, and test its security when applied to a serialized implementation of AES-128 using Test Vector Leakage Assessment (TVLA). Our results show that the required number of traces to break AES-128 implemented using our proposed countermeasure is significantly larger than the implementations using simple one-round shuffling. Furthermore, the new shuffling method has significantly lower overhead of around 1.3 times, as compared to other side-channel countermeasures such as masking that have an overhead of approximately two times.
Sikhar Patranabis, Debapriya Basu Roy, Praveen Kumar Vadnala, Debdeep Mukhopadhyay, Santosh Ghosh
ICCD2
2016 SmashClean: A hardware level mitigation to stack smashing attacks in OpenRISC
abstract
Buffer overflow and stack smashing have been one of the most popular software based vulnerabilities in literature. There have been multiple works which have used these vulnerabilities to induce powerful attacks to trigger malicious code snippets or to achieve privilege escalation. In this work, we attempt to implement hardware level security enforcement to mitigate such attacks on OpenRISC architecture. We have analyzed the given exploits [5] in detail and have identified two major vulnerabilities in the exploit codes: memory corruption by non-secure memcpy() and return address modification by buffer overflow. We have individually addressed each of these exploits and have proposed a combination of compiler and hardware level modification to prevent them. The advantage of having hardware level protection against these attacks provides reliable security against the popular software level countermeasures.
Manaar Alam, Debapriya Basu Roy, Sarani Bhattacharya, Vidya Govindan, Rajat Subhra Chakraborty, Debdeep Mukhopadhyay
MEMOCODE2
2015 Integrated Sensor: A Backdoor for Hardware Trojan Insertions?
abstract
Embedded system face a serious threat from physical attacks when applied in critical applications. Therefore, modern systems have several integrated sensors to detect potential threats. In this paper, we put forward a new issue where these sensors can open other security loopholes. We demonstrate that sensors, which are deployed to prevent faults, can be exploited to insert effective and almost zero-overhead hardware Trojans. Two case studies are presented on Xilinx Virtex-5 FPGA. The first case study exploits the in-build temperature sensor of Virtex-5 system monitors while the other exploits a user deployed sensor. Both the sensor can be used to trigger a powerful Trojan with minimal and at times zero overhead.
Xuan Thuy Ngo, Zakaria Najm, Shivam Bhasin, Debapriya Basu Roy, Jean-Luc Danger, Sylvain Guilley
DSD4
2015 From theory to practice of private circuit: A cautionary note
abstract
Private circuits, from their publication, have been really popular among the researchers. They also form the basis for provable masking schemes. There are several works which try to improve the results of bit-level private circuits based on 2-input gates for the combinational logic. However, strangely, no practical side-channel analysis of private circuits has been presented so far, which is the focus of the present paper. In this paper, we have tried to identify the `ambush' or hidden dangers in the implementation of private circuits, which can compromise its security in practical scenarios. We have implemented block cipher SIMON with private circuit and have performed side-channel analysis on it. The result shows that, in practice, there is significant amount of information leakage which can be exploited by adversaries. Some leakage comes from practical optimization applied by standard CAD tools, if they restructure the netlists. But even with immutable netlists, we identify leakage caused by a kind of glitch known as early evaluation. Lastly, we demonstrate how to translate theoretically secure private circuit to practically secure private circuit with added overhead, by clocking every combinational gate. Leakage detection tests are applied to attest the security of considered variants of private circuits.
Debapriya Basu Roy, Shivam Bhasin, Sylvain Guilley, Jean-Luc Danger, Debdeep Mukhopadhyay
ICCD1
2015 ECC on Your Fingertips: A Single Instruction Approach for Lightweight ECC Design in GF(p)
Debapriya Basu Roy, Poulami Das 0003, Debdeep Mukhopadhyay
SAC1
2014 Tile Before Multiplication: An Efficient Strategy to Optimize DSP Multiplier for Accelerating Prime Field ECC for NIST Curves
abstract
High speed DSP blocks present in the modern FPGAs can be used to implement prime field multiplication to accelerate Elliptic Curve scalar multiplication in prime fields. However, compared to logic slices, DSP blocks are scarce resources, hence its usage needs to be optimized. The asymmetric 25 × 18 signed multipliers in FPGAs open a new paradigm for multiplier design, where operand decomposition becomes equivalent to a tiling problem. Previous literature has reported that for asymmetric multiplier, it is possible to generate a tiling (known as non-standard tiling) which requires less number of DSP blocks compared to standard tiling, generated by school book algorithm. In this paper, we propose a generic technique for such tiling generation and generate this tiling for field multiplication in NIST specified curves. We compare our technique with standard school book algorithm to highlight the improvement. The acceleration in ECC scalar multiplication due to the optimized field multiplier is experimentally validated for P-256. The impact of this accelerated scalar multiplication is shown for the key encapsulation algorithm PSEC-KEM (Provably Secure Key Encapsulation Mechanism).
Debapriya Basu Roy, Debdeep Mukhopadhyay, Masami Izumi, Junko Takahashi
DAC1
2013 Role of power grid in side channel attack and power-grid-aware secure design
abstract
Side-channel attack (SCA) is a method in which an attacker aims at extracting secret information from crypto chips by analyzing physical parameters (e.g. power). SCA has emerged as a serious threat to many mathematically unbreakable cryptography systems. From an attacker's point of view, the difficulty of mounting SCA largely depends on Signal-to-Noise Ratio (SNR) of the side-channel information. It has been shown that SNR primarily depends on algorithmic and circuit-level implementation, measurement noise, as well as device thermal noise. However, to the best of our knowledge, there has not been any study on the effect of power delivery network (PDN) on SCA resistance. We note that the PDN plays a significant role in SNR of measured supply current. Furthermore, SCA resistance strongly depends on the operating frequency due to RLC structure of a power grid. In this paper, we analyze the effect of power grid on SCA and provide quantitative results to demonstrate the frequency-dependent SCA resistance due to PDN-induced noise. This property can potentially be exploited by an attacker to facilitate the attack by operating a device at favorable frequency points. On the other hand, from a designer's perspective, one can explore countermeasures to secure the device at all operating frequencies while minimizing the design overhead. Based on this observation, we propose a frequency-dependent noise-injection based compensation technique to efficiently protect against SCA. Simulation results using realistic PDN model as well as experimental measurements using FPGA test board validate the observations on role of PDN in SCA and the efficacy of the proposed compensation approach.
Xinmu Wang, Wen Yueh, Debapriya Basu Roy, Seetharam Narasimhan, Yu Zheng 0011, Saibal Mukhopadhyay, Debdeep Mukhopadhyay, Swarup Bhunia
DAC3