Rashmi S. Agrawal 0001

dblp:202/9483 · also Rashmi Agrawal 0001 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
9since 2021 · last 2025
0000-0002-9461-0170ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 10 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FIDESlib: A Fully-Fledged Open-Source FHE Library for Efficient CKKS on GPUs
abstract
Word-wise Fully Homomorphic Encryption (FHE) schemes, such as CKKS, are gaining significant traction due to their ability to provide post-quantum-resistant, privacypreserving approximate computing-an especially desirable feature in the Machine-Learning-as-a-Service (MLaaS) paradigm. In this work, we introduce FIDESlib, the first open-source server-side CKKS GPU library that is fully interoperable with well-established client-side OpenFHE operations. Unlike other existing open-source GPU libraries, FIDESlib provides the first implementation featuring heavily optimized GPU kernels for all CKKS primitives, including bootstrapping. Our library also integrates robust benchmarking and testing, ensuring it remains adaptable to further optimization. Comparing our scheme against Phantom (the previously top open-source CKK library, we show that FIDESlib offers superior performance and scalability. For bootstrapping, FIDESlib achieves no less than$70 \times$speedup over the AVX-optimized OpenFHE implementation. FIDESlib is available on Github11https://github.com/CAPS-UMU/FIDESlib.
Carlos Agulló-Domingo, Óscar Vera-López, Seyda Nur Güzelhan, Lohit Daksha, Aymane El Jerari, Kaustubh Shivdikar, Rashmi S. Agrawal 0001, David R. Kaeli, Ajay Joshi, José L. Abellán
ISPASS7
2024 HEAP: A Fully Homomorphic Encryption Accelerator with Parallelized Bootstrapping
abstract
Fully homomorphic encryption (FHE) is a cryptographic technology with the potential to revolutionize data privacy by enabling computation on encrypted data. Lately, the CKKS FHE scheme has become quite popular because it can process real numbers. However, CKKS computing is not pervasive yet because it is resource-intensive both in terms of compute and memory, and is multiple orders of magnitude slower than computing on unencrypted data. The recent algorithmic and hardware optimizations to accelerate CKKS computing are promising, but CKKS computing continues to underperform due to an expensive operation known as bootstrapping. While there have been several efforts to accelerate bootstrapping, it continues to remain the main performance bottleneck. One of the reasons for this performance bottleneck is that unlike the non-bootstrapping parts of CKKS computing the bootstrapping algorithm is inherently sequential and exhibits interdependencies among the data. To address this challenge, in this paper, we introduce HEAP an accelerator that uses a hybrid scheme-switching approach. HEAP uses the CKKS scheme for the non-bootstrapping steps, but switches to the TFHE scheme when performing the bootstrapping step of the CKKS scheme. The hybrid approach transitions to the TFHE scheme by extracting coefficients from a single RLWE ciphertext to represent multiple LWE ciphertexts. We incorporate the bootstrapping function into the TFHE BlindRotate operation and simultaneously apply the BlindRotate operation to all LWE ciphertexts. A parallelized execution of bootstrapping is then feasible because there are no data dependencies between distinct LWE ciphertexts. With our approach, we require smaller-sized bootstrapping keys leading to about $18 \times$ less amount of data to be read from the main memory for the keys. In addition, we introduce a variety of hardware optimizations in HEAP—from modular arithmetic level to NTT and BlindRotate datapath optimizations. The approach in HEAP is agnostic of the hardware and can be mapped to any system with multiple compute nodes. To evaluate HEAP, we implemented it in RTL and mapped it to a single FPGA system and an eight-FPGA system. Our comprehensive evaluation of HEAP for the bootstrapping operation shows a $15.39 \times$ improvement when compared to FAB. Similarly, evaluation of HEAP for the logistic regression model training shows $14.71 \times$ and $11.57 \times$ improvement when compared to FAB and FAB-2 implementations, respectively.
Rashmi S. Agrawal 0001, Anantha P. Chandrakasan, Ajay Joshi
ISCA1
2023 FAB: An FPGA-based Accelerator for Bootstrappable Fully Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) offers protection to private data on third-party cloud servers by allowing computations on the data in encrypted form. To support general-purpose encrypted computations, all existing FHE schemes require an expensive operation known as "bootstrapping". Unfortunately, the computation cost and the memory bandwidth required for bootstrapping add significant overhead to FHE-based computations, limiting the practical use of FHE.In this work, we propose FAB, an FPGA-based accelerator for bootstrappable FHE. Prior FPGA-based FHE accelerators have proposed hardware acceleration of basic FHE primitives for impractical parameter sets without support for bootstrapping. FAB, for the first time ever, accelerates bootstrapping (along with basic FHE primitives) on an FPGA for a secure and practical parameter set. The key contribution of this work is the architecture of a balanced FAB design, which is not memory bound. In our design, we leverage recent algorithms for bootstrapping while being cognizant of the compute and memory constraints of our FPGA. In addition, we use a minimal number of functional units for computing, operate at a low frequency, leverage high data rates to and from main memory, utilize the limited on-chip memory effectively, and perform careful operation scheduling.We evaluate FAB using a single Xilinx Alveo U280 FPGA and by scaling it to a multi-FPGA system consisting of eight such FPGAs. For bootstrapping a fully-packed ciphertext, while operating at 300MHz, FAB outperforms existing state-of-the-art CPU and GPU implementations by 213× and 1.5× respectively. Our target FHE application is training a logistic regression model over encrypted data. For logistic regression model training scaled to 8 FPGAs on the cloud, FAB outperforms a CPU and GPU by 456× and 9.5× respectively, providing practical performance at a fraction of the ASIC design cost.
Rashmi S. Agrawal 0001, Leo de Castro, Guowei Yang 0005, Chiraag Juvekar, Rabia Tugce Yazicigil, Anantha P. Chandrakasan, Vinod Vaikuntanathan, Ajay Joshi
HPCA1
2023 MAD: Memory-Aware Design Techniques for Accelerating Fully Homomorphic Encryption
abstract
Cloud computing has made it easier for individuals and companies to get access to large compute and memory resources. However, it has also raised privacy concerns about the data that users share with the remote cloud servers. Fully homomorphic encryption (FHE) offers a solution to this problem by enabling computations over encrypted data. Unfortunately, all known constructions of FHE require a noise term for security, and this noise grows during computation. To perform unlimited computations on the encrypted data, we need to perform a periodic noise reduction step known as bootstrapping. This bootstrapping operation is memory-bound as it requires several GBs of data. This leads to orders of magnitude increase in the time required for operating on encrypted data as compared to unencrypted data.
Rashmi S. Agrawal 0001, Leo de Castro, Chiraag Juvekar, Anantha P. Chandrakasan, Vinod Vaikuntanathan, Ajay Joshi
MICRO1
2023 GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) enables the processing of encrypted data without decrypting it. FHE has garnered significant attention over the past decade as it supports secure outsourcing of data processing to remote cloud services. Despite its promise of strong data privacy and security guarantees, FHE introduces a slowdown of up to five orders of magnitude as compared to the same computation using plaintext data. This overhead is presently a major barrier to the commercial adoption of FHE.
Kaustubh Shivdikar, Yuhui Bao, Rashmi S. Agrawal 0001, Michael Tian Shen, Gilbert Jonatan, Evelio Mora, Alexander Ingare, Neal Livesay, José L. Abellán, John Kim 0001, Ajay Joshi, David R. Kaeli
MICRO3
2023 RISE: RISC-V SoC for En/Decryption Acceleration on the Edge for Homomorphic Encryption
abstract
Today, edge devices commonly connect to the cloud to use its storage and computing capabilities. This leads to security and privacy concerns about user data. Homomorphic encryption (HE) is a promising solution to address the data privacy problem as it allows arbitrarily complex computations on encrypted data without ever needing to decrypt it. While there has been a lot of work on accelerating HE computations in the cloud, small attention has been paid to the message-to-ciphertext and ciphertext-to-message conversion operations on the edge. In this work, we profile the edge-side conversion operations, and our analysis shows that during conversion error sampling, encryption and decryption operations are the bottlenecks. To overcome these bottlenecks, we present RISE, an area and energy-efficient RISC-V system-on-chip (SoC). RISE leverages an efficient and lightweight pseudorandom number generator (PRNG) core and combines it with fast sampling techniques to accelerate the error sampling operations. To accelerate the encryption and decryption operations, RISE uses scalable data-level parallelism to implement the number theoretic transform (NTT) operation, the main bottleneck within the encryption and decryption operations. In addition, RISE saves area by implementing a unified en/decryption datapath, and efficiently exploits techniques like memory reuse and data reordering to utilize a minimal amount of ON-chip memory. We evaluate RISE using a complete RTL design containing a RISC-V processor interfaced with our accelerator. Our analysis reveals that for message-to-ciphertext and ciphertext-to-message conversions, using RISE leads up to$5986.99\times $and$1164.1\times $more energy-efficient solution, respectively, than when using just the RISC-V processor.
Zahra Azad, Guowei Yang 0005, Rashmi S. Agrawal 0001, Daniel Ruelas-Petrisko, Michael B. Taylor, Ajay Joshi
IEEE Trans. Very Large Scale Integr. Syst.3
2022 Efficient FPGA-based ECDSA Verification Engine for Permissioned Blockchains
abstract
Permissioned blockchain platforms heavily depend on cryptography to provide a layer of trust within the blockchain network, thus verification of cryptographic signatures often becomes the bottleneck. ECDSA is the most commonly used cryptographic scheme in permissioned blockchains. In this work, we propose an efficient implementation of ECDSA signature verification on FPGA, in order to improve performance of permissioned blockchains that aim to use FPGA-based hardware accelerators. We propose several optimizations for modular arithmetic (e.g., custom multipliers and fast modular reduction) and point arithmetic (e.g., reduced number of point double and addition operations, and optimal width NAF representation). Based on these optimized modular and point arithmetic modules, we propose an ECDSA verification engine that can be used by any application for fast verification of signatures. We further optimize our ECDSA verification engine for Hyperledger Fabric (one of the most widely used blockchain platforms) by moving carefully selected operations to a precomputation block, thus simplifying the critical path of ECDSA signature verification. Our ECDSA engine running at 250MHz on Xilinx Alveo U250 accelerator card can perform a verification in$760\mu s$with a throughput of 1, 315 verifications/sec, which is$\sim 2.5$times faster than state-of-the-art FPGA-based implementations. Our Hyperledger Fabric-specific ECDSA engine can perform a verification in$368\mu s$, and 2, 717 verifications/sec. With 10 engines, Hyperledger Fabric can achieve a throughput of 7, 520 transactions/sec.
Rashmi S. Agrawal 0001, Haris Javaid
ASAP1
2022 Efficient FPGA-based ECDSA Verification Engine For Permissioned Blockchains
abstract
Permissioned blockchain platforms heavily depend on cryptography to provide a layer of trust within the blockchain network, thus verification of cryptographic signatures often becomes the bottleneck. ECDSA is the most commonly used cryptographic scheme in permissioned blockchains. In this work, we propose an efficient implementation of ECDSA signature verification on FPGA, in order to improve performance of permissioned blockchains that aim to use FPGA-based hardware accelerators. We propose several optimizations for modular arithmetic (e.g., custom multipliers and fast modular reduction) and point arithmetic (e.g., reduced number of point double and addition operations, and optimal width NAF representation). Based on these optimized modular and point arithmetic modules, we propose an ECDSA verification engine that can be used by any application for fast verification of signatures. We further optimize our ECDSA verification engine for Hyperledger Fabric (one of the most widely used blockchain platforms) by moving carefully selected operations to a precomputation block, thus simplifying the critical path of ECDSA signature verification. Our ECDSA engine running at 250MHz on Xilinx Alveo U250 accelerator card can perform a verification in 760 microseconds with a throughput of 1,315 verifications/second, which is ~2.5 times faster than state-of-the-art FPGA-based implementations. Our Hyperledger Fabric-specific ECDSA engine can perform a verification in 368 microseconds, and 2,717 verifications/second. With 10 engines, Hyperledger Fabric can achieve a throughput of 7,520 transactions/second.
Rashmi S. Agrawal 0001, Haris Javaid
FPGA1
2022 RACE: RISC-V SoC for En/decryption Acceleration on the Edge for Homomorphic Computation
abstract
As more and more edge devices connect to the cloud to use its storage and compute capabilities, they bring in security and data privacy concerns. Homomorphic Encryption (HE) is a promising solution to maintain data privacy by enabling computations on the encrypted user data in the cloud. While there has been a lot of work on accelerating HE computation in the cloud, little attention has been paid to optimize the en/decryption on the edge. Therefore, in this paper, we present RACE, a custom-designed area- and energy-efficient SoC for en/decryption of data for HE. Owing to similar operations in en/decryption, RACE unifies the en/decryption datapath to save area. RACE efficiently exploits techniques like memory reuse and data reordering to utilize minimal amount of on-chip memory. We evaluate RACE using a complete RTL design containing a RISC-V processor and our unified accelerator. Our analysis shows that, for the end-to-end en/decryption, using RACE leads to, on average, 48 × to 39729 × (for a wide range of security parameters) more energy-efficient solution than purely using a processor.
Zahra Azad, Guowei Yang 0005, Rashmi S. Agrawal 0001, Daniel Ruelas-Petrisko, Michael B. Taylor, Ajay Joshi
ISLPED3
2020 Design-flow Methodology for Secure Group Anonymous Authentication
abstract
In heterogeneous distributed systems, computing devices and software components often come from different providers and have different security, trust, and privacy levels. In many of these systems, the need frequently arises to (i) control the access to services and resources granted to individual devices or components in a context-aware manner and (ii) establish and enforce data sharing policies that preserve the privacy of the critical information on end users. In essence, the need is to authenticate and anonymize an entity or device simultaneously, two seemingly contradictory goals. The design challenge is further complicated by potential security problems, such as man-in-the-middle attacks, hijacked devices, and counterfeits. In this work, we present a system design flow for a trustworthy group anonymous authentication protocol (GAAP), which not only fulfills the desired functionality for authentication and privacy, but also provides strong security guarantees.
Rashmi S. Agrawal 0001, Lake Bu, Eliakin Del Rosario, Michel A. Kinsy
DATE1
2020 Fast Arithmetic Hardware Library For RLWE-Based Homomorphic Encryption
abstract
With billions of devices connected over the internet, the rise of sensor-based electronic devices have led to cloud computing being used as a commodity technology service. These sensor-based devices are often small and limited by power, storage, or compute capabilities, and hence, they achieve these capabilities via cloud services. However, this gives rise to data privacy issues as sensitive data is stored and computed over the cloud, which at most times, is a shared resource. Homomorphic encryption can be used along with cloud services to perform computations on encrypted data, guaranteeing data privacy. While about a decade’s work on improving homomorphic encryption has ensured its practicality, it is still several magnitudes slower than expected, making it expensive and infeasible to use. In this work, we propose a first-of-its-kind FPGA-based arithmetic hardware library that focuses on accelerating the key arithmetic operations involved in Ring Learning with Error (RLWE) based homomorphic encryption. We design and implement the FPGAbased Residue Number System (RNS), Chinese Remainder Theorem (CRT), modulo inverse and modulo reduction operations as a first step. For all of these operations, we include a hardware cost efficient serial, and a fast parallel implementation in the library. A modular and parameterized design approach helps in easy customization, provides flexibility to extend these operations for use in most homomorphic encryption applications, and fits well into emerging FPGA-equipped cloud architectures.
Rashmi S. Agrawal 0001, Lake Bu, Michel A. Kinsy
FCCM1
2020 Towards Programmable All-Digital True Random Number Generator
abstract
Random number generator (RNG) is a core component in many applications such as scientific research, testing and diagnosis, gaming, and cryptosystems (e.g., obfuscation, encryption, and authentication). Although, there are various RNG designs targeting specific application goals such as low-power, high-throughput, stronger security guarantees, a universal programmable RNG design has remained elusive. Indeed, it is a challenge to have only one RNG unit in a system with multiple compute modules with different randomness requirements. In this work, we aim to provide a practical solution to this design challenge by proposing a multi-purpose true random number generator (TRNG), which can be configured in real time to generate random sequences with different requirements. Such a programmable TRNG is able to supply random bits to multiple modules with different demands. The proposed TRNG is a highly convenient multi-purpose hardware primitive that can be deployed in many designs as it provides a tunable physical entropy source and a dynamic cost-performance trade-off.
Rashmi S. Agrawal 0001, Lake Bu, Eliakin Del Rosario, Michel A. Kinsy
ACM Great Lakes Symposium on VLSI1
2020 Quantum-Proof Lightweight McEliece Cryptosystem Co-processor Design
abstract
Due to the rapid advances in the development of quantum computers and their susceptibility to errors, there is a renewed interest in error correction algorithms. In particular, error correcting code-based cryptosystems have reemerged as a highly desirable coding technique. This is due to the fact that most classical asymmetric cryptosystems will fail in the quantum computing era. However, code-based cryptosystems are still secure against quantum computers, since the decoding of linear codes remains NP-hard even on these computing systems. One such code-based cryptosystem was proposed by McEliece. The classic McEliece cryptosystem uses binary Goppa code, which is known for its good code rate and error correction capability. However, its key generation and decoding procedures have a high computation complexity. In this work, we propose the design of a public-key encryption and decryption coprocessor based on a new variant of the McEliece cryptosystem. This co-processor takes advantage of non-binary Orthogonal Latin Square Code to achieve much smaller computation complexity and key size. We also propose a hardware-cost efficient, fully-parameterized FPGA-based implementation of the co-processor to perform fast encoding and decoding operations. When compared to an existing classic McEliece cryptosystem, we observe a speed up of about 3.3 ×.
Rashmi S. Agrawal 0001, Lake Bu, Michel A. Kinsy
ICCD1
2019 Open-Source FPGA Implementation of Post-Quantum Cryptographic Hardware Primitives
abstract
The following topics are dealt with: field programmable gate arrays; learning (artificial intelligence); convolutional neural nets; logic design; system-on-chip; power aware computing; cryptography; cloud computing; reconfigurable architectures; neural nets.
Rashmi S. Agrawal 0001, Lake Bu, Alan Ehret, Michel A. Kinsy
FPL1