Homer Gamil

dblp:198/8757 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-3256-783XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Fine-Grained Parallelization of FHE Workloads in Multi-GPU Systems
Homer Gamil, Michail Maniatakos
ASP-DAC1
2026 Big Integer Parallel Stream Modular Multiplier With Variable Bit-Widths
abstract
In this paper, we present a new modular multiplier design that offers flexibility regarding the operand sizes it processes in parallel. The multiplier can efficiently compute different sizes using the same ASIC hardware, enabling parallel computations for smaller sizes, for example a 1024-bit instantiation of our multiplier can perform either one 1024-bit, sixteen 64-bit, or four 256-bit multiplications, etc. This capability is particularly valuable in accelerating a plethora of cryptosystems, such as RSA, ECC, or Fully Homomorphic Encryption, using the same ASIC hardware, since operand sizes can vary depending on the security parameters and the application requirements. The multiplier can be used in conjunction with software methods for parallelization. For instance, our multiplier enables users to employ both RNS and non-RNS versions of FHE using a single hardware accelerator. We implement our multiplier in hardware and demonstrate its efficiency compared to state-of-theart Montgomery designs, while offering the additional advantage of parallel processing flexibility
Oleg Mazonka, Eduardo Chielle, Mohammed Nabeel Thari Moopan, Homer Gamil, Michail Maniatakos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Coala: Coalescion-Based Acceleration of Polynomial Multiplication for GPU Execution
abstract
In this study, we introduce Coala, a novel framework designed to enhance the performance of finite field transformations for GPU environments. We have developed a GPU-optimized version of the Discrete Galois Transformation (DGT), a variant of the Number Theoretic Transform (NTT). We introduce a novel data access pattern scheme specifically engineered to enable coalesced accesses, significantly enhancing the efficiency of data transfers between global and shared memory. This enhancement not only boosts execution efficiency but also optimizes the interaction with the GPU's memory architecture. Additionally, Coala presents a comprehensive framework that optimizes the allocation of computational tasks across the GPU's architecture and execution kernels, thereby maximizing the use of GPU resources. Lastly, we provide a flexible method to adjust security levels and polynomial sizes through the incorporation of an in-kernel RNS method, and a flexible parameter generation approach. Comparative analysis against current state-of-the-art techniques reveals significant improvements. We observe performance gains of 2.82′ − 17.18′ against other DGT works on GPUs for different parameters, achieved concurrently with equal or lesser memory utilization.
Homer Gamil, Oleg Mazonka, Michail Maniatakos
DATE1
2024 MCS-NTT: Multi-Chip System Design for NTT Acceleration
abstract
Hardware implementations of Number Theoretic Transform (NTT), especially ASIC designs, have provided significant speed improvements for lattice-based cryptography schemes used by Post-Quantum Cryptography (PQC) and Fully Homo-morphic Encryption (FHE). While most of the existing solutions are tailored for fixed polynomial degrees and modulus sizes, both parameters can vary considerably depending on the application and scheme. Toward this end, our paper introduces MCS-NTT, the first hardware architecture for NTT acceleration that is based on a multi-chip-system (MCS) design approach. Our proposed solution provides scalability to existing NTT accelerators by seamlessly integrating multiple accelerator units around an FPGA-based centralized unit. This configuration effectively establishes a customized star network tailored to meet specific use cases. The experimental results indicate that MCS-NTT offers considerable flexibility with better performance metrics.
Mohammed Nabeel Thari Moopan, Homer Gamil, Johann Knechtel, Michail Maniatakos
VLSI-SoC2
2024 Silicon-Proven ASIC Design for the Polynomial Operations of Fully Homomorphic Encryption
abstract
In this work, we elaborate on our endeavors to design, implement, fabricate, and post-silicon validate CoFHEE 1, a co-processor for low-level polynomial operations targeting Fully Homomorphic Encryption execution. With a compact design area of 12mm2, CoFHEE features ASIC implementations of fundamental polynomial operations, including polynomial addition and subtraction, Hadamard product, and Number Theoretic Transform, which underlie most higher-level FHE primitives. CoFHEE is capable of natively supporting polynomial degrees of up to n = 214 with a coefficient size of 128 bits, and has been fabricated and silicon-verified using 55nm CMOS technology. To evaluate it, we conduct performance and power experiments on our chip, and compare it to state-of-the-art software implementations and other ASIC designs.
Mohammed Nabeel Thari Moopan, Homer Gamil, Deepraj Soni, Mohammed Ashraf, Mizan Abraha Gebremichael, Eduardo Chielle, Ramesh Karri, Mihai Sanduleanu, Michail Maniatakos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Coupling bit and modular arithmetic for efficient general-purpose fully homomorphic encryption
abstract
Fully Homomorphic Encryption (FHE) enables computation directly on encrypted data. This property is desirable for outsourced computation of sensitive data as it relies solely on the underlying security of the cryptosystem and not in access control policies. Even though FHE is still significantly slower than unencrypted computation, practical times are possible for applications easily representable as low-order polynomials, since most FHE schemes support modular addition and multiplication over ciphertexts. If, however, an application cannot be expressed with low-order polynomials, then Boolean logic must be emulated. This bit-level arithmetic enables any computation to be performed homomorphically. Nevertheless, as it runs on top of the natively supported modular arithmetic, it has poor performance, which hinders its use in the majority of scenarios. In this work, we propose Bridging, a technique that allows conversion from bit-level to modular arithmetic and vice-versa. This enables the use of the comprehensive computation provided by bit-level arithmetic and the performance of modular arithmetic within the same application. Experimental results show that Bridging can lead to 1-2 orders of magnitude performance improvement for tested benchmarks and two real-world applications: URL denylisting and genotype imputation. Bridging performance comes from two factors: reduced number of operations and smaller multiplicative depth.
Eduardo Chielle, Oleg Mazonka, Homer Gamil, Michail Maniatakos
ACM Trans. Embed. Comput. Syst.3
2023 CoFHEE: A Co-processor for Fully Homomorphic Encryption Execution
abstract
In this paper, we present the blueprint of a specialized co-processor for Fully Homomorphic Encryption, dubbed CoFHEE. With a small design area of$12mm^{2}$, CoFHEE incorporates ASIC implementations of fundamental polynomial operations, such as polynomial addition and subtraction, Hadamard product, and Number Theoretic Transform, which are underneath all higher-level FHE primitives. CoFHEE has native support of polynomial degrees of up to$n=2^{14}$with a coefficient size of 128 bits. We evaluate our chip with performance and power experiments and compare it against state-of-the-art software implementations and other ASIC designs. A more elaborate description of the CoFHEE design can be found in [1].
Mohammed Nabeel Thari Moopan, Deepraj Soni, Mohammed Ashraf, Mizan Abraha Gebremichael, Homer Gamil, Eduardo Chielle, Ramesh Karri, Mihai Sanduleanu, Michail Maniatakos
DATE5
2023 RPU: The Ring Processing Unit
abstract
Ring-Learning-with-Errors (RLWE) has emerged as the foundation of many important techniques for improving security and privacy, including homomorphic encryption and post-quantum cryptography. While promising, these techniques have received limited use due to their extreme overheads of running on general-purpose machines. In this paper, we present a novel vector Instruction Set Architecture (ISA) and microarchitecture for accelerating the ring-based computations of RLWE. The ISA, named B512, is developed to meet the needs of ring processing workloads while balancing high-performance and general-purpose programming support. Having an ISA rather than fixed hardware facilitates continued software improvement post-fabrication and the ability to support the evolving workloads. We then propose the ring processing unit (RPU), a high-performance, modular implementation of B512. The RPU has native large word modular arithmetic support, capabilities for very wide parallel processing, and a large capacity highbandwidth scratchpad to meet the needs of ring processing. We address the challenges of programming the RPU using a newly developed SPIRAL backend. A configurable simulator is built to characterize design tradeoffs and quantify performance. The best performing design was implemented in RTL and used to validate simulator performance. In addition to our characterization, we show that a RPU using 20.5mm2of GF12nm can provide a speedup of 1485× over a CPU running a 64k, 128-bit NTT, a core RLWE workload.
Deepraj Soni, Negar Neda, Naifeng Zhang, Benedict Reynwar, Homer Gamil, Benjamin Heyman, Mohammed Nabeel Thari Moopan, Ahmad Al Badawi, Yuriy Polyakov, Kellie Canida, Massoud Pedram, Michail Maniatakos, David Cousins, Franz Franchetti, Matthew French, Andrew G. Schmidt, Brandon Reagen
ISPASS5
2022 Accelerating Fully Homomorphic Encryption by Bridging Modular and Bit-Level Arithmetic
abstract
The dramatic increase of data breaches in modern computing platforms has emphasized that access control is not sufficient to protect sensitive user data. Recent advances in cryptography allow end-to-end processing of encrypted data without the need for decryption using Fully Homomorphic Encryption (FHE). Such computation however, is still orders of magnitude slower than direct (unencrypted) computation. Depending on the underlying cryptographic scheme, FHE schemes can work natively either at bit-level using Boolean circuits, or over integers using modular arithmetic. Operations on integers are limited to addition/subtraction and multiplication. On the other hand, bit-level arithmetic is much more comprehensive allowing more operations, such as comparison and division. While modular arithmetic can emulate bit-level computation, there is a significant cost in performance. In this work, we propose a novel method, dubbed bridging, that blends faster and restricted modular computation with slower and comprehensive bit-level computation, making them both usable within the same application and with the same cryptographic scheme instantiation. We introduce and open source C++ types representing the two distinct arithmetic modes, offering the possibility to convert from one to the other. Experimental results show that bridging modular and bit-level arithmetic computation can lead to 1--2 orders of magnitude performance improvement for tested synthetic benchmarks, as well as one real-world FHE application: a genotype imputation case study.
Eduardo Chielle, Oleg Mazonka, Homer Gamil, Michail Maniatakos
ICCAD3
2021 Real-time Private Membership Test using Homomorphic Encryption
abstract
With the ever increasing volume of private data residing on the cloud, privacy is becoming a major concern. Often times, sensitive information is leaked during a querying process between a client and an online server hosting a database; The query may leak information about the element the client is looking up, while sensitive details about the contents of its database can leak on the server side. The ability to check if an element is included in a database while maintaining both the client's and the server's privacy is known as the Private Membership Test. In this context, we propose a method to privately query a database with computational complexity O(1) using Bloom filters and Homomorphic Encryption. The proposed methodology also enables post-encryption insertions and deletions without requiring a new setup. Experimental results show that our proposed solution has practical setup, insertion and deletion times for databases of up to a few million entries, with constant query time less than 0.3$s$, considering a false positive rate lower than 10−3. We instantiate our methodology for a URL denylisting service, and demonstrate that it can provide solid security guarantees without affecting the user experience.
Eduardo Chielle, Homer Gamil, Michail Maniatakos
DATE2
2020 Muon-Ra: Quantum random number generation from cosmic rays
abstract
True Random Number Generators (TRNGs) are the cornerstone of modern cryptographic applications. In this work, we present the first quantum1random number generator based on muon detection. The proposed implementation utilizes silicon photomultipliers and plastic scintillators to convert the time interval between crossing muons to random bits. Compared to the state-of-the-art, this design operates using a passive entropy source, scaling down its power consumption significantly. Additionally, the proposed muon-based TRNG can be fully integrated in modern computer hardware, making it suitable for low-power embedded device applications. We evaluate the proposal on its throughput and ability to pass standard randomness tests. Our method is successful in passing the NIST STS SP 800-22 and Dieharder evaluations. Finally, the implementation is compared to other well-established methods of generating random numbers.1We use the term “quantum” to denote the utilization of elementary particles as the output generation source, and not necessarily their properties, similar to related work [1], [2].
Homer Gamil, Pranav Mehta, Eduardo Chielle, Adriano Di Giovanni, Mohammed Nabeel Thari Moopan, Francesco Arneodo, Michail Maniatakos
IOLTS1