Eduardo Chielle

dblp:119/3679 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-1938-912XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 5 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 CHEHAB RL: Learning to Optimize Fully Homomorphic Encryption Computations
abstract
Fully Homomorphic Encryption (FHE) enables computations directly on encrypted data, but its high computational cost remains a significant barrier. Writing efficient FHE code is a complex task requiring cryptographic expertise, and finding the optimal sequence of program transformations is often intractable. In this paper, we propose CHEHAB RL, a novel framework that leverages deep reinforcement learning (RL) to automate FHE code optimization. Instead of relying on predefined heuristics or combinatorial search, our method trains an RL agent to learn an effective policy for applying a sequence of rewriting rules to automatically vectorize scalar FHE code while reducing instruction latency and noise growth. The proposed approach supports the optimization of both structured and unstructured code. To train the agent, we synthesize a diverse dataset of computations using a large language model (LLM). We integrate our proposed approach into the CHEHAB FHE compiler and evaluate it on a suite of benchmarks, comparing its performance against Coyote, a state-of-the-art vectorizing FHE compiler. The results show that our approach generates code that is 5.3× faster in execution, accumulates 2.54× less noise, while the compilation process itself is 27.9× faster than Coyote (geometric means).
Bilel Sefsaf, Abderraouf Dandani, Abdessamed Seddiki, Arab Mohammed, Eduardo Chielle, Michail Maniatakos, Riyadh Baghdadi
ASPLOS (2)5
2026 CHEHAB: Automatic Compiler Code Optimization for Fully Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) enables computations to be performed directly on encrypted data without requiring decryption, providing strong privacy guarantees. However, FHE remains computationally expensive, and writing efficient FHE programs is a complex, error-prone, and time-consuming task that demands significant cryptographic expertise. Programmers are often unaware of available optimizations, and applying them manually requires substantial effort. In this paper, we present CHEHAB, a compiler that automatically vectorizes scalar code, optimizes it, and generates highly efficient FHE programs. CHEHAB supports both structured and unstructured code and takes as input programs written in a domain-specific language embedded in C++. It relies on a Term Rewriting System (TRS) based on equality saturation to simplify and transform programs. CHEHAB targets two key challenges in FHE compilation: (1) automatic vectorization of scalar code, and (2) reduction of instruction execution latency and ciphertext noise growth. By leveraging equality saturation, CHEHAB explores a large optimization space to reduce instruction count and circuit depth while improving vector utilization. Experimental evaluation on a set of representative kernels shows that CHEHAB outperforms Coyote, a state-of-the-art vectorizing compiler for FHE. On average, CHEHAB generates code that is 7.38× faster at runtime, incurs 2.49× less accumulated noise, and achieves 251× faster compilation time. CHEHAB is released as an open-source compiler to support reproducibility and further research in FHE compilation.
Abdessamed Seddiki, Arab Mohammed, Zakaria Hebbal, Aimad Chabounia, Eduardo Chielle, Karima Benatchba, Yacine Challal, Djamel Eddine Menacer, Michail Maniatakos, Riyadh Baghdadi
CC5
2026 Big Integer Parallel Stream Modular Multiplier With Variable Bit-Widths
abstract
In this paper, we present a new modular multiplier design that offers flexibility regarding the operand sizes it processes in parallel. The multiplier can efficiently compute different sizes using the same ASIC hardware, enabling parallel computations for smaller sizes, for example a 1024-bit instantiation of our multiplier can perform either one 1024-bit, sixteen 64-bit, or four 256-bit multiplications, etc. This capability is particularly valuable in accelerating a plethora of cryptosystems, such as RSA, ECC, or Fully Homomorphic Encryption, using the same ASIC hardware, since operand sizes can vary depending on the security parameters and the application requirements. The multiplier can be used in conjunction with software methods for parallelization. For instance, our multiplier enables users to employ both RNS and non-RNS versions of FHE using a single hardware accelerator. We implement our multiplier in hardware and demonstrate its efficiency compared to state-of-theart Montgomery designs, while offering the additional advantage of parallel processing flexibility
Oleg Mazonka, Eduardo Chielle, Mohammed Nabeel Thari Moopan, Homer Gamil, Michail Maniatakos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Recurrent Private Set Intersection for Unbalanced Databases with Cuckoo Hashing and Leveled FHE
Eduardo Chielle, Michail Maniatakos
NDSS1
2024 Optimizing Ciphertext Management for Faster Fully Homomorphic Encryption Computation
abstract
Fully Homomorphic Encryption (FHE) is the pin-nacle of privacy-preserving outsourced computation as it enables meaningful computation to be performed in the encrypted domain without the need for decryption or back-and-forth communication between the client and service provider. Nevertheless, FHE is still orders of magnitude slower than unencrypted computation, which hinders its widespread adoption. In this work, we propose Furbo, a plug-and-play framework that can act as middleware between any FHE compiler and any FHE library. Our proposal employs smart ciphertext memory management and caching techniques to reduce data movement and computation, and can be applied to FHE applications without modifications to the underlying code. Experimental results using Microsoft SEAL as the base FHE library and focusing on privacy-preserving Machine Learning as a Service show up to 2x performance improvement in the fully-connected layers, and up to 24x improvement in the convolutional layers without any code change.
Eduardo Chielle, Oleg Mazonka, Michail Maniatakos
DATE1
2024 Silicon-Proven ASIC Design for the Polynomial Operations of Fully Homomorphic Encryption
abstract
In this work, we elaborate on our endeavors to design, implement, fabricate, and post-silicon validate CoFHEE 1, a co-processor for low-level polynomial operations targeting Fully Homomorphic Encryption execution. With a compact design area of 12mm2, CoFHEE features ASIC implementations of fundamental polynomial operations, including polynomial addition and subtraction, Hadamard product, and Number Theoretic Transform, which underlie most higher-level FHE primitives. CoFHEE is capable of natively supporting polynomial degrees of up to n = 214 with a coefficient size of 128 bits, and has been fabricated and silicon-verified using 55nm CMOS technology. To evaluate it, we conduct performance and power experiments on our chip, and compare it to state-of-the-art software implementations and other ASIC designs.
Mohammed Nabeel Thari Moopan, Homer Gamil, Deepraj Soni, Mohammed Ashraf, Mizan Abraha Gebremichael, Eduardo Chielle, Ramesh Karri, Mihai Sanduleanu, Michail Maniatakos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Coupling bit and modular arithmetic for efficient general-purpose fully homomorphic encryption
abstract
Fully Homomorphic Encryption (FHE) enables computation directly on encrypted data. This property is desirable for outsourced computation of sensitive data as it relies solely on the underlying security of the cryptosystem and not in access control policies. Even though FHE is still significantly slower than unencrypted computation, practical times are possible for applications easily representable as low-order polynomials, since most FHE schemes support modular addition and multiplication over ciphertexts. If, however, an application cannot be expressed with low-order polynomials, then Boolean logic must be emulated. This bit-level arithmetic enables any computation to be performed homomorphically. Nevertheless, as it runs on top of the natively supported modular arithmetic, it has poor performance, which hinders its use in the majority of scenarios. In this work, we propose Bridging, a technique that allows conversion from bit-level to modular arithmetic and vice-versa. This enables the use of the comprehensive computation provided by bit-level arithmetic and the performance of modular arithmetic within the same application. Experimental results show that Bridging can lead to 1-2 orders of magnitude performance improvement for tested benchmarks and two real-world applications: URL denylisting and genotype imputation. Bridging performance comes from two factors: reduced number of operations and smaller multiplicative depth.
Eduardo Chielle, Oleg Mazonka, Homer Gamil, Michail Maniatakos
ACM Trans. Embed. Comput. Syst.1
2023 CoFHEE: A Co-processor for Fully Homomorphic Encryption Execution
abstract
In this paper, we present the blueprint of a specialized co-processor for Fully Homomorphic Encryption, dubbed CoFHEE. With a small design area of$12mm^{2}$, CoFHEE incorporates ASIC implementations of fundamental polynomial operations, such as polynomial addition and subtraction, Hadamard product, and Number Theoretic Transform, which are underneath all higher-level FHE primitives. CoFHEE has native support of polynomial degrees of up to$n=2^{14}$with a coefficient size of 128 bits. We evaluate our chip with performance and power experiments and compare it against state-of-the-art software implementations and other ASIC designs. A more elaborate description of the CoFHEE design can be found in [1].
Mohammed Nabeel Thari Moopan, Deepraj Soni, Mohammed Ashraf, Mizan Abraha Gebremichael, Homer Gamil, Eduardo Chielle, Ramesh Karri, Mihai Sanduleanu, Michail Maniatakos
DATE6
2022 Accelerating Fully Homomorphic Encryption by Bridging Modular and Bit-Level Arithmetic
abstract
The dramatic increase of data breaches in modern computing platforms has emphasized that access control is not sufficient to protect sensitive user data. Recent advances in cryptography allow end-to-end processing of encrypted data without the need for decryption using Fully Homomorphic Encryption (FHE). Such computation however, is still orders of magnitude slower than direct (unencrypted) computation. Depending on the underlying cryptographic scheme, FHE schemes can work natively either at bit-level using Boolean circuits, or over integers using modular arithmetic. Operations on integers are limited to addition/subtraction and multiplication. On the other hand, bit-level arithmetic is much more comprehensive allowing more operations, such as comparison and division. While modular arithmetic can emulate bit-level computation, there is a significant cost in performance. In this work, we propose a novel method, dubbed bridging, that blends faster and restricted modular computation with slower and comprehensive bit-level computation, making them both usable within the same application and with the same cryptographic scheme instantiation. We introduce and open source C++ types representing the two distinct arithmetic modes, offering the possibility to convert from one to the other. Experimental results show that bridging modular and bit-level arithmetic computation can lead to 1--2 orders of magnitude performance improvement for tested synthetic benchmarks, as well as one real-world FHE application: a genotype imputation case study.
Eduardo Chielle, Oleg Mazonka, Homer Gamil, Michail Maniatakos
ICCAD1
2022 Fast and Compact Interleaved Modular Multiplication Based on Carry Save Addition
abstract
Improving fully homomorphic encryption computation by designing specialized hardware is an active topic of research. The most prominent encryption schemes operate on long polynomials requiring many concurrent modular multiplications of very big numbers. Thus, it is crucial to use many small and efficient multipliers. Interleaved and Montgomery iterative multipliers are the best candidates for the task. Interleaved designs, however, suffer from longer latency as they require a number comparison within each iteration; Montgomery designs, on the other hand, need extra conversion of the operands or the result. In this work, we propose a novel hardware design that combines the best of both worlds: Exhibiting the carry save addition of Montgomery designs without the need for any domain conversions. Experimental results demonstrate improved latency-area product efficiency by up to 47% when compared to the standard Interleaved multiplier for large arithmetic word sizes.
Oleg Mazonka, Eduardo Chielle, Deepraj Soni, Michail Maniatakos
ICCAD2
2022 E3X: Encrypt-Everything-Everywhere ISA eXtensions for Private Computation
abstract
The rapid increase of recent privacy attacks has significantly decreased trust on behalf of the users. A root cause to these problems is that modern computer architectures have always been designed for performance, while security protections are traditionally addressed reactively. Practical security protections, such as Intel SGX, rely on processing unencrypted data in the architectural state, which leaves them exposed to software attacks (e.g., SGXpectre). This work revisits the traditional computation stack and introduces a novel computation paradigm, where data is never decrypted in the architectural state. Through our architecture, data are protected with symmetric or asymmetric encryption and the programmer manipulates them directly in the encrypted domain. To increase performance, we exploit data locality by introducing decryption caches in the microarchitectural state. Our proposal addresses all abstraction levels in the computation stack: from microarchitecture to library support for high-level programming. The proposed architecture is instantiated through new assembly instructions, registers and functional units operating on large integers. In our evaluation, we extend the OpenRISC 1000 architecture and develop open-source libraries for C++. As a case study, we employ data-oblivious benchmarks and observe that for benchmarks with high temporal locality, our architecture can achieve comparable performance to processing unencrypted data.
Eduardo Chielle, Nektarios Georgios Tsoutsos, Oleg Mazonka, Michail Maniatakos
IEEE Trans. Dependable Secur. Comput.1
2021 Real-time Private Membership Test using Homomorphic Encryption
abstract
With the ever increasing volume of private data residing on the cloud, privacy is becoming a major concern. Often times, sensitive information is leaked during a querying process between a client and an online server hosting a database; The query may leak information about the element the client is looking up, while sensitive details about the contents of its database can leak on the server side. The ability to check if an element is included in a database while maintaining both the client's and the server's privacy is known as the Private Membership Test. In this context, we propose a method to privately query a database with computational complexity O(1) using Bloom filters and Homomorphic Encryption. The proposed methodology also enables post-encryption insertions and deletions without requiring a new setup. Experimental results show that our proposed solution has practical setup, insertion and deletion times for databases of up to a few million entries, with constant query time less than 0.3$s$, considering a false positive rate lower than 10−3. We instantiate our methodology for a URL denylisting service, and demonstrate that it can provide solid security guarantees without affecting the user experience.
Eduardo Chielle, Homer Gamil, Michail Maniatakos
DATE1
2020 Muon-Ra: Quantum random number generation from cosmic rays
abstract
True Random Number Generators (TRNGs) are the cornerstone of modern cryptographic applications. In this work, we present the first quantum1random number generator based on muon detection. The proposed implementation utilizes silicon photomultipliers and plastic scintillators to convert the time interval between crossing muons to random bits. Compared to the state-of-the-art, this design operates using a passive entropy source, scaling down its power consumption significantly. Additionally, the proposed muon-based TRNG can be fully integrated in modern computer hardware, making it suitable for low-power embedded device applications. We evaluate the proposal on its throughput and ability to pass standard randomness tests. Our method is successful in passing the NIST STS SP 800-22 and Dieharder evaluations. Finally, the implementation is compared to other well-established methods of generating random numbers.1We use the term “quantum” to denote the utilization of elementary particles as the output generation source, and not necessarily their properties, similar to related work [1], [2].
Homer Gamil, Pranav Mehta, Eduardo Chielle, Adriano Di Giovanni, Mohammed Nabeel Thari Moopan, Francesco Arneodo, Michail Maniatakos
IOLTS3
2018 PHYLAX: Snapshot-based profiling of real-time embedded devices via JTAG interface
abstract
Real-time embedded systems play a significant role in the functionality of critical infrastructure. Legacy microprocessor-based embedded systems, however, have not been developed with security in mind. Applying traditional security mechanisms in such systems is challenging due to computing constraints and/or real-time requirements. Their typical 20-30 year lifespan further exacerbates the problem. In this work, we propose PHYLAX, a plug-and-play solution to detect intrusions in already installed embedded devices. PHYLAX is an external monitoring tool which does not require code instrumentation. Also, our tool adapts and prioritizes intrusion detection based on the requirements of the underlying infrastructure (power grid, chemical factory, etc.) as well as the computing capabilities of the target embedded system (CPU model, memory size, etc.). PHYLAX can be employed on any legacy device which incorporates a JTAG interface. As a case study, we present the inclusion of PHYLAX on a power grid recloser controller.
Charalambos Konstantinou, Eduardo Chielle, Michail Maniatakos
DATE2
2015 Application-Based Analysis of Register File Criticality for Reliability Assessment in Embedded Microprocessors
Felipe Restrepo-Calle, Sergio Cuenca-Asensi, Antonio Martínez-Álvarez, Eduardo Chielle, Fernanda Lima Kastensmidt
J. Electron. Test.4