VLDB 2026 Research / reviewers in the wild / expert
Nir Drucker
dblp:179/7421
· DBLP profile ↗
17ranked-venue papers
10as first author
9since 2021 · last 2025
0000-0002-7273-4797ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 3 first-author · 6 since 2021Theory of computation · 4 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FHENDI: A Near-DRAM Accelerator for Compiler-Generated Fully Homomorphic Encryption ApplicationsabstractFully homomorphic encryption (FHE) is a powerful cryptographic technique that enables computation on encrypted data without needing to decrypt it. It has broad applications in scenarios where sensitive data needs to be processed in the cloud or in other untrusted environments. FHE applications are both compute- and memory-intensive, owing to expensive operations on large data. While prior works address the challenges of efficient compute using dedicated hardware, expensive memory transfers still remain a major limiting factor. In this work, we propose a hierarchical near-DRAM processing (NDP) solution for FHE applications, called FHENDI, that harnesses the massive DRAM bank bandwidth. We observe various data access patterns in FHE that reveal distinct levels of parallelism: element-wise, limb-wise, coefficient-wise, and ciphertext-wise. FHENDI exploits these levels of parallelism to map FHE operations and data onto different hierarchies of our design, while addressing three major challenges with NDP for FHE: (i) the lack of bank-to-bank communication support, (ii) limited die-to-die bandwidth, and (iii) large memory access latencies. We resolve the first problem through a novel, conflict-free mapping algorithm built atop localized permutation networks that enables efficient element-wise and butterfly operations in FHE. The second problem is addressed by pipelining the execution of parallel bootstrap operations observed in compiled FHE workloads. Finally, we hide the memory access latency behind computation latency by exploiting a dual-banking scheme and subarray-level parallelism (SLP) of the DRAM banks. We evaluate FHENDI using representative workloads in the domains of privacy-preserving machine learning inference on CNNs and Transformers, database range query, and sorting, that are obtained using a compiler framework called HElayers. We compare FHENDI with a server-class CPU and GPU running the state-of-the-art HEaaN library, and an FHE accelerator ASIC, and report mean speedups of $2145.8 \times, 118.29 \times$, and $2.45 \times$, respectively. Yongmo Park, Aporva Amarnath, Subhankar Pal, Karthik Swaminathan, Alper Buyuktosunoglu, Hayim Shaul, Ehud Aharoni, Nir Drucker, Wei Lu 0003, Omri Soceanu, Pradip Bose |
HPCA | 8 |
| 2024 | Converting Transformers to Polynomial Form for Secure Inference Over Homomorphic EncryptionabstractDesigning privacy-preserving DL solutions is a major challenge within the AI community. Homomorphic Encryption (HE) has emerged as one of the most promising approaches in this realm, enabling the decoupling of knowledge between a model owner and a data owner. Despite extensive research and application of this technology, primarily in CNNs, applying HE on transformer models has been challenging because of the difficulties in converting these models into a polynomial form. We break new ground by introducing the first polynomial transformer, providing the first demonstration of secure inference over HE with full transformers. This includes a transformer architecture tailored for HE, alongside a novel method for converting operators to their polynomial equivalent. This innovation enables us to perform secure inference on LMs and ViTs with several datasts and tasks. Our techniques yield results comparable to traditional models, bridging the performance gap with transformers of similar scale and underscoring the viability of HE for state-of-the-art applications. Finally, we assess the stability of our models and conduct a series of ablations to quantify the contribution of each model component. Our code is publicly available. Itamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov, Omri Soceanu, Lior Wolf |
ICML | 3 |
| 2024 | BLEACH: Cleaning Errors in Discrete Computations Over CKKS
Nir Drucker, Guy Moshkowich, Tomer Pelleg, Hayim Shaul |
J. Cryptol. | 1 |
| 2023 | Poster: Efficient AES-GCM Decryption Under Homomorphic EncryptionabstractComputation delegation to untrusted third-party while maintaining data confidentiality is possible with homomorphic encryption (HE). However, in many cases, the data was encrypted using another cryptographic scheme such as AES-GCM. Hybrid encryption (a.k.a Transciphering) is a technique that allows moving between cryptosystems, which currently has two main drawbacks: 1) lack of standardization or bad performance of symmetric decryption under FHE; 2) lack of input data integrity. Ehud Aharoni, Nir Drucker, Gilad Ezov, Eyal Kushnir, Hayim Shaul, Omri Soceanu |
CCS | 2 |
| 2023 | Tutorial-HEPack4ML '23: Advanced HE Packing Methods with Applications to MLabstractOutsourcing computations over sensitive data to a third-party cloud environment should often rely on dedicated privacy-preserving solutions in order to adhere to privacy regulations such as the GDPR [7]. One solution that gained great attention is fully homomorphic encryption (FHE), a cryptographic method that allows performing different types of computation on encrypted data. Still, writing a non-interactive FHE code that evaluates complex functions is a task that is mostly left to experts. Otherwise, the resulted code may become very slow and even impractical. Ehud Aharoni, Nir Drucker, Hayim Shaul |
CCS | 2 |
| 2023 | Efficient Pruning for Machine Learning Under Homomorphic Encryption
Ehud Aharoni, Moran Baruch, Pradip Bose, Alper Buyuktosunoglu, Nir Drucker, Subhankar Pal, Tomer Pelleg, Kanthi K. Sarpatwar, Hayim Shaul, Omri Soceanu, Roman Vaculín |
ESORICS (4) | 5 |
| 2023 | HeLayers: A Tile Tensors Framework for Large Neural Networks on Encrypted DataabstractPrivacy-preserving solutions enable companies to offload confidential data to third-party services while fulfilling their government regulations. To accomplish this, they leverage various cryptographic techniques such as Homomorphic Encryption (HE), which allows performing computation on encrypted data. Most HE schemes work in a SIMD fashion, and the data packing method can dramatically affect the running time and memory costs. Finding a packing method that leads to an optimal performant implementation is a hard task. We present a simple and intuitive framework that abstracts the packing decision for the user. We explain its underlying data structures and optimizer, and propose a novel algorithm for performing 2D convolution operations. We used this framework to implement an inference operation over an encrypted HE-friendly AlexNet neural network with large inputs, which runs in around five minutes, several orders of magnitude faster than other state-of-the-art non-interactive HE solutions. Ehud Aharoni, Allon Adir, Moran Baruch, Nir Drucker, Gilad Ezov, Ariel Farkash, Lev Greenberg, Ramy Masalha, Guy Moshkowich, Dov Murik, Hayim Shaul, Omri Soceanu |
Proc. Priv. Enhancing Technol. | 4 |
| 2021 | Fast polynomial inversion for post quantum QC-MDPC cryptographyabstractNew post-quantum Key Encapsulation Mechanism (KEM) designs, evaluated as part of the NIST PQC standardization Project, pose challenging tradeoffs between communication bandwidth and computational overheads. Several KEM designs evaluated in Round-2 of the project are based on QC-MDPC codes. BIKE-2 uses the smallest communication bandwidth, but its key generation requires a costly polynomial inversion. In this paper, we provide details on the optimized polynomial inversion algorithm for QC-MDPC codes (originally proposed in the conference version of this work). This algorithm makes the runtime of BIKE-2 key generation tolerable. It brings a speedup of 11.4× over the commonly used NTL library, and 83.5× over OpenSSL. We achieve additional speedups by leveraging the latest Intel's Vector-PCLMULQDQ instructions, 14.3× over NTL and 103.9× over OpenSSL. Our algorithm and implementation were the reason that BIKE team chose BIKE-2 as the only scheme for its Round-3 specification (now called BIKE). Nir Drucker, Shay Gueron, Dusan Kostic |
Inf. Comput. | 1 |
| 2021 | Selfie: reflections on TLS 1.3 with PSK
Nir Drucker, Shay Gueron |
J. Cryptol. | 1 |
| 2020 | QC-MDPC Decoders with Several Shades of Gray
Nir Drucker, Shay Gueron, Dusan Kostic |
PQCrypto | 1 |
| 2019 | Fast constant time implementations of ZUC-256 on x86 CPUsabstractZUC-256 is a Pseudo Random Number Generator (PRNG) that is proposed as a successor of ZUC-128. Similarly to ZUC-128 that is incorporated in the 128-EEA3 and 128-EIA3 encryption and integrity algorithms, ZUC-256 is designed to offer 256-bit security and to be incorporated in the upcoming encryption and authentication algorithm in 5G technologies. In this context software optimizations of ZUC-256 are desired. This paper proposes several ZUC-256 optimizations for x86 processors, especially, modern processors that have efficient AVX vectorization. Surprisingly, we also show that AES-NI can also be used for ZUC-256 and help creating constant-time implementations. Our results show speedup of up to 4.5 x(per key stream) when computational tasks are parallelized efficiently. Nir Drucker, Shay Gueron |
CCNC | 1 |
| 2019 | Cyclic-routing of Unmanned Aerial Vehicles
Nir Drucker, Hsi-Ming Ho, Joël Ouaknine, Michal Penn, Ofer Strichman |
J. Comput. Syst. Sci. | 1 |
| 2018 | Fast multiplication of binary polynomials with the forthcoming vectorized VPCLMULQDQ instructionabstractPolynomial multiplication over binary fields \mathbbF2nis a common primitive, used for example by current cryptosystems such as AES-GCM (with n=128). It also turns out to be a primitive for other cryptosystems, that are being designed for the Post Quantum era, with values n ≫ 128. Examples from the recent submissions to the NIST Post-Quantum Cryptography project, are BIKE, LEDAKem, and GeMSS, where the performance of the polynomial multiplications, is significant. Therefore, efficient polynomial multiplication over F2n, with large n, is a significant emerging optimization target. Anticipating future applications, Intel has recently announced that its future architecture (codename “Ice Lake”) will introduce a new vectorized way to use the current VPCLMULQDQ instruction. In this paper, we demonstrate how to use this instruction for accelerating polynomial multiplication. Our analysis shows a prediction for at least 2× speedup for multiplications with polynomials of degree 512 or more. Nir Drucker, Shay Gueron, Vlad Krasnov |
ARITH | 1 |
| 2018 | The Comeback of Reed Solomon CodesabstractDistributed storage systems utilize erasure codes to reduce their storage costs while efficiently handling failures. Many of these codes (e. g., Reed-Solomon (RS) codes) rely on Galois Field (GF) arithmetic, which is considered to be fast when the field characteristic is 2. Nevertheless, some developments in the field of erasure codes offer new efficient techniques that require mostly XOR operations, and are thus faster than GF operations. Recently, Intel announced [1] that its future architecture (codename “Ice Lake”) will introduce new set of instructions called Galois Field New Instruction (GF-NI). These instructions allow software flows to perform vector and matrix multiplications over GF (28) on the wide registers that are available on the AVX512 architectures. In this paper, we explain the functionality of these instructions, and demonstrate their usage for some fast computations in GF(28). We also use the Intel®Intelligent Storage Acceleration Library (ISA-L) in order to estimate potential future improvement for erasure codes that are based on RS codes. Our results predict ≈ 1.4× speedup for vectorized multiplication, and 1.83× speedup for the actual encoding. Nir Drucker, Shay Gueron, Vlad Krasnov |
ARITH | 1 |
| 2018 | Cryptosystems with a multi prime composite modulusabstractMulti-Prime (MP)RSA is an RSA construction in which the public modulus is a product of more than two primes, and its private key operations can be accelerated by using the Chinese Reminder Theorem (CRT). While MPRSA has been studied extensively, only limited information is found for other MP constructions, such as Paillier cryptosystem. This paper shows how to extend the security proofs for Quadratic Residue Problem (QRP), Higher Residuosity Problem (HRP) and Decisional Composite Residuosity Problem (DCRP), formulated for a two-primes modulus, to a MP setting. For the Paillier cryptosystem, we demonstrate how this technique can speed up decryption by more than 17x. Shay Gueron, Nir Drucker |
CCNC | 2 |
| 2017 | Paillier-encrypted databases with fast aggregated queriesabstractThe proliferating usage of cloud environments to store databases poses new challenges. Traditional encryption protects the user's data privacy, but prevents the server from executing computations on behalf of the user (client). By contrast, Partially Homomorphic Encryption schemes, such as the Paillier cryptosystem, facilitate some server queries but involve heavy computations that make them relatively slow. This paper shows a simple performance optimization for Paillier encryption. It significantly reduces the server side workload and can be deployed by the server unilaterally, while remaining transparent to the client. Our optimization trades modular multiplications with cheaper Montgomery Multiplications, by converting the database to a favourable format. We explore several techniques to accelerate the relevant Montgomery multiplications on current and future modern processor architectures, and demonstrate the resulting speed-ups by comparing to the current method implemented via the OpenSSL library. For example, on the latest Intel processor (Architecture Codename Skylake) our method speeds up aggregated queries by a factor of 4×. Nir Drucker, Shay Gueron |
CCNC | 1 |
| 2016 | Cyclic Routing of Unmanned Aerial Vehicles
Nir Drucker, Michal Penn, Ofer Strichman |
CPAIOR | 1 |