EDBT 2026 Demo / reviewers in the wild / expert
Erdinç Öztürk
dblp:24/4753
· DBLP profile ↗
19ranked-venue papers
5as first author
6since 2021 · last 2022
0000-0003-0456-9213ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 6 since 2021Security and privacy · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | An Accelerated GPU Library for Homomorphic Encryption Operations of BFV SchemeabstractThis paper presents an accelerated and parallelized GPU implementation for homomorphic encryption operations of the Brakerski-Fan-Vercauteren (BFV) scheme. We improved the run-time performance by optimizing homomorphic multiplication, relinearization, rotation, and addition using Number Theoretic Transform (NTT) and Barrett Reduction and utilizing a Compute Unified Device Architecture (CUDA). To the best of our knowledge, this implementation performs the fastest homomorphic operations in the literature. We used the Simple Encrypted Arithmetic Library (SEAL) version 3.6.6 BFV scheme for implementation on a GPU. Our implementation achieved $13.39\times, 47.01\times, 39.6\times$, and $33.71\times$ speedup compared to SEAL running on CPU for addition, multiplication, relinearization, and rotation, respectively for a modulus size of 438-bits and ring degree of 16,384. For the same modulus size and ring degree, this implementation performed one homomorphic multiplication in 1 ms, a relinearization operation in 0.4 ms, a rotation in 0.5 ms, and an addition in 0.017 ms, which demonstrates significant performance improvement over state-of-the-art. Enes Recep Türkoglu, Ali Sah Özcan, Can Ayduman, Ahmet Can Mert, Erdinç Öztürk, Erkay Savas |
ISCAS | 5 |
| 2022 | Machine learning-based load distribution and balancing in heterogeneous database management systemsabstractSummary For dynamic and continuous data analysis, conventional OLTP systems are slow in performance. Today's cutting‐edge high‐performance computing hardware, such as GPUs, has been used as accelerators for data analysis tasks, which traditionally leverage CPUs on classical database management systems (DBMS). When CPUs and GPUs are used together, the architectural heterogeneity, that is, leveraging hardware with different performance characteristics jointly, creates complex problems that need careful treatment for performance optimization. Load distribution and balancing are crucial problems for DBMSs working on heterogeneous architectures. In this work, focusing on a hybrid, CPU‐GPU database management system to process users' queries, we propose heuristical and machine‐learning‐based (ML‐based) load distribution and balancing models. In more detail, we employ multiple linear regression (MLR), random forest (RF), and Adaboost (Ada) models to dynamically decide the processing unit for each incoming query based on the response time predictions on both CPU and GPU. The ML‐based models outperformed the other algorithms, as well as the CPU and GPU‐only running modes with up to 27%, 29%, and 40%, respectively, in overall performance (response time) while answering intense real‐life working scenarios. Finally, we propose to use a hybrid load‐balancing model that would be more efficient than the models we tested in this work. Anes Abdennebi, Anil Elakas, Fatih Tasyaran, Erdinç Öztürk, Kamer Kaya, Sinan Yildirim |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | An Extensive Study of Flexible Design Methods for the Number Theoretic TransformabstractEfficient lattice-based cryptosystems operate with polynomial rings with the Number Theoretic Transform (NTT) to reduce the computational complexity of polynomial multiplication. NTT has therefore become a major arithmetic component (thus computational bottleneck) in various cryptographic constructions like hash functions, key-encapsulation mechanisms, digital signatures, and homomorphic encryption. Although there exist several hardware designs in prior work for NTT, they all are isolated design instances fixed for specific NTT parameters or parallelization level. This article provides an extensive study of flexible design methods for NTT implementation. To that end, we evaluate three cases: (1) parametric hardware design, (2) high-level synthesis (HLS) design approach, and (3) design for software implementation compiled on soft-core processors, where all are targeted on reconfigurable hardware devices. We evaluate the designs that implement multiple NTT parameters and/or processing elements, demonstrate the design details for each case, and provide a fair comparison with each other and prior work. On a Xilinx Virtex-7 FPGA, compared to HLS and processor-based methods, the results show that the parametric hardware design is on average$4.4\times$and$73.9\times$smaller and$22.5\times$and$19.3\times$faster, respectively. Surprisingly, HLS tools can yield less efficient solutions than processor-based approaches in some cases. Ahmet Can Mert, Emre Karabulut, Erdinç Öztürk, Erkay Savas, Aydin Aysu |
IEEE Trans. Computers | 3 |
| 2022 | Low-Latency ASIC Algorithms of Modular Squaring of Large Integers for VDF EvaluationabstractThis article is an attempt in quest of the fastest hardware algorithms for the computation of the evaluation component of verifiable delay functions (VDFs),$a^{2^T} \bmod N$, proposed for use in various distributed protocols, in which no party is assumed to compute it significantly faster than other participants. To this end, we propose a class of modular squaring algorithms suitable for low-latency ASIC implementations. The proposed algorithms aim to achieve highest levels of parallelization that have not been explored in previous works in the literature, which usually pursue more balanced optimization of speed and area. For this, we utilize redundant representations of integers and introduce three modular squaring algorithms that work with integers in redundant forms: i) Montgomery algorithm, ii) memory-based algorithm and iii) direct reduction algorithm for fixed moduli. All algorithms enable$O(\log k)$depth circuit implementations, where$k$is the bit-size of the modulus$N$in the VDF function. We analyze and compare gate level-circuits of the proposed algorithms and provide estimates for their critical path delay and gate count. Ahmet Can Mert, Erdinç Öztürk, Erkay Savas |
IEEE Trans. Computers | 2 |
| 2022 | Efficient number theoretic transform implementation on GPU for homomorphic encryption
Özgün Özerk, Can Elgezen, Ahmet Can Mert, Erdinç Öztürk, Erkay Savas |
J. Supercomput. | 4 |
| 2021 | A Hardware Accelerator for Polynomial Multiplication Operation of CRYSTALS-KYBER PQC SchemeabstractPolynomial multiplication is one of the most time-consuming operations utilized in lattice-based post-quantum cryptography (PQC) schemes. CRYSTALS-KYBER is a lattice-based key encapsulation mechanism (KEM) and it was recently announced as one of the four finalists at round three in NIST's PQC Standardization. Therefore, efficient implementations of polynomial multiplication operation are crucial for highperformance CRYSTALS-KYBER applications. In this paper, we propose three different hardware architectures (lightweight, balanced, high-performance) that implement the NTT, Inverse NTT (INTT) and polynomial multiplication operations for the CRYSTALS-KYBER scheme. The proposed architectures include a unified butterfly structure for optimizing polynomial multiplication and can be utilized for accelerating the key generation, encryption and decryption operations of CRYSTALS-KYBER. Our high-performance hardware with 16 butterfly units shows up to 112×, 132× and 109× improved performance for NTT, INTT and polynomial multiplication, respectively, compared to the high-speed software implementations on Cortex-M4. Ferhat Yaman, Ahmet Can Mert, Erdinç Öztürk, Erkay Savas |
DATE | 3 |
| 2020 | A Flexible and Scalable NTT Hardware : Applications from Homomorphically Encrypted Deep Learning to Post-Quantum CryptographyabstractThe Number Theoretic Transform (NTT) enables faster polynomial multiplication and is becoming a fundamental component of next-generation cryptographic systems. NTT hardware designs have two prevalent problems related to design-time flexibility. First, algorithms have different arithmetic structures causing the hardware designs to be manually tuned for each setting. Second, applications have diverse throughput/area needs but the hardware have been designed for a fixed, pre-defined number of processing elements. This paper proposes a parametric NTT hardware generator that takes arithmetic configurations and the number of processing elements as inputs to produce an efficient hardware with the desired parameters and throughput. We illustrate the employment of the proposed design in two applications with different needs: A homomorphically encrypted deep neural network inference (CryptoNets) and a post-quantum digital signature scheme (qTESLA). We propose the first NTT hardware acceleration for both applications on FPGAs. Compared to prior software and high-level synthesis solutions, the results show that our hardware can accelerate NTT up to 28× and 48×, respectively. Therefore, our work paves the way for high-level, automated, and modular design of next-generation cryptographic hardware solutions. Ahmet Can Mert, Emre Karabulut, Erdinç Öztürk, Erkay Savas, Michela Becchi, Aydin Aysu |
DATE | 3 |
| 2020 | Design and Implementation of Encryption/Decryption Architectures for BFV Homomorphic Encryption SchemeabstractFully homomorphic encryption (FHE) is a technique that allows computations on encrypted data without the need for decryption and it provides privacy in various applications such as privacy-preserving cloud computing. In this article, we present two hardware architectures optimized for accelerating the encryption and decryption operations of the Brakerski/Fan-Vercauteren (BFV) homomorphic encryption scheme with high-performance polynomial multipliers. For proof of concept, we utilize our architectures in a hardware/software codesign accelerator framework, in which encryption and decryption operations are offloaded to an FPGA device, while the rest of operations in the BFV scheme are executed in software running on an off-the-shelf desktop computer. Specifically, our accelerator framework is optimized to accelerate Simple Encrypted Arithmetic Library (SEAL), developed by the Cryptography Research Group at Microsoft Research. The hardware part of the proposed framework targets the XILINX VIRTEX7 FPGA device, which communicates with its software part via a peripheral component interconnect express (PCIe) connection. For proof of concept, we implemented our designs targeting 1024degree polynomials with 8-bit and 32-bit coefficients for plaintext and ciphertext, respectively. The proposed framework achieves almost 12× and 7× latency speedups, including I/O operations for the offloaded encryption and decryption operations, respectively, compared to their pure software implementations. Ahmet Can Mert, Erdinç Öztürk, Erkay Savas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Design and Implementation of a Fast and Scalable NTT-Based Polynomial Multiplier ArchitectureabstractIn this paper, we present an optimized FPGA implementation of a novel, fast and highly parallelized NTT-based polynomial multiplier architecture, which proves to be effective as an accelerator for lattice-based homomorphic cryptographic schemes. As I/O operations are as time-consuming as NTT operations during homomorphic computations in a host processor/accelerator setting, instead of achieving the fastest NTT implementation possible on the target FPGA, we focus on a balanced time performance between the NTT and I/O operations. Even with this goal, we achieved the fastest NTT implementation in literature, to the best of our knowledge. For proof of concept, we utilize our architecture in a framework for Fan-Vercauteren (FV) homomorphic encryption scheme, utilizing a hardware/software co-design approach, in which polynomial multiplication operations are offloaded to the accelerator via PCIe bus while the rest of operations in the FV scheme are executed in software running on an off-the-shelf desktop computer. Specifically, our framework is optimized to accelerate Simple Encrypted Arithmetic Library (SEAL), developed by the Cryptography Research Group at Microsoft Research, for the FV encryption scheme, where large degree polynomial multiplications are utilized extensively. The hardware part of the proposed framework targets Xilinx Virtex-7 FPGA device and the proposed framework achieves almost 11x latency speedup for the offloaded operations compared to their pure software implementations. We achieved a throughput of almost 800K polynomial multiplications per second, for polynomials of degree 1024 with 32-bit coefficients. Ahmet Can Mert, Erdinç Öztürk, Erkay Savas |
DSD | 2 |
| 2017 | A Custom Accelerator for Homomorphic Encryption ApplicationsabstractAfter the introduction of first fully homomorphic encryption scheme in 2009, numerous research work has been published aiming at making fully homomorphic encryption practical for daily use. The first fully functional scheme and a few others that have been introduced has been proven difficult to be utilized in practical applications, due to efficiency reasons. Here, we propose a custom hardware accelerator, which is optimized for a class of reconfigurable logic, for Lopez-Alt, Tromer and Vaikuntanathan's somewhat homomorphic encryption based schemes. Our design is working as a co-processor which enables the operating system to offload the most compute-heavy operations to this specialized hardware. The core of our design is an efficient hardware implementation of a polynomial multiplier as it is the most compute-heavy operation of our target scheme. The presented architecture can compute the product of very-large polynomials in under 6.25 ms which is 102 times faster than its software implementation. In case of accelerating homomorphic applications; we estimate the per block homomorphic AES as 442 ms which is 28.5 and 17 times faster than the CPU and GPU implementations, respectively. In evaluation of Prince block cipher homomorphically, we estimate the performance as 52 ms which is 66 times faster than the CPU implementation. Erdinç Öztürk, Yarkin Doröz, Erkay Savas, Berk Sunar |
IEEE Trans. Computers | 1 |
| 2015 | Accelerating LTV Based Homomorphic Encryption in Reconfigurable Hardware
Yarkin Doröz, Erdinç Öztürk, Erkay Savas, Berk Sunar |
CHES | 2 |
| 2015 | Accelerating Fully Homomorphic Encryption in HardwareabstractWe present a custom architecture for realizing the Gentry-Halevi fully homomorphic encryption (FHE) scheme. This contribution presents the first full realization of FHE in hardware. The architecture features an optimized multi-million bit multiplier based on the Schonhage Strassen multiplication algorithm. Moreover, a number of optimizations including spectral techniques as well as a precomputation strategy is used to significantly improve the performance of the overall design. When synthesized using 90 nm technology, the presented architecture achieves to realize the encryption, decryption, and recryption operations in 18.1 msec, 16.1 msec, and 3.1 sec, respectively, and occupies a footprint of less than 30 million gates. Yarkin Doröz, Erdinç Öztürk, Berk Sunar |
IEEE Trans. Computers | 2 |
| 2013 | Evaluating the Hardware Performance of a Million-Bit MultiplierabstractIn this work we present the first full and complete evaluation of a very large multiplication scheme in custom hardware. We designed a novel architecture to realize a million-bit multiplication architecture based on the Schönhage-Strassen Algorithm and the Number Theoretical Transform (NTT). The construction makes use of an innovative cache architecture along with processing elements customized to match the computation and access patterns of the FFT-based recursive multiplication algorithm. When synthesized using a 90nm TSMC library operating at a frequency of 666 MHz, our architecture is able to compute the product of integers in excess of a million bits in 7.74 milliseconds. Estimates show that the performance of our design matches that of previously reported software implementations on a high-end 3 Ghz Intel Xeon processor, while requiring only a tiny fraction of the area. Yarkin Doröz, Erdinç Öztürk, Berk Sunar |
DSD | 2 |
| 2008 | Unclonable Lightweight Authentication Scheme
Ghaith Hammouri, Erdinç Öztürk, Berk Birand, Berk Sunar |
ICICS | 2 |
| 2008 | Physical unclonable function with tristate buffersabstractThe lack of robust tamper-proofing techniques in security applications has provided attackers the ability to virtually circumvent mathematically strong cryptographic primitives by directly attacking the hardware. Consequently, physical tamper-proofing has emerged as an essential element in secure system design. To this end, physical unclonable functions (PUF) have been proposed to build tamper proof hardware and thereby create secure data storages. A delay based PUF scheme was proposed which uses the intrinsic process variations to randomize switches and thereby implement a pseudorandom function family. When tampered with, the device experiences a change in its internal physical parameters. This will alter the pseudorandom function family. Therefore, rendering an attack unsuccessful. In this paper, we propose a delay based PUF circuit that is constructed from tristate buffers. The proposed PUF circuit consumes less power and requires less area. Furthermore, we develop a linear delay model which turns out to be identical to that of switch based PUFs. Hence, the nonlinearization techniques proposed to secure switch based PUFs can directly be applied to secure tristate PUFs as well. Erdinç Öztürk, Ghaith Hammouri, Berk Sunar |
ISCAS | 1 |
| 2008 | Towards Robust Low Cost Authentication for Pervasive DevicesabstractLow cost devices such as RFIDs, sensor network nodes, and smartcards are crucial for building the next generation pervasive and ubiquitous networks. The inherent power and footprint limitations of such networks, prevent us from employing standard cryptographic techniques for authentication which were originally designed to secure high end systems with abundant power. Furthermore, the sharp increase in the number, diversity and strength of physical attacks which directly target the implementation may have devastating consequences in a network setting creating a single point of failure. A compromised node may leak a master key, or may give the attacker an opportunity for injecting faulty messages. In this paper we present a lightweight challenge response authentication scheme based on noisy physical unclonable functions (PUF) that allows for extremely efficient implementations. Furthermore, the inherent properties of PUFs provide cryptographically strong tamper resilience. In a network setting this means that a tampered device will no longer authenticate and in a sense will be isolated from the network. Erdinç Öztürk, Ghaith Hammouri, Berk Sunar |
PerCom | 1 |
| 2008 | A tamper-proof and lightweight authentication scheme
Ghaith Hammouri, Erdinç Öztürk, Berk Sunar |
Pervasive Mob. Comput. | 2 |
| 2007 | Tate Pairing with Strong Fault Resiliency
Erdinç Öztürk, Gunnar Gaubatz, Berk Sunar |
FDTC | 1 |
| 2004 | Low-Power Elliptic Curve Cryptography Using Scaled Modular Arithmetic
Erdinç Öztürk, Berk Sunar, Erkay Savas |
CHES | 1 |