EDBT 2026 Demo / reviewers in the wild / expert
Ismail San
dblp:72/10309
· DBLP profile ↗
13ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0003-3005-1813ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 4 since 2021Security and privacy · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 1,024-FPGA DES Supercomputer on the AWS CloudabstractABSTRACT We present a 1,024‐FPGA DES supercomputer accelerator that is automatically compiled from a single‐threaded sequential DES key search application by means of our High‐Level Synthesis compiler. Our 1,024‐FPGA supercomputer is deployed on several Amazon Web Services (AWS) EC2 F1 instance platforms from different AWS regions. Consequently, it can be considered the first multi‐chip application‐specific supercomputer that is scattered to multiple geographically distributed data centers around the world. Furthermore, invoking our 1,024‐FPGA DES supercomputer is functionally identical to invoking the single‐threaded sequential DES application the supercomputer accelerator is compiled from. Our 1,024‐FPGA supercomputer achieves 3.016E+12 keys/sec and it performs 5,286,000 times better than an AWS EC2 m5.8xlarge Xeon x86 machine executing the original sequential application with a performance of 5.706 E+5 keys/sec. Kemal Ebcioglu, Batuhan Bulut, Atakan Dogan, Gürhan Küçük, Ismail San |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | IMpress: Large Integer Multiplication Expression Rewriting for FPGA HLSabstractLarge integer multiplication is becoming a major challenge for FPGA-based acceleration of many cryptographic applications. Existing techniques for decomposing and optimizing large integer multiplication bring about nontrivial trade-offs between different resource types as well as performance. In this work, we regard determining the level and order of multiplication decomposition as a phase ordering problem, which is a notable problem in compiler optimization. Our framework, IMpress, leverages equality saturation to automatically produce a wide range of equivalent integer multiplication expressions corresponding to various hardware implementations. We devise constrained and multi-objective extraction techniques to automatically choose the optimal expressions based on the resource requirements of a given application. IMpress automatically translates extracted integer multiplication expressions into behavioral descriptions in C++ and initiates FPGA compilation through high-level synthesis. IMpress offers significant control over resource utilization and balance, and it increases the maximum number of instances of cryptographic applications on FPGA. Ecenur Ustun, Ismail San, Cunxi Yu, Zhiru Zhang |
FCCM | 2 |
| 2022 | Deep reinforcement learning-based autonomous parking design with neural network compute acceleratorsabstractAbstract We describe the design and implementation of an autonomous prototype vehicle which finds an empty parking slot in a parking area, and parks itself in the empty parking slot, using neural networks based on deep reinforcement learning (RL). To perform an autonomous parking procedure for our prototype vehicle, two different artificial neural networks (ANNs) are trained using a deep RL Algorithm in a simulation environment and embedded into the computing platform of the prototype car. One of the ANNs enables the vehicle to drive autonomously in the parking environment. At the same time, an image processing algorithm is used to determine whether a parking slot is empty. When the image processing algorithm finds a suitable parking slot, a different ANN is activated and performs a safe parking procedure. However, ANN‐based machine learning techniques require high processing power and impose a high computational burden on embedded CPU and GPU platforms. To alleviate the computational burden, one can achieve higher performance and less power consumption using an application‐specific hardware design, where logic resources are fully exploited according to the algorithm of interest, in an energy‐efficient manner. In this article, hardware accelerators for our ANN models are designed and generated via the Vivado high‐level synthesis (HLS) tool, targeting an ARM based programmable SoC platform, ZedBoard. Our ANN accelerators have achieved a speedup of 17x as compared to an ARM software implementation. For deeper fully‐connected layers used in deep RL‐based solutions, function‐level parallelism (Vivado's dataflow) is employed to improve the computational efficiency. Our proposed stage‐level description for fully connected layers outperforms recent studies in terms of computation time. Alican Özeloglu, Ismihan Gül Gürbüz, Ismail San |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Highly Parallel Multi-FPGA System Compilation from Sequential C/C++ Code in the AWS CloudabstractWe present a High Level Synthesis compiler that automatically obtains a multi-chip accelerator system from a single-threaded sequential C/C++ application. Invoking the multi-chip accelerator is functionally identical to invoking the single-threaded sequential code the multi-chip accelerator is compiled from. Therefore, software development for using the multi-chip accelerator hardware is simplified, but the multi-chip accelerator can exhibit extremely high parallelism. We have implemented, tested, and verified our push-button system design model on multiple field-programmable gate arrays (FPGAs) of the Amazon Web Services EC2 F1 instances platform, using, as an example, a sequential-natured DES key search application that does not have any DOALL loops and that tries each candidate key in order and stops as soon as a correct key is found. An 8- FPGA accelerator produced by our compiler achieves 44,600 times better performance than an x86 Xeon CPU executing the sequential single-threaded C program the accelerator was compiled from. New features of our compiler system include: an ability to parallelize outer loops with loop-carried control dependences, an ability to pipeline an outer loop without fully unrolling its inner loops, and fully automated deployment, execution and termination of multi-FPGA application-specific accelerators in the AWS cloud, without requiring any manual steps. Kemal Ebcioglu, Ismail San |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2017 | Detecting hardware Trojans in unspecified functionality through solving satisfiability problemsabstractFor modern complex designs it is impossible to fully specify design behavior, and only feasible to verify functionally meaningful scenarios. Hardware Trojans modifying only unspecified functionality are not possible to detect using existing verification methodologies and Trojan detection strategies. We propose a detection methodology for these Trojans by 1) precisely defining “suspicious” unspecified functionality in terms of information leakage, and 2) formulating detection as a satisfiability problem that can take advantage of the recent advances in both boolean and satisfiability modulo theory (SMT) solvers. The formulated detection procedure can be applied to a gate-level design using commercial equivalence checking tools, or directly to the Verilog/VHDL code by reasoning about the satisfiability of SMT expressions built from traversing the data-flow graph. We demonstrate the effectiveness of our approach on an adder coprocessor and a UART communication controller infected with Trojans which process information leaked from the on-chip bus during idle cycles using signals with only partially specified behavior. Nicole Fern, Ismail San, Kwang-Ting Cheng |
ASP-DAC | 2 |
| 2017 | A low-area unified hardware architecture for the AES and the cryptographic hash function Grøstl
Nuray At, Jean-Luc Beuchat, Eiji Okamoto, Ismail San, Teppei Yamazaki |
J. Parallel Distributed Comput. | 4 |
| 2017 | Hiding Hardware Trojan Communication Channels in Partially Specified SoC Bus FunctionalityabstractOn-chip bus implementations must be bug-free and secure to provide the functionality and performance required by modern system-on-a-chip (SoC) designs. Regardless of the specific topology and protocol, bus behavior is never fully specified, meaning there exist cycles/conditions where some bus signals are irrelevant, and ignored by the verification effort. We highlight the susceptibility of current bus implementations to Hardware Trojans hiding in this partially specified behavior, and present a model for creating a covert Trojan communication channel between SoC components for any bus topology and protocol. By only altering existing bus signals during the period where their behaviors are unspecified, the Trojan channel is very difficult to detect. We give Trojan channel circuitry specifics for AMBA AXI4 and advanced peripheral bus (APB), then create a simple system comprised of several master and slave units connected by an AXI4-Lite interconnect to quantify the overhead of the Trojan channel and illustrate the ability of our Trojans to evade a suite of protocol compliance checking assertions from ARM. We also create an SoC design running a multiuser Linux OS to demonstrate how a Trojan communication channel can allow an unprivileged user access to root-user data. We then outline several detection strategies for this class of Hardware Trojan. Nicole Fern, Ismail San, Çetin Kaya Koç, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Hardware Trojans in incompletely specified on-chip bus systems
Nicole Fern, Ismail San, Çetin Kaya Koç, Kwang-Ting Cheng |
DATE | 2 |
| 2016 | Trojans modifying soft-processor instruction sequences embedded in FPGA bitstreamsabstractReconfigurable platforms such as FPGAs and CPLDs are used to implement flexible and lightweight embedded systems often using soft-processors and a fixed instruction sequence stored in block memories. The bitstream format is proprietary for most vendors, however, in this work we demonstrate how to identify and extract block memory contents within the bitstream, allowing an adversary to learn and possibly modify the fixed instruction sequence. Manipulating the instruction sequence by inserting a Trojan in the bitstream as opposed to in the RTL code allows an adversary to bypass many verification steps. Moreover, the proposed Trojans only add extra instructions to the sequence to leak secret information, and do not change the original program behavior, making them virtually impossible to detect using functional tests. We present a case study where a Trojan is injected into a MIPS AES encryption program to leak internal state information by adding extra instructions from the available ones without changing the original program behavior. Ismail San, Nicole Fern, Çetin Kaya Koç, Kwang-Ting Cheng |
FPL | 1 |
| 2016 | Efficient paillier cryptoprocessor for privacy-preserving data miningabstractAbstract Paillier cryptosystem is extensively utilized as a homomorphic encryption scheme to ensure privacy requirements in many privacy‐preserving data mining schemes. However, overall performance of the applications employing Paillier cryptosystem intrinsically degrades because of modular multiplications and exponentiation operations performed by the cryptosystem. In this study, we investigate how to tackle with such performance degradation because of Paillier cryptosystem. We first exploit parallelism among the operations in the cryptosystem and interleaving among independent operations. Then, we develop hardware realization of our scheme using field‐programmable gate arrays. As a case study, we evaluate our cryptoprocessor for a well‐known privacy‐preserving set intersection protocol. We demonstrate how the proposed cryptoprocessor responds promising performance for hard real‐time privacy‐preserving data mining applications. Copyright © 2016 John Wiley & Sons, Ltd. Ismail San, Nuray At, Ibrahim Yakut, Huseyin Polat 0001 |
Secur. Commun. Networks | 1 |
| 2014 | Improving the computational efficiency of modular operations for embedded systems
Ismail San, Nuray At |
J. Syst. Archit. | 1 |
| 2012 | On Increasing the Computational Efficiency of Long Integer Multiplication on FPGAabstractThis paper presents a compact hardware architecture for long integer multiplication and proposes a strategy to increase the computational efficiency of the Karatsuba algorithm on FPGA. The presented architecture aims to provide an efficient and compact architecture to be used where long integer multiplication is definitely required such as Cryptography, especially Public Key Cryptography (PKC), Coding theory, DSP and many more. There are several studies in the literature related to increase the efficiency of multiplication, especially in public key cryptography. From our point of view, the main advantage of this method over other existing methods is that recursive utilization of hardware resources with tight scheduling brings better performance with smaller logic area. Our coprocessor is also suitable for multiplications of polynomials in GF(p) and GF(2k). Our method achieves highest available frequency of FPGA. We compare our hardware performance figures for different bit width multiplication with other reported studies. The results show that our architecture combines performance with small area size. Ismail San, Nuray At |
TrustCom | 1 |
| 2011 | Compact Hardware Architecture for Hummingbird Cryptographic AlgorithmabstractHummingbird is an ultra-lightweight cryptographic algorithm aiming at resource-constrained devices. In this paper, we present an enhanced hardware implementation of the Hummingbird cryptographic algorithm that is based on the memory blocks embedded within Spartan-3 FPGAs. The enhancement is not only from the introduction of the coprocessor approach but also from the employment of serialized data processing principles. Due to the compactness of the proposed architecture, remaining reconfigurable area in FPGAs can be used for other purposes. Comparisons to the other reported FPGA implementation of the Hummingbird cryptographic algorithm indicate that the proposed architecture outperforms the previous work in terms of both efficiency and area. We remark that our architecture can also be used as stand-alone although it is built via coprocessor approach. Ismail San, Nuray At |
FPL | 1 |