VLDB 2026 Research / reviewers in the wild / expert
Kris Gaj
dblp:02/1286
· DBLP profile ↗
80ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0002-5050-8748ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 64 · 8 first-author · 9 since 2021Security and privacy · 16 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lightweight Champions of the World: Side-Channel Resistant Open Hardware for Finalists in the NIST Lightweight Cryptography Standardization ProcessabstractCryptographic competitions have played a significant role in stimulating the development and release of open hardware for cryptography. The primary reason was the focus of standardization organizations and other contest organizers on transparency and fairness of hardware benchmarking, which could be achieved only with all source code made available for public scrutiny. Consequently, the number and quality of open-source hardware implementations developed during subsequent major competitions, such as AES, SHA-3, and CAESAR, have steadily increased. However, most of these implementations were still quite far from being used in future products due to the lack of countermeasures against side-channel analysis (SCA). In this article, we discuss the first coordinated effort at developing SCA-resistant open hardware for all finalists of a cryptographic standardization process. The developed hardware is then evaluated by independent labs for information leakage and resilience to selected attacks. Our target included the 10 finalists of the NIST lightweight cryptography standardization process. The authors’ contributions included formulating detailed requirements, publicizing the submissions, matching open hardware with suitable SCA-evaluation labs, developing a subset of all implementations, serving as one of the six evaluation labs, performing field-programmable gate array benchmarking of all protected and unprotected implementations, and summarizing results in the comprehensive report. Our results confirm that NIST made the right decision in selecting Ascon as a future lightweight cryptography standard. They also indicate that at least three other algorithms, Xoodyak, TinyJAMBU, and ISAP, were very strong competitors and outperformed Ascon in at least one of the evaluated performance metrics. Kamyar Mohajerani, Luke Beckwith, Abubakr Abdulgadir, Jens-Peter Kaps, Kris Gaj |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2023 | A High-Performance Hardware Implementation of the LESS Digital Signature Scheme
Luke Beckwith, Robert Wallace, Kamyar Mohajerani, Kris Gaj |
PQCrypto | 4 |
| 2023 | High-Speed Hardware Architectures and FPGA Benchmarking of CRYSTALS-Kyber, NTRU, and SaberabstractPost-Quantum Cryptography (PQC) has emerged as a response of the cryptographic community to the danger of attacks performed using quantum computers. All PQC schemes can be implemented in software and hardware using conventional (non-quantum) computing systems. PQC is the biggest revolution in cryptography since the invention of public-key schemes in the mid-1970 s. Lattice-based key exchange schemes have emerged as leading candidates in the NIST PQC standardization process due to their relatively short public keys and ciphertexts. This paper presents novel high-speed hardware architectures for four lattice-based Key Encapsulation Mechanisms (KEMs) representing three NIST PQC finalists: NTRU (with two distinct variants, NTRU-HPS and NTRU-HRSS), CRYSTALS-Kyber, and Saber. We benchmark these candidates in terms of their performance and resource utilization in today's FPGAs. Our best architectures outperform the best designs from other groups reported to date in terms of the area-time product by factors ranging from 1.01 to 2.88, depending on the algorithm and security level. Additionally, our study demonstrates that CRYSTALS-Kyber and Saber have very similar hardware performance. Both outperform NTRU in terms of execution time by a factor 36-62 for key generation and 3-7 for decapsulation, assuming the same security level. Viet Ba Dang, Kamyar Mohajerani, Kris Gaj |
IEEE Trans. Computers | 3 |
| 2023 | Engineering Practical Rank-Code-Based Cryptographic Schemes on Embedded Hardware. A Case Study on ROLLOabstractIn this paper, we investigate the practical performance of rank-code based cryptography on FPGA platforms by presenting a case study on the quantum-safe KEM scheme based on LRPC codes called ROLLO, which was among NIST post-quantum cryptography standardization round-2 candidates. Specifically, we present an FPGA implementation of the encapsulation and decapsulation operations of the ROLLO KEM scheme with some variations to the original specification. The design is fully parameterized, using code-generation scripts to support a wide range of parameter choices for security levels specified in ROLLO. At the core of the ROLLO hardware, we presented a generic approach for hardware-based Gaussian elimination, which can process both non-singular and singular matrices. Previous works on hardware-based Gaussian elimination can only process non-singular ones. However, a plethora of cryptosystems, for instance, quantum-safe key encapsulation mechanisms based on rank-metric codes, ROLLO and RQC, which are among NIST post-quantum cryptography standardization round-2 candidates, require performing Gaussian elimination for random matrices regardless of the singularity. To the best of our knowledge, this work is the first hardware implementation for rank-code-based cryptographic schemes. The experimental results suggest rank-code-based schemes can be highly efficient. Jingwei Hu 0001, Wen Wang 0007, Kris Gaj, Huaxiong Wang |
IEEE Trans. Computers | 3 |
| 2022 | Session details: Session 1A: Hardware SecurityabstractNo abstract available. Kris Gaj |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | Session details: Session 2A: Hardware SecurityabstractNo abstract available. Kris Gaj |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | Session details: Session 5A: Hardware SecurityabstractNo abstract available. Kris Gaj |
ACM Great Lakes Symposium on VLSI | 1 |
| 2021 | Hardware Benchmarking of Round 2 Candidates in the NIST Lightweight Cryptography Standardization ProcessabstractTwenty five Round 2 candidates in the NIST Lightweight Cryptography (LWC) process have been implemented in hardware by groups from all over the world. All implementations compliant with the LWC Hardware API, proposed in 2019, have been submitted for hardware benchmarking to George Mason University's LWC benchmarking team. The received submissions were first verified for correct functionality and compliance with the hardware API's specification. Then, the execution times in clock cycles, as a function of input sizes, have been determined using behavioral simulation. The compatibility of all implementations with FPGA toolsets from three major vendors, Xilinx, Intel, and Lattice Semiconductor was verified. Optimized values of the maximum clock frequency and resource utilization metrics, such as the number of look-up tables (LUTs) and flip-flops (FFs), were obtained by running optimization tools, such as Minerva, ATHENa, and Xeda. The raw post-place and route results were then converted into values of the corresponding throughputs for long, medium-size, and short inputs. The results were presented in the form of easy to interpret graphs and tables, demonstrating the relative performance of all investigated algorithms. An effort was made to make the entire process as transparent as possible and results easily reproducible by other groups. Kamyar Mohajerani, Richard Haeussler, Rishub Nagpal, Farnoud Farahmand, Abubakr Abdulgadir, Jens-Peter Kaps, Kris Gaj |
DATE | 7 |
| 2021 | High-Performance Hardware Implementation of CRYSTALS-DilithiumabstractMany currently deployed public-key cryptosystems are based on the difficulty of the discrete logarithm and integer factorization problems. However, given an adequately sized quantum computer, these problems can be solved in polynomial time as a function of the key size. Due to the future threat of quantum computing to current cryptographic standards, alternative algorithms that remain secure under quantum computing are being evaluated for future use. One such algorithm is CRYSTALS-Dilithium, a lattice-based digital signature scheme, which is a finalist in the NIST Post Quantum Cryptography (PQC) competition. As a part of this evaluation, high-performance implementations of these algorithms must be investigated. This work presents a high-performance implementation of CRYSTALS-Dilithium targeting FPGAs. In particular, we present a design that achieves the best latency for an FPGA implementation to date. We also compare our results with the most-relevant previous work on hardware implementations of NIST Round 3 post-quantum digital signature candidates. Luke Beckwith, Duc Tri Nguyen, Kris Gaj |
FPT | 3 |
| 2021 | Side-channel Resistant Implementations of a Novel Lightweight Authenticated Cipher with Application to Hardware SecurityabstractLightweight authenticated ciphers are crucial in many resource-constrained applications, including hardware security. To protect Intellectual Property (IPs) from theft and reverse-engineering, multiple obfuscation methods have been developed. An essential component of such schemes is the need for secrecy and authenticity of the obfuscation keys. Such keys may need to be exchanged through the unprotected channels, and their recovery attempted using side-channel attacks. However, the use of the current AES-GCM standard to protected key exchange requires a substantial area and power overhead. NIST is currently coordinating a standardization process to select lightweight algorithms for resource-constrained applications. Although security against cryptanalysis is paramount, cost, performance, and resistance to side-channel attacks are among the most important selection criteria. Since the cost of protection against side-channel attacks is a function of the algorithm, quantifying this cost is necessary for estimating its cost and performance in real-world applications. In this work, we investigate side-channel resistant lightweight implementations of an authenticated cipher TinyJAMBU, one of ten finalists in the current NIST LWC standardization process. Our results demonstrate that these implementations achieve robust security against side-channel attacks while keeping the area and power consumption significantly lower than it is possible using the current standards. Abubakr Abdulgadir, Sammy Lin, Farnoud Farahmand, Jens-Peter Kaps, Kris Gaj |
ACM Great Lakes Symposium on VLSI | 5 |
| 2021 | Fast NEON-Based Multiplication for Lattice-Based NIST Post-quantum Cryptography Finalists
Duc Tri Nguyen, Kris Gaj |
PQCrypto | 2 |
| 2020 | Sampling from Discrete Distributions in Combinational Hardware with Application to Post-Quantum CryptographyabstractRandom values from discrete distributions are typically generated from uniformly-random samples. A common technique is to use a cumulative distribution table (CDT) lookup for inversion sampling, but it is also possible to use Boolean functions to map a uniformly-random bit sequence into a value from a discrete distribution. This work presents a methodology for deriving such functions for any discrete distribution, encoding them in VHDL for implementation in combinational hardware, and (for moderate precision and sample space size) confirming the correctness of the produced distribution. The process is demonstrated using a discrete Gaussian distribution with a small sample space, but it is applicable to any discrete distribution with fixed parameters. Results are presented for sampling schemes from several submissions to the NIST PQC standardization process, comparing this method to CDT lookups on a Xilinx Artix-7 FPGA. The process produces compact solutions for distributions up to moderate size and precision. Michael X. Lyons, Kris Gaj |
DATE | 2 |
| 2020 | A Multiplatform Parallel Approach for Lattice Sieving Algorithms
Michal Andrzejczak, Kris Gaj |
ICA3PP (1) | 2 |
| 2020 | Special Session: The Recent Advance in Hardware Implementation of Post-Quantum CryptographyabstractThe recent advancement in quantum technology has initiated a new round of cryptosystem innovation, i.e., the emergence of Post-Quantum Cryptography (PQC). This new class of cryptographic schemes is intended to be mathematically resistant against any known attacks using quantum computers, but, at the same time, be fully implementable using traditional semiconductor technology. The National Institutes of Standards and Technology (NIST) has already started the PQC standardization process, and the initial pool of 69 submissions has been reduced to 26 Round 2 candidates. Echoing the pace of the PQC "revolution," this paper gives a detailed and thorough introduction to recent advances in the hardware implementation of PQC schemes, including challenges, new implementation methods, and novel hardware architectures. Specifically, we have: (i) described the challenges and rewards of implementing PQC in hardware; (ii) presented the novel methodology for the design-space exploration of PQC implementations using high-level synthesis (HLS); (iii) introduced a new underexplored PQC scheme (binary Ring-Learning-with-Errors), as well as its novel hardware implementation for possible lightweight applications. The overall content delivered by this paper could serve multiple purposes: (i) provide useful references for the potential learners and the interested public; (ii) introduce new areas and directions for potential research to the VTS community; (iii) facilitate the PQC standardization process and the exploration of related new ways of implementing cryptography in existing and emerging applications. Jiafeng Xie, Kanad Basu, Kris Gaj, Ujjwal Guin |
VTS | 3 |
| 2019 | Software/Hardware Codesign of the Post Quantum Cryptography Algorithm NTRUEncrypt Using High-Level Synthesis and Register-Transfer Level Design MethodologiesabstractWhen quantum computers become scalable and reliable, they are likely to break all public-key cryptography standards, such as RSA and Elliptic Curve Cryptography. The projected threat of quantum computers has led the U.S. National Institute of Standards and Technology (NIST) to an effort aimed at replacing existing public-key cryptography standards with new quantum-resistant alternatives. In December 2017, 69 candidates were accepted by NIST to Round 1 of the NIST Post-Quantum Cryptography (PQC) standardization process. NTRUEncrypt is one of the most well-known PQC algorithms that has withstood cryptanalysis. The speed of NTRUEncrypt in software, especially on embedded software platforms, is limited by the long execution time of its primary operation, polynomial multiplication. In this paper, we investigate speeding up NTRUEncrypt using software/hardware codesign on a Xilinx Zynq UltraScale+ multiprocessor system-on-chip (MPSoC). Polynomial multiplication is implemented in the Programmable Logic (PL) of Zynq using two approaches: traditional Register-Transfer Level (RTL) and High-Level Synthesis (HLS). The remaining operations of NTRUEncrypt are executed in software on the Processing System (PS) of Zynq, using the bare-metal mode. The speed-up of our software/hardware codesigns vs. purely software implementations is determined experimentally and analyzed in the paper. The results are reported for the RTL-based and HLS-based hardware accelerators, and compared to the best available software implementation, included in the NIST submission package. The speed-ups for encryption were 2.4 and 3.9, depending on the selected parameter set. For decryption, the corresponding speed-ups were 4.0 and 6.8. In addition, for the polynomial multiplication operation itself, the speed up was in excess of 75. Our code for the NTRUEncrypt polynomial multiplier accelerator is being made open-source for further evaluation on multiple software/hardware platforms. Farnoud Farahmand, Duc Tri Nguyen, Viet Ba Dang, Ahmed Ferozpuri, Kris Gaj |
FPL | 5 |
| 2019 | Evaluating the Potential for Hardware Acceleration of Four NTRU-Based Key Encapsulation Mechanisms Using Software/Hardware Codesign
Farnoud Farahmand, Viet Ba Dang, Duc Tri Nguyen, Kris Gaj |
PQCrypto | 4 |
| 2019 | COMA: Communication and Obfuscation Management Architecture
Kimia Zamiri Azar, Farnoud Farahmand, Hadi Mardani Kamali, Shervin Roshanisefat, Houman Homayoun, William Diehl, Kris Gaj, Avesta Sasan |
RAID | 7 |
| 2018 | Improved Lightweight Implementations of CAESAR Authenticated CiphersabstractAuthenticated ciphers offer potential benefits to resource-constrained devices in the Internet of Things (IoT). The CAESAR competition seeks optimal authenticated ciphers based on several criteria, including performance in resource-constrained (i.e., low-area, low-power, and low-energy) hardware. Although the competition specified a "lightweight"? use case for Round 3, most hardware submissions to Round 3 were not lightweight implementations, in that they employed architectures optimized for best throughput-to-area (TP/A) ratio, and used the Pre- and PostProcessor modules from the CAESAR Hardware (HW) Development Package designed for highspeed applications. In this research, we provide true lightweight implementations of selected ciphers (ACORN, NORX, CLOC-AES, SILC-AES, and SILC-LED). These implementations use an improved version of the CAESAR HW Development Package designed for lightweight applications, and are fully compliant with the CAESAR HW Application Programming Interface for Authenticated Ciphers. Our lightweight implementations achieve an average of 55% reduction in area and 40% reduction in power compared to their corresponding high-speed versions. Although the average energy per bit of lightweight ciphers increases by a factor of 3.6, the lightweight version of NORX actually uses 47% less energy per bit than its corresponding high-speed implementation. Farnoud Farahmand, William Diehl, Abubakr Abdulgadir, Jens-Peter Kaps, Kris Gaj |
FCCM | 5 |
| 2018 | Face-off Between the CAESAR Lightweight Finalists: ACORN vs. AsconabstractAuthenticated ciphers potentially provide resource savings and security improvements over the joint use of secret-key ciphers and message authentication codes. The CAESAR competition aims to choose the most suitable authenticated ciphers for several categories of applications, including a lightweight use case, for which the primary criteria are performance in resource-constrained devices, and ease of protection against side channel attacks (SCA). In March 2018, two of the candidates from this category, ACORN and Ascon, were selected as CAESAR contest finalists. In this research, we compare two SCA-resistant FPGA implementations of ACORN and Ascon, where one set of implementations has area consumption nearly equivalent to the defacto standard AES-GCM, and the other set has throughput (TP) close to that of AES-GCM. The results show that protected implementations of ACORN and Ascon, with area consumption less than but close to AES-GCM, have 23.3 and 2.5 times, respectively, the TP of AES-GCM. Likewise, implementations of ACORN and Ascon with TP greater than but close to AES-GCM, consume 18% and 74% of the area, respectively, of AES-GCM. Farnoud Farahmand, Abubakr Abdulgadir, Jens-Peter Kaps, Kris Gaj |
FPT | 4 |
| 2018 | A High-Speed Constant-Time Hardware Implementation of NTRUEncrypt SVESabstractIn this paper, we present a high-speed constant-time hardware implementation of NTRUEncrypt Short Vector Encryption Scheme (SVES), fully compliant with the IEEE 1363.1 Standard Specification for Public Key Cryptographic Techniques Based on Hard Problems over Lattices. Our implementation follows an earlier proposed Post-Quantum Cryptography (PQC) Hardware Application Programming Interface (API), which facilitates its fair comparison with implementations of other PQC schemes. The paper contains the detailed flow and block diagrams, timing analysis, as well as results in terms of latency (in clock cycles), maximum clock frequency, and resource utilization in modern high-performance Field Programmable Gate Arrays (FPGAs). Our design takes full advantage of the ability to parallelize the major operation of NTRU, polynomial multiplication, in hardware. As a result, the execution time bottleneck shifts to the hash function, SHA-256, which is sequential in nature and as a result cannot be easily sped up in hardware. The obtained FPGA results for NTRU Encrypt SVES are compared with the equivalent results for Classic McEliece, a competing, well-established Post-Quantum Cryptography encryption scheme, with a long history of unsuccessful attempts at breaking. Our code for NTRUEncrypt SVES is being made open-source to speed-up further design-space exploration and benchmarking on multiple hardware platforms. Farnoud Farahmand, Umar Sharif, Kevin Briggs, Kris Gaj |
FPT | 4 |
| 2018 | Challenges and Rewards of Implementing and Benchmarking Post-Quantum Cryptography in HardwareabstractPractical quantum computers have been recently selected as one of 10 breakthrough technologies of 2017 by the MIT Technology Review. Although various fields of human activity, such as chemistry, medicine, and materials science, are likely to be dramatically affected by practical quantum computers, the most likely immediate impact will take place in the area of cryptography and cyber security. As a result of this potential threat, a new field of science has emerged, called Post-Quantum Cryptography (PQC). PQC is devoted to the design and analysis of cryptographic algorithms that are resistant against any known attacks using quantum computers, but by themselves can be implemented using classical computing platforms, based on traditional modern semiconductor technologies. In this paper, we provide an overview and motivation for the PQC, NIST Standardization Effort, cryptographic competitions, and hardware benchmarking of candidates in cryptographic contests. Five major families of PQC schemes, code-, hash-, isogeny-, lattice-, and multivariate-based, are shortly introduced. The challenges of fair and comprehensive hardware benchmarking of PQC submissions are highlighted, together with the possible ways of overcoming these difficulties, such as the use of a common API, development packages, specialized libraries, and high-level synthesis. Kris Gaj |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Benchmarking the Capabilities and Limitations of SAT Solvers in Defeating Obfuscation SchemesabstractIn this paper, we investigate the strength of six different SAT solvers in attacking various obfuscation schemes. Our investigation revealed that Glucose and Lingeling SAT solvers are generally suited for attacking small-to-midsize obfuscated circuits, while the MapleGlucose, if the system is not memory bound, is best suited for attacking mid-to-difficult obfuscation methods. Our experimental result indicates that when dealing with extremely large circuits and very difficult oufuscation problems, the SAT solver may be memory bound, and Lingeling, for having the most memory efficient implementation, is the best suited solver for such problems. Additionally, our investigation revealed that SAT solver execution times may vary widely across different SAT solvers. Hence, when testing the hardness of an obfuscation methods, although the increase in difficulty could be verified by one SAT solver, the pace of increase in difficulty is dependent on the choice of a SAT solver. Shervin Roshanisefat, Harshith K. Thirumala, Kris Gaj, Houman Homayoun, Avesta Sasan |
IOLTS | 3 |
| 2017 | Analysis and Inner-Round Pipelined Implementation of Selected Parallelizable CAESAR Competition CandidatesabstractIn this paper, we have first characterized candidates of the Competition for Authenticated Encryption, Security, Applicability, and Robustness (CAESAR) from the point of view of their suitability for parallel processing of multiple blocks of associated data, message, and ciphertext. Then, we have chosen seven candidates from the Round 2 and Round 3 submissions, namely SCREAM, AES-COPA, Minalpher, OCB, AES-OTR, COLM, and Deoxys. We first obtained the initial estimates of the maximum clock frequency, throughput, area, and critical path for the high-speed Basic Iterative Architecture of each of the above candidates. Then, we implemented a two-stage inner-round pipelining for all the aforementioned algorithms in order to improve the frequency and throughput by reducing the critical path and processing multiple blocks of data simultaneously. We targeted the largest available FPGA in the student version of Xilinx ISE, i.e., Xilinx Virtex 6 XC6VLX75T-3FF784. Our results have demonstrated the improvement in the clock frequency and throughput by a factor varying from x1.28 for OCB to x1.84 for SCREAM, and the change in the throughput to area ratio (with area expressed using LUTs) by a factor varying from x0.93 for Minalpher to x1.72 for SCREAM. Sanjay Deshpande, Kris Gaj |
DSD | 2 |
| 2017 | Comparison of hardware and software implementations of selected lightweight block ciphersabstractLightweight block ciphers are an important topic of research in the context of the Internet of Things (IoT). Current cryptographic contests and standardization efforts seek to benchmark lightweight ciphers in both hardware and software. Although there have been several benchmarking studies of both hardware and software implementations of lightweight ciphers, direct comparison of hardware and software implementations is difficult due to differences in metrics, measures of effectiveness, and implementation platforms. In this research, we facilitate this comparison by use of a custom lightweight reconfigurable processor. We implement six ciphers, AES, SIMON, SPECK, PRESENT, LED and TWINE, in hardware using register transfer level (RTL) design, and in software using the custom reconfigurable processor. Both hardware and software implementations are instantiated in identical Xilinx Kintex-7 FPGAs, which enables direct comparison of throughput, area, throughput-to-area (TP/A) ratio, power, and energy. Results show that TWINE and AES have the highest TP/A ratios for hardware and software implementations, respectively, assuming an area target of 300–450 LUTs. In terms of direct comparison, software implementations on tailored reconfigurable processers generally use less power — especially where reconfigurable instruction set extensions are permitted. However, custom hardware implementations have higher throughput and energy-efficiency than software implementations on the same platform. William Diehl, Farnoud Farahmand, Panasayya Yalla, Jens-Peter Kaps, Kris Gaj |
FPL | 5 |
| 2017 | Comparing the cost of protecting selected lightweight block ciphers against differential power analysis in low-cost FPGAsabstractLightweight block ciphers are an important topic in the Internet of Things (IoT), since they provide moderate security, while requiring fewer resources than AES. Ongoing cryptographic contests and standardization efforts evaluate lightweight block ciphers on their resistance to power analysis side channel attack (SCA), and the ability to apply countermeasures. While some ciphers have been individually evaluated, a large scale comparison of resistance to side channel attack and formulation of the relative cost of implementing countermeasures is difficult, since researchers typically use varied architectures, optimization strategies, technologies, and evaluation techniques. In this research we leverage the t-test leakage detection methodology and an open-source side channel analysis suite (FOBOS) to compare FPGA implementations of AES, SIMON, SPECK, PRESENT, LED, and TWINE, using a choice of architecture targeted to optimize throughput-to-area (TP/A) ratio, for resistance to differential power analysis (DPA). We then apply an equivalent level of protection to the above ciphers using 3-share threshold implementations (TI), and verify improved resistance to DPA. We find that SIMON has the highest TP/A ratio of protected versions, followed by PRESENT, TWINE, LED, AES, and SPECK. However, PRESENT uses the least energy in terms of nJ-per-bit. William Diehl, Abubakr Abdulgadir, Jens-Peter Kaps, Kris Gaj |
FPT | 4 |
| 2017 | Toward a new HLS-based methodology for FPGA benchmarking of candidates in cryptographic competitions: The CAESAR contest case studyabstractThe increasing number of candidates competing in cryptographic contests has made hardware benchmarking using the traditional Register-Transfer Level (RTL) methodology too inefficient and potentially unfair, especially at the early stages of the competitions. In this paper, we propose supplementing, and eventually replacing, this traditional RTL methodology with the use of High-Level Synthesis (HLS) tools. We apply our proposed HLS-based approach to FPGA benchmarking in the ongoing CAESAR contest, by comparing and ranking 16 authenticated ciphers, including the current standard, AES-GCM, and the primary variants of 13 Round 3 CAESAR candidates. After a careful survey of available HLS tools, we chose Xilinx Vivado HLS as our primary benchmarking tool. Our study has demonstrated high correlation between the rankings of the evaluated algorithms, obtained using both investigated methodologies. In particular, after applying HLS, the algorithm rankings in terms of two major performance metrics — throughput and throughput to area ratio — have either remained unchanged or have been affected only for algorithms with very similar RTL performance. Ekawat Homsirikamol, Kris Gaj |
FPT | 2 |
| 2017 | Selection of an error-correcting code for FPGA-based physical unclonable functionsabstractThis paper explores error-correcting codes for fuzzy extractor applications with Physical Unclonable Functions. We investigate BCH codes and compare them to convolutional codes using criteria of remaining entropy, probability of decoder failure, and hardware requirements. Parallel BCH coding is analyzed with a comprehensive search performed to find the smallest BCH code which satisfies the criteria in a parallel design to produce 128, 192, and 256-bit keys. A convolutional code is selected for comparison against the BCH codes found in this analysis. Application of the selected codes to a fuzzy extractor design is analyzed. Hardware requirements for FPGA implementations of each code is compared, with a BCH decoder design implemented for Artix-7 and Spartan-6 FPGA families. We find that a (127, 22, 47) parallel BCH code or (2, 1, 12) convolutional code is capable of performing as well as a single large BCH code, while requiring fewer FPGA resources when block RAMs can be leveraged. The convolutional code additionally requires the least amount of PUF ID bits. Brian Jarvis, Kris Gaj |
FPT | 2 |
| 2016 | Hybrid STT-CMOS designs for reverse-engineering preventionabstractThis paper presents a rigorous step towards design-for-assurance by introducing a new class of logically reconfigurable design resilient to design reverse engineering. Based on the non-volatile spin transfer torque (STT) magnetic technology, we introduce a basic set of non-volatile reconfigurable Look-Up-Table (LUT) logic components (NV-STT-based LUTs). STT-based LUT with significantly different set of characteristics compared to CMOS provides new opportunities to enhance design security yet makes it challenging to remain highly competitive with custom CMOS or even SRAM-based LUT in terms of power, performance and area. To address these challenges, we propose several algorithms to select and replace custom CMOS gates with reconfigurable STT-based LUTs during design implementation such that the functionality of STT-based components and therefore the entire design cannot be determined in any manageable time, rendering any design reverse engineering attack ineffective. Our study conducted on a large number of standard circuit benchmarks concludes significant resiliency of hybrid STT-CMOS circuits against various types of attacks. Furthermore, the selection algorithms on average have a small impact of less than 3%, 8%, and 3% on design parametric constraints including performance, power and area, respectively. Theodore Winograd, Hassan Salmani, Hamid Mahmoodi, Kris Gaj, Houman Homayoun |
DAC | 4 |
| 2016 | RTL Implementations and FPGA Benchmarking of Three Authenticated Ciphers Competing in CAESAR Round TwoabstractAuthenticated ciphers are cryptographic transformations which combine the functionality of confidentiality, integrity, and authentication. This research uses register transfer-level (RTL) design to describe selected authenticated ciphers using a hardware description language (HDL), verifies their proper operation through functional simulation, and implements them on target FPGAs. The authenticated ciphers chosen for this research are the CAESAR Round Two variants of SCREAM, POET, and Minalpher. Ciphers are discussed from an engineering standpoint, and are compared and contrasted in terms of design features. To ensure conformity and standardization in evaluation, all three candidates are implemented with an identical version of the CAESAR Hardware API for authenticated ciphers. Functionally correct implementations of all three ciphers are realized, and results are compared against each other and previous results in terms of throughput, area, and throughput-to-area (T/A) ratio. SCREAM is found to have the highest T/A ratio of these three ciphers in the Virtex-6 FPGA, while Minalpher has the highest T/A ratio in the Virtex-7 FPGA. William Diehl, Kris Gaj |
DSD | 2 |
| 2016 | Implementation of a Boolean Masking Scheme for the SCREAM CipherabstractMasking is a proven countermeasure to protect physical cryptographic implementations against power analysis side-channel attacks, such as differential power analysis (DPA). Boolean masking is one of several types of masking schemes that can be added to a cipher to increase its security. However, implementing a secure and efficient Boolean masking scheme across all components of a cipher, including non-linear transformations, can be challenging. In this research, a 1st order Boolean masking scheme is applied to the SCREAM authenticated cipher, a CAESAR Round Two candidate. The non-masked and masked versions of the full authenticated cipher are implemented in the Virtex-6 FPGA and compared in terms of throughput, area, and throughput-to-area (T/A) ratio. The SCREAM block cipher is then compared to a masked version of AES to determine the relative costs of masking among the two ciphers. The results show that the T/A ratio of the masked SCREAM full authenticated cipher is only 50% of the T/A ratio of the non-masked version, and that the masking cost of the SCREAM block cipher is roughly equal to that of an equivalently-masked version of the AES block cipher. William Diehl, Kris Gaj |
DSD | 2 |
| 2016 | High-Speed RTL Implementations and FPGA Benchmarking of Three Authenticated Ciphers Competing in CAESAR Round TwoabstractAuthenticated ciphers are cryptographic transformations which combine the functionality of confidentiality, integrity, and authentication. This research uses register transfer-level (RTL) design to describe selected authenticated ciphers using a hardware description language (HDL), verifies their proper operation through functional simulation, and implements them on target FPGAs -- the Xilinx Virtex-6 and Virtex-7. The authenticated ciphers chosen for this research are the CAESAR Round Two variants of SCREAM, POET, and Minalpher. To ensure standardization in evaluation, all three candidates are implemented with an identical version of a universal hardware API for authenticated ciphers. Results are compared against each other in terms of performance, area, and throughput-to-area (TP/A) ratio. SCREAM is found to have the highest TP/A ratio of these three ciphers. William Diehl, Kris Gaj |
FCCM | 2 |
| 2016 | Hardware-software codesign of RSA for optimal performance vs. flexibility trade-offabstractPublic-key cryptosystems such as RSA have been widely used to secure digital data in many commercial systems. Modular arithmetic on large operands used during modular exponentiation makes RSA computationally challenging. Traditionally, software implementations of these algorithms provided the highest flexibility but lacked performance. On the contrary, custom hardware accelerators provided the highest performance but lacked flexibility and adaptability to changing algorithms, parameters, and key sizes. In this paper, we present a hardware/software codesign of RSA cryptosystem that improves performance, while retaining flexibility. We adopted Xilinx Zynq-7000 SoC platform, which integrates a dual-core ARM Cortex-A9 processing system along with Xilinx programmable logic. The software part of our implementation is based on RELIC library (Efficient Library for Cryptography). The performance vs. flexibility trade-off is investigated, and the speed-up of our codesign implementation vs. the purely software implementation of RSA on the same platform is reported. Our results show a speedup of up to 57 times when compared with the software implementation for 2048-bit operand size. We also propose a generic model for HW/SW codesign focused on flexibility with comparable performance to existing HW/SW implementations. Umar Sharif, Rabia Shahid, Kris Gaj, Marcin Rogawski |
FPL | 3 |
| 2015 | Using Facebook for Image SteganographyabstractBecause Facebook is available on hundreds of millions of desktop and mobile computing platforms around the world and because it is available on many different kinds of platforms (from desktops and laptops running Windows, Unix, or OS X to hand held devices running iOS, Android, or Windows Phone), it would seem to be the perfect place to conduct steganography. On Facebook, information hidden in image files will be further obscured within the millions of pictures and other images posted and transmitted daily. Facebook is known to alter and compress uploaded images so they use minimum space and bandwidth when displayed on Facebook pages. The compression process generally disrupts attempts to use Facebook for image steganography. This paper explores a method to minimize the disruption so JPEG images can be used as steganography carriers on Facebook. Jason Hiney, Tejas Dakve, Krzysztof Szczypiorski, Kris Gaj |
ARES | 4 |
| 2014 | ICEPOLE: High-Speed, Hardware-Oriented Authenticated Encryption
Pawel Morawiecki, Kris Gaj, Ekawat Homsirikamol, Krystian Matusiewicz, Josef Pieprzyk, Marcin Rogawski, Marian Srebrny, Marcin Wójcik |
CHES | 2 |
| 2014 | A novel modular adder for one thousand bits and more using fast carry chains of modern FPGAsabstractIn this paper a novel, low-latency family of high-radix Parallel Prefix Network adders and modular adders has been proposed. This family efficiently takes advantage of fast carry chains of modern FPGAs. The implementation results reveal that these adders have great potential for efficient implementation of modular addition with the long integers used in various public key cryptography schemes. Marcin Rogawski, Ekawat Homsirikamol, Kris Gaj |
FPL | 3 |
| 2013 | FPGA PUF Based on Programmable LUT DelaysabstractStrong and efficient techniques are required for chip authentication and secret key generation by integrated circuits (IC). This paper presents a novel approach toward an FPGA friendly Ring Oscillator (RO) based Physical Unclonable Function (PUF). In this design the internal variations of FPGA Look-Up Tables are exploited to generate a PUF response. Statistical tests were performed to study the strength of this PUF. Moreover, stability is compared with the state of the art reported in literature to date. Our design has been tested on 31 Spartan-3e devices and the results are promising with inter-device Hamming distance of 48.3%, Uniformity 50.13%, Bit-aliasing 51.8%, Reliability 97.88%, and Steadiness 99.5%. Furthermore, we also analyzed the frequencies to extract the random variation offered by our design. Bilal Habib, Kris Gaj, Jens-Peter Kaps |
DSD | 2 |
| 2012 | A High-Speed Unified Hardware Architecture for AES and the SHA-3 Candidate GrøstlabstractThe NIST competition for developing the new cryptographic hash standard SHA-3 is currently in the third round. One of the five remaining candidates, Grøstl, is inspired by the Advanced Encryption Standard. This unique feature can be exploited in a large variety of practical applications. In order to have a better picture of the Grøstl-AES computational efficiency (high-level scheduling, internal pipelining, resource sharing, etc.), we designed a high-speed coprocessor for Grøstl-based HMAC and AES in the counter mode. This coprocessor offers high-speed computations of both authentication and encryption with relatively small penalty in terms of area and speed when compared to the authentication (original Grøstl circuitry) functionality only. From our perspective, the main advantage of Grøstl over other finalists is the fact that its hardware hardware architecture naturally accommodates AES at the cost of a small area overhead. Marcin Rogawski, Kris Gaj |
DSD | 2 |
| 2012 | Option space exploration using distributed computing for efficient benchmarking of FPGA cryptographic modulesabstractBenchmarking of digital designs targeting FPGAs is a time intensive and challenging process. Benchmarking results depend on a myriad of variables beyond the properties inherent to the designs being evaluated, encompassing the tools, tool options, FPGA families, and languages used. In this paper we will be discussing enhancements made to the ATHENa benchmarking tool to utilize distributed computing as well as option space exploration techniques to increase the efficiency of the pre-existing process. The capabilities of our environment are demonstrated using two example designs from the SHA-3 cryptographic hash function competition, BLAKE and JH. Benjamin Y. Brewster, Ekawat Homsirikamol, Rajesh Velegalati, Kris Gaj |
FPT | 4 |
| 2011 | Throughput vs. Area Trade-offs in High-Speed Architectures of Five Round 3 SHA-3 Candidates Implemented Using Xilinx and Altera FPGAs
Ekawat Homsirikamol, Marcin Rogawski, Kris Gaj |
CHES | 3 |
| 2011 | Cryptographic Contests: Toward Fair and Comprehensive Benchmarking of Cryptographic Algorithms in Hardware (Abstract)abstractA fair comparison of functionally equivalent digital systems is a challenging and non-trivial task. Objective difficulties include lack of standard interfaces, influence of tools and their options, the dependence of the obtained results on the time spent for optimization, etc. In cryptography, there is a strong need for such a fair evaluation, associated with the way new cryptographic standards are being developed, namely through open competitions of algorithms submitted by research groups from all over the world. Such competitions included for example the AES contest in the U.S., the NESSIE and eSTREAM competitions in Europe, and the CRYPTREC project in Japan. At this point, the focus of attention of the entire cryptographic community is on the SHA-3 contest for a new cryptographic hash function standard, organized by NIST. In this talk I will analyze typical evaluation pitfalls and objective challenges facing the evaluators of cryptographic algorithms from the point of view of performance in hardware. I will present practical benchmarking methodologies and tools that facilitate overcoming these difficulties, and can be used to fairly compare competing algorithms, hardware architectures, development platforms, languages, and tools. I will also discuss the remaining challenges and difficulties worth exploring in the future, and the ways of generalizing experiences gained from benchmarking cryptographic hardware, and applying them to other domains, such as communications and digital signal processing. Kris Gaj |
DSD | 1 |
| 2011 | A Configurable Ring-Oscillator-Based PUF for Xilinx FPGAsabstractDevadas has first proposed the notion of Silicon Physical Unclonable Function (sPUF), which takes advantage of delay variations of wires and gates. A Ring-Oscillator-Based PUF (RO PUF) is one possible implementation of an sPUF. One disadvantage of RO PUFs is that they require one pair of ring oscillators per bit of output. Therefore, in order to collect enough output bits for a safe security level, a large number of ring oscillators is needed. Configurable PUFs may help solving this problem. In 2009, Maiti introduced a configurable RO PUF to improve RO PUF reliability, where each RO is implemented in one configurable logic block (CLB) by using lookup tables (LUTs) and dedicated multiplexers. In this paper we analyze Maiti's configurable RO PUFs and propose improvements to generate more output bits, by utilizing latches as well as the resource mentioned above. Experimental results demonstrate that our improved method outputs more bits than Maiti's configurable RO PUFs and the original RO PUFs, while using the same amount of area. Jens-Peter Kaps, Kris Gaj |
DSD | 3 |
| 2011 | Use of embedded FPGA resources in implementations of 14 round 2 SHA-3 candidatesabstractIn this paper, we present results of a comprehensive study devoted to the optimization of FPGA implementations of modern cryptographic hash functions using embedded FPGA resources, such as Digital Signal Processing (DSP) units and Block Memories. Fifteen hash functions, including the current American hash standard SHA-2 and 14 candidates for the new hash standard SHA-3, have been included in our investigation. Our methodology involves implementing, characterizing, and comparing all algorithms with a focus on minimizing the amount of reconfigurable logic resources, and achieving a better balance between the use of reconfigurable logic resources and embedded resources in four FPGA families, representing major low-cost and high-performance families of Xilinx and Altera. Rabia Shahid, Umar Sharif, Marcin Rogawski, Kris Gaj |
FPT | 4 |
| 2011 | Hardware architectures for algebra, cryptology, and number theory
Kris Gaj, Rainer Steinwandt |
Integr. | 1 |
| 2011 | New Hardware Architectures for Montgomery Modular Multiplication AlgorithmabstractMontgomery modular multiplication is one of the fundamental operations used in cryptographic algorithms, such as RSA and Elliptic Curve Cryptosystems. At CHES 1999, Tenca and Koç proposed the Multiple-Word Radix-2 Montgomery Multiplication (MWR2MM) algorithm and introduced a now-classic architecture for implementing Montgomery multiplication in hardware. With parameters optimized for minimum latency, this architecture performs a single Montgomery multiplication in approximately 2n clock cycles, where n is the size of operands in bits. In this paper, we propose two new hardware architectures that are able to perform the same operation in approximately n clock cycles with almost the same clock period. These two architectures are based on precomputing partial results using two possible assumptions regarding the most significant bit of the previous word. These two architectures outperform the original architecture of Tenca and Koç in terms of the product latency times area by 23 and 50 percent, respectively, for several most common operand sizes used in cryptography. The architecture in radix-2 can be extended to the case of radix-4, while preserving a factor of two speedup over the corresponding radix-4 design by Tenca, Todorov, and Koç from CHES 2001. Our optimization has been verified by modeling it using Verilog-HDL, implementing it on Xilinx Virtex-II 6000 FPGA, and experimentally testing it using SRC-6 reconfigurable computer. Miaoqing Huang, Kris Gaj, Tarek A. El-Ghazawi |
IEEE Trans. Computers | 2 |
| 2010 | Fair and Comprehensive Methodology for Comparing Hardware Performance of Fourteen Round Two SHA-3 Candidates Using FPGAs
Kris Gaj, Ekawat Homsirikamol, Marcin Rogawski |
CHES | 1 |
| 2010 | ATHENa - Automated Tool for Hardware EvaluatioN: Toward Fair and Comprehensive Benchmarking of Cryptographic Hardware Using FPGAsabstractA fair comparison of functionally equivalent digital system designs targeting FPGAs is a challenging and time consuming task. The results of the comparison depend on the inherent properties of competing algorithms, as well as on selected hardware architectures, implementation techniques, FPGA families, languages and tools. In this paper, we introduce an open-source environment, called ATHENa for fair, comprehensive, automated, and collaborative hardware benchmarking of algorithms belonging to the same class. As our first goal, we select the benchmarking of algorithms belonging to the area of cryptography. Algorithms from this area have been shown to achieve significant speed-ups and security gains compared to software when implemented in FPGAs. The capabilities of our environment are demonstrated using three examples: two different hardware architectures of the current cryptographic hash function standard, SHA-256, and one architecture of a candidate for the new standard, Fugue. All source codes, testbenches, and configuration files necessary to repeat experiments described in this paper are made available through the project web site. Kris Gaj, Jens-Peter Kaps, Venkata Amirineni, Marcin Rogawski, Ekawat Homsirikamol, Benjamin Y. Brewster |
FPL | 1 |
| 2010 | Area-Time Efficient Implementation of the Elliptic Curve Method of Factoring in Reconfigurable Hardware for Application in the Number Field SieveabstractA novel portable hardware architecture of the Elliptic Curve Method of factoring, designed and optimized for application in the relation collection step of the Number Field Sieve, is described and analyzed. A comparison with an earlier proof-of-concept design by Pelzl et al. has been performed, and a substantial improvement has been demonstrated in terms of both the execution time and the area-time product. The ECM architecture has been ported across five different families of FPGA devices in order to select the family with the best performance to cost ratio. A timing comparison with the highly optimized software implementation, GMP-ECM, has been performed. Our results indicate that low-cost families of FPGAs, such as Spartan-3 and Spartan-3E, offer at least an order of magnitude improvement over the same generation of microprocessors in terms of the performance to cost ratio, without the use of embedded FPGA resources, such as embedded multipliers. Kris Gaj, Soonhak Kwon, Patrick Baier, Paul Kohlbrenner, Hoang Le, Mohammed Khaleeluddin, Ramakrishna Bachimanchi, Marcin Rogawski |
IEEE Trans. Computers | 1 |
| 2009 | Reconfigurable Computing Approach for Tate Pairing Cryptosystems over Binary FieldsabstractTate-pairing-based cryptosystems, because of their ability to be used in multiparty identity-based key management schemes, have recently emerged as an alternative to traditional public key cryptosystems. Due to the inherent parallelism of the existing pairing algorithms, high performance can be achieved via hardware realizations. Three schemes for Tate pairing computations have been proposed in the literature: cubic elliptic, binary elliptic, and binary hyperelliptic. In this paper, we propose a new FPGA-based architecture of the Tate-pairing-based computation over binary fields. Even though our field sizes are larger than in the architectures based on cubic elliptic curves or binary hyperelliptic curves with the same security strength, nevertheless fewer multiplications in the underlying field need to be performed. As a result, the computational latency for a pairing computation has been reduced, and our implementation runs 2-20 times faster than the equivalent implementations of other pairing-based schemes at the same level of security strength. Furthermore, we ported our pairing designs for eight field sizes ranging from 239 to 557 bits to the reconfigurable computer, SGI Altix 4700 supported by Silicon Graphics, Inc., and performance and cost are demonstrated. Chang Shu 0003, Soonhak Kwon, Kris Gaj |
IEEE Trans. Computers | 3 |
| 2008 | Memory security management for reconfigurable embedded systemsabstractThe constrained operating environments of many FPGA-based embedded systems require flexible security that can be configured to minimize the impact on FPGA area and power consumption. In this paper, a security approach for external memory in FPGA-based embedded systems that exploits FPGA configurability is presented. Our FPGA-based security core provides both confidentiality and integrity for data stored externally to an FPGA which is accessed by a processor on the FPGA chip. The benefits of our security core are demonstrated using four embedded applications implemented on a Stratix II device. Each application requires a collection of tasks with varying memory security requirements. Our security core is used in conjunction with a NIOS II soft processor running the MicroC/OS II operating system. An average memory and energy savings of about 64%and 16%, respectively, is achieved for the four applications versus a non-configurable, uniform security approach. Romain Vaslin, Guy Gogniat, Jean-Philippe Diguet, Russell Tessier, Deepak Unnikrishnan, Kris Gaj |
FPT | 6 |
| 2008 | Portable library development for reconfigurable computing systems: A case study
Proshanta Saha, Esam El-Araby, Miaoqing Huang, Mohamed Taher, Sergio López-Buedo, Tarek A. El-Ghazawi, Chang Shu 0003, Kris Gaj, Alan Michalski, Duncan A. Buell |
Parallel Comput. | 8 |
| 2006 | Implementing the Elliptic Curve Method of Factoring in Reconfigurable Hardware
Kris Gaj, Soonhak Kwon, Patrick Baier, Paul Kohlbrenner, Hoang Le, Mohammed Khaleeluddin, Ramakrishna Bachimanchi |
CHES | 1 |
| 2006 | FPGA accelerated tate pairing based cryptosystems over binary fieldsabstractTate pairing based cryptosystems have recently emerged as an alternative to traditional public key cryptosystems because of their ability to be used in multi-party identity-based key management schemes. Due to the inherent parallelism of the existing pairing algorithms, high performance can be achieved via hardware realizations. Three schemes for Tate pairing computations have been proposed in the literature: cubic elliptic, binary elliptic, and binary hyperelliptic. For our implementation we have chosen the binary elliptic case because of the simple underlying algorithms and efficient binary arithmetic. In this paper, we propose a new FPGA-based architecture of the Tate pairing-based computation over the binary fields F2239and F2283. Even though our field sizes are larger than in the architectures based on cubic elliptic curves or binary hyperelliptic curves with the same security strength, nevertheless fewer multiplications in the underlying field need to performed. As a result, the computational latency for a pairing computation has been reduced, and our implementation runs 10-to-20 times faster than the equivalent implementations of other pairing-based schemes at the same level of security strength. At the same time, an improvement in the product of latency by area by a factor between 12 and 46 for an equivalent type of implementation has been achieved Chang Shu 0003, Soonhak Kwon, Kris Gaj |
FPT | 3 |
| 2006 | M03 - Reconfigurable supercomputingabstractThe synergistic advances in high-performance computing and reconfigurable computing, based on field programmable gate arrays (FPGAs), has resulted in hybrid parallel systems of microprocessors and FPGAs. Such systems support both fine-grain and coarse-grain parallelism, and can dynamically tune their architecture to fit various applications. Programming these systems can be quite challenging as programming of FPGA devices can involve hardware design. This tutorial will introduce the field of reconfigurable supercomputing and its advances in systems, programming, applications and tools. Reconfigurable system developments at SRC, Cray, SGI, and Star Bridge will be highlighted, and case studies including full application developments will be presented along with the live demonstrations. This tutorial will be the first to show scalability studies for real-life applications over entire HPRC systems. This will reveal the tremendous promise held by this class of architectures in performance, power and cost improvements. Challenges that remain will be also discussed. Tarek A. El-Ghazawi, Duncan A. Buell, Volodymyr V. Kindratenko, Kris Gaj |
SC | 4 |
| 2005 | Reconfigurable computers: an empirical analysis (abstract only)abstractReconfigurable Computers are parallel systems that are designed around multiple general-purpose processors and multiple field programmable gate array (FPGA) chips. These systems can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In this work we conduct an experimental study using one of the state-of-the-art reconfigurable computers and a representative set of applications to assess the field, uncover the challenges, propose solutions, and conceive a realistic evolution path. We consider issues of concern including performance/cost. We also consider productivity in the sense of development, compiling, running, and system reliability. It will be shown that for some applications, the performance/cost can be orders of magnitude better than conventional computers. It will be also shown that programming such machines may still require some hardware knowledge, similar to hardware knowledge computer programmers must acquire to write scalable programs. Tarek A. El-Ghazawi, Kris Gaj, Nikitas A. Alexandridis, Allen Michalski, Osman Devrim Fidanci, Mohamed Taher, Esam El-Araby, Esmail Chitalwala, Proshanta Saha |
FPGA | 2 |
| 2005 | Image processing library for reconfigurable computers (abstract only)abstractReconfigurable Computers (RCs) are parallel systems that are designed around multiple general-purpose processors and multiple field programmable gate array (FPGA) chips. These systems can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. RCs have proposed very high processing capabilities for computationally intensive applications such as Image Processing. This is due to the inherently parallel operation paradigm of the FPGA hardware.In this paper we present the design and implementation of image processing kernels for RCs. This library of kernels have been tested and verified for performance on one of the state-of-the-art reconfigurable computers, SRC-6E. This paper shows that RCs are between 8 to 400 times faster than comparable Pentiums for image based tasks. Mohamed Taher, Esam El-Araby, Tarek A. El-Ghazawi, Kris Gaj |
FPGA | 4 |
| 2005 | High-Throughput Reconfigurable Computing: A Design Study of an IDEA Encryption Cryptosystem on the SRC-6e Reconfigurable ComputerabstractThe combination of traditional microprocessors workstations and hardware-reconfigurable field programmable gate arrays (FPGAs) has developed a new class of workstations known as reconfigurable computers, with several examples demonstrating significant speedups compared to standalone PC workstations alone. Several platforms implement PC-FPGA communication using common PC peripheral interface buses such as PCI-X. A new approach from SRC Computers implements a highspeed communication interface that increases the throughput compared to PCI interfaces. This paper demonstrates an efficient high-throughput implementation of IDEA encryption using the SRC platform. SRC design choices that influence both throughput and area are evaluated. Detailed analyses of FPGA resource utilizations, data transfer and reconfiguration overheads for the SRC system are provided, and a comparison between SRC and a public domain software implementation of IDEA are given. Allen Michalski, Kris Gaj, Duncan A. Buell |
FPL | 2 |
| 2005 | A System-Level Design Methodology for Reconfigurable Computing Applications
Esam El-Araby, Tarek A. El-Ghazawi, Kris Gaj |
FPT | 3 |
| 2005 | Implementation of EAX Mode of Operation for FPGA Bitstream Encryption and Authentication
Milind M. Parelkar, Kris Gaj |
FPT | 2 |
| 2005 | Low Latency Elliptic Curve Cryptography Accelerators for NIST Curves Over Binary Fields
Chang Shu 0003, Kris Gaj, Tarek A. El-Ghazawi |
FPT | 2 |
| 2005 | Secure Partial Reconfiguration of FPGAs
Amir Sheikh Zeineddini, Kris Gaj |
FPT | 2 |
| 2004 | Efficient Linear Array for Multiplication in GF(2m) Using a Normal Basis for Elliptic Curve Cryptography
Soonhak Kwon, Kris Gaj, Chang Hoon Kim, Chun Pyo Hong |
CHES | 2 |
| 2004 | A 1 Gbit/s Partially Unrolled Architecture of Hash Functions SHA-1 and SHA-512
Roar Lien, Tim Grembowski, Kris Gaj |
CT-RSA | 3 |
| 2004 | Implementation of elliptic curve cryptosystems over GF(2n) in optimal normal basis on a reconfigurable computerabstractDuring the last few years, a considerable effort has been devoted to the development of reconfigurable computers, machines that are based on the close interoperation of traditional microprocessors and Field Programmable Gate Arrays. Several prototype machines of this type have been designed, and demonstrated significant speed-ups compared to conventional workstations for computationally intensive problems, such as codebreaking. In this paper, we demonstrate an efficient implementation of Elliptic Curve scalar multiplication over GF(2 n ) in Optimal Normal Basis, using one of the leading reconfigurable computers available on the market, SRC-6E. We show how the hardware architecture and programming model of this reconfigurable computer has influenced the choice of the optimum program partitioning scheme. The detailed analysis of the control, data transfer, and reconfiguration overheads is given in the paper. The end-to-end speed-ups in the range from 895 to 1300 compared to the microprocessor implementation are demonstrated depending on the chosen partitioning scheme. Sashisu Bajracharya, Chang Shu 0003, Kris Gaj, Tarek A. El-Ghazawi |
FPGA | 3 |
| 2004 | An embedded true random number generator for FPGAsabstractField Programmable Gate Arrays (FPGAs) are an increasingly popular choice of platform for the implementation of cryptographic systems. Until recently, designers using FPGAs had less than optimal choices for a source of truly random bits. In this paper we extend a technique that uses on-chip jitter and PLLs to a much larger class of FPGAs that do not contain PLLs. Our design uses only the Configurable Logic Blocks (CLBs) common to all FPGAs, and has a self-testing capability. Using the intrinsic jitter contained in digital circuits, we produce random bits at speeds of up to 0.5 Mbits/second with good statistical characteristics. We discuss the engineering challenges of extracting random bits from digital circuits, and we report the results of running standard statistical tests (NIST) on the output generated by our system. Paul Kohlbrenner, Kris Gaj |
FPGA | 2 |
| 2004 | Implementation of Elliptic Curve Cryptosystems over GF(2n) in Optimal Normal Basis on a Reconfigurable Computer
Sashisu Bajracharya, Chang Shu 0003, Kris Gaj, Tarek A. El-Ghazawi |
FPL | 3 |
| 2004 | Reconfigurable hardware implementation of mesh routing in number field sieve factorizationabstractFactorization of large numbers has been a constant source of interest in cryptanalysis. The fastest known algorithm for factoring large numbers is the number field sieve (NFS). The two most time consuming phases of NFS are sieving and matrix step. We propose an efficient way of implementing the matrix step in reconfigurable hardware. Our solution is based on the mesh-routing method proposed by Lenstra et al. We determine the practical size of a partial mesh that can fit in one FFGA device, Xilinx Virtex II XC2V6000. We further extrapolate the computation time for the case of a square systolic array of FFGAs for 512-bit and 1024-bit numbers' factorization. We demonstrate that for practical sizes of numbers used in cryptography, 1024 bits, the matrix step of factorization can be performed using 1024 Virtex II FFGAs in less than 40 days. Sashisu Bajracharya, Deapesh Misra, Kris Gaj, Tarek A. El-Ghazawi |
FPT | 3 |
| 2004 | Effective system and performance benchmarking for reconfigurable computersabstractApplications running on a reconfigurable computer can be divided into two major categories: computationally intensive and input/output intensive. In the first case, the input and output are limited, and therefore the performance of the reconfigurable computer depends primarily on the power of the FPGAs, and the capability to exploit parallelism available in a given application. In the second case, the execution time is dominated by input/output, and therefore, an application cannot process data faster than the speed of its slowest input/output channel. The focus of This work is on developing micro-benchmarks to characterize the behavior of various communication channels within reconfigurable computers for the second class of applications. The paper defines a system of 'paper and pencil' micro-benchmarks for the measurement of maximum throughput and minimum latency in the communication between various components of a generic reconfigurable system. The results help to dynamically characterize a reconfigurable machine and the SRC 6E reconfigurable computer is used as a test case to validate the proposed model. Esmail Chitalwala, Tarek A. El-Ghazawi, Kris Gaj, Nikitas A. Alexandridis, Daniel S. Poznanovic |
FPT | 3 |
| 2004 | Wavelet spectral dimension reduction of hyperspectral imagery on a reconfigurable computerabstractHyperspectral imagery, by definition, provides valuable remote sensing observations at hundreds of frequency bands. Conventional image classification (interpretation) methods may not be used without dimension reduction preprocessing. Automatic wavelet reduction has been proven to yield better or comparable classification accuracy, while achieving substantial computational savings. However, the large hyperspectral data volumes remain to present a challenge for traditional processing techniques. Reconfigurable computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. We investigate the potential of using RCs for on-board, i.e. aboard airborne/spaceborne carriers, preprocessing of hyperspectral imagery by prototyping for the first time the automatic wavelet dimension reduction algorithm. Our investigation exploits the fine and coarse grain parallelism provided by the RCs and has been experimentally verified on one of the state-of the art reconfigurable platforms, SRC-6E. An order of magnitude speedup over traditional processing techniques has been reported. Esam El-Araby, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart, Kris Gaj |
FPT | 4 |
| 2004 | System-Level Parallelism and Throughput Optimization in Designing Reconfigurable Computing ApplicationsabstractSummary form only given. Reconfigurable computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In a large class of applications, the total I/O time is comparable or even greater than the computations time. As a result, the rate of the DMA transfer between the microprocessor memory and the on-board memory of the FPGA-based processor becomes the performance bottleneck. We perform a theoretical and experimental study of this specific performance limitation. The mathematical formulation of the problem has been experimentally verified on the state-of-the art reconfigurable platform, SRC-6E. We demonstrate and quantify the possible solution to this problem that exploits the system-level parallelism within reconfigurable machines. Esam El-Araby, Mohamed Taher, Kris Gaj, Tarek A. El-Ghazawi, David Caliga, Nikitas A. Alexandridis |
IPDPS | 3 |
| 2004 | A performance study of job management systemsabstractAbstract Job Management Systems (JMSs) efficiently schedule and monitor jobs in parallel and distributed computing environments. Therefore, they are critical for improving the utilization of expensive resources in high‐performance computing systems and centers, and an important component of Grid software infrastructure. With many JMSs available commercially and in the public domain, it is difficult to choose an optimum JMS for a given computing environment. In this paper, we present the results of the first empirical study of JMSs reported in the literature. Four commonly used systems, LSF, PBS Pro, Sun Grid Engine/CODINE, and Condor were considered. The study has revealed important strengths and weaknesses of these JMSs under different operational conditions. For example, LSF was shown to exhibit excellent throughput for a wide range of job types and submission rates. Alternatively, CODINE appeared to outperform other systems in terms of the average turn‐around time for small jobs, and PBS appeared to excel in terms of turn‐around time for relatively larger jobs. Copyright © 2004 John Wiley & Sons, Ltd. Tarek A. El-Ghazawi, Kris Gaj, Nikitas A. Alexandridis, Frederic Vroman, Nguyen Nguyen 0003, Jacek R. Radzikowski, Preeyapong Samipagdi, Suboh A. Suboh |
Concurr. Pract. Exp. | 2 |
| 2003 | Very Compact FPGA Implementation of the AES Algorithm
Pawel Chodowiec, Kris Gaj |
CHES | 2 |
| 2003 | Facts and Myths of Enigma: Breaking Stereotypes
Kris Gaj, Arkadiusz Orlowski |
EUROCRYPT | 1 |
| 2003 | IPsec-Protected Transport of HDTV over IP
Peter Bellows, Jaroslav Flidr, Ladan Gharai, Colin Perkins, Pawel Chodowiec, Kris Gaj |
FPL | 6 |
| 2003 | An Implementation Comparison of an IDEA Encryption Cryptosystem on Two General-Purpose Reconfigurable Computers
Allen Michalski, Kris Gaj, Tarek A. El-Ghazawi |
FPL | 2 |
| 2003 | Exploiting system-level parallelism in the application development on a reconfigurable computerabstractReconfigurable Computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In a large class of applications, the total I/O time is comparable or even greater than the computations time. As a result, the rate of the DMA transfer between the microprocessor memory and the on-board memory becomes the performance bottleneck even on RCs. In this paper, we perform a theoretical and experimental study of this specific performance limitation for the state-of-the art reconfigurable platform, SRC-6E. We demonstrate and quantify the possible solution to this problem that exploits the system-level parallelism within the reconfigurable machine. Esam El-Araby, Mohamed Taher, Kris Gaj, Tarek A. El-Ghazawi, David Caliga, Nikitas A. Alexandridis |
FPT | 3 |
| 2003 | Implementation of Elliptic Curve Cryptosystems on a reconfigurable computerabstractDuring the last few years, a considerable effort has been devoted to the development of reconfigurable computers, machines that are based on the close interoperation of traditional microprocessors and Field Programmable Gate Arrays (FPGAs). Several prototype machines of this type have been designed, and demonstrated significant speedups compared to conventional workstations for computationally intensive problems, such as codebreaking. Nevertheless, the efficient use and programming of such machines is still an unresolved problem. In this paper, we demonstrate an efficient implementation of an Elliptic Curve scalar multiplication over GF(2/sup m/), using one of the leading reconfigurable computers available on the market, SRC-6E. We show how the hardware architecture and programming model of this reconfigurable computer has influenced the choice of the algorithm partitioning strategy for this application. A detailed analysis of the control, data transfer, and reconfiguration overheads is given in the paper, together with the performance comparison of our implementation against an optimized microprocessor implementation. Nghi Nguyen, Kris Gaj, David Caliga, Tarek A. El-Ghazawi |
FPT | 2 |
| 2002 | Comparative Analysis of the Hardware Implementations of Hash Functions SHA-1 and SHA-512
Tim Grembowski, Roar Lien, Kris Gaj, Nghi Nguyen, Peter Bellows, Jaroslav Flidr, Tom Lehman, Brian Schott |
ISC | 3 |
| 2001 | Fast Implementation and Fair Comparison of the Final Candidates for Advanced Encryption Standard Using Field Programmable Gate Arrays
Kris Gaj, Pawel Chodowiec |
CT-RSA | 1 |
| 2001 | Fast implementations of secret-key block ciphers using mixed inner- and outer-round pipeliningabstractThe new design methodology for secret-key block ciphers, based on introducing an optimum number of pipeline stages inside of a cipher round is presented and evaluated. This methodology is applied to five well-known modern ciphers, Triple DES, Rijndael, RC6, Serpent, and Twofish, with the goal to first obtain the architecture with the optimum throughput to area ratio, and then the architecture with the highest possible throughput. All ciphers are modeled in VHDL, and implemented using Xilinx Virtex FPGA devices. It is demonstrated that all investigated ciphers can operate with similar maximum clock frequencies, in the range from 95 to 131 MHz, limited only by the delay of a single CLB layer and delays of interconnects. Rijndael, RC6, Twofish, and Serpent achieve throughputs in the range from 12.1 Gbit/s to 16.8 Gbit/s; and Triple DES achieves the throughput of 7.5 Gbit/s. Because of the optimum speed to cost ratio, the proposed architecture seems to be very well suited for practical implementations of secret-key block ciphers using both FPGAs and custom ASICs. We also show that using this architecture for comparing hardware performance of secret-key block ciphers, such as AES candidates, operating in non-feedback cipher modes, leads to the more prudent and fairer analysis than comparisons based on other types of pipelined architectures. Pawel Chodowiec, Po Khuon, Kris Gaj |
FPGA | 3 |
| 2001 | Experimental Testing of the Gigabit IPSec-Compliant Implementations of Rijndael and Triple DES Using SLAAC-1V FPGA Accelerator Board
Pawel Chodowiec, Kris Gaj, Peter Bellows, Brian Schott |
ISC | 2 |