EDBT 2026 Demo / reviewers in the wild / expert
Hwajeong Seo
dblp:121/4554
· DBLP profile ↗
48ranked-venue papers
14as first author
9since 2021 · last 2025
0000-0003-0069-9061ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 23 · 8 first-author · 1 since 2021Systems, architecture and hardware · 17 · 6 first-author · 6 since 2021Theory of computation · 3 · 1 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Implementations of AIMer Post-Quantum Signature Scheme for Low-End to High-End IoT DevicesabstractPost-quantum cryptography (PQC) is becoming increasingly critical for securing Internet of Things (IoT) applications against potential threats posed by quantum computers. However, cryptographic schemes based on mathematically hard problems, such as lattice-based constructions, may face reduced security margins due to ongoing cryptanalytic advancements. In IoT use cases requiring long-term deployment without guaranteed secure updates, symmetric-based signature schemes, such as$\mathsf {SPHINCS+}$, whose security relies solely on symmetric primitives, are considered robust alternatives. Among these,$\textsf {AIMer}$, selected as a finalist in the KpqC competition, is emerging as a promising candidate. This article presents optimized software implementations of$\textsf {AIMer}$for various IoT platforms, ranging from low-end to high-end devices. We propose memory-optimized and time–memory tradeoff implementations for resource-constrained ARM Cortex-M4 devices, and single instruction multiple data (SIMD)-accelerated implementations for AArch64 and AVX2 architectures. Furthermore, we conduct a comprehensive performance evaluation comparing our optimized implementations with other PQC signature schemes. Our results show that$\textsf {AIMer}$achieves significantly smaller signature sizes and faster key generation and signing compared to$\mathsf {SPHINCS+}$across all security levels, making it a viable option for long-term IoT deployment. Jihoon Kwon, Sangyub Lee 0002, ByeongHak Lee, Hwajeong Seo |
IEEE Internet Things J. | 4 |
| 2025 | Optimizing AES-GCM on 32-Bit ARM Cortex-M4 Microcontrollers: Fixslicing and FACE-Based ApproachabstractAdvanced Encryption Standard (AES) in Galois/Counter Mode (GCM) delivers both confidentiality and integrity, yet poses performance and security challenges on resource-limited microcontrollers. In this article, we present an optimized AES-GCM implementation for the 32-bit ARM Cortex-M4 that combines the Fixslicing AES approach with the FACE (Fast AES-CTR Encryption) strategy, significantly reducing redundant computations in AES-CTR. We further examine two GHASH implementations, a 4-bit table-based approach and a Karatsuba-based constant-time variant, to balance speed, memory usage, and resistance to timing attacks. Our evaluations on an STM32F4 microcontroller show that the Fixslicing and FACE method reduces the AES-128 GCTR cycle counts by up to 19.41%, while the Table-based GHASH achieves nearly double the speed of its Karatsuba counterpart. These results confirm that with the right mix of bit-slicing optimizations, counter-mode caching, and lightweight polynomial multiplication, secure and efficient AES-GCM can be obtained even on low-power embedded devices. Hwajeong Seo |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Optimized Quantum Circuit Implementation of Payoff FunctionabstractLarge-scale quantum computers that can execute practical quantum algorithms have the potential to solve complex problems that are currently challenging for classical computers. This involves converting these problems into a form that can be processed by quantum circuits, a crucial process that requires minimizing quantum resources like qubit count, gate count, and circuit depth. Our work focuses on implementing and optimizing the foundational task of quantum finance, known as option pricing, as a quantum circuit. This enables the utilization of quantum computing benefits, within the financial domain. Specifically, we implement and optimize the function fK(S) = max(S−K, 0). Taking into consideration the significant trade-offs between qubit count and circuit depth, we have developed quantum circuits for the optimized implementation of the fK(S). Our work incorporates various optimization techniques for the circuit, such as selecting the optimal adder, optimizing the S−K operation, parallelization, and qubit reuse. Furthermore, we offer various versions of our quantum circuits for the fK(S), each featuring different adders and Toffoli decompositions, thereby providing flexibility for a wide range of use cases. Sejin Lim, Kyungbae Jang, Anubhab Baksi, Anupam Chattopadhyay, Hwajeong Seo |
VLSI-SoC | 7 |
| 2023 | Look-up the Rainbow: Table-based Implementation of Rainbow Signature on 64-bit ARMv8 ProcessorsabstractThe Rainbow Signature Scheme is one of the finalists in the National Institute of Standards and Technology (NIST) Post-Quantum Cryptography (PQC) standardization competition, but failed to win because it has lack of stability in the parameter selection. It is the only signature candidate based on a multivariate quadratic equation. Rainbow signatures have smaller signature sizes compared with other post-quantum cryptography candidates. However, they require expensive tower-field based polynomial multiplications. In this article, we propose an efficient implementation of Rainbow signatures using a look-up table–based multiplication method. The polynomial multiplications in Rainbow signatures are performed on the 𝔽 16 field, which is divided into sub-fields 𝔽 4 and 𝔽 2 under the tower-field method. To accelerate the multiplication process on target processors, we propose a look-up table–based tower-field multiplication technique. In 𝔽 16 , all values are expressed in 4-bit data format and can be implemented using a 256-byte look-up table access. The implementation uses the TBL and TBX instructions of the 64-bit ARMv8 target processor. For Rainbow III and Rainbow V, they are computed on the 𝔽 256 field using an additional 16-byte table instead of creating a new look-up table. The proposed technique uses the vector registers of 64-bit ARMv8 processors and can calculate 16 result values with a single instruction. We also proposed implementations that are resistant to timing attacks. There are two types of implementations. The first one is the cache side-attack resistant implementation, which utilizes the 128-byte cache lines of the M1 processor. In this implementation, cache misses do not occur, and cache hits always occur. The second type is the constant-time implementation. This method takes a step-by-step approach to finding the required look-up table value and ensures that the same number of accesses is made regardless of which look-up table value is called. This implementation is designed to be constant-time, meaning it does not leak timing information. Our experiments on modern Apple M1 processors showed up to 428.73× and 114.16× better performance for finite field multiplications and Rainbow signatures schemes, respectively, compared with previous reference implementations. To the best of our knowledge, this proposed Rainbow implementation is the first optimized Rainbow implementation for 64-bit ARMv8 processors. Hyeokdong Kwon, Minjoo Sim, Wai-Kong Lee, Hwajeong Seo |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2022 | DPCrypto: Acceleration of Post-Quantum Cryptography Using Dot-Product Instructions on GPUsabstractModern NVIDIA GPU architectures offer dot-product instructions (DP2A and DP4A), with the aim of accelerating machine learning and scientific computing applications. These dot-product instructions allow the computation of multiply-and-add instructions in a single clock cycle, effectively achieving higher throughput compared to conventional 32-bit integer units. In this paper, we show that the dot-product instruction can also be used to accelerate matrix-multiplication and polynomial convolution operations, which are widely used in post-quantum lattice-based cryptographic schemes. In particular, we propose a highly optimized implementation of FrodoKEM wherein the matrix-multiplication is accelerated by the dot-product instruction. We also present specially designed data structures that allow an efficient implementation of Saber key-encapsulation mechanism, utilizing the dot-product instruction to speed-up the polynomial convolution. The proposed FrodoKEM implementation achieves$4.37\times $higher throughput than the state-of-the-art implementation on a V100 GPU. This paper also presents the first implementation of Saber on GPU platforms, achieving 124,418, 120,463, and 31,658 key exchanges per second on RTX3080, V100, and T4 GPUs, respectively. Since matrix-multiplication and polynomial convolution operations are the most time-consuming operations in lattice-based cryptographic schemes, we strongly believe that the proposed methods can be beneficial to other KEM and signatures schemes based on lattices. Wai-Kong Lee, Hwajeong Seo, Seong Oun Hwang, Ramachandra Achar, Angshuman Karmakar, Jose Maria Bermudo Mera |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | External Reviewers ARITH 2021abstractThe conference offers a note of thanks and lists its reviewers. Karim Bigou, Mojtaba Bisheh-Niasar, Luís Fiolhais, Rogerio Paludo, Hwajeong Seo |
ARITH | 5 |
| 2021 | Kyber on ARM64: Compact Implementations of Kyber on 64-Bit ARM Cortex-A Processors
Pakize Sanal, Emrah Karagoz, Hwajeong Seo, Reza Azarderakhsh, Mehran Mozaffari Kermani |
SecureComm (2) | 3 |
| 2021 | Supersingular Isogeny Key Encapsulation (SIKE) Round 2 on ARM Cortex-M4abstractWe present the first practical software implementation of Supersingular Isogeny Key Encapsulation (SIKE) round 2, targeting NIST's 1, 2, 3, and 5 security levels on 32-bit ARM Cortex-M4 microcontrollers. The proposed library introduces a new speed record of all SIKE Round 2 protocols with reasonable memory consumption on the low-end target platform. We achieved this record by adopting several state-of-the-art engineering techniques as well as highly-optimized hand-crafted assembly implementation of finite field arithmetic. In particular, we carefully redesign the previous optimized implementations of finite field arithmetic on the 32-bit ARM Cortex-M4 platform and propose a set of novel techniques which are explicitly suitable for SIKE primes. The benchmark result on STM32F4 Discovery board equipped with 32-bit ARM Cortex-M4 microcontrollers shows that entire key encapsulation and decapsultation over SIKEp434 take about 184 million clock cycles (i.e., 1.09 seconds @168 MHz). In contrast to the previous optimized implementation of the isogeny-based key exchange on low-end 32-bit ARM Cortex-M4, our performance evaluation shows feasibility of using SIKE mechanism on the low-end platform. In comparison to the most of the post-quantum candidates, SIKE requires an excessive number of arithmetic operations, resulting in significantly slower timings. However, its small key size makes this scheme as a promising candidate on low-end microcontrollers in the quantum era by ensuring the lower energy consumption for key transmission than other schemes. Hwajeong Seo, Mila Anastasova, Amir Jalali, Reza Azarderakhsh |
IEEE Trans. Computers | 1 |
| 2021 | SIKE in 32-bit ARM Processors Based on Redundant Number System for NIST Level-IIabstractWe present an optimized implementation of the post-quantum Supersingular Isogeny Key Encapsulation (SIKE) for 32-bit ARMv7-A processors supporting NEON engine (i.e., SIMD instruction). Unlike previous SIKE implementations, finite field arithmetic is efficiently implemented in a redundant representation, which avoids carry propagation and pipeline stall. Furthermore, we adopted several state-of-the-art engineering techniques as well as hand-crafted assembly implementation for high performance. Optimized implementations are ported to Microsoft SIKE library written in “a non-redundant representation” and evaluated in high-end 32-bit ARMv7-A processors, such as ARM Cortex-A5, A7, and A15. A full key-exchange execution of SIKEp503 is performed in about 109 million cycles on ARM Cortex-A15 processors (i.e., 54.5 ms @2.0 GHz), which is about 1.58× faster than previous state-of-the-art work presented in CHES’18. Hwajeong Seo, Pakize Sanal, Reza Azarderakhsh |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2020 | A High-Speed Public-Key Signature Scheme for 8-b IoT-Constrained DevicesabstractMore than 98% of all microprocessors sold worldwide are used in embedded devices, and will continue to accelerate due to the emerging applications in Internet of Things (IoT). Cryptographic algorithms based on multivariate quadratic (MQ) equations are suitable for low-cost IoT-constrained devices since they require only modest computational resources. In this article, we describe the design and implementations of a new public-key signature scheme based on MQ equations highly optimized for practicability. The novel scheme is obtained as a result of the combination of solvable systems of quadratic equations and sparse polynomials. We provide a security analysis of our scheme against known algebraic attacks and derive a concrete parameter. Performance of our scheme on an 8-b AVR microprocessor is the fastest among known signature schemes: signing of our scheme is about 36.4× faster than that of ECDSA-256 (NIST P-256), the most widely deployed international standard signature scheme. The signing of our scheme is about 8.6× and 11.0× faster than those of Rainbow and BLISS-BI, respectively. The secret size of our scheme has reduced by a factor of 87% compared to Rainbow. Our scheme requires signatures of 80 B which is comparable to ECDSA-256 of 64 B. We also implement a protected version of our scheme on the 8-b microprocessor to prevent the current side-channel attacks presented in CHES 2018. Even our protected version is the fastest singing among the known signature schemes. Kyung-Ah Shim, Cheol-Min Park, Namhun Koo, Hwajeong Seo |
IEEE Internet Things J. | 4 |
| 2020 | Efficient Software Implementation of Ring-LWE Encryption on IoT ProcessorsabstractEmbedded processors have been widely used for building up Internet of Things (IoT) platforms, in which the security issue is becoming critical. This paper studies efficient techniques of lattice-based cryptography on these processors and presents the first implementation of ring-LWE encryption on ARM NEON and MSP430 architectures. For ARM NEON architecture, we propose a vectorized version of Iterative Number Theoretic Transform (NTT) for high-speed computation of polynomial multiplication on ARM NEON platforms and a 32-bit variant of SAMS2 technique for fast reduction. For MSP430 architecture, we propose an optimized SWAMS2 reduction technique, which consists of five different basic operations, including Shifting, Swapping, Addition, and two Multiplication-Subtractions. Regarding of the sampling from the discrete Gaussian distribution, we adopt Knuth-Yao sampler, accompanied with optimized methods such as Look-Up Table (LUT) and byte-scanning. Subsequently, a full-fledged implementation of Ring-LWE is presented by both taking advantage of our proposed method and previous optimization techniques re-designed for desired platforms. Our ring-LWE implementation of encryption/decryption at a classical security level of 128 bits requires only 149:4k=32:8k clock cycles on ARM NEON, and 2126:3k=244:5k clock cycles on MSP430. These results are roughly 7 times faster than the fastest ECC implementation on desired platforms with same security level. Zhe Liu 0001, Reza Azarderakhsh, Howon Kim 0001, Hwajeong Seo |
IEEE Trans. Computers | 4 |
| 2020 | Four$\mathbb {Q}$Q on Embedded Devices with Strong Countermeasures Against Side-Channel AttacksabstractThis work deals with the energy-efficient, high-speed and high-security implementation of elliptic curve scalar multiplication, elliptic curve Diffie-Hellman (ECDH) key exchange and elliptic curve digital signatures on embedded devices using FourQ and incorporating strong countermeasures to thwart a wide variety of side-channel attacks. First, we set new speed records for constant-time curve-based scalar multiplication, DH key exchange and digital signatures at the 128-bit security level with implementations targeting 8, 16 and 32-bit microcontrollers. For example, our software computes a static ECDH shared secret in ~6.9 million cycles (or 0.86 seconds @8 MHz) on a low-power 8-bit AVR microcontroller which, compared to the fastest Curve25519 and genus-2 Kummer implementations on the same platform, offers 2× and 1.4× speedups, respectively. Similarly, it computes the same operation in ~495 thousand cycles on a 32-bit ARM Cortex-M4 microcontroller, achieving a factor-1.9 speedup when compared to the fastest Curve25519 implementation targeting another Cortex-M4 platform. A similar speed performance is observed in the case of digital signatures. Second, we engineer a set of side-channel countermeasures taking advantage of FourQ's rich arithmetic and propose a secure implementation that offers protection against a wide range of sophisticated side-channel attacks, including differential power analysis (DPA). Despite the use of strong countermeasures, the experimental results show that our FourQ software is still efficient enough to outperform implementations of Curve25519 that only protect against timing attacks. Finally, we perform a differential power analysis evaluation of our software running on an ARM Cortex-M4, and report that no leakage was detected with up to 10 million traces. These results demonstrate the potential of deploying FourQ on low-power applications such as protocols forthe Internet of Things. Zhe Liu 0001, Patrick Longa, Geovandro C. C. F. Pereira, Oscar Reparaz, Hwajeong Seo |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2020 | Montgomery Multiplication for Public Key Cryptography on MSP430XabstractFor traditional public key cryptography and post-quantum cryptography, such as elliptic curve cryptography and supersingular isogeny key encapsulation, modular multiplication is the most performance-critical operation among basic arithmetic of these cryptographic schemes. For this reason, the execution timing of such cryptographic schemes, which may highly determine that the service availability for low-end microprocessors (e.g., 8-bit AVR, 16-bit MSP430X, and 32-bit ARM Cortex-M), mainly relies on the efficiency of modular multiplication on target embedded processors. In this article, we present new optimal modular multiplication techniques based on the interleaved Montgomery multiplication on 16-bit MSP430X microprocessors, where the multiplication part is performed in a hardware multiplier and the reduction part is performed in a basic arithmetic logic unit (ALU) with the optimal modular multiplication routine, respectively. This two-step approach is effective for the special modulus of NIST curves, SM2 curves, and supersingular isogeny key encapsulation. We further optimized the Montgomery reduction by using techniques for “Montgomery-friendly” prime. This technique significantly reduces the number of partial products. To demonstrate the superiority of the proposed implementation of Montgomery multiplication, we applied the proposed method to the NIST P-256 curve, of which the implementation improves the previous modular multiplication operation by 23.6% on 16-bit MSP430X microprocessors and to the SM2 curve as well (first implementation on 16-bit MSP430X microcontrollers). Moreover, secure countermeasures against timing attack and simple power analysis are also applied to the scalar multiplication of NIST P-256 and SM2 curves, which achieve the 8,582,338 clock cycles (0.53 seconds@16 MHz) and 10,027,086 clock cycles (0.62 seconds@16 MHz), respectively. The proposed Montgomery multiplication is a generic method that can be applied to other cryptographic schemes and microprocessors with minor modifications. Hwajeong Seo, Kyuhwang An, Hyeokdong Kwon |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2019 | SIKE Round 2 Speed Record on ARM Cortex-M4
Hwajeong Seo, Amir Jalali, Reza Azarderakhsh |
CANS | 1 |
| 2019 | Compact Implementations of HIGHT Block Cipher on IoT PlatformsabstractRecent lightweight block cipher competition (FELICS Triathlon) evaluates efficient implementations of block ciphers for Internet of things (IoT) environment. In the competition, the implementation of HIGHT block cipher achieved the most efficient lightweight block cipher, in terms of code size (ROM), memory (RAM), and execution time. In this paper, we further investigate lightweight features of HIGHT block cipher and present the optimized implementations of both software and hardware for low-end IoT platforms, including resource-constrained devices (8-bit AVR and 32-bit ARM Cortex-M3) and application-specific integrated circuit (ASIC). By using proposed optimization methods, the implemented HIGHT block cipher shows better performance compared to previous state-of-the-art implementations. Bohun Kim, Junghoon Cho, Byungjun Choi, Jongsun Park 0001, Hwajeong Seo |
Secur. Commun. Networks | 5 |
| 2019 | Memory-Efficient Implementation of Elliptic Curve Cryptography for the Internet-of-ThingsabstractIn this paper, we present memory-efficient and scalable implementations of NIST standardized elliptic curves P-256, P-384 and P-521 on three ARMv6-M processors (i.e. Cortex-M0, M0+, and M1). Specifically, we propose a refined approach to perform the Multiply-ACcumulate (MAC) operation using hardware multiplier provided by ARMv6-M processor, and a compact doubling routine for multi-precision squaring that executes both doubling and partial product operations in an efficient way. We demonstrate that the proposed squaring implementation achieves a speed up of 28 percent compared to the same operation employed in Micro-ECC. Then, we reduce one modular reduction in co-Z conjugate point addition by using lazy reduction and special form representation (CD-AB, EF-AB), which further reduces the execution time of both P-256 and P-384 implementations. Finally, we propose scalable implementations of ECC scalar multiplication on ARMv6-M processors that are widely used for Internet of Things applications. Zhe Liu 0001, Hwajeong Seo, Aniello Castiglione, Kim-Kwang Raymond Choo, Howon Kim 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2019 | Lightweight Implementations of NIST P-256 and SM2 ECC on 8-bit Resource-Constraint Embedded DeviceabstractElliptic Curve Cryptography (ECC) now is one of the most important approach to instantiate asymmetric encryption and signature schemes, which has been extensively exploited to protect the security of cyber-physical systems. With the advent of the Internet of Things (IoT), a great deal of constrained devices may require software implementations of ECC operations. Under this circumstances, the SM2, a set of public key cryptographic algorithms based on elliptic curves published by Chinese Commercial Cryptography Administration Office, was standardized at ISO in 2017 to enhance the cyber-security. However, few research works on the implementation of SM2 for constrained devices have been conducted. In this work, we fill this gap and propose our efficient, secure, and compact implementation of scalar multiplication on a 256-bit elliptic curve recommended by the SM2, as well as a comparison implementation of scalar multiplication on the same bit-length elliptic curve recommended by NIST. We re-design some existent techniques to fit the low-end IoT platform, namely 8-bit AVR processors, and our implementations evaluated on the desired platform show that the SM2 algorithms have competitive efficiency and security with NIST, which would work well to secure the IoT world. Lu Zhou 0002, Chunhua Su, Hwajeong Seo |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2019 | IoT-NUMS: Evaluating NUMS Elliptic Curve Cryptography for IoT PlatformsabstractIn 2015, NIST held a workshop calling for new candidates for the next generation of elliptic curves to replace the almost two-decade old NIST curves. Nothing Upon My Sleeves (NUMS) curves are among the potential candidates presented in the workshop. Here, we present the first implementation of the NUMS256, NUMS379, and NUMS384 curves on two types of embedded devices. The implementations, which exhibit regular, constant-time execution to protect against timing and simple side-channel attacks, set new speed records and advance the state-of-the-art of curve-based (without endomorphism) scalar multiplication on 8-bit AVR and 32-bit ARM11 microcontrollers. For example, our NUMS256 implementation computes a scalar multiplication in ~1.4 million cycles on a low-power 32-bit ARM11 microcontroller using mixed C and assembly language. These results demonstrate the potential of deploying IoT-NUMS on constrained and low-power applications such as protocols for the Internet of Things. Zhe Liu 0001, Hwajeong Seo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Hybrid approach of parallel implementation on CPU-GPU for high-speed ECDSA verification
Hwajeong Seo, Hyeokchan Kwon, Hyunsoo Yoon |
J. Supercomput. | 2 |
| 2018 | Secure GCM implementation on AVR
Zhe Liu 0001, Hwajeong Seo, Chien-Ning Chen, Yasuyuki Nogami, Taehwan Park, Jongseok Choi, Howon Kim 0001 |
Discret. Appl. Math. | 2 |
| 2018 | Efficient Parallel Implementation of Matrix Multiplication for Lattice-Based Cryptography on Modern ARM ProcessorabstractRecently, various types of postquantum cryptography algorithms have been proposed for the National Institute of Standards and Technology’s Postquantum Cryptography Standardization competition. Lattice-based cryptography, which is based on Learning with Errors, is based on matrix multiplication. A large-size matrix multiplication requires a long execution time for key generation, encryption, and decryption. In this paper, we propose an efficient parallel implementation of matrix multiplication and vector addition with matrix transpose using ARM NEON instructions on ARM Cortex-A platforms. The proposed method achieves performance enhancements of 36.93%, 6.95%, 32.92%, and 7.66%. The optimized method is applied to the Lizard. CCA key generation step enhances the performance by 7.04%, 3.66%, 7.57%, and 9.32% over previous state-of-the-art implementations. Taehwan Park, Hwajeong Seo, Junsub Kim, Haeryong Park, Howon Kim 0001 |
Secur. Commun. Networks | 2 |
| 2018 | Compact Software Implementation of Public-Key Cryptography on MSP430XabstractOn the low-end embedded processors, the implementations of Elliptic Curve Cryptography (ECC) are considered to be a challenging task due to the limited computation power and storage of the low-end embedded processors. Particularly, the multi-precision multiplication and squaring operations are the most expensive operations for ECC implementations. In order to enhance the performance, many works presented efficient multiplication and squaring routines on the target devices. Recent works show that 128-bit security level ECC is available within a second and this is practically fast enough for IoT services. However, previous approaches missed the other important storage issues (i.e., program size, ROM). Considering that the embedded processors only have a few KB ROM, we need to pay attention to the compact ROM size with reasonable performance. In this article, we present very compact and generic implementations of multiplication and squaring operations on the 16-bit MSP430X processors for the ECC. The implementations utilize the new 32-bit multiplier and advanced multiplication and squaring routines. Since the proposed routines are generic, the arbitrary length of operand is available with high-speed and small code size. With proposed multiplication and squaring routines, we implemented Curve25519 on the MSP430X processors. The scalar multiplication is performed within 6,666,895 clock cycles and 4,054 bytes. Compared with previous works based on the speed-optimized version, our memory-efficient version reduces the code size by 59.8%, sacrificing the execution timing by 20.5%. Hwajeong Seo |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2018 | Compact Implementations of ARX-Based Block Ciphers on IoT ProcessorsabstractIn this article, we present implementations for Addition, Rotation, and eXclusive-or (ARX)-based block ciphers, including LEA and HIGHT, on IoT devices, including 8-bit AVR, 16-bit MSP, 32-bit ARM, and 32-bit ARM-NEON processors. We optimized 32-/8-bitwise ARX operations for LEA and HIGHT block ciphers by considering variations in word size, the number of general purpose registers, and the instruction set of the target IoT devices. Finally, we achieved the most compact implementations of LEA and HIGHT block ciphers. The implementations were fairly evaluated through the Fair Evaluation of Lightweight Cryptographic Systems framework, and implementations won the competitions in the first and the second rounds. Hwajeong Seo, Ilwoong Jeong, Jung-Keun Lee, Woo-Hwan Kim |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2018 | Secure IoT framework and 2D architecture for End-To-End security
Jongseok Choi, Youngjin In, Changjun Park, Seonhee Seok, Hwajeong Seo, Howon Kim 0001 |
J. Supercomput. | 5 |
| 2018 | Erratum to: Secure IoT framework and 2D architecture for End-To-End security
Jongseok Choi, Youngjin In, Changjun Park, Seonhee Seok, Hwajeong Seo, Howon Kim 0001 |
J. Supercomput. | 5 |
| 2017 | Multiprecision Multiplication on ARMv8abstractMultiplication of large integers is a fundamental operation for public key cryptography. In contemporary public key cryptography, the sizes of integers are typically from more than one hundred bits to even several thousands of bits. Because these sizes exceed the bit widths of all general-purpose processors, these multiplications must be performed with a multiprecision multiplication algorithm which splits the operation into multiple partial products and accumulation steps. To ensure efficiency, multiprecision multiplication algorithms must be designed with special care and optimized for the instruction sets of specific processors. Consequently, developing efficient multiprecision multiplication algorithms and optimizing them for specific platforms has been an active research topic. In this paper, we optimize multiprecision multiplication and squaring specifically for the 64-bit ARMv8 processors which are widely used, for example, in modern smart phones and tablets. We combine the subtractive Karatsuba algorithm, operand-scanning techniques (for multiplication) and sliding-block-doubling methods (for squaring) to accelerate the performance of the 256-bit multiprecision multiplication and squaring by 7.6% and 7.0% compared to the OpenSSL implementations. We focus particularly on the multiprecision multiplications that are required in elliptic curve cryptography. Our implementation supports general elliptic curves of various sizes and all source codes are available in public domain. Zhe Liu 0001, Kimmo Järvinen 0001, Weiqiang Liu 0001, Hwajeong Seo |
ARITH | 4 |
| 2017 | Four \mathbb Q on Embedded Devices with Strong Countermeasures Against Side-Channel Attacks
Zhe Liu 0001, Patrick Longa, Geovandro C. C. F. Pereira, Oscar Reparaz, Hwajeong Seo |
CHES | 5 |
| 2017 | Implementing RSA for sensor nodes in smart cities
Lirong Qiu, Zhe Liu 0001, Geovandro C. C. F. Pereira, Hwajeong Seo |
Pers. Ubiquitous Comput. | 4 |
| 2017 | On Emerging Family of Elliptic Curves to Secure Internet of Things: ECC Comes of AgeabstractLightweight Elliptic Curve Cryptography (ECC) is a critical component for constructing the security system of Internet of Things (IoT). In this paper, we define an emerging family of lightweight elliptic curves to meet the requirements on some resource-constrained devices. We present the design of a scalable, regular, and highly-optimized ECC library for both MICAz and Tmote Sky nodes, which supports both widely-used key exchange and signature schemes. Our parameterized implementation of elliptic curve group arithmetic supports pseudo-Mersenne prime fields at different security levels with two optimized-specific designs: the high-speed version (HS) and the memory-efficient (ME) version. The former design achieves record times for computation of cryptographic schemes at roughly$80\sim 128$-bit security levels, while the latter implementation only requires half of the code size of the current best implementation. We also describe our efforts to evaluate the energy consumption and harden our library against some basic side-channel attacks, e.g., timing attacks and simple power analysis (SPA) attacks. Zhe Liu 0001, Xinyi Huang 0001, Muhammad Khurram Khan, Hwajeong Seo, Lu Zhou 0002 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2017 | Efficient Elliptic Curve Cryptography for Embedded DevicesabstractMany resource-constrained embedded devices, such as wireless sensor nodes, require public key encryption or a digital signature, which has induced plenty of research on efficient and secure implementation of elliptic curve cryptography (ECC) on 8-bit processors. In this work, we study the suitability of a special class of finite fields, called optimal prime fields (OPFs), for a “lightweight” ECC implementation with a view toward high performance and security. First, we introduce a highly optimized arithmetic library for OPFs that includes two implementations for each finite field arithmetic operation, namely a performance-optimized version and a security-optimized variant. The latter is resistant against simple power analysis attacks in the sense that it always executes the same sequence of instructions, independent of the operands. Based on this OPF library, we then describe a performance-optimized and a security-optimized implementation of scalar multiplication on the elliptic curve over OPFs at several security levels. The former uses the Gallant-Lambert-Vanstone method on twisted Edwards curves and reaches an execution time of 3.14M cycles (over a 160-bit OPF) on an 8-bit ATmega128 processor, whereas the latter is based on a Montgomery curve and executes in 5.53M cycles. Zhe Liu 0001, Jian Weng 0001, Hwajeong Seo |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | High-Performance Ideal Lattice-Based Cryptography on 8-Bit AVR MicrocontrollersabstractOver recent years lattice-based cryptography has received much attention due to versatile average-case problems like Ring-LWE or Ring-SIS that appear to be intractable by quantum computers. In this work, we evaluate and compare implementations of Ring-LWE encryption and the bimodal lattice signature scheme (BLISS) on an 8-bit Atmel ATxmega128 microcontroller. Our implementation of Ring-LWE encryption provides comprehensive protection against timing side-channels and takes 24.9ms for encryption and 6.7ms for decryption. To compute a BLISS signature, our software takes 317ms and 86ms for verification. These results underline the feasibility of lattice-based cryptography on constrained devices. Zhe Liu 0001, Thomas Pöppelmann, Tobias Oder, Hwajeong Seo, Sujoy Sinha Roy, Tim Güneysu, Johann Großschädl, Howon Kim 0001, Ingrid Verbauwhede |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2016 | Key Update at Train Stations: Two-Layer Dynamic Key Update Scheme for Secure Train Communications
Sang-Yoon Chang, Shaoying Cai, Hwajeong Seo, Yih-Chun Hu |
SecureComm | 3 |
| 2016 | A Synthesis of Multi-Precision Multiplication and Squaring Techniques for 8-Bit Sensor Nodes: State-of-the-Art Research and Future Challenges
Zhe Liu 0001, Hwajeong Seo, Howon Kim 0001 |
J. Comput. Sci. Technol. | 2 |
| 2016 | A fast ARX model-based image encryption schemeabstractThis paper proposes a novel ARX model-based image encryption scheme that uses addition, rotation, and XOR as its confusion and diffusion mechanism instead of S-Box and permutation as in SP networks. The confusion property of the proposed scheme is satisfied by rotation and XOR with chaotic sequences generated from two logistic maps. Unlike classical image encryption schemes that adopt S-Box or permutation of the entire plain image, the diffusion property is satisfied using addition operations. The proposed scheme exhibits good performance on correlation coefficients (horizontal, vertical and diagonal), Shannon’s entropy and NPCR (Number of Pixels Change Rate). Furthermore, simulation results indicate that its time complexity is 9.2 times more efficient than the fastest algorithm(Yang’s algorithm). Jongseok Choi, Seonhee Seok, Hwajeong Seo, Howon Kim 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Efficient arithmetic on ARM-NEON and its application for high-speed RSA implementationabstractAbstract Advanced modern processors support single instruction, multiple data instructions (e.g., Intel‐AVX and ARM‐NEON) and a massive body of research on vector‐parallel implementations of modular arithmetic, which are crucial components for modern public‐key cryptography ranging from Rivest, Shamir, and Adleman (RSA), ElGamal, Digital Signature Algorithm (DSA), and elliptic curve cryptography, have been conducted. In this paper, we introduce a novel double operand scanning method to speed up multi‐precision squaring with non‐redundant representations on single instruction, multiple data architecture where the part of the operands are doubled to compute the squaring operation without read‐after‐write dependencies between source and destination variables. Afterwards, Karatsuba algorithm is applied to both multiplication and squaring operations. For modular multiplication, separated Montgomery algorithm is chosen. Finally, the Rivest, Shamir, and Adleman (RSA) implementations outperform the best‐known results on the ARM‐NEON platforms. Copyright © 2017 John Wiley & Sons, Ltd. Hwajeong Seo, Zhe Liu 0001, Johann Großschädl, Howon Kim 0001 |
Secur. Commun. Networks | 1 |
| 2016 | Binary field multiplication on ARMv8abstractAbstract In this paper, we show efficient implementations of binary field multiplication over ARMv8. We exploit an advanced 64‐bit polynomial multiplication (PMULL) supported by ARMv8 and conduct multiple levels of asymptotically faster Karatsuba multiplication for polynomial multiplication. Finally, our method completed binary field multiplication within 57 and 153 clock cycles for B‐251 and B‐571 cases, respectively. Proposed method improves the speed‐performance by a factor of 4.5 times than previous techniques on same target platform. Copyright © 2016 John Wiley & Sons, Ltd. Hwajeong Seo, Zhe Liu 0001, Yasuyuki Nogami, Jongseok Choi, Howon Kim 0001 |
Secur. Commun. Networks | 1 |
| 2016 | Hybrid Montgomery ReductionabstractIn this article, we present a hybrid method to improve the performance of the Montgomery reduction by taking advantage of the Karatsuba technique. We divide the Montgomery reduction into two sub-parts, including one for the conventional Montgomery reduction and the other one for Karatsuba-aided multiplication. This approach reduces the multiplication complexity of n -limb Montgomery reduction from θ( n 2 + n ) to asymptotic complexity θ (7 n 2 /8 + n ). Our practical implementation results over an 8-bit microcontroller also show performance enhancements by 11%. Hwajeong Seo, Zhe Liu 0001, Yasuyuki Nogami, Jongseok Choi, Howon Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Efficient Implementation of NIST-Compliant Elliptic Curve Cryptography for 8-bit AVR-Based Sensor NodesabstractIn this paper, we introduce a highly optimized software implementation of standards-compliant elliptic curve cryptography (ECC) for wireless sensor nodes equipped with an 8-bit AVR microcontroller. We exploit the state-of-the-art optimizations and propose novel techniques to further push the performance envelope of a scalar multiplication on the NIST P-192 curve. To illustrate the performance of our ECC software, we develope the prototype implementations of different cryptographic schemes for securing communication in a wireless sensor network, including elliptic curve Diffie–Hellman (ECDH) key exchange, the elliptic curve digital signature algorithm (ECDSA), and the elliptic curve Menezes–Qu–Vanstone (ECMQV) protocol. We obtain record-setting execution times for fixed-base, point variable-base, and double-base scalar multiplication. Compared with the related work, our ECDH key exchange achieves a performance gain of roughly 27% over the best previously published result using the NIST P-192 curve on the same platform, while our ECDSA performs twice as fast as the ECDSA implementation of the well-known TinyECC library. We also evaluate the impact of Karatsuba’s multiplication technique on the overall execution time of a scalar multiplication. In addition to offering high performance, our implementation of scalar multiplication has a highly regular execution profile, which helps to protect against certain side-channel attacks. Our results show that NIST-compliant ECC can be implemented efficiently enough to be suitable for resource-constrained sensor nodes. Zhe Liu 0001, Hwajeong Seo, Johann Großschädl, Howon Kim 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | Efficient Implementation of ECDH Key Exchange for MSP430-Based Wireless Sensor NetworksabstractPublic-Key Cryptography (PKC) is an indispensable building block of modern security protocols, and, thus, essential for secure communication over insecure networks. Despite a significant body of research devoted to making PKC more "lightweight," it is still commonly perceived that software implementations of PKC are computationally too expensive for practical use in ultra-low power devices such as wireless sensor nodes. In the present paper we aim to challenge this perception and present a highly-optimized implementation of Elliptic Curve Cryptography (ECC) for the TI MSP430 series of 16-bit microcontrollers. Our software is inspired by MoTE-ECC and supports scalar multiplication on two families of elliptic curves, namely Montgomery and twisted Edwards curves. However, in contrast to MoTE-ECC, we use pseudo-Mersenne prime fields as underlying algebraic structure to facilitate inter-operability with existing ECC implementations. We introduce a novel "zig-zag" technique for multiple-precision squaring on the MSP430 and assess its execution time. Similar to MoTE-ECC, we employ the Montgomery model for variable-base scalar multiplications and the twisted Edwards model if the base point is fixed (e.g. to generate an ephemeral key pair). Our experiments show that the two scalar multiplications needed to perform an ephemeral ECDH key exchange can be accomplished in 4.88 million clock cycles altogether (using a 159-bit prime field), which sets a new speed record for ephemeral ECDH on a 16-bit processor. We also describe the curve generation process and analyze the execution time of various field and point arithmetic operations on curves over a 159-bit and a 191-bit pseudo-Mersenne prime field. Zhe Liu 0001, Hwajeong Seo, Xinyi Huang 0001, Johann Großschädl |
AsiaCCS | 2 |
| 2015 | Efficient Ring-LWE Encryption on 8-Bit AVR Processors
Zhe Liu 0001, Hwajeong Seo, Sujoy Sinha Roy, Johann Großschädl, Howon Kim 0001, Ingrid Verbauwhede |
CHES | 2 |
| 2015 | Montgomery multiplication and squaring for Optimal Prime Fields
Hwajeong Seo, Zhe Liu 0001, Yasuyuki Nogami, Jongseok Choi, Howon Kim 0001 |
Comput. Secur. | 1 |
| 2015 | Performance evaluation of twisted Edwards-form elliptic curve cryptography for wireless sensor nodesabstractAbstract Wireless sensor networks (WSNs) pose a number of unique security challenges that demand innovation in several areas including the design of cryptographic primitives and protocols. Despite recent progress, the efficient implementation of Elliptic Curve Cryptography (ECC) for WSNs is still a very active research topic, and techniques to further reduce the time and energy cost of ECC are eagerly sought. This paper presents an optimized ECC implementation that we developed from scratch to comply with the severe resource constraints of 8‐bit sensor nodes such as the MICAz and IRIS motes. Our ECC software uses Optimal Prime Fields as underlying algebraic structure and supports two different families of elliptic curves, namely, Weierstraß‐form and twisted Edwards‐form curves. Due to the combination of efficient field arithmetic and fast group operations, we achieve an execution time of 5.3·106clock cycles for a full 160‐bit scalar multiplication on an 8‐bit ATmega128 microcontroller, which is more than three times faster than the widely used TinyECC library. Our implementation also shows that the energy cost of scalar multiplication on a MICAz (or IRIS) mote amounts to just 17.34mJ when using a twisted Edwards curve over a 160‐bit Optimal Prime Field. This result further demonstrates the advantage of special family of elliptic curves for resource‐constrained environments. Copyright © 2015 John Wiley & Sons, Ltd. Zhe Liu 0001, Hwajeong Seo, Qiuliang Xu |
Secur. Commun. Networks | 2 |
| 2015 | Karatsuba-Block-Comb technique for elliptic curve cryptography over binary fieldsabstractEfficient implementation of elliptic curve cryptography on resource-constrained microcontroller is considered to be one of the hot and challenging research topics because of the limited computing power and storages of target platforms and high computational costs of elliptic curve cryptography. In this paper, we focus on enhancing the performance of scalar multiplication over GF2m by suggesting a new technique for speeding up the performance of multiplication, called Karatsuba-Block-Comb KBC multiplication. KBC method combines both the advantages of Karatsuba algorithm and Block-Comb method. This technique replaces the part of expensive Block-Comb binary field multiplications with several cheap additions by following Karatsuba rule. In case of squaring, we describe an optimized squaring algorithm with 8-bit look-up table that is significantly faster than previous works with 4-bit look-up table. Both of the proposed approaches improve the best known results by a factor of 24.6% and 16.8% 160-bit operand over 8-bit AVR processor Atmel Corporation, San Jose, CA, USA, respectively. Finally, we realize the scalar multiplication over GF2163, which only requires 0.29s for a full scalar multiplication when the processor runs at 7.37MHz. This result outperforms the previous best implementation by a factor of 9.3%. The research results presented in this paper prove that it is also possible to achieve high performance over binary fields by combing the algorithm with sub-quadratic complexity. Furthermore, we suggest constant time KBC method. Block-Comb method does not provide constant time, and look-up table method is also vulnerable to memory address side channel attack. However, our method is establishing the scalar multiplication in 0.35s with high security against both attacks. Copyright © 2015 John Wiley & Sons, Ltd. Hwajeong Seo, Zhe Liu 0001, Jongseok Choi, Howon Kim 0001 |
Secur. Commun. Networks | 1 |
| 2015 | Optimized Karatsuba squaring on 8-bit AVR processorsabstractAbstract Multi‐precision squaring is one of the performance‐critical operations for implementation of elliptic curve cryptography. This paper continues the line of research on high‐speed multi‐precision squaring on embedded processors. In particular, we present an optimized Karatsuba squaring method for 8‐bit AVR processors. We compute the multiplication part with the fastest Karatsuba multiplication, and then the remaining two squaring parts are conducted with the fastest sliding block doubling squaring. As a result, The proposed method sets the new speed records for multi‐precision squaring, improving the execution time by up to 8.49% compared with the best known works. Copyright © 2015 John Wiley & Sons, Ltd. Hwajeong Seo, Zhe Liu 0001, Jongseok Choi, Howon Kim 0001 |
Secur. Commun. Networks | 1 |
| 2014 | Reverse Product-Scanning Multiplication and Squaring on 8-Bit AVR Processors
Zhe Liu 0001, Hwajeong Seo, Johann Großschädl, Howon Kim 0001 |
ICICS | 2 |
| 2014 | Binary and prime field multiplication for public key cryptography on embedded microprocessorsabstractEmbedded microprocessors are used in a wide variety of platforms, including Radio frequency identification RFID systems, sensor networks, and smartphones. Unfortunately, as practical use of microprocessors has increased, so have the security problems associated with them. Although public key cryptography PKC can mitigate these problems, standard implementations of PKC also impose a steep computational cost on resource-constrained devices. To reduce this cost, researchers have proposed alternative implementations that accelerate multiprecision multiplication, the most expensive operation involved in PKC. In this paper, we focus on a further optimization of this same operation, using several innovative methods: carry-once, optimized multiplication and accumulation MAC, unbalanced comb, and optimized comb-window. These methods yield further performance improvements of 2%, 17%, 4.5%, and 9.5%, respectively, on representative modern microprocessors including ATmega128 and MSP430. Copyright © 2013 John Wiley & Sons, Ltd. Hwajeong Seo, Yeoncheol Lee, Taehwan Park, Howon Kim 0001 |
Secur. Commun. Networks | 1 |
| 2013 | Efficient Implementation of NIST-Compliant Elliptic Curve Cryptography for Sensor Nodes
Zhe Liu 0001, Hwajeong Seo, Johann Großschädl, Howon Kim 0001 |
ICICS | 2 |
| 2013 | Performance enhancement of TinyECC based on multiplication optimizationsabstractABSTRACT Because wireless sensor network (WSN), which is composed of a large number of low‐cost and resource‐constrained devices, communicates on the basis of wireless protocols such as IEEE 802.15.4, ZigBee, and DASH‐7, it is easily vulnerable to eavesdropping, illegal modification, privacy infringement and denial‐of‐service attacks. These attacks destroy the data integrity, confidentiality, and authentication of the basic WSN security requirements and then the reliability and security of the WSN‐based applications are deteriorated. There have been many research efforts to make secure WSN environments. Among these efforts, TinyECC is one of outstanding works. It provides several security protocols such as Elliptic Curve Diffie–Hellman, Elliptic Curve Digital Signature Algorithm, and Elliptic Curve Integrated Encryption Scheme, based on the Elliptic Curve Cryptography (ECC). TinyECC is basically a well‐written TinyOS‐based code and is optimized to resource‐constrained environments. The Barrett reduction, hybrid multiplication, and several optimization techniques are also used for high performance even with low‐energy consumption. However, the hybrid multiplication technique used in TinyECC is known to be not suitable for 16‐bit processor, MSP430, which is a familiar processor for sensor node. This is due to the fact that the MSP 430 processor does not provide enough number of registers for hybrid multiplication techniques. Because the multiplication operation over the finite field is a major operation of the ECC, it causes a high latency of multiplication operations and eventually degrades the performance of the ECC operation. In this paper, we propose a novel multiplication operation based on the cached operands and reordered partial products. The proposed method shows that the latency of the polynomial multiplication, which is the core operation of the ECC, is 6% smaller than previously known results. Copyright © 2012 John Wiley & Sons, Ltd. Hwajeong Seo, Kyung-Ah Shim, Howon Kim 0001 |
Secur. Commun. Networks | 1 |