Wai-Kong Lee

dblp:124/4146 · DBLP profile ↗
← Back
31ranked-venue papers
10as first author
21since 2021 · last 2025
0000-0003-4659-8979ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 6 first-author · 9 since 2021Computer networks · 10 · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AuGQ: Augmented quantization granularity to overcome accuracy degradation for sub-byte quantized deep neural networks
Ahmed Mujtaba, Wai-Kong Lee, ByoungChul Ko, Hyung Jin Chang, Seong Oun Hwang
Appl. Intell.2
2025 RISC32-A: A Low-Power Asynchronous IoT Processor on FPGA With Adaptive Pipeline Structure
abstract
Field-programmable gate array (FPGA)-based sensor nodes are gaining popularity for Internet of Things (IoT) applications due to their flexible hardware reconfigurability. RISC32 is a recently proposed synchronous IoT processor, targeting the FPGA-based sensor nodes. However, its dynamic energy consumption is relatively high as its circuit components often activate with each tick of the global clock, irrespective of the actual need. RISC32-LP presented various power reduction techniques to enhance its energy efficiency, but it still necessitates the use of a constantly switching global clock in many parts of the system for synchronization. In view of that, this research work proposed a novel asynchronous processor design (RISC32-A) to significantly minimize the switching events in RISC32, thus lowering its overall dynamic energy consumption. Besides, this research work also proposed an adaptive asynchronous pipeline structure that allows selective pipeline stages to be dynamically skipped, merged, split, and stalled during the program run-time for optimal performance-energy tradeoffs. A novel pipeline stage skipping solution was introduced, which considers various instruction types and prevents wasteful data movement for the early-completed instructions. Additionally, a two-phase collapsible register handshake component with normally closed flip-flops was presented for enhanced dynamic power savings. Finally, a new pipeline halting solution was proposed to eliminate the decoding and forwarding hardware overheads found in the prior works. Experimental results show that the proposed RISC32-A can achieve an average reduction of$\approx 83.71\%$and$\approx 65.98\%$in dynamic energy consumption when compared to its synchronous counterpart (RISC32) and its optimal low-power synchronous counterpart (RISC32-LP), respectively.
Min-An Yong, Kai Ming Mok, Wai-Kong Lee, Shen-Khang Teoh, Denis Chee-Keong Wong
IEEE Internet Things J.3
2025 A compact and flexible FPGA accelerator for regular and octave convolutional neural networks
Jin-Chuan See, Hui-Fuang Ng, Hung-Khoon Tan, Jing-Jing Chang, Wai-Kong Lee
Neural Comput. Appl.5
2025 A signature scheme constructed from zero knowledge argument of knowledge for the subgraph isomorphism problem
Chii Liang Ng, Denis Chee-Keong Wong, Gek Ling Chia, Bok-Min Goi, Wai-Kong Lee, Wun-She Yap
Theor. Comput. Sci.5
2024 Efficient TMVP-Based Polynomial Convolution on GPU for Post-Quantum Cryptography Targeting IoT Applications
abstract
Recently proposed lattice-based cryptography algorithms can be used to protect the IoT communication against the threat from quantum computers, but they are computationally heavy. In particular, polynomial convolution is one of the most time-consuming operations in lattice-based cryptography. To achieve efficient implementation, the Number Theoretic Transform (NTT) algorithm is an ideal choice, but it has certain limitations on the parameters, which not all lattice-based schemes can employ directly. Hence, alternative techniques are proposed to accelerate polynomial convolution on lattice-based schemes that cannot utilize the NTT directly. In this paper, we propose a parallel Toeplitz matrix-vector product (TMVP) version to accelerate the polynomial convolution in PQC algorithms implemented it on a graphics processing unit (GPU). This is the first time a TMVP parallel version has been proposed and experimented on different GPU cores (i.e., CUDA-cores and Tensor-cores). The effectiveness of the proposed solution is validated on Saber (the NIST post-quantum standardization finalist) and Sable (an improved version of Saber) schemes. Experimental results show that TMVP-based polynomial convolution using CUDA-cores fails to exhibit a significant enhancement compared to the schoolbook CUDA-core method already proposed by Hafeez et al. 2023. However, when the TMVP technique is applied to Tensor-cores, it outperformed state-of-the-art implementations. The proposed Tensor-core approach outperformed the schoolbook Tensor-core method by up to 1.21×, and outperformed the dot-product-instructions method (Lee et al. 2022) by up to 3.63×. The proposed TMVP Tensor-cores is also faster than the TMVP CUDA-cores method by 13.76×.
Muhammad Asfand Hafeez, Wai-Kong Lee, Angshuman Karmakar, Seong Oun Hwang
IEEE Internet Things J.2
2024 High Throughput Lattice-Based Signatures on GPUs: Comparing Falcon and Mitaka
abstract
The US National Institute of Standards and Technology initiated a standardization process for post-quantum cryptography in 2017, with the aim of selecting key encapsulation mechanisms and signature schemes that can withstand the threat from emerging quantum computers. In 2022, Falcon was selected as one of the standard signature schemes, eventually attracting effort to optimize the implementation of Falcon on various hardware architectures for practical applications. Recently, Mitaka was proposed as an alternative to Falcon, allowing parallel execution of most of its operations. These recent advancements motivate us to develop high throughput implementations of Falcon and Mitaka signature schemes on Graphics Processing Units (GPUs), a massively parallel architecture widely available on cloud service platforms. In this paper, we propose the first parallel implementation of Falcon on various GPUs. An iterative version of the sampling process in Falcon, which is also the most time-consuming Falcon operation, was developed. This allows us to implement Falcon signature generation without relying on expensive recursive function calls on GPUs. In addition, we propose a parallel random samples generation approach to accelerate the performance of Mitaka on GPUs. We evaluate our implementation techniques on state-of-the-art GPU architectures (RTX 3080, A100, T4 and V100). Experimental results show that our Falcon-512 implementation achieves 58,595 signatures/second and 2,721,562 verifications/second on an A100 GPU, which is$20.03\times$and$29.51\times$faster than the highly optimized AVX2 implementation on CPU. Our Mitaka implementation achieves 161,985 signatures/second and 1,421,046 verifications/second on the same GPU. Due to the adoption of a parallelizable sampling process, Mitaka signature generation enjoys$\approx 2$–$20 \times$higher throughput than Falcon on various GPUs. The high throughput signature generation and verification achieved by this work can be very useful in various emerging applications, including the Internet of Things.
Wai-Kong Lee, Raymond K. Zhao, Ron Steinfeld, Amin Sakzad, Seong Oun Hwang
IEEE Trans. Parallel Distributed Syst.1
2023 Efficient, Error-Resistant NTT Architectures for CRYSTALS-Kyber FPGA Accelerators
abstract
The dawn of cost-effective miniaturised satellites is currently attracting venture capital in a never seen before ratio to launch mega-constellations of satellites for a diverse range of applications. These satellites are vulnerable to attacks by high-capability cyber-criminals (including quantum enabled adversaries), due to the critical data they transmit. Additionally, space missions have long lifespan and a long lead time in terms of development process, requiring a pre-emptive outlook to ensuring their safety. In 2016, National Institute of Standards and Technology (NIST) initiated the competition to standardise the post-quantum cryptography (PQC) schemes, announcing the first portfolio of chosen schemes in 2022. This work targets the only public key exchange (PKE) scheme among the winners of the NIST-PQC standardisation process, CRYSTALS-Kyber, and implements its core bottleneck operation, i.e., number theoretic transform (NTT) extensively used for the polynomial multiplication. To avoid data corruption due to space based radiations, a novel error-resistant model for NTT is presented based on hybrid protection mechanisms, i.e., the use of hamming codes for detection and correction of errors in the twiddle factors and the use of parity computed for all NTT coefficients for error detection. Benchmarking error protection overheads on a Xilinx Virtex-7 FPGA reports 16.4% and 10.8% degradation on the hardware efficiency when the hamming codes for twiddle factors and parity bit for NTT coefficients are used to mitigate errors, respectively. A total of 29.2% area overhead is benchmarked when compared to the standard unprotected NTT implementations.
Safiullah Khan, Ayesha Khalid, Ciara Rafferty, Yasir Ali Shah, Máire O'Neill, Wai-Kong Lee, Seong Oun Hwang
VLSI-SoC6
2023 High Throughput Acceleration of Scabbard Key Exchange and Key Encapsulation Mechanism Using Tensor Core on GPU for IoT Applications
abstract
High throughput key encapsulations and decapsulations are needed by Internet of Things (IoT) applications in order to simultaneously process a multitude of small data in secure communication. In this article, we present two novel techniques for accelerating the implementation of polynomial convolution on a graphics processing unit (GPU), utilizing advanced Tensor cores, which benefit the performance of key encapsulations. First, a polynomial restructuring technique is proposed to allow several polynomials with distinct public keys to be processed in a single communication cycle. This is an improvement compared to the previous work by Lee et al. Next, we observe that polynomial convolution in some key encapsulation mechanisms contains reduction patterns that are not friendly to parallel implementation. We propose separating the multiplication and reduction processes so they can be parallelized independently. To verify the effectiveness of our proposed techniques, we applied it to two key-encapsulation mechanisms from the Scabbard post-quantum key-encapsulation mechanism suite and evaluate their performance. Experimental results show that polynomial convolution using Tensor cores is$1.05\times $faster (for the Florete scheme) and$3.6\times $faster (for the Sable scheme) than using compute unified device architecture core-based multiplication with conventional cores on a GPU. The Tensor cores-based encapsulations and decapsulations are faster than a reference implementation on a CPU supporting AVX2 by more than$5.6\times $and$6.4\times $, respectively, for the Florete scheme and$8.3\times $and$13.3\times $faster, respectively, for the Sable scheme. This shows that the proposed techniques can achieve significantly higher throughput for key exchange and encapsulation mechanisms, which are important for securing IoT applications.
Muhammad Asfand Hafeez, Wai-Kong Lee, Angshuman Karmakar, Seong Oun Hwang
IEEE Internet Things J.2
2023 Area-Time Efficient Implementation of NIST Lightweight Hash Functions Targeting IoT Applications
abstract
To mitigate cybersecurity breaches, secure communication is crucial for the Internet of Things (IoT) environment. Data integrity is one of the most significant characteristics of security, which can be achieved by employing cryptographic hash functions. In view of the demand from IoT applications, the National Institute of Standards and Technology (NIST) initiated a standardization process for lightweight hash functions. This work presents field-programmable gate array (FPGA) implementations and carefully worked out optimizations of four Round-3 finalists in the NIST standardization process. A novel compact PHOTON-Beetle implementation is proposed wherein the underlying matrix multiplication is executed in serialized fashion to achieve a small hardware footprint. Sparkle implementations are carried out by implementing the ARX-box in serialized, parallelized, and hybrid approaches. For Ascon and Xoodyak, the proposed implementations compute certain permutation rounds in one clock cycle in order to explore the tradeoff between computation time and hardware area. As a result, this work achieves the smallest hardware footprint for PHOTON-Beetle consuming an area$3.4 \times $smaller than state-of-the-art implementations. Ascon and Xoodyak are implemented in a flexible manner that achieves throughput-to-area (TP/A) ratios$1.8 \times $and$3.9 \times $higher, respectively, compared to implementations found in the literature. In addition, we propose the first FPGA implementations for the Sparkle hash function. These efficient implementations provide guidelines for choosing a suitable architecture for applications in demand that can be employed in the IoT environment to achieve data integrity for various applications.
Safiullah Khan, Wai-Kong Lee, Angshuman Karmakar, Jose Maria Bermudo Mera, Abdul Majeed 0001, Seong Oun Hwang
IEEE Internet Things J.2
2023 KaratSaber: New Speed Records for Saber Polynomial Multiplication Using Efficient Karatsuba FPGA Architecture
abstract
SABER is a round 3 candidate in the NIST Post-Quantum Cryptography Standardization process. Polynomial convolution is one of the most computationally intensive operation in Saber Key Encapsulation Mechanism, that can be performed through widely explored algorithms like the schoolbook polynomial multiplication algorithm (SPMA) and Number Theoretic Transform (NTT). While SPMA multiplier has a slow latency performance, the NTT-based multiplier usually requires large hardware. In this work, we propose KaratSaber, an optimized Karatsuba polynomial multiplier architecture with a balanced hardware efficiency (throughput-per-slice, TPS) compared to NTT and SPMA based designs. KaratSaber employs several techniques for an efficient design: a parallel grid input technique for efficient pre-processing stage in Karatsuba-based polynomial multiplier, a novel instruction code result-mapping technique catering the negacyclic operations improves the post-processing stage efficiency, a double multiplicand shifter-based multiplier doubles the throughput at the multiplication stage. Combining these three techniques, the proposed KaratSaber architecture is 7.47 × faster compared to the state-of-the-art SPMA Saber architecture at the expense of 4.96 × additional hardware resources; making KaratSaber 46.04% more area-time efficient. When compared to LWRPro, a recent Karatsuba Saber architecture, KaratSaber architecture achieves a 2.11 × higher throughput by only utilizing 1.92 × additional hardware; thus gaining a 10.44% improvement in area-time efficiency.
Zheng-Yan Wong, Denis Chee-Keong Wong, Wai-Kong Lee, Kai Ming Mok, Wun-She Yap, Ayesha Khalid
IEEE Trans. Computers3
2023 Cryptensor: A Resource-Shared Co-Processor to Accelerate Convolutional Neural Network and Polynomial Convolution
abstract
Practical deployment of convolutional neural network (CNN) and cryptography algorithm on constrained devices are challenging due to the huge computation and memory requirement. Developing separate hardware accelerator for AI and cryptography incur large area consumption, which is not desirable in many applications. This article proposes a viable solution to this issue by expressing the CNN and cryptography as generic-matrix-multiplication (GEMM) operations and map them to the same accelerator for reduced hardware consumption. A novel systolic tensor array (STA) design was proposed to reduce the data movement, effectively reducing the operand registers by$2\times $. Two novel techniques, input layer extension and polynomial factorization, are proposed to mitigate the under-utilization issue found in existing STA architecture. Additionally, the tensor processing element (TPE) is fused using DSP unit to reduce the look-up table (LUT) and flip-flops (FFs) consumption for implementing multipliers. On top of that, a novel memory efficient factorization technique is proposed to allow computation of polynomial convolution on the same STA. Experimental results show that Cryptensor achieved 21.6% better throughput for VGG-16 implementation on XC7Z020 FPGA; up to$8.40\times $better-energy efficiency compared to existing ResNet-18 implementation on XC7Z045 FPGA. Cryptensor can also flexibly support multiple security levels in NTRU scheme, with no additional hardware. The proposed hardware unifies the computation of two different domains that are critical for IoT applications, which greatly reduces the hardware consumption on edge nodes.
Jin-Chuan See, Hui-Fuang Ng, Hung-Khoon Tan, Jing-Jing Chang, Kai Ming Mok, Wai-Kong Lee, Chih-Yang Lin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 Look-up the Rainbow: Table-based Implementation of Rainbow Signature on 64-bit ARMv8 Processors
abstract
The Rainbow Signature Scheme is one of the finalists in the National Institute of Standards and Technology (NIST) Post-Quantum Cryptography (PQC) standardization competition, but failed to win because it has lack of stability in the parameter selection. It is the only signature candidate based on a multivariate quadratic equation. Rainbow signatures have smaller signature sizes compared with other post-quantum cryptography candidates. However, they require expensive tower-field based polynomial multiplications. In this article, we propose an efficient implementation of Rainbow signatures using a look-up table–based multiplication method. The polynomial multiplications in Rainbow signatures are performed on the 𝔽 16 field, which is divided into sub-fields 𝔽 4 and 𝔽 2 under the tower-field method. To accelerate the multiplication process on target processors, we propose a look-up table–based tower-field multiplication technique. In 𝔽 16 , all values are expressed in 4-bit data format and can be implemented using a 256-byte look-up table access. The implementation uses the TBL and TBX instructions of the 64-bit ARMv8 target processor. For Rainbow III and Rainbow V, they are computed on the 𝔽 256 field using an additional 16-byte table instead of creating a new look-up table. The proposed technique uses the vector registers of 64-bit ARMv8 processors and can calculate 16 result values with a single instruction. We also proposed implementations that are resistant to timing attacks. There are two types of implementations. The first one is the cache side-attack resistant implementation, which utilizes the 128-byte cache lines of the M1 processor. In this implementation, cache misses do not occur, and cache hits always occur. The second type is the constant-time implementation. This method takes a step-by-step approach to finding the required look-up table value and ensures that the same number of accesses is made regardless of which look-up table value is called. This implementation is designed to be constant-time, meaning it does not leak timing information. Our experiments on modern Apple M1 processors showed up to 428.73× and 114.16× better performance for finite field multiplications and Rainbow signatures schemes, respectively, compared with previous reference implementations. To the best of our knowledge, this proposed Rainbow implementation is the first optimized Rainbow implementation for 64-bit ARMv8 processors.
Hyeokdong Kwon, Minjoo Sim, Wai-Kong Lee, Hwajeong Seo
ACM Trans. Embed. Comput. Syst.4
2022 High Throughput Implementation of Post-quantum Key Encapsulation and Decapsulation on GPU for Internet of Things Applications
abstract
[J1C2 Presentation Abstract at IEEE SERVICES 2022 for IEEE Transactions on Services Computing DOI 10.1109/TSC.2021.3103956]
Wai-Kong Lee, Seong Oun Hwang
SERVICES1
2022 RISC32-LP: Low-Power FPGA-Based IoT Sensor Nodes With Energy Reduction Program Analyzer
abstract
Field-programmable gate array (FPGA)-based sensor nodes are popular for their flexible design approach and field reconfigurability. RISC32 is one of the recent Internet of Things (IoT) processors proposed for the development of FPGA-based sensor nodes, and it includes the ability to reconfigure the microarchitecture on the fly in order to reduce dynamic energy consumption. However, such a method does not minimize static energy consumption, which is important in FPGA-based systems. In this work, clock gating (CG) and dynamic voltage and frequency scaling (DVFS) are applied to further reduce the energy consumption in RISC32. In the research presented here, we implemented a new software called the energy reduction program analyzer to estimate the parameters that configure a sensor node to achieve minimum energy consumption, targeting the typical IoT application scenario. Experimental results show that the low-power techniques applied in this work (RISC32-LP) can reduce energy consumption by 47%, compared to the standard RISC32 processor.
Beng-Liong Tan, Kai Ming Mok, Jing-Jing Chang, Wai-Kong Lee, Seong Oun Hwang
IEEE Internet Things J.4
2022 DPCrypto: Acceleration of Post-Quantum Cryptography Using Dot-Product Instructions on GPUs
abstract
Modern NVIDIA GPU architectures offer dot-product instructions (DP2A and DP4A), with the aim of accelerating machine learning and scientific computing applications. These dot-product instructions allow the computation of multiply-and-add instructions in a single clock cycle, effectively achieving higher throughput compared to conventional 32-bit integer units. In this paper, we show that the dot-product instruction can also be used to accelerate matrix-multiplication and polynomial convolution operations, which are widely used in post-quantum lattice-based cryptographic schemes. In particular, we propose a highly optimized implementation of FrodoKEM wherein the matrix-multiplication is accelerated by the dot-product instruction. We also present specially designed data structures that allow an efficient implementation of Saber key-encapsulation mechanism, utilizing the dot-product instruction to speed-up the polynomial convolution. The proposed FrodoKEM implementation achieves$4.37\times $higher throughput than the state-of-the-art implementation on a V100 GPU. This paper also presents the first implementation of Saber on GPU platforms, achieving 124,418, 120,463, and 31,658 key exchanges per second on RTX3080, V100, and T4 GPUs, respectively. Since matrix-multiplication and polynomial convolution operations are the most time-consuming operations in lattice-based cryptographic schemes, we strongly believe that the proposed methods can be beneficial to other KEM and signatures schemes based on lattices.
Wai-Kong Lee, Hwajeong Seo, Seong Oun Hwang, Ramachandra Achar, Angshuman Karmakar, Jose Maria Bermudo Mera
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 High Throughput Implementation of Post-Quantum Key Encapsulation and Decapsulation on GPU for Internet of Things Applications
abstract
Internet of Things (IoT) sensor nodes are placed ubiquitously to collect information, which is then vulnerable to malicious attacks. For instance, adversaries can perform side channel attack on the sensor nodes to recover the symmetric key for encrypting IoT data. Refreshing the symmetric key frequently can reduce the risk of compromised keys. However, the number of sensor nodes connected to the gateway and cloud server is massive. Refreshed symmetric keys need to be sent to gateway devices and cloud server frequently with a secure key encapsulation mechanism (KEM), which is time-consuming. In this article, novel and efficient implementation techniques are proposed to accelerate Kyber, a post-quantum KEM, on a Graphics Processing Unit (GPU). Fully parallel implementation of number theoretic transform (NTT) with combined levels is presented, which is 2.65× faster than state-of-the-art result on a GPU. Other proposed techniques include parallel rejection sampling, central binomial distribution with coalesced memory access and parallel fine-grain AES-256. These techniques enable high throughput performance with 162760 encapsulations/second and 107631 decapsulations/second on an RTX2060 GPU. This is also the first fine grain implementation of post-quantum KEM (Kyber) on a GPU, which can be used to offer key encapsulation/decapsulation as a service to reduce the burden on IoT systems.
Wai-Kong Lee, Seong Oun Hwang
IEEE Trans. Serv. Comput.1
2021 Novel Postquantum MQ-Based Signature Scheme for Internet of Things With Parallel Implementation
abstract
Internet of Things (IoT) is a paradigm shifting technology that enables many innovative applications in the near future. Proactive measures are required to protect such architecture from cyber attacks. One of the most important security issues in this architecture is the authentication of edge nodes, which can be resolved through the deployment of digital signatures. However, existing standardized digital signatures are vulnerable to attacks from quantum computers, which can be unsafe in the near future. In this article, we propose a new signature scheme based on multivariate polynomials with efficient key and signature sizes, which is resistant to quantum computer attacks. The proposed scheme is also very friendly to parallel implementation, enabling efficient deployment of edge nodes authentication at high throughput. When implemented on a GPU device, the proposed scheme can generate 113 signatures/s and verify 120 signatures/s, which is 12.56× and 10.00× faster than a serial implementation in CPU.
Sedat Akleylek, Meryem Soysaldi, Wai-Kong Lee, Seong Oun Hwang, Denis Chee-Keong Wong
IEEE Internet Things J.3
2021 Scalable and Efficient Hardware Architectures for Authenticated Encryption in IoT Applications
abstract
Internet of Things (IoT) is a key enabling technology, wherein sensors are placed ubiquitously to collect and exchange information with their surrounding nodes. Due to the inherent interconnectivity, IoT devices are vulnerable to cybersecurity attacks. To mitigate these vulnerabilities, cryptographic primitives can be employed, but they require significant computation, which restricts their adoption in IoT. Moreover, IoT systems have diverse requirements, ranging from high-throughput (TP) to the area constrained. This makes it hard to deploy appropriate security measures in a systematic manner. To address these issues, three generic implementation strategies (unrolled, round-based, and serialized) are proposed for developing highly efficient hardware architectures. They are applicable to all authenticated encryption schemes and are lightweight and fast, compared to conventional public key encryption. In this article, Ascon is implemented as an example based on those three strategies: 1) the unrolled architecture achieves TP of 766.9 Mb/s (Ascon-128) and 1389.2 Mb/s (Ascon-128a), which are suitable for high-throughput IoT applications; 2) the round-based architecture achieves 0.153 (Ascon-128) and 0.244 (Ascon-128a) TP-to-area ratio, which are, respectively, 73.8% and 40.2% better than state-of-the-art results; and 3) a novel serialized implementation technique is proposed wherein the substitution-box (S-box) is processed in multiple-bit-per-cycle, in contrast to the conventional one-bit-per-cycle approach. The TP of the two-bits-per-clock-cycle implementation is increased by 230.8% with only 36.8% additional hardware area. The proposed strategies allow us to scale the number of rounds (round-based) and bits-per-clock-cycle (serialized) to meet differing requirements in TP and area which are demonstrated for smart city IoT applications.
Safiullah Khan, Wai-Kong Lee, Seong Oun Hwang
IEEE Internet Things J.2
2021 GPU-Accelerated Adaptive PCBSO Mode-Based Hybrid RLA for Sparse LU Factorization in Circuit Simulation
abstract
LU factorization is extensively used in engineering and scientific computations for solution of large set of linear equations. Particularly, circuit simulators rely heavily on sparse version of LU factorization for solution involving circuit matrices. One of the recent advances in this field is exploiting the emerging computing platform of graphics processing units (GPUs) for parallel and sparse LU factorization. In this article, following contributions are made to advance the state of the art in hybrid right-looking algorithm (RLA): 1) a novel GPU kernel based on parallel column and block size optimization (PCBSO) is developed for adaptively allocating the block size while optimizing the number of columns for parallel execution based on the size of their associated submatrices at every level. The proposed approach helps to minimize the resource contention and to improve the computational performance and 2) an algorithm is developed to enable the execution of the new adaptive mode with dynamic parallelism. Also, a comprehensive performance comparison using a set of benchmark circuit examples is presented. The results indicate that, the proposed advancements can improve the results of state-of-the-art right looking sparse LU factorization in GPU by$1.54\times $(Arithmetic Mean).
Wai-Kong Lee, Ramachandra Achar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Accelerating number theoretic transform in GPU platform for fully homomorphic encryption
Jia-Zheng Goey, Wai-Kong Lee, Bok-Min Goi, Wun-She Yap
J. Supercomput.2
2021 Parallel implementation of Nussbaumer algorithm and number theoretic transform on a GPU platform: application to qTESLA
Wai-Kong Lee, Sedat Akleylek, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi, Seong Oun Hwang
J. Supercomput.1
2020 A Parameter-Free Vibration Analysis Solution for Legacy Manufacturing Machines' Operation Tracking
abstract
Despite the fact that the revolution of Industry 4.0 has started almost a decade ago, there are still many yesteryear's manufacturing machines that are still currently in operation in many small and medium enterprises (SME) factories. These legacy manufacturing machines are built without computing power and Internet connectivity. Therefore, the process of gathering operational information of such systems is often done manually. This article aims to automatically track these machines' operation status via the vibration produced by these machines, by using a retrofit Internet-of-Things (IoT) approach that attaches wireless vibration sensors onto legacy manufacturing machines to capture the vibration of the machines. One of the challenges of the proposed retrofit approach is to interpret the meaning of the vibration without any prior knowledge of the machine's vibration and also without the privilege to interrupt the manufacturing process to produce data sets with labels. Although there are many existing works that capture and analyze vibration, they very often only focus on fault diagnosis and prognosis. Also, many of these vibration analysis techniques are not parameter free; i.e., parameters need to be fine-tuned according to the data. The contribution of this article is the proposal of a parameter-free vibration analysis technique to cluster and classify the type of vibrations produced by a machine. Experiments, which were carried out in a limestone processing factory on real industrial machineries, show that the proposed technique is able to track the operation status of a 3-speed industrial exhaust fan with an average accuracy of 98.6% (worst case 95.5%) and standard uncertainty of 1.06%.
Boon-Yaik Ooi, Woan Lin Beh, Wai-Kong Lee, Shervin Shirmohammadi
IEEE Internet Things J.3
2020 Optimized IoT Cryptoprocessor Based on QC-MPDC Key Encapsulation Mechanism
abstract
The key encapsulation mechanism (KEM) is an important cryptographic tool to protect communication in the Internet of Things (IoT). In the near future, classical algorithms used to construct KEMs, such as RSA and elliptic curve cryptography, will be vulnerable to attacks from quantum computers. Recently, Yamada et al. proposed the quasicyclic medium density parity check (QC-MDPC) KEM, which is considered one of the most advanced code-based cryptosystems to resist quantum attacks. In this article, an optimized implementation of QC-MDPC KEM for IoT applications is presented. Our main contributions are threefold: 1) the fastest QC-MDPC McEliece decryption in field-programmable gate array (FPGA); 2) the first QC-MDPC KEM implementation in FPGA; and 3) the first iteration count attack-resistant QC-MDPC decoder in FPGA. To improve the decryption speed, we introduce a novel customized rotation engine (CRE) and incorporated several recent techniques reported in the literature, including adaptive threshold and Hamming weight estimation. The best-achieved throughput in our implementation on Xilinx Virtex 7 FPGA is 12.7% faster than the state-of-the-art result reported by Heyse et al. The proposed CRE was then integrated with QC-MDPC KEM to produce a fast and secure KEM. Furthermore, to prevent timing attacks demonstrated recently, a constant-time implementation of the QC-MDPC McEliece decoder was presented.
Jun-Hoe Phoon, Wai-Kong Lee, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi, Raphael C.-W. Phan
IEEE Internet Things J.2
2020 Area-Time-Efficient Code-Based Postquantum Key Encapsulation Mechanism on FPGA
abstract
Postquantum cryptography attracts a lot of attention from the research community recently due to the emergence threat from quantum computer toward the conventional cryptographic schemes. In view of that, NIST had initiated the standardization process in 2017. Bit flipping key encapsulation (BIKE) designed by Aragon et al. is one of the promising code-based schemes among the round-3 candidates. BIKE utilizes a quasi-cyclic medium density parity check (QC-MDPC) code and incorporates a few variants derived from the McEliece, Niederreiter, and Ouroboros schemes. In this article, we present efficient and constant time implementation of BIKEI and BIKE-III in field-programmable gate array (FPGA), which has the best area-time efficiency so far. We proposed modification to the original one-round bit flipping algorithm to achieve more area-time-efficient decoding in hardware, which achieved latency of 464.73 and 556.52 μs for BIKE-I and BIKE-III, respectively, in Virtex-7. A pipelined key encapsulation architecture is proposed to speedup the key encapsulation of BIKE-I and BIKE-III, achieving the latency of 146.47 and 153.25 μs on the same FPGA platform. Considering the Artix-7 FPGA platform, our combined key generation and encapsulation module for BIKE-I is also three more area-time efficient compared with the state-of-the-art BIKE-I implementation by Aragon et al.
Jun-Hoe Phoon, Wai-Kong Lee, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Accelerating Number Theoretic Transform in GPU Platform for qTESLA Scheme
Wai-Kong Lee, Sedat Akleylek, Wun-She Yap, Bok-Min Goi
ISPEC1
2019 Terabit encryption in a second: Performance evaluation of block ciphers in GPU with Kepler, Maxwell, and Pascal architectures
abstract
Summary With the emergence of IoT and cloud computing technologies, massive data are generated from various applications everyday and communicated through the Internet. Secure communication is essential to protect these data from malicious attacks. Block ciphers are one mechanism to offer such protection but unfortunately involve intensive computations that can be performance bottlenecks to the servers, especially when the data center needs to handle thousands of concurrent transactions. In this paper, we investigate the feasibility of the GPU as an accelerator to perform high‐speed encryption in server environments. We present optimized implementations of a conventional block cipher (AES) and new lightweight block ciphers (LEA, Chaskey, SIMON, SPECK, and SIMECK) across three new GPU architectures (Kepler, Maxwell, and Pascal). For AES, we improve the fine‐grain implementation by utilizing the warp shuffle instruction available in these three new GPU architectures, which yield a 6%‐16% improvement over the previous implementations. For LEA, Chaskey, SIMON, SPECK, and SIMECK, we first analyze why they cannot have efficient fine‐grain implementations in the GPU and then present our optimization techniques, which are able to achieve impressive encryption speeds of 1.912, 637, 1.485, 2.291, and 1.478 Tb/s, respectively, in GTX1080.
Wai-Kong Lee, Bok-Min Goi, Raphael C.-W. Phan
Concurr. Comput. Pract. Exp.1
2019 Blockchain based searchable encryption for electronic health record sharing
Lanxiang Chen, Wai-Kong Lee, Chin-Chen Chang 0001, Kim-Kwang Raymond Choo
Future Gener. Comput. Syst.2
2019 Signature Gateway: Offloading Signature Generation to IoT Gateway Accelerated by GPU
abstract
The emergence of Internet of Things (IoT) brings us the possibility to form a well connected network for ubiquitous sensing, intelligent analysis, and timely actuation, which opens up many innovative applications in our daily life. To secure the communication between sensor nodes, gateway devices and cloud servers, cryptographic algorithms (e.g., digital signature, block cipher, and hash function) are widely used. Although cryptographic algorithms are effective in preventing malicious attacks, they involve heavy computation that may not be executed efficiently in resource constraint sensor nodes. In particular, the authentication of a sensor node is usually performed through a digital signature (e.g., RSA and elliptic curve cryptography), which can be slow when executed on a microcontroller. In this paper, an IoT architecture that offloads the digital signature generation to a nearby signature gateway equipped with graphic processing unit (GPU) accelerator are proposed. The communication process for signature offloading, together with optimized implementation techniques for RSA in signature gateway, are also presented in this paper. We have evaluated two different ways to implement modular exponentiation in RSA, namely residue number system and multiprecision montgomery multiplication (MPMM). The experimental results show that our RSA implementation using MPMM is 10.1% faster than the best RSA implementation in GPU. Our proposed IoT architecture with signature gateway can successfully reduce the burden of sensor nodes to generate signatures, at the same time preserve the ability to authenticate the sensor nodes.
Chin-Chen Chang 0001, Wai-Kong Lee, Yanjun Liu 0002, Bok-Min Goi, Raphael C.-W. Phan
IEEE Internet Things J.2
2019 Hierarchical gated recurrent neural network with adversarial and virtual adversarial training on text classification
Hoon-Keng Poon, Wun-She Yap, Yee-Kai Tee, Wai-Kong Lee, Bok-Min Goi
Neural Networks4
2018 ArchCam: Real time expert system for suspicious behaviour detection in ATM site
Wai-Kong Lee, Chun Farn Leong, Weng-Kin Lai, Lee Kien Leow, Thiah-Huat Yap
Expert Syst. Appl.1
2018 Dynamic GPU Parallel Sparse LU Factorization for Fast Circuit Simulation
Wai-Kong Lee, Ramachandra Achar, Michel S. Nakhla
IEEE Trans. Very Large Scale Integr. Syst.1