Denis Chee-Keong Wong

dblp:231/8592 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-7985-3449ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Theory of computation · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cross-plane color image encryption based on two-dimensional sine-henon map and genetic algorithm
Kuan-Wai Wong, Bok-Min Goi, Wun-She Yap, Denis Chee-Keong Wong, Guodong Ye
Signal Process. Image Commun.4
2025 RISC32-A: A Low-Power Asynchronous IoT Processor on FPGA With Adaptive Pipeline Structure
abstract
Field-programmable gate array (FPGA)-based sensor nodes are gaining popularity for Internet of Things (IoT) applications due to their flexible hardware reconfigurability. RISC32 is a recently proposed synchronous IoT processor, targeting the FPGA-based sensor nodes. However, its dynamic energy consumption is relatively high as its circuit components often activate with each tick of the global clock, irrespective of the actual need. RISC32-LP presented various power reduction techniques to enhance its energy efficiency, but it still necessitates the use of a constantly switching global clock in many parts of the system for synchronization. In view of that, this research work proposed a novel asynchronous processor design (RISC32-A) to significantly minimize the switching events in RISC32, thus lowering its overall dynamic energy consumption. Besides, this research work also proposed an adaptive asynchronous pipeline structure that allows selective pipeline stages to be dynamically skipped, merged, split, and stalled during the program run-time for optimal performance-energy tradeoffs. A novel pipeline stage skipping solution was introduced, which considers various instruction types and prevents wasteful data movement for the early-completed instructions. Additionally, a two-phase collapsible register handshake component with normally closed flip-flops was presented for enhanced dynamic power savings. Finally, a new pipeline halting solution was proposed to eliminate the decoding and forwarding hardware overheads found in the prior works. Experimental results show that the proposed RISC32-A can achieve an average reduction of$\approx 83.71\%$and$\approx 65.98\%$in dynamic energy consumption when compared to its synchronous counterpart (RISC32) and its optimal low-power synchronous counterpart (RISC32-LP), respectively.
Min-An Yong, Kai Ming Mok, Wai-Kong Lee, Shen-Khang Teoh, Denis Chee-Keong Wong
IEEE Internet Things J.5
2025 A signature scheme constructed from zero knowledge argument of knowledge for the subgraph isomorphism problem
Chii Liang Ng, Denis Chee-Keong Wong, Gek Ling Chia, Bok-Min Goi, Wai-Kong Lee, Wun-She Yap
Theor. Comput. Sci.2
2024 To Deploy New or to Deploy More?: An Online SFC Deployment Scheme at Network Edge
abstract
Service Function Chaining (SFC) dynamically links multiple Virtual Network Functions (VNFs) to provide flexible and scalable network services for network entities and users. Implementing SFCs at the network edge provides instant VNF service yet is confined by the limited edge resources. Existing strategies suggest either to deploy new VNFs for diverse service provision or to deploy more installed VNFs for reliable service provision. However, these one-sided optimizations fail to realize comprehensive improvements in the network service quality. To this end, the motivation of this paper is to consider a more comprehensive SFC deployment plan to provide more efficient network services. In this paper, we propose DeepSFC, an online SFC deployment scheme at network edge. Our DeepSFC considers the impact of resource allocations and deployment locations on the average latency of overall service requests. It realizes an elegant trade-off between the diversity and the availability of SFCs by adopting the Deep Reinforcement Learning (DRL) method. To be specific, we first determine the type and number of VNFs that need to be deployed. Thereafter, we optimize the deployment locations of these chosen VNFs in the service chain, considering the impact of dynamic bandwidth in the real network. For more general scenarios wherein users’ service requirements change or the deployed server crashes, we further relocate the VNF deployment with the joint consideration of performance degradation and migration cost. Evaluation results show that DeepSFC outperforms its competitors in various experimental settings and responds the requests with lower average latency.
Zongyang Yuan, Lailong Luo, Deke Guo, Denis Chee-Keong Wong, Geyao Cheng, Bangbang Ren, Qianzhen Zhang
IEEE Internet Things J.4
2024 Cryptanalysis of an image encryption scheme based on two-point diffusion strategy and Henon map
Kuan-Wai Wong, Wun-She Yap, Bok-Min Goi, Denis Chee-Keong Wong, Guodong Ye
J. Inf. Secur. Appl.4
2023 KaratSaber: New Speed Records for Saber Polynomial Multiplication Using Efficient Karatsuba FPGA Architecture
abstract
SABER is a round 3 candidate in the NIST Post-Quantum Cryptography Standardization process. Polynomial convolution is one of the most computationally intensive operation in Saber Key Encapsulation Mechanism, that can be performed through widely explored algorithms like the schoolbook polynomial multiplication algorithm (SPMA) and Number Theoretic Transform (NTT). While SPMA multiplier has a slow latency performance, the NTT-based multiplier usually requires large hardware. In this work, we propose KaratSaber, an optimized Karatsuba polynomial multiplier architecture with a balanced hardware efficiency (throughput-per-slice, TPS) compared to NTT and SPMA based designs. KaratSaber employs several techniques for an efficient design: a parallel grid input technique for efficient pre-processing stage in Karatsuba-based polynomial multiplier, a novel instruction code result-mapping technique catering the negacyclic operations improves the post-processing stage efficiency, a double multiplicand shifter-based multiplier doubles the throughput at the multiplication stage. Combining these three techniques, the proposed KaratSaber architecture is 7.47 × faster compared to the state-of-the-art SPMA Saber architecture at the expense of 4.96 × additional hardware resources; making KaratSaber 46.04% more area-time efficient. When compared to LWRPro, a recent Karatsuba Saber architecture, KaratSaber architecture achieves a 2.11 × higher throughput by only utilizing 1.92 × additional hardware; thus gaining a 10.44% improvement in area-time efficiency.
Zheng-Yan Wong, Denis Chee-Keong Wong, Wai-Kong Lee, Kai Ming Mok, Wun-She Yap, Ayesha Khalid
IEEE Trans. Computers2
2021 Freeness Problem for Matrix Semigroups of Parikh Matrices
abstract
Since the undecidability of the mortality problem for 3 × 3 matrices over integers was proved using the Post Correspondence Problem, various studies on decision problems of matrix semigroups have emerged. The freeness problem in particular has received much attention but decidability remains open even for 2 × 2 upper triangular matrices over nonnegative integers. Parikh matrices are upper triangular matrices introduced as a generalization of Parikh vectors and have become useful tools in studying of subword occurrences. In this work, we focus on semigroups of Parikh matrices and study the freeness problem in this context.
Wen Chean Teh, Adrian Atanasiu, Denis Chee-Keong Wong
Fundam. Informaticae3
2021 Novel Postquantum MQ-Based Signature Scheme for Internet of Things With Parallel Implementation
abstract
Internet of Things (IoT) is a paradigm shifting technology that enables many innovative applications in the near future. Proactive measures are required to protect such architecture from cyber attacks. One of the most important security issues in this architecture is the authentication of edge nodes, which can be resolved through the deployment of digital signatures. However, existing standardized digital signatures are vulnerable to attacks from quantum computers, which can be unsafe in the near future. In this article, we propose a new signature scheme based on multivariate polynomials with efficient key and signature sizes, which is resistant to quantum computer attacks. The proposed scheme is also very friendly to parallel implementation, enabling efficient deployment of edge nodes authentication at high throughput. When implemented on a GPU device, the proposed scheme can generate 113 signatures/s and verify 120 signatures/s, which is 12.56× and 10.00× faster than a serial implementation in CPU.
Sedat Akleylek, Meryem Soysaldi, Wai-Kong Lee, Seong Oun Hwang, Denis Chee-Keong Wong
IEEE Internet Things J.5
2021 Parallel implementation of Nussbaumer algorithm and number theoretic transform on a GPU platform: application to qTESLA
Wai-Kong Lee, Sedat Akleylek, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi, Seong Oun Hwang
J. Supercomput.3
2020 Optimized IoT Cryptoprocessor Based on QC-MPDC Key Encapsulation Mechanism
abstract
The key encapsulation mechanism (KEM) is an important cryptographic tool to protect communication in the Internet of Things (IoT). In the near future, classical algorithms used to construct KEMs, such as RSA and elliptic curve cryptography, will be vulnerable to attacks from quantum computers. Recently, Yamada et al. proposed the quasicyclic medium density parity check (QC-MDPC) KEM, which is considered one of the most advanced code-based cryptosystems to resist quantum attacks. In this article, an optimized implementation of QC-MDPC KEM for IoT applications is presented. Our main contributions are threefold: 1) the fastest QC-MDPC McEliece decryption in field-programmable gate array (FPGA); 2) the first QC-MDPC KEM implementation in FPGA; and 3) the first iteration count attack-resistant QC-MDPC decoder in FPGA. To improve the decryption speed, we introduce a novel customized rotation engine (CRE) and incorporated several recent techniques reported in the literature, including adaptive threshold and Hamming weight estimation. The best-achieved throughput in our implementation on Xilinx Virtex 7 FPGA is 12.7% faster than the state-of-the-art result reported by Heyse et al. The proposed CRE was then integrated with QC-MDPC KEM to produce a fast and secure KEM. Furthermore, to prevent timing attacks demonstrated recently, a constant-time implementation of the QC-MDPC McEliece decoder was presented.
Jun-Hoe Phoon, Wai-Kong Lee, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi, Raphael C.-W. Phan
IEEE Internet Things J.3
2020 Cryptanalysis of genetic algorithm-based encryption scheme
Kuan-Wai Wong, Wun-She Yap, Denis Chee-Keong Wong, Raphael C.-W. Phan, Bok-Min Goi
Multim. Tools Appl.3
2020 Area-Time-Efficient Code-Based Postquantum Key Encapsulation Mechanism on FPGA
abstract
Postquantum cryptography attracts a lot of attention from the research community recently due to the emergence threat from quantum computer toward the conventional cryptographic schemes. In view of that, NIST had initiated the standardization process in 2017. Bit flipping key encapsulation (BIKE) designed by Aragon et al. is one of the promising code-based schemes among the round-3 candidates. BIKE utilizes a quasi-cyclic medium density parity check (QC-MDPC) code and incorporates a few variants derived from the McEliece, Niederreiter, and Ouroboros schemes. In this article, we present efficient and constant time implementation of BIKEI and BIKE-III in field-programmable gate array (FPGA), which has the best area-time efficiency so far. We proposed modification to the original one-round bit flipping algorithm to achieve more area-time-efficient decoding in hardware, which achieved latency of 464.73 and 556.52 μs for BIKE-I and BIKE-III, respectively, in Virtex-7. A pipelined key encapsulation architecture is proposed to speedup the key encapsulation of BIKE-I and BIKE-III, achieving the latency of 146.47 and 153.25 μs on the same FPGA platform. Considering the Artix-7 FPGA platform, our combined key generation and encapsulation module for BIKE-I is also three more area-time efficient compared with the state-of-the-art BIKE-I implementation by Aragon et al.
Jun-Hoe Phoon, Wai-Kong Lee, Denis Chee-Keong Wong, Wun-She Yap, Bok-Min Goi
IEEE Trans. Very Large Scale Integr. Syst.3