Weijia Wang 0003

dblp:30/6437-3 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0001-6982-2537ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 20 · 6 first-author · 13 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hardware masking with buffer chain
abstract
Abstract Side-channel attacks pose a major threat to cryptographic implementations, as they can exploit physical leakages to recover secret information. Masking is one of the most widely adopted countermeasures, aiming to protect sensitive intermediate values by randomization. However, when deployed in hardware, masking faces the challenges from the glitch leakage, which can easily undermine the security integrity of masking techniques and affect the foundational independent assumptions upon which they are based. To address the intricacies of hardware implementation, circuit separation using registers has emerged as a straightforward method. In this paper, we investigate low-latency hardware masking by exploring the use of buffers (rather than registers) to prevent glitch propagation. Rather than directly inserting buffers into the circuit path, our approach involves employing a chain of buffers to generate signals that serve as controls, thereby synchronizing blocks that require sequential computation. This significantly reduces the power consumption of the shielding circuit while also decreasing latency within the circuit.
Guofeng Qin, Chun Guo 0002, Hao Cheng 0009, Weijia Wang 0003
Cybersecur.6
2026 More Practical and Robust: Enhancing Simple Power Analysis on Cryptosystems With Double Clustering
abstract
The widespread use of public key cryptographic algorithms in embedded devices has made them a primary target for side-channel analysis. Clustering-based Simple Power Analysis (SPA) poses a significant threat to public key implementations by inferring secret keys through the identification of distinguishable patterns in side-channel information. However, traditional clustering-based SPA methods are highly sensitive—even to non-key-dependent patterns—thereby limiting their robustness and practical applicability. To address these limitations, this paper proposes a double clustering method that enhances the flexibility, accuracy, and robustness of clustering-based SPA. By progressively adjusting the the number of clusters, the method adaptively identifies optimal clustering configurations, mitigating the need for fixed assumptions and improving resistance to noise and other interfering factors. Experiments covering multiple cryptographic algorithms, hardware platforms, and countermeasure settings demonstrate that the proposed method consistently outperforms traditional clustering-based SPA methods.
Annyu Liu, Weijia Wang 0003, An Wang 0001
IEEE Internet Things J.3
2025 Tighter Security Notions for a Modular Approach to Private Circuits
Juelin Zhang, Yu Yu 0001, Weijia Wang 0003
EUROCRYPT (8)4
2025 Thorough Power Analysis on Falcon Gaussian Samplers and Practical Countermeasure
Xiuhan Lin, Shiduo Zhang, Yang Yu 0008, Weijia Wang 0003, Qidi You, Ximing Xu 0003, Xiaoyun Wang 0001
PKC (1)4
2025 Finding More Hints - Improved Power Analysis Attacks on Dilithium
Tianfu Zhang, Yu Yu 0001, Weijia Wang 0003
IEEE Trans. Inf. Forensics Secur.7
2024 Leakage-Resilient Circuit Garbling
abstract
Due to the ubiquitous requirements and performance leap in the past decade, it has become feasible to execute garbling and secure computations in settings sensitive to side-channel attacks, including smartphones, IoTs and dedicated hardwares, and the possibilities have been demonstrated by recent works. To maintain security in the presence of a moderate amount of leaked information about internal secrets, we investigate leakage-resilient garbling. We augment the classical privacy, obliviousness and authenticity notions with leakages of the garbling function, and define their leakage-resilience analogues. We examine popular garbling schemes and unveil additional side-channel weaknesses due to wire label reuse and XOR leakages. We then incorporate the idea of label refreshing into the GLNP garbling scheme of Gueron et al. and propose a variant GLNPLR that provably satisfies our leakage-resilience definitions. Performance comparison indicates that GLNPLR is 60X (using AES-NI) or 5X (without AES-NI) faster than the HalfGates garbling with second order side-channel masking, for garbling AES circuit when the bandwidth is 2Gbps.
Chun Guo 0002, François-Xavier Standaert, Weijia Wang 0003, Xiao Wang 0012
CCS5
2024 Fast Fourier Transform and Gaussian Sampling Instructions Designed for FALCON Digital Signature
Tao-Yun Wang, Shuai-Yu Chen, Lu Li 0006, Weijia Wang 0003
Inscrypt (2)4
2024 Improved Masking Multiplication with PRGs and Its Application to Arithmetic Addition
abstract
At Eurocrypt 2020, Coron et al. proposed a masking technique allowing the use of random numbers from pseudo‐random generators (PRGs) to largely reduce the use of expansive true‐random generators (TRNGs). For security against d probes, they describe a construction using 2 d PRGs, each of which is fed with at most 2 d random variables in a finite field, resulting in a randomness requirement of . In this paper, we improve the technique on multiple frontiers. On the theoretical level, we push the limits of the randomness requirement by providing an improved masking multiplication using only d PRGs, each of which is fed with d random variables, saving more than half random bits. On the practical level, considering that the masking of arithmetic addition usually requires more randomness (than multiplication), we apply the technique to the algorithm proposed at FSE 2015 that is a very efficient scheme performing arithmetic addition modulo 2 w . It significantly reduces the randomness cost of masked arithmetic addition, and further advocates the advantage of masking with PRGs. Furthermore, we apply our masking scheme to the Speck , XTEA , and Sparkle , and provide the first (to the best of our knowledge) higher order masked implementations for the ciphers using ARX structure.
Qian Sui, Fanjie Ji, Chun Guo 0002, Weijia Wang 0003
IET Inf. Secur.5
2024 Compact Instruction Set Extensions for Kyber
abstract
Kyber is the only post-quantum cryptography (PQC) key encapsulation mechanism in the National Institute of Standards and Technology PQC project. This brief investigates the design of compact instruction set extensions (ISEs) for Kyber. We focus on implementing number-theoretic transform (NTT) and propose a hardware design of the modular multiplication based on an optimized$k^{2}$-reduction. Compared to other works, our design is more compact since the optimized$k^{2}$-reduction comprises multiplications with significantly smaller multipliers than Montgomery reduction and Barrett reduction. Then, we integrate the$k^{2}$-reduction into an instruction for the butterfly transformation. We also propose auxiliary instructions that can switch the half words between two registers to facilitate the rearranging coefficients in NTT. To showcase the advantage of the instructions, we implement the ISEs in a chip design for the Hummingbird E203 core. Compared to the software implementation on RISC-V with assembly code, our co-design implementations for NTT show a speedup by a factor of 2.6. Besides, the area overhead is 93 LUTs and 1 DSP without any additional resources of FFs and RAMs using Artix-7 FPGA, which is more compact than previous software–hardware co-designs of Kyber.
Lu Li 0006, Guofeng Qin, Yang Yu 0008, Weijia Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 ISA Extensions of Shuffling Against Side-Channel Attacks
abstract
Shuffling is a time-randomized countermeasure against side-channel attacks. To achieve effective protections, shuffling is usually combined with other countermeasures, such as the masking. It requires the shuffling to be as efficient as possible. In this work, we describe an instruction set extensions (ISEs) for shuffling countermeasure. Our ISEs focuses on the generation of random permutations, which is the most difficult part to deploy the shuffling in microprocessors. The Thorp shuffling is implemented in hardware, enabling the instruction to generate random permutations. We design new ISEs compatible to the RISC-V standard instruction set format. Then, we present applications of our ISEs by giving two combinations of shuffling and masking, which can be regarded as promising software–hardware co-designs of side-channel countermeasures. At last, we embed the ISEs to the RISC-V core called tinyriscv, and evaluate the silicon overhead and the side-channel security of the shuffled masked AND operation. The evaluation shows that the new instruction can significantly improve the security of masking countermeasures.
Jiayun Zhou, Guofeng Qin, Lu Li 0006, Chun Guo 0002, Weijia Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Compact Instruction Set Extensions for Dilithium
abstract
Post-quantum cryptography is considered to provide security against both traditional and quantum computer attacks. Dilithium is a digital signature algorithm that derives its security from the challenge of finding short vectors in lattices. It has been selected as one of the standardizations in the NIST post-quantum cryptography project. Hardware-software co-design is a commonly adopted implementation strategy to address various implementation challenges, including limited resources, high performance, and flexibility requirements. In this study, we investigate using compact instruction set extensions (ISEs) for Dilithium, aiming to improve software efficiency with low hardware overheads. To begin with, we propose tightly coupled accelerators that are deeply integrated into the RISC-V processor. These accelerators target the most computationally demanding components in resource-constrained processors, such as polynomial generation, Number Theoretic Transform (NTT), and modular arithmetic. Next, we design a set of custom instructions that seamlessly integrate with the RISC-V base instruction formats, completing the accelerators in a compact manner. Subsequently, we implement our ISEs in a chip design for the Hummingbird E203 core and conduct performance benchmarks for Dilithium utilizing these ISEs. Additionally, we evaluate the resource consumption of the ISEs on FPGA and ASIC technologies. Compared to the reference software implementation on the RISC-V core, our co-design demonstrates a remarkable speedup factor ranging from 6.95 to 9.96. This significant improvement in performance is achieved by incorporating additional hardware resources, specifically, a 35% increase in LUTs, a 14% increase in FFs, 7 additional DSPs, and no additional RAM. Furthermore, compared to the state-of-the-art approach, our work achieves faster speed performance with a reduced circuit cost. Specifically, the usage of additional LUTs, FFs, and RAMs is reduced by 47.53%, 50.43%, and 100%, respectively. On ASIC technology, our approach demonstrates 12, 412 cell counts. Our co-design provides a better tradeoff implementation on speed performance and circuit overheads.
Lu Li 0006, Guofeng Qin, Shuaiyu Chen, Weijia Wang 0003
ACM Trans. Embed. Comput. Syst.5
2023 Towards Minimizing Non-linearity in Type-II Generalized Feistel Networks
Chun Guo 0002, Weijia Wang 0003
CANS3
2023 Improved Power Analysis Attacks on Falcon
Shiduo Zhang, Xiuhan Lin, Yang Yu 0008, Weijia Wang 0003
EUROCRYPT (4)4
2023 Shorter Linkable Ring Signature Based on Middle-Product Learning with Errors Problem
abstract
Abstract DualRing is a novel generic construction introduced by Yuen et al. (CRYPTO’21), which can transform a special kind of (Type-T*) canonical identification scheme to a ring signature scheme. Compared with the classical approaches, this method can get a shorter signature. In this paper, we construct a new middle-product learning with errors (MPLWE)-based ring signature scheme by using this framework. Specifically, we propose a new MPLWE-based identification scheme, which is compatible with the DualRing, then we obtain a ring signature scheme by using DualRing framework. We also show how to achieve linkability from this ring signature by using a collision resistant hash function. In the end, we provide available parameter options for our (linkable) ring signature scheme. Under these parameters, the signature size of our linkable ring signature is $2-40 \times $ shorter (depending on the ring size) than the previous MPLWE-based scheme by Das et al. (Africacrypt’19).
Hao Lin 0012, Shifeng Sun 0001, Joseph K. Liu, Weijia Wang 0003
Comput. J.5
2023 Bit-Sliced Implementation of SM4 and New Performance Records
abstract
SM4 is a popular block cipher issued by the Office of State Commercial Cryptography Administration (OSCCA) of China. In this paper, we use the bit‐slicing technique that has been shown as a powerful strategy to achieve very fast software implementations of SM4. We investigate optimizations on two frontiers. First, we present a more efficient bit‐sliced representation for SM4, which enables running 64 blocks in parallel with 256‐bit registers. Second, we describe an optimized algorithm for data form transformations, also allowing efficient implementations of SM4 under Counter (CTR) mode and Galois/Counter mode. The above optimizations contribute to a significant performance gain on one core compared with the state‐of‐the‐art results. This work is an extension of the conference paper at Inscrypt 2022, awarded the best paper award.
Lu Li 0006, Chun Guo 0002, Meiqin Wang 0001, Weijia Wang 0003
IET Inf. Secur.5
2022 How Fast Can SM4 be in Software?
Chun Guo 0002, Weijia Wang 0003
Inscrypt4
2022 SAND: an AND-RX Feistel lightweight block cipher supporting S-box-based security evaluations
Yanhong Fan 0001, Ling Sun 0001, Meiqin Wang 0001, Weijia Wang 0003, Chun Guo 0002
Des. Codes Cryptogr.8
2021 Forced Independent Optimized Implementation of 4-Bit S-Box
Yanhong Fan 0001, Weijia Wang 0003, Zhihu Li, Siu-Ming Yiu, Meiqin Wang 0001
ACISP2
2020 Packed Multiplication: How to Amortize the Cost of Side-Channel Masking?
Weijia Wang 0003, Chun Guo 0002, François-Xavier Standaert, Yu Yu 0001, Gaëtan Cassiers
ASIACRYPT (1)1
2019 Side-Channel Analysis for the Authentication Protocols of CDMA Cellular Networks
Chi Zhang 0061, Dawu Gu, Weijia Wang 0003, Xiangjun Lu, Zheng Guo 0001, Haining Lu
J. Comput. Sci. Technol.4
2019 Provable Order Amplification for Code-Based Masking: How to Avoid Non-Linear Leakages Due to Masked Operations
abstract
Code-based masking schemes have been shown to provide higher theoretical security guarantees than Boolean masking. In particular, one interesting feature put forward at CARDIS 2016 and then analyzed at CARDIS 2017 was the socalled security order amplification: under the assumption that the leakage function is linear, it guarantees that an implementation performing only linear operations will have a security order in the bounded moment leakage model larger than d - 1, where d is the number of shares. The main question regarding this feature is its practical relevance. First of all, concrete block ciphers do not only perform linear operations. Second, it may be that actual leakage functions are not perfectly linear (raising questions regarding what happens when one deviates from such assumptions). In this paper, we show that the issue of only linear operations can be provably avoided and that it is possible to obtain security order amplification for any functionality to implement. We then show that (not so) slightly non-linear leakage functions do not annihilate the nice properties (i.e., that the code-based schemes we consider remain interesting compared to the Boolean masking). We conclude with a performance evaluation of the proposals, showing that the performance overheads are moderate for a reasonable number of shares (we studied when the number of the shares d = 2,3,4). In additional, our results could be specified to the case of provable security for low entropy masking, which can be considered as a side bonus of our contributions. We give some preliminary results on how to construct the low entropy masking schemes with provable high security order against linear leakage.
Weijia Wang 0003, Yu Yu 0001, François-Xavier Standaert
IEEE Trans. Inf. Forensics Secur.1
2018 Similar operation template attack on RSA-CRT as a case study
Xiangjun Lu, Yang Li 0022, Lei Wang 0031, Weijia Wang 0003, Haihua Gu, Zheng Guo 0001, Dawu Gu
Sci. China Inf. Sci.6
2018 Ridge-Based DPA: Improvement of Differential Power Analysis For Nanoscale Chips
abstract
Differential power analysis (DPA), as a very practical type of side-channel attacks, has been widely studied and used for the security analysis of cryptographic implementations. However, as the development of the chip industry leads to smaller technologies, the leakage of cryptographic implementations in nanoscale devices tends to be nonlinear (i.e., leakages of intermediate bits are no longer independent) and unpredictable. These phenomena make some existing side-channel attacks not perfectly suitable, i.e., decreasing their performance and making some common used prior power models (e.g., Hamming weight) to be much less respected in practice. To solve the above issues, we introduce the regularization process from statistical learning to the area of side-channel attack and propose the ridge-based DPA. We also apply the cross-validation technique to search for the most suitable value of the parameter for our new attack methods. In addition, we present theoretical analyses to deeply investigate the properties of ridge-based DPA for nonlinear leakages. We evaluate the performance of ridge-based DPA in both simulation-based and practical experiments, comparing to the state-to-the-art DPAs. The results confirm the theoretical analysis. Further, our experiments show the robustness of ridge-based DPA to cope with the difference between the leakages of profiling and exploitation power traces. Therefore, by showing a good adaptability to the leakage of the nanoscale chips, the ridge-based DPA is a good alternative to the state-to-the-art ones.
Weijia Wang 0003, Yu Yu 0001, François-Xavier Standaert, Zheng Guo 0001, Dawu Gu
IEEE Trans. Inf. Forensics Secur.1
2017 Trace Augmentation: What Can Be Done Even Before Preprocessing in a Profiled SCA?
Sihang Pu, Yu Yu 0001, Weijia Wang 0003, Zheng Guo 0001, Dawu Gu
CARDIS3
2017 Ridge-Based Profiled Differential Power Analysis
Weijia Wang 0003, Yu Yu 0001, François-Xavier Standaert, Dawu Gu, Chi Zhang 0061
CT-RSA1
2016 Inner Product Masking for Bitslice Ciphers and Security Order Amplification for Linear Leakages
Weijia Wang 0003, François-Xavier Standaert, Yu Yu 0001, Sihang Pu, Zheng Guo 0001, Dawu Gu
CARDIS1
2015 Evaluation and Improvement of Generic-Emulating DPA Attacks
Weijia Wang 0003, Yu Yu 0001, Zheng Guo 0001, François-Xavier Standaert, Dawu Gu
CHES1