EDBT 2026 Demo / reviewers in the wild / expert
Bei Wang 0013
dblp:08/6391-13
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-4760-3362ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Efficient Fully-Pipelined Hardware Architecture for Optimized Sparse Polynomial Multiplication in CRYSTALS-Dilithium
Bei Wang 0013, Zeren Zhu, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2026 | Optimized NTT Architecture Based on the Plantard Algorithm for ML-KEM and ML-DSAabstractModular multiplication is a vital operation in the Number Theoretic Transform (NTT), significantly enhancing polynomial multiplication in Post-Quantum Cryptography (PQC). The design efficiency of modular multiplication directly influences the computational performance of polynomial computation units. This work marks the first hardware-oriented improvement of the Plantard algorithm, optimizing the NTT architecture. We modify the Plantard algorithm and propose three innovative enhanced versions tailored for lattice-based cryptography (LBC). By employing pre-processed twiddle factors for result correction and eliminating an additional constant multiplication, we greatly simplify the computation steps. Based on these improvements, we further design a lightweight BRAM-free iterative NTT and a high-speed Multi-path Delay Commutator (MDC) pipelined NTT, both targeting the ML-KEM and ML-DSA parameter sets. Implementation results on the Xilinx Artix-7 platform demonstrate that our Plantard_preω design reduces slice usage by 22% to 53% and delay by 41.1% to 44.3% compared to existing algorithms like Barrett and K2RED. Furthermore, our iterative NTT design achieves the minimal area-time product (ATP) among state-of-the-art implementations, with reductions of 61.3% and 40.8% in ENS, and frequency increases of 72.7% and 145.4% for ML-KEM and ML-DSA, respectively. The pipelined NTT design also reduces delay by 23.7% and ATP by 4.9%, showcasing the compactness and superior performance of our approach. Bei Wang 0013, Ziying Ni, Mengxue Li, Fei Lyu 0002, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 2 |
| 2025 | An Efficient Hardware Implementation of Improved Plantard Mod-Multiplication for Lattice-Based CryptographyabstractThe modular multiplication (mod-multiplication) algorithm is an essential operation in lattice-based cryptography (LBC) that utilizes the Number Theoretical Transform (NTT) for polynomial multiplication. An efficient mod-multiplication algorithm determines the computational efficiency and performance of the entire polynomial multiplier/NTT computation unit. In this manuscript, we propose an improved Plantard modmultiplication algorithm for the NTT in Kyber and Dilithium, which not only reduces one multiplication but also eliminates the post-processing operations compared with the original Plantard algorithm. Additionally, we design an optimized hardware implementation for the improved Plantard mod-multiplication algorithm. Based on the Xilinx Artix-7 platform, when compared with state-of-the-art designs, our improved Plantard algorithm reduces the number of slices by 22.2%∼53.3% for Kyber and 18.4%∼35.4% for Dilithium, while boosting hardware efficiency by 43.9%∼55.9% for Kyber and 33.7%∼54.4% for Dilithium. Overall, our improved Plantard algorithm shows significant advantages in resource consumption and computational speed. Mengxue Li, Bei Wang 0013, Fei Lyv, Weiqiang Liu 0001, Yijun Cui |
ISCAS | 3 |
| 2025 | Ultra-compact and Side-channel Resistant Design of FIFO-based NTT Core for PQCsabstractCryptographic algorithms like CRYSTALS-Kyber and Dilithium might be insecure with their naive implementation facing side-channel attacks (SCA). This work presents a compact implementation of Number Theoretic Transform (NTT) with shuffling countermeasure against power analysis attacks (PA). At first, a compact FIFO-only Shuffler module is presented to perform group-wise first-index randomization (FIR). A modified butterfly (BF) unit using optimized modulus reduction is then promoted to restore the misaligned data flow, which is critical for forming efficient shuffle pattern. The shuffler module and BF unit are then used in a pipelined BRAM-free NTT baseline. Through efficient shuffling, the proposed design maintains compactness akin to its baseline while enhancing robust SCA resistance with a permutation space of up to 2494. Compared to its prior state-of-the-art designs, the proposed NTT core presents an improvement of 53.2% in area-time trade-off while offering ×1.7 times more bits of randomness to improve hardware security. Jiatong Tian, Yijun Cui, Ziying Ni, Bei Wang 0013, Fei Lyv, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2025 | A Lightweight and Efficient BRAM-free NTT Unit for Crystals-DilithiumabstractDuring the standardization of post-quantum cryptography by the National Institute of Standards and Technology (NIST), the lattice-based Crystals-Dilithium algorithm was selected as the standardized digital signature scheme. This work designs a lightweight and efficient BRAM-free Number Theoretic Transform (NTT) unit, which is a major bottleneck for Crystals-Dilithium. Firstly, we propose an improved parallel modular multiplication based on the K-RED algorithm, effectively reducing resource consumption and shortening the critical path. Furthermore, a BRAM-free iterative NTT architecture is designed, utilizing three first-in-first-out (FIFO) buffers to store intermediate data. Evaluated on the Xilinx Artix-7 and Zynq UltraScale+ platforms, our proposed NTT architecture presents the best hardware efficiency with less resource consumption. Experimental results show that our design is 36.7%-80.9% reduced in terms of resource consumption and is 31.5%-89.9% better in terms of hardware efficiency compared with state-of- the-art works. Junjie Zhong, Bei Wang 0013, Zeren Zhu, Weiqiang Liu 0001, Yijun Cui |
ISCAS | 2 |
| 2025 | High-Performance Hardware Implementation of Crystals-Dilithium Based on Improved MDC-NTTabstractThe growing threat of quantum computing to traditional cryptographic systems has necessitated the development of robust post-quantum algorithms. Crystal-Dilithium, recently standardized by NIST after a three-round competition, is a leading lattice-based digital signature algorithm designed to meet this need. However, conventional hardware implementations of Dilithium often suffer from inefficiencies and performance bottlenecks. To address these weaknesses, this work presents an optimized hardware design for Dilithium across all security levels. The proposed design features a parallel modular multiplication unit, and an enhanced scaling method to reduce bit width and minimize calibration. Additionally, an improved radix-2 Multipath Delay Commutator Number Theoretic Transform (MDC-NTT) and pipelined parallelization using FIFO and BRAM-based buffers are integrated to maximize operating frequency. Evaluated on the Xilinx Artix-7 platform, our implementation achieves a peak frequency of 191 MHz, delivering speedups of 26.3%, 32.5% and 29.6% for key generation, signature generation and signature verification respectively, compared with state-of-the-art works at the highest security level, along with superior hardware efficiency. Yijun Cui, Junjie Zhong, Bei Wang 0013, Tianyu Xu 0002, Chenghua Wang, Weiqiang Liu 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Lattice-based Multi-Stage Secret Sharing 3D Secure Encryption SchemeabstractWith the widespread deployment of three-dimensional (3D) models in industry and daily life, protecting the security of this data becomes crucial. Additionally, three-dimensional (3D) models may be distributed to users with varying security levels, necessitating distinct visualizations for each user. Recent research proposes 3D model encryption method that facilitates distinct visualizations post-decryption through hierarchical decryption. However, this method permits the decryption of 3D models at varying visual security levels based on user privileges. It has potential security vulnerabilities concerning key management and simultaneously limits its capacity to address diverse user requirements. To address this, a multi-stage secret sharing mechanism is integrated into the existing hierarchical encryption framework to bolster the security of hierarchical keys. When combined with lattice-based cryptography techniques, it ensures that only users with adequate shares can decrypt the corresponding 3D model hierarchy, achieving distinct visual effects while maintaining secret key security under diverse user needs. Experimental results demonstrate that the scheme effectively enhances security while maintaining data integrity and availability. Yinghao Wu, Bei Wang 0013, Yijun Cui |
TrustCom | 5 |