Jinwei Pu

dblp:429/1143 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0004-2771-5409ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 S2PUF: A <1.77E-8 BER Schmitt Effect SRAM PUF with Load Tilt Masking for Resource-Constrained systems
Xijun Huang, Jinwei Pu, You Meng
ISCAS3
2026 An Area-Efficient ML-DSA Accelerator With Interleaved and Dynamic Execution
abstract
The lattice-based digital signature algorithm CRYSTALS-Dilithium has been standardized as ML-DSA following the NIST post-quantum cryptography (PQC) competition. Due to the high computational complexity and data interaction in ML-DSA, its hardware implementation faces problems of large area overhead and low efficiency. This work invokes multiple optimizations to achieve an area-efficient hardware accelerator for ML-DSA. Specifically, we employ a multi-clock strategy in the architecture to maximize module performance and introduce interleaved and dynamic execution techniques to improve parallelism and minimize wait delays between modules. Moreover, several optimized modules are designed to improve the efficiency of the proposed architecture, including a multimodal polynomial arithmetic module, a unified SHA-3 module with preload functionality, and a BRAM-based Usehint module. Compared to the state-of-the-art, our implementation achieves an improvement of$1.21 \sim 4.62 \times $,$1.25 \sim 5.31 \times $, and$1.37 \sim 5.49 \times $in area-time product (ATP) metric at three security strengths, respectively. This establishes our design as the most area-efficient ML-DSA accelerator to date.
Jinwei Pu, Jiliang Zhang 0002
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 A Low-Cost Local Masking Radix-4 NTT Against Soft-Analytical Side-Channel Attacks
abstract
The number theoretic transform (NTT) is essential for accelerating polynomial multiplication in lattice-based cryptography. However, it is vulnerable to soft-analytical side-channel attacks (SASCAs). Although local masking countermeasure provides theoretical resistance against such attacks, its direct implementation in Radix-4 NTT architecture leads to more than a 4 times increase in modular multiplications, resulting in substantial hardware overhead. To address this challenge, we propose the modular multiplication parallel mask sharing (MMPMS) scheme, which optimizes the modular multiplication parallelism of the Radix-4 butterfly units and shares random twiddle factors, thereby achieving a balance between hardware overhead and security. Then, we construct a complete local masking NTT/INTT algorithm and efficiently implement it on the Artix-7 field-programmable gate array (FPGA). Experimental results show that compared with the state-of-the-art local masking NTT, our scheme reduces the equivalent area and ATP overhead by more than 8.24 times and 6.74 times, respectively. In addition, a nonspecifict-test analysis indicates no significant side-channel leakage.
Congwei Chen, Jinwei Pu, Jianxiong Zhang 0003, Jiaying Liao, Ruidian Zhan, Yun Chen 0004, Shuting Cai
IEEE Trans. Very Large Scale Integr. Syst.2
2026 Highly Reliable RRAM-Based Physical Unclonable Function With Auto-Write Technique
abstract
Physical unclonable functions (PUFs) have attracted significant attention for high-security applications, particularly in the Internet of Things (IoT). While resistive random access memory (RRAM)-based PUF technology offers excellent energy efficiency and integration density, its reliability remains a major challenge. This work proposes a highly reliable RRAM-based PUF that utilizes the 4T2R cross-coupled PUF cell to extract the response. Additionally, an automatic write-back technique is introduced to enhance the initial resistance mismatch in RRAM, thereby improving the PUF reliability. Experimental results demonstrate that the proposed PUF achieves outstanding performance, with near-ideal uniqueness (50.02%), uniformity (50.11%), and robust randomness (passing NIST 800-22 tests). The design exhibits robust reliability across a wide operating range (from$- 40~^{\circ }$C to$120~^{\circ }$C and 0.7 to 1.4 V), achieving a bit error rate (BER) of$\lt 4.89\times 10^{-7}$. Furthermore, the energy consumption per bit is optimized to 4.8 fJ at 1.1 V,$25~^{\circ }$C, showing its suitability for energy-constrained applications.
Jinwei Pu, Shuwen Xin, Aolin Wang, Zhengxun Lai, You Meng
IEEE Trans. Very Large Scale Integr. Syst.3