EDBT 2026 Demo / reviewers in the wild / expert
Wen Wang 0007
dblp:29/4680-7
· DBLP profile ↗
12ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0002-3019-2336ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Security and privacy · 4 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Engineering Practical Rank-Code-Based Cryptographic Schemes on Embedded Hardware. A Case Study on ROLLOabstractIn this paper, we investigate the practical performance of rank-code based cryptography on FPGA platforms by presenting a case study on the quantum-safe KEM scheme based on LRPC codes called ROLLO, which was among NIST post-quantum cryptography standardization round-2 candidates. Specifically, we present an FPGA implementation of the encapsulation and decapsulation operations of the ROLLO KEM scheme with some variations to the original specification. The design is fully parameterized, using code-generation scripts to support a wide range of parameter choices for security levels specified in ROLLO. At the core of the ROLLO hardware, we presented a generic approach for hardware-based Gaussian elimination, which can process both non-singular and singular matrices. Previous works on hardware-based Gaussian elimination can only process non-singular ones. However, a plethora of cryptosystems, for instance, quantum-safe key encapsulation mechanisms based on rank-metric codes, ROLLO and RQC, which are among NIST post-quantum cryptography standardization round-2 candidates, require performing Gaussian elimination for random matrices regardless of the singularity. To the best of our knowledge, this work is the first hardware implementation for rank-code-based cryptographic schemes. The experimental results suggest rank-code-based schemes can be highly efficient. Jingwei Hu 0001, Wen Wang 0007, Kris Gaj, Huaxiong Wang |
IEEE Trans. Computers | 2 |
| 2023 | Scalable and Conflict-Free NTT Hardware Accelerator Design: Methodology, Proof, and ImplementationabstractNumber theoretic transform (NTT) is useful for the acceleration of polynomial multiplication, which is the main performance bottleneck in the next-generation cryptographic schemes. Different NTT-based cryptographic algorithms have different security settings. The diverse application scenarios introduce different cost-performance tradeoffs and hardware constraints. Motivated by the emerging demand for more versatile NTT hardware accelerators, we propose a new design methodology that can generate area-efficient and high-performance NTT accelerators for any length and modulus of NTT polynomials and single processing element (PE) or PE array with a varying number of layers. The proposed NTT accelerator architecture pivots on a conflict-free memory access pattern for adaptation to different combinations of security and PE array configuration parameters. The proposed memory access pattern is formally proved to be conflict-free for any parametric configurations. The criterion for read-after-write conflict without pipeline stall is also established. Our proposed design methodology can produce NTT accelerators with single PE or multilayer PE array for different polynomial size and modulus, with hardware area and computational efficiency comparable to accelerators customized for a fixed set of parameters. Our proposed methodology produces parameterized accelerator with higher scalability than the existing parameterized accelerator design. On average, the accelerators generated by our proposed method are 71.4% more area-time efficient. Up to 30.7% area-time reduction over the most area-time efficient state-of-the-art scalable NTT accelerator can be achieved for the same security parameters. Jianan Mu, Wen Wang 0007, Yizhong Hu, Chip-Hong Chang, Junfeng Fan, Jing Ye 0001, Yuan Cao 0003, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | An Efficient Full Hardware Implementation of Extended Merkle Signature SchemeabstractThis paper presents a full hardware implementation of the eXtended Merkle Signature Scheme (XMSS), a NIST approved and IETF RFC specified post-quantum cryptography (PQC) algorithm. An optimized node traversal is proposed to enable efficient memory utilization without compromising the computational latency of the L-tree and Merkle tree construction, which are two key components used for the compression of the Winternitz One-Time Signature (WOTS) public key in XMSS. The computation of the authentication path during signature generation has also been significantly sped up by our proposed hardware implementation of the Buchmann, Dahmen, and Schneider (BDS) algorithm. Our implementation has completely avoided the use of block random-access memory, which is known to be vulnerable to side-channel attacks. The memory requirement has been highly optimized for implementation with small flip-flop chains and register counters as pointers for fast data access. To the best of our knowledge, this is the first full hardware implementation of all threekey generation,signingandverificationoperations of XMSS. The design has been prototyped and evaluated on a 28 nm FPGA platform to demonstrate its performance improvements over the most efficient software and hardware/software co-design methods reported to date. Specifically, it increases the computational efficiency of the best reported XMSS implementation forkey generationandsignature generationby about 20% and 50%, respectively. It can also run at 10% higher clock speed than the fastest hardware implementation ofsignature verificationin FPGA with 8% lower hardware resource utilization. Yuan Cao 0003, Yanze Wu, Wen Wang 0007, Jing Ye 0001, Chip-Hong Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | The Cost to Break SIKE: A Comparative Hardware-Based Analysis with AES and SHA-3
Patrick Longa, Wen Wang 0007, Jakub Szefer |
CRYPTO (3) | 2 |
| 2020 | Optimization Space Exploration of Hardware Design for CRYSTALS-KYBERabstractPublic key cryptography is important in the global communication digital infrastructure. However, the emergence of quantum computer and Shor algorithm has greatly threatened the security of public key cryptography. The CRYSTALS-KYBER, as a lattice-based KEM algorithm, passed three rounds of a global solicitation for post-quantum cryptography algorithms held by the National Institute of Standards and Technology (NIST). This paper explores the implementation and optimization space of hardware design according to CRYSTALS-KYBER algorithm. We analyze its software code and try different strategies to optimize the hardware implementation, and conduct comparative analysis in terms of area and speed. The experimental results show that the performance can be greatly improved by moderately optimizing the loops. In comparison with optimal results of the work [12], our optimizations improve the performance by up to 74.6% for encapsulation algorithm and 54.4% for decapsulation algorithm. Zhiteng Chao, Jing Ye 0001, Wen Wang 0007, Yuan Cao 0003, Xiaowei Li 0001, Huawei Li 0001 |
ATS | 4 |
| 2020 | Pipeline-aware Logic Deduplication in High-Level Synthesis for Post-Quantum Cryptography AlgorithmsabstractWith the technical advance of quantum computers that can solve intractable problems for conventional computers, many of the currently used public-key cryptosystems become vulnerable. Recently proposed post-quantum cryptography (PQC) is secure against both classical and quantum computers, but existing embedded systems such as smart card can not easily support the PQC algorithms due to their much larger key sizes and more complex arithmetics. To accelerate the PQC algorithms, embedded systems have to embed the PQC hardware blocks, which can lead to huge hardware design costs. Although High-Level Synthesis (HLS) helps significantly reduce the design costs, current HLS frameworks produce inefficient hardware design for the PQC algorithms in terms of area and performance. This work analyzes common features of the PQC algorithms and proposes a new pipeline-aware logic deduplication method in HLS. The proposed method shares commonly invoked logic across hardware design while considering load balancing in pipeline and resolving dynamic memory accesses. This work implements FPGA hardware design of seven PQC algorithms in the round 2 candidates from the National Institute of Standards and Technology (NIST) PQC standardization process. Compared to commercial HLS framework, the proposed method achieves an area-delay-product reduction by 34.5%. Changsu Kim 0004, Yongwoo Lee 0001, Shinnung Jeong, Wen Wang 0007, Jakub Szefer, Hanjun Kim 0001 |
FPGA | 4 |
| 2020 | ASIC Accelerator in 28 nm for the Post-Quantum Digital Signature Scheme XMSSabstractThis paper presents the first 28 nm ASIC implementation of an accelerator for the post-quantum digital signature scheme XMSS. In particular, this paper presents an architecture for a novel, pipelined XMSS Leaf accelerator for accelerating the most compute-intensive step in the XMSS algorithm. This paper then presents the ASIC designs for both an existing non-pipelined accelerator architecture and the novel, pipelined XMSS Leaf accelerator. In addition, the performance of the 28 nm ASIC is compared to the same designs on 28 nm Artix-7 FPGA. The novel pipelined XMSS Leaf accelerator is 25% faster compared to the non-pipelined version in the ASIC, and both accelerator architectures have a 10 × lower power consumption than on the FPGA. The evaluation shows that the pipelining increases the frequency by 1.7× on the FPGA but only 1.2× on the ASIC, due to the critical path in the ASIC being in the memory. The non-pipelined XMSS Leaf accelerator is shown to have a significantly better area-delay and energy-delay metric on the ASIC, while the pipelined accelerator wins out in these metrics on the FPGA. Consequently, this work shows the different architectural decisions that need to be made between FPGA and ASIC designs, when selecting how to best implement post-quantum cryptographic accelerators in hardware. Prashanth Mohan, Wen Wang 0007, Bernhard Jungk, Ruben Niederhagen, Jakub Szefer, Ken Mai |
ICCD | 2 |
| 2019 | XMSS and Embedded Systems
Wen Wang 0007, Bernhard Jungk, Julian Wälde, Shuwen Deng, Naina Gupta 0001, Jakub Szefer, Ruben Niederhagen |
SAC | 1 |
| 2018 | Post-Quantum Cryptography on FPGAs: The Niederreiter Cryptosystem: Extended AbstractabstractOur invited presentation will give an introduction to major hardware building blocks needed to implement code-based cryptographic systems. We will present details of a modern, FPGA-based, constant-time implementation of the Niederreiter cryptosystem using binary Goppa codes, including modules for encryption, decryption, and key generation. The presentation will also include a brief summary of other existing implementations of code-based cryptographic systems and it will present research challenges for implementing such systems efficiently. Wen Wang 0007, Jakub Szefer, Ruben Niederhagen |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | FPGA-Based Niederreiter Cryptosystem Using Binary Goppa Codes
Wen Wang 0007, Jakub Szefer, Ruben Niederhagen |
PQCrypto | 1 |
| 2017 | FPGA-based Key Generator for the Niederreiter Cryptosystem Using Binary Goppa Codes
Wen Wang 0007, Jakub Szefer, Ruben Niederhagen |
CHES | 1 |
| 2016 | Design and implementation of open-source SATA III core for Stratix V FPGAsabstractSATA is the de-facto standard computer interface that connects a host, typically a computing device, to a persistent storage device, such as a hard drive or solid-state drive. In order for FPGA-based designs to be able to leverage the variety of persistent storage devices, a SATA core is needed. Over time, the SATA standard has been revised to provide greater bandwidth, with SATA III being the newest version of the standard. In this paper, we are the first to present a SATA III core designed for Altera Stratix V FPGAs. Our implementation is written using Verilog, and tested using an industry-standard SATA protocol analyzer. We evaluate the performance of our SATA core by measuring the throughput of random and sequential read and write operations using various hard drives and solid-state drives. In addition, we compare the complexity of our SATA III core implementation with those of the older SATA I and II open-source implementations, and show that SATA III is still feasible, using only about 11% of Stratix V FPGA resources. Sumedh Guha, Wen Wang 0007, Shafeeq Ibraheem, Mahesh Balakrishnan 0001, Jakub Szefer |
FPT | 2 |