EDBT 2026 Demo / reviewers in the wild / expert
Yazheng Tu
dblp:318/5058
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-1624-9500ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEDAR: A Compact and Efficient Decoder Architecture for RS-RM Code in HQC
Yazheng Tu, Tianyou Bao, Jiafeng Xie |
ISCAS | 1 |
| 2025 | HSPA: High-Throughput Sparse Polynomial Multiplication for Code-based Post-Quantum CryptographyabstractIncreasing attention has been paid to code-based post-quantum cryptography (PQC) schemes, e.g., HQC (Hamming Quasi-Cyclic) and BIKE (Bit Flipping Key Encapsulation), since they’ve been selected as the fourth-round National Institute of Standards and Technology (NIST) PQC standardization candidates. Though sparse polynomial multiplication is one of the critical components for HQC and BIKE, hardware-implemented high-performance sparse polynomial multiplier is rarely reported in the literature (due to its high-dimension and sparsity of polynomials involved in the computation). Based on this consideration, in this article, we propose two novel H igh-throughput S parse P olynomial multiplication A ccelerators (HSPA) for the mentioned two code-based PQC schemes. Specifically, we have designed the two accelerators based on two different implementation strategies targeting potential applications with different resource availability, i.e., one accelerator deploys a memory-based structure for computation while the other does not need memory usage. We have proposed three layers of coherent interdependent efforts to obtain the proposed accelerators. First, we have proposed two implementation strategies to execute the targeted sparse polynomial multiplication, i.e., a new parallel segment based accumulation (PSA) approach and a novel permutating-with-power (PWP)-based method. Then, the proposed two hardware accelerators are presented with detailed structural descriptions. Finally, field-programmable gate array (FPGA)-based implementation is presented to showcase the superior performance of the proposed accelerators. A proper comparison is also carried out to confirm the efficiency of the proposed designs. For instance, the proposed accelerator (using memory-based structure) has 56.84% and 80.25% less area-delay product (ADP) than the existing memory-based design (an extended high-speed version) on the UltraScale+ device, respectively, for n =17,669 and ω =75 (HQC) and n = 12,323 and ω =142 (BIKE). The proposed design strategy fits well with the two targeted code-based PQC schemes, which can be extended further to construct high-performance hardware cryptoprocessors. We hope the results of this work will be useful for the ongoing NIST PQC standardization process. Pengzhou He, Yazheng Tu, Tianyou Bao, Çetin Kaya Koç, Jiafeng Xie |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2025 | EMINEM: Efficient FPGA Implementation of Mixed-RadIx NTT Hardware AccElerators for NIST Post-QuantuM Cryptography Falcon, Dilithium, and HAWKabstractThe advent of quantum computing poses a significant threat to modern cryptography. To address this challenge, the National Institute of Standards and Technology (NIST) has initiated the Post-Quantum Cryptography (PQC) standardization process, and several algorithms have been selected (a few are still under consideration in the additional standardization process). Among these schemes, lattice-based PQC has emerged as a promising approach and garnered substantial attention from the implementation community, especially on the hardware platforms. Notably, the Field-Programmable Gate Array (FPGA) has gained considerable attention as a convenient platform for hardware implementation, not only from NIST but also from the research community, as reflected by the related recommendations from NIST and the number of works reported recently. This work follows the existing trend of developing novel FPGA implementations for PQC. It is worth mentioning that the polynomial multiplication in these NIST lattice-based PQC algorithms can be implemented with Number Theoretic Transform (NTT) for efficiency. Nevertheless, there remains a lack of novel and universal NTT methods for polynomial multiplication at different sizes. For instance, for \(n=512\) (F alcon and HAWK), the existing works are mostly limited to the Radix-2 NTT (other methods like Radix-4 or Radix-8 cannot be directly applied). To fill the research gap, this article presents a novel design framework, i.e., Efficient Mixed-RadIx NTT hardware accElerators for NIST post-quantuM cryptography (EMINEM) , specially tailored for targeted schemes. Our design leverages Radix-4 for polynomial sizes of 256 and 1,024, while introducing a hybrid Radix-2/4 strategy for NTT of length 512 and achieving comparable performance to pure Radix-4 at other lengths. In total, our contributions include: (i) a generic Radix-4/Mixed-Radix NTT algorithm is proposed for \(n=256\) , 512, and 1,024; (ii) an efficient NTT hardware accelerator is designed with the help of a new memory access pattern and some optimization techniques; (iii) two types of butterfly architectures are developed to obtain pure Radix-4 time complexity and low resource usage, respectively; (iv) a detailed implementation and comparison showcase the superior performance of the proposed design strategy. Overall, the proposed strategy enables the efficient deployment of Mixed-Radix NTT in targeted NIST schemes, surpassing the limitations of the conventional Radix-2 approach for 512-length NTT designs. The proposed design offers a significant advancement in the field, facilitating efficient FPGA acceleration of PQC standards. Yazheng Tu, Jiafeng Xie |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2025 | SCOPE: Schoolbook-Originated Novel Polynomial Multiplication Accelerators for NTRU-Based PQCabstractTheNth-degree truncated polynomial ring units (NTRUs)-based postquantum cryptography (PQC) has drawn significant attention from the research communities, e.g., the National Institute of Standards and Technology (NIST) PQC standardization process selected algorithm Fast Fourier lattice-based compact (Falcon). Following the research trend, efficient hardware accelerator design for polynomial multiplication (an important component of the NTRU-based PQC) is crucial. Unlike the commonly used number theoretic transform (NTT) method, in this article, we have presented a novel SChoolbook-Originated Polynomial multiplication accElerators (SCOPE) design framework. Overall, we have proposed the schoolbook-based method in an innovative format to implement the targeted polynomial multiplication, first through a schoolbook-variant version and then through a Toeplitz matrix-vector product (TMVP)-based approach. Four layers of coherent and interdependent efforts have been carried out: 1) a novel lookup table (LUT)-based point-wise multiplier is proposed along with a related modular reduction technique to obtain optimal implementation; 2) a new hardware accelerator is introduced for the targeted polynomial multiplication, deploying the proposed point-wise multiplier; 3) the proposed architecture is extended to a TMVP-based polynomial multiplication accelerator; and 4) the efficiency of the proposed accelerators is demonstrated through implementation and comparison. Finally, the proposed design strategy is also extended to another NTRU-based scheme and other schoolbook- and toom-cook-based polynomial multiplications (used in other PQC), and obtains the same superior performance. We hope that the outcome of this research can impact the ongoing NIST PQC standardization process and related full-hardware implementation work for schemes like Falcon. Yazheng Tu, Shi Bai 0001, Jinjun Xiong, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Efficient Implementation of Ring-Binary-LWE-based Lightweight PQC Accelerator on the FPGA PlatformabstractPost-quantum cryptography (PQC) has gained sub-stantial attention from various communities recently. Along with the ongoing National Institute of Standards and Technology (NIST) PQC standardization process that targets the general-purpose PQC algorithms, the research community is also looking for efficient lightweight PQC schemes. Among this direction of efforts, Ring-Binary-Learning-with-Errors (RBLWE)-based encryption scheme (RBLWE-ENC) is regarded as a promising lightweight PQC fitting Internet-of-Things (IoT) and edge computing applications. As hardware implementation for PQC algorithms has become one of the major advances in the field, in this paper, we follow this trend to present an efficient implementation of RBLWE-ENC lightweight accelerator on the field-programmable gate array (FPGA) platform. Overall, we have demonstrated three coherent interdependent stages of efforts: (i) we have presented detailed derivation processes to formulate the proposed algorithmic operation; (ii) we have then implemented the proposed algorithm into a desired hardware accelerator; and (iii) we provided thorough complexity analysis and comparison to showcase the superior performance of the proposed accelerator over the state-of-the-art designs, e.g., the proposed accelerator with$v=8$has at least 66.67% less area-time complexities than the existing ones (Virtex-7 FPGA). We hope the outcome of this work can facilitate lightweight PQC development. Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
FCCM | 3 |
| 2023 | LOCS: LOw-Latency and ConStant-Timing Implementation of Fixed-Weight Sampler for HQCabstractPost-quantum cryptography (PQC) has drawn significant attention from various communities recently and one of the recent advances is the hardware acceleration of PQC algorithms. While Hamming Quasi-Cyclic (HQC) is one of the recently announced National Institute of Standards and Technology (NIST) fourth-round PQC standardization candidates, very few related hardware implementation works have been reported, particularly lacking solid works on important components such as the sampler. As a fixed-weight sparse vector sampler with constant-time operation is critical to the hardware HQC accelerator, in this paper, we present a novel hardware-implemented LOw-latency and ConStant-timing fixed-weight sampler (LOCS). In total, we have proposed three stages of efforts. First of all, a new algorithm for efficient realization of the fixed-weight sparse vector generation based on Fisher-Yates shuffle algorithm is proposed. Then, we have innovatively designed the algorithm into a new hardware sampler: LOCS. Finally, we have conducted a thorough comparison to showcase the efficiency of the proposed sampler, e.g., the proposed LOCS involves 66.7% less latency time than the state-of-the-art design$(n=17,669)$while remaining constant-time operation. To the authors' best knowledge, this is the first hardware-implemented pure constant-time (no failure probability) fixed-weight sampler for HQC. Pengzhou He, Yazheng Tu, Jiafeng Xie |
ISCAS | 2 |
| 2023 | COPMA: Compact and Optimized Polynomial Multiplier Accelerator for High-Performance Implementation of LWR-Based PQCabstractThe rapid progress in quantum computing has initiated a new round of cryptographic innovation, that is, developing postquantum cryptography (PQC) to resist attacks from well-established quantum computers. In this brief, we propose a novel compact and optimized polynomial multiplier accelerator (COPMA) for high-performance implementation of learning-with-rounding (LWR)-based PQC. As not many LWR-based PQC schemes are available in the literature, we have just used Saber, the National Institute of Standards and Technology (NIST) third-round PQC standardization finalist, as a typical case study example. First of all, we have formulated the polynomial multiplication, the major component of Saber, into a novel “subpolynomial”-based processing format for compact computation (yet has the potential for fast operation). Then, we have designed the proposed algorithm into an area-efficient polynomial multiplication hardware accelerator with high-frequency operational capability. Finally, we have verified the efficiency of the developed COPMA and have deployed it to build a cryptoprocessor. The implementation and analysis demonstrate the superior performance of the proposed COPMA. The proposed strategy is highly efficient and can be extended to build other PQC hardware accelerators. Pengzhou He, Yazheng Tu, Tianyou Bao, Leonel Sousa, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | KINA: Karatsuba Initiated Novel Accelerator for Ring-Binary-LWE (RBLWE)-Based Post-Quantum CryptographyabstractAlong with the National Institute of Standards and Technology (NIST) post-quantum cryptography (PQC) standardization process, lightweight PQC-related research, and development have also gained substantial attention from the research community. Ring-binary-learning-with-errors (RBLWE), a ring variant of binary-LWE (BLWE), has been used to build a promising lightweight PQC scheme for emerging Internet-of-Things (IoT) and edge computing applications, namely the RBLWE-based encryption scheme (RBLWE-ENC). The parameter settings of RBLWE-ENC, however, are not in favor of deploying typical fast algorithms like number theoretic transform (NTT). Following this direction, in this work, we propose a Karatsuba initiated novel accelerator (KINA) for efficient implementation of RBLWE-ENC. Overall, we have made several coherent interdependent stages of efforts to carry out the proposed work: 1) we have innovatively used the Karatsuba algorithm (KA) to derive the major arithmetic operation of RBLWE-ENC into a new form for high-performance operation; 2) we have then effectively mapped the proposed algorithm into an efficient hardware accelerator with the help of a number of optimization techniques; and 3) we have also provided detailed complexity analysis and implementation comparison to demonstrate the superior performance of the proposed KINA, e.g., the proposed design with$u=2$involves 64.71% higher throughput and 15.37% less area-delay product (ADP) than the state-of-the-art design for$n=512$(Virtex-7). The proposed KINA offers flexible processing speed and is suitable for high-performance applications like IoT servers. This work is expected to be useful for lightweight PQC development. Pengzhou He, Yazheng Tu, Jiafeng Xie, H. S. Jacinto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | LEAP: Lightweight and Efficient Accelerator for Sparse Polynomial Multiplication of HQCabstractThe Hamming quasi-cyclic (HQC) code-based encryption scheme is one of the fourth-round algorithms selected by the National Institute of Standards and Technology (NIST) postquantum cryptography (PQC) standardization process. However, very few hardware implementations have been reported for HQC to date. In this brief, we propose a novel Lightweight and Efficient Accelerator for sparse Polynomial multiplication (LEAP) of HQC, compatible with different parameters, on the field-programmable gate array (FPGA) platform. First, we give a mathematical derivation process for the sparse polynomial multiplication deployed in HQC. Then, we explain the proposed hardware structure in detail. Finally, we present the FPGA implementation results to confirm the efficiency of the proposed LEAP, for example, the proposed design for hqc-192 has at least 31.03% less area-delay product (ADP) than the existing design. LEAP can be extended further to construct efficient HQC cryptoprocessors. Yazheng Tu, Pengzhou He, Çetin Kaya Koç, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | HPMA-Saber: High-Performance Polynomial Multiplication Accelerator for KEM SaberabstractThe recent research in post-quantum cryptography (PQC) field has gradually switched to efficient implementation of PQC algorithms on hardware platforms. As polynomial multiplication is typically one of the critical operations within lattice-based PQC, its hardware acceleration has drawn significant attention from the research community recently. We propose a high-speed processing strategy to construct a new High-performance Polynomial Multiplication Accelerator (HPMA) for key encapsulation mechanism (KEM) Saber. Firstly, we have given a detailed mathematical derivation to obtain a low-latency processing algorithm for Saber polynomial multiplication. Then, we have innovatively used the derived the proposed algorithm to construct a new structure HPMA for FPGA implementation. Lastly, we have demonstrated the superior performance of the proposed HPMA-Saber by comparing with state-of-the-art works. The proposed design strategy is highly efficient and the obtained results can be useful for the PQC research community. Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
ICCD | 3 |
| 2022 | Hardware Implementation of High-Performance Polynomial Multiplication for KEM SaberabstractRecent advances in quantum computing have initiated a new round of cryptosystem innovation as the existing public-key cryptosystems are proven to be vulnerable to quantum attacks. Several types of cryptographic algorithms have been proposed for possible post-quantum cryptography (PQC) candidates and the lattice-based key encapsulation mechanism (KEM) Saber is one of the most promising algorithms. Noticing that the polynomial multiplication over ring is the key arithmetic operation of KEM Saber, in this paper, we propose a novel strategy for efficient implementation of polynomial multiplication on the hardware platform. First of all, we present the proposed mathematical derivation process for polynomial multiplication. Then, the proposed hardware structure is provided. Finally, field-programmable gate array (FPGA) based implementation results are obtained, and it is shown that the proposed design has better performance than the existing ones. The proposed polynomial multiplication can be further deployed to construct efficient hardware cryptoprocessors for KEM Saber. Yazheng Tu, Pengzhou He, Chiou-Yng Lee, Danai Chasaki, Jiafeng Xie |
ISCAS | 1 |
| 2022 | Efficient Hardware Arithmetic for Inverted Binary Ring-LWE Based Post-Quantum CryptographyabstractRing learning-with-errors(RLWE)-based encryption scheme is a lattice-based cryptographic algorithm that constitutes one of the most promising candidates for Post-Quantum Cryptography (PQC) standardization due to its efficient implementation and low computational complexity.Binary Ring-LWE (BRLWE) is a new optimized variant of RLWE, which achieves smaller computational complexity and higher efficient hardware implementations. In this paper, two efficient architectures based onLinear-Feedback Shift Register(LFSR) for the arithmetic used inInverted Binary Ring-LWE (InvBRLWE)-based encryption scheme are presented, namely the operation of$A\cdot B+C$over the polynomial ring$\mathbb {Z}_{q}/(x^{n}+1)$. The first architecture optimizes the resource usage for major computation and has a novel input processing setup to speed up the overall processing latency with minimized input loading cycles. The second architecture deploys an innovative serial-in serial-out processing format to reduce the involved area usage further yet maintains a regular input loading time-complexity. Experimental results show that the architectures presented here improve the complexities obtained by competing schemes found in the literature, e.g., involving 71.23% less area-delay product than recent designs. Both architectures are highly efficient in terms of area-time complexities and can be extended for deploying in different lightweight application environments. José Luis Imaña, Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |