EDBT 2026 Demo / reviewers in the wild / expert
Tianyou Bao
dblp:234/3397
· DBLP profile ↗
18ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0003-3321-5123ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 6 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trident: Efficient FPGA Acceleration of XMSS Tree in Post-Quantum Signature Scheme SLH-DSAabstractThe emergence of quantum computing poses significant threats to conventional cryptographic systems, necessitating the efficient hardware acceleration of Post-Quantum Cryptography (PQC), especially on the Field-Programmable Gate Array (FPGA) platforms. SPHINCS+, recently standardized by NIST (National Institute of Standards and Technology) as SLH-DSA (Stateless Hash-Based Digital Signature Algorithm), represents the only hash-based digital signature scheme. Its practical deployment, however, is restricted by computationally intense operations, particularly in the eXtended Merkle Signature Scheme (XMSS) tree, where WOTS+ (Winternitz One-Time Signature Plus) public key generation consumes the majority of signature generation cycles. With this background, this paper presents Trident, an innovative FPGA-based hardware accelerator that addresses critical performance and resource challenges in XMSS of SLH-DSA. First, we propose a triangle hash unit architecture that enables parallel execution of up to three hash operations simultaneously, directly addressing the computational bottleneck in XMSS tree construction and WOTS+ chain operations. Second, we develop an optimized memory caching scheme that reduces on-chip memory requirements via intermediate value management. Third, we implement the Trident on FPGAs and comprehensively evaluate it across all parameter sets at multiple security levels, i.e., up to 8.6× improvement in signature generation and up to 5.4× speed-up in verification operations. Extended Hypertree evaluation shows a 34.6× area-delay product (ADP) improvement on UltraScale+ FPGA for SLH-DSA-128s. This Trident represents a significant advancement toward practical SLH-DSA deployment in FPGA environments. Tianyou Bao, Joshua Ennis, Kirill Morozov, Jiafeng Xie |
FCCM | 1 |
| 2026 | CEDAR: A Compact and Efficient Decoder Architecture for RS-RM Code in HQC
Yazheng Tu, Tianyou Bao, Jiafeng Xie |
ISCAS | 2 |
| 2025 | Efficient Post-Quantum Cryptographic Hardware for Healthcare ApplicationsabstractIn light of rapid progress in quantum computing, Post-Quantum Cryptography (PQC) for healthcare/health-monitoring devices has gained substantial attention from the community recently as the existing cryptosystems are proven to be vulnerable to quantum attacks. Following the National Institute of Standards and Technology (NIST) PQC standardization process, this paper seeks to develop an efficient hardware implementation for Number Theoretic Transform (NTT) (a key component for NIST-selected PQC) on healthcare devices for emerging applications. Specifically, we target to design an efficient NTT on a resource-constrained FPGA platform to simulate the potential healthcare device application scenario. The proposed design architecture, implementation results, and comparison are provided to demonstrate the efficiency of the proposed design. Discussions and future research are also provided. We hope the outcome of this work can impact the field of PQC for healthcare/health-monitoring applications. Samuel Coulon, Tianyou Bao, Jiafeng Xie |
ISCAS | 2 |
| 2025 | HSPA: High-Throughput Sparse Polynomial Multiplication for Code-based Post-Quantum CryptographyabstractIncreasing attention has been paid to code-based post-quantum cryptography (PQC) schemes, e.g., HQC (Hamming Quasi-Cyclic) and BIKE (Bit Flipping Key Encapsulation), since they’ve been selected as the fourth-round National Institute of Standards and Technology (NIST) PQC standardization candidates. Though sparse polynomial multiplication is one of the critical components for HQC and BIKE, hardware-implemented high-performance sparse polynomial multiplier is rarely reported in the literature (due to its high-dimension and sparsity of polynomials involved in the computation). Based on this consideration, in this article, we propose two novel H igh-throughput S parse P olynomial multiplication A ccelerators (HSPA) for the mentioned two code-based PQC schemes. Specifically, we have designed the two accelerators based on two different implementation strategies targeting potential applications with different resource availability, i.e., one accelerator deploys a memory-based structure for computation while the other does not need memory usage. We have proposed three layers of coherent interdependent efforts to obtain the proposed accelerators. First, we have proposed two implementation strategies to execute the targeted sparse polynomial multiplication, i.e., a new parallel segment based accumulation (PSA) approach and a novel permutating-with-power (PWP)-based method. Then, the proposed two hardware accelerators are presented with detailed structural descriptions. Finally, field-programmable gate array (FPGA)-based implementation is presented to showcase the superior performance of the proposed accelerators. A proper comparison is also carried out to confirm the efficiency of the proposed designs. For instance, the proposed accelerator (using memory-based structure) has 56.84% and 80.25% less area-delay product (ADP) than the existing memory-based design (an extended high-speed version) on the UltraScale+ device, respectively, for n =17,669 and ω =75 (HQC) and n = 12,323 and ω =142 (BIKE). The proposed design strategy fits well with the two targeted code-based PQC schemes, which can be extended further to construct high-performance hardware cryptoprocessors. We hope the results of this work will be useful for the ongoing NIST PQC standardization process. Pengzhou He, Yazheng Tu, Tianyou Bao, Çetin Kaya Koç, Jiafeng Xie |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | CHIRP: Compact and High-Performance FPGA Implementation of Unified Hardware Accelerators for Ring-Binary-LWE-based PQCabstractPost-quantum cryptography (PQC) has drawn significant attention from the hardware design research community, especially on field-programmable gate array (FPGA) platforms. In line with this trend, in this article, we present a novel FPGA-based PQC design work (CHIRP), i.e., Compact and high-Performance FPGA implementation of unified accelerators for Ring-Binary-Learning-with-Errors (RBLWE)-based PQC, a promising lightweight PQC suited for related applications like Internet-of-Things. The proposed accelerators offer flexibility across the available two security levels, thus expanding their application potential. In total, we presented four distinct hardware accelerators tailored to different performance and resource requirements, ranging from resource-constrained devices to high-throughput applications. Our innovation encompasses three key efforts: (i) we derived four optimized algorithms for RBLWE-ENC’s unified operation (covering the available two security levels), allowing flexible switching of security sizes while boosting calculations; (ii) we then presented the four novel accelerators (CHIRP) targeting FPGA platforms, featuring dedicated hardware structures; (iii) we finally conducted a comprehensive evaluation to validate the efficiency of the proposed accelerators on various FPGA devices. Compared to the existing unified design, the proposed accelerator demonstrated up to 91.4% reduction in area-delay product (ADP) on the Straix-V device. Even when compared with the state-of-the-art single security designs, the proposed accelerator (best version) obtains much better resource usage and ADP performance while unified operation (flexibly switching between two security levels) is considered on both AMD-Xilinx and Intel devices. We anticipate the findings of this research will foster advancements in FPGA implementation techniques for lightweight PQC development. Tianyou Bao, Pengzhou He, Daisuke Fujimoto, Yuichi Hayashi, Jiafeng Xie |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2025 | High-Performance Instruction-Set Hardware Accelerator for Ring-Binary-LWE-Based Lightweight PQCabstractRecent advances in hardware acceleration for postquantum cryptography (PQC) have also switched to lightweight PQC. Apart from the traditional hardware design methodology, instruction-set accelerator for PQC represents a new design trend but has not been explored on lightweight PQC. To fill the research gap, in this work, we present a novel instruction-set acceleration of a ring-binary-learning-with-errors (RBLWEs)-based PQC (RBLWE-based encryption (ENC), a promising lightweight scheme) on field-programmable gate array (FPGA). Key efforts include: 1) derivation of an algorithm for the major operation of RBLWE-ENC to facilitate instruction-set acceleration; 2) development of the instruction-set accelerator, including the polynomial multiplication core designed from the derived algorithm; and 3) evaluation to showcase the efficiency of the proposed design. The proposed accelerator is efficient and complete in cryptographic operations, which can help further lightweight PQC development. Pengzhou He, Tianyou Bao, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Invited Paper: Enhancing Privacy-Preserving Computing with Optimized CKKS Encryption: A Hardware Acceleration ApproachabstractThe widespread adoption of the FHE (Fully Homomorphic Encryption) CKKS (Cheon-Kim-Kim-Song) encryption algorithm is stunted by its slow software implementation, driving the need for efficient hardware acceleration solutions. This paper presents a novel approach to address the challenge by optimizing a segment of the CKKS encryption and decryption process. Using the Residue Number System (RNS), the inputs are encoded into 32-bit representations and processed through the number theoretic transform (NTT). This innovative strategy reduces the utilities of the hardware resources and also accelerates calculations. This unified calculation methodology is designed to adapt inputs of diverse bit widths, seamlessly decoding them back into their original format after the CRT calculation, thereby significantly enhancing the throughput of calculations. We hope this work advances privacy-preserving computing in resource-constrained environments. Tianyou Bao, Pengzhou He, Jiafeng Xie |
ICCAD | 1 |
| 2024 | AEKA: FPGA Implementation of Area-Efficient Karatsuba Accelerator for Ring-Binary-LWE-Based Lightweight PQCabstractLightweight PQC-related research and development have gradually gained attention from the research community recently. Ring-Binary-Learning-with-Errors (RBLWE)-based encryption scheme (RBLWE-ENC), a promising lightweight PQC based on small parameter sets to fit related applications (but not in favor of deploying popular fast algorithms like number theoretic transform). To solve this problem, in this article, we present a novel implementation of hardware acceleration for RBLWE-ENC based on Karatsuba algorithm, particularly on the field-programmable gate array (FPGA) platform. In detail, we have proposed an area-efficient Karatsuba Accelerator (AEKA) for RBLWE-ENC, based on three layers of innovative efforts. First of all, we reformulate the signal processing sequence within the major arithmetic component of the KA-based polynomial multiplication for RBLWE-ENC to obtain a new algorithm. Then, we have designed the proposed algorithm into a new hardware accelerator with several novel algorithm-to-architecture mapping techniques. Finally, we have conducted thorough complexity analysis and comparison to demonstrate the efficiency of the proposed accelerator, e.g., it involves 62.5% higher throughput and 60.2% less area-delay product (ADP) than the state-of-the-art design for n =512 (Virtex-7 device, similar setup). The proposed AEKA design strategy is highly efficient on the FPGA devices, i.e., small resource usage with superior timing, which can be integrated with other necessary systems for lightweight-oriented high-performance applications (e.g., servers). The outcome of this work is also expected to generate impacts for lightweight PQC advancement. Tianyou Bao, Pengzhou He, Jiafeng Xie, H. S. Jacinto |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2024 | TINA: TMVP-Initiated Novel Accelerator for Lightweight Ring-LWE-Based PQCabstractPostquantum cryptography (PQC) has recently garnered significant attention across various communities. Alongside the ongoing standardization process for general-purpose PQC algorithms by the National Institute of Standards and Technology (NIST), the research community is actively exploring the realm of lightweight PQC schemes. A ring-binary-learning-with-error (RBLWE)-based encryption scheme (RBLWE-ENC) is a promising lightweight PQC candidate suitable for Internet-of-Things (IoT) and edge computing applications. The parameters of the RBLWE-ENC, however, do not favor deploying typical fast algorithms, such as number-theoretic transform (NTT). In this article, therefore, we propose to design aToeplitz matrix-vector product (TMVP)-initiatednovelaccelerator (TINA) for RBLWE-ENC. We innovatively used TMVP (a subquadratic-complexity fast algorithm for polynomial multiplication) to derive the significant arithmetic operation of RBLWE-ENC into a new form for high-performance operation. This novel formulation culminates in the development of a comprehensive accelerator known as TINA. Through implementation and comparative analysis, we demonstrate the efficiency gains achieved by our proposed accelerator. To the authors’ best knowledge, this is the first report on the TMVP strategy-initiated RBLWE-ENC accelerator. The findings of this work are expected to provide valuable references in the ongoing advancement of lightweight PQC development. Tianyou Bao, Pengzhou He, Shi Bai 0001, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | FELIX: FPGA-Based Scalable and Lightweight Accelerator for Large Integer Extended GCDabstractThe extended greatest common divisor (XGCD) computation is a critical component in various cryptographic applications and algorithms, including both pre-and postquantum cryptosystems. In addition to computing the greatest common divisor (GCD) of two integers, the XGCD also produces Bézout coefficients$b_a$and$b_b$which satisfy$\mathrm{GCD}(a,b) = a\times b_a + b\times b_b$. In particular, computing the XGCD for large integers is of significant interest. Most recently, XGCD computation between 6479-bit integers is required for solving$N$th-degree truncated polynomial ring unit (NTRU) trapdoors in Falcon, a National Institute of Standards and Technology (NIST)-selected postquantum digital signature scheme. To this point, existing literature has primarily focused on exploring software-based implementations for XGCD. The few existing high-performance hardware architectures require significant hardware resources and may not be desirable for practical usage, and the lightweight architectures suffer from poor performance. To fill the research gap, this work proposes a novel FPGA-based scalable and lightweight accelerator for large integer XGCD (FELIX). First, a new algorithm suitable for scalable and lightweight computation of XGCD is proposed. Next, a hardware accelerator (FELIX) is presented, including both constant-and variable-time versions. Finally, a thorough evaluation is carried out to showcase the efficiency of the proposed FELIX. In certain configurations, FELIX involves 81% less equivalent area-time product (eATP) than the state-of-the-art design for 1024-bit integers, and achieves a 95% reduction in latency over the software for 6479-bit integers (Falcon parameter set) with reasonable resource usage. Overall, the proposed FELIX is highly efficient, scalable, lightweight, and suitable for very large integer computation, making it the first such XGCD accelerator in the literature (to the best of our knowledge). Samuel Coulon, Tianyou Bao, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Efficient Implementation of Ring-Binary-LWE-based Lightweight PQC Accelerator on the FPGA PlatformabstractPost-quantum cryptography (PQC) has gained sub-stantial attention from various communities recently. Along with the ongoing National Institute of Standards and Technology (NIST) PQC standardization process that targets the general-purpose PQC algorithms, the research community is also looking for efficient lightweight PQC schemes. Among this direction of efforts, Ring-Binary-Learning-with-Errors (RBLWE)-based encryption scheme (RBLWE-ENC) is regarded as a promising lightweight PQC fitting Internet-of-Things (IoT) and edge computing applications. As hardware implementation for PQC algorithms has become one of the major advances in the field, in this paper, we follow this trend to present an efficient implementation of RBLWE-ENC lightweight accelerator on the field-programmable gate array (FPGA) platform. Overall, we have demonstrated three coherent interdependent stages of efforts: (i) we have presented detailed derivation processes to formulate the proposed algorithmic operation; (ii) we have then implemented the proposed algorithm into a desired hardware accelerator; and (iii) we provided thorough complexity analysis and comparison to showcase the superior performance of the proposed accelerator over the state-of-the-art designs, e.g., the proposed accelerator with$v=8$has at least 66.67% less area-time complexities than the existing ones (Virtex-7 FPGA). We hope the outcome of this work can facilitate lightweight PQC development. Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
FCCM | 2 |
| 2023 | FPGA Implementation of Compact Hardware Accelerators for Ring-Binary-LWE-based Post-quantum CryptographyabstractPost-quantum cryptography (PQC) has recently drawn substantial attention from various communities owing to the proven vulnerability of existing public-key cryptosystems against the attacks launched from well-established quantum computers. The Ring-Binary-Learning-with-Errors (RBLWE), a variant of Ring-LWE, has been proposed to build PQC for lightweight applications. As more Field-Programmable Gate Array (FPGA) devices are being deployed in lightweight applications like Internet-of-Things (IoT) devices, it would be interesting if the RBLWE-based PQC can be implemented on the FPGA with ultra-low complexity and flexible processing. However, thus far, limited information is available for such implementations. In this article, we propose novel RBLWE-based PQC accelerators on the FPGA with ultra-low implementation complexity and flexible timing. We first present the process of deriving the key operation of the RBLWE-based scheme into the proposed algorithmic operation. The corresponding hardware accelerator is then efficiently mapped from the proposed algorithm with the help of algorithm-to-architecture implementation techniques and extended to obtain higher-throughput designs. The final complexity analysis and implementation results (on a variety of FPGAs) show that the proposed accelerators have significantly smaller area-time complexities than the state-of-the-art designs. Overall, the proposed accelerators feature low implementation complexity and flexible processing, making them desirable for emerging FPGA-based lightweight applications. Pengzhou He, Tianyou Bao, Jiafeng Xie, Moeness G. Amin |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | COPMA: Compact and Optimized Polynomial Multiplier Accelerator for High-Performance Implementation of LWR-Based PQCabstractThe rapid progress in quantum computing has initiated a new round of cryptographic innovation, that is, developing postquantum cryptography (PQC) to resist attacks from well-established quantum computers. In this brief, we propose a novel compact and optimized polynomial multiplier accelerator (COPMA) for high-performance implementation of learning-with-rounding (LWR)-based PQC. As not many LWR-based PQC schemes are available in the literature, we have just used Saber, the National Institute of Standards and Technology (NIST) third-round PQC standardization finalist, as a typical case study example. First of all, we have formulated the polynomial multiplication, the major component of Saber, into a novel “subpolynomial”-based processing format for compact computation (yet has the potential for fast operation). Then, we have designed the proposed algorithm into an area-efficient polynomial multiplication hardware accelerator with high-frequency operational capability. Finally, we have verified the efficiency of the developed COPMA and have deployed it to build a cryptoprocessor. The implementation and analysis demonstrate the superior performance of the proposed COPMA. The proposed strategy is highly efficient and can be extended to build other PQC hardware accelerators. Pengzhou He, Yazheng Tu, Tianyou Bao, Leonel Sousa, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Work-in-Progress: High-Performance Systolic Hardware Accelerator for RBLWE-based Post-Quantum CryptographyabstractRing-Binary-Learning-with-Errors (RBLWE)-based post-quantum cryptography (PQC) is a promising scheme suitable for lightweight applications. This paper presents an efficient hardware systolic accelerator for RBLWE-based PQC, targeting high-performance applications. We have briefly given the algorithmic background for the proposed design. Then, we have transferred the proposed algorithmic operation into a new systolic accelerator. Lastly, field-programmable gate array (FPGA) implementation results have confirmed the efficiency of the proposed accelerator. Tianyou Bao, José Luis Imaña, Pengzhou He, Jiafeng Xie |
CODES+ISSS | 1 |
| 2022 | Ultra Low-Complexity Implementation of Binary Ring-LWE based Post-Quantum Cryptography on FPGA PlatformabstractPost-quantum cryptography (PQC) has drawn substantial attention from various communities recently since the existing public-key cryptosystems are proven to be vulnerable to the attacks launched from well-established quantum computers. The Ring-Learning-with-Errors (Ring-LWE)-based encryption scheme is an important lattice-based PQC. The binary Ring-LWE (BRLWE), a variant of Ring-LWE, has been proposed to build lightweight PQC with much smaller complexity for resource-constrained applications. So far, however, limited reports have been released on the ultra low-complexity implementations of the BRLWE-based scheme (especially on the hardware platform). Therefore, in this paper, we propose to obtain a novel implementation of the BRLWE-based PQC on the Field-Programmable Gate Array (FPGA) platform with ultra low-complexity. We have proposed three layers of coherent interdependent efforts: (i) motivation and the related derivation for the key operation of the BRLWE-based scheme are presented first; (ii) the corresponding hardware structure is then efficiently mapped from the proposed algorithm with the help of algorithm-architecture co-implementation techniques; (iii) the final complexity analysis and FPGA implementation results show that the proposed designs have significantly smaller area-time complexities than the state-of-the-art designs. Overall, the proposed designs possess multiple unique features and thus are desirable for emerging lightweight applications. Jiafeng Xie, Pengzhou He, Tianyou Bao |
FPGA | 3 |
| 2022 | HPMA-Saber: High-Performance Polynomial Multiplication Accelerator for KEM SaberabstractThe recent research in post-quantum cryptography (PQC) field has gradually switched to efficient implementation of PQC algorithms on hardware platforms. As polynomial multiplication is typically one of the critical operations within lattice-based PQC, its hardware acceleration has drawn significant attention from the research community recently. We propose a high-speed processing strategy to construct a new High-performance Polynomial Multiplication Accelerator (HPMA) for key encapsulation mechanism (KEM) Saber. Firstly, we have given a detailed mathematical derivation to obtain a low-latency processing algorithm for Saber polynomial multiplication. Then, we have innovatively used the derived the proposed algorithm to construct a new structure HPMA for FPGA implementation. Lastly, we have demonstrated the superior performance of the proposed HPMA-Saber by comparing with state-of-the-art works. The proposed design strategy is highly efficient and the obtained results can be useful for the PQC research community. Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
ICCD | 2 |
| 2022 | Efficient Hardware Arithmetic for Inverted Binary Ring-LWE Based Post-Quantum CryptographyabstractRing learning-with-errors(RLWE)-based encryption scheme is a lattice-based cryptographic algorithm that constitutes one of the most promising candidates for Post-Quantum Cryptography (PQC) standardization due to its efficient implementation and low computational complexity.Binary Ring-LWE (BRLWE) is a new optimized variant of RLWE, which achieves smaller computational complexity and higher efficient hardware implementations. In this paper, two efficient architectures based onLinear-Feedback Shift Register(LFSR) for the arithmetic used inInverted Binary Ring-LWE (InvBRLWE)-based encryption scheme are presented, namely the operation of$A\cdot B+C$over the polynomial ring$\mathbb {Z}_{q}/(x^{n}+1)$. The first architecture optimizes the resource usage for major computation and has a novel input processing setup to speed up the overall processing latency with minimized input loading cycles. The second architecture deploys an innovative serial-in serial-out processing format to reduce the involved area usage further yet maintains a regular input loading time-complexity. Experimental results show that the architectures presented here improve the complexities obtained by competing schemes found in the literature, e.g., involving 71.23% less area-delay product than recent designs. Both architectures are highly efficient in terms of area-time complexities and can be extended for deploying in different lightweight application environments. José Luis Imaña, Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2018 | A New Optimized Queueing Model with Compensation and Buffer
Tianyou Bao, Hanxing Hu, Shiang Li, Changjiang Zhang |
BIBM | 1 |