EDBT 2026 Demo / reviewers in the wild / expert
Jiafeng Xie
dblp:64/8798
· DBLP profile ↗
54ranked-venue papers
15as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 13 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trident: Efficient FPGA Acceleration of XMSS Tree in Post-Quantum Signature Scheme SLH-DSAabstractThe emergence of quantum computing poses significant threats to conventional cryptographic systems, necessitating the efficient hardware acceleration of Post-Quantum Cryptography (PQC), especially on the Field-Programmable Gate Array (FPGA) platforms. SPHINCS+, recently standardized by NIST (National Institute of Standards and Technology) as SLH-DSA (Stateless Hash-Based Digital Signature Algorithm), represents the only hash-based digital signature scheme. Its practical deployment, however, is restricted by computationally intense operations, particularly in the eXtended Merkle Signature Scheme (XMSS) tree, where WOTS+ (Winternitz One-Time Signature Plus) public key generation consumes the majority of signature generation cycles. With this background, this paper presents Trident, an innovative FPGA-based hardware accelerator that addresses critical performance and resource challenges in XMSS of SLH-DSA. First, we propose a triangle hash unit architecture that enables parallel execution of up to three hash operations simultaneously, directly addressing the computational bottleneck in XMSS tree construction and WOTS+ chain operations. Second, we develop an optimized memory caching scheme that reduces on-chip memory requirements via intermediate value management. Third, we implement the Trident on FPGAs and comprehensively evaluate it across all parameter sets at multiple security levels, i.e., up to 8.6× improvement in signature generation and up to 5.4× speed-up in verification operations. Extended Hypertree evaluation shows a 34.6× area-delay product (ADP) improvement on UltraScale+ FPGA for SLH-DSA-128s. This Trident represents a significant advancement toward practical SLH-DSA deployment in FPGA environments. Tianyou Bao, Joshua Ennis, Kirill Morozov, Jiafeng Xie |
FCCM | 4 |
| 2026 | CEDAR: A Compact and Efficient Decoder Architecture for RS-RM Code in HQC
Yazheng Tu, Tianyou Bao, Jiafeng Xie |
ISCAS | 3 |
| 2026 | Hardware Acceleration for Zero-Knowledge Proof: Recent Advances and Challenges
Pengzhou He, Debapriya Basu Roy, Jiafeng Xie |
VTS | 3 |
| 2026 | SoK: Can Fully Homomorphic Encryption Support General AI Computation? A Functional and Cost AnalysisabstractArtificial intelligence (AI) increasingly powers sensitive applications in domains such as healthcare and finance, relying on both extit{linear operations} (e.g., matrix multiplications in large language models) and extit{non-linear operations} (e.g., sorting in retrieval-augmented generation). Fully homomorphic encryption (FHE) has emerged as a promising tool for privacy-preserving computation, but it remains unclear whether existing methods can support the full spectrum of AI workloads that combine these operations. In this SoK, we ask: extit{Can FHE support general AI computation?} We provide both a functional analysis and a cost analysis. First, we categorize ten distinct FHE approaches and evaluate their ability to support general computation. We then identify three promising candidates and benchmark workloads that mix linear and non-linear operations across different bit lengths and SIMD parallelization settings. Finally, we evaluate five real-world, privacy-sensitive AI applications that instantiate these workloads. Our results quantify the costs of achieving general computation in FHE and offer practical guidance on selecting FHE methods that best fit specific AI application requirements. Our codes are available at https://github.com/UCF-ML-Research/FHE-AI-Generality. Wei Zhang 0076, Mengxin Zheng, Minxuan Zhou, Yushun Dong, Dongjie Wang 0001, Jiafeng Xie, David Mohaisen, Hongyi Wu, Qian Lou |
Proc. Priv. Enhancing Technol. | 10 |
| 2025 | LEAF: Lightweight and Efficient Hardware Accelerator for Signature Verification of FALCONabstractAlong with the National Institute of Standards and Technology (NIST) post-quantum cryptography (PQC) standardization process, efficient hardware acceleration for PQC has become a priority. Among the NIST-selected PQC digital signature schemes, FALCON shows great promise due to its compact key sizes and efficient Signature Verification procedure. However, FALCON is regarded as highly computationally complex, and as a result, few works for hardware acceleration of FALCON can be found in the literature, where the few existing ones only target high-performance. To fill the gap, this paper presents a Lightweight and Efficient hardware accelerator for the Signature Verification portion of FALCON (LEAF), specifically for resource-constrained applications. We propose an efficient design strategy, including a novel data dependence flow, to maximize the utilization of very small resources for all arithmetic procedures. Then, the proposed full-hardware LEAF is built, containing an ultra-lightweight number theoretical transform (NTT) core with a novel twiddle factor access pattern. Finally, we conduct a thorough evaluation to demonstrate the efficiency of LEAF. To the best of our knowledge, this is the first lightweight and meanwhile most resource-efficient FALCON Signature Verification full-hardware accelerator in the literature, offering 65% and 66% less aggregate resource usage and achieving 24% and 14% less equivalent area-time product (eATP), compared to the state-of-the-art for FALCON-512 and FALCON-1024, respectively. We hope that this work can spur further research in the field. Samuel Coulon, Jinjun Xiong, Jiafeng Xie |
ICCAD | 3 |
| 2025 | Efficient Post-Quantum Cryptographic Hardware for Healthcare ApplicationsabstractIn light of rapid progress in quantum computing, Post-Quantum Cryptography (PQC) for healthcare/health-monitoring devices has gained substantial attention from the community recently as the existing cryptosystems are proven to be vulnerable to quantum attacks. Following the National Institute of Standards and Technology (NIST) PQC standardization process, this paper seeks to develop an efficient hardware implementation for Number Theoretic Transform (NTT) (a key component for NIST-selected PQC) on healthcare devices for emerging applications. Specifically, we target to design an efficient NTT on a resource-constrained FPGA platform to simulate the potential healthcare device application scenario. The proposed design architecture, implementation results, and comparison are provided to demonstrate the efficiency of the proposed design. Discussions and future research are also provided. We hope the outcome of this work can impact the field of PQC for healthcare/health-monitoring applications. Samuel Coulon, Tianyou Bao, Jiafeng Xie |
ISCAS | 3 |
| 2025 | ISCAS Guest Editorial Special Issue Based on the 2025 IEEE International Symposium on Circuits and Systems
Jiafeng Xie, Yuan Du, Xinmiao Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | HSPA: High-Throughput Sparse Polynomial Multiplication for Code-based Post-Quantum CryptographyabstractIncreasing attention has been paid to code-based post-quantum cryptography (PQC) schemes, e.g., HQC (Hamming Quasi-Cyclic) and BIKE (Bit Flipping Key Encapsulation), since they’ve been selected as the fourth-round National Institute of Standards and Technology (NIST) PQC standardization candidates. Though sparse polynomial multiplication is one of the critical components for HQC and BIKE, hardware-implemented high-performance sparse polynomial multiplier is rarely reported in the literature (due to its high-dimension and sparsity of polynomials involved in the computation). Based on this consideration, in this article, we propose two novel H igh-throughput S parse P olynomial multiplication A ccelerators (HSPA) for the mentioned two code-based PQC schemes. Specifically, we have designed the two accelerators based on two different implementation strategies targeting potential applications with different resource availability, i.e., one accelerator deploys a memory-based structure for computation while the other does not need memory usage. We have proposed three layers of coherent interdependent efforts to obtain the proposed accelerators. First, we have proposed two implementation strategies to execute the targeted sparse polynomial multiplication, i.e., a new parallel segment based accumulation (PSA) approach and a novel permutating-with-power (PWP)-based method. Then, the proposed two hardware accelerators are presented with detailed structural descriptions. Finally, field-programmable gate array (FPGA)-based implementation is presented to showcase the superior performance of the proposed accelerators. A proper comparison is also carried out to confirm the efficiency of the proposed designs. For instance, the proposed accelerator (using memory-based structure) has 56.84% and 80.25% less area-delay product (ADP) than the existing memory-based design (an extended high-speed version) on the UltraScale+ device, respectively, for n =17,669 and ω =75 (HQC) and n = 12,323 and ω =142 (BIKE). The proposed design strategy fits well with the two targeted code-based PQC schemes, which can be extended further to construct high-performance hardware cryptoprocessors. We hope the results of this work will be useful for the ongoing NIST PQC standardization process. Pengzhou He, Yazheng Tu, Tianyou Bao, Çetin Kaya Koç, Jiafeng Xie |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | CHIRP: Compact and High-Performance FPGA Implementation of Unified Hardware Accelerators for Ring-Binary-LWE-based PQCabstractPost-quantum cryptography (PQC) has drawn significant attention from the hardware design research community, especially on field-programmable gate array (FPGA) platforms. In line with this trend, in this article, we present a novel FPGA-based PQC design work (CHIRP), i.e., Compact and high-Performance FPGA implementation of unified accelerators for Ring-Binary-Learning-with-Errors (RBLWE)-based PQC, a promising lightweight PQC suited for related applications like Internet-of-Things. The proposed accelerators offer flexibility across the available two security levels, thus expanding their application potential. In total, we presented four distinct hardware accelerators tailored to different performance and resource requirements, ranging from resource-constrained devices to high-throughput applications. Our innovation encompasses three key efforts: (i) we derived four optimized algorithms for RBLWE-ENC’s unified operation (covering the available two security levels), allowing flexible switching of security sizes while boosting calculations; (ii) we then presented the four novel accelerators (CHIRP) targeting FPGA platforms, featuring dedicated hardware structures; (iii) we finally conducted a comprehensive evaluation to validate the efficiency of the proposed accelerators on various FPGA devices. Compared to the existing unified design, the proposed accelerator demonstrated up to 91.4% reduction in area-delay product (ADP) on the Straix-V device. Even when compared with the state-of-the-art single security designs, the proposed accelerator (best version) obtains much better resource usage and ADP performance while unified operation (flexibly switching between two security levels) is considered on both AMD-Xilinx and Intel devices. We anticipate the findings of this research will foster advancements in FPGA implementation techniques for lightweight PQC development. Tianyou Bao, Pengzhou He, Daisuke Fujimoto, Yuichi Hayashi, Jiafeng Xie |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2025 | EMINEM: Efficient FPGA Implementation of Mixed-RadIx NTT Hardware AccElerators for NIST Post-QuantuM Cryptography Falcon, Dilithium, and HAWKabstractThe advent of quantum computing poses a significant threat to modern cryptography. To address this challenge, the National Institute of Standards and Technology (NIST) has initiated the Post-Quantum Cryptography (PQC) standardization process, and several algorithms have been selected (a few are still under consideration in the additional standardization process). Among these schemes, lattice-based PQC has emerged as a promising approach and garnered substantial attention from the implementation community, especially on the hardware platforms. Notably, the Field-Programmable Gate Array (FPGA) has gained considerable attention as a convenient platform for hardware implementation, not only from NIST but also from the research community, as reflected by the related recommendations from NIST and the number of works reported recently. This work follows the existing trend of developing novel FPGA implementations for PQC. It is worth mentioning that the polynomial multiplication in these NIST lattice-based PQC algorithms can be implemented with Number Theoretic Transform (NTT) for efficiency. Nevertheless, there remains a lack of novel and universal NTT methods for polynomial multiplication at different sizes. For instance, for \(n=512\) (F alcon and HAWK), the existing works are mostly limited to the Radix-2 NTT (other methods like Radix-4 or Radix-8 cannot be directly applied). To fill the research gap, this article presents a novel design framework, i.e., Efficient Mixed-RadIx NTT hardware accElerators for NIST post-quantuM cryptography (EMINEM) , specially tailored for targeted schemes. Our design leverages Radix-4 for polynomial sizes of 256 and 1,024, while introducing a hybrid Radix-2/4 strategy for NTT of length 512 and achieving comparable performance to pure Radix-4 at other lengths. In total, our contributions include: (i) a generic Radix-4/Mixed-Radix NTT algorithm is proposed for \(n=256\) , 512, and 1,024; (ii) an efficient NTT hardware accelerator is designed with the help of a new memory access pattern and some optimization techniques; (iii) two types of butterfly architectures are developed to obtain pure Radix-4 time complexity and low resource usage, respectively; (iv) a detailed implementation and comparison showcase the superior performance of the proposed design strategy. Overall, the proposed strategy enables the efficient deployment of Mixed-Radix NTT in targeted NIST schemes, surpassing the limitations of the conventional Radix-2 approach for 512-length NTT designs. The proposed design offers a significant advancement in the field, facilitating efficient FPGA acceleration of PQC standards. Yazheng Tu, Jiafeng Xie |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2025 | High-Performance Instruction-Set Hardware Accelerator for Ring-Binary-LWE-Based Lightweight PQCabstractRecent advances in hardware acceleration for postquantum cryptography (PQC) have also switched to lightweight PQC. Apart from the traditional hardware design methodology, instruction-set accelerator for PQC represents a new design trend but has not been explored on lightweight PQC. To fill the research gap, in this work, we present a novel instruction-set acceleration of a ring-binary-learning-with-errors (RBLWEs)-based PQC (RBLWE-based encryption (ENC), a promising lightweight scheme) on field-programmable gate array (FPGA). Key efforts include: 1) derivation of an algorithm for the major operation of RBLWE-ENC to facilitate instruction-set acceleration; 2) development of the instruction-set accelerator, including the polynomial multiplication core designed from the derived algorithm; and 3) evaluation to showcase the efficiency of the proposed design. The proposed accelerator is efficient and complete in cryptographic operations, which can help further lightweight PQC development. Pengzhou He, Tianyou Bao, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | SCOPE: Schoolbook-Originated Novel Polynomial Multiplication Accelerators for NTRU-Based PQCabstractTheNth-degree truncated polynomial ring units (NTRUs)-based postquantum cryptography (PQC) has drawn significant attention from the research communities, e.g., the National Institute of Standards and Technology (NIST) PQC standardization process selected algorithm Fast Fourier lattice-based compact (Falcon). Following the research trend, efficient hardware accelerator design for polynomial multiplication (an important component of the NTRU-based PQC) is crucial. Unlike the commonly used number theoretic transform (NTT) method, in this article, we have presented a novel SChoolbook-Originated Polynomial multiplication accElerators (SCOPE) design framework. Overall, we have proposed the schoolbook-based method in an innovative format to implement the targeted polynomial multiplication, first through a schoolbook-variant version and then through a Toeplitz matrix-vector product (TMVP)-based approach. Four layers of coherent and interdependent efforts have been carried out: 1) a novel lookup table (LUT)-based point-wise multiplier is proposed along with a related modular reduction technique to obtain optimal implementation; 2) a new hardware accelerator is introduced for the targeted polynomial multiplication, deploying the proposed point-wise multiplier; 3) the proposed architecture is extended to a TMVP-based polynomial multiplication accelerator; and 4) the efficiency of the proposed accelerators is demonstrated through implementation and comparison. Finally, the proposed design strategy is also extended to another NTRU-based scheme and other schoolbook- and toom-cook-based polynomial multiplications (used in other PQC), and obtains the same superior performance. We hope that the outcome of this research can impact the ongoing NIST PQC standardization process and related full-hardware implementation work for schemes like Falcon. Yazheng Tu, Shi Bai 0001, Jinjun Xiong, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | LAMP: Efficient Implementation of Lightweight Accelerator for Polynomial MultiPlication, From Falcon to RBLWE-ENCabstractPost-quantum cryptography (PQC) has drawn significant attention from the hardware design research community. In particular, efficient implementation for major components of PQC algorithms like polynomial multiplication has been a hot topic recently. Following this trend, in this paper, we propose a novel hardware-implemented Lightweight Accelerator for the large integer polynomial MultiPlication (LAMP) used in PQC schemes. Specifically, we target the polynomial multiplications not bound by fixed fast algorithms like Number Theoretic Transform (NTT), i.e., Falcon (one of the National Institute of Standards and Technology (NIST) selected PQC algorithms) and Ring Binary Learning-with-Errors based encryption scheme (RBLWE-ENC, a promising lightweight PQC scheme). Overall, we have carried out three layers of innovative efforts. (i) A new lightweight computation strategy for the targeted polynomial multiplication is proposed; (ii) The new accelerator is then designed based on the proposed algorithm (applicable for both targeted schemes); (iii) A thorough evaluation process is carried out to showcase the superior performance of the proposed accelerator over the competing designs, e.g., at least 21.2% less area-delay product (ADP) when LAMP is used for RBLWE-ENC (on Virtex-7 device). The proposed work is efficient and interesting, and we hope this outcome can facilitate PQC development. Pengzhou He, Ben Mongirdas, Çetin Kaya Koç, Jiafeng Xie |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | Invited Paper: Enhancing Privacy-Preserving Computing with Optimized CKKS Encryption: A Hardware Acceleration ApproachabstractThe widespread adoption of the FHE (Fully Homomorphic Encryption) CKKS (Cheon-Kim-Kim-Song) encryption algorithm is stunted by its slow software implementation, driving the need for efficient hardware acceleration solutions. This paper presents a novel approach to address the challenge by optimizing a segment of the CKKS encryption and decryption process. Using the Residue Number System (RNS), the inputs are encoded into 32-bit representations and processed through the number theoretic transform (NTT). This innovative strategy reduces the utilities of the hardware resources and also accelerates calculations. This unified calculation methodology is designed to adapt inputs of diverse bit widths, seamlessly decoding them back into their original format after the CRT calculation, thereby significantly enhancing the throughput of calculations. We hope this work advances privacy-preserving computing in resource-constrained environments. Tianyou Bao, Pengzhou He, Jiafeng Xie |
ICCAD | 3 |
| 2024 | SMALL: Scalable Matrix OriginAted Large Integer PoLynomial Multiplication Accelerator for Lattice-Based Post-Quantum Cryptography
Jiafeng Xie, Pengzhou He, Samira Carolina Oliva Madrigal, Çetin Kaya Koç |
WAIFI | 1 |
| 2024 | AEKA: FPGA Implementation of Area-Efficient Karatsuba Accelerator for Ring-Binary-LWE-Based Lightweight PQCabstractLightweight PQC-related research and development have gradually gained attention from the research community recently. Ring-Binary-Learning-with-Errors (RBLWE)-based encryption scheme (RBLWE-ENC), a promising lightweight PQC based on small parameter sets to fit related applications (but not in favor of deploying popular fast algorithms like number theoretic transform). To solve this problem, in this article, we present a novel implementation of hardware acceleration for RBLWE-ENC based on Karatsuba algorithm, particularly on the field-programmable gate array (FPGA) platform. In detail, we have proposed an area-efficient Karatsuba Accelerator (AEKA) for RBLWE-ENC, based on three layers of innovative efforts. First of all, we reformulate the signal processing sequence within the major arithmetic component of the KA-based polynomial multiplication for RBLWE-ENC to obtain a new algorithm. Then, we have designed the proposed algorithm into a new hardware accelerator with several novel algorithm-to-architecture mapping techniques. Finally, we have conducted thorough complexity analysis and comparison to demonstrate the efficiency of the proposed accelerator, e.g., it involves 62.5% higher throughput and 60.2% less area-delay product (ADP) than the state-of-the-art design for n =512 (Virtex-7 device, similar setup). The proposed AEKA design strategy is highly efficient on the FPGA devices, i.e., small resource usage with superior timing, which can be integrated with other necessary systems for lightweight-oriented high-performance applications (e.g., servers). The outcome of this work is also expected to generate impacts for lightweight PQC advancement. Tianyou Bao, Pengzhou He, Jiafeng Xie, H. S. Jacinto |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | TINA: TMVP-Initiated Novel Accelerator for Lightweight Ring-LWE-Based PQCabstractPostquantum cryptography (PQC) has recently garnered significant attention across various communities. Alongside the ongoing standardization process for general-purpose PQC algorithms by the National Institute of Standards and Technology (NIST), the research community is actively exploring the realm of lightweight PQC schemes. A ring-binary-learning-with-error (RBLWE)-based encryption scheme (RBLWE-ENC) is a promising lightweight PQC candidate suitable for Internet-of-Things (IoT) and edge computing applications. The parameters of the RBLWE-ENC, however, do not favor deploying typical fast algorithms, such as number-theoretic transform (NTT). In this article, therefore, we propose to design aToeplitz matrix-vector product (TMVP)-initiatednovelaccelerator (TINA) for RBLWE-ENC. We innovatively used TMVP (a subquadratic-complexity fast algorithm for polynomial multiplication) to derive the significant arithmetic operation of RBLWE-ENC into a new form for high-performance operation. This novel formulation culminates in the development of a comprehensive accelerator known as TINA. Through implementation and comparative analysis, we demonstrate the efficiency gains achieved by our proposed accelerator. To the authors’ best knowledge, this is the first report on the TMVP strategy-initiated RBLWE-ENC accelerator. The findings of this work are expected to provide valuable references in the ongoing advancement of lightweight PQC development. Tianyou Bao, Pengzhou He, Shi Bai 0001, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | FELIX: FPGA-Based Scalable and Lightweight Accelerator for Large Integer Extended GCDabstractThe extended greatest common divisor (XGCD) computation is a critical component in various cryptographic applications and algorithms, including both pre-and postquantum cryptosystems. In addition to computing the greatest common divisor (GCD) of two integers, the XGCD also produces Bézout coefficients$b_a$and$b_b$which satisfy$\mathrm{GCD}(a,b) = a\times b_a + b\times b_b$. In particular, computing the XGCD for large integers is of significant interest. Most recently, XGCD computation between 6479-bit integers is required for solving$N$th-degree truncated polynomial ring unit (NTRU) trapdoors in Falcon, a National Institute of Standards and Technology (NIST)-selected postquantum digital signature scheme. To this point, existing literature has primarily focused on exploring software-based implementations for XGCD. The few existing high-performance hardware architectures require significant hardware resources and may not be desirable for practical usage, and the lightweight architectures suffer from poor performance. To fill the research gap, this work proposes a novel FPGA-based scalable and lightweight accelerator for large integer XGCD (FELIX). First, a new algorithm suitable for scalable and lightweight computation of XGCD is proposed. Next, a hardware accelerator (FELIX) is presented, including both constant-and variable-time versions. Finally, a thorough evaluation is carried out to showcase the efficiency of the proposed FELIX. In certain configurations, FELIX involves 81% less equivalent area-time product (eATP) than the state-of-the-art design for 1024-bit integers, and achieves a 95% reduction in latency over the software for 6479-bit integers (Falcon parameter set) with reasonable resource usage. Overall, the proposed FELIX is highly efficient, scalable, lightweight, and suitable for very large integer computation, making it the first such XGCD accelerator in the literature (to the best of our knowledge). Samuel Coulon, Tianyou Bao, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Efficient Implementation of Ring-Binary-LWE-based Lightweight PQC Accelerator on the FPGA PlatformabstractPost-quantum cryptography (PQC) has gained sub-stantial attention from various communities recently. Along with the ongoing National Institute of Standards and Technology (NIST) PQC standardization process that targets the general-purpose PQC algorithms, the research community is also looking for efficient lightweight PQC schemes. Among this direction of efforts, Ring-Binary-Learning-with-Errors (RBLWE)-based encryption scheme (RBLWE-ENC) is regarded as a promising lightweight PQC fitting Internet-of-Things (IoT) and edge computing applications. As hardware implementation for PQC algorithms has become one of the major advances in the field, in this paper, we follow this trend to present an efficient implementation of RBLWE-ENC lightweight accelerator on the field-programmable gate array (FPGA) platform. Overall, we have demonstrated three coherent interdependent stages of efforts: (i) we have presented detailed derivation processes to formulate the proposed algorithmic operation; (ii) we have then implemented the proposed algorithm into a desired hardware accelerator; and (iii) we provided thorough complexity analysis and comparison to showcase the superior performance of the proposed accelerator over the state-of-the-art designs, e.g., the proposed accelerator with$v=8$has at least 66.67% less area-time complexities than the existing ones (Virtex-7 FPGA). We hope the outcome of this work can facilitate lightweight PQC development. Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
FCCM | 4 |
| 2023 | PasCore: A Chinese Overlapping Relation Extraction Model Based on Global Pointer Annotation StrategyabstractRecent work for extracting relations from texts has achieved excellent performance. However, existing studies mainly focus on simple relation extraction, these methods perform not well on overlapping triple problem because the tags of shared entities would conflict with each other. Especially, overlapping entities are common and indispensable in Chinese. To address this issue, this paper proposes PasCore, which utilizes a global pointer annotation strategy for overlapping relation extraction in Chinese. PasCore first obtains the sentence vector via general pre-training model encoder, and uses classifier to predicate relations. Subsequently, it uses global pointer annotation strategy for head entity annotation, which uses global tags to label the start and end positions of the entities. Finally, PasCore integrates the relation, head entity and its type to mark the tail entity. Furthermore, PasCore performs conditional layer normalization to fuse features, which connects all stages and greatly enriches the association between relations and entities. Experimental results on both Chinese and English real-world datasets demonstrate that PasCore outperforms strong baselines on relation extraction and, especially, shows superior performance on overlapping relation extraction. Peng Wang 0004, Jiafeng Xie, Xiye Chen |
IJCAI | 2 |
| 2023 | LOCS: LOw-Latency and ConStant-Timing Implementation of Fixed-Weight Sampler for HQCabstractPost-quantum cryptography (PQC) has drawn significant attention from various communities recently and one of the recent advances is the hardware acceleration of PQC algorithms. While Hamming Quasi-Cyclic (HQC) is one of the recently announced National Institute of Standards and Technology (NIST) fourth-round PQC standardization candidates, very few related hardware implementation works have been reported, particularly lacking solid works on important components such as the sampler. As a fixed-weight sparse vector sampler with constant-time operation is critical to the hardware HQC accelerator, in this paper, we present a novel hardware-implemented LOw-latency and ConStant-timing fixed-weight sampler (LOCS). In total, we have proposed three stages of efforts. First of all, a new algorithm for efficient realization of the fixed-weight sparse vector generation based on Fisher-Yates shuffle algorithm is proposed. Then, we have innovatively designed the algorithm into a new hardware sampler: LOCS. Finally, we have conducted a thorough comparison to showcase the efficiency of the proposed sampler, e.g., the proposed LOCS involves 66.7% less latency time than the state-of-the-art design$(n=17,669)$while remaining constant-time operation. To the authors' best knowledge, this is the first hardware-implemented pure constant-time (no failure probability) fixed-weight sampler for HQC. Pengzhou He, Yazheng Tu, Jiafeng Xie |
ISCAS | 3 |
| 2023 | Monocular Road Planar Parallax EstimationabstractEstimating the 3D structure of the drivable surface and surrounding environment is a crucial task for assisted and autonomous driving. It is commonly solved either by using 3D sensors such as LiDAR or directly predicting the depth of points via deep learning. However, the former is expensive, and the latter lacks the use of geometry information for the scene. In this paper, instead of following existing methodologies, we propose Road Planar Parallax Attention Network (RPANet), a new deep neural network for 3D sensing from monocular image sequences based on planar parallax, which takes full advantage of the omnipresent road plane geometry in driving scenes. RPANet takes a pair of images aligned by the homography of the road plane as input and outputs a γ map (the ratio of height to depth) for 3D reconstruction. The γ map has the potential to construct a two-dimensional transformation between two consecutive frames. It implies planar parallax and can be combined with the road plane serving as a reference to estimate the 3D structure by warping the consecutive frames. Furthermore, we introduce a novel cross-attention module to make the network better perceive the displacements caused by planar parallax. To verify the effectiveness of our method, we sample data from the Waymo Open Dataset and construct annotations related to planar parallax. Comprehensive experiments are conducted on the sampled dataset to demonstrate the 3D reconstruction accuracy of our approach in challenging scenarios. Haobo Yuan, Wei Sui, Jiafeng Xie, Lefei Zhang, Qian Zhang 0009 |
IEEE Trans. Image Process. | 4 |
| 2023 | FPGA Implementation of Compact Hardware Accelerators for Ring-Binary-LWE-based Post-quantum CryptographyabstractPost-quantum cryptography (PQC) has recently drawn substantial attention from various communities owing to the proven vulnerability of existing public-key cryptosystems against the attacks launched from well-established quantum computers. The Ring-Binary-Learning-with-Errors (RBLWE), a variant of Ring-LWE, has been proposed to build PQC for lightweight applications. As more Field-Programmable Gate Array (FPGA) devices are being deployed in lightweight applications like Internet-of-Things (IoT) devices, it would be interesting if the RBLWE-based PQC can be implemented on the FPGA with ultra-low complexity and flexible processing. However, thus far, limited information is available for such implementations. In this article, we propose novel RBLWE-based PQC accelerators on the FPGA with ultra-low implementation complexity and flexible timing. We first present the process of deriving the key operation of the RBLWE-based scheme into the proposed algorithmic operation. The corresponding hardware accelerator is then efficiently mapped from the proposed algorithm with the help of algorithm-to-architecture implementation techniques and extended to obtain higher-throughput designs. The final complexity analysis and implementation results (on a variety of FPGAs) show that the proposed accelerators have significantly smaller area-time complexities than the state-of-the-art designs. Overall, the proposed accelerators feature low implementation complexity and flexible processing, making them desirable for emerging FPGA-based lightweight applications. Pengzhou He, Tianyou Bao, Jiafeng Xie, Moeness G. Amin |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | COPMA: Compact and Optimized Polynomial Multiplier Accelerator for High-Performance Implementation of LWR-Based PQCabstractThe rapid progress in quantum computing has initiated a new round of cryptographic innovation, that is, developing postquantum cryptography (PQC) to resist attacks from well-established quantum computers. In this brief, we propose a novel compact and optimized polynomial multiplier accelerator (COPMA) for high-performance implementation of learning-with-rounding (LWR)-based PQC. As not many LWR-based PQC schemes are available in the literature, we have just used Saber, the National Institute of Standards and Technology (NIST) third-round PQC standardization finalist, as a typical case study example. First of all, we have formulated the polynomial multiplication, the major component of Saber, into a novel “subpolynomial”-based processing format for compact computation (yet has the potential for fast operation). Then, we have designed the proposed algorithm into an area-efficient polynomial multiplication hardware accelerator with high-frequency operational capability. Finally, we have verified the efficiency of the developed COPMA and have deployed it to build a cryptoprocessor. The implementation and analysis demonstrate the superior performance of the proposed COPMA. The proposed strategy is highly efficient and can be extended to build other PQC hardware accelerators. Pengzhou He, Yazheng Tu, Tianyou Bao, Leonel Sousa, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | KINA: Karatsuba Initiated Novel Accelerator for Ring-Binary-LWE (RBLWE)-Based Post-Quantum CryptographyabstractAlong with the National Institute of Standards and Technology (NIST) post-quantum cryptography (PQC) standardization process, lightweight PQC-related research, and development have also gained substantial attention from the research community. Ring-binary-learning-with-errors (RBLWE), a ring variant of binary-LWE (BLWE), has been used to build a promising lightweight PQC scheme for emerging Internet-of-Things (IoT) and edge computing applications, namely the RBLWE-based encryption scheme (RBLWE-ENC). The parameter settings of RBLWE-ENC, however, are not in favor of deploying typical fast algorithms like number theoretic transform (NTT). Following this direction, in this work, we propose a Karatsuba initiated novel accelerator (KINA) for efficient implementation of RBLWE-ENC. Overall, we have made several coherent interdependent stages of efforts to carry out the proposed work: 1) we have innovatively used the Karatsuba algorithm (KA) to derive the major arithmetic operation of RBLWE-ENC into a new form for high-performance operation; 2) we have then effectively mapped the proposed algorithm into an efficient hardware accelerator with the help of a number of optimization techniques; and 3) we have also provided detailed complexity analysis and implementation comparison to demonstrate the superior performance of the proposed KINA, e.g., the proposed design with$u=2$involves 64.71% higher throughput and 15.37% less area-delay product (ADP) than the state-of-the-art design for$n=512$(Virtex-7). The proposed KINA offers flexible processing speed and is suitable for high-performance applications like IoT servers. This work is expected to be useful for lightweight PQC development. Pengzhou He, Yazheng Tu, Jiafeng Xie, H. S. Jacinto |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | LEAP: Lightweight and Efficient Accelerator for Sparse Polynomial Multiplication of HQCabstractThe Hamming quasi-cyclic (HQC) code-based encryption scheme is one of the fourth-round algorithms selected by the National Institute of Standards and Technology (NIST) postquantum cryptography (PQC) standardization process. However, very few hardware implementations have been reported for HQC to date. In this brief, we propose a novel Lightweight and Efficient Accelerator for sparse Polynomial multiplication (LEAP) of HQC, compatible with different parameters, on the field-programmable gate array (FPGA) platform. First, we give a mathematical derivation process for the sparse polynomial multiplication deployed in HQC. Then, we explain the proposed hardware structure in detail. Finally, we present the FPGA implementation results to confirm the efficiency of the proposed LEAP, for example, the proposed design for hqc-192 has at least 31.03% less area-delay product (ADP) than the existing design. LEAP can be extended further to construct efficient HQC cryptoprocessors. Yazheng Tu, Pengzhou He, Çetin Kaya Koç, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | Work-in-Progress: High-Performance Systolic Hardware Accelerator for RBLWE-based Post-Quantum CryptographyabstractRing-Binary-Learning-with-Errors (RBLWE)-based post-quantum cryptography (PQC) is a promising scheme suitable for lightweight applications. This paper presents an efficient hardware systolic accelerator for RBLWE-based PQC, targeting high-performance applications. We have briefly given the algorithmic background for the proposed design. Then, we have transferred the proposed algorithmic operation into a new systolic accelerator. Lastly, field-programmable gate array (FPGA) implementation results have confirmed the efficiency of the proposed accelerator. Tianyou Bao, José Luis Imaña, Pengzhou He, Jiafeng Xie |
CODES+ISSS | 4 |
| 2022 | Ultra Low-Complexity Implementation of Binary Ring-LWE based Post-Quantum Cryptography on FPGA PlatformabstractPost-quantum cryptography (PQC) has drawn substantial attention from various communities recently since the existing public-key cryptosystems are proven to be vulnerable to the attacks launched from well-established quantum computers. The Ring-Learning-with-Errors (Ring-LWE)-based encryption scheme is an important lattice-based PQC. The binary Ring-LWE (BRLWE), a variant of Ring-LWE, has been proposed to build lightweight PQC with much smaller complexity for resource-constrained applications. So far, however, limited reports have been released on the ultra low-complexity implementations of the BRLWE-based scheme (especially on the hardware platform). Therefore, in this paper, we propose to obtain a novel implementation of the BRLWE-based PQC on the Field-Programmable Gate Array (FPGA) platform with ultra low-complexity. We have proposed three layers of coherent interdependent efforts: (i) motivation and the related derivation for the key operation of the BRLWE-based scheme are presented first; (ii) the corresponding hardware structure is then efficiently mapped from the proposed algorithm with the help of algorithm-architecture co-implementation techniques; (iii) the final complexity analysis and FPGA implementation results show that the proposed designs have significantly smaller area-time complexities than the state-of-the-art designs. Overall, the proposed designs possess multiple unique features and thus are desirable for emerging lightweight applications. Jiafeng Xie, Pengzhou He, Tianyou Bao |
FPGA | 1 |
| 2022 | HPMA-Saber: High-Performance Polynomial Multiplication Accelerator for KEM SaberabstractThe recent research in post-quantum cryptography (PQC) field has gradually switched to efficient implementation of PQC algorithms on hardware platforms. As polynomial multiplication is typically one of the critical operations within lattice-based PQC, its hardware acceleration has drawn significant attention from the research community recently. We propose a high-speed processing strategy to construct a new High-performance Polynomial Multiplication Accelerator (HPMA) for key encapsulation mechanism (KEM) Saber. Firstly, we have given a detailed mathematical derivation to obtain a low-latency processing algorithm for Saber polynomial multiplication. Then, we have innovatively used the derived the proposed algorithm to construct a new structure HPMA for FPGA implementation. Lastly, we have demonstrated the superior performance of the proposed HPMA-Saber by comparing with state-of-the-art works. The proposed design strategy is highly efficient and the obtained results can be useful for the PQC research community. Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
ICCD | 4 |
| 2022 | FastRE: Towards Fast Relation Extraction with Convolutional Encoder and Improved Cascade Binary Tagging FrameworkabstractRecent work for extracting relations from texts has achieved excellent performance. However, most existing methods pay less attention to the efficiency, making it still challenging to quickly extract relations from massive or streaming text data in realistic scenarios. The main efficiency bottleneck is that these methods use a Transformer-based pre-trained language model for encoding, which heavily affects the training speed and inference speed. To address this issue, we propose a fast relation extraction model (FastRE) based on convolutional encoder and improved cascade binary tagging framework. Compared to previous work, FastRE employs several innovations to improve efficiency while also keeping promising performance. Concretely, FastRE adopts a novel convolutional encoder architecture combined with dilated convolution, gated unit and residual connection, which significantly reduces the computation cost of training and inference, while maintaining the satisfactory performance. Moreover, to improve the cascade binary tagging framework, FastRE first introduces a type-relation mapping mechanism to accelerate tagging efficiency and alleviate relation redundancy, and then utilizes a position-dependent adaptive thresholding strategy to obtain higher tagging accuracy and better model generalization. Experimental results demonstrate that FastRE is well balanced between efficiency and performance, and achieves 3-10$\times$ training speed, 7-15$\times$ inference speed faster, and 1/100 parameters compared to the state-of-the-art models, while the performance is still competitive. Our code is available at \url{https://github.com/seukgcode/FastRE}. Peng Wang 0004, Jiafeng Xie, Qiqing Luo |
IJCAI | 4 |
| 2022 | Hardware Implementation of High-Performance Polynomial Multiplication for KEM SaberabstractRecent advances in quantum computing have initiated a new round of cryptosystem innovation as the existing public-key cryptosystems are proven to be vulnerable to quantum attacks. Several types of cryptographic algorithms have been proposed for possible post-quantum cryptography (PQC) candidates and the lattice-based key encapsulation mechanism (KEM) Saber is one of the most promising algorithms. Noticing that the polynomial multiplication over ring is the key arithmetic operation of KEM Saber, in this paper, we propose a novel strategy for efficient implementation of polynomial multiplication on the hardware platform. First of all, we present the proposed mathematical derivation process for polynomial multiplication. Then, the proposed hardware structure is provided. Finally, field-programmable gate array (FPGA) based implementation results are obtained, and it is shown that the proposed design has better performance than the existing ones. The proposed polynomial multiplication can be further deployed to construct efficient hardware cryptoprocessors for KEM Saber. Yazheng Tu, Pengzhou He, Chiou-Yng Lee, Danai Chasaki, Jiafeng Xie |
ISCAS | 5 |
| 2022 | Certificateless signature schemes in Industrial Internet of Things: A comparative survey
Syed Sajid Ullah, Ihsan Ali, Jiafeng Xie, Venkata N. Inukollu |
Comput. Commun. | 4 |
| 2022 | AFIA: ATPG-Guided Fault Injection Attack on Secure Logic Locking
Yadi Zhong, Ayush Jain 0002, M. Tanjidur Rahman, Navid Asadizanjani, Jiafeng Xie, Ujjwal Guin |
J. Electron. Test. | 5 |
| 2022 | Efficient Hardware Arithmetic for Inverted Binary Ring-LWE Based Post-Quantum CryptographyabstractRing learning-with-errors(RLWE)-based encryption scheme is a lattice-based cryptographic algorithm that constitutes one of the most promising candidates for Post-Quantum Cryptography (PQC) standardization due to its efficient implementation and low computational complexity.Binary Ring-LWE (BRLWE) is a new optimized variant of RLWE, which achieves smaller computational complexity and higher efficient hardware implementations. In this paper, two efficient architectures based onLinear-Feedback Shift Register(LFSR) for the arithmetic used inInverted Binary Ring-LWE (InvBRLWE)-based encryption scheme are presented, namely the operation of$A\cdot B+C$over the polynomial ring$\mathbb {Z}_{q}/(x^{n}+1)$. The first architecture optimizes the resource usage for major computation and has a novel input processing setup to speed up the overall processing latency with minimized input loading cycles. The second architecture deploys an innovative serial-in serial-out processing format to reduce the involved area usage further yet maintains a regular input loading time-complexity. Experimental results show that the architectures presented here improve the complexities obtained by competing schemes found in the literature, e.g., involving 71.23% less area-delay product than recent designs. Both architectures are highly efficient in terms of area-time complexities and can be extended for deploying in different lightweight application environments. José Luis Imaña, Pengzhou He, Tianyou Bao, Yazheng Tu, Jiafeng Xie |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Efficient Implementation of Finite Field Arithmetic for Binary Ring-LWE Post-Quantum Cryptography Through a Novel Lookup-Table-Like MethodabstractThe recent advance in the post-quantum cryptography (PQC) field has gradually shifted from the theory to the implementation of the cryptosystem, especially on the hardware platforms. Following this trend, in this paper, we aim to present efficient implementations of the finite field arithmetic (key component) for the binary Ring-Learning-with-Errors (Ring-LWE) PQC through a novel lookup-table (LUT)-like method. In total, we have carried out four stages of interdependent efforts: (i) an algorithm-hardware co-design driven derivation of the proposed LUT-like method is provided detailedly for the key arithmetic of the BRLWE scheme; (ii) the proposed hardware architecture is then presented along with the internal structural description; (iii) we have also presented a novel hybrid size structure suitable for flexible operation, which is the first report in the literature; (iv) the final implementation and comparison processes have also been given, demonstrating that our proposed structures deliver significant improved performance over the state-of-the-art solutions. The proposed designs are highly efficient and are expected to be employed in many emerging applications. Jiafeng Xie, Pengzhou He, Wujie Wen |
DAC | 1 |
| 2021 | CROP: FPGA Implementation of High-Performance Polynomial Multiplication in Saber KEM based on Novel Cyclic-Row Oriented Processing StrategyabstractThe rapid advancement in quantum technology has initiated a new round of post-quantum cryptography (PQC) related exploration. The key encapsulation mechanism (KEM) Saber is an important module lattice-based PQC, which has been selected as one of the PQC finalists in the ongoing National Institute of Standards and Technology (NIST) standardization process. On the other hand, however, efficient hardware implementation of KEM Saber has not been well covered in the literature. In this paper, therefore, we propose a novel cyclic-row oriented processing (CROP) strategy for efficient implementation of the key arithmetic operation of KEM Saber, i.e., the polynomial multiplication. The proposed work consists of three layers of interdependent efforts: (i) first of all, we have formulated the main operation of KEM Saber into desired mathematical forms to be further developed into CROP based algorithms, i.e., the basic version and the advanced higher-speed version; (ii) then, we have followed the proposed CROP strategy to innovatively transfer the derived two algorithms into desired polynomial multiplication structures with the help of a series of algorithm-architecture co-implementation techniques; (iii) finally, detailed complexity analysis and implementation results have shown that the proposed polynomial multiplication structures have better area-time complexities than the state-of-the-art solutions. Specifically, the field-programmable gate array (FPGA) implementation results show that the proposed design, e.g., the basic version has at least less 11.2% area-delay product (ADP) than the best competing one (Cyclone V device). The proposed high-performance polynomial multipliers offer not only efficient operation for output results delivery but also possess low-complexity feature brought by CROP strategy. The outcome of this work is expected to provide useful references for further development and standardization process of KEM Saber. Jiafeng Xie, Pengzhou He, Chiou-Yng Lee |
ICCD | 1 |
| 2020 | Efficient Subquadratic Space Complexity Digit-Serial Multipliers over GF(2m) based on Bivariate Polynomial Basis RepresentationabstractDigit-serial finite field multipliers over GF(2m) with subquadratic space complexity are critical components to many applications such as elliptic curve cryptography. In this paper, we propose a pair of novel digit-serial multipliers based on bivariate polynomial basis (BPB). Firstly, we have proposed a novel digit-serial BPB multiplication algorithm based on a new decomposition strategy. Secondly, the proposed algorithm is properly mapped into a pair of pipelined and non-pipelined digit-serial multipliers. Lastly, through the detailed complexity analysis and comparison, the proposed designs are found to have less area-time complexities than the competing ones. Chiou-Yng Lee, Jiafeng Xie |
ASP-DAC | 2 |
| 2020 | Special Session: The Recent Advance in Hardware Implementation of Post-Quantum CryptographyabstractThe recent advancement in quantum technology has initiated a new round of cryptosystem innovation, i.e., the emergence of Post-Quantum Cryptography (PQC). This new class of cryptographic schemes is intended to be mathematically resistant against any known attacks using quantum computers, but, at the same time, be fully implementable using traditional semiconductor technology. The National Institutes of Standards and Technology (NIST) has already started the PQC standardization process, and the initial pool of 69 submissions has been reduced to 26 Round 2 candidates. Echoing the pace of the PQC "revolution," this paper gives a detailed and thorough introduction to recent advances in the hardware implementation of PQC schemes, including challenges, new implementation methods, and novel hardware architectures. Specifically, we have: (i) described the challenges and rewards of implementing PQC in hardware; (ii) presented the novel methodology for the design-space exploration of PQC implementations using high-level synthesis (HLS); (iii) introduced a new underexplored PQC scheme (binary Ring-Learning-with-Errors), as well as its novel hardware implementation for possible lightweight applications. The overall content delivered by this paper could serve multiple purposes: (i) provide useful references for the potential learners and the interested public; (ii) introduce new areas and directions for potential research to the VTS community; (iii) facilitate the PQC standardization process and the exploration of related new ways of implementing cryptography in existing and emerging applications. Jiafeng Xie, Kanad Basu, Kris Gaj, Ujjwal Guin |
VTS | 1 |
| 2019 | Embracing Systolic: Super Systolization of Large-Scale Circulant Matrix-vector Multiplication on FPGA with Subquadratic Space ComplexityabstractThe recent advance in artificial intelligence (AI) technology has led to a new round of systolic structure innovation. Many AI accelerators have employed systolic structure to realize the core large-scale matrix-vector multiplication for high-performance processing, which has a complexity of $o(n^2)$ for matrix size of $n\times n$ (difficult to be implemented on the field-programmable gate array (FPGA) platform). To overcome this drawback, in this paper, we propose a super systolization strategy to implement the core circulant matrix-vector multiplication into a systolic structure with subquadratic space complexity. The proposed effort is carried out through two stages of coherent interdependent efforts: (i) a novel matrix-vector multiplication algorithm based on Toeplitz matrix-vector product (TMVP) approach is proposed to obtain subquadratic space complexity; (ii) a series of optimization techniques are introduced to map the proposed algorithm into desired systolic structure. Finally, detailed complexity analysis and comparison have been conducted to prove the efficiency of the proposed strategy. The proposed strategy is highly efficient and can be extended in many neural network based hardware implementation platforms. Jiafeng Xie, Chiou-Yng Lee |
FPGA | 1 |
| 2019 | LSM: Novel Low-Complexity Unified Systolic Multiplier over Binary Extension FieldabstractUnified (hybrid field-size) systolic multiplier over GF(2m) (binary extension field) has attracted significant attentions from research communities recently as it can be used in reconfigurable cryptographic processors. In this paper, we present a novel Low-complexity unified Systolic Multiplier (LSM) design strategy to address the above mentioned challenge. First, a novel multiplication algorithm for unified implementation is proposed first with detailed mathematical derivation. Then, an efficient systolic structure is then presented based on the proposed algorithm (with the help of several mapping techniques). Finally, a thorough complexity analysis is given to verify the superior performance of the proposed design (the proposed structure involves better complexities than the newly reported design). Jiafeng Xie, Chiou-Yng Lee |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | Efficient Scalable Three Operand Multiplier Over GF(2^m) Based on Novel Decomposition StrategyabstractIt is expected that efficient scalable three operand multiplier (STOM) over GF(2m) (polynomial basis) generally can provide quite a number of superior benefits such as low-complexity and flexibility on processing bits and thus is very ideal to many applications like elliptic curve cryptography and pairing cryptography. The actual efficient hardware implementation of STOM, however, is still not covered in the literature. Based on this consideration, in this paper, we propose a novel decomposition strategy based design scheme to obtain efficient STOM on hardware platforms. First of all, a novel decomposition strategy (summarized as Toeplitz Matrix Oriented Karatsuba Algorithm, TMOKA) based STOM is presented with detailed mathematical derivation. Then, the proposed STOM structure is introduced along with a number of optimization techniques. Finally, the complexity analysis and comparison have been given to confirm the efficiency of the proposed STOM, e.g., the proposed structure with scalable digit-size of 64 has at least 50.8% less area-delay product (ADP) than the TOM employing a newly reported finite field multiplier ([8]) on the FPGA platform. The proposed STOM can thus be extended and employed in many cryptographic applications. Chiou-Yng Lee, Jiafeng Xie |
ICCD | 2 |
| 2019 | Low-Complexity Systolic Multiplier for GF(2m) using Toeplitz Matrix-Vector Product MethodabstractLow-complexity systolic multipliers for GF(2m) are required in several high-performance cryptographic systems. In this paper, we propose a novel design strategy to derive efficient systolic multiplier for GF(2m) based on Toeplitz Matrix-Vector Product (TMVP) approach. The proposed work is carried out through two coherent interdependent stages. (i) A novel multiplication algorithm based on TMVP method to obtain subquadratic space complexity is proposed first. (ii) The proposed algorithm is then mapped unto to a novel and efficient architecture which is optimized further to derive a low-complexity systolic structure. The complexity analysis and comparison show that the proposed design outperforms the existing work. The proposed design can thus be used in many practical cryptosystems. Jiafeng Xie, Chiou-Yng Lee, Pramod Kumar Meher |
ISCAS | 1 |
| 2019 | Predicting Future Instance Segmentation with Contextual Pyramid ConvLSTMsabstractDespite the remarkable progress in instance segmentation, the problem of predicting future instance segmentation remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting pyramid features to represent unobserved future frames. However, they mainly predict features for each pyramid level independently, and ignore the underlying structural relationship between features of different levels. Jiangxin Sun, Jiafeng Xie, Jianfang Hu, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001 |
ACM Multimedia | 2 |
| 2019 | Novel Systolization of Subquadratic Space Complexity Multipliers Based on Toeplitz Matrix-Vector Product ApproachabstractSystolic finite field multiplier over GF(2m), because of its superior features such as high throughput and regularity, is highly desirable for many demanding cryptosystems. On the other side, however, obtaining high-performance systolic multiplier with relatively low hardware cost is still a challenging task due to the fact that the systolic structure usually involves large area complexity. Based on this consideration, in this paper, we propose to carry out two novel coherent interdependent efforts. First, a new digit-serial multiplication algorithm based on polynomial basis over binary field (GF(2m)) is proposed. Novel Toeplitz matrix-vector product (TMVP)-based decomposition strategy is employed to derive an efficient subquadratic space complexity. Second, The proposed algorithm is then innovatively mapped into a low-complexity systolic multiplier, which involves less area-time complexities than the existing ones. A series of resource optimization techniques also has been applied on the multiplier which optimizes further the proposed design (it is the first report on digit-serial systolic multiplier based on TMVP approach covering all irreducible polynomials, to the best of our knowledge). The following complexity analysis and comparison confirm the efficiency of the proposed multiplier, that is, it has lower area-delay product (ADP) than the existing ones. The extension of the proposed multiplier for bit-parallel implementation is also considered in this paper. Jeng-Shyang Pan 0001, Chiou-Yng Lee, Anissa Sghaier, Medien Zeghid, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | Novel Bit-Parallel and Digit-Serial Systolic Finite Field Multipliers Over $GF(2^m)$ Based on Reordered Normal BasisabstractEfficient implementation of finite field multipliers based on a reordered normal basis (RNB) is highly desirable in the current/emerging cryptosystems since it offers almost free realization of squaring operation. Therefore, in this paper, we propose novel bit-parallel and digit-serial finite field multipliers over GF(2m) based on RNB. By efficient transformation of the core multiplication algorithm using a unique circular shifting feature, we have derived an efficient algorithm for low-complexity systolic mapping. Both bit-parallel and digit-serial structures of the multipliers are then obtained and optimized to enhance the area-time efficiency. We have also utilized the unique feature of the proposed multiplication algorithm to obtain the systolic multipliers by Karatsubalike decomposition. Detailed analysis and comparison show the superior performance of the proposed implementation. For example, the proposed regular and Karatsuba-based bit-parallel designs involve at least 48.4% less area-delay product (ADP) and 42.2% less power-delay product (PDP) than the best existing ones (37.7% and 55.3% less ADP and PDP on field-programmable gate array platform), respectively. The proposed multipliers, because of their lower area-time complexities, can be used for efficient realization of cryptographic applications. Jiafeng Xie, Chiou-Yng Lee, Pramod Kumar Meher, Zhi-Hong Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Improving Fast Segmentation With Teacher-Student Learning
Jiafeng Xie, Bing Shuai, Jianfang Hu, Wei-Shi Zheng 0001 |
BMVC | 1 |
| 2018 | Low Area-Delay Complexity Digit-Level Parallel-In Serial-Out Multiplier Over GF(2m) Based on Overlap-Free Karatsuba AlgorithmabstractOverlap-free Karatsuba algorithm (OFKA) is one of the Karatsuba algorithms (KAs) which can be employed to reduce the computation complexity of fast/high-precision polynomial based multiplication (the space complexity of the product can be reduced from O(m2) to O(mlog23)). Meanwhile, digit-level (DL) finite field multipliers over GF(2m), generally can be categorized as DL parallel-in serial-out (DL-PISO) and DL serial-in parallel-out (DL-SIPO) styles, have gained substantial attentions in cryptographic related applications (such as elliptic curve cryptosystem (ECC)) recently due to their efficient performance in area-delay tradeoffs. In this paper, aim at deriving an efficient DL-PISO multiplier with sub-quadratic space complexity, we present a novel polynomial basis multiplication through a combination of three coherent interdependent efforts. First of all, a novel bivariate polynomial multiplication algorithm using two steps of reduction is presented. Then, a new DL-PISO polynomial basis multiplication algorithm using OFKA over GF(2m) is introduced as well as its corresponding structure. Finally, the complexity and comparison are detailed given to confirm the efficiency of the proposed DL-PISO multiplier over the existing DL-PISO and DL-SIPO designs, i.e., the proposed one has lower area-delay product (ADP) when compared with the competing ones. The proposed DL-PISO multiplier is highly regular and hence can be applied in many resource-constrained environments. Chiou-Yng Lee, Jiafeng Xie |
ICCD | 2 |
| 2018 | Reliable Inversion in GF(28) With Redundant Arithmetic for Secure Error Detection of Cryptographic ArchitecturesabstractIn secure cryptographic primitives, such as block ciphers, the reliability of hardware implementations needs to be closely considered because faults in the hardware implementations can potentially reduce or impact on the underlying security. In this paper, we present approaches to detect errors in hardware implementations of the inversion in GF(28). The proposed approaches are based on both nonredundant and redundant arithmetic, utilizing normal basis (nonredundant) and two redundant Galois field representations, i.e., polynomial ring representation and redundantly represented basis through tower fields. To the best of our knowledge, this is the first work focusing on the error detection architectures for redundant arithmetic-based inversion in GF(28). The presented signature-based schemes in this paper are general and can be applied to block ciphers with 8-bit S-boxes, such as Camellia, SMS4, the advanced encryption standard, and CLEFIA. We present the results of error simulations and application-specific integrated circuit implementations to demonstrate the utility of the presented schemes. Based on the specific implementation's security/reliability objectives and the overhead/degradation tolerance for implementation/performance metrics, one can fine-tune and tailor the proposed work to achieve more reliable inversions in GF(28). Mehran Mozaffari Kermani, Amir Jalali, Reza Azarderakhsh, Jiafeng Xie, Kim-Kwang Raymond Choo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | FPGA Realization of Low Register Systolic All-One-Polynomial Multipliers Over $GF(2^{m})$ and Their Applications in Trinomial MultipliersabstractSystolic all-one-polynomial (AOP) multipliers usually suffer from the problem of high register complexity, especially in field-programmable gate array (FPGA) platforms where the register resources are not that abundant. In this paper, we have shown that the AOP-based systolic multipliers can easily achieve low register-complexity implementations and the proposed architectures can be employed as computation cores to derive efficient implementations of systolic Montgomery multipliers based on trinomials. First, we propose a novel data broadcasting scheme in which the register complexity involved within existing AOP-based systolic multipliers is significantly reduced. We have found out that the modified AOP-based structure can be packed as a standard computation core. Next, we propose a novel Montgomery multiplication algorithm that can fully employ the proposed AOP-based computation core. The proposed Montgomery algorithm employs a novel precomputed-modular operation, and the systolic structures based on this algorithm fully inherit the advantages brought from the AOP-based core (low register complexity, low critical-path delay, and low latency) except some marginal hardware overhead brought by a precomputation unit. The proposed architectures are then implemented by Xilinx ISE 14.1 and it is shown that compared with the existing designs, the proposed designs achieve at least 61.8% and 47.6% less area-delay product and power-delay product than the best of competing designs, respectively. Pingxiuqi Chen, Shaik Nazeem Basha, Mehran Mozaffari Kermani, Reza Azarderakhsh, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2017 | Low-Complexity Digit-Level Systolic Gaussian Normal Basis MultiplierabstractNormal basis multiplication over GF(2m) is widely used in various applications such as elliptic curve cryptography. As a special class of normal basis with low complexity, Gaussian normal basis (GNB) has received considerable attention recently. In this paper, we propose a novel decomposition algorithm to develop a digit-level (DL) low-complexity systolic structure for GNB multiplication over GF(2m). First, we propose two algorithms separately to achieve a systolic GNB multiplier with low critical path delay and low register complexity. Next, we present the corresponding structure according to the proposed algorithm (combination of previous two proposed algorithms). Compared with the existing systolic DL GNB multipliers (through both the theoretical and application-specific integrated circuit comparison), the proposed multiplier achieves significantly less area-delay product (ADP), e.g., for a systolic structure of digit size of 8 for GF(2409), the proposed structure has 12.3% less ADP compared to the best of the existing designs, for the same digit size. Qiliang Shao, Zhenji Hu, Shaobo Chen, Pingxiuqi Chen, Jiafeng Xie |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2013 | Hardware-Efficient Realization of Prime-Length DCT Based on Distributed ArithmeticabstractThis paper presents an efficient decomposition scheme for hardware-efficient realization of discrete cosine transform (DCT) based on distributed arithmetic. We have proposed an efficient design for the implementation of cyclic convolution based on a group distributed arithmetic (GDA) technique where the read-only memory size could be reduced over the existing GDA-based design. The proposed structure for DCT implementation, based on the new decomposition scheme and proposed design of GDA-based cyclic convolution, involves significantly less area complexity than the existing one. For example, to implement the DCT of transform length N = 17, the proposed design needs a lookup table of 128 words, while the existing design for N = 16 requires a lookup table of 256 words. From the synthesis results, it is found that proposed design involves significantly less area, gives higher throughput, and consumes less power compared to the existing designs of nearly the same or lower lengths. Jiafeng Xie, Pramod Kumar Meher |
IEEE Trans. Computers | 1 |
| 2013 | Low Latency Systolic Montgomery Multiplier for Finite Field $GF(2^{m})$ Based on PentanomialsabstractIn this paper, we present a low latency systolic Montgomery multiplier over$GF(2^{m})$based on irreducible pentanomials. An efficient algorithm is presented to decompose the multiplication into a number of independent units to facilitate parallel processing. Besides, a novel so-called “pre-computed addition” technique is introduced to further reduce the latency. The proposed design involves significantly less area-delay and power-delay complexities compared with the best of the existing designs. It has the same or shorter critical-path and involves nearly one-fourth of the latency of the other in case of the National Institute of Standards and Technology recommended irreducible pentanomials. Jiafeng Xie, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Low-Complexity Multiplier for GF(2m) Based on All-One PolynomialsabstractThis paper presents an area-time-efficient systolic structure for multiplication overGF(2m) based on irreducible all-one polynomial (AOP). We have used a novel cut-set retiming to reduce the duration of the critical-path to one XOR gate delay. It is further shown that the systolic structure can be decomposed into two or more parallel systolic branches, where the pair of parallel systolic branches has the same input operand, and they can share the same input operand registers. From the application-specific integrated circuit and field-programmable gate array synthesis results we find that the proposed design provides significantly less area-delay and power-delay complexities over the best of the existing designs. Jiafeng Xie, Pramod Kumar Meher |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Low-latency area-delay-efficient systolic multiplier over GF(2m) for a wider class of trinomials using parallel register sharingabstractSystolic structures for finite field multiplication involve large number of registers for parallel implementation, while bit-serial implementations require a large computation time, which increases along with the order of the field. In this paper, we present a novel scheme for the decomposition of the multiplication over GF(2m) based on irreducible trinomials into several independent units that facilitates maximal resister sharing and low-latency parallel implementation. It is shown that the proposed design involves significantly less area-delay complexity compared with the best of the corresponding existing systolic designs, and could be used for a wider class of trinomials. Jiafeng Xie, Pramod Kumar Meher |
ISCAS | 1 |