EDBT 2026 Demo / reviewers in the wild / expert
Yang Su 0003
dblp:17/686-3
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-3619-7936ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Low Multiplicative Depth Polynomial Evaluation Architectures for Homomorphic Encrypted DataabstractIn order to reduce the multiplicative depth required by high-order polynomial evaluation for homomorphic encrypted data, we propose two novel and low multiplicative depth polynomial evaluation algorithms: an x4-step nesting algorithm based on parity extraction (x4-NAPE) and a parallel algorithm with high data utilization (PAHU). Compared with the conventional Schoolbook method and Horner's rule, x4-NAPE can reduce about 50% of the ciphertext-ciphertext multiplication (CCM) and 75% of the multiplicative depth in the ideal limiting case. While PAHU does not reduce CCM, it can achieve a logarithmic growth of multiplicative depth, with the increase of polynomial length, the growth rate is much slower than x4-NAPE, Schoolbook method, and Horner's rule. Moreover, the hardware architectures of the above polynomial evaluation algorithms for homomorphic encrypted data are investigated and proposed. The proposed hardware architectures are assessed on the FPGA-based reconfigurable hardware platform for FHE named ReMCA. The assessment results demonstrate that under a fixed upper limit of multiplicative depth, our proposed architectures of x4-NAPE and PAHU support 2.67× and 14.3× the range of polynomial lengths of Schoolbook and Horner's rule, respectively. For polynomial evaluation of the same length, compared with architectures of Schoolbook and Horner's rule, our proposed architectures of x4-NAPE and PAHU can achieve up to 1.13× improvement in execution time, up to 50% reduction in multiplicative depth, and up to 2.39× improvement in the depth-time product. Jianfei Wang 0003, Fahong Zhang 0002, Yishuo Meng, Yang Su 0003, Chen Yang 0005 |
ASP-DAC | 5 |
| 2025 | A Reconfigurable and Area-Efficient Polynomial Multiplier Using a Novel In-Place Constant-Geometry NTT/INTT and Conflict-Free Memory Mapping SchemeabstractOut-of-place constant-geometry (CG) NTT usually has a simple and uniform memory access pattern. However, out-of-place CG NTT always requires ping-pong memory, resulting in a memory capacity requirement of$2N$. Therefore, we propose a novel radix-4 in-place CG (IPCG) NTT/INTT that reduces the capacity requirement from$2N$to N. An area-efficient and dynamical reconfigurable polynomial multiplier (RAEPM) based on IPCG NTT is proposed to speed up polynomial multiplication over rings. In RAEPM, a Barrett modular multiplier using area-efficient radix-4 booth multiplier is designed to reduce area. In addition, an odd-bank buffer structure is proposed to achieve conflict-free memory mapping independent of polynomial length N and NTT/INTT stage. Moreover, we also proposed an efficient modular reduction for specific numbers and introduced a division equivalent method to eliminate the odd number modular reduction and odd number division in addressing. RAEPM is implemented on Xilinx VC709 FPGA and runs at 294MHz clock frequency. Compared with the prior pure NTT accelerators, under the same parameters, RAEPM achieves a decrease of 39.02%$\sim ~57.63$% in area-time complexity of equivalent LUT, and a decrease of 15.97%$\sim ~49.24$% in area-time complexity of equivalent FF. Compared with the prior NTT-based polynomial multipliers, under the same parameters, RAEPM achieves a decrease of 35.48%$\sim ~90.81$% in area-time complexity of equivalent LUT, and a decrease of 24.24%$\sim ~88.41$% in area-time complexity of equivalent FF. Jianfei Wang 0003, Chen Yang 0005, Yishuo Meng, Fahong Zhang 0002, Siwei Xiang, Yang Su 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | A Scalable and Efficient Architecture for Binary Polynomial Multiplication in BIKE Utilizing Inter-/Inner-Wise Sparsity and Block-by-Block PipelineabstractEfficient binary polynomial multiplication (BPM) implementations are crucial for the practical deployment of bit flipping key encapsulation (BIKE) postquantum cryptography (PQC) due to its computation-intensive nature. To speed up BPM, this brief proposes a scalable and efficient architecture. The proposed architecture employs a novel blockwise sparsity algorithm, which segments sparse polynomials into blocks and leverages interblock and inner block sparsity to eliminate invalid computations, thereby significantly reducing computational operations. Moreover, a scalable block-by-block pipeline structure, along with a multibank random access memory (RAM) for sparse polynomials, is designed to effectively process blocks, resulting in substantial enhancement in performance. Experimental results on Xilinx Artix-7 Field-Programmable Gate Arrays (FPGAs) demonstrate significant performance superiority on the proposed architecture, compared with existing approaches. Across different bandwidth settings of 16, 32, 64, or 128, our design can achieve$4.5\times \sim 35.1\times $,$4.9\times \sim 78.8\times $,$2.5\times \sim 112.7\times $, and$0.5\times \sim 164.2\times $speedup, respectively. Compared with state-of-the-art works, our design achieves$2.8\times \sim 152.0\times $improvements in area efficiency. Jianfei Wang 0003, Yishuo Meng, Fahong Zhang 0002, Yang Su 0003, Chen Yang 0005 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A Scalable and Efficient NTT/INTT Architecture Using Group-Based Pairwise Memory Access and Fast Interstage ReorderingabstractPolynomial multiplication is a significant bottleneck in mainstream postquantum cryptography (PQC) schemes. To speed it up, number theoretic transform (NTT) is widely used, which decreases the time complexity from${O}(n^{2})$to$O[n\log _{2}(n)]$. However, it is challenging to ensure optimal hardware efficiency in conjunction with scalability. This brief proposes a novel pipelined NTT/inverse-NTT (INTT) architecture on field-programmable gate array (FPGA). A group-based pairwise memory access (GPMA) scheme is proposed, and a scratchpad and reordering unit (SRU) is designed to form an efficient dataflow that simplifies control units and achieves almost$n/2$processing cycles on average for n-point NTT. Moreover, our architecture can support varying parameters. Compared to the state-of-the-art works, our architecture achieves up to$4.8\times $latency improvements and up to$4.3\times $improvements on area time product (ATP). Yushu Yang, Jianfei Wang 0003, Yang Su 0003, Chen Yang 0005 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | A High-Throughput Toom-Cook-4 Polynomial Multiplier for Lattice-Based Cryptography Using a Novel Winograd-Schoolbook AlgorithmabstractPolynomial multiplication over rings is a significant bottleneck of ring learning with error (RLWE)-based encryption. To speed it up, the number theoretic transform (NTT) and Toom-Cook-4 (TC4) are commonly used algorithms. Compared with NTT, TC4 is less restrictive and more flexible. However, there is a large opportunity at the algorithm level to improve the Schoolbook algorithm and postprocessing of TC4. Therefore, we propose a novel and efficient Winograd-Schoolbook algorithm that reduces multiplication by 29.1% (N = 256). We also propose a fused and low-density postprocessing that simplifies the algorithm flow and reduces multiplication by 56.25%. In total, these two-part improvements reduce the multiplication of TC4 by 32.47%. A high throughput and efficiency TC4 polynomial multiplier (TCMW) is proposed to speed up polynomial multiplication over rings. In TCMW, a highly parallel full pipelined structure without data waiting between modules is designed to make the parallelism of each module match and avoid the storage of intermediate results. In addition, based on the improved TC4, data buffers with data reuse, elementwise vector multiplication (EWVM) arrays, and efficient interpolation arrays are all designed to improve the performance and efficiency of TCMW. Implemented on the Xilinx VC709 field programmable gate array (FPGA) platform, TCMW can perform a TC4-based$256\times 256$polynomial multiplication over rings with an unrestricted modulus (as long as its factors do not contain 3 or 5) every$1.89~\mu $s at a 385 MHz clock frequency. Compared with prior designs of TC4, under the same conditions, the throughput of TCMW achieves an improvement of$1.91\times \,\,\sim \,\,7.71\times $, and the efficiency of LUT and DSP achieve improvements of$1.31\times \,\,\sim \,\,3.67\times $and$1.87\times \,\,\sim \,\,4.92\times $, respectively. Jianfei Wang 0003, Chen Yang 0005, Fahong Zhang 0002, Yishuo Meng, Siwei Xiang, Yang Su 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | TCPM: A Reconfigurable and Efficient Toom-Cook-Based Polynomial Multiplier Over Rings Using a Novel Compressed Postprocessing AlgorithmabstractPolynomial multiplication over rings is a significant bottleneck of ring learning with error (RLWE)-based encryption. To speed it up, three algorithms are widely used, i.e., number theoretic transform (NTT), Schoolbook, and Toom-Cook. Compared with Schoolbook and NTT, Toom-Cook can achieve a better trade-off between performance and flexibility. However, in Toom-Cook postprocessing, there are many redundant steps and calculations that have not been eliminated. Therefore, we propose an efficient, compressed, and fused Toom-Cook postprocessing algorithm that reduces the number of steps and at least 33.33% of the arithmetic operations of postprocessing. A highly reconfigurable and efficient Toom-Cook-based polynomial multiplier (TCPM) is proposed to speed up polynomial multiplication over rings. In TCPM, a high-throughput and efficient heterogeneous processing element (PE) array is designed to exploit the parallelism of Toom-Cook, and based on the compressed algorithm, the PE array for postprocessing is scaled down. In addition, as it is provided with a reconfigurable evaluation module, a flexible polynomial data storage module and a universal PE array, TCPM can efficiently map and execute Toom-Cook-2, 3, and 4 on a unified hardware architecture. Implemented on the Xilinx VC709 field-programmable gate array (FPGA) platform, TCPM can perform a Toom-Cook-4-based$256\times256$polynomial multiplication over rings with a modulus of a power of two or a prime every 3.28$\mu \text{s}$at a 360-MHz clock frequency. It achieves a$2.47\times $to$50.11\times $speedup compared with the previous designs. Jianfei Wang 0003, Chen Yang 0005, Fahong Zhang 0002, Yishuo Meng, Yang Su 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | ReMCA: A Reconfigurable Multi-Core Architecture for Full RNS Variant of BFV Homomorphic EvaluationabstractFully homomorphic encryption (FHE) allows arbitrary computation on encrypted data and thus has potential in privacy-preserving computing. However, efficiency is still the bottleneck. In this paper we present an area-efficient and highly unified reconfigurable multi-core architecture (named ReMCA) for full Residue Number System (RNS) variant of Fan-Vercauteren variant of Brakerski’s scheme (RNS-BFV), which employs a variable number of reconfigurable processing elements (PEs) and RNS channels. The PE unit can be flexibly configured as NTT, INTT or modular multiplier, thereby avoiding the need of other extra computational units. To reduce the computational complexity, ReMCA merges the pre/post-processing into NTT/INTT and unifies the read/write structure of NTT and INTT. Also, a conflict-free memory access pattern that doesn’t need separate bit-reversal operation is proposed to optimize the memory access. Furthermore, targeting different computational requirements, a unified hardware architecture mapping model and data memory organization model are introduced, and all the computing units that RNS-BFV involved are optimized and mapped on ReMCA. ReMCA is evaluated on a Xilinx Virtex-7 FPGA platform. Running at 250MHz, it can perform 2260 homomorphic multiplication per second. When normalized to the same parameter set, the throughput and Area-Time-Products (ATPs) of ReMCA achieve$1.45\times \sim 5.51\times $and$1.58\times \sim 5.12\times $improvements. Yang Su 0003, Bai-Long Yang, Chen Yang 0005, Songyin Zhao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | A Highly Unified Reconfigurable Multicore Architecture to Speed Up NTT/INTT for Homomorphic Polynomial MultiplicationabstractThe ring learning with error (RLWE)-based fully homomorphic encryption (FHE) scheme has become one of the most promising FHE schemes. However, its performance is limited by the homomorphic multiplication, especially the polynomial multiplication which occupies major computing resources. Therefore, efficient implementation of polynomial multiplication is crucial for high-performance FHE applications. In this article, we present an area-efficient and highly unified reconfigurable multicore number theoretic transform (NTT)/inverse NTT (INTT) architecture (named MCNA), which employs NTT and INTT for polynomial multiplier with a variable number of reconfigurable processing elements. To reduce latency, MCNA merges the preprocessing and postprocessing into the constant-geometry NTT and INTT, respectively. Also, a reconfigurable modular multiplier based on digital signal processor (DSP) is proposed to speed up the modular multiplication. In order to avoid designing independent memory access pattern for INTT, a unified read/write structure of NTT/INTT is presented. Furthermore, a novel memory access pattern named “cyclic-sharing” is proposed to reduce 25% memory capacity. MCNA is evaluated on a Xilinx Virtex-7 field-programmable gate array (FPGA) platform. Running at 250-MHz clock frequency, the throughput of MCNA for NTT/INTT achieves$2.78\times \sim 9.32\times $improvements in comparison to prior works, while the area efficiency of lookup table (LUT) and flip-flop (FF) is improved by$1.25\times \sim 4.79\times $. For polynomial multiplication, the throughput of MCNA achieves$3.73\times \sim 7.69\times $enhancements, as well as$1.13\times \sim 14.8\times $area efficiency improvements. Yang Su 0003, Bai-Long Yang, Chen Yang 0005, Zepeng Yang, Yi-Wei Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |