EDBT 2026 Demo / reviewers in the wild / expert
Ruidian Zhan
dblp:269/5110
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-1918-3375ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NTT-LSU: Tightly Coupled Architecture for Efficient NTT Implementation on RISC-V ProcessorabstractPolynomial multiplication is one of the most computationally intensive operations in lattice-based cryptographic systems, directly impacting overall computational efficiency. Although the Number Theoretic Transform (NTT) reduces the time complexity of polynomial multiplication fromO(n2) toO(nlogn), the varying parameter requirements of different lattice algorithms limit the generality of hardware designs. To address this issue, we propose an innovative hardware architecture that tightly couples the NTT unit with the Load and Store Unit (LSU) in the pipeline of a RISC-V processor. We also propose a hybrid width data path method that effectively reduces data transfer time. Compared to previous designs, our architecture minimizes data transfer latency while enhancing computational flexibility and scalability. Specifically, we have customized NTT-related instructions to support Inverse Number Theoretic Transform (INTT) and various NTT parameter configurations. Experimental results demonstrate that our solution significantly shortens the NTT computation cycle, achieving over 10× speedup compared to software implementations. In comparison to existing solutions, our architecture exhibits superior area-time product (ATP) performance. Yinqiao Zhao, Zilong Xie, Ruidian Zhan, Xiaoming Xiong, Yun Chen 0004, Shuting Cai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | An FPGA-Efficient CNN Accelerator for Hybrid Model Compression With Scalable Bit-Serial and Bit-Parallel MACabstractMixed-precision quantization and unstructured pruning have emerged as two effective compression techniques, demonstrating great potential in reducing model size and computational cost in the deployment of convolutional neural networks (CNNs). However, their joint deployment still faces two major challenges: 1) The former introduces heterogeneous bit-widths in the bit-level, while the latter results in irregular sparsity in the value-level; their fundamentally incompatible data representations and computation patterns require two distinct types of hardware overhead to process them separately, which severely limits hardware execution efficiency. 2) Jointly applying both techniques often leads to notable accuracy degradation. In this paper, we propose a novel compression perspective that reinterprets zero-values generated by unstructured pruning as multiple consecutive 0-bits. We further introduce column-based bit-level sparsity, which provides a unified representation for weights after mixed-precision quantization and unstructured pruning, requiring only a single type of hardware overhead. Based on these techniques, we develop a hybrid compression framework that jointly optimizes model size, accuracy, and hardware implementation. Our method achieves weight/activation precision of 2.13b/4.06b on VGG16, delivering 7.40$\times $compression and 2.74$\times $speedup with 0.92% accuracy loss compared to the 8b baseline. Compared to state-of-the-art accelerators, our design achieves 1.12$\times $-6.23$\times $and 1.31$\times $-6.60$\times $improvements in energy efficiency and LUT efficiency when deploying VGG16 and ResNet50. Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Heng Mai, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | A Low-Cost Local Masking Radix-4 NTT Against Soft-Analytical Side-Channel AttacksabstractThe number theoretic transform (NTT) is essential for accelerating polynomial multiplication in lattice-based cryptography. However, it is vulnerable to soft-analytical side-channel attacks (SASCAs). Although local masking countermeasure provides theoretical resistance against such attacks, its direct implementation in Radix-4 NTT architecture leads to more than a 4 times increase in modular multiplications, resulting in substantial hardware overhead. To address this challenge, we propose the modular multiplication parallel mask sharing (MMPMS) scheme, which optimizes the modular multiplication parallelism of the Radix-4 butterfly units and shares random twiddle factors, thereby achieving a balance between hardware overhead and security. Then, we construct a complete local masking NTT/INTT algorithm and efficiently implement it on the Artix-7 field-programmable gate array (FPGA). Experimental results show that compared with the state-of-the-art local masking NTT, our scheme reduces the equivalent area and ATP overhead by more than 8.24 times and 6.74 times, respectively. In addition, a nonspecifict-test analysis indicates no significant side-channel leakage. Congwei Chen, Jinwei Pu, Jianxiong Zhang 0003, Jiaying Liao, Ruidian Zhan, Yun Chen 0004, Shuting Cai |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2026 | Efficient FPGA Acceleration for 4-bit CNNs via Quantization-Induced Structured Sparsity and LUT-Based MultiplicationabstractN:M structured sparsity is key to convolutional neural network (CNN) compression and acceleration, but two challenges remain. From the algorithm perspective, prior works have mainly focused on 8-bit quantized models, where N:M sparsity yields limited hardware efficiency. From the hardware perspective, the cost differences across N:M sparsity have not been analyzed. To address these issues, we present a unified algorithm–hardware co-design framework for 4-bit CNN acceleration. We show that 4-bit quantization induces over 80% zero weights and strongly structured sparsity, with over 95% of weight groups satisfying 4:8, 8:16, or 16:32 patterns. We propose a pruning-after-quantization (PAQ) algorithm that enforces strict N:M sparsity with minimal accuracy loss. We also analyze the hardware overhead of activation fetch units (AFUs) under different N:M sparsity patterns (4:8, 8:16, 16:32), revealing that the 4:8 AFU reduces look-up table (LUT) cost by up to 66.7% compared to 16:32. Finally, we introduce a 4-bit LUT-based sign-magnitude multiplier (LBSMM) requiring only 11 LUT6 resources, outperforming existing multipliers. Integrated on a Xilinx VCU118 field-programmable gate array (FPGA), our accelerator achieves$2.51\times $–$12.89\times $improvements in equivalent LUT efficiency over SOTA designs. The implementations of the PAQ algorithm and the RTL of LBSMM are available athttps://github.com/haden-01/PAQ-and-LBSMM.git Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | AES software and hardware system co-design for resisting side channel attacksabstractAbstract The threat of side‐channel attacks poses a significant risk to the security of cryptographic algorithms. To counter this threat, we have designed an AES system capable of defending against such attacks, supporting AES‐128, AES‐192, and AES‐256 encryption standards. In our system, the CPU oversees the AES hardware via the AHB bus and employs true random number generation to provide secure random inputs for computations. The hardware implementation of the AES S‐box utilizes complex domain inversion techniques, while intermediate data is shielded using full‐time masking. Furthermore, the system incorporates double‐path error detection mechanisms to thwart fault propagation. Our results demonstrate that the system effectively conceals key power information, providing robust resistance against CPA attacks, and is capable of detecting injected faults, thereby mitigating fault‐based attacks. Liguo Dong, Xinliang Ye, Libin Zhuang, Ruidian Zhan, M. Shamim Hossain |
Expert Syst. J. Knowl. Eng. | 4 |