Zewen Ye

dblp:330/1619 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-3623-3554ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HTCNN: High-Throughput Batch CNN Inference With Homomorphic Encryption
abstract
Homomorphic Encryption (HE) technology allows for processing encrypted data, breaking through data isolation barriers and providing a promising solution for privacy-preserving computation. The integration of HE technology into Convolutional Neural Network (CNN) inference shows potential in addressing privacy issues in identity verification, medical imaging diagnosis, and various other applications. The CKKS HE algorithm stands out as a popular option for homomorphic CNN inference due to its capability to handle real number computations. However, challenges such as computational delays and resource overhead present significant obstacles to the practical implementation of homomorphic CNN inference, largely due to the complex nature of HE operations. In addition, current methods for speeding up homomorphic CNN inference primarily address individual images or large batches of input images, lacking a solution for efficiently processing a moderate number of input images with fast homomorphic inference capabilities. In response to these challenges, we introduce a novel leveled homomorphic CNN inference scheme aimed at reducing latency and improving throughput using the CKKS scheme. Our proposed inference strategy involves mapping multiple inputs to a set of ciphertext by exploiting the sliding window properties of convolutions to utilize CKKS's inherent Single-Instruction-Multiple-Data (SIMD) capability. To mitigate the delay associated with homomorphic CNN inference, we introduce optimization techniques, including mask-weight merging, rotation multiplexing, stride convolution segmentation, and folding rotations. The efficacy of our homomorphic inference scheme is demonstrated through evaluations carried out on the MNIST and CIFAR-10 datasets. Specifically, results from the MNIST dataset on a single CPU thread show that inference for 163 images can be completed in 10.4 seconds with an accuracy of 98.95%, which is a 6.9× throughput improvement over state-of-the-art works. Comparative analysis with existing methodologies highlights the superior performance of our proposed inference scheme in terms of latency, throughput, communication overhead, and memory utilization.
Zewen Ye, Tianyu Wang 0037, Tianshun Huang, Yonggen Li, Chengxuan Wang, Ray C. C. Cheung, Kejie Huang
IEEE Trans. Dependable Secur. Comput.1
2025 MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
abstract
Vector quantization(VQ) is a hardware-friendly DNN compression method that can reduce the storage cost and weight-loading datawidth of hardware accelerators. However, conventional VQ techniques lead to significant accuracy loss because the important weights are not well preserved. To tackle this problem, a novel approach called MVQ is proposed, which aims at better approximating important weights with a limited number of codewords. At the algorithm level, our approach removes the less important weights through N:M pruning and then minimizes the vector clustering error between the remaining weights and codewords by the masked k-means algorithm. Only distances between the unpruned weights and the codewords are computed, which are then used to update the codewords. At the architecture level, our accelerator implements vector quantization on an EWS (Enhanced weight stationary) CNN accelerator and proposes a sparse systolic array design to maximize the benefits brought by masked vector quantization.
Shuaiting Li, Chengxuan Wang, Juncan Deng, Zeyu Wang 0010, Zewen Ye, Zongsheng Wang, Haibin Shen, Kejie Huang
ASPLOS (1)5
2025 FastViT: Real-Time Linear Attention Accelerator for Dense Predictions of Vision Transformer (ViT)
abstract
The commercial success of generative artificial intelligence (GenAI) has driven an exponential surge in demand for real-time inference in Vision Transformer (ViT) applications, including latency-sensitive domains in autonomous driving, medical imaging and computational photography. This paper introduces FastViT, a high-performance and energy-efficient hardware accelerator for emerging kernel function-based linear attention mechanisms. By leveraging cost-efficient multiplication, mixed-precision quantisation and optimised data flow, FastViT improves real-time performance for high-resolution dense prediction tasks. Compared to existing approaches, experiments demonstrate that FastViT achieves higher throughput and energy efficiency while maintaining negligible accuracy degradation and balanced resource allocation. In the future, we will improve its scalability for next-generation hardware equipped with advanced DSP cores.
Zhuoheng Ran, Zewen Ye, Chong Wu 0007, Ray C. C. Cheung, Hong Yan 0001
ISCAS2
2025 Titan-I: An Open-Source, High Performance RISC-V Vector Core
abstract
Vector processing has evolved from early systems like the CDC STAR-100 and Cray-1 to modern ISAs like ARM's Scalable Vector Extension (SVE) and RISC-V Vector (RVV) extensions.However, scaling vector processing for contemporary workloads presents challenges due to overheads in traditional architectures.We introduce Titan-I (T1), an out-of-order (OoO) RVV architecture designed
Jiuyang Liu, Qinjun Li, Yunqian Luo, Jiongjia Lu, Shupei Fan, Jianhao Ye, Yanqi Yang, Zewen Ye, Yuhang Zeng, Wei Cong, Xuecheng Zou, Mingyu Gao 0001
MICRO11
2025 High-Radix/Mixed-Radix NTT Multiplication Algorithm/Architecture Co-Design Over Fermat Modulus
abstract
Polynomial multiplication using Number Theoretic Transform (NTT) is crucial in lattice-based post-quantum cryptography (PQC) and fully homomorphic encryption (FHE), with modulusqsignificantly affecting performance. Fermat moduli of the form$2^{2^{n}} + 1$, such as 65537, offer efficiency gains due to simplified modular reduction and powers-of-2 twiddle factors in NTT. While Fermat moduli have been directly applied or explored for incorporation into existing schemes, Fermat NTT-based polynomial multiplication designs remain underexplored in fully exploiting the benefits of Fermat moduli. This work presents a high-radix/mixed-radix NTT architecture tailored for Fermat moduli, which improves the utilization of the powers-of-2 twiddle factors in large transform sizes. In most cases, our design achieves a 30%–85% reduction in DSP area-time product (ATP) and a 70%–100% reduction in BRAM ATP compared to state-of-the-art designs with smaller or equivalent modulus, while maintaining competitive LUT and FF ATP, underscoring the potential of Fermat NTT-based polynomial multipliers in lattice-based cryptography.
Yile Xing, Guangyan Li, Zewen Ye, Ryan W. L. Luk, Donald Donglong Chen, Hong Yan 0001, Ray C. C. Cheung
IEEE Trans. Computers3
2025 PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum Processor
abstract
Post-quantum cryptography (PQC) has rapidly evolved in response to the emergence of quantum computers, with the US National Institute of Standards and Technology (NIST) selecting four finalist algorithms for PQC standardization in 2022, including the Falcon digital signature scheme. Hawk is currently the only lattice-based candidate in NIST Round 2 additional signatures. Falcon and Hawk are based on the NTRU lattice, offering compact signatures, fast generation, and verification suitable for deployment on resource-constrained Internet-of-Things (IoT) devices. Despite the popularity of ML-DSA and ML-KEM, research on NTRU-based schemes has been limited due to their complex algorithms and operations. Falcon and Hawk's performance remains constrained by the lack of parallel execution in crucial operations like the Number Theoretic Transform (NTT) and Fast Fourier Transform (FFT), with data dependency being a significant bottleneck. This paper enhances NTRU-based schemes Falcon and Hawk through hardware/software co-design on a customized Single-Instruction-Multiple-Data (SIMD) processor, proposing new SIMD hardware units and instructions to expedite these schemes along with software optimizations to boost performance. Our NTT optimization includes a novel layer merging technique for SIMD architecture to reduce memory accesses, and the use of modular algorithms (Signed Montgomery and Improved Plantard) targets various modulus data widths to enhance performance. We explore applying layer merging to accelerate fixed-point FFT at the SIMD instruction level and devise a dual-issue parser to streamline assembly code organization to maximize dual-issue utilization. A System-on-chip (SoC) architecture is devised to improve the practical application of the processor in real-world scenarios. Evaluation on 28$nm$technology and field programmable gate array (FPGA) platform shows that our design and optimizations can increase the performance of Hawk signature generation and verification by over 7$\times$.
Zewen Ye, Junhao Huang 0001, Tianshun Huang, Yudan Bai, Guangyan Li, Donald Donglong Chen, Ray C. C. Cheung, Kejie Huang
IEEE Trans. Computers1
2025 RVSLH: Acceleration of Postquantum Standard SLH-DSA With Customized RISC-V Processor
abstract
Postquantum cryptography (PQC) has developed quickly in response to the rise of quantum computers. The US National Institute of Standards and Technology (NIST) recently released three PQC standards, one of which is the hash-based standard stateless hash-based digital signature standard (SLH-DSA), built on SPHINCS+ selected during the NIST Round 3 submissions. Despite its potential, SLH-DSA’s performance is hindered by inefficient execution in the hash function and extensive memory accesses, with the data dependency of the hash function presenting a notable bottleneck. This brief aims to enhance the efficiency of SHAKE-based SLH-DSA schemes using hardware/software co-design on a customized RISC-V processor. We incorporate tightly coupled hardware units and instructions on RISC-V to expedite SLH-DSA, coupled with memory optimizations to enhance overall performance. The contributions of this brief are twofold. First, our design introduces customized single-instruction-multiple-data (SIMD) instructions and corresponding computation hardware units to accelerate Keccak, the essential operation of SHAKE256. In addition, our design streamlines hash operation memory accesses by reorganizing memory space. The proposed processor’s performance is evaluated on Artix-7 field-programmable gate array (FPGA) and 28-nm technology application-specific integrated circuit (ASIC), demonstrating approximately 15 times acceleration compared with the baseline design. Furthermore, it surpasses recent state-of-the-art works in terms of the performance-power–area product.
Zewen Ye, Xin Li 0177, Chuhui Wang, Ray C. C. Cheung, Kejie Huang
IEEE Trans. Very Large Scale Integr. Syst.1
2024 A Folded Computation-in-Memory Accelerator for Fast Polynomial Multiplication in BIKE
Chuhui Wang, Zewen Ye, Haibin Shen, Kejie Huang
Euro-Par (2)2
2024 ProgramGalois: A Programmable Generator of Radix-4 Discrete Galois Transformation Architecture for Lattice-Based Cryptography
abstract
Lattice-based cryptography (LBC) has been established as a prominent research field, with particular attention on post-quantum cryptography (PQC) and fully homomorphic encryption (FHE). As the implementing bottleneck of PQC and FHE, number theoretic transform (NTT) has been extensively studied. However, current works struggled with scalability, hindering their adaptation to various parameters, such as bit width and polynomial length. In this article, we proposed a novel Discrete Galois Transformation (DGT) algorithm utilizing the radix-4 variant to achieve a higher level of parallelism to the existing NTT. Furthermore, to implement the efficient radix-4 DGT adapting more LBCs, we proposed a set of scalable building blocks, including a modified Barrett modular multiplier accepting arbitrary modulus with only one integer multiplier, a radix-4 DGT butterfly unit, and a stream permutation network. The proposed modules are implemented on the Xilinx Virtex-7 and U250 FPGA to evaluate resource utilization and performance. Lastly, a design space exploration framework is proposed to generate optimized radix-4 DGT hardware constrained by polynomial and platform parameters. The sensitivity analysis showcases the generated hardware’s performance and scalability. The implementation results on the Xilinx Virtex-7 and U250 FPGA show significant performance improvements over the state-of-the-art works, which reached at least 35%, 192%, and 68% area-time product improvements in terms of LUTs, BRAMs, and DSPs, respectively.
Guangyan Li, Zewen Ye, Donald Donglong Chen, Wangchen Dai, Gaoyu Mao, Kejie Huang, Ray C. C. Cheung
ACM Trans. Reconfigurable Technol. Syst.2