Tianshun Huang

dblp:389/5359 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 85% Processor architecture and microarchitecture · 8% Integrated circuit design · 8%
Network and information security
2 papers
Cryptographic primitives and cryptanalysis · 90% Privacy and data protection · 10%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis › homomorphic encryption › fully homomorphic encryption
CKKS
1.012026
HTCNN: High-Throughput Batch CNN Inference With Homomorphic Encryption · IEEE Trans. Dependable Secur. Comput. 2026
Cryptographic primitives and cryptanalysis
homomorphic encryption
1.012026
HTCNN: High-Throughput Batch CNN Inference With Homomorphic Encryption · IEEE Trans. Dependable Secur. Comput. 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.012026
HTCNN: High-Throughput Batch CNN Inference With Homomorphic Encryption · IEEE Trans. Dependable Secur. Comput. 2026
Hardware accelerators and domain-specific architectures › security accelerator
privacy-preserving inference accelerator
1.012026
HTCNN: High-Throughput Batch CNN Inference With Homomorphic Encryption · IEEE Trans. Dependable Secur. Comput. 2026
Cryptographic primitives and cryptanalysis
post-quantum cryptography
0.912025
PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum Processor · IEEE Trans. Computers 2025
Hardware accelerators and domain-specific architectures
cryptographic accelerator
0.912025
PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum Processor · IEEE Trans. Computers 2025
Privacy and data protection › privacy-preserving machine learning
privacy-preserving machine learning inference
0.312026
HTCNN: High-Throughput Batch CNN Inference With Homomorphic Encryption · IEEE Trans. Dependable Secur. Comput. 2026
Processor architecture and microarchitecture › SIMD
SIMD processor
0.312025
PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum Processor · IEEE Trans. Computers 2025
Integrated circuit design
system-on-chip
0.312025
PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum Processor · IEEE Trans. Computers 2025

Methods — techniques the papers use, named apart from their topics

stride convolution segmentation · 2.0sliding window convolution · 2.0rotation multiplexing · 2.0mask-weight merging · 2.0folding rotations · 2.0SIMD packing · 2.0layer merging · 1.7improved plantard · 1.7fast fourier transform · 1.7number-theoretic transform · 0.9number theoretic transform · 0.9
YearPublicationVenuePosition
2026 HTCNN: High-Throughput Batch CNN Inference With Homomorphic Encryption
abstract
Homomorphic Encryption (HE) technology allows for processing encrypted data, breaking through data isolation barriers and providing a promising solution for privacy-preserving computation. The integration of HE technology into Convolutional Neural Network (CNN) inference shows potential in addressing privacy issues in identity verification, medical imaging diagnosis, and various other applications. The CKKS HE algorithm stands out as a popular option for homomorphic CNN inference due to its capability to handle real number computations. However, challenges such as computational delays and resource overhead present significant obstacles to the practical implementation of homomorphic CNN inference, largely due to the complex nature of HE operations. In addition, current methods for speeding up homomorphic CNN inference primarily address individual images or large batches of input images, lacking a solution for efficiently processing a moderate number of input images with fast homomorphic inference capabilities. In response to these challenges, we introduce a novel leveled homomorphic CNN inference scheme aimed at reducing latency and improving throughput using the CKKS scheme. Our proposed inference strategy involves mapping multiple inputs to a set of ciphertext by exploiting the sliding window properties of convolutions to utilize CKKS's inherent Single-Instruction-Multiple-Data (SIMD) capability. To mitigate the delay associated with homomorphic CNN inference, we introduce optimization techniques, including mask-weight merging, rotation multiplexing, stride convolution segmentation, and folding rotations. The efficacy of our homomorphic inference scheme is demonstrated through evaluations carried out on the MNIST and CIFAR-10 datasets. Specifically, results from the MNIST dataset on a single CPU thread show that inference for 163 images can be completed in 10.4 seconds with an accuracy of 98.95%, which is a 6.9× throughput improvement over state-of-the-art works. Comparative analysis with existing methodologies highlights the superior performance of our proposed inference scheme in terms of latency, throughput, communication overhead, and memory utilization.
Zewen Ye, Tianyu Wang 0037, Tianshun Huang, Yonggen Li, Chengxuan Wang, Ray C. C. Cheung, Kejie Huang
IEEE Trans. Dependable Secur. Comput.3
2025 PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum Processor
abstract
Post-quantum cryptography (PQC) has rapidly evolved in response to the emergence of quantum computers, with the US National Institute of Standards and Technology (NIST) selecting four finalist algorithms for PQC standardization in 2022, including the Falcon digital signature scheme. Hawk is currently the only lattice-based candidate in NIST Round 2 additional signatures. Falcon and Hawk are based on the NTRU lattice, offering compact signatures, fast generation, and verification suitable for deployment on resource-constrained Internet-of-Things (IoT) devices. Despite the popularity of ML-DSA and ML-KEM, research on NTRU-based schemes has been limited due to their complex algorithms and operations. Falcon and Hawk's performance remains constrained by the lack of parallel execution in crucial operations like the Number Theoretic Transform (NTT) and Fast Fourier Transform (FFT), with data dependency being a significant bottleneck. This paper enhances NTRU-based schemes Falcon and Hawk through hardware/software co-design on a customized Single-Instruction-Multiple-Data (SIMD) processor, proposing new SIMD hardware units and instructions to expedite these schemes along with software optimizations to boost performance. Our NTT optimization includes a novel layer merging technique for SIMD architecture to reduce memory accesses, and the use of modular algorithms (Signed Montgomery and Improved Plantard) targets various modulus data widths to enhance performance. We explore applying layer merging to accelerate fixed-point FFT at the SIMD instruction level and devise a dual-issue parser to streamline assembly code organization to maximize dual-issue utilization. A System-on-chip (SoC) architecture is devised to improve the practical application of the processor in real-world scenarios. Evaluation on 28$nm$technology and field programmable gate array (FPGA) platform shows that our design and optimizations can increase the performance of Hawk signature generation and verification by over 7$\times$.
Zewen Ye, Junhao Huang 0001, Tianshun Huang, Yudan Bai, Guangyan Li, Donald Donglong Chen, Ray C. C. Cheung, Kejie Huang
IEEE Trans. Computers3