Antian Wang

dblp:272/2552 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-2271-486XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HEDWIG: Homomorphic Encryption Accelerator Design Using BFV-HPS With HiGh-Speed Fixed-Point Approximation
Antian Wang, Weihang Tan, Zhenyu Xu 0007, Tao Wei 0001, Caiwen Ding, Keshab K. Parhi, Yingjie Lao
FPGA1
2024 HERMES: Homomorphic Encryption over Residual Number System for Multi-level EvaluationS
abstract
Homomorphic encryption enables computations on the ciphertext to preserve data privacy. However, its practical deployment has been hindered by the significant computational overhead compared to the plaintext computations. In response to this challenge, we present HERMES, a novel hardware acceleration system designed to explore the computation flow of the CKKS homomorphic encryption bootstrapping process. Among the major contributions of our proposed architecture, we first analyze the properties of the CKKS computation data flow and propose a new scheduling strategy by partitioning the computation modules into general-purpose and special-purpose modular computation modules to allow smaller resource consumption and flexible scheduling. The computation modules are also reconfigurable to reduce the memory access overhead during the intermediate computation. We also optimize the CKKS computation dataflow to improve the regularity with reduced control overhead.
Antian Wang, Keshab K. Parhi, Yingjie Lao
ICCAD1
2024 PaReNTT: Low-Latency Parallel Residue Number System and NTT-Based Long Polynomial Modular Multiplication for Homomorphic Encryption
abstract
High-speed long polynomial multiplication is important for applications in homomorphic encryption (HE) and lattice-based cryptosystems. This paper addresses low-latency hardware architectures for long polynomial modular multiplication using the number-theoretic transform (NTT) and inverse NTT (iNTT). Parallel NTT and iNTT architectures are proposed to reduce the number of clock cycles to process the polynomials. Chinese remainder theorem (CRT) is used to decompose the modulus into multiple smaller moduli. Our proposed architecture, namely PaReNTT, makes three novel contributions. First, cascaded parallel NTT and iNTT architectures are proposed such that any buffer requirement for permuting the product of the NTTs before it is input to the iNTT is eliminated. This is achieved by using different folding sets for the NTTs and iNTT. Second, a novel approach to expand the set of feasible special moduli is presented where the moduli can be expressed in terms of a few signed power-of-two terms. Third, novel architectures for pre-processing for computing residual polynomials using the CRT and post-processing for combining the residual polynomials are proposed. These architectures significantly reduce the area consumption of the pre-processing and post-processing steps. The proposed long modular polynomial multiplications are ideal for applications that require low latency and high sample rate such as in the cloud, as these feed-forward architectures can be pipelined at arbitrary levels. Pipelining and latency tradeoffs are also investigated. Compared to a prior design, the proposed architecture reduces latency by a factor of 49.2, and the area-time products (ATP) for the lookup table and DSP, ATP(LUT) and ATP(DSP), respectively, by 89.2% and 92.5%. Specifically, we show that for$n=4096$and a 180-bit coefficient, the proposed 2-parallel architecture requires 6.3 Watts of power while operating at 240 MHz, with 6 moduli, each of length 30 bits, using Xilinx Virtex Ultrascale+ FPGA.
Weihang Tan, Sin-Wei Chiu, Antian Wang, Yingjie Lao, Keshab K. Parhi
IEEE Trans. Inf. Forensics Secur.3
2023 NNTesting: Neural Network Fault Attacks Detection Using Gradient-Based Test Vector Generation
abstract
Recent studies have shown Neural Networks (NNs) are highly vulnerable to fault attacks. This work proposes a novel defensive framework, NNTesting, for detecting the fault attack and recovering the model. We first leverage gradient-based optimization to generate a set of high-quality Test Vectors (TVs) that effectively differentiate faulty profile models and further optimize the TV set by reducing the TVs through compression. The selected final TV set is then used to recover the model. The effectiveness of the proposed method is comprehensively evaluated on a wide range of models across various benchmark datasets. For instance, we successfully generate more than thousands of TV candidates using a gradient-based generation method. After compression, we achieve up to 94.76% detection success rate with only 140 TVs on the CIFAR-10 dataset.
Antian Wang, Bingyin Zhao, Weihang Tan, Yingjie Lao
DAC1
2023 High-Speed VLSI Architectures for Modular Polynomial Multiplication via Fast Filtering and Applications to Lattice-Based Cryptography
abstract
This paper presents a low-latency hardware accelerator for modular polynomial multiplication for lattice-based post-quantum cryptography and homomorphic encryption applications. The proposed novel modular polynomial multiplier exploits the fast finite impulse response (FIR) filter architecture to reduce the computational complexity of the schoolbook modular polynomial multiplication. We also extend this structure to fast$M$-parallel architectures while achieving low-latency, high-speed, and full hardware utilization. We comprehensively evaluate the performance of the proposed architectures under various polynomial settings as well as in the Saber scheme for post-quantum cryptography as a case study. The experimental results show that our proposed modular polynomial multiplier reduces the computation time and area-time product, respectively, compared to the state-of-the-art designs.
Weihang Tan, Antian Wang, Xinmiao Zhang 0001, Yingjie Lao, Keshab K. Parhi
IEEE Trans. Computers2
2021 A Task-Oriented Dialogue Architecture via Transformer Neural Language Models and Symbolic Injection
abstract
Recently, transformer language models have been applied to build both task-and non-taskoriented dialogue systems.Although transformers perform well on most of the NLP tasks, they perform poorly on context retrieval and symbolic reasoning.Our work aims to address this limitation by embedding the model in an operational loop that blends both natural language generation and symbolic injection.We evaluated our system on the multi-domain DSTC8 data set and reported joint goal accuracy of 75.8% (ranked among the first half positions), intent accuracy of 97.4% (which is higher than the reported literature), and a 15% improvement for success rate compared to a baseline with no symbolic injection.These promising results suggest that transformer language models can not only generate proper system responses but also symbolic representations that can further be used to enhance the overall quality of the dialogue management as well as serving as scaffolding for complex conversational reasoning.
Oscar J. Romero, Antian Wang, John Zimmerman, Aaron Steinfeld, Anthony Tomasic
SIGDIAL2
2021 NoPUF: A Novel PUF Design Framework Toward Modeling Attack Resistant PUFs
abstract
With the rapid development and globalization of the semiconductor industry, hardware security has emerged as a critical concern. New attacking and tampering methods are continuously challenge current hardware protection methods. Combating these powerful attacks is of great importance in securing hardware devices. This paper proposes a novel framework to protect Physical Unclonable Function (PUF) against modeling attacks, denominated as Noisy PUF (NoPUF). NoPUF exploits structural unpredictability to improve overall security. We present several PUF architectures under the proposed framework that could reconfigure a conventional reliable PUF to a noisy PUF. The reconfigured PUF becomes inherently unreliable and hence achieves a higher resistance against modeling attacks. Moreover, since only a small portion of the Challenge-Response Pairs (CRPs) are required for authentication, the designer can use the information obtained from the initial reliable PUF configuration to find CRPs, which are still reliable in the noisy PUF configuration for authentication. Exploiting such information asymmetry between designer and attacker is the nexus of the proposed NoPUF design methodology. Experimental results show that we can achieve a maximum attacker and designer accuracy difference of 44.79% for a 64-stage NoPUF candidate architecture while ensuring high reliability for selected challenges.
Antian Wang, Weihang Tan, Yuejiang Wen, Yingjie Lao
IEEE Trans. Circuits Syst. I Regul. Pap.1