Huaizhi Zhang

dblp:177/2015 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating Scalability Challenges in LUT-Based Neural Networks via Pruning Optimisations
abstract
Modern deep neural networks heavily rely on a large number of multiply-accumulate operations, which constitute the predominant computational cost. To address this, Look-Up Table (LUT)-based matrix multiplications have emerged as a promising alternative for reducing the computational cost and time of the multiply-accumulate operations in a neural network. However, the LUT-based neural network still faces the scalability challenge due to the inherent limitations of LUT-based matrix multiplication. To mitigate these scalability limitations, this paper proposes a scalable and energy-efficient LUT-based approximate matrix multiplication unit (LUT-MU) constituting the basic component of the neural networks by integrating a pruning strategy on the MADDNESS algorithm, a LUT-based matrix multiplication methodology. With increasing problem size and precision demands in matrix multiplication, our proposed LUT-MU architecture effectively constrains resource expansion. The case study shows that deploying our LUT-MU in neural network architectures, including fully connected layers (MNIST) and ResNets (CIFAR-10, ImageNet)—on XCZU7EV and XCZU19EG FPGAs, produces up to 1.6× throughput improvement and 4.2× energy efficiency gains over mainstream CUDA-based network implementations, and 1.8× energy efficiency compared to leading quantised neural network implementations, with moderate impact on accuracy. Compared to original MADDNESS-based neural networks, our LUT-MU shows 1.3 to 2.6× resource savings based on various resolution configuration settings of MADDNESS.
Xuqi Zhu, Huaizhi Zhang, Chandrajit Pal, Sangeet Saha, Klaus D. McDonald-Maier, Xiaojun Zhai
IEEE Trans. Computers2
2025 Late Breaking Results: Approximated LUT-Based Neural Networks for FPGA Accelerated Inference
abstract
This work presents LUT-MU, an approximated LUT-based Matrix Multiplication (MM) architecture designed for FPGA-based Neural Network (NN) inference across. The proposed architecture maximises the utilisation of on-chip memory bandwidth through dedicated memory distribution and pipeline design, addressing performance limitations inherent to LUT-based MM. Experimental evaluation demonstrates that LUT-MU achieves a four-fold improvement in NN inference throughput whilst reducing hardware resource consumption by 80% with only a 5% decline in accuracy. These results validate that our optimisation approach successfully resolves the performance constraints caused by the limited arithmetic intensity and memory bandwidth, enabling the LUT-MU to serve as a foundation for efficient NN acceleration systems.
Xuqi Zhu, Huaizhi Zhang, Tamim M. Al-Hasan, Klaus D. McDonald-Maier, Xiaojun Zhai
DATE3
2025 Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic Systems
abstract
Integrating Large Language Models (LLMs) into modern robotic systems presents significant computational and energy constraint challenges, particularly for human-centered robotic applications. This paper presents a novel hardware optimization technique for deploying LLMs on resource-constrained embedded devices, achieving an up to 77% reduction in computational latency through an FPGA implementation in comparison to other popular embedded computing devices (e.g., CPU and GPUs). Additionally, we demonstrate our methodology by deploying a LLaMA 2-7B model on a Unitree Go2 robotic dog integrated with the proposed FPGA platform. The proposed optimization framework preserves real-time interaction capabilities while significantly reducing computational and energy overhead, facilitating efficient natural language processing for human-robot interaction in safety-critical and dynamic environments. Experimental results demonstrate that the FPGA-based LLaMA 2-7B implementation achieves up to 6.06-fold and 1.95-fold higher throughput compared to baseline CPU and GPU implementations while maintaining comparable inference accuracy. Furthermore, the proposed FPGA design surpasses existing state-of-the-art FPGA implementations, delivering a 30% improvement in computational efficiency.
Huaizhi Zhang, Tamim M. Al-Hasan, Xuqi Zhu, Weiyong Si, Klaus D. McDonald-Maier, Xiaojun Zhai
IROS1
2025 A Privacy-aware Quantilisation Approach for Efficient Edge Deep Learning Accelerator
abstract
Data privacy is one of the key concerns in machine learning model applications at the edge, especially in sensitive domains such as the healthcare sector. Here adversaries may exploit Membership Inference Attacks (MIAs) to determine if particular data points were used as part of datasets from the model’s training set, potentially leading to further data leakage issues. Although privacy preservation techniques like differential privacy (DP) can mitigate such risks during the training phase, this often results in degradation of model accuracy, making them less suitable for cloud training and edge deployment paradigms. For AI edge applications, existing research works for designing neural network accelerators primarily prioritize computational performance and power efficiency as their design target. In this paper, we introduce a novel privacy-aware quantilisation approach for deep learning accelerators and analyse the tradeoff between computational efficiency and privacy protection. The proposed system allows to adjust privacy constraints through tunable parameters, enabling flexible deployment on edge devices while meeting privacy and performance constraints. We have evaluated the proposed design on an AMD VCK190 board using a range of hypothetical MIA benchmarks. The results demonstrate that the proposed approach can effectively reduce the success rate of MIA attacks across multiple performance metrics.
Huaizhi Zhang, Xuqi Zhu, Klaus D. McDonald-Maier, Xiaojun Zhai
ISCAS1
2023 Poseidon: Practical Homomorphic Encryption Accelerator
abstract
With the development of the important solution for privacy computing, the explosion of data size and computing intensity in Fully Homomorphic Encryption (FHE) has brought enormous challenges to the hardware design. In this paper, we propose a practical FHE accelerator - "Poseidon", which focuses on improving the hardware resource and bandwidth consumption. Poseidon supports complex FHE operations like Bootstrapping, Keyswitch, Rotation and so on, under limited FPGA resources. It refines these operations by abstracting five key operators: Modular Addition (MA), Modular Multiplication (MM), Number Theoretic Transformation (NTT), Automorphsim and Shared Barret Reduction (SBT). These operators are combined and reused to implement higher-level FHE operations. To utilize the FPGA resources more efficiently and improve the parallelism, we adopt the radix-based NTT algorithm and propose HFAuto, an optimized automorphism implementation suitable for FPGA. Then, we design the hardware accelerator based on the optimized key operators and HBM to maximize computational efficiency. We evaluate Poseidon with four domain-specific FHE benchmarks on Xilinx Alveo U280 FPGA. Empirical results show that the efficient reuse of the operator cores and on-chip storage enables superior performance compared with the state-of-the-art GPU, FPGA and accelerator ASICs. We highlight the following results: (1) up to 370× speedup over CPU for the basic operations of FHE; (2) up to 1300×/52× speedup over CPU and the FPGA solution for the key operators; (3) up to 10.6×/8.7× speedup over GPU and the ASIC solution for the FHE benchmark.
Yinghao Yang 0001, Huaizhi Zhang, Shengyu Fan, Mingzhe Zhang 0005, Xiaowei Li 0001
HPCA2
2018 Topic detection and tracking on heterogeneous information
abstract
Given the proliferation of social media and the abundance of news feeds, a substantial amount of real-time content is distributed through disparate sources, which makes it increasingly difficult to glean and distill useful information. Although combining heterogeneous sources for topic detection has gained attention from several research communities, most of them fail to consider the interaction among different sources and their intertwined temporal dynamics. To address this concern, we studied the dynamics of topics from heterogeneous sources by exploiting both their individual properties (including temporal features) and their inter-relationships. We first implemented a heterogeneous topic model that enables topic–topic correspondence between the sources by iteratively updating its topic–word distribution. To capture temporal dynamics, the topics are then correlated with a time-dependent function that can characterise its social response and popularity over time. We extensively evaluate the proposed approach and compare to the state-of-the-art techniques on heterogeneous collection. Experimental results demonstrate that our approach can significantly outperform the existing ones.
Long Chen 0008, Huaizhi Zhang, Joemon M. Jose, Hai-Tao Yu 0003, Yashar Moshfeghi, Peter Triantafillou
J. Intell. Inf. Syst.2
2016 Probabilistic Topic Modelling with Semantic Graph
Long Chen 0008, Joemon M. Jose, Hai-Tao Yu 0003, Fajie Yuan, Huaizhi Zhang
ECIR5