Zehuan Zhang

dblp:322/1825 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0006-0607-060XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EDSSC: An Efficient FPGA-based Accelerator for Dynamic Sparse Spectral Clustering
abstract
This paper presents a hardware-software co-design for efficient dynamic sparse spectral clustering. We introduce a heterogeneous streaming architecture that replaces dense SVD with a sparse iterative solver. Key contributions include: (1) dynamic graph generation to minimize memory footprint; (2) tile-centric dataflow to maximize sparse matrix reuse; and (3) an INT8 mixed-precision datapath. Implemented on Xilinx VCK190, our design achieves up to 22× speedup and 39× higher energy efficiency than an RTX 3090 GPU.
Zhengyan Liu, Ce Guo 0002, Zehuan Zhang, Qiang Liu 0011, Wayne Luk
FCCM3
2026 CODESCA: Co-Design for Spectral Clustering Acceleration
abstract
Abstract: Spectral clustering is powerful but limited by O(N³) complexity. We present CODESCA, a co-design on Xilinx VCK190. By offloading sparse graph construction to the host and utilizing a quantization-aware block power iteration engine on FPGA, CODESCA achieves 22× speedup over CPU and 9× better throughput-per-watt than RTX 3080 GPU, enabling efficient edge data mining.
Zhengyan Liu, Ce Guo 0002, Zehuan Zhang, Qiang Liu 0011, Wayne Luk
FPGA3
2024 Accelerating MRI Uncertainty Estimation with Mask-Based Bayesian Neural Network
abstract
Accurate and reliable Magnetic Resonance Imaging (MRI) analysis is particularly important for adaptive radio-therapy, a recent medical advance capable of improving cancer diagnosis and treatment. Recent studies have shown that IVIM-NET, a deep neural network (DNN), can achieve high accuracy in MRI analysis, indicating the potential of deep learning to enhance diagnostic capabilities in healthcare. However, IVIM-NET does not provide calibrated uncertainty information needed for reliable and trustworthy predictions in healthcare. Moreover, the expensive computation and memory demands of IVIM-NET reduce hardware performance, hindering widespread adoption in realistic scenarios. To address these challenges, this paper proposes an algorithm-hardware co-optimization flow for high-performance and reliable MRI analysis. At the algorithm level, a transformation design flow is introduced to convert IVIM-NET to a mask-based Bayesian Neural Network (BayesNN), facilitating reliable and efficient uncertainty estimation. At the hardware level, we propose an FPGA-based accelerator with several hardware optimizations, such as mask-zero skipping and operation reordering. Experimental results demonstrate that our co-design approach can satisfy the uncertainty requirements of MRI analysis, while achieving 7.5 times and 32.5 times speedup on an Xilinx VU13P FPGA compared to GPU and CPU implementations with reduced power consumption.
Zehuan Zhang, Matej Genci, Hongxiang Fan, Andreas Wetscherek, Wayne Luk
ASAP1
2024 Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
abstract
The increasing deployment of artificial intelligence (AI) for critical decision-making amplifies the necessity for trustworthy AI, where uncertainty estimation plays a pivotal role in ensuring trustworthiness. Dropout-based Bayesian Neural Networks (BayesNNs) are prominent in this field, offering reliable uncertainty estimates. Despite their effectiveness, existing dropout-based BayesNNs typically employ a uniform dropout design across different layers, leading to suboptimal performance. Moreover, as diverse applications require tailored dropout strategies for optimal performance, manually optimizing dropout configurations for various applications is both error-prone and labor-intensive. To address these challenges, this paper proposes a novel neural dropout search framework that automatically optimizes both the dropout-based BayesNNs and their hardware implementations on FPGA. We leverage one-shot supernet training with an evolutionary algorithm for efficient dropout optimization. A layer-wise dropout search space is introduced to enable the automatic design of dropout-based BayesNNs with heterogeneous dropout configurations. Extensive experiments demonstrate that our proposed framework can effectively find design configurations on the Pareto frontier. Compared to manually-designed dropout-based BayesNNs on GPU, our search approach produces FPGA designs that can achieve up to 33× higher energy efficiency. Compared to state-of-the-art FPGA designs of BayesNN, the solutions from our approach can achieve higher algorithmic performance and energy efficiency.
Zehuan Zhang, Hongxiang Fan, Hao Mark Chen, Lukasz Dudziak, Wayne Luk
DAC1
2021 Knowledge-Based Multiple Lightweight Attribute Networks for Zero-Shot Learning
Zehuan Zhang, Qiang Liu 0011, Difei Guo
ICONIP (5)1