Zhengyan Liu

dblp:168/4579 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EDSSC: An Efficient FPGA-based Accelerator for Dynamic Sparse Spectral Clustering
abstract
This paper presents a hardware-software co-design for efficient dynamic sparse spectral clustering. We introduce a heterogeneous streaming architecture that replaces dense SVD with a sparse iterative solver. Key contributions include: (1) dynamic graph generation to minimize memory footprint; (2) tile-centric dataflow to maximize sparse matrix reuse; and (3) an INT8 mixed-precision datapath. Implemented on Xilinx VCK190, our design achieves up to 22× speedup and 39× higher energy efficiency than an RTX 3090 GPU.
Zhengyan Liu, Ce Guo 0002, Zehuan Zhang, Qiang Liu 0011, Wayne Luk
FCCM1
2026 CODESCA: Co-Design for Spectral Clustering Acceleration
abstract
Abstract: Spectral clustering is powerful but limited by O(N³) complexity. We present CODESCA, a co-design on Xilinx VCK190. By offloading sparse graph construction to the host and utilizing a quantization-aware block power iteration engine on FPGA, CODESCA achieves 22× speedup over CPU and 9× better throughput-per-watt than RTX 3080 GPU, enabling efficient edge data mining.
Zhengyan Liu, Ce Guo 0002, Zehuan Zhang, Qiang Liu 0011, Wayne Luk
FPGA1
2026 Self-weighted low-rank representation for multivariate compositional data
Zhengyan Liu
Neural Networks1
2025 SDTA: An Efficient Sparse DNN Training Accelerator with Data Hierarchical Pre-fetching and Dynamic Scheduling
abstract
Recently, training deep neural networks (DNNs) on edge devices has attracted much attention due to its strong adaptability and avoidance of private data transmission. However, limited computational, storage, and energy resources pose significant challenges for edge devices. The structural and computational redundancies in DNNs create opportunities for sparse training through model pruning and zero-computation skipping. Although feasible, the sparse training accelerator design encounters common issues, such as redundant data duplication and unbalanced workloads, caused by irregular sparsity. To address these issues, this paper proposes a sparse DNN training accelerator, SDTA, together with a hierarchical pre-fetching buffer and a dynamic scheduler to achieve high design efficiency. The SDTA is deployed on the FPGA XCVU3P platform. Compared to the prior FPGA-based accelerators and the GPU, SDTA improves the energy efficiency by up to 2.29×, the storage utilization efficiency by up to 7.37×, and the computational efficiency by up to 1.9×. Compared to the dense accelerator, it achieves a speedup of up to 5.88×, while ensuring model accuracy.
Mengting Wang, Yuntao Han, Yingchang Mao, Peng Shao, Zhengyan Liu, Qiang Liu 0011
ISCAS5
2025 Self-weighted subspace clustering with adaptive neighbors
Zhengyan Liu
Neurocomputing1
2025 Locality-constrained double-layer structure scaled simplex multi-view subspace clustering
Zhengyan Liu
Vis. Comput.1
2024 An Efficient FPGA-based Depthwise Separable Convolutional Neural Network Accelerator with Hardware Pruning
abstract
Convolutional neural networks (CNNs) have been widely deployed in computer vision tasks. However, the computation and resource intensive characteristics of CNN bring obstacles to its application on embedded systems. This article proposes an efficient inference accelerator on Field Programmable Gate Array (FPGA) for CNNs with depthwise separable convolutions. To improve the accelerator efficiency, we make four contributions: (1) an efficient convolution engine with multiple strategies for exploiting parallelism and a configurable adder tree are designed to support three types of convolution operations; (2) a dedicated architecture combined with input buffers is designed for the bottleneck network structure to reduce data transmission time; (3) a hardware padding scheme to eliminate invalid padding operations is proposed; and (4) a hardware-assisted pruning method is developed to support online tradeoff between model accuracy and power consumption. Experimental results show that for MobileNetV2 the accelerator achieves 10× and 6× energy efficiency improvement over the CPU and GPU implementation, and 302.3 frames per second and 181.8 GOPS performance that is the best among several existing single-engine accelerators on FPGAs. The proposed hardware-assisted pruning method can effectively reduce 59.7% power consumption at the accuracy loss within 5%.
Zhengyan Liu, Qiang Liu 0011, Shun Yan, Ray C. C. Cheung
ACM Trans. Reconfigurable Technol. Syst.1
2021 An FPGA-based MobileNet Accelerator Considering Network Structure Characteristics
abstract
Convolutional neural networks (CNNs) have been widely deployed in computer vision tasks. However, the computation and resource intensive characteristics of CNN bring obstacles to its application on embedded systems. MobileNet, as a representative of compact models, can reduce the amount of parameters and computation. A high-performance inference accelerator on FPGA for MobileNet is proposed in this paper. With respect to the three types of convolution operations, multiple parallel strategies are exploited and the corresponding hardware structures such as input buffer and configurable adder tree are designed. With respect to the bottleneck block, a dedicated architecture is proposed to reduce data transmission time. In addition, a hardware padding scheme to improve the efficiency of padding is proposed. The accelerator implemented on Virtex-7 FPGA reaches 70.8% Top-1 accuracy under 8-bit quantization. The accelerator achieves 302.3 FPS and 181.8 GOPS, which obtains 22.7x, 3.9x and 1.4x speedup compared to the implementations in Snapdragon 821 CPU, i7-6700HQ CPU and GTX 960M GPU, respectively.
Shun Yan, Zhengyan Liu, Chenglong Zeng, Qiang Liu 0011, Bowen Cheng, Ray C. C. Cheung
FPL2
2017 Multi-Robot Task Allocation Based on Cloud Ant Colony Algorithm
Zhengyan Liu, Fuxiao Tan
ICONIP (4)2