EDBT 2026 Demo / reviewers in the wild / expert
Zhendong Zheng
dblp:214/1507
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-6486-3062ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UDP: A Universal DSP Packing Framework for Low-bitwidth MAC Acceleration on FPGAsabstractLow-bitwidth multiply-accumulate (MAC) operations are fundamental to efficient hardware acceleration of recent neural networks. Existing approaches often fail to fully exploit the potential of digital signal processing (DSP) blocks in FPGAs. They struggle to balance the high-bitwidth computational capabilities of DSPs with the low-precision quantization requirements of neural networks. DSP packing consolidates multiple low-precision operations into a single DSP unit, significantly enhancing MAC efficiency. However, current DSP packing solutions suffer from limited support for continuous operations, poor adaptation to different neural network computation patterns, and complex software-hardware deployment workflows. These issues result in insufficient utilization of DSP resources, limiting the overall effectiveness of acceleration. Jundong Wu, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
FPGA | 2 |
| 2026 | HE-DeepFM: An FHE Inference System for CTR Prediction with Efficient FM InteractionsabstractScoring models such as click-through rate (CTR) prediction underpin recommendation and advertising systems, but their features are highly sensitive, making plaintext cloud inference risky. Fully homomorphic encryption (FHE) enables inference directly on ciphertexts, yet homomorphic computation is expensive and bootstrapping often dominates end-to-end latency. We present HE-DeepFM, a FHE inference system for CTR prediction. We first design HE-FM, a homomorphic-friendly Factorization Machine (FM) operator that exploits CKKS SIMD packing to compute second-order interactions efficiently, thereby avoiding the naive O(F2) cost of pairwise feature interactions. Building on HE-FM, HE-DeepFM reduces bootstrapping under the same FHE budget, and can eliminate it for small model configurations. We implement HE-DeepFM with Orion and Lattigo and evaluate it on real-world datasets. On Criteo, HE-DeepFM reduces bootstrapping from 12 to 4 and achieves up to 2.85× end-to-end speedup while maintaining prediction quality comparable to the baseline. Qiyue Su, Hang Gu, Zhiguang Wang, Zhendong Zheng, Qianyu Cheng, Lei Gong 0003, Chao Wang 0003 |
SIGIR | 5 |
| 2026 | Out-of-Memory Graph Processing Acceleration via Algorithmic-Hardware Codesign on FPGAsabstractEmerging FPGA devices are widely used in graph processing and achieve high performance by exploiting fine-grained parallelism and high-bandwidth memory (HBM) subsystems with dozens of channels. However, the graph size increases rapidly and often exceeds the capacity of the FPGA device or one of the memory channels, incurring out-of-memory (OoM) faults and performance degradation. Existing approaches focus on improving the performance of on-device full graph processing, ignoring the support for hyper-scale graphs and cross-iteration acceleration. In this work, we propose SoGraph, a graph processing architecture for large graph acceleration on HBM-equipped FPGAs, cooperating with the host. To minimize PCIe traffic during the processing, we adopt three hardware-oriented optimizations: 1) activeness-driven subgraph scheduler, which minimizes on-device graph resident set; 2) lightweight state-aware processing extension in edge-centric paradigm; 3) bit-level update analyzer for to-be-transfer subgraph reduction. The results show that the prototype on Alveo U280 FPGA achieves an average of 22.3x performance speedup over the modified state-of-the-art FPGA design and 1.3x device energy efficiency over the GPU solution. Qianyu Cheng, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Huaping Chen 0001, Xuehai Zhou |
IEEE Trans. Computers | 3 |
| 2026 | LORA: A Latency-Oriented Recurrent Architecture for Large Language Model on Multi-FPGA Platform With Communication OptimizationabstractThe remarkable performance of Large Language Models (LLMs) has driven their widespread deployment in data centers to support diverse user-facing applications. However, the rapidly growing computational and storage demands of these models have made single-device deployment increasingly impractical. Prior research on LLM inference has primarily addressed this challenge through algorithmic optimizations such as quantization or by integrating customized hardware acceleration frameworks. As model parameters continue to scale, multi-device deployment has become a necessary approach for enabling efficient LLM inference. Nevertheless, constructing low-latency multi-device platforms for LLMs inference using available FPGA or GPU accelerators remains constrained by inefficient synchronization schemes or limited compute intensity in current architectures. Furthermore, existing solutions often lack co-optimized designs that effectively integrate communication with computation. To address these limitations, this paper proposes LORA, a low-latency end-to-end LLMs acceleration platform utilizing multiple FPGAs. Firstly, we optimize the synchronization timing within the LLMs to minimize storage, computation, and BRAM overhead. Secondly, we tightly couple communication and computation through techniques such as pipeline overlapping and input data packing. Next, we deploy homogeneous accelerators on each FPGA device, leveraging a recurrent architecture to further reduce inference latency. Finally, we apply FPGA-specific optimizations and conduct performance modeling and analysis of the acceleration framework to select optimal deployment parameters for various computational tasks. Implemented on Xilinx Alveo U280 FPGAs, LORA-F and LORA-Q achieve average speedups of 14.4× and 32.6×, respectively, compared to NVIDIA V100 GPUs when running modern LLMs. Compared with existing multi-FPGA accelerator platforms, LORA-F and LORA-Q demonstrate average performance improvements of up to 2.6× and 4.3×, respectively. Zhendong Zheng, Qianyu Cheng, Wenqi Lou, Lei Gong 0003, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | MoE-Sched: Enabling Efficient FPGA Deployment of Mixture-of-Experts Vision Transformers via Coordinated SchedulingabstractCompared to traditional Vision Transformers (ViT), Mixture-of-Experts ViTs (MoE-ViTs) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, field-programmable gate array (FPGA)-based accelerators for MoE-ViTs emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of efficiently managing the highly dynamic and sparse computation patterns, resulting in suboptimal tradeoffs between resource utilization and performance. To address the inefficiencies in deploying MoE-ViTs on FPGAs, we present MoE-Sched, a novel end-to-end accelerator that embraces a scheduling-centric design philosophy. Rather than optimizing isolated kernels, MoE-Sched coordinates multilevel scheduling, from fine-grained intrakernel streaming to module reuse and multidie mapping, to holistically balance latency, bandwidth (BW), and resource usage. We further integrate a hardware-aware quantization scheme tailored for streaming attention and sparse expert execution, preserving accuracy while minimizing overhead. Experimental results demonstrate that our accelerator achieves nearly 100 frames/s on M3ViT-tiny, a$3.13\times $improvement in throughput, and over 75% energy reduction compared to state-of-the-art (SOTA) FPGA MoE accelerators, while maintaining less than 1% accuracy loss across vision benchmarks. Our implementation will be open-sourced. Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Hermes: An FPGA-based NTT Accelerator Supporting Various Lengths for HHEabstractHybrid Homomorphic Encryption (HHE) scheme integrates two types of Fully Homomorphic Encryption (FHE), arithmetic FHE and logic FHE to enhance the performance and scalability of privacy-preserving computations. However, the performance of HHE mainly depends on the efficiency of the Number Theoretic Transform (NTT). Accordingly, this paper introduces Hermes, an FPGA-based NTT accelerator for HHE. We have designed a cross-scheme-friendly NTT architecture that supports NTT of varying lengths through the reuse of NTT units. Experimental results demonstrate that our proposed architecture achieves high hardware utilization and increases throughput by 1.3× compared to existing state-of-the-art approaches across various NTT lengths. Hang Gu, Qianyu Cheng, Jinao Li, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
CODES+ISSS | 5 |
| 2025 | Late Breaking Results: A Fast Nearest Neighbor Search Acceleration for 3D Point CloudabstractThis paper presents FastNN, a novel accelerator architecture for efficient K-Nearest Neighbors (KNN) search in point clouds. FastNN leverages a locality-sensitive E2LSH partitioning method and a precomparator module to significantly reduce the candidate search space and minimize the number of Euclidean distance calculations. Compared to octree-based partitioning methods, our approach reduces candidate points by 58.57% to 86.17% and achieves a $10.04 \times$ acceleration in processing throughput relative to the BitNN comparator subsystem. The proposed design effectively enhances search throughput, resource utilization, and precision, highlighting its potential for accelerating KNN search on FPGA platforms. Jinao Li, Qianyu Cheng, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
DAC | 4 |
| 2025 | CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
Jiale Dong, Wenqi Lou, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
Euro-Par (2) | 5 |
| 2025 | UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGAabstractCompared to traditional Vision Transformers (ViT), Mixture-of-Experts Vision Transformers (MoE-ViT) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, FPGA-based accelerators for MoE-ViT emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of fully exploring the design space, leading to suboptimal trade-offs between resource utilization and performance. To overcome this problem, we introduce UbiMoE, a novel end-to-end FPGA accelerator tailored for MoE-ViT. Leveraging the unique computational and memory access patterns of MoE-ViTs, we develop a latency-optimized streaming attention kernel and a resource-efficient reusable linear kernel, effectively balancing performance and resource consumption. To further enhance design efficiency, we propose a two-stage heuristic search algorithm that optimally tunes hardware parameters for various FPGA resource constraints. Compared to state-of-the-art (SOTA) FPGA designs, UbiMoE achieves 1.34× and 3.35× throughput improvements for MoE-ViT on Xilinx ZCU102 and Alveo U280 platforms, respectively, while enhancing energy efficiency by 1.75× and 1.54×. Our implementation is available at https://github.com/DJ000011/UbiMoE. Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
ISCAS | 3 |
| 2024 | Ph.D. Project: Achieving Low-Latency Acceleration on Multi-FPGA for GPT ApplicationabstractThis paper proposes a latency-oriented recurrent architecture for GPT on multi-FPGA with communication optimization. We devise an efficient communication scheme that overlaps part of the computation and communication delay to improve the latency and scalability of our platform. Then, we deploy recurrent structures on each FPGA to accelerate the different phases of GPT. A preliminary experiment shows that our method can reduce the synchronization overhead and increase the computing intensity, resulting in an average 11.8 x speedup over NVIDIA V100 GPU and 3.0x speedup over existing multi-FPGA accelerator appliance on the GPT-2 model. Zhendong Zheng, Chao Wang 0003 |
FCCM | 1 |
| 2024 | SoGraph: A State-Aware Architecture for Out-of-Memory Graph Processing on HBM-Equipped FPGAsabstractEmerging FPGA devices are widely used in graph processing and achieve high performance by exploiting fine-grained parallelism and high-bandwidth memory (HBM) sub-systems with dozens of channels. However, the graph size increases rapidly and often exceeds the capacity of the FPGA device or one of the memory channels, incurring out-of-memory (OoM) faults and performance degradation. Existing approaches focus on improving the performance of on-device full graph processing, ignoring the support for hyper-scale graphs and cross-iteration acceleration. In this work, we propose SoGraph, a graph processing architecture for large graph acceleration on HBM-equipped FPGAs, cooperating with the host. To minimize PCIe traffic during the processing, we adopt three hardware-oriented optimizations: 1) activeness-driven subgraph scheduler, which minimizes on-device graph resident set; 2) lightweight state-aware processing extension in edge-centric paradigm; 3) bit-level update analyzer for to-be-transfer subgraph reduction. The results show that the prototype on Alveo U280 FPGA achieves an average of 3.18x performance speedup over the modified state-of-the-art FPGA design and 1.3x energy efficiency over the GPU solution. Qianyu Cheng, Zhendong Zheng, Tianhao Jiang, Cheng Tang 0004, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
FPL | 2 |
| 2024 | LORA: A Latency-Oriented Recurrent Architecture for GPT Model on Multi-FPGA Platform with Communication OptimizationabstractLarge Language Models (LLMs) have been widely deployed in data centers to provide various services, among which the most representative is the Generative Pre-trained Transformer (GPT). The GPT model has heavy memory and computing overhead, and its inference process has two stages with distinct computing characteristics: Prefill and Decode. Utilizing existing GPUs and FPGA accelerators to construct a platform for deploying GPT in data centers faces the challenges of needing more effective synchronization schemes or structures with higher computational intensity. This paper proposes LORA, a low latency end-to-end GPT acceleration platform utilizing multiple FPGAs. Firstly, we optimize the synchronization timing of the GPT model to reduce the computation and communication overhead. Secondly, we devise some efficient synchronization steps for specific layers of the GPT model that overlap part of the computation and communication delay to improve the latency of our platform. Finally, we deploy recurrent structures on each FPGA to accelerate the different stages of the GPT model. Implemented on the Xilinx Alveo U280 FPGAs, LORA achieves an average $11.1 \times$ speedup over NVIDIA V100 GPUs on the modern GPT-2 model. Compared to the existing multi-FPGA accelerator appliance, LORA shows performance improvements of up to $4 \times$ and $2.7 \times$ in the Prefill and Decode stages. Zhendong Zheng, Qianyu Cheng, Lei Gong 0003, Xianglan Chen, Cheng Tang 0004, Chao Wang 0003, Xuehai Zhou |
FPL | 1 |
| 2023 | A graph-based framework to integrate semantic object/land-use relationships for urban land-use mapping with case studies of Chinese citiesabstractUrban land-use types, such as residential and administration, can be inferred through semantic objects and their relationships. Point of interest (POI) data can serve as the semantic objects for urban land-use mapping. However, the previous POI-based approaches have rarely considered the relationships between the semantic objects in the urban land-use mapping, and three main challenges remain: 1) the lack of paired semantic object/land-use samples; 2) the lack of a unified model for semantic objects and the relationships between sematic objects and urban land use; and 3) the difficulty of automatically learning semantic object/land-use mapping relationships. In this paper, to address these issues, a graph-based urban land-use mapping framework integrating semantic object/land-use relationships (GOLR) is proposed. Based on open-source area of interest (AOI) and POI data, an urban object/land-use (UOLU) dataset covering 34 cities in China was built. To model the spatial and mapping relationships, the semantic objects and their relationships are used to jointly build an urban land-use graph. The mapping from semantic objects to urban land use can then be learned by the urban land-use graph isomorphic network (ULGIN) model. Finally, the GOLR framework was applied to obtain accurate land-use mapping results for multiple Chinese cities. Yanfei Zhong, Yinhe Liu, Zhendong Zheng |
Int. J. Geogr. Inf. Sci. | 4 |
| 2022 | Work-in-Progress: BloCirNN: An Efficient Software/hardware Codesign Approach for Neural Network Accelerators with Block-Circulant MatrixabstractNowadays, the scale of deep neural networks is getting larger and larger. These large-scale deep neural networks are both compute and memory intensive. To overcome these problems, we use block-circulant weight matrices and Fast Fourier Transform (FFT) to compress model and optimize computation. Compared to weight pruning, this method does not suffer from irregular networks. The main contributions of this paper include the implementation of a convolution module and a fully-connected module with High-Level Synthesis (HLS), deployment and performance test on FPGA platform. We use AlexNet as a case study, which demonstrates our design is more efficient than the FPGA2016. Yunji Qin, Lei Gong 0003, Zhendong Zheng, Chao Wang 0003 |
CODES+ISSS | 3 |
| 2022 | Domain Adaptation via a Task-Specific Classifier Framework for Remote Sensing Cross-Scene ClassificationabstractThe scene classification of high spatial resolution (HSR) imagery involves labeling an HSR image with a specific high-level semantic class according to the composition of the semantic objects and their spatial relationships. As such, scene classification has attracted increased attention in recent years, and many different algorithms have now been proposed for the cross-scene classification task. However, the recently proposed scene classification methods based on deep convolutional neural networks (CNNs) still suffer from domain shift problems, because of the training data and validation data not following the assumption of independent and identical distributions. The employment of generative adversarial networks has been found to be an effective way to bridge the domain shift/gap. However, the existing cross-scene classification methods do not use the classification information in the target domain, and the domain classifier is task-independent for different scene classification tasks. In this article, to solve this problem, domain adaptation via a task-specific classifier (DATSNET) framework is proposed for HSR image scene classification. Task-specific classifiers and minimizing and maximizing, “ i.e., minimaxing,” of the classifier discrepancy are integrated in the DATSNET framework. The task-specific classifiers are proposed to align the distributions of the source domain features and target domain features by utilizing task-specific decision boundaries in the target domain. In order to align the two task-specific classifiers’ feature distributions, minimaxing the defined discrepancy between the different classifiers in an adversarial manner is proposed to obtain better task-specific classifier boundaries in the target domain and a better-aligned feature distribution in both domains. The experimental results obtained with different remote sensing cross-scene classification tasks demonstrate that the proposed method achieves a significantly improved performance compared with the other state-of-the-art remote sensing cross-scene classification algorithms. Zhendong Zheng, Yanfei Zhong, Ailong Ma |
IEEE Trans. Geosci. Remote. Sens. | 1 |