EDBT 2026 Demo / reviewers in the wild / expert
Zhihan Zhang 0004
dblp:245/8608-4
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0000-8864-9143ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Single-Step Hardware-Aware Neural Network Quantization With Mixed PrecisionabstractQuantization is a neural network compression technique that effectively improves the deployment performance on inference hardware. Fixed-point quantization methods use the same bit-width for all layers in the network, which leads to difficulties in balancing compression rate and accuracy loss. Therefore, mixed-precision quantization has recently been proposed. The major challenge of the mixed-precision quantization is to select the quantization bit-widths of each layer in network to simultaneously meet the requirements of minimizing accuracy loss and hardware resource consumption. In this paper, we present a Single-Step Hardware-Aware Quantization (SHQ) method. It can calculate the resource consumption of hardware such as Field Programmable Gate Arrays (FPGAs) before actual deployment and find effective quantization schemes by just single step, different from common software-hardware two-step approaches with large workload and time-consuming. Genetic algorithm is combined with SHQ to search quantization schemes with low hardware resource usage and high accuracy. In addition, the correlation between hardware resource cost and qualitative indicators of proxy signal has been analyzed to bring computer-aided design insights of neural network accelerators. Experiments of full-pipeline accelerator deployment on FPGAs platforms show that our approach saves 39% of Digital Signal Processors (DSP) and 59% of Block Random Access Memory (BRAM) usage on MobileNet compared to full 8-bit quantization, while the accuracy drops by only 0.56%. The source code about our method can be found at this link:https://github.com/hujie369/SHQ. Jie Hu 0044, Zhihan Zhang 0004, Qunkang Meng, Qijun Huang, Hao Wang 0046, Sheng Chang 0003 |
IEEE Trans. Computers | 2 |
| 2026 | MetaAccel: A High-Performance and Agile Accelerator Design Framework With Multi Clock Domain Optimization for Complex CNNabstractEdge computing for artificial intelligence (AI) has become a new focus today. At the edge, the growing complexity and diversity of AI models has made FPGA, which has shorter development cycles, a good choice. Traditional AI accelerators on FPGA are mainly based on the Compute Engine (CE) architecture, suffering from low resource utilization and suboptimal speed. In contrast, the pipeline architecture achieves higher performance through its algorithm-structure-aware feature and fully on-chip data flow. However, customized designs and large bandwidth demands bring new challenges to its development agility and memory utilization, while high-performance acceleration for complex neural networks is still hard. In this article, we proposed MetaAccel, a novel fully pipelined accelerator. It uses two clock domains to manage data scheduling and calculation, significantly improving computing resource efficiency and on-chip memory utilization. Besides that, we built a hyperparameter-driven resource estimation model that can match the most appropriate design solutions for specific network structures. Based on this architecture, processing method for networks with complex branch structures and various operations is given, which makes MetaAccel suitable for Convolutional Neural Networks (CNNs) in different fields, such as image classification, object detection, and image segmentation. For typical networks, MetaAccel can achieve a throughput of more than 0.7TOPS and a DSP efficiency of up to 2.0GOPS/DSP and outperforms previous FPGA work in other metrics, showing its advantages in complex CNNs’ acceleration. Yuxian Jiang, Zhihan Zhang 0004, Qunkang Meng, Hao Wang 0046, Qijun Huang, Sheng Chang 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | A High-Intensity Solution of Hardware Accelerator for Sparse and Redundant Computations in Semantic Segmentation ModelsabstractThe rapid development of artificial intelligence (AI) has met people’s personalized needs. However, with the increase of data capacities and computing requirements, the imbalance between large-scale data transmission and limited network bandwidth has become increasingly prominent. To improve the speed of embedded system, real-time intelligent computing is gradually moving from the cloud to the edge. Traditional FPGA-based AI accelerators mainly utilize PE architecture, but the low computing throughput and resource utilization make it difficult to meet the power requirement of edge AI application scenarios such as image segmentation. In recent years, AI accelerators based on streaming architecture have become a trend, and it is necessary to customize high-performance streaming accelerators for specific segmentation algorithms. In this paper, we design a high-intensity pixel-level fully pipelined accelerator with customized strategies to eliminate the sparse and redundant computations in specific algorithms of semantic segmentation, which significantly improve the accelerator’s computing throughput and hardware resources utilization. On Xilinx FPGA, our acceleration of two typical semantic segmentation networks-ESPNet and DeepLabV3, achieves optimized throughputs of 171.3 GOPS and 1324.8 GOPS, and computing efficiency of 9.26 and 9.01, respectively. It provides the possibility of hardware deployment in real-time application with high computing intensity. Yuxian Jiang, Zhihan Zhang 0004, Hao Wang 0046, Sheng Chang 0003 |
IEEE Trans. Computers | 4 |
| 2025 | PEDSA: High-Throughput Pipeline-Based FPGA Accelerator for Convolutional Encoder-Decoder Segmentation NetworksabstractIn the era of artificial intelligence (AI), rapidly growing data and computing demands stimulate a shift toward more intelligent processing at the edge in the Internet of Things (IoT). AI application scenarios, such as image segmentation, pose new challenges to the computing capability of edge hardware, which cannot be solved by traditional AI accelerators with the traditional processing element (PE) architecture. Recently, the streaming architecture has received more attention due to its higher performance. To improve the throughput of edge platforms for segmentation tasks, customizing streaming accelerators for segmentation models is now necessary. Based on these motivations, we proposed pipelined encoder-decoder segmentation model accelerator (PEDSA), a fully pipelined streaming accelerator for convolutional encoder-decoder segmentation networks. PEDSA maps all the layers in the network into a pixel-level pipeline. All operations, especially upsampling (unpooling, deconvolution, etc.) which involves complex data rearrangement, are integrated into a regular, concise, and fully on-chip data flow. On Xilinx field-programmable gate arrays (FPGAs), our acceleration of SegNet-Basic and U-Net reached performances of 2676.47 and 7646.29 GOPS, respectively, outperforming previous accelerators for this kind of network. This work provides new ideas for the deployment of segmentation algorithms at the edge. Yuxian Jiang, Zhihan Zhang 0004, Hao Wang 0046, Sheng Chang 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | A High-Performance Pixel-Level Fully Pipelined Hardware Accelerator for Neural NetworksabstractThe design of convolutional neural network (CNN) hardware accelerators based on a single computing engine (CE) architecture or multi-CE architecture has received widespread attention in recent years. Although this kind of hardware accelerator has advantages in hardware platform deployment flexibility and development cycle, it is still limited in resource utilization and data throughput. When processing large feature maps, the speed can usually only reach 10 frames/s, which does not meet the requirements of application scenarios, such as autonomous driving and radar detection. To solve the above problems, this article proposes a full pipeline hardware accelerator design based on pixel. By pixel-by-pixel strategy, the concept of the layer is downplayed, and the generation method of each pixel of the output feature map (Ofmap) can be optimized. To pipeline the entire computing system, we expand each layer of the neural network into hardware, eliminating the buffers between layers and maximizing the effect of complete connectivity across the entire network. This approach has yielded excellent performance. Besides that, as the pixel data stream is a fundamental paradigm in image processing, our fully pipelined hardware accelerator is universal for various CNNs (MobileNetV1, MobileNetV2 and FashionNet) in computer vision. As an example, the accelerator for MobileNetV1 achieves a speed of 4205.50 frames/s and a throughput of 4787.15 GOP/s at 211 MHz, with an output latency of 0.60 ms per image. This extremely shorts processing time and opens the door for AI's application in high-speed scenarios. Zhihan Zhang 0004, Jie Hu 0044, Qunkang Meng, Hao Wang 0046, Qijun Huang, Sheng Chang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |