EDBT 2026 Demo / reviewers in the wild / expert
Duoli Zhang
dblp:55/8455
· DBLP profile ↗
11ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-9515-2829ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Falcon Signature Verification Accelerator Using Area-Efficient NTT Architecture With Simplified Barrett Modular MultiplierabstractFalcon is a lattice-based post-quantum digital-signature scheme standardized by U.S. National Institute of Standards and Technology (NIST). As for its hardware implementation, signature verification is the critical operation, in which the polynomial multiplier and the hash generator are computationally intensive. Existing hardware implementations face an inherent resource–performance tradeoff in the multiplier, and the overall verification latency remains high. In this article, we present an accelerator for Falcon signature verification using area-efficient number-theoretic transform (NTT) architecture with simplified Barrett modular multiplier. The precomputed constant of Barrett modular-multiplication unit is adjusted from 21 845 to 21 840. This modification maintains computation correctness while lowering resource consumption and improving throughput, which serves as an efficient building block for NTT/inverse NTT (INTT) operations. Moreover, a fully pipelined, nonstored radix-2 multipath delay commutator (R2-MDC) architecture is adopted for the NTT/INTT, which effectively reduces the delay and achieves a more balanced resource–performance tradeoff. Both the polynomial multiplier and the hash generator are parallelized to improve performance. A complete Falcon-1024 signature verification accelerator is implemented on FPGA platform. Experimental results show that, for the two-stage NTT inside the multiplier, the proposed design attains the lowest area–time product (ATP) compared to state-of-the-art designs, and the ATP of the entire Falcon-1024 verification procedure is reduced by 50%. Hongran Hu, Yangyi Chu, Gaoming Du, Duoli Zhang, Zhenmin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2026 | An Efficient Accelerator for Dehazing Neural Network Based on Physical Perception Model and Cross-Scale Pixel AttentionabstractHaze reduces visibility, hindering real-time image processing applications. Although deep learning-based dehazing algorithms can significantly enhance image quality, their high computational complexity and storage demands make them challenging to deploy on resource-constrained hardware platforms. To address this, we propose an efficient dehazing neural network featuring physical perception model and cross-scale pixel attention (EP-CSANet), together with its hardware accelerator. In terms of algorithm, we design an efficient physical perception model (EPM) based on the atmospheric scattering model (ASM), combined with adaptive channel attention (ACA), which accurately approximates atmospheric light (A) and transmittance [t(x)], and effectively adapts to various environmental conditions, significantly improving dehazing accuracy. Furthermore, to enhance multiscale information interaction and preserve fine-grained texture details, we introduce cross-scale pixel attention (CSPA), which utilizes a dual-branch approach to extract high receptive-field information while maintaining fine-grained textures. In terms of hardware, we design a dedicated FPGA-based dehazing acceleration architecture. Through an efficient 16-stage pipeline, we achieve a dehazing rate of 127 frames per second (fps) for$640\times 480$images, meeting real-time processing requirements. In addition, by incorporating sparse convolution optimization techniques, we significantly improve resource utilization: LUTs increased by 31.8%, FFs by 27.9%, and DSPs by 22.3%. Experimental results demonstrate that EP-CSANet, using only 2361 parameters on the SOTS dataset, outperforms other dehazing algorithms. The source code is publicly available athttps://github.com/netflymachine/EP-CSANet Gaoming Du, Zhenmin Li, Duoli Zhang |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2026 | ParaPM: Efficient Hardware Accelerator for Postquantum Signature With High-Performance Polynomial MultiplierabstractDuring the Institute of Standards and Technology (NIST) postquantum cryptography standardization process, the lattice-based Dilithium scheme was selected as one of the three third-round finalists for digital signature algorithms. Although numerous hardware implementations of Dilithium have been proposed, there remains substantial room for performance optimization, particularly in terms of computational speed. In this work, we target the two most time-consuming operations, namely the coefficient generation and polynomial multiplication. We propose a high-speed hardware architecture that fully exploits hardware parallelism. Our design introduces a fully pipelined radix-2 multipath delay commutator (R2MDC) structure supporting both NTT and inverse NTT (INTT) modes, a throughput-matched polynomial multiplier enabling on-the-fly pointwise multiplication (PWM) without buffering, and a low-resource Keccak. These optimizations collectively eliminate the throughput mismatch bottleneck and maximize hardware utilization through task-level parallelism. We implement coefficient generation, signing, and verification for three security levels on the A7 and Z7 platforms. Experimental results demonstrate that our current design achieves the optimal area-time product (ATP) across all three security levels. Gaoming Du, Yongsheng Yin, Duoli Zhang, Zhenmin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | A Low-Latency Polynomial Multiplier Accelerator for CRYSTALS-Dilithium Digital SignatureabstractIn the post-quantum signature CRYSTALS-Dilithium algorithm, polynomial multiplication accounts for 32% of the total computation Latency. It is necessary to design a high-performance polynomial multiplication module. In this paper, we have proposed a low-Latency polynomial multiplier. First, we insert registers in the Radix-2 Multi-path Delay Commutator (R2MDC) structure to increase clock frequency. Additionally, to further enhance the computational clock frequency, we have designed a five-stage pipelined Butterfly unit. Secondly, we have proposed a fully pipelined polynomial multiplier that supports polynomial point-wise multiplication during NTT/INTT transformations to save a significant number of cycles. We also designed a configurable Polynomial Pointwise Multiplication (PPM) module that supports calculations with three different security levels. Our polynomial multiplier structure is implemented on the K7 model of FPGA. Compared to the currently fastest design in terms of computational speed, we have saved between 13% and 39% of the computation latency. Gaoming Du, Zhuo Chen 0037, Zhenmin Li, Duoli Zhang |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | A Memory-Constraint-Aware List Scheduling Algorithm for Memory-Constraint Heterogeneous Muti-Processor SystemabstractAn effective scheduling algorithm is vital for the execution efficiency of applications on Heterogeneous Muti-Processor System (HMPS), especially Memory-Constraint Heterogeneous Muti-Processor System (MCHMPS). Stringent local and external memory constraints have significant impact on the execution performance of applications executed on MCHMPS, predictability is also a critical factor for task scheduling on MCHMPS. Therefore, a novel list scheduling algorithm termed Memory-constraint-aware Improved Predict Priority and Optimistic Processor Selection Scheduling (MIPPOSS), essentially a heuristic search optimization algorithm, is proposed in this paper. In MIPPOSS, a predictive approach is applied for task prioritization and processor selection, and a novel memory-constraint-aware approach is employed in the processor selection phase. MIPPOSS has polynomial complexity and produces better results for application scheduling on target architecture. Randomly generated DAGs and 3 real-world applications experiments, including Cybershake, LIGO, and Montage, show that MIPPOSS outperforms the other five competing algorithms by a large margin. Duoli Zhang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | A Real-Time Effective Fusion-Based Image Defogging Architecture on FPGAabstractFoggy weather reduces the visibility of photographed objects, causing image distortion and decreasing overall image quality. Many approaches (e.g., image restoration, image enhancement, and fusion-based methods) have been proposed to work out the problem. However, most of these defogging algorithms are facing challenges such as algorithm complexity or real-time processing requirements. To simplify the defogging process, we propose a fusional defogging algorithm on the linear transmission of gray single-channel. This method combines gray single-channel linear transform with high-boost filtering according to different proportions. To enhance the visibility of the defogging image more effectively, we convert the RGB channel into a gray-scale single channel without decreasing the defogging results. After gray-scale fusion, the data in the gray-scale domain should be linearly transmitted. With the increasing real-time requirements for clear images, we also propose an efficient real-time FPGA defogging architecture. The architecture optimizes the data path of the guided filtering to speed up the defogging speed and save area and resources. Because the pixel reading order of mean and square value calculations are identical, the shift register in the box filter after the average and the computation of the square values is separated from the box filter and put on the input terminal for sharing, saving the storage area. What’s more, using LUTs instead of the multiplier can decrease the time delays of the square value calculation module and increase efficiency. Experimental results show that the linear transmission can save 66.7% of the total time. The architecture we proposed can defog efficiently and accurately, meeting the real-time defogging requirements on 1920 × 1080 image size. Gaoming Du, Jiting Wu, Hongfang Cao, Kun Xing, Zhenmin Li, Duoli Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2019 | Efficient Softmax Hardware Architecture for Deep Neural NetworksabstractDeep neural network (DNN) has become a pivotal machine learning and object recognition technology in the big data era. The softmax layer is one of the key component layers for completing multi-classification tasks. However, the softmax layer contains complex exponential and division operations, resulting in low accuracy and long critical paths in hardware accelerator design. In order to solve the above issues, we present a softmax hardware architecture with proper accuracy, good trade-off and strong expansibility. We summarize the classification rules of neural network and balance the calculation accuracy between resource consumption. On this basis, we proposed an exponential calculation unit based on the group lookup table, and improve a natural logarithmic calculation unit based on the Maclaurin series and the data preprocessing scheme matching them. The experimental results show that the softmax hardware architecture proposed in this paper can achieve the calculation accuracy of 3 decimal fraction and the classification accuracy of $99.01%$. Theoretically, it can accomplish the classification task of infinite categories. Gaoming Du, Zhenmin Li, Duoli Zhang, Yongsheng Yin |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | NR-MPA: Non-Recovery Compression Based Multi-Path Packet-Connected-Circuit Architecture of Convolution Neural Networks AcceleratorabstractConvolution Neural Networks (CNNs) involve massive data to be calculated and stored. To meet the challenges above, parallel hardware accelerators consisting of hundreds of Processing Elements (PEs) arranged as a many-core systemon-chip, connected by a Network-on-Chip (NoC) are proposed, which achieve high throughput exploiting parallel PE array. However, most of existing accelerators focus on only one aspect, such as compute structure of PE and data movement overhead above NoC, which causes the throughout, area and latency of the accelerator not fully optimized. In this paper, we propose an efficient general purpose CNN accelerator including both compute based on Non-Recovery Compression (NRC) method and data movement by novel Multi-Paths Packet Connection Circuit (MP-PCC). NRC can save computation time due to zero multiplier through shift decoding in PE and improve power efficiency by saving a large number of data transmission. MPPCC, evolved from Packet Connection Circuit, supports single and multicast transmission modes at the same time, and changes the multicast (X, Y) routing algorithm to multicast Y algorithm to improve the transmission efficiency. The proposed architecture which was implemented on Xilinx FPGA achieves 17.7x faster computation speed and 2.2x fewer memory accesses compared with the state-of-the-art method. Gaoming Du, Zhenwen Yang, Zhenmin Li, Duoli Zhang, Yongsheng Yin, Zhonghai Lu |
ICCD | 4 |
| 2019 | SSS: Self-aware System-on-chip Using a Static-dynamic Hybrid MethodabstractNetwork-on-Chip (NoC) has become the de facto communication standard for multi-core or many-core System-on-Chip (SoC) due to its scalability and flexibility. However, an important factor in NoC design is temperature, which affects the overall performance of SoC—decreasing circuit frequency, increasing energy consumption, and even shortening chip lifetime. In this article, we propose SSS, a self-aware SoC using a static-dynamic hybrid method that combines dynamic mapping and static mapping to reduce the hotspot temperature for NoC-based SoCs. First, we propose monitoring and thermal modeling for self-state sensoring. Then, in static mapping stage, we calculate the optimal mapping solutions under different temperature modes using the discrete firefly algorithm to help self-decision making. Finally, in dynamic mapping stage, we achieve dynamic mapping through configuring NoC and SoC sentient units for self-optimizing. Experimental results show that SSS has substantially reduced the peak temperature by up to 37.52%. The FPGA prototype proves the effectiveness and smartness of SSS in reducing hotspot temperature. Gaoming Du, Guanyu Liu, Zhenmin Li, Duoli Zhang, Minglun Gao, Zhonghai Lu |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2017 | On the Accuracy of Stochastic Delay Bound for Network on ChipabstractDelay bound guarantee in network on chip (NoC) is important for hard real-time applications, and deterministic network calculus (DNC) is a effective tool for delay bound modeling. But for soft real-time applications, delay bound derivation using DNC is often over-pessimistic, resulting in too much chip area (e.g., router buffer) and power consumption; stochastic network calculus (SNC), on the contrary, improves the delay bound accuracy by providing stochastic service curves. Existing service models assume that contention takes place as long as there exist contention flows from different input channels requesting the same output channel. These models only consider flow paths in flows contention analyzing. We have observed that, beyond flow path contentions, the arrival rate also has deep influence on the flow contention in NoC, consequently affecting delay bound. In this paper, we further analyze the intrinsic factors affecting the flow contention, and propose a stochastic analytic model of per-flow delay bound to improve the calculation accuracy, according to both path and arrival rate. Within this model, the end-to-end delay bound is evaluated based on SNC. Experimental results show that our proposed model is both effective and accurate. Gaoming Du, Zhenmin Li, Guanyu Liu, Duoli Zhang |
NOCS | 5 |
| 2010 | Application-level pipelining on Hierarchical NoCabstractMultiprocessor System-on-Chip is a promising solution for the high performance Embedded System. This paper is based on an independent research about Hierarchical NoC (Network-on-chip). By integrating 16 ARM cores in the FPGA board, we can bring out the four-channel fade-in and fade-out for real-time streaming media. We present two parallel models for our multiprocessor. One is fine-grained parallelization, with which the speed-up is 7.6, the other module is coarse-grained parallelization, with which the speed-up is higher than 9.2. Hongbing Pan, Li Li 0003, Minglun Gao, Ning Hou, Gaoming Du, Duoli Zhang |
ISCAS | 8 |