VLDB 2026 Research / reviewers in the wild / expert
Gaoming Du
dblp:99/8454
· DBLP profile ↗
18ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0002-6730-3167ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 9 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Falcon Signature Verification Accelerator Using Area-Efficient NTT Architecture With Simplified Barrett Modular MultiplierabstractFalcon is a lattice-based post-quantum digital-signature scheme standardized by U.S. National Institute of Standards and Technology (NIST). As for its hardware implementation, signature verification is the critical operation, in which the polynomial multiplier and the hash generator are computationally intensive. Existing hardware implementations face an inherent resource–performance tradeoff in the multiplier, and the overall verification latency remains high. In this article, we present an accelerator for Falcon signature verification using area-efficient number-theoretic transform (NTT) architecture with simplified Barrett modular multiplier. The precomputed constant of Barrett modular-multiplication unit is adjusted from 21 845 to 21 840. This modification maintains computation correctness while lowering resource consumption and improving throughput, which serves as an efficient building block for NTT/inverse NTT (INTT) operations. Moreover, a fully pipelined, nonstored radix-2 multipath delay commutator (R2-MDC) architecture is adopted for the NTT/INTT, which effectively reduces the delay and achieves a more balanced resource–performance tradeoff. Both the polynomial multiplier and the hash generator are parallelized to improve performance. A complete Falcon-1024 signature verification accelerator is implemented on FPGA platform. Experimental results show that, for the two-stage NTT inside the multiplier, the proposed design attains the lowest area–time product (ATP) compared to state-of-the-art designs, and the ATP of the entire Falcon-1024 verification procedure is reduced by 50%. Hongran Hu, Yangyi Chu, Gaoming Du, Duoli Zhang, Zhenmin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | An Efficient Accelerator for Dehazing Neural Network Based on Physical Perception Model and Cross-Scale Pixel AttentionabstractHaze reduces visibility, hindering real-time image processing applications. Although deep learning-based dehazing algorithms can significantly enhance image quality, their high computational complexity and storage demands make them challenging to deploy on resource-constrained hardware platforms. To address this, we propose an efficient dehazing neural network featuring physical perception model and cross-scale pixel attention (EP-CSANet), together with its hardware accelerator. In terms of algorithm, we design an efficient physical perception model (EPM) based on the atmospheric scattering model (ASM), combined with adaptive channel attention (ACA), which accurately approximates atmospheric light (A) and transmittance [t(x)], and effectively adapts to various environmental conditions, significantly improving dehazing accuracy. Furthermore, to enhance multiscale information interaction and preserve fine-grained texture details, we introduce cross-scale pixel attention (CSPA), which utilizes a dual-branch approach to extract high receptive-field information while maintaining fine-grained textures. In terms of hardware, we design a dedicated FPGA-based dehazing acceleration architecture. Through an efficient 16-stage pipeline, we achieve a dehazing rate of 127 frames per second (fps) for$640\times 480$images, meeting real-time processing requirements. In addition, by incorporating sparse convolution optimization techniques, we significantly improve resource utilization: LUTs increased by 31.8%, FFs by 27.9%, and DSPs by 22.3%. Experimental results demonstrate that EP-CSANet, using only 2361 parameters on the SOTS dataset, outperforms other dehazing algorithms. The source code is publicly available athttps://github.com/netflymachine/EP-CSANet Gaoming Du, Zhenmin Li, Duoli Zhang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2026 | ParaPM: Efficient Hardware Accelerator for Postquantum Signature With High-Performance Polynomial MultiplierabstractDuring the Institute of Standards and Technology (NIST) postquantum cryptography standardization process, the lattice-based Dilithium scheme was selected as one of the three third-round finalists for digital signature algorithms. Although numerous hardware implementations of Dilithium have been proposed, there remains substantial room for performance optimization, particularly in terms of computational speed. In this work, we target the two most time-consuming operations, namely the coefficient generation and polynomial multiplication. We propose a high-speed hardware architecture that fully exploits hardware parallelism. Our design introduces a fully pipelined radix-2 multipath delay commutator (R2MDC) structure supporting both NTT and inverse NTT (INTT) modes, a throughput-matched polynomial multiplier enabling on-the-fly pointwise multiplication (PWM) without buffering, and a low-resource Keccak. These optimizations collectively eliminate the throughput mismatch bottleneck and maximize hardware utilization through task-level parallelism. We implement coefficient generation, signing, and verification for three security levels on the A7 and Z7 platforms. Experimental results demonstrate that our current design achieves the optimal area-time product (ATP) across all three security levels. Gaoming Du, Yongsheng Yin, Duoli Zhang, Zhenmin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Image encryption/decryption accelerator based on Fast Cosine Number Transform
Zhenmin Li, Gaoming Du |
Integr. | 5 |
| 2024 | A Low-Latency Polynomial Multiplier Accelerator for CRYSTALS-Dilithium Digital SignatureabstractIn the post-quantum signature CRYSTALS-Dilithium algorithm, polynomial multiplication accounts for 32% of the total computation Latency. It is necessary to design a high-performance polynomial multiplication module. In this paper, we have proposed a low-Latency polynomial multiplier. First, we insert registers in the Radix-2 Multi-path Delay Commutator (R2MDC) structure to increase clock frequency. Additionally, to further enhance the computational clock frequency, we have designed a five-stage pipelined Butterfly unit. Secondly, we have proposed a fully pipelined polynomial multiplier that supports polynomial point-wise multiplication during NTT/INTT transformations to save a significant number of cycles. We also designed a configurable Polynomial Pointwise Multiplication (PPM) module that supports calculations with three different security levels. Our polynomial multiplier structure is implemented on the K7 model of FPGA. Compared to the currently fastest design in terms of computational speed, we have saved between 13% and 39% of the computation latency. Gaoming Du, Zhuo Chen 0037, Zhenmin Li, Duoli Zhang |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | A fast hardware accelerator for nighttime fog removal based on image fusion
Tianyi Lv, Gaoming Du, Zhenmin Li, Peiyi Teng |
Integr. | 2 |
| 2022 | A BNN Accelerator Based on Edge-skip-calculation Strategy and Consolidation Compressed TreeabstractBinarized neural networks (BNNs) and batch normalization (BN) have already become typical techniques in artificial intelligence today. Unfortunately, the massive accumulation and multiplication in BNN models bring challenges to field-programmable gate array (FPGA) implementations, because complex arithmetics in BN consume too much computing resources. To relax FPGA resource limitations and speed up the computing process, we propose a BNN accelerator architecture based on consolidation compressed tree scheme by combining both XNOR and accumulation operation of the low bit into a systematic one. During the compression process, we adopt 0-padding (not ±1) to achieve no-accuracy-loss from software modeling to hardware implementation. Moreover, we introduce shift-addition-BN free binarization technique to shorten the delay path and optimize on-chip storage. To sum up, we drastically cut down the hardware consumption while maintaining great speed performance with the same model complexity as the previous design. We evaluate our accelerator on MNIST and CIFAR-10 dataset and implement the whole system on the ARTIX-7 100T FPGA with speed performance of 2052.65 GOP/s and area efficiency of 70.15 GOPS/KLUT. Gaoming Du, Bangyi Chen, Zhenmin Li, Zhenxing Tu, Shenya Wang, Qinghao Zhao, Yongsheng Yin |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2021 | A Real-Time Effective Fusion-Based Image Defogging Architecture on FPGAabstractFoggy weather reduces the visibility of photographed objects, causing image distortion and decreasing overall image quality. Many approaches (e.g., image restoration, image enhancement, and fusion-based methods) have been proposed to work out the problem. However, most of these defogging algorithms are facing challenges such as algorithm complexity or real-time processing requirements. To simplify the defogging process, we propose a fusional defogging algorithm on the linear transmission of gray single-channel. This method combines gray single-channel linear transform with high-boost filtering according to different proportions. To enhance the visibility of the defogging image more effectively, we convert the RGB channel into a gray-scale single channel without decreasing the defogging results. After gray-scale fusion, the data in the gray-scale domain should be linearly transmitted. With the increasing real-time requirements for clear images, we also propose an efficient real-time FPGA defogging architecture. The architecture optimizes the data path of the guided filtering to speed up the defogging speed and save area and resources. Because the pixel reading order of mean and square value calculations are identical, the shift register in the box filter after the average and the computation of the square values is separated from the box filter and put on the input terminal for sharing, saving the storage area. What’s more, using LUTs instead of the multiplier can decrease the time delays of the square value calculation module and increase efficiency. Experimental results show that the linear transmission can save 66.7% of the total time. The architecture we proposed can defog efficiently and accurately, meeting the real-time defogging requirements on 1920 × 1080 image size. Gaoming Du, Jiting Wu, Hongfang Cao, Kun Xing, Zhenmin Li, Duoli Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Efficient Softmax Hardware Architecture for Deep Neural NetworksabstractDeep neural network (DNN) has become a pivotal machine learning and object recognition technology in the big data era. The softmax layer is one of the key component layers for completing multi-classification tasks. However, the softmax layer contains complex exponential and division operations, resulting in low accuracy and long critical paths in hardware accelerator design. In order to solve the above issues, we present a softmax hardware architecture with proper accuracy, good trade-off and strong expansibility. We summarize the classification rules of neural network and balance the calculation accuracy between resource consumption. On this basis, we proposed an exponential calculation unit based on the group lookup table, and improve a natural logarithmic calculation unit based on the Maclaurin series and the data preprocessing scheme matching them. The experimental results show that the softmax hardware architecture proposed in this paper can achieve the calculation accuracy of 3 decimal fraction and the classification accuracy of $99.01%$. Theoretically, it can accomplish the classification task of infinite categories. Gaoming Du, Zhenmin Li, Duoli Zhang, Yongsheng Yin |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | NR-MPA: Non-Recovery Compression Based Multi-Path Packet-Connected-Circuit Architecture of Convolution Neural Networks AcceleratorabstractConvolution Neural Networks (CNNs) involve massive data to be calculated and stored. To meet the challenges above, parallel hardware accelerators consisting of hundreds of Processing Elements (PEs) arranged as a many-core systemon-chip, connected by a Network-on-Chip (NoC) are proposed, which achieve high throughput exploiting parallel PE array. However, most of existing accelerators focus on only one aspect, such as compute structure of PE and data movement overhead above NoC, which causes the throughout, area and latency of the accelerator not fully optimized. In this paper, we propose an efficient general purpose CNN accelerator including both compute based on Non-Recovery Compression (NRC) method and data movement by novel Multi-Paths Packet Connection Circuit (MP-PCC). NRC can save computation time due to zero multiplier through shift decoding in PE and improve power efficiency by saving a large number of data transmission. MPPCC, evolved from Packet Connection Circuit, supports single and multicast transmission modes at the same time, and changes the multicast (X, Y) routing algorithm to multicast Y algorithm to improve the transmission efficiency. The proposed architecture which was implemented on Xilinx FPGA achieves 17.7x faster computation speed and 2.2x fewer memory accesses compared with the state-of-the-art method. Gaoming Du, Zhenwen Yang, Zhenmin Li, Duoli Zhang, Yongsheng Yin, Zhonghai Lu |
ICCD | 1 |
| 2019 | CPCA: An efficient wireless routing algorithm in WiNoC for cross path congestion awareness
Jianhua Li 0003, Chenglong Sun, Huaguo Liang, Gaoming Du |
Integr. | 6 |
| 2019 | SSS: Self-aware System-on-chip Using a Static-dynamic Hybrid MethodabstractNetwork-on-Chip (NoC) has become the de facto communication standard for multi-core or many-core System-on-Chip (SoC) due to its scalability and flexibility. However, an important factor in NoC design is temperature, which affects the overall performance of SoC—decreasing circuit frequency, increasing energy consumption, and even shortening chip lifetime. In this article, we propose SSS, a self-aware SoC using a static-dynamic hybrid method that combines dynamic mapping and static mapping to reduce the hotspot temperature for NoC-based SoCs. First, we propose monitoring and thermal modeling for self-state sensoring. Then, in static mapping stage, we calculate the optimal mapping solutions under different temperature modes using the discrete firefly algorithm to help self-decision making. Finally, in dynamic mapping stage, we achieve dynamic mapping through configuring NoC and SoC sentient units for self-optimizing. Experimental results show that SSS has substantially reduced the peak temperature by up to 37.52%. The FPGA prototype proves the effectiveness and smartness of SSS in reducing hotspot temperature. Gaoming Du, Guanyu Liu, Zhenmin Li, Duoli Zhang, Minglun Gao, Zhonghai Lu |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2017 | SSS: self-aware system-on-chip using static-dynamic hybrid method (work-in-progress)abstractNetwork on chip has become the de facto communication standard for multi-core or many-core system on chip, due to its scalability and flexibility. However, temperature is an important factor in NoC design, which affects the overall performance of SoC---decreasing circuit frequency, increasing energy consumption, and even shortening chip lifetime. In this paper, we propose SSS, a self-aware SoC using a static-dynamic hybrid method, which combines dynamic mapping and static mapping to reduce the hot-spots temperature for NoC based SoCs. First, we propose monitoring the thermal distribution for self-state sensoring. Then, in static mapping stage, we calculate the optimal mapping solutions under different temperature modes using discrete firefly algorithm to help self-decision making. Finally, in dynamic mapping stage, we achieve dynamic mapping through configuring NoC and SoC sentient unit for self-optimizing. Experimental results show SSS can reduce the peak temperature by up to 30.64%. FPGA prototype shows the effectiveness and smartness of SSS in reducing hot-spots temperature. Gaoming Du, Shibi Ma, Zhenmin Li, Zhonghai Lu, Minglun Gao |
CASES | 1 |
| 2017 | On the Accuracy of Stochastic Delay Bound for Network on ChipabstractDelay bound guarantee in network on chip (NoC) is important for hard real-time applications, and deterministic network calculus (DNC) is a effective tool for delay bound modeling. But for soft real-time applications, delay bound derivation using DNC is often over-pessimistic, resulting in too much chip area (e.g., router buffer) and power consumption; stochastic network calculus (SNC), on the contrary, improves the delay bound accuracy by providing stochastic service curves. Existing service models assume that contention takes place as long as there exist contention flows from different input channels requesting the same output channel. These models only consider flow paths in flows contention analyzing. We have observed that, beyond flow path contentions, the arrival rate also has deep influence on the flow contention in NoC, consequently affecting delay bound. In this paper, we further analyze the intrinsic factors affecting the flow contention, and propose a stochastic analytic model of per-flow delay bound to improve the calculation accuracy, according to both path and arrival rate. Within this model, the end-to-end delay bound is evaluated based on SNC. Experimental results show that our proposed model is both effective and accurate. Gaoming Du, Zhenmin Li, Guanyu Liu, Duoli Zhang |
NOCS | 1 |
| 2016 | OLITS: An Ohm's Law-like traffic splitting model based on congestion prediction
Gaoming Du, Yanghao Ou, Zhonghai Lu, Minglun Gao |
DATE | 1 |
| 2014 | An analytical model for worst-case reorder buffer size of multi-path minimal routing NoCsabstractReorder buffers are often needed in multi-path routing networks-on-chips (NoCs) to guarantee in-order packet delivery. However, the buffer sizes are usually over-dimensioned, due to lack of worst-case analysis, leading to unnecessary larger area overhead. Based on network calculus, we propose an analysis framework for the worst-case reorder buffer size in multi-path minimal routing NoCs. Experiments with synthetic traffic and an industry case show that our method can effectively explore the traffic splitting space, as well as the mapping effects in terms of reorder buffer size with a maximum improvement of 36.50%. Gaoming Du, Zhonghai Lu, Minglun Gao |
NOCS | 1 |
| 2010 | Application-level pipelining on Hierarchical NoCabstractMultiprocessor System-on-Chip is a promising solution for the high performance Embedded System. This paper is based on an independent research about Hierarchical NoC (Network-on-chip). By integrating 16 ARM cores in the FPGA board, we can bring out the four-channel fade-in and fade-out for real-time streaming media. We present two parallel models for our multiprocessor. One is fine-grained parallelization, with which the speed-up is 7.6, the other module is coarse-grained parallelization, with which the speed-up is higher than 9.2. Hongbing Pan, Li Li 0003, Minglun Gao, Ning Hou, Gaoming Du, Duoli Zhang |
ISCAS | 7 |
| 2007 | On the Implementation of Virtual Array Using Configuration Plane
Yongsheng Yin, Li Li 0003, Minglun Gao, Gaoming Du, Yu-Kun Song |
APPT | 4 |