Wu Gao

dblp:91/8935 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0007-0968ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A 181.9-dB FoMSNDR Power/Bandwidth Scalable Fully Dynamic Discrete-Time Zoom ADC Using Gain Boost Floating Inverter Amplifiers
Zhengyu Ren, Tianlong Gao, Wu Gao
ISCAS5
2026 Accuracy evaluation of classification models on partially labeled datasets
Huanjie Tao, Wu Gao
Expert Syst. Appl.2
2026 SpikeGate-YOLO: Spiking Object Detection With Dynamic Gating and Multigranularity Fusion
abstract
Spiking Neural Networks (SNNs) transmit information via discrete spike events, offering advantages in energy efficiency and computational cost. However, current SNN-based object detectors suffer from limited feature expression and inefficient fusion due to temporal sparsity and the asynchronous nature of spike features. To address these challenges, we propose SpikeGate-YOLO, a spiking object detection architecture optimized for spike-driven processing. Specifically, we introduce the Reparam-Spike Gating (RSG) block to enhance feature expressiveness while maintaining computational efficiency. We also design the Spike Multi-Granularity Difference-aware Feature Harmonizer (SpikeMDFH), which improves multi-scale feature fusion through dynamic attention and biologically inspired gating mechanisms, preserving spike sparsity. Experiments on both the COCO and Gen1 datasets show that SpikeGate-YOLO achieves state-of-the-art results, reaching 63.4% mAP@50 and 46.3% mAP@50:95 on COCO, and 69.3% mAP@50 and 42.9% mAP@50:95 on Gen1. These results confirm the effectiveness of our architecture in overcoming spike-specific limitations in feature representation and fusion for object detection.
Qiang Niu, Chen Zhao 0009, Shizhou Zhang, Wu Gao
IEEE Internet Things J.5
2026 An Energy-Efficient In-Sensor Computing Architecture With In-Pixel Spiking Neuron and Event-Driven Computation for Low-Energy X-Ray Imaging: Modeling and Evaluation
abstract
The use of neural networks for processing X-ray detector data has become a key development trend in high-energy physics and medical imaging. As detector arrays grow, handling the resulting massive data volumes poses significant challenges in terms of system bandwidth and power consumption. To address the bandwidth bottleneck and high analog-to-digital conversion (ADC) power consumption in conventional X-ray imaging systems, we propose an in-sensor computing architecture, SpikeX, for accelerating spiking neural networks (SNN). First, a photon-counting analog front end is directly integrated with pixel-level SNN neurons, enabling the first convolutional layer to be computed using spike signals. This eliminates the need for analog-to-digital conversion, significantly reducing power and bandwidth requirements. Next, an event-driven on-chip SNN accelerator performs efficient temporal inference, enabling an end-to-end pipeline from signal acquisition to image classification. Furthermore, detector non-idealities can be absorbed by the SNN model during training. To systematically investigate these effects, we develop a signal processing algorithm, enabling effective correction of detector non-idealities. To validate the proposed architecture, we implemented the SpikeX using a 180 nm CMOS process, featuring a$28\times 28$pixel array and 32 Leaky Integrate-and-Fire (LIF) processing units. Simulation results show that, at a clock frequency of 50 MHz, the architecture achieves an energy efficiency of 1.88 TOPS/W, significantly outperforming comparable state-of-the-art designs.
De Xu, Zhaoqi Miao, Guanhong Zheng, Musheer Abdullah, Shengbo Lin, Wu Gao
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 Real-Time Road Damage Detection Using an Optimized YOLOv9s-Fusion in IoT Infrastructure
abstract
In IoT-enabled smart infrastructure, accurate and real-time road damage detection is crucial for enhancing road safety and optimizing maintenance processes. However, detecting road damage in complex and dynamic environments presents significant challenges, such as varying lighting conditions, diverse damage types, and the need for fast processing to enable real-time decision-making. This study introduces an advanced approach utilizing the YOLOv9s-Fusion model to overcome these challenges. Leveraging the RDD2022 dataset, which comprises 1976 annotated images of road damage from China, we employ comprehensive data preprocessing to create optimal conditions for model training. The YOLOv9s-Fusion model integrates innovative features, including a Transformer-based auxiliary module and enhanced feature extraction layers, specifically designed to detect fine-grained damage patterns accurately. Experimental results demonstrate that the model outperforms existing approaches, achieving notable improvements in mean average precision (mAP) and F1-score. Ablation studies further validate the impact of our modifications, highlighting the model’s robustness in real-time detection across diverse conditions. This IoT-centric approach sets a new standard for autonomous road damage detection, significantly advancing vehicle navigation and smart infrastructure management capabilities.
Khan Muhammad 0001, Mohammad S. Obaidat, Khalid Mahmood 0002, Balqies Sadoun, Hafiz Muhammad Sanaullah Badar, Wu Gao
IEEE Internet Things J.6
2024 Real-Time Road Damage Detection and Infrastructure Evaluation Leveraging Unmanned Aerial Vehicles and Tiny Machine Learning
abstract
Road damage detection (RDD) through computer vision and deep learning techniques can ensure the safety of vehicles and humans on the roads. Integrating unmanned aerial vehicles (UAVs) in RDD and infrastructure evaluation (IE) has also emerged as a key enabler, contributing significantly to data acquisition and real-time monitoring of road damages such as potholes, cracks, and surface anomalies, facilitating proactive maintenance and improved road conditions. These UAVs are low-powered and resource-constrained devices that work autonomously to perform pattern detection and decision-making leveraging tiny machine learning (Tiny ML) algorithms. These Tiny ML algorithms are designed to run on edge devices, IoT devices, UAVs, etc. In this study, the RDD2022 dataset collected using UAVs and dashboard cameras of vehicles was utilized to train pure and mixed models that exhibit class instance imbalance in certain classes which is addressed by implementing data augmentation as a regularization technique. State-of-the-art two-stage detectors; Faster R-CNN ResNet101 and one-stage detectors; SSD MobileNet V1 FPN, YOLOv5, and Efficientdet D1 are employed. The results indicate that the two-stage detector achieved an impressive mAP of 88.49% overall and 96.62% for focused classes. Notably, the state-of-the-art Efficientdet D1 approach achieved a competitive mAP of 86.47% overall and 95.12% for focused classes, with significantly lower computational cost. These findings highlight the potential of advanced object detection techniques, particularly Efficientdet D1, to enhance the accuracy and efficiency of RDD systems, thereby improving passenger safety and overall performance.
Khan Muhammad 0001, Mohammad S. Obaidat, Khalid Mahmood 0002, Dania Batool, Hafiz Muhammad Sanaullah Badar, Muhammad Aamir 0002, Wu Gao
IEEE Internet Things J.7
2022 A Survey of GPU Multitasking Methods Supported by Hardware Architecture
abstract
The ability to support multitasking becomes more and more important in the development of graphic processing unit (GPU). GPU multitasking methods are classified into three types: temporal multitasking, spatial multitasking, and simultaneous multitasking (SMK). This article first introduces the features of some commercial GPU architectures to support multitasking and the common metrics used for evaluating the performance of GPU multitasking methods, and then reviews the GPU multitasking methods supported by hardware architecture (i.e., hardware GPU multitasking methods). The main problems of each type of hardware GPU multitasking methods to be solved are illustrated. Meanwhile, the key idea of each previous hardware GPU multitasking method is introduced. In addition, the characteristics of hardware GPU multitasking methods belonging to the same type are compared. This article also gives some valuable suggestions for the future research. An enhanced GPU simulator is needed to bridge the gap between academia and industry. In addition, it is promising to expand the research space with machine learning technologies, advanced GPU architectural innovations, 3D stacked memory, etc. Because most previous GPU multitasking methods are based on NVIDIA GPUs, this article focuses on NVIDIA GPU architecture, and uses NVIDIA's terminology. To our knowledge, this article is the first survey about hardware GPU multitasking methods. We believe that our survey can help the readers gain insights into the research field of hardware GPU multitasking methods.
Chen Zhao 0009, Wu Gao, Feiping Nie 0001, Huiyang Zhou
IEEE Trans. Parallel Distributed Syst.2
2021 A Resource-Efficient Parallel Connected Component Labeling Algorithm and Its Hardware Implementation
abstract
Connected Component labeling (CCL) is usually time-consuming, so a dedicated hardware accelerator of CCL is essential in the embedded vision and multimedia system. In this paper, we propose a single-scan resource-efficient parallel CCL algorithm. Our CCL method scans two adjacent rows simultaneously to extract runs and detect equivalent runs; an equivalent label set is used to resolve equivalences. After each row scan, the finished objects are output in time, and the freed memory resources are reused to reduce memory requirements. Both pixel-based labeled image (PLI) and run-based labeled image (RLI) can be generated by our CCL method. In addition, the steps of our CCL method are executed concurrently to improve labeling performance. The hardware architecture based on our CCL method is implemented with Verilog. The evaluation results illustrate that our CCL architecture can label more than 40 2048 × 1536 benchmark images per second on average, and outperforms previous CCL architectures in terms of labeling performance or memory resource consumption.
Chen Zhao 0009, Wu Gao, Feiping Nie 0001
IEEE Trans. Multim.2
2020 Fair and cache blocking aware warp scheduling for concurrent kernel execution on GPU
Chen Zhao 0009, Wu Gao, Feiping Nie 0001, Fei Wang 0008, Huiyang Zhou
Future Gener. Comput. Syst.2
2020 A Memory-Efficient Hardware Architecture for Connected Component Labeling in Embedded System
abstract
In this paper, we introduce a hardware architecture to accelerate connected component labeling (CCL) for embedded systems. The proposed CCL architecture scans the given binary image only once, and during the raster scan, the input binary image is compressed with run-length encoding to extract runs. The equivalences between runs are resolved efficiently by merging equivalent label lists, and all the intermediate data generated while labeling the images are stored in on-chip memory to avoid frequent access to off-chip memory. The finished connected components are determined and then output directly to free onchip memory resources early. These freed memory resources can be reused, which saves memory. Our CCL architecture is implemented with Verilog, and a quantitative comparison of memory cost shows that the proposed CCL architecture is memory-efficient and requires significantly fewer memory resources compared to other methods. In addition, our CCL architecture can process more than 25 2048 × 1536 benchmark images per second when it works at 300 MHz.
Chen Zhao 0009, Wu Gao, Feiping Nie 0001
IEEE Trans. Circuits Syst. Video Technol.2