Bocai Wu

dblp:301/0966 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Uncertainty-informed prototype contrastive learning for cross-scene hyperspectral image classification
Kai Xu 0008, Zhuou Zhu, Bocai Wu, Chengcheng Fan
Knowl. Based Syst.3
2025 Enhancing Remote Sensing Scene Classification With Hy-MSDA: A Hybrid CNN-Transformer for Multisource Domain Adaptation
abstract
Multisource unsupervised domain adaptation (MUDA) has demonstrated its effectiveness in improving model performance for remote sensing (RS) scene classification, particularly in cases where the target-domain lacks labeled data. However, most current methods based on convolutional neural networks (CNNs) or Transformers fail to fully exploit the features within each source domain, which benefits classification accuracy. Additionally, these methods frequently overlook the varying levels of similarity between domains, limiting the potential of MUDA. To alleviate this limitation, we propose Hy-MSDA, a hybrid CNN-Transformer with consistent learning and dynamic weighting for MUDA, which fully explores and utilizes valuable information from multiple sources at both the feature and decision levels. To achieve a better alignment of categories among domains, the consistency learning module aims to learn domain-invariant features, maintaining consistency in global high-level features and rich semantic information at the feature level. At the decision level, the dynamic weighting strategy balances the contribution of each source domain by considering their varying importance levels. These two components are closely interconnected and mutually reinforce each other, resulting in improved cross-domain collaboration of Hy-MSDA for domain adaptation. Experimental results on the RS classification datasets clearly demonstrate that Hy-MSDA outperforms the state-of-the-art methods, achieving a significant enhancement of 2%–7% in classification accuracy. Moreover, experiments on RS segmentation datasets reveal that Hy-MSDA performs comparably to supervised methods, further validating the effectiveness of Hy-MSDA in segmentation tasks. The source code is available athttps://github.com/phaeton2017/Hy-MSDA.
Kai Xu 0008, Zhou Zhu, Chengcheng Fan, Bocai Wu
IEEE Trans. Geosci. Remote. Sens.5
2024 TransGA-Net: Integration Transformer With Gradient-Aware Feature Aggregation for Accurate Cloud Detection in Remote Sensing Imagery
abstract
Significant progress has been achieved in cloud detection using deep convolutional neural networks, leading to a substantial increase in detection accuracy. However, discriminating between actual clouds and cloud-like objects, as well as accurately identifying thin clouds and cloud boundaries, have been persistent challenges in the cloud detection. To address these issues, an Integration Transformer with gradient-aware feature aggregation (TransGA-Net) is proposed in this paper. This framework employs Transformers as encoders, significantly enhancing the modeling capabilities for global features and long-term dependencies. Additionally, a novel Gradient-Aware Feature Aggregation Module is proposed, focused on gradient variations within cloud boundary regions and utilizing gradient information to guide feature aggregation. The performance of TransGA-Net is validated on Landsat-8, GF-2, and Sentinel-2 datasets. In general, experiments show TransGA-Net surpasses the highest performance achieved by the other methods by 0.6%, 1.3%, and 0.7% in OA on the three datasets respectively, demonstrating that proposed method are particularly effective in identifying areas with thin cloud cover and distinguishing between clouds and cloud-like objects. The source code is available at https://github.com/phaeton2017/transga-net.
Kai Xu 0008, Xiaoyuan Deng, Anling Wang, Bocai Wu
IEEE Geosci. Remote. Sens. Lett.5
2024 SARGap: A Full-Link General Decoupling Automatic Pruning Algorithm for Deep Learning-Based SAR Target Detectors
abstract
Synthetic aperture radar (SAR) target detectors based on deep learning have difficulty finding a good balance between accuracy and speed. Current pruning methods are usually used for backbone consistent pruning and seldom directly for the whole structure of deep learning target detectors; therefore, for edge-end applications, this article proposes a new full-link general automatic pruning algorithm for SAR target detectors, referred to as SARGap. First, SARGap automatically analyzes the network structure by creating a dependency graph, divides the pair-coupled network structure into the same group, and prunes the same channel for the same group of network structures so that the algorithm can be applied to a variety of complex target detectors. Second, an automatic pruning rate search method (APRS) is designed to search for the optimal pruning rate of each group of network structures in the target detector. Finally, to find a good balance between precision and speed in the automatic search of the pruning rate, a multiobjective optimization loss function (MOOL) is constructed as the APRS objective function. A series of experiments based on SSDD and HRSID, two large-scale SAR target detection datasets, are carried out to prove the superiority of this method. Using Yolov5s as the baseline, SARGap can compress parameters by 84.29%/82.86% and flops by 80.50%/81.93% on two datasets with almost no loss of accuracy. In addition, SARGap can be applied to any deep learning target detector and match hardware computing resources to achieve optimal full-link pruning.
Jingqian Yu, Jie Chen 0035, Huiyao Wan, Yice Cao, Zhixiang Huang, Yingsong Li 0001, Bocai Wu, Baidong Yao
IEEE Trans. Geosci. Remote. Sens.8
2023 SARNas: A Hardware-Aware SAR Target Detection Algorithm via Multiobjective Neural Architecture Search
abstract
Most of the existing deep learning-based SAR target detection algorithms rely on manual experience to repeatedly adjust structures and parameters to design models suitable for specific scenarios or tasks. The implementation of the above methods is complicated, the design efficiency is low, and it is difficult to ensure the balance between accuracy and complexity. We innovatively propose a hardware-aware SAR target detection algorithm via multiobjective neural architecture search (NAS), referred to as SARNas. First, we design a flexible and efficient search space, a supernet search strategy and a subnet contribution evaluation strategy. Furthermore, we construct a new NAS loss function, called SARMI-Loss, to guide the learning of a SAR object detector that balances accuracy and computational complexity. Our SAR-Nas method can address the resource limitations of edge devices and automatically search for the optimal SAR target detector in an end-to-end manner for any deep learning-based SAR baseline model. A series of comparative experiments on three SAR image object detection datasets (SSDD, HRSID and MSAR) demonstrate the superiority of our method. The experimental results with YOLOV5 as the benchmark model show that the detection accuracy of the target detection networks automatically found by using the SARNas method on the SSDD, HRSID, and MSAR datasets can reach 98.5%, 92.8%, and 91.8% in mean average precision (mAP) with only 2.31M, 1.99M, 2.21M parameters, respectively. The number of model parameters is reduced by 88.9%, 90.46%, and 68.5%, respectively, and the inference speed is increased by 51.6%, 46.1%, and 13.9% without losing accuracy.
Wentian Du, Jie Chen 0035, Chaochen Zhang, Po Zhao, Huiyao Wan, Yice Cao, Zhixiang Huang, Yingsong Li 0001, Bocai Wu
IEEE Trans. Geosci. Remote. Sens.10
2023 Orientation Detector for Ship Targets in SAR Images Based on Semantic Flow Feature Alignment and Gaussian Label Matching
abstract
To address the challenges in synthetic aperture radar (SAR) ship target detection, this paper proposes a SAR ship small target orientation detector named FADet based on semantic flow feature alignment and Gaussian label matching. First, to solve the feature misalignment problem caused by feature extraction downsampling and residual connections, we introduce the FAM module into FPN, which automatically aligns deep and shallow fine-grained semantics information through semantic flow alignment. Second, due to the scattering characteristics of SAR imaging, the boundary information of SAR targets is not obvious, we combining attention mechanisms design an adaptive boundary enhancement module to enhance the target boundary information. Finally, to solve the problem that small targets have difficulty matching positive samples under IOU rules, we design a label matching strategy based on Gaussian distribution. This matching strategy can still learn regression information when two boxes do not intersect. Based on the SSDD+ and RSDD-SAR datasets, the effectiveness of each module in FADet is verified by ablation experiments. Additionally, through comparison experiments with the latest orientation detection methods, FADet achieves a good compromise between accuracy and inference speed. The AP50 and AP75 on the SSDD+ and RSDD-SAR is 91.03, 59.94 and 90.78, 59.91 respectively, and the FPS is 19.83.
Huiyao Wan, Jie Chen 0035, Zhixiang Huang, Wentian Du, Feng Xu 0001, Feng Wang 0022, Bocai Wu
IEEE Trans. Geosci. Remote. Sens.7
2023 HRLE-SARDet: A Lightweight SAR Target Detection Algorithm Based on Hybrid Representation Learning Enhancement
abstract
In recent years, deep learning has been widely used in remote sensing, especially in the field of synthetic aperture radar (SAR) image target detection. However, all of these deep learning models continue increasing the network’s depth and width without maintaining a good balance between accuracy and speed. Therefore, in this article, we propose a hybrid representation learning-enhanced SAR target detection algorithm based on the unique features of SAR images from a lightweight perspective called HRLE-SARDet. First, we design a lightweight and scattering feature extraction backbone that is more suitable for SAR image data. Second, for the multiscale feature discrepancy, we design a new multiscale feature fusion neck. Next, to better extract the scattering information from small targets of SAR images and improve the detection accuracy, we design a lightweight hybrid representation learning enhancement module. Finally, to better fit target detection for SAR image datasets, we redesign a more flexible loss function, which allows for an easy adjustment of the importance of polynomial bases according to the target task and dataset. Extensive experimental results on three SAR image ship target datasets (SSDD, AIR-SARShip-2.0, and HRSID) and a newly released large multiclass target SAR dataset (MSAR-1.0) show that our HRLE-SARDet achieves 98.4%, 79.2%, 92.5%, and 88.4% mean average precision (mAP) with only 1.09 M parameters and 2.5 G floating-point operations (FLOPs) on the SSDD, AIR-SARShip-2.0, HRSID, and MSAR-1.0 datasets, respectively, which is an excellent performance.
Jie Chen 0035, Zhixiang Huang, Jianming Lv, Honglin Luo, Bocai Wu, Yingsong Li 0001, Paulo S. R. Diniz
IEEE Trans. Geosci. Remote. Sens.7
2023 FPT: Fine-Grained Detection of Driver Distraction Based on the Feature Pyramid Vision Transformer
abstract
According to the surveys of the World Health Organization, distracted driving is one of main causes of road traffic accidents. To improve road traffic safety, real-time detection of drivers’ driving behavior is very important for the development of highly reliable Advanced Driver Assistance System (ADAS). At present, the deep learning architecture based on a Convolutional Neural Network (CNN) has disadvantages such as large number of parameters and weak global feature extraction ability. Therefore, this paper proposes an innovative driver distraction detection model based on the fusion of a transformer and a CNN, referred to as FPT, which is the first exploration in the field of driver distraction detection. First, we introduce the latest Twins transformer as a benchmark. Then, we design residual embedding to replace block embedding, which can further integrate the convolutional neural network with Transformer and improve the feature extraction ability. In addition, the Multilayer Perceptron (MLP) module with a large parameter occupancy rate in the original transformer structure is replaced with a lightweight group convolution module to reduce computational complexity. Finally, a cross-entropy loss function for label smoothing is designed to guide network learning with significantly differentiated features. Comparison results on two large-scale driver distraction detection datasets show that the proposed FPT offers a better compromise between computational cost and performance compared to the state-of-the-art CNN and Transformer architectures.
Jie Chen 0035, Zhixiang Huang, Bing Li 0033, Jianming Lv, Jingmin Xi, Bocai Wu, Jun Zhang 0034, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.7
2022 Geometric Auto-Calibration of SAR Images Utilizing Constraints of Symmetric Geometry
abstract
Synthetic aperture radar (SAR) has evolved into an essential Earth observation technique. Geometric quality is one of its most fundamental elements to ensure subsequent applications. Excellent localization accuracy requires comprehensive error compensation and calibration. While traditional calibration work is time-consuming and depends on the test sites. Besides, the available calibration test sites are scarce and unevenly distributed for certain SAR satellites. It is, hence, necessary to research how to make geometric calibration without ground control points (GCPs) for SAR. This study proposes a method utilizing the constraints of symmetric geometry that could cancel out plane error symmetrically, and then calculates calibration constants combined with external DEM. Experiments demonstrate the calibration accuracy of the proposed way is approximately 0.3 m along range direction when all constraints are met, and the discrepancy between the proposed method and that using GCPs are analyzed. Generally, the method can achieve fast, labor-saving, and normalized calibration work for SAR images in the absence of available control data.
Kai Xu 0008, Guo Zhang 0001, Bocai Wu
IEEE Geosci. Remote. Sens. Lett.6
2022 AFSar: An Anchor-Free SAR Target Detection Algorithm Based on Multiscale Enhancement Representation Learning
abstract
Unlike optical images, synthetic aperture radar (SAR) images have unique characteristics, such as few samples, strong scattering, sparseness, multiple scales, complex interference and background, and inconspicuous target edge contour information. Current SAR target detection algorithms have difficulty in balancing accuracy and speed, and the performance of these algorithms is relatively limited, thus making it difficult to deploy practical applications. To this end, this article proposes AFSar, an innovative anchor-free SAR target detection algorithm based on multiscale enhancement representation learning. First, we introduce the latest anchor-free architecture YOLOX as the basic framework. Second, to reduce the computational complexity of the model and to improve the ability of multiscale feature extraction, we redesigned the lightweight backbone, namely, MobileNetV2S. Furthermore, we propose an attention enhancement PAN module, called CSEMPAN, which highlights the unique strong scattering characteristics of SAR targets by integrating channel and spatial attention mechanisms. Finally, in view of the multiscale and strong sparse characteristics of SAR targets, we propose a new target detection head, namely, ESPHead. ESPHead extracts the features of targets with different scales by using dilated convolution with different dilated rates, so as to enhance the detection ability of the model for targets with different scales. The results of ablation experiments on the SSDD dataset show that the mAP of our algorithm reaches 0.977, while the Flops is only 9.86 G, achieving state of the art.
Huiyao Wan, Jie Chen 0035, Zhixiang Huang, Runfan Xia, Bocai Wu, Baidong Yao, Mengdao Xing
IEEE Trans. Geosci. Remote. Sens.5
2022 FSODS: A Lightweight Metalearning Method for Few-Shot Object Detection on SAR Images
abstract
At present, few-shot object detection research in the field of optical remote sensing images has been conducted, but few-shot object detection in the field of SAR images have rarely been explored. To this end, this paper proposes a lightweight meta-learning-based SAR image few-shot object detection method, which improves the accuracy and speed of SAR image few-shot object detection from a more balanced perspective. First, we introduce the latest FSODM method in optical remote sensing as a benchmark framework. Second, a lightweight meta-feature extractor named DarknetS is designed to enhance the feature representation of SAR images and improve detection timeliness. Furthermore, we build a new aggregation module called AggregationS, which encodes support features and query features into the same feature subspace via a novel transformer encoder. This module design can better extract the correlation and saliency between different classes in the support set, improve the detection accuracy of the query set, and enhance the detection generalization performance of new classes. Finally, we built several real-world SAR image few-shot object detection datasets to verify the effectiveness of the method. Experimental results show that FSODS can achieve a better object detection performance compared to the baseline model under the condition that only a small amount of labelled data is required for new classes of SAR image objects.
Jie Chen 0035, Zhixiang Huang, Huiyao Wan, Pei Chang, Baidong Yao, Bocai Wu, Mengdao Xing
IEEE Trans. Geosci. Remote. Sens.8
2021 Fine-Grained Detection of Driver Distraction Based on Neural Architecture Search
abstract
In the future, vehicles will be equipped with increasingly advanced interactive intelligent electronic devices, which will induce drivers to conduct secondary tasks, thereby leading to distractions. Therefore, the detection and early warning of driver distraction are essential for improving driving safety and pose an important challenge in intelligent transportation systems. Previous studies used traditional machine learning and deep learning transfer models, which have the disadvantages of complicated and time-consuming manual feature engineering, strong subjectivity, and weak generalization performance. In this paper, we propose a fine-grained detection method for driver distraction based on neural architecture search. First, we design an automatic construction algorithm for deep convolutional neural networks based on neural architecture search, which automatically searches for the optimal deep convolutional neural network architecture without human involvement. In addition, we fuse driver-related multisource perception information, use an automatically constructed deep convolutional neural network to extract high-dimensional mapping features, and implement fine-grained detection of various types of driver distraction states. The results on a large-scale multimodal driver distraction dataset demonstrate that our method can efficiently search an optimal deep convolutional neural network, which can quickly converge, and can accurately detect the considered types of driver distraction states, the average detection accuracy reaches 99.7796%; moreover, it has satisfactory robustness.
Jie Chen 0035, Zhixiang Huang, Xiaohui Guo, Bocai Wu
IEEE Trans. Intell. Transp. Syst.5