Libo Yao

dblp:189/3536 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0001-8110-0211ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Oriented Decoupling Target Detection Method for SAR Image Based on Multi-Channel Localization and Soft Thresholding
abstract
Synthetic Aperture Radar (SAR) images are crucial for maritime vessel detection; however, challenges such as blurred ship edges, strong land scattering interference, and angular regression mismatches across varying target sizes hinder accurate rotational localization. In this paper, an oriented decoupling target detection method (R-MCLST) is proposed to address these issues. The method integrates three key modules: a multi-channel positioning module (MC-PM) that employs distributed average pooling and additional coordinate channels to enhance orientation awareness; a soft threshold-based multilayer perceptron (ST-MLP) that effectively mitigates background interference while robustly extracting complex features; and a Gaussian distribution-based prediction box (GD-BPB) that transforms rotated bounding box encoding into a two-dimensional Gaussian distribution using KL divergence for adaptive parameter adjustment. Experimental evaluations on the R-SSDD and MR-HRSID datasets demonstrate that R-MCLST achieves superior performance, with the R-SSDD dataset yielding AP50 of 87.48%, AP75 of 34.96%, AP of 41.15%, and AR of 46.52%, and the MR-HRSID dataset yielding AP50 of 61.59%, AP75 of 4.96%, AP of 19.13%, and AR of 22.24%. Comparative analyses confirm that the proposed method outperforms current state-of-the-art networks in accurately localizing rotating targets under challenging SAR imaging conditions.
Gui Gao, Gang Yang 0006, Libo Yao, Xi Zhang 0028, Gaosheng Li
IEEE Trans. Circuits Syst. Video Technol.4
2025 An Oriented Ship Detection Method of Remote Sensing Image With Contextual Global Attention Mechanism and Lightweight Task-Specific Context Decoupling
abstract
Ship detection in remote sensing images has been attracting a lot of attention due to its great application value in both military and civilian fields. However, ships in high-resolution remote sensing images are characterized by the remarkable features of multiscale, arbitrary orientation, and dense arrangement, which is a great challenge for fast and accurate target detection. In order to solve problems, we propose a YOLOV5-based oriented ship detection method of remote sensing images with contextual global attention mechanism and lightweight task-specific context decoupling (CGTC-RYOLO) in this article. First, a cross-stage partial context transformer (CSP-COT) module is introduced to capture global contextual spatial relations using multihead self-attention (MHSA) to verify their implications in implicit dependencies. Second, we propose an angle classification prediction branch in the YOLOV5 head network for detecting targets in any direction and design probability and distribution loss function (PrfoIoU) to optimize the regression effect. Third, the lightweight task-specific context decoupling (LTSCODE) for target detection is employed to replace the original head in the YOLOV5 model, which is used to solve the accuracy problem caused by YOLOV5’s hybridization of classification and localization. Ablation experiments demonstrate the importance and effectiveness of each module. Compared with the benchmark model, the CGTC-RYOLO has the 5.9%, 3.7%, and 4.3% mAP improvements on the DOTA-ship dataset, the HRSC2016 dataset, and the UCAS-AOD dataset, respectively. Moreover, the model’s generalization is also validated. Compared with state-of-the-artmethods, the CGTC-RYOLO can achieve better accuracy and fewer parameters.
Gui Gao, Gang Yang 0006, Libo Yao, Xi Zhang 0028, Heng-Chao Li 0001, Gaosheng Li
IEEE Trans. Geosci. Remote. Sens.5
2025 A Multibranch Embedding Network With Bi-Classifier for Few-Shot Ship Classification of SAR Images
abstract
Ship classification in synthetic aperture radar (SAR) images is a challenge in the field of ocean monitoring. On the one hand, there are few labeled samples in SAR remote sensing ship datasets, and a commonly used single classification criterion cannot effectively represent the distribution of categories. On the other hand, the small size of the SAR ship and the inconspicuous appearance characteristics lead to the fact that the SAR ship samples are with less discriminative information; therefore, the rich feature space of a ship cannot be effectively obtained, which increases the difficulty of target distinguishability. A multibranch embedding network with bi-classifier (MBEN-BC) model was proposed to address these problems and for few-shot SAR ship classification. First, the MBEN module was utilized to extract the multiscale feature map spatial information of the input image at multiple levels and establish cross-channel information interaction so as to obtain discriminative features at the local and global levels, which effectively enriched the feature space. Then, the BC module was constructed to represent the image features from the image level and descriptor level, respectively, and the two classification criteria were presented to promote a more compact distribution of similar samples in the feature space in order to effectively represent the distribution of categories with a small number of labeled samples. Experimental validation was carried out using the FUSAR-Ship, Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset, and OPENSAR-Ship dataset, and the MBEN-BC method achieved superior performance and good generalization ability compared to the current popular and state-of-the-art few-shot methods.
Gui Gao, Meixiang Wang, Libo Yao, Xi Zhang 0028, Heng-Chao Li 0001, Gaosheng Li
IEEE Trans. Geosci. Remote. Sens.4
2024 METSM: Multiobjective energy-efficient task scheduling model for an edge heterogeneous multiprocessor system
Qiangqiang Jiang, Xu Xin, Libo Yao, Bo Chen 0015
Future Gener. Comput. Syst.3
2024 RepSViT: An Efficient Vision Transformer Based on Spiking Neural Networks for Object Recognition in Satellite On-Orbit Remote Sensing Images
abstract
The role of on-orbit computing for satellites is transitioning from being a backup measure to becoming a primary key function. However, the limited computing resources available on satellites make it difficult to deploy advanced models with large parameters. Additionally, satellite on-orbit computing requires high speed and accuracy, posing significant challenges for developing suitable models. To overcome these challenges, we propose an efficient vision transformer, RepSViT, for satellite on-orbit computing. The RepSViT introduces Spiking neural networks (SNNs) with high biological plausibility, event-driven property and low power consumption into the field of remote sensing image processing and satellite on-orbit computing for the first time and incorporates structural reparameterization. Specifically, we design a dynamic dilated spiking convolution (D2SC) based on SNNs to improve the feature extraction capability and efficiency of RepSViT. We also develop a spiking guided attention module (SGAM) to make RepSViT pay more attention to object-related features with lower computational costs. Furthermore, we design an efficient coupled fine–coarse-grained block (ECFC) to enhance the model’s capability in extracting coarse and fine-grained features. To ensure effective feature extraction, inference speed and reduced computational costs, we design a reparameterized feed-forward network (RepFFN). RepSViT achieves an inference latency of 8.33 ms and a recognition accuracy of 95% on an embedded GPU, utilizing 3.77 million parameters and consuming 0.6 GFLOPs computational costs.
Yanhua Pang, Libo Yao, Chengguo Dong, Qinglei Kong, Bo Chen 0015
IEEE Trans. Geosci. Remote. Sens.2
2024 Cog-Net: A Cognitive Network for Fine-Grained Ship Classification and Retrieval in Remote Sensing Images
abstract
In light of the escalating volume of high-quality remote sensing ship images, the imperative task is to effectively classify and retrieve such images from extensive remote sensing archives. While prior research has yielded promising outcomes in ship classification, a comprehensive framework catering to both ship classification and retrieval remains absent. Additionally, prevailing studies neglect the critical aspect of model interpretability, merely furnishing predicted results devoid of a transparent reasoning process. This opacity, coupled with the high-stakes nature of outcomes, significantly impedes the safe utilization of these models. To address these problems, this paper introduces the cognitive network (Cog-Net), an inherently interpretable model tailored for fine-grained ship classification and retrieval in remote sensing images. Cog-Net imitates the reasoning process employed by domain experts, navigating fromperceptiontocognitionduring decision-making. The initial stage incorporates the Causal Multi-grained Feature Learning (CMFL) module, mirroring the humanperceptualprocess by identifying salient regions of a ship object within entire images as references for visual concept learning. Subsequently, the second stage introduces the Visual Concept Learning (VCL) module, imitates the humancognitiveprocess by learning basis visual concepts for each ship category and generating predictions through interpretable reasoning based on these basis visual concepts. Furthermore, to facilitate experimentation, a novel dataset, Fine-grained Ship Remote Sensing Image Slices (FGSRSI-23), comprising 23 fine-grained ship subcategories, is constructed. Extensive experiments are conducted, encompassing our FGSRSI-23 dataset alongside two publicly available datasets, FGSC-23 and FGSCR-42. Results attest to the competitiveness of Cog-Net in both ship image classification and retrieval tasks, offering a transparent and interpretable reasoning process for predicted outcomes.
Libo Yao
IEEE Trans. Geosci. Remote. Sens.3
2024 Confusion Region Mining for Crowd Counting
abstract
Existing works mainly focus on crowd and ignore the confusion regions which contain extremely similar appearance to crowd in the background, while crowd counting needs to face these two sides at the same time. To address this issue, we propose a novel end-to-end trainable confusion region discriminating and erasing network called CDENet. Specifically, CDENet is composed of two modules of confusion region mining module (CRM) and guided erasing module (GEM). CRM consists of basic density estimation (BDE) network, confusion region aware bridge and confusion region discriminating network. The BDE network first generates a primary density map, and then the confusion region aware bridge excavates the confusion regions by comparing the primary prediction result with the ground-truth density map. Finally, the confusion region discriminating network learns the difference of feature representations in confusion regions and crowds. Furthermore, GEM gives the refined density map by erasing the confusion regions. We evaluate the proposed method on four crowd counting benchmarks, including ShanghaiTech Part_A, ShanghaiTech Part_B, UCF_CC_50, and UCF-QNRF, and our CDENet achieves superior performance compared with the state-of-the-arts.
Jiawen Zhu 0003, Wenda Zhao 0003, Libo Yao, You He 0002, Maodi Hu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.3
2023 Dual Attention Feature Fusion for Visible-Infrared Object Detection
Limin Shi, Libo Yao, Lubin Weng
ICANN (7)3
2023 A Bi-Prototype BDC Metric Network With Lightweight Adaptive Task Attention for Few-Shot Fine-Grained Ship Classification in Remote Sensing Images
abstract
Fine-grained ship classification in optical remote sensing images is a major challenge in the ocean observation field, elaborated as follows: First, the cost of acquiring ship images is expensive. Obtaining numerous labeled samples is difficult, resulting in the poor generalization ability of training models. Second, the features of ship target cannot be accurately obtained owing to complex background interference. Third, inter-class similarity and intra-class diversity among different ships render ship classification difficult. In this study, we propose LATA-BP-BDC: a bi-prototype Brownian Distance Covariance (BDC) metric network with lightweight adaptive task attention (LATA) for few-shot fine-grained ship classification. First, the LATA module is used to generate 3-dimensional (3D) weights, which can effectively reduce complex background interference and improve the adaptive capturing ability of target features without including additional network operators. Second, we input target features into the BDC metric module and output the BDC matrices to represent image information. Because the similarity between two images can be calculated as the corresponding BDC matrices distance, the improvement of the relevance of similar targets can be realized. Finally, we use the bi-prototype module to generate highly accurate prototypes, further calibrating information differences between images, which enhances the correlation between the same category samples and separability between different categories samples. Consequently, this process effectively reduces the influence of large intra-class appearance variation and small inter-class appearance variation. We perform validation using two fine-grained datasets, FGSCR and CUB. Compared with state-of-the-art methods, the LATA-BP-BDC achieves a superior performance and has good generalization for fine-grained few-shot classification.
Gui Gao, Libo Yao, Jia Liu 0055, Dingfeng Duan
IEEE Trans. Geosci. Remote. Sens.3
2023 Depth-Distilled Multi-Focus Image Fusion
abstract
Homogeneous regions, which are smooth areas that lack blur clues to discriminate if they are focused or non-focused. Therefore, they bring a great challenge to achieve high accurate multi-focus image fusion (MFIF). Fortunately, we observe that depth maps are highly related to focus and defocus, containing a preponderance of discriminative power to locate homogeneous regions. This offers the potential to provide additional depth cues to assist MFIF task. Taking depth cues into consideration, in this paper, we propose a new depth-distilled multi-focus image fusion framework, namely D2MFIF. In D2MFIF, depth-distilled model (DDM) is designed for adaptively transferring the depth knowledge into MFIF task, gradually improving MFIF performance. Moreover, multi-level fusion mechanism is designed to integrate multi-level decision maps from intermediate outputs for improving the final prediction. Visually and quantitatively experimental results demonstrate the superiority of our method over several state-of-the-art methods.
Fan Zhao 0005, Wenda Zhao 0003, Huimin Lu 0001, Yong Liu 0017, Libo Yao, Yu Liu 0005
IEEE Trans. Multim.5
2022 Image-Scale-Symmetric Cooperative Network for Defocus Blur Detection
abstract
Defocus blur detection (DBD) for natural images is a challenging vision task especially in the presence of homogeneous regions and gradual boundaries. In this paper, we propose a novel image-scale-symmetric cooperative network (IS2CNet) for DBD. On one hand, in the process of image scales from large to small, IS2CNet gradually spreads the recept of image content. Thus, the homogeneous region detection map can be optimized gradually. On the other hand, in the process of image scales from small to large, IS2CNet gradually feels the high-resolution image content, thereby gradually refining transition region detection. In addition, we propose a hierarchical feature integration and bi-directional delivering mechanism to transfer the hierarchical feature of previous image scale network to the input and tail of the current image scale network for guiding the current image scale network to better learn the residual. The proposed approach achieves state-of-the-art performance on existing datasets.Codes and results are available at:https://github.com/wdzhao123/IS2CNet.
Fan Zhao 0006, Huimin Lu 0001, Wenda Zhao 0003, Libo Yao
IEEE Trans. Circuits Syst. Video Technol.4
2022 Generalizable Crowd Counting via Diverse Context Style Learning
abstract
Existing crowd counting approaches predominantly perform well on the training-testing protocol. However, due to large style discrepancies not only among images but also within a single image, they suffer from obvious performance degradation when applied to unseen domains. In this paper, we aim to design a generalizable crowd counting framework which is trained on a source domain but can generalize well on the other domains. To reach this, we propose a gated ensemble learning framework. Specifically, we first propose a diverse fine-grained style attention model to help learn discriminative content feature representations, allowing for exploiting diverse features to improve generalization. We then introduce a channel-level binary gating ensemble model, where diverse feature prior, input-dependent guidance and density grade classification constraint are implemented, to optimally select diverse content features to participate in the ensemble, taking advantage of their complementary while avoiding redundancy. Extensive experiments show that our gating ensemble approach achieves superior generalization performance among four public datasets. Codes are publicly available athttps://github.com/wdzhao123/DCSL.
Wenda Zhao 0003, Yu Liu 0005, Huimin Lu 0001, Cong'an Xu, Libo Yao
IEEE Trans. Circuits Syst. Video Technol.6
2022 Feature Balance for Fine-Grained Object Classification in Aerial Images
abstract
Fine-grained object classification (FGOC) focuses on identifying subcategories of objects, which is crucial in military and civilian. Existing FGOC methods primarily focus on high-resolution aerial images, limiting their application on low-resolution (LR) FGOC that is a more realistic setting, especially on resource-constrained satellite devices. It is more challenging to deal with LR FGOC since objects’ details are blurred or missing. Addressing this issue, we make the first attempt to explore LR FGOC and propose a novel pipeline based on two technical insights: 1) feature balance strategy discriminatively integrates super-resolution weak and strong detailed presentations into coarse features of LR aerial images, achieving a feature balance to avoid that the weak detailed presentations are inhibited by the strong ones and 2) iterative interaction mechanism alternately refines feature details of the discriminative ship regions and optimizes the performance of FGOC. Moreover, we build a low-resolution fine-grained object (LFS) dataset to promote further study and evaluation. Extensive experiments on the proposed LFS dataset and the other three object datasets of DOTA, FS23, and HRSC2016 demonstrate that our method outperforms state-of-the-art algorithms. Dataset and code are publicly available athttps://github.com/wdzhao123/FBNet.
Wenda Zhao 0003, Tingting Tong, Libo Yao, Yu Liu 0005, Cong'an Xu, You He 0002, Huchuan Lu
IEEE Trans. Geosci. Remote. Sens.3
2021 Information Fusion of GF-1 and GF-4 Satellite Imagery for Ship Surveillance
abstract
Gaofen-4 (GF-4) satellite is the first high-resolution geostationary orbit (GEO) optical satellite in the world, which has appealing potential in wide-area maritime surveillance and can provide dynamic information of ships as its high revisit time. However, its spatial resolution is not high enough to obtain attributes of ships, while Gaofen-1 (GF-1) satellite in low earth orbit (LEO) can extract rich features, as its high spatial resolution. Therefore, the integration of GF-1 and GF-4 satellites can provide a better maritime situational awareness. The aim of the paper is to provide a feasible architecture of data fusion of GF-1 and GF-4 imagery for maritime surveillance, and a novel ship association method using multi-level information in order to improve the association accuracy, and experimental results verify the effectiveness of our method.
Yong Liu 0017, Pengyu Guo, Mingjiang Ji, Libo Yao
IGARSS5
2019 GF-4 Satellite and Automatic Identification System Data Fusion for Ship Tracking
abstract
GF-4 satellite, as the first geostationary orbit optical remote sensing satellite with medium resolution, has the abilities to point and stare at a particular sea area for maritime surveillance. The integration of GF-4 image sequence and automatic identification system (AIS) reports has the appealing potential to provide a better maritime situational awareness by tracking ships, especially noncooperative targets. The aim of this letter is to provide a track-level fusion architecture for GF-4 and AIS heterogeneous data based on the characteristics of different data sources to form a consolidated and comprehensive maritime picture. The proposed fusion strategy is tested in the specific area of the East China Sea, and experimental results verify the effectiveness of our method.
Yong Liu 0017, Libo Yao
IEEE Geosci. Remote. Sens. Lett.2
2017 Track filtering for space-based maritime surveillance in geographic coordinates
abstract
In order to filter tracks of ship targets for space-based maritime surveillance using electronic reconnaissance satellites, an extended Kalman filter (EKF) algorithm in geographic coordinates is proposed in this paper. Firstly, different methods of Dead reckoning (DR) are analysed in different coordinate systems. Then, the formula of EKF based on middle latitude sailing is derived. Finally, satellite-based automatic identification system (AIS) data and simulation data are used to perform track prediction and track filtering respectively, which verifies the validity of the proposed filtering method based on middle latitude sailing.
Yong Liu 0017, Libo Yao
FUSION3
2016 Fusion detection of ship targets in low resolution multi-spectral images
abstract
Aiming at ship detection in multi-spectral images at low resolution, this paper proposes a new method for ship detection based on fusion detection which combines spectral feature with thermal feature. Firstly, it selects infrared band instead of visible band image to detect cloud according to size feature. With cloud available, it does segmentation work in thermal infrared image for the removal of cloud pixels. Then, the result is mapped to the IR images and cloud masking is completed. Next, the fusion of two kinds of images using wavelet transform is adopted. At last, the fused image is used to detect ships and morphological operations are used to discriminate ships. The experiment result on multi-spectral data of Landsat 8 shows that the proposed method which is robust against clutter can detect ships effectively.
Yong Liu 0017, Libo Yao
IGARSS2