Longbao Wang

dblp:01/7779 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-6164-1253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021Computer networks · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MSTNet: Multi-Scale Contextual Analysis Network for Semantic Segmentation of Remote Sensing Images
abstract
ABSTRACT With the rapid development of deep learning, research on semantic segmentation of remote sensing images has made significant progress. However, there are common problems in remote sensing images, such as large‐scale differences between different types of objects and unbalanced sample numbers, which leads to poor semantic segmentation results, especially for small targets and rare types of objects. To address these challenges, a remote sensing image semantic segmentation method based on multi‐scale contextual information analysis named MSTNet is innovatively proposed. Its core design includes the semantic information enhancement module (SIE) of feature adaptive clustering, which strengthens the feature expression of different categories through adaptive clustering to alleviate sample imbalance; the weighted feature fusion module (WFF) adaptively aggregates cross‐level features, cooperates with the multi‐scale context enhancement module (MSCE), combines convolution and Transformer operations, and deeply mines local and global contexts to cope with scale changes; in addition, the network also contains a pixel space feature optimisation module (SFEM) to enhance spatial details. Experiments on the UAVid, LoveDA, Potsdam and Vaihingen datasets show that MSTNet significantly improves the ability to handle scale changes and imbalance problems and reaches advanced levels in key indicators such as OA, mIoU, and mF1, proving that MSTNet achieves competitive or even better performance.
Longbao Wang, Mingxuan Wang, Xiaoliang Luo, Lvchun Wang, Chong Long
IET Image Process.1
2025 DSSNet: An Anchor-Free Rotated Object Detection Network With Dynamic Sample Selection for Remote Sensing Images
abstract
ABSTRACT Object detection in remote sensing imagery requires precise localisation and identification of targets under challenging conditions. Facing the challenges of arbitrary target orientations, wide‐scale variations, dense distributions, and small objects in remote sensing object detection, anchor‐based methods suffer from inadequate rotated target representation using rectangular boxes. This necessitates excessive angle‐specific anchors, leading to heavy computational overhead, severe sample imbalance, and slow speeds unsuitable for mobile deployment. To address these accuracy‐efficiency trade‐offs, we propose DSSNet: an anchor‐free rotated object detection network with dynamic sample selection for remote sensing images. DSSNet replaces traditional backbones with the parameter‐efficient ConvNeXt‐T and utilises an FPN for accelerated multi‐scale feature extraction. During prediction, it employs a shape‐adaptive selection strategy combined with a contour point quality assessment strategy to dynamically refine target contour points, enabling real‐time rotated object detection. The efficacy of DSSNet has been thoroughly validated through benchmark comparisons on diverse datasets. On the DOTA dataset, DSSNet clearly outperforms baseline methods in detection performance, achieving a mean Average Precision (mAP) of 76.97% and the fastest detection speed of 26.2 frames per second (FPS).
Longbao Wang, Yongheng Yu, Xiaoliang Luo, Lvchun Wang, Yican Shen, Zhijun Zhou, Hongmin Gao 0001
IET Image Process.1
2025 MKGFA: Multimodal Knowledge Graph Construction and Fact-Assisted Reasoning for VQA
abstract
Knowledge-based visual question answering relies on open-ended external knowledge and a fine-grained comprehension of both the visual content of images and semantic information. Existing methods for utilizing knowledge have the following limitations: (1) Language pre-training methods output answers in the form of plain text, which only understand shallow visual content; (2) The knowledge retrieved by image objects as labels is represented as first-order logic, making it difficult to infer complex questions. To address the above problems, this paper integrates visual-textual multimodal information, accumulates domain-specific and external multi-modal knowledge, introduces and supplements external objective facts, and proposes a multimodal knowledge graph construction and fact-assisted reasoning network (MKGFA). The network consists of three parts: the multimodal knowledge graph construction module (MKGC), the objective fact-assisted reasoning module (FAR), and the answer inference module. The MKGC engages in the coarse-to-fine-grained learning of triplet representations for multimodal knowledge units. The FAR establishes deep cross-modal relations between visual objects and factual words for correlating real answers. The answer inference module makes the final decision based on the results of both. Among them, the former two modules employ a pre-training and fine-tuning strategy, systematically accumulating foundational and domain-specific knowledge. Compared with the state-of-the-arts, MKGFA achieves 1.09% and 0.7% higher accuracy on the two challenging OKVQA and KRVQA datasets, respectively. The experimental results demonstrate the complementary advantages of the integration of the two modules.
Longbao Wang, Libing Zhang, Shufang Xu, Hongmin Gao 0001
Int. J. Comput. Intell. Appl.1
2024 CSFFNet: Lightweight cross-scale feature fusion network for salient object detection in remote sensing images
abstract
Abstract Salient object detection (SOD), one of the most important applications in the field of computer vision, aims to extract the most visually appealing regions of scenes. However, the improvement of the accuracy of existing salient object detection in optical remote sensing images (ORSI‐SOD) is usually accompanied by an increase of network complexity, which affects the application of these models. Motivated by this, a novel lightweight edge‐supervised neural network for ORSI‐SOD is proposed, named CSFFNet. Specifically, the backbone (ResNet34) is first lightened by feature encoding module (FEM), building a lightweight subnet for feature extraction. Then, in the transformer‐based feature pyramid enhancement module (FPEM), the convolutional features obtained in the FEM are enhanced by long‐distance dependence to obtain multi‐scale features containing rich saliency cues. Based on this, the feature fusion module (FFM) is designed to capture cross‐scale long‐range dependencies and effectively fuse high‐level semantic information with low‐level detail information. Thus, the increase in network complexity due to multi‐level decoding is avoided. Finally, the segmentation results are optimized by using salient edges as auxiliary information, which effectively improves the contrast and completeness of the results. Experimental results on two public datasets demonstrate that the lightweight CSFFNet achieves competitive or even better performance compared with state‐of‐the‐art methods.
Longbao Wang, Chong Long, Xin Li 0090, Xiaodan Tang, Zhipeng Bai, Hongmin Gao 0001
IET Image Process.1
2024 Rotated points for object detection in remote sensing images
abstract
Abstract Object detection in remote sensing images poses great challenges due to the dense distribution, arbitrary orientation, and aspect ratio variations of objects. Most of the existing methods rely on aligned convolutional features, which fail to capture the geometric information of objects effectively and result in the inconsistency between the classification score and localization accuracy. Moreover, densely packed objects suffer from spatial feature aliasing caused by the intersection of reception fields between objects. To address this issue, a deformable convolution‐based method named rotated points is proposed, which consists of two modules: a point set loss module and a high‐quality sample assignment module. The point set loss module can extract geometric features of objects in arbitrary directions with fine‐grained point sets for feature representation and introduce outlier penalties to penalize outlier points. The high‐quality sample assignment module measures the classification and localization ability, orientation quality, and point‐wise correlation of point sets comprehensively to enhance the consistency of classification and regression significantly. Experiments on the DOTA and FAIR1M datasets demonstrate that the proposed method achieves significant improvements over the benchmark model.
Longbao Wang, Yican Shen, Hongmin Gao 0001
IET Image Process.1
2024 A CBAM-GAN-based method for super-resolution reconstruction of remote sensing image
abstract
Abstract As satellite imagery technology advances, remote sensing plays an increasingly prominent role in modern society. Nevertheless, the limitations of existing imaging sensors and complex atmospheric conditions constrain the quality of raw remote sensing data, posing challenges for interpretation and noise reduction. Super‐resolution technology focuses on enhancing low‐quality, low‐resolution remote sensing images. In this study, we introduce a method that utilizes a high‐order degradation model to generate low‐resolution remote sensing images. We employ a Generative Adversarial Network with a Convolutional Block Attention Module (CBAM‐GAN) to enhance these images, reducing noise interference and improving texture and feature display. Our approach outperforms other methods on the UCMerced‐LandUse, WHU‐RS19, and AID datasets. Specifically, it raises SSIM index scores to 0.9443, 0.8928, and 0.8633, respectively, exceeding baselines by 1.31%, 0.19%, and 1.30%. The MOS index also improves to 3.98, 3.96, and 3.83, respectively, representing a 2.31%, 8.20%, and 2.96% gain over the baseline. Our reconstruction produces superior results, demonstrating the effectiveness of our proposed method.
Longbao Wang, Xin Li 0090, Hongmin Gao 0001
IET Image Process.1
2022 Adaptive sparse ternary gradient compression for distributed DNN training in edge computing
abstract
Abstract In edge computing, though distributed training of Deep Neural Networks (DNNs) is expected to exchange massive gradients between parameter servers and working nodes, the high communication cost constrains the training speed. To break this limitation, gradient compression algorithms expect the ultimate compression ratio at the expense of the accuracy of the trained model. Therefore, new gradient compression techniques are necessary to ensure both communication efficiency and model accuracy. This paper introduces a novel technique—an Adaptive Sparse Ternary Gradient Compression (ASTC) scheme, which relies on the number of gradients in model layers to compress gradients. ASTC establishes the model compression selection criterion by gradients’ amount, compresses the network layer that meets the model’s standard, evaluates the gradients’ importance based on entropy to adaptively perform sparse compression, and finally conducts ternary quantization compression and a lossless code scheme on sparse gradients. Using public datasets (MNIST, CIFAR-10, Tiny ImageNet) and deep learning models (CNN, LeNet5, ResNet18) for experimental evaluation, we exhibit excellent results that the training efficiency of ASTC is about 1.6 times, 1.37 times, and 1.1 times higher than that of Top-1, AdaComp, and SBC, respectively. Furthermore, ASTC can be improved by an average of about $$1.9\%$$ 1.9% in training accuracy compared with the above approaches.
Yingchi Mao, Jun Wu 0001, Xuesong Xu, Longbao Wang
CCF Trans. High Perform. Comput.4
2015 Low-complexity soft-interference cancellation turbo equalisation for multi-input-multi-output systems with multilevel modulations
abstract
This study presents a low‐complexity soft‐interference cancellation equaliser (SICE) for the turbo detection of multiple‐input–multiple‐output systems operating in time dispersive channels. The SICE contains three time‐invariant linear filters: a feedforward filter, a causal feedback filter and an anti‐causal feedback filter. The feedforward filter is designed to suppress the intersymbol interference because of channel time dispersion and the multiplexing interference from multiple transmit antennas. The causal (or anti‐causal) feedback filter is developed to remove the residual interference caused by the symbols transmitted before (or after) the symbol under detection. The filters are designed by analysing the statistical properties of soft decisions. The performance of the proposed SICE is verified through both extrinsic information transfer (EXIT) chart analysis and computer simulations. The EXIT chart analysis shows that, because of the inclusion of the anti‐causal soft decision, the SICE performance approaches the ideal matched filter bound as the iteration progresses. Consequently, the proposed SICE achieves significant performance gains over conventional equalisers.
Jingxian Wu 0001, Longbao Wang, Chengshan Xiao
IET Commun.2
2013 Low complexity soft-interference cancellation turbo equalization for MIMO systems with multilevel modulations
abstract
This paper presents a low complexity soft-interference cancellation equalizer (SICE) for the turbo detection of multiple-input multiple-output (MIMO) systems operating in time dispersive channels. The SICE contains three time-invariant linear filters: a feedforward filter, a causal feedback filter and an anti-causal feedback filter. The feedforward filter is designed to suppress the intersymbol interference (ISI) due to time dispersive channels and the multiplexing interference from multiple transmit antennas. The causal (or anti-causal) feedback filter is developed to remove the residual interference caused by the symbols transmitted before (or after) the symbol under detection. The performance of the proposed SICE is verified through both extrinsic information transfer chart (EXIT) analysis and computer simulations. The analytical and simulation results demonstrated that the inclusion of the anti-causal soft decision during SICE is critical to the system performance. The EXIT chart analysis shows that the SICE performance approaches the ideal matched filter bound as the iteration progresses.
Jingxian Wu 0001, Longbao Wang, Chengshan Xiao
GLOBECOM2
2013 Low-complexity turbo detection for single-carrier low-density parity-check-coded multiple-input multiple-output underwater acoustic communications
abstract
ABSTRACT A low‐complexity turbo detection scheme is proposed for single‐carrier multiple‐input multiple‐output (MIMO) underwater acoustic (UWA) communications using low‐density parity‐check (LDPC) channel coding. The low complexity of the proposed detection algorithm is achieved in two aspects: first, the frequency‐domain equalization technique is adopted, and it maintains a low complexity irrespective of the highly dispersive UWA channels; second, the computation of the soft equalizer output, in the form of extrinsic log‐likelihood ratio, is performed with an approximating method, which further reduces the complexity. Moreover, attributed to the LDPC decoding, the turbo detection converges within only a few iterations. The proposed turbo detection scheme has been used for processing real‐world data collected in two different undersea trials: WHOI09 and ACOMM09. Experimental results show that it provides robust detection for MIMO UWA communications with different modulations and different symbol rates, at different transmission ranges. Copyright © 2011 John Wiley & Sons, Ltd.
Longbao Wang, Jun Tao 0004, Chengshan Xiao
Wirel. Commun. Mob. Comput.1