Xiaoliang Qian

dblp:31/467 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-4328-6411ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Domain adaptation for remote sensing image semantic segmentation with prototype-driven domain disentangle alignment
Xiufei Zhang, Yuanwei Liu, Xiaoliang Qian, Gong Cheng 0003, Xiwen Yao
Pattern Recognit.3
2025 Incorporating Multiscale Context and Task-Consistent Focal Loss into Oriented Object Detection
abstract
Oriented object detection (OOD) in remote sensing images (RSIs) aims to precisely localize and identify objects with arbitrary orientations. Two-stage OOD methods attract lots of interest due to their superior accuracy, however, they still face two major problems. First of all, the misclassification problem frequently occurs because the majority of classification strategies solely relies on the features of proposals. Secondly, most of loss functions cannot simultaneously concentrate on hard samples and boost the consistency between identification and localization, which restricts the further improvement of OOD models. To address the first problem, the multi-scale context (MSC) is incorporated into a two-stage OOD model in this paper. Specifically,Ncontextual branches are added to predict the class confidence score (CCS) of each proposal and itsNenlarged proposals which contain the MSC, and the final CCS of each proposal is determined by the mean value of aboveN+ 1 CCSs. To tackle the second problem, a task-consistent focal (TF) loss is proposed. The TF loss employs the difficulty of localization as the weight of classification loss, and the difficulty of identification is used as the weight of regression loss. Concentrating on hard samples and synchronous optimization of classification and regression can be achieved by minimizing the TF loss. The ablation studies show the validity of MSC, TF and their combination. The comparison with popular OOD models demonstrates the superior performance of our model on the DOTA and DIOR-R datasets. The source code can be obtained from https://github.com/qxlzengli/MSC-TF.
Xiaoliang Qian, Qingqing Jian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.1
2025 IPS-YOLO: Iterative Pseudo-Fully Supervised Training of YOLO for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images only requires image-level labels, greatly reducing the cost of manual annotations. Recently, the pseudo-fully supervised object detection (pseudo-FSOD) models give better performance than traditional WSOD models based on multiple instance learning. However, the existing pseudo-FSOD models still have two problems that need to be addressed. Firstly, existing models tend to focus on the salient parts of object rather than the whole object. Secondly, the pseudo ground truth (PGT) cannot be continuously improved during the training process. For the first problem, a filtering and weighted synthesis guided by category confidence score (FWSC) strategy is proposed to generate PGT. The FWSC strategy firstly removes the instances with low category confidence score (CCS), then, to cover the whole object as much as possible, any two instances are weighted and synthesized according to their CCSs if they have very high spatial overlap, otherwise, the non-maximum suppression (NMS) operation is conducted. For the second problem, an iterative refinement (IR) scheme of PGT is proposed. Specifically, the PGT instances produced by the FWSC strategy are firstly used to train a YOLO model, then a filtering and weighted synthesis guided by IoU (FWSI) strategy employs the detection results inferred from the trained YOLO model to refine the PGT instances, and above two steps can be repeated multiple times. Furthermore, three improvement strategies are proposed to enhance the traditional baseline WSOD model in proposals generation, the selection of augmented samples, and the definition of pseudo-labels, respectively. The ablation studies demonstrate the effectiveness of FWSC strategy, IR scheme of PGT, and the three improvement strategies of baseline model. The comparisons with popular WSOD models show that our model gives the best results on the NWPU VHR-10. v2 and DIOR datasets.The source codes have been released at https://github.com/qxlzengli/IPS-YOLO.
Xiaoliang Qian, Baihui Zhang, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.1
2024 A color edge extraction method on color image
Qinge Wu, Zhichao Song, Yingbo Lu, Lintao Zhou, Xiaoliang Qian
Multim. Tools Appl.6
2024 Complete and Invariant Instance Classifier Refinement for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images is used to detect high-value objects by utilizing image-level labels. However, the current models still have two problems. Firstly, the misclassification of neighboring instances is easily occurred because the one-hot label is assigned to all of seed instances and their neighboring instances. Secondly, the supervisory information of each instance classifier refinement (ICR) branch is generated from the predicted class score of upper ICR branch rather than the real label, thus the prediction mistake of each ICR branch will be accumulated with the propagation of supervisory information. To address the first problem, a complete definition of pseudo soft label (CPSL) of instances is proposed to directly train each ICR branch, where the CPSL of seed instances is defined according to the predicted class scores of upper ICR branch, and the CPSL of other instances are determined by the spatial distance weighted feature similarity between them and seed instances. To handle the second problem, an invariant multiple instance learning (IMIL) scheme is proposed to indirectly train each ICR branch by using the real image-level labels. Furthermore, the affine transformations of original image are incorporated into the baseline model to enhance the invariance of our model. The ablation studies verify the effectiveness of CPSL, IMIL and their combination. The quantitative comparisons with popular methods show that the 73.63% (31.08%) mAP and 79.88% (57.52%) CorLoc of our method is the best on the NWPU VHR-10.v2 (DIOR) dataset, and the qualitative comparisons intuitively demonstrate it again.
Xiaoliang Qian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.1
2024 Attention Erasing and Instance Sampling for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) trains detectors by only weak labels, aiming to save the burden of expensive bounding box-level annotations. Most previous efforts formulate WSOD as a multiple instance learning (MIL) problem, which is prone to detect discriminative object parts and miss object instances. This article proposes an attention erasing and instance sampling (AE-IS) approach to alleviate the above problems. Concretely, we first apply an attention erasing (AE) scheme to the WSOD model to hide the most discriminative region for capturing the integral extent of the object. Then, we employ an intersection-over-union (IoU)-balanced sampling component toward mining more object instances. Moreover, an instance reweighted loss (IRL) is designed to learn a larger portion of object instances, thereby further enhancing the performance of the object detector. Experimental results demonstrate that our method significantly improves the baseline approach by great margins and achieves competitive performance with the state-of-the-art algorithms on the NWPU VHR-10.v2 (72.0% mAP, 76.1% CorLoc) and DIOR (29.1% mAP, 55.9% CorLoc) datasets. The source code will be available athttps://github.com/XuanX/AE-IS.
Gong Cheng 0003, Xiaoxu Feng, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Multiple Instance Complementary Detection and Difficulty Evaluation for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images (RSIs) has attracted lots of attention because it solely employs image-level labels to drive the model training. Most of the WSOD methods incline to mine salient object as positive instance, and the less salient objects are considered as negative instances, which will cause the problem of missing instances. In addition, the quantity of hard and easy instances is usually imbalanced, consequently, the cumulative loss of a large amount of easy instances dominates the training loss, which limits the upper bound of WSOD performance. To handle the first problem, a complementary detection network (CDN) is proposed, which consists of a complementary multiple instance detection network (CMIDN) and a complementary feature learning (CFL) module. The CDN can capture robust complementary information from two basic multiple instance detection networks (MIDNs) and mine more object instances. To handle the second problem, an instance difficulty evaluation metric named instance difficulty score (IDS) is proposed, which is employed as the weight of each instance in the training loss. Consequently, the hard instances will be assigned larger weights according to the IDS, which can improve the upper bound of WSOD performance. The ablation experiments demonstrate that our method significantly increases the baseline method by large margins, i.e. 23.6% (10.2%) mAP and 32.4% (13.1%) CorLoc gains on the NWPU VHR-10.v2 (DIOR) dataset. Our method obtains 58.1% (26.7%) mAP and 72.4% (47.9%) CorLoc on the NWPU VHR-10.v2 (DIOR) dataset, which achieves better performance compared with seven advanced WSOD methods.
Yu Huo 0001, Xiaoliang Qian, Chao Li 0072, Wei Wang 0245
IEEE Geosci. Remote. Sens. Lett.2
2023 Mining High-Quality Pseudoinstance Soft Labels for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection in remote sensing image (RSI) is still a challenge because of the lack of instance-level labels, and many existing methods have two problems. Firstly, most of the existing methods usually mine the pseudo ground truth (PGT) instances solely relying on proposal class scores (PCS). Actually, the reliability of PCS is not enough because of the bird’s eye view imaging and large-scale chaotic background of RSIs, and the instances with high PCS incline to cover the discriminative region rather than the whole object. Secondly, the existing methods assign a one-hot label to each instance, and the label of PGT instance is copied to its neighbor instances, which induces the misclassification problem to some extent. Actually, the probability that the neighbor instances contain the object with the same category is smaller than the PGT instance. For the first problem, the proposal quality score (PQS) is proposed for mining high-quality PGT instances, which contains PCS and dual-context projection score (DCPS). The DCPS is calculated through semantic segmentation, and is employed to measure the completeness that each proposal covers an object. For the second problem, a pseudo soft label assignment (PSLA) strategy is proposed to assign more precise soft label for each instance, where the soft label is determined by the spatial distance between each instance and its nearest PGT instance. The ablation study validates the effectiveness of the PQS and PSLA. The comprehensive comparisons with other WSOD methods on three popular benchmarks show the excellent performance of our method.
Xiaoliang Qian, Yu Huo 0001, Gong Cheng 0003, Chenyang Gao, Xiwen Yao, Wei Wang 0245
IEEE Trans. Geosci. Remote. Sens.1
2023 Building a Bridge of Bounding Box Regression Between Oriented and Horizontal Object Detection in Remote Sensing Images
abstract
Oriented object detection (OOD) aims to precisely detect the objects with arbitrary orientation in remote sensing images. Up to now, most of bounding box regression (BBR) losses for OOD are transferred from horizontal object detection (HOD) methods, however, the transferring requires lots of professional knowledge and experiences for designers, consequently, many excellent BBR losses for HOD have not been transferred to OOD. To accelerate the research progress of BBR loss for OOD, a unified transferring strategy (UTS) is proposed to facilitate the transferring of BBR loss from HOD to OOD. The UTS proposes that the BBR of oriented bounding box (OBB) can be converted into the joint BBR of its horizontal smallest enclosing rectangle (HSER) and two offsets, so the BBR loss in HOD can be easily transferred to OOD by using HSER as a bridge. Following the UTS, a BBR loss named Rotated-IoU (RIoU) loss is designed for OOD by transferring an advanced BBR loss in HOD, which can be considered as an example to show how to transfer. On the basis of RIoU loss, a focal rotated-IoU (FRIoU) loss is proposed to assign larger weights to hard samples in the BBR. The comparisons with other BBR losses show that the RIoU and FRIoU losses can give better performance. The ablation study shows that giving more attention to hard samples in BBR is effective. The comparisons with many advanced methods demonstrate that the combinations of baseline methods and FRIoU loss achieve state-of-the-art performance on the DOTA and DIOR-R datasets.
Xiaoliang Qian, Baokun Wu, Gong Cheng 0003, Xiwen Yao, Wei Wang 0245, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Co-Saliency Detection Guided by Group Weakly Supervised Learning
abstract
The detection results of many existing co-saliency detection methods are easily interfered by the unrelated salient objects, which have similar appearance characteristics to co-salient objects. Therefore, mining the inter-saliency cues which contain the common category information of multiple related images is the core of co-saliency detection. To address above concern, a novel group weakly supervised learning induced co-saliency detection (GWSCoSal) model is proposed in this paper. First of all, a novel group class activation maps (GCAM) network is constructed and trained through a group weakly supervised learning scheme, which adopts the common category of a group of related images as the ground truth. The GCAM produced by the trained GCAM network are considered as the inter-saliency cues, which can only highlight the regions covered by the objects with common category. Afterwards, the GCAM are integrated into a feature pyramid networks (FPN) based backbone trained by the pixel-level labels to infer the co-saliency maps. The group weakly supervised and the pixel-level learning are jointly implemented for end-to-end training of GWSCoSal model. The comprehensive comparisons with 13 state-of-the-art methods demonstrate that, our GWSCoSal model can detect the co-salient objects more accurately under the condition of being interfered by the similar unrelated salient objects, and the overall performance of which has achieved the level of state-of-the-art methods. The ablation study of our GWSCoSal model validates the effectiveness of proposed GCAM network.
Xiaoliang Qian, Yinfeng Zeng, Wei Wang 0245, Qiuwen Zhang
IEEE Trans. Multim.1
2022 Guiding Clean Features for Object Detection in Remote Sensing Images
abstract
Recently, object detection has gained significant progress in remote sensing images. Nevertheless, we conclude two defects in remote sensing image object detection. At first, most methods rely on feature pyramid, but the features of different levels would influence each other when we use top-down operation. Second, the traditional label assignment strategy cannot assign suitable labels, as it adopts the fixed intersection over union (IoU) threshold to divide positive samples and negative samples during training. According to the problems we pointed out, a simple yet effective framework is employed to eliminate these two limitations. It integrates two novel components: aware feature pyramid network (AFPN) and group assignment strategy (GAS). AFPN is to mitigate the adverse effects caused by the first problem. Specifically, it learns a vector for the higher level features in the feature pyramid to obtain clean features. As for the second limitation, we recommend a new label assignment strategy named GAS to address this problem. Samples will be grouped according to their overlaps with ground truth, and then, they are assigned to positive or negative labels in each group. Extensive experiments are conducted on the large-scale object detection dataset DIOR and DOTA. With the newly introduced two key components, our model significantly improves the detection accuracy. Without bells and whistles, our proposed method achieves 2.0% and 1.9% higher mean average precision (mAP) than Faster R-CNN with FPN when using ResNet50 and ResNet101 as the backbones, respectively. Finally, we obtain 73.3% mAP on the DIOR dataset without any tricks. Our code is available athttps://github.com/hm-better/dior_detect.
Gong Cheng 0003, Hailong Hong, Xiwen Yao, Xiaoliang Qian, Lei Guo 0002
IEEE Geosci. Remote. Sens. Lett.5
2021 Small target recognition method on weak features
Qinge Wu, Ziming An, Xiaoliang Qian
Multim. Tools Appl.4
2021 Two-Stream Encoder GAN With Progressive Training for Co-Saliency Detection
abstract
The recent end-to-end co-saliency models have good performance, however, they cannot express the semantic consistency among a group of images well and usually require many co-saliency labels. To this end, a two-stream encoder generative adversarial network (TSE-GAN) with progressive training is proposed in this paper. In the pre-training stage, the salient object detection generative adversarial networks (SOD-GAN) and classification network (CN) are separately trained by the salient object detection (SOD) datasets and co-saliency datasets with only category labels to learn the intra-saliency and preliminary inter-saliency cues and alleviate the problem of insufficient co-saliency labels. In the second training stage, the backbone of TSE-GAN is inherited from the trained SOD-GAN, the encoder of trained SOD-GAN (SOD-Encoder) is used to extract intra-saliency features, the group-wise semantic encoder (GS-Encoder) is constructed by the multi-level group-wise category features extracted from CN for extracting inter-saliency features with better semantic consistency, the TSE-GAN constructed by incorporating the GS-Encoder into SOD-GAN is trained on co-saliency datasets for co-saliency detection. The comprehensive comparisons with 13 state-of-the-art methods demonstrate the effectiveness of proposed method.
Xiaoliang Qian, Gong Cheng 0003, Xiwen Yao, Liying Jiang
IEEE Signal Process. Lett.1
2020 Novel visual tracking approach via ant lion optimiser
abstract
Ant lion optimiser (ALO) is a new nature‐inspired swarm intelligence optimisation algorithm that mimics the hunting mechanism of antlions in nature. ALO has been proved to have the merits of high exploitation and convergence speed benefiting from adaptive boundary shrinking mechanism and elitism. In this work, visual tracking is expressed as searching for object in whole search space by interaction between antlions and ants. A novel ALO‐based visual tracking framework is proposed and the adaptation and sensitivity of the parameters in ALO are discussed to improve tracking performance. In addition, considering that ALO tracker needs a lot of iteration consumption, kernel correlation filter with deep feature is integrated into the ALO tracking framework (ALOKCF) to improve track efficiency. Extensive experimental results prove that the ALO tracker is very competitive compared to other trackers, especially for abrupt motion tracking. At the same time, two visual tracking benchmarks are used to verify ALOKCF tracker achieves state‐of‐the‐art performance.
Huanlong Zhang, Zeng Gao, Jie Zhang 0066, Xiankai Lu, Jian Chen 0038, Guohao Nie, Xiaoliang Qian
IET Image Process.7
2020 Micro-cracks detection of solar cells surface via combining short-term and long-term deep features
Xiaoliang Qian, Jing Li 0067, Jinde Cao, Yuanyuan Wu 0002, Wei Wang 0245
Neural Networks1
2019 Performance Comparison of Two Pooling Strategies for Remote Sensing Image Scene Classification
abstract
With the advances of convolutional neural networks (CNNs), the accuracy of remote sensing image scene classification has been greatly boosted thanks to the powerful features extracted through CNNs. Although significant success has been achieved, most of existing methods are dominated by the use of fully-connected CNN features. This paper focuses on the performance comparison of two kinds of novel pooling strategies, including generalized max pooling (GMP) and taskdriven pooling (TDP), for remote sensing image scene classification. To this end, an off-the-shelf CNN model is used as backbone network to extract multi-scale convolutional features. Then, GMP and TDP are respectively adopted to obtain globally pooled features. Finally, scene classification is performed with support vector machine (SVM). In the experiment, we evaluate the performance of these two kinds of pooling schemes on a widely-used scene classification benchmark data set. The experimental results show that (i) using pooled CNN convolutional features can obtain better results than using fully-connected CNN features and (ii) TDP is slightly better than GMP.
Maoxiong Wu, Gong Cheng 0003, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001, Lei Guo 0002
IGARSS4
2018 A Novel Visual Tracking Method Based on Moth-Flame Optimization Algorithm
Huanlong Zhang, Xiujiao Zhang, Xiaoliang Qian, Yibin Chen
PRCV (4)3
2018 Extended cuckoo search-based kernel correlation filter for abrupt motion tracking
abstract
Kernelised correlation filter (KCF)‐based trackers have recently attracted considerable attention due to their exciting accuracy and efficiency. Numerous improvements have been made later for coping with scales variation or partial occlusion etc . However, when there is an abrupt motion between the consecutive image frames, these trackers would face failure. To alleviate the problem, the authors present an extended cuckoo search (CS)‐based KCF tracker (called ECSKCF). At first, the extended CS algorithm is constructed by the Simplex method (SM). CS has obvious capability in global search while the SM has exceptional advantage in local search. Based on ECS method, motion prediction is transformed to globally search for optimal position intending to enhance the quality of base image. Then, combined ECS with Gaussian distribution, a hybrid motion model is introduced to KCF framework, which has the capability of capturing abrupt motion. Finally, a unified framework is designed to track smooth or abrupt motion simultaneously. Extensive experimental results in both quantitative and qualitative measures demonstrate the effectiveness of the authors’ proposed method for abrupt motion tracking.
Huanlong Zhang, Xiujiao Zhang, Yong Wang 0032, Xiaoliang Qian, Yanfeng Wang 0002
IET Comput. Vis.4
2017 Fast depth map mode decision based on depth-texture correlation and edge classification for 3D-HEVC
Qiuwen Zhang, Kunqiang Huang, Xiaoliang Qian, Yong Gan
J. Vis. Commun. Image Represent.5
2014 Image visual attention computation and application via the learning of object attributes
Junwei Han 0001, Ling Shao 0001, Xiaoliang Qian, Gong Cheng 0003, Jungong Han
Mach. Vis. Appl.4
2013 Optimal contrast based saliency detection
Xiaoliang Qian, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002
Pattern Recognit. Lett.1
2013 An Object-Oriented Visual Saliency Detection Framework Based on Sparse Coding Representations
abstract
Saliency detection aims at quantitatively predicting attended locations in an image. It may mimic the selection mechanism of the human vision system, which processes a small subset of a massive amount of visual input while the redundant information is ignored. Motivated by the biological evidence that the receptive fields of simple cells in V1 of the vision system are similar to sparse codes learned from natural images, this paper proposes a novel framework for saliency detection by using image sparse coding representations as features. Unlike many previous approaches dedicated to examining the local or global contrast of each individual location, this paper develops a probabilistic computational algorithm by integrating objectness likelihood with appearance rarity. In the proposed framework, image sparse coding representations are yielded through learning on a large amount of eye-fixation patches from an eye-tracking dataset. The objectness likelihood is measured by three generic cues called compactness, continuity, and center bias. The appearance rarity is inferred by using a Gaussian mixture model. The proposed paper can serve as a basis for many techniques such as image/video segmentation, retrieval, retargeting, and compression. Extensive evaluations on benchmark databases and comparisons with a number of up-to-date algorithms demonstrate its effectiveness.
Junwei Han 0001, Xiaoliang Qian, Lei Guo 0002, Tianming Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3