Wei Wang 0245

dblp:35/7092-245 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-8770-3862ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Incorporating Multiscale Context and Task-Consistent Focal Loss into Oriented Object Detection
abstract
Oriented object detection (OOD) in remote sensing images (RSIs) aims to precisely localize and identify objects with arbitrary orientations. Two-stage OOD methods attract lots of interest due to their superior accuracy, however, they still face two major problems. First of all, the misclassification problem frequently occurs because the majority of classification strategies solely relies on the features of proposals. Secondly, most of loss functions cannot simultaneously concentrate on hard samples and boost the consistency between identification and localization, which restricts the further improvement of OOD models. To address the first problem, the multi-scale context (MSC) is incorporated into a two-stage OOD model in this paper. Specifically,Ncontextual branches are added to predict the class confidence score (CCS) of each proposal and itsNenlarged proposals which contain the MSC, and the final CCS of each proposal is determined by the mean value of aboveN+ 1 CCSs. To tackle the second problem, a task-consistent focal (TF) loss is proposed. The TF loss employs the difficulty of localization as the weight of classification loss, and the difficulty of identification is used as the weight of regression loss. Concentrating on hard samples and synchronous optimization of classification and regression can be achieved by minimizing the TF loss. The ablation studies show the validity of MSC, TF and their combination. The comparison with popular OOD models demonstrates the superior performance of our model on the DOTA and DIOR-R datasets. The source code can be obtained from https://github.com/qxlzengli/MSC-TF.
Xiaoliang Qian, Qingqing Jian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.3
2025 IPS-YOLO: Iterative Pseudo-Fully Supervised Training of YOLO for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images only requires image-level labels, greatly reducing the cost of manual annotations. Recently, the pseudo-fully supervised object detection (pseudo-FSOD) models give better performance than traditional WSOD models based on multiple instance learning. However, the existing pseudo-FSOD models still have two problems that need to be addressed. Firstly, existing models tend to focus on the salient parts of object rather than the whole object. Secondly, the pseudo ground truth (PGT) cannot be continuously improved during the training process. For the first problem, a filtering and weighted synthesis guided by category confidence score (FWSC) strategy is proposed to generate PGT. The FWSC strategy firstly removes the instances with low category confidence score (CCS), then, to cover the whole object as much as possible, any two instances are weighted and synthesized according to their CCSs if they have very high spatial overlap, otherwise, the non-maximum suppression (NMS) operation is conducted. For the second problem, an iterative refinement (IR) scheme of PGT is proposed. Specifically, the PGT instances produced by the FWSC strategy are firstly used to train a YOLO model, then a filtering and weighted synthesis guided by IoU (FWSI) strategy employs the detection results inferred from the trained YOLO model to refine the PGT instances, and above two steps can be repeated multiple times. Furthermore, three improvement strategies are proposed to enhance the traditional baseline WSOD model in proposals generation, the selection of augmented samples, and the definition of pseudo-labels, respectively. The ablation studies demonstrate the effectiveness of FWSC strategy, IR scheme of PGT, and the three improvement strategies of baseline model. The comparisons with popular WSOD models show that our model gives the best results on the NWPU VHR-10. v2 and DIOR datasets.The source codes have been released at https://github.com/qxlzengli/IPS-YOLO.
Xiaoliang Qian, Baihui Zhang, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.4
2024 Complete and Invariant Instance Classifier Refinement for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images is used to detect high-value objects by utilizing image-level labels. However, the current models still have two problems. Firstly, the misclassification of neighboring instances is easily occurred because the one-hot label is assigned to all of seed instances and their neighboring instances. Secondly, the supervisory information of each instance classifier refinement (ICR) branch is generated from the predicted class score of upper ICR branch rather than the real label, thus the prediction mistake of each ICR branch will be accumulated with the propagation of supervisory information. To address the first problem, a complete definition of pseudo soft label (CPSL) of instances is proposed to directly train each ICR branch, where the CPSL of seed instances is defined according to the predicted class scores of upper ICR branch, and the CPSL of other instances are determined by the spatial distance weighted feature similarity between them and seed instances. To handle the second problem, an invariant multiple instance learning (IMIL) scheme is proposed to indirectly train each ICR branch by using the real image-level labels. Furthermore, the affine transformations of original image are incorporated into the baseline model to enhance the invariance of our model. The ablation studies verify the effectiveness of CPSL, IMIL and their combination. The quantitative comparisons with popular methods show that the 73.63% (31.08%) mAP and 79.88% (57.52%) CorLoc of our method is the best on the NWPU VHR-10.v2 (DIOR) dataset, and the qualitative comparisons intuitively demonstrate it again.
Xiaoliang Qian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.3
2023 Multiple Instance Complementary Detection and Difficulty Evaluation for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images (RSIs) has attracted lots of attention because it solely employs image-level labels to drive the model training. Most of the WSOD methods incline to mine salient object as positive instance, and the less salient objects are considered as negative instances, which will cause the problem of missing instances. In addition, the quantity of hard and easy instances is usually imbalanced, consequently, the cumulative loss of a large amount of easy instances dominates the training loss, which limits the upper bound of WSOD performance. To handle the first problem, a complementary detection network (CDN) is proposed, which consists of a complementary multiple instance detection network (CMIDN) and a complementary feature learning (CFL) module. The CDN can capture robust complementary information from two basic multiple instance detection networks (MIDNs) and mine more object instances. To handle the second problem, an instance difficulty evaluation metric named instance difficulty score (IDS) is proposed, which is employed as the weight of each instance in the training loss. Consequently, the hard instances will be assigned larger weights according to the IDS, which can improve the upper bound of WSOD performance. The ablation experiments demonstrate that our method significantly increases the baseline method by large margins, i.e. 23.6% (10.2%) mAP and 32.4% (13.1%) CorLoc gains on the NWPU VHR-10.v2 (DIOR) dataset. Our method obtains 58.1% (26.7%) mAP and 72.4% (47.9%) CorLoc on the NWPU VHR-10.v2 (DIOR) dataset, which achieves better performance compared with seven advanced WSOD methods.
Yu Huo 0001, Xiaoliang Qian, Chao Li 0072, Wei Wang 0245
IEEE Geosci. Remote. Sens. Lett.4
2023 Mining High-Quality Pseudoinstance Soft Labels for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection in remote sensing image (RSI) is still a challenge because of the lack of instance-level labels, and many existing methods have two problems. Firstly, most of the existing methods usually mine the pseudo ground truth (PGT) instances solely relying on proposal class scores (PCS). Actually, the reliability of PCS is not enough because of the bird’s eye view imaging and large-scale chaotic background of RSIs, and the instances with high PCS incline to cover the discriminative region rather than the whole object. Secondly, the existing methods assign a one-hot label to each instance, and the label of PGT instance is copied to its neighbor instances, which induces the misclassification problem to some extent. Actually, the probability that the neighbor instances contain the object with the same category is smaller than the PGT instance. For the first problem, the proposal quality score (PQS) is proposed for mining high-quality PGT instances, which contains PCS and dual-context projection score (DCPS). The DCPS is calculated through semantic segmentation, and is employed to measure the completeness that each proposal covers an object. For the second problem, a pseudo soft label assignment (PSLA) strategy is proposed to assign more precise soft label for each instance, where the soft label is determined by the spatial distance between each instance and its nearest PGT instance. The ablation study validates the effectiveness of the PQS and PSLA. The comprehensive comparisons with other WSOD methods on three popular benchmarks show the excellent performance of our method.
Xiaoliang Qian, Yu Huo 0001, Gong Cheng 0003, Chenyang Gao, Xiwen Yao, Wei Wang 0245
IEEE Trans. Geosci. Remote. Sens.6
2023 Building a Bridge of Bounding Box Regression Between Oriented and Horizontal Object Detection in Remote Sensing Images
abstract
Oriented object detection (OOD) aims to precisely detect the objects with arbitrary orientation in remote sensing images. Up to now, most of bounding box regression (BBR) losses for OOD are transferred from horizontal object detection (HOD) methods, however, the transferring requires lots of professional knowledge and experiences for designers, consequently, many excellent BBR losses for HOD have not been transferred to OOD. To accelerate the research progress of BBR loss for OOD, a unified transferring strategy (UTS) is proposed to facilitate the transferring of BBR loss from HOD to OOD. The UTS proposes that the BBR of oriented bounding box (OBB) can be converted into the joint BBR of its horizontal smallest enclosing rectangle (HSER) and two offsets, so the BBR loss in HOD can be easily transferred to OOD by using HSER as a bridge. Following the UTS, a BBR loss named Rotated-IoU (RIoU) loss is designed for OOD by transferring an advanced BBR loss in HOD, which can be considered as an example to show how to transfer. On the basis of RIoU loss, a focal rotated-IoU (FRIoU) loss is proposed to assign larger weights to hard samples in the BBR. The comparisons with other BBR losses show that the RIoU and FRIoU losses can give better performance. The ablation study shows that giving more attention to hard samples in BBR is effective. The comparisons with many advanced methods demonstrate that the combinations of baseline methods and FRIoU loss achieve state-of-the-art performance on the DOTA and DIOR-R datasets.
Xiaoliang Qian, Baokun Wu, Gong Cheng 0003, Xiwen Yao, Wei Wang 0245, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Co-Saliency Detection Guided by Group Weakly Supervised Learning
abstract
The detection results of many existing co-saliency detection methods are easily interfered by the unrelated salient objects, which have similar appearance characteristics to co-salient objects. Therefore, mining the inter-saliency cues which contain the common category information of multiple related images is the core of co-saliency detection. To address above concern, a novel group weakly supervised learning induced co-saliency detection (GWSCoSal) model is proposed in this paper. First of all, a novel group class activation maps (GCAM) network is constructed and trained through a group weakly supervised learning scheme, which adopts the common category of a group of related images as the ground truth. The GCAM produced by the trained GCAM network are considered as the inter-saliency cues, which can only highlight the regions covered by the objects with common category. Afterwards, the GCAM are integrated into a feature pyramid networks (FPN) based backbone trained by the pixel-level labels to infer the co-saliency maps. The group weakly supervised and the pixel-level learning are jointly implemented for end-to-end training of GWSCoSal model. The comprehensive comparisons with 13 state-of-the-art methods demonstrate that, our GWSCoSal model can detect the co-salient objects more accurately under the condition of being interfered by the similar unrelated salient objects, and the overall performance of which has achieved the level of state-of-the-art methods. The ablation study of our GWSCoSal model validates the effectiveness of proposed GCAM network.
Xiaoliang Qian, Yinfeng Zeng, Wei Wang 0245, Qiuwen Zhang
IEEE Trans. Multim.3
2020 Micro-cracks detection of solar cells surface via combining short-term and long-term deep features
Xiaoliang Qian, Jing Li 0067, Jinde Cao, Yuanyuan Wu 0002, Wei Wang 0245
Neural Networks5