VLDB 2026 Research / reviewers in the wild / expert
Xiaoxu Feng
dblp:253/1916
· DBLP profile ↗
20ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0003-2639-2447ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-perspective filter pruning via diversity and independence collaboration
Chenyang Gao, Qinglong Cao, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003 |
Pattern Recognit. | 4 |
| 2026 | DecoupleNet: Domain-specific task decoupling network for low-light image enhancement
Peiliang Huang, Xianmin Chen, Xiaoxu Feng, Qiangqiang Wang, Dingwen Zhang, Longfei Han, Junwei Han 0001 |
Pattern Recognit. | 3 |
| 2026 | Retinex-RAWMamba: Bridging Demosaicing and Denoising for Low-Light RAW Image EnhancementabstractLow-light image enhancement, particularly in cross-domain tasks such as mapping from the raw domain to the sRGB domain, remains a significant challenge. Many deep learning-based methods have been developed to address this issue and have shown promising results in recent years. However, single-stage methods, which attempt to unify the complex mapping across both domains, leading to limited denoising performance. In contrast, existing two-stage approaches typically overlook the characteristic of demosaicing within the Image Signal Processing (ISP) pipeline, leading to color distortions under varying lighting conditions, especially in low-light scenarios. To address these issues, we propose a novel Mamba-based method customized for low light RAW images, called RAWMamba, to effectively handle raw images with different CFAs. Furthermore, we introduce a Retinex Decomposition Module (RDM) grounded in Retinex prior, which decouples illumination from reflectance to facilitate more effective denoising and automatic non-linear exposure correction, reducing the effect of manual linear illumination enhancement. By bridging demosaicing and denoising, better enhancement for low light RAW images is achieved. Experimental evaluations conducted on public datasets SID and MCR demonstrate that our proposed RAWMamba achieves state-of-the-art performance on cross-domain mapping. The code is available at https://github.com/Cynicarlos/RetinexRawMamba. Xianmin Chen, Longfei Han, Peiliang Huang, Xiaoxu Feng, Dingwen Zhang, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Cross-Modality Domain Adaptation Based on Semantic Graph Learning: From Optical to SAR ImagesabstractSynthetic aperture radar (SAR) imaging provides a distinct advantage in scene understanding due to its capability for all-weather data acquisition. However, in comparison to easily annotated optical remote sensing images, the lower imaging quality of SAR images presents significant challenges in obtaining manually annotated training data, which poses substantial issues for SAR image analysis. In this paper, we employ the domain adaptation (DA) that leverages labeled optical images to better understand unlabeled SAR images. Global feature alignment as a method for DA has demonstrated effectiveness in transferring knowledge, yet it faces challenges in cross-modality adaptation from optical remote sensing to SAR images due to their differing imaging mechanisms. With distinct visual features between optical and SAR images, the semantic dependency is difficult to construct, which results in low-quality pseudo-label assignment for SAR images. To address the above issue, we propose a semantic graph learning framework to comprehensively align the global features of optical remote sensing and SAR images by modeling the cross-modality semantics and generating high-quality pseudo-labels. It can be applied for SAR scene classification and object detection when only optical remote sensing images are labeled. Specifically, a cross-modality semantic graph alignment (CSGA) module is constructed to model and align the second-order semantic dependencies by aggregating cross-modality visual semantic information. Then, an uncertainty-based robust pseudo-label generation (URPG) module is designed to generate pseudo-labels for effective semantic alignment and self-training by modeling the uncertainty of pseudo-labels for each SAR image. Comprehensive experiments show that our proposed method outperforms the state-of-the-art methods on scene classification (NWPU-RESISC45→WHU-SAR6, MLRSNet→NWPU-SAR6, MLRSNet→NWPU-SAR6, and NWPU-RESISC45→NWPU-SAR6) and object detection (MASATI-ship→SSDD, MVSRD→SARDet-vehicle, and DIOR-airplane→SAR-airplane) tasks. The code and datasets are publicly accessible at https://github.com/XZhang878/SGLF. Xiufei Zhang, Zhongling Huang, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DIMA: Digging Into Multigranular Archetype for Fine-Grained Object DetectionabstractFine-grained remote sensing object detection aims at precisely locating objects and determining the fine-level categories. This task is exceptionally challenging due to the substantial interclass similarity, presenting difficulties in capturing discriminative features. We attribute this to the absence of essential information that can serve as supervision for the learning. This involves comprehensive visual patterns of objects and intrinsic relationships of multigranular features. In this article, we propose a novel scheme dubbed as digging into the multigranular archetype (DIMA) for fine-grained remote sensing object detection. In detail, we first design a simple yet effective frequency-aware representation supplement (FARS) mechanism learning from original images and their auxiliary frequency counterparts simultaneously. The FARS introduces high- and low-frequency representations to reinforce a range of visual cues, such as particular regions associated with the former and contours of objects related to the latter. Then, we further devise a module named hierarchical classification paradigm (HCP), which constructs the interhierarchy relationships between coarse and fine-level representations and then exploits them to guide fine-grained feature enhancement. HCP eventually selects and boosts samples that are hard to discriminate by keeping consistency in multilevels. Our method can be easily integrated into prevailing oriented object detectors and brings consistent performance improvements across these detectors. Notably, our method combined with oriented RCNN (ORCNN) achieves 44.44% (+3.62%) on the FAIR1M and 91.0% (+6.9%) on the MAR20. Moreover, thoughtful discussions about qualitative results and rich visualizations are provided to intuitively underscore the superiority of our approach. The source code is available athttps://github.com/chengjc2019/DIMA. Jiacheng Cheng 0001, Xiwen Yao, Xuguang Yang, Xiaoxu Feng, Gong Cheng 0003, Xiankai Huang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Attention Erasing and Instance Sampling for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) trains detectors by only weak labels, aiming to save the burden of expensive bounding box-level annotations. Most previous efforts formulate WSOD as a multiple instance learning (MIL) problem, which is prone to detect discriminative object parts and miss object instances. This article proposes an attention erasing and instance sampling (AE-IS) approach to alleviate the above problems. Concretely, we first apply an attention erasing (AE) scheme to the WSOD model to hide the most discriminative region for capturing the integral extent of the object. Then, we employ an intersection-over-union (IoU)-balanced sampling component toward mining more object instances. Moreover, an instance reweighted loss (IRL) is designed to learn a larger portion of object instances, thereby further enhancing the performance of the object detector. Experimental results demonstrate that our method significantly improves the baseline approach by great margins and achieves competitive performance with the state-of-the-art algorithms on the NWPU VHR-10.v2 (72.0% mAP, 76.1% CorLoc) and DIOR (29.1% mAP, 55.9% CorLoc) datasets. The source code will be available athttps://github.com/XuanX/AE-IS. Gong Cheng 0003, Xiaoxu Feng, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Learning an Invariant and Equivariant Network for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD) is of increasing importance in the community of computer vision as its extensive applications and low manual cost. Most of the advanced WSOD approaches build upon an indefinite and quality-agnostic framework, leading to unstable and incomplete object detectors. This paper attributes these issues to the process of inconsistent learning for object variations and the unawareness of localization quality and constructs a novel end-to-end Invariant and Equivariant Network (IENet). It is implemented with a flexible multi-branch online refinement, to be naturally more comprehensive-perceptive against various objects. Specifically, IENet first performs label propagation from the predicted instances to their transformed ones in a progressive manner, achieving affine-invariant learning. Meanwhile, IENet also naturally utilizes rotation-equivariant learning as a pretext task and derives an instance-level rotation-equivariant branch to be aware of the localization quality. With affine-invariance learning and rotation-equivariant learning, IENet urges consistent and holistic feature learning for WSOD without additional annotations. On the challenging datasets of both natural scenes and aerial scenes, we substantially boost WSOD to new state-of-the-art performance. The codes have been released at: https://github.com/XiaoxFeng/IENet. Xiaoxu Feng, Xiwen Yao, Hui Shen 0005, Gong Cheng 0003, Bin Xiao 0002, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Learning to Assess Image Quality Like an ObserverabstractHuman observers are the ultimate receivers and evaluators of the image visual information and have powerful perception ability of visual quality with short-term global perception and long-term regional observation. Thus, it is natural to design an image quality assessment (IQA) computational model to act like an observer for accurately predicting the human perception of image quality. Inspired by this, here, we propose a novel observer-like network (OLN) to perform IQA by jointly considering the global glimpsing information and local scanning information. Specifically, the OLN consists of a global distortion perception (GDP) module and a local distortion observation (LDO) module. The GDP module is designed to mimic the observer's global perception of image quality through performing classification of images' distortion categories and levels. Simultaneously, to simulate the human local observation behavior, the LDO module attempts to gather the long-term regional observation information of the distorted images by continuously tracing the human scanpath in the observer-like scanning manner. By leveraging the bilinear pooling layer to collaborate the short-term global perception with the long-term regional observation, our network precisely predicts the quality scores of distorted images, such as human observers. Comprehensive experiments on the public datasets powerfully demonstrate that the proposed OLN achieves state-of-the-art performance. Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Weakly Supervised Rotation-Invariant Aerial Object Detection NetworkabstractObject rotation is among longstanding, yet still unexplored, hard issues encountered in the task of weakly supervised object detection (WSOD) from aerial images. Existing predominant WSOD approaches built on regular CNNs which are not inherently designed to tackle object rotations without corresponding constraints, thereby leading to rotation-sensitive object detector. Meanwhile, current solutions have been prone to fall into the issue with unsTable detectors, as they ignore lower-scored instances and may regard them as backgrounds. To address these issues, in this paper, we construct a novel end-to-end weakly supervised Rotation-Invariant aerial object detection Network (RINet). It is implemented with a flexible multi-branch online detector refinement, to be naturally more rotation-perceptive against oriented objects. Specifically, RINet first performs label propagating from the predicted instances to their rotated ones in a progressive refinement manner. Meanwhile, we propose to couple the predicted in-stance labels among different rotation-perceptive branches for generating rotation-consistent supervision and mean-while pursuing all possible instances. With the rotation-consistent supervisions, RINet enforces and encourages consistent yet complementary feature learning for WSOD without additional annotations and hyper-parameters. On the challenging NWPU VHR-10.v2 and DIOR datasets, extensive experiments clearly demonstrate that we significantly boost existing WSOD methods to a new state-of-the-art performance. The code will be available at: https://github.com/XiaoxFeng/RINet. Xiaoxu Feng, Xiwen Yao, Gong Cheng 0003, Junwei Han 0001 |
CVPR | 1 |
| 2022 | A spatial feature adaptive network for text detection
Qingsong Tang, Xiaoxu Feng, Xiangde Zhang |
Multim. Tools Appl. | 2 |
| 2022 | P-CNN: Part-Based Convolutional Neural Networks for Fine-Grained Visual CategorizationabstractThis paper proposes an end-to-end fine-grained visual categorization system, termed Part-based Convolutional Neural Network (P-CNN), which consists of three modules. The first module is a Squeeze-and-Excitation (SE) block, which learns to recalibrate channel-wise feature responses by emphasizing informative channels and suppressing less useful ones. The second module is a Part Localization Network (PLN) used to locate distinctive object parts, through which a bank of convolutional filters are learned as discriminative part detectors. Thus, a group of informative parts can be discovered by convolving the feature maps with each part detector. The third module is a Part Classification Network (PCN) that has two streams. The first stream classifies each individual object part into image-level categories. The second stream concatenates part features and global feature into a joint feature for the final classification. In order to learn powerful part features and boost the joint feature capability, we propose a Duplex Focal Loss used for metric learning and part classification, which focuses on training hard examples. We further merge PLN and PCN into a unified network for an end-to-end training process via a simple training technique. Comprehensive experiments and comparisons with state-of-the-art methods on three benchmark datasets demonstrate the effectiveness of our proposed method. Junwei Han 0001, Xiwen Yao, Gong Cheng 0003, Xiaoxu Feng, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Self-Guided Proposal Generation for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD) in remote sensing images remains a challenging task when learning object detectors with only image-level labels. As we know, object proposal generation plays a crucial role in WSOD. At present, the proposal generation of most existing WSOD methods mainly relies on heuristic strategies such as selective search and Edge Boxes. However, the proposals obtained by the above methods cannot well cover the entire objects, severely hindering the performance of WSOD. To address this issue, this paper proposes a Self-guided Proposal Generation approach, termed SPG. It can be easily implemented with most WSOD methods in a unified framework. To this end, we first introduce a confidence propagation approach to obtain the objectness confidence map for each image, which on the one hand highlights informative object locations, and on the other hand aggregates discriminative feature representation by combining the objectness confidence map with the deep features. Then, the proposal generation is implemented by mining informative regions as proposals on the objectness confidence map. Extensive evaluations on two challenging datasets demonstrate that our SPG significantly improves the baseline methods Online Instance Classifier Refinement (OICR) and Min-Entropy Latent Model (MELM) by large margins (for OICR: 15.86% mAP and 12.89% CorLoc gains on the NWPU VHR-10.v2 dataset, 3.65% mAP and 4.87% CorLoc gains on the DIOR dataset; for MELM: 20.51% mAP and 23.54% CorLoc gains on the NWPU VHR-10.v2 dataset, 7.11% mAP and 4.96% CorLoc gains on the DIOR dataset) and achieves the state-of-the-art results compared with existing methods. Gong Cheng 0003, Weining Chen, Xiaoxu Feng, Xiwen Yao, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SAENet: Self-Supervised Adversarial and Equivariant Network for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised object detection (WSOD) in remote sensing images (RSIs) remains a challenge when learning a subtle object detection model with only image-level annotations. Most works tend to optimize the detection model via exploiting the most contributed region, thereby to be dominated by the most discriminative part of an object. Meanwhile, these methods ignore the consistency across different spatial transformations of the same image and always label them with different classes, which introduces potential ambiguities. To tackle these challenges, we propose a unique self-supervised adversarial and equivariant network (SAENet) and aim at learning complementary and consistent visual patterns for WSOD in RSIs. To this end, an adversarial dropout–activation block is first designed to facilitate the entire object detector via adaptively hiding the discriminative parts and highlighting the instance-related regions. Besides, we further introduce a flexible self-supervised transformation equivariance mechanism on each potential instance from multiple spatial transformations to obtain spatially consistent self-supervisions. Accordingly, the obtained supervisions can be leveraged to pursue a more robust and spatially consistent object detector. Comprehensive experiments on the challenging LEarning, VIsion and Remote sensing Laboratory (LEVIR), NorthWestern Polytechnical University (NWPU) VHR-10.v2, and detection in optical RSIs (DIOR) datasets validate that SAENet outperforms the previous state-of-the-art works and achieves 46.2%, 60.7%, and 27.1% mAP, respectively. Xiaoxu Feng, Xiwen Yao, Gong Cheng 0003, Jungong Han, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Scale-Aware Detailed Matching for Few-Shot Aerial Image Semantic SegmentationabstractFew-shot semantic segmentation, aiming to segment query images with a few annotated support samples, has drawn increasing attention. Most existing few-shot methods leverage the single prototype obtained from global average pooling to represent all support information and further use the extracted prototype to segment the query images in a matching manner. Although promising results for natural images have been reported, these methods cannot be directly applied on aerial images. The main reason comes from that the extracted single support prototype can only provide a coarse guidance for matching between query and support images and could not handle the large variance of objects’ appearances and scales. To deal with these challenges on aerial images, we propose a scale-aware few-shot semantic segmentation network to perform detailed matching with multiple prototypes. More specifically, the detailed matching module is first constructed to compute the pixel-level similarity between the query features and the extracted multiple support prototypes for providing more accurate parsing guidance. Subsequently, to address the problem of scale imbalance, the scale-aware focal loss is designed to dynamically down-weight the loss assigned to large well-parsed objects and focus training on tiny hard-parsed objects. To facilitate the reproducible research on the task of few-shot semantic segmentation in aerial images, we further provide a few-shot segmentation benchmark iSAID-$5^{\mathrm {i}}$constructed from the large-scale iSAID dataset[1]. Comprehensive experiments and comparisons with the state-of-the-art few-shot segmentation methods on the iSAID-$5^{\mathrm {i}}$dataset clearly demonstrate the superiority of our proposed method. The code and dataset are available athttps://github.com/caoql98/SDM. Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | R²IPoints: Pursuing Rotation-Insensitive Point Representation for Aerial Object DetectionabstractAnchor-free aerial object detection methods have recently attracted much attention due to their simplicity and efficiency. However, the performance is still unsatisfactory due to the following two main limitations. On the one hand, the anchor-free detector employs ordinary convolution layers with axis-aligned receptive fields to extract object features, resulting in lacking internal mechanisms to handle the rotation variance. On the other hand, the detector sacrifices much semantic information to achieve faster detection, leading to the inability to deal with objects’ high inter-class similarity and intra-class diversity. To address these issues, in this paper, we present a unique anchor-free detector, termed Rotation-Insensitive Point Representation (R2IPoints), of which a set of category-aware points are employed to encode the spatial and semantic information of the arbitrary-oriented objects. Specifically, we first devise a Stacked Rotation convolution Module (SRM) to encourage the learning of rotation-insensitive point representation by adaptively modelling orientation-agnostic interdependencies over stochastically rotated features. Meanwhile, we further introduce a Class-specific Semantic enhancement Module (CSM). It performs category-aware semantic activation to recalibrate features, thus enabling the point representation to be aware of object categories. Through jointly optimizing the two proposed modules in an end-to-end manner, R2IPoints could simultaneously generate rotation-insensitive and category-aware point representation. Extensive experiments on the challenging DIOR and DOTA datasets demonstrate the superiority of the proposed method. We achieve 72.7% mAP on DIOR and 74.34% mAP on DOTA, surpassing the baseline method of +2.4% mAP and +2.49% mAP, respectively. The code is available at https://github.com/shnew/R2IPoints. Xiwen Yao, Hui Shen 0005, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | DFENet for Domain Adaptation-Based Remote Sensing Scene ClassificationabstractDomain adaptation scene classification refers to the task of scene classification where the training set (called source domain) has different distributions from the test set (called target domain). Although remarkable results have been reported, the misalignment of source and target domain features still remains a big challenge when the large intraclass variances of remote sensing images encounter the insufficient exploration of discriminative feature representations for both domains. To address this challenge, a novel domain feature enhancement network (DFENet) is proposed to adaptively enhance the discriminative ability of the learned features for dealing with the domain variances of scene classification. Specifically, an adaptive context-aware feature refinement (CAFR) module is first designed to automatically recalibrate global and local features by explicitly modeling interdependencies between the channel and spatial for each domain. Then, a multilevel adversarial dropout (MAD) module is further designed to strengthen the generalization capability of our network by adaptively reconfiguring the sparsity of the feature level and decision level in the target domain. The cooperation of CAFR module and MAD module formulates a unique DFENet that can be learned in an end-to-end manner. Comprehensive experiments show that our proposed method is better than state-of-the-art methods on Merced$\to $RSSCN7, AID$\to $RSSCN7, NWPU$\to $RSSCN7, RSSCN7$\to $Merced, RSSCN7$\to $AID, and RSSCN7$\to $NWPU datasets. Xiufei Zhang, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | TCANet: Triple Context-Aware Network for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised object detection (WSOD) in remote sensing images (RSI) plays an essential role in RSI understanding applications. Currently, predominant works are inclined to first activate the most discriminative region and then pursue the whole object by analyzing the context information of the activated region. However, the most discriminative region usually only covers a small crucial part. Besides, many same-class instances often appear in adjacent locations. In such a case, treating proposals of large spatial overlap as the same-class instances not only introduces potential ambiguities but also misleads the detection model to recognize multiple adjacent instances as one object instance. To address these challenges, a novel triple context-aware network (TCANet) is proposed to learn complementary and discriminative visual patterns for WSOD in RSIs. Specifically, a global context-aware enhancement (GCAE) module is first designed to activate the features of the whole object by capturing the global visual scene context. Then, a dual-local context residual (DLCR) module is further developed to capture the instance-level discriminative cues by leveraging the semantic discrepancy of the local context. Furthermore, an effective adaptive-weighted refinement loss is integrated into the DLCR module to reduce the ambiguities in the label propagating process. The collaboration of GCAE and DLCR formulates a unique TCANet that can be learned in an end-to-end manner. Comprehensive experiments are carried out on the challenging NWPU VHR-10.v2 and DIOR data sets. We achieve a 58.8% mAP and a 25.8% mAP on the NWPU VHR-10.v2 and DIOR data sets, respectively, which both significantly outperform the state of the arts. Xiaoxu Feng, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Automatic Weakly Supervised Object Detection From High Spatial Resolution Remote Sensing Images via Dynamic Curriculum LearningabstractIn this article, we focus on tackling the problem of weakly supervised object detection from high spatial resolution remote sensing images, which aims to learn detectors with only image-level annotations, i.e., without object location information during the training stage. Although promising results have been achieved, most approaches often fail to provide high-quality initial samples and thus are difficult to obtain optimal object detectors. To address this challenge, a dynamic curriculum learning strategy is proposed to progressively learn the object detectors by feeding training images with increasing difficulty that matches current detection ability. To this end, an entropy-based criterion is firstly designed to evaluate the difficulty for localizing objects in images. Then, an initial curriculum that ranks training images in ascending order of difficulty is generated, in which easy images are selected to provide reliable instances for learning object detectors. With the gained stronger detection ability, the subsequent order in the curriculum for retraining detectors is accordingly adjusted by promoting difficult images as easy ones. In such way, the detectors can be well prepared by training on easy images for learning from more difficult ones and thus gradually improve their detection ability more effectively. Moreover, an effective instance-aware focal loss function for detector learning is developed to alleviate the influence of positive instances of bad quality and meanwhile enhance the discriminative information of class-specific hard negative instances. Comprehensive experiments and comparisons with state-of-the-art methods on two publicly available data sets demonstrate the superiority of our proposed method. Xiwen Yao, Xiaoxu Feng, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Progressive Contextual Instance Refinement for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised learning has been attracting much attention due to its broad applications, which only requires image-level annotations to indicate whether there exist objects in the images. Currently, most of the existing weakly supervised object detection (WSOD) methods are inclined to seek only one top-scoring object instance per image from noisy proposals to train the corresponding object detector. However, more than one same-class instances often exist in the large-scale, cluttered remote sensing images. Thus, selecting only one top-scoring proposal usually results in highlighting the most representative part of an object rather than the whole object, which may cause learning a suboptimal object detector by losing much important information. To address this problem, a novel end-to-end progressive contextual instance refinement (PCIR) method is proposed to perform WSOD. Specifically, a dual-contextual instance refinement (DCIR) strategy is designed to divert the focus of the detection network from the local distinct part to the whole object and further to other potential instances by leveraging both local and global context information. Benefiting from DCIR, a progressive proposal self-pruning (PPSP) strategy is further developed to mitigate the influence of the complex background by dynamically rejecting the negative training proposals. Comprehensive experiments on the challenging NWPU VHR-10.v2 and DIOR data sets clearly demonstrate that the proposed method can significantly boost the detection accuracy compared with the state of the arts. Xiaoxu Feng, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Rotation-Invariant Latent Semantic Representation Learning for Object Detection in VHR Optical Remote Sensing ImagesabstractObject detection in very high resolution (VHR) optical remote sensing images is a fundamental yet challenging problem for the field of remote sensing image analysis. The detection performance is heavily dependent on the representation capability of the extracted features. Recently, convolutional neural networks (CNNs) have made a breakthrough for various applications in nature images. However, it is problematic to directly apply CNN to perform object detection in VHR optical remote sensing images due to the problem of object rotation variations. To address this issue, a novel rotation invariant probabilistic Latent Semantic Analysis (RI-pLSA) model is proposed to learn latent semantic representations for object detection. This is achieved by imposing a rotation-invariant regularization term on the objective function of pLSA to enforce the learned representation from all rotations of the same sample to be as consistent as possible. Additionally, the proposed RI-pLSA model takes the CNN features as input, which generates more powerful semantic representation for object detection. Comprehensive experiments on a publicly available ten-class object detection dataset demonstrate the superiority and effectiveness of our method compared with state-of-the-arts. Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002 |
IGARSS | 2 |