Xiwen Yao

dblp:163/0541 · DBLP profile ↗
← Back
72ranked-venue papers
11as first author
57since 2021 · last 2026
0000-0002-7466-7428ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 7 first-author · 25 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 7 · 7 since 2021
YearPublicationVenuePosition
2026 UQ-ViT: Harmonizing Extreme Activations with Hardware-Friendly Uniform Quantization in Vision Transformers
abstract
Post-Training Quantization enables efficient Vision Transformer (ViTs) deployment with a small calibration data, and its prevalent use of uniform quantization harnesses AI accelerator matrix cores for high-speed inference. However, the application of uniform quantization is fundamentally challenged by the extreme non-uniformity of activation distributions.Specifically, the power-law nature of post-Softmax attention scores and the significant inter-channel variance in post-GELU activations create a dilemma for conventional quantization, as it struggles to preserve critical high-magnitude values without sacrificing overall precision. To resolve this core conflict, we introduce UQ-ViT (Uniform Quantization for Vision Transformers), a novel uniform quantization framework designed to reconcile high precision with hardware efficiency. Central to UQ-ViT are two operators: Dynamic Elimination of Maximum (DeMax) and Normalization Quantization (NormQuant). DeMax is a quantization operator for post-Softmax attention scores that utilizes uniform quantization. It dynamically eliminates and preserves dominant values, effectively mitigating quantization loss from the extreme values in the power-law distribution. NormQuant utilizes a per-channel quantization strategy during quantization and reverts to a per-tensor format for dequantization, achieving both high accuracy and computational efficiency. Crucially, it is applicable to any linear layer, enabling effective quantization of post-GELU activations in ViTs. Through extensive experiments on various ViTs and vision tasks, including image classification, object detection, and instance segmentation, we demonstrate that our proposed approach outperforms existing methods, achieving superior accuracy while ensuring hardware friendliness.
Tao Jiang 0002, Yucheng Jiang, Xiwen Yao, Gong Cheng 0003, Junwei Han 0001
AAAI3
2026 From Human Pragmatic Language Skills to Conversational Agent Design: A Systematic Review of Transfer Strategies
abstract
While conversational agents’ (CAs) semantic and syntactic capabilities have advanced, their pragmatic skills, using language appropriately in context, have emerged as a critical focus in practical applications. Hence, scholars integrate conversational skills derived from human-human interaction into CA designs. However, existing research mainly adopts an empirical approach and focuses on specific CA deployment, making it challenging to identify overarching patterns or develop a comprehensive methodology for transferring human pragmatic skills to CA design. Thus, we conducted a systematic review of 85 studies from primary databases (e.g., ACM, IEEE, etc.), focusing on designing CAs with human-derived conversational skills. We identified skill categories (verbal, paralinguistic, nonverbal), transfer strategies (from dialog data, theories, and via co-design), implementations, and evaluation metrics. We consolidated these insights into a four-stage design process: human skill exploration, definition, transfer, and iterative evaluation. Future research can leverage this to design CAs that achieve conversational goals through contextually appropriate language use.
Jiaxiong Hu, Xiwen Yao, Danxuan Liang, Dongjie Yang, Dingdong Liu, Junze Li, Yuanhao Zhang, Xiaojuan Ma
CHI2
2026 GestuProp: 3D Virtual Reality Prop Generation with Co-Speech Gestures
abstract
Virtual Reality (VR) has been widely adopted in domains such as gaming, education, and healthcare, where 3D props play a central role in enabling immersive interaction. With the advancement of generative AI, 3D props can now be created rapidly; however, little research has explored how gestures and speech can be integrated to support prop generation. To address this gap, we introduce GestuProp, a VR prop generation system driven by co-speech gestures. Building on a formative study with 30 participants, we proposed a gesture design space and developed the VR system GestuProp. We then conducted a user study with 14 participants, which showed that GestuProp demonstrates good usability and favorable user experiences, while also revealing how object categories influence gesture use and interaction. These findings highlight the potential of gesture–speech synergy to advance prop generation in VR.
Zhihao Yao 0004, Xiwen Yao, Haowei Xiong, Yuanling Feng, Qirui Sun, Yijie Guo, Haipeng Mi
CHI2
2026 LumiBite: An In-the-Wild Technology Probe Exploring Personalized bottom-up Lighting Lunchbox for Enhanced Dining Experiences
abstract
Food perception is a multisensory experience shaped by environmental cues such as ambient lighting. Previous studies have demonstrated that lighting can impact how satisfied we feel while dining. However, many of these studies were conducted in controlled laboratory settings with standardized meals, overlooking how lighting interacts with personal dietary choices and diverse dining contexts. This paper introduces LumiBite, a portable lighting system integrated into a lunchbox, designed for dining environments to enable personalized lighting adjustments during meals. Through a seven-day in-the-wild study with six participants, where they freely chose when, what, and where to eat, we explored the feasibility of deploying LumiBite and investigated how user agency in customizing lighting settings impacts dining satisfaction, sensory perception and dietary behaviors. Our findings demonstrate that LumiBite not only enhances food aesthetics but also shifts users from passive consumers to active meal curators. The study highlights key challenges, including cultural dining practices and ambient light interference, and offers actionable design principles for creating context-aware, culturally sensitive dining technologies.
Haiqing Xu 0001, Xiwen Yao, Sixuan Wu, Jung Hyun Bae, Zhifan Guo, Dian Lv, Zhihao Yao 0004, HyunJoo Oh 0001, Alexander Travis Adams
TEI2
2026 Relaxed Knowledge Distillation
Xiwen Yao, Xuguang Yang, Gong Cheng 0003, Junwei Han 0001
Int. J. Comput. Vis.2
2026 Beyond Support Samples: Incorporating Unlabeled Queries for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) often struggles with the intra-class diversity issue between query and support images, caused by the category-biased information provided by limited annotated support images for matching objects. While increasing the number of annotated support images could mitigate this bias, it is impractical within the few-shot learning framework. Therefore, our proposed Unlabeled Query Integration Few-Shot Segmentation (UQI-FSS) tackles this challenge by incorporating unlabeled query images into the learning paradigm. This approach aims to achieve a more comprehensive category representation, which is essential to enhance segmentation accuracy in various scenarios. However, integrating unlabeled query images directly requires careful management to prevent the dilution of vital information from the annotated support set. To address this issue, we present an Unlabeled Query Integration Network (UQINet), which adaptively extracts beneficial and suppresses detrimental information from the unlabeled query images. Specifically, we first introduce an Information Bridging Module to close the gap between support and unlabeled query features, generating a pseudo-support set enriched with additional category data. Next, we introduce a Query Fusion Module to incorporate query information from both prototype and pixel levels into the pseudo-support features, thus improving their adaptability to the query. Finally, we propose an Adaptive Selection Module to select effective category information and combine the pseudo-support features, thereby activating target objects for precise segmentation prediction. Experimental results show considerable performance improvements over previous methods on various FSS benchmarks. The application of this method to four challenging scenarios further underscores its versatility and practical value.
Yuanwei Liu, Nian Liu 0002, Tao Jiang 0002, Xiwen Yao, Junwei Han 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Dual-perspective filter pruning via diversity and independence collaboration
Chenyang Gao, Qinglong Cao, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003
Pattern Recognit.3
2026 Domain adaptation for remote sensing image semantic segmentation with prototype-driven domain disentangle alignment
Xiufei Zhang, Yuanwei Liu, Xiaoliang Qian, Gong Cheng 0003, Xiwen Yao
Pattern Recognit.5
2026 Prototype Decoupled Knowledge Distillation
Yuanwei Liu, Nian Liu 0002, Xiwen Yao, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 TPTAF: Task-Prior Tripartite Attention for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to integrate complementary information to produce informative fused images for visual perception and downstream tasks. However, mainstream methods often focus on fusion reconstruction, while losses from different tasks may introduce competing optimization demands, making it difficult to balance visual quality and detection performance. To address this issue, we propose task-prior tripartite attention for infrared and visible image fusion (TPTAF), which introduces detection semantics as task-prior guidance for representation-level cross-modal interaction. TPTAF employs a decoupled encoder to organize infrared and visible features into structural and discriminative spaces, enabling cross-modal layout consistency and modality-specific detail preservation to be modeled with different roles. Meanwhile, detection semantics extracted from hybrid infrared-visible features are transformed into task priors and integrated with the decoupled representations through tripartite attention. In this interaction, structural cues stabilize the fusion layout, while task-prior guidance regulates the selection of discriminative details toward target-related evidence. For joint optimization, uncertainty-weighted learning further balances fusion and detection losses, reducing the dependence on manually assigned loss weights. Experiments on M3FD, RoadScene, AVMS, and MSRS demonstrate that TPTAF maintains stable fusion quality across different data distributions while improving the utility of fused images for downstream object detection and semantic segmentation evaluation. The code will be made available at https://github.com/Duxeno/TPTAF.
Xuyang Du, Xiwen Yao, Ankang Zang, Gong Cheng 0003
IEEE Trans. Image Process.2
2025 A Card-based Co-Design Toolkit for Exploring Smart Material Applications with Multiple Stakeholders: A Case Study on Automotive Interior Design
abstract
Smart materials have garnered significant attention in both academia and industry, yet identifying pragmatically impactful applications still requires contributions from multiple stakeholders, including researchers, designers, and industry professionals.Although previous research has explored novel technical approaches or user-centered applications of smart materials, this study focuses on how to stimulate effective dialogue among stakeholders to explore impactful smart material applications.
Tianyu Yu 0001, Yao Lu 0038, Kejin Yu, Xiwen Yao, Wenjing Deng, Xueqing Li 0005, Yue Yang 0005, Yijie Guo, Guanhong Liu, Haipeng Mi
Conference on Designing Interactive Systems5
2025 Tactile Data Comics: Combining Step-by-step Presentation of Tactile Graphics with Verbal Narration for the Blind and Visually Impaired
Ruoting Sun, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu 0003
ASSETS4
2025 Exploring the Design of LLM-based Agent in Enhancing Self-disclosure Among the Older Adults
Yijie Guo, Ruhan Wang, Zhenhan Huang, Tongtong Jin, Xiwen Yao, Yuanling Feng, Haipeng Mi
CHI5
2025 Miau-BOT: A Cat Paw Companion Robot for Tactile Emotional Support When Driving
abstract
This study presents Miau-BOT, an emotional companion robot that improves the driver's experience and enhances driving safety through dynamic tactile feedback. Miau-BOT assesses the driver's emotional and physical state in real-time using physiological data, and based on the driver's level of fatigue and emotional changes, it provides personalized emotional support by inflating and deflating cat paw pads and claws at different frequencies. Through emotional transfer, Miau-BOT establishes an emotional connection between humans and robots in the driving context, demonstrating the potential of emotional design to enhance both driving safety and comfort.
Xinjun Gong, Yueyao Jiang, Yut Wong, Xiwen Yao
HRI5
2025 Not All Tokens Matter All The Time: Dynamic Token Aggregation Towards Efficient Detection Transformers
abstract
The substantial computational demands of detection transformers (DETRs) hinder their deployment in resource-constrained scenarios, with the encoder consistently emerging as a critical bottleneck. A promising solution lies in reducing token redundancy within the encoder. However, existing methods perform static sparsification while ignoring the varying importance of tokens across different levels and encoder blocks for object detection, leading to suboptimal sparsification and performance degradation. In this paper, we propose **Dynamic DETR** (**Dynamic** token aggregation for **DE**tection **TR**ansformers), a novel strategy that leverages inherent importance distribution to control token density and performs multi-level token sparsification. Within each stage, we apply a proximal aggregation paradigm for low-level tokens to maintain spatial integrity, and a holistic strategy for high-level tokens to capture broader contextual information. Furthermore, we propose center-distance regularization to align the distribution of tokens throughout the sparsification process, thereby facilitating the representation consistency and effectively preserving critical object-specific patterns. Extensive experiments on canonical DETR models demonstrate that Dynamic DETR is broadly applicable across various models and consistently outperforms existing token sparsification methods.
Jiacheng Cheng 0001, Xiwen Yao, Junwei Han 0001
ICML2
2025 Bridge the Intra-Class Gap: K-Shot Multi-Scale Intermediate Prototype Mining Transformer for Few-Shot Semantic Segmentation
abstract
Few-shot segmentation (FSS) aims to accurately segment target objects in a query image using only a limited number of annotated support images. Existing approaches typically follow a paradigm that directly leverages category information from the support set to identify target objects in the query. However, these methods often ignore the category information gap between query and support images, leading to suboptimal performance when faced with images containing objects exhibiting significant intra-class diversity. To address this issue, we propose a novel framework that introduces intermediate prototypes to capture both deterministic information from the support images and adaptive knowledge from the query at multiple scales. Our framework, named the K-shot Multi-scale Intermediate Prototype Mining Transformer (KMIPMT), is based on the Transformer architecture and learns intermediate prototypes in an iterative manner, where each KMIPMT layer propagates category information from both K-shot support features and multi-scale query features to intermediate prototypes. This information is then utilized to activate the query feature map. Through repeated iterations, both intermediate prototypes and the query feature are progressively enhanced, and the final refined query feature is used for generating precise segmentation predictions. Despite its simplicity, our method achieves remarkable performance gains on standard benchmarks, including PASCAL-$5^{i}$5i, COCO-$20^{i}$20i, and FSS-1000, setting new state-of-the-art results. Furthermore, we explore several practical and challenging extensions of our method, including 3D point cloud FSS, zero-shot segmentation, weak-label FSS, and cross-domain FSS. These extensions showcase the versatility and effectiveness of our proposed KMIPMT framework across different domains and scenarios.
Yuanwei Liu, Nian Liu 0002, Tao Jiang 0002, Xiwen Yao, Rao Muhammad Anwer, Hisham Cholakkal, Junwei Han 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 NTRENet++: Unleashing the Power of Non-Target Knowledge for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment the target object under the condition of a few annotated samples. However, current studies on FSS primarily concentrate on extracting information related to the object, resulting in inadequate identification of ambiguous regions, particularly in non-target areas, including the background (BG) and Distracting Objects (DOs). Intuitively, to alleviate this problem, we propose a novel framework, namely NTRENet++, to explicitly mine and eliminate BG and DO regions in the query. First, we introduce a BG Mining Module (BGMM) to extract BG information and generate a comprehensive BG prototype from all images. For this purpose, a BG mining loss is formulated to supervise the learning of BGMM, utilizing only the known target object segmentation ground truth. Subsequently, based on this BG prototype, we employ a BG Eliminating Module to filter out the BG information from the query and obtain a BG-free result. Following this, the target information is utilized in the target matching module to generate the initial segmentation result. Finally, a DO Eliminating Module is proposed to further mine and eliminate DO regions, based on which we can obtain a BG and DO-free target object segmentation result. Moreover, we present a prototypical-pixel contrastive learning algorithm to enhance the model’s capability to differentiate the target object from DOs. Extensive experiments conducted on both PASCAL-5i and COCO-20i datasets demonstrate the effectiveness of our approach despite its simplicity. Additionally, we extend our method to the few-shot video object segmentation task and achieve improved performance on a baseline model, demonstrating its generalization ability. Code is available athttps://github.com/LIUYUANWEI98/NTRENet++.
Yuanwei Liu, Nian Liu 0002, Hisham Cholakkal, Rao Muhammad Anwer, Xiwen Yao, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Incorporating Multiscale Context and Task-Consistent Focal Loss into Oriented Object Detection
abstract
Oriented object detection (OOD) in remote sensing images (RSIs) aims to precisely localize and identify objects with arbitrary orientations. Two-stage OOD methods attract lots of interest due to their superior accuracy, however, they still face two major problems. First of all, the misclassification problem frequently occurs because the majority of classification strategies solely relies on the features of proposals. Secondly, most of loss functions cannot simultaneously concentrate on hard samples and boost the consistency between identification and localization, which restricts the further improvement of OOD models. To address the first problem, the multi-scale context (MSC) is incorporated into a two-stage OOD model in this paper. Specifically,Ncontextual branches are added to predict the class confidence score (CCS) of each proposal and itsNenlarged proposals which contain the MSC, and the final CCS of each proposal is determined by the mean value of aboveN+ 1 CCSs. To tackle the second problem, a task-consistent focal (TF) loss is proposed. The TF loss employs the difficulty of localization as the weight of classification loss, and the difficulty of identification is used as the weight of regression loss. Concentrating on hard samples and synchronous optimization of classification and regression can be achieved by minimizing the TF loss. The ablation studies show the validity of MSC, TF and their combination. The comparison with popular OOD models demonstrates the superior performance of our model on the DOTA and DIOR-R datasets. The source code can be obtained from https://github.com/qxlzengli/MSC-TF.
Xiaoliang Qian, Qingqing Jian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.4
2025 IPS-YOLO: Iterative Pseudo-Fully Supervised Training of YOLO for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images only requires image-level labels, greatly reducing the cost of manual annotations. Recently, the pseudo-fully supervised object detection (pseudo-FSOD) models give better performance than traditional WSOD models based on multiple instance learning. However, the existing pseudo-FSOD models still have two problems that need to be addressed. Firstly, existing models tend to focus on the salient parts of object rather than the whole object. Secondly, the pseudo ground truth (PGT) cannot be continuously improved during the training process. For the first problem, a filtering and weighted synthesis guided by category confidence score (FWSC) strategy is proposed to generate PGT. The FWSC strategy firstly removes the instances with low category confidence score (CCS), then, to cover the whole object as much as possible, any two instances are weighted and synthesized according to their CCSs if they have very high spatial overlap, otherwise, the non-maximum suppression (NMS) operation is conducted. For the second problem, an iterative refinement (IR) scheme of PGT is proposed. Specifically, the PGT instances produced by the FWSC strategy are firstly used to train a YOLO model, then a filtering and weighted synthesis guided by IoU (FWSI) strategy employs the detection results inferred from the trained YOLO model to refine the PGT instances, and above two steps can be repeated multiple times. Furthermore, three improvement strategies are proposed to enhance the traditional baseline WSOD model in proposals generation, the selection of augmented samples, and the definition of pseudo-labels, respectively. The ablation studies demonstrate the effectiveness of FWSC strategy, IR scheme of PGT, and the three improvement strategies of baseline model. The comparisons with popular WSOD models show that our model gives the best results on the NWPU VHR-10. v2 and DIOR datasets.The source codes have been released at https://github.com/qxlzengli/IPS-YOLO.
Xiaoliang Qian, Baihui Zhang, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.5
2025 Cross-Modality Domain Adaptation Based on Semantic Graph Learning: From Optical to SAR Images
abstract
Synthetic aperture radar (SAR) imaging provides a distinct advantage in scene understanding due to its capability for all-weather data acquisition. However, in comparison to easily annotated optical remote sensing images, the lower imaging quality of SAR images presents significant challenges in obtaining manually annotated training data, which poses substantial issues for SAR image analysis. In this paper, we employ the domain adaptation (DA) that leverages labeled optical images to better understand unlabeled SAR images. Global feature alignment as a method for DA has demonstrated effectiveness in transferring knowledge, yet it faces challenges in cross-modality adaptation from optical remote sensing to SAR images due to their differing imaging mechanisms. With distinct visual features between optical and SAR images, the semantic dependency is difficult to construct, which results in low-quality pseudo-label assignment for SAR images. To address the above issue, we propose a semantic graph learning framework to comprehensively align the global features of optical remote sensing and SAR images by modeling the cross-modality semantics and generating high-quality pseudo-labels. It can be applied for SAR scene classification and object detection when only optical remote sensing images are labeled. Specifically, a cross-modality semantic graph alignment (CSGA) module is constructed to model and align the second-order semantic dependencies by aggregating cross-modality visual semantic information. Then, an uncertainty-based robust pseudo-label generation (URPG) module is designed to generate pseudo-labels for effective semantic alignment and self-training by modeling the uncertainty of pseudo-labels for each SAR image. Comprehensive experiments show that our proposed method outperforms the state-of-the-art methods on scene classification (NWPU-RESISC45→WHU-SAR6, MLRSNet→NWPU-SAR6, MLRSNet→NWPU-SAR6, and NWPU-RESISC45→NWPU-SAR6) and object detection (MASATI-ship→SSDD, MVSRD→SARDet-vehicle, and DIOR-airplane→SAR-airplane) tasks. The code and datasets are publicly accessible at https://github.com/XZhang878/SGLF.
Xiufei Zhang, Zhongling Huang, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Bidirectional Reciprocative Information Communication for Few-Shot Semantic Segmentation
abstract
Existing few-shot semantic segmentation methods typically rely on a one-way flow of category information from support to query, ignoring the impact of intra-class diversity. To address this, drawing inspiration from cybernetics, we introduce a Query Feedback Branch (QFB) to propagate query information back to support, generating a query-related support prototype that is more aligned with the query. Subsequently, a Query Amplifier Branch (QAB) is employed to amplify target objects in the query using the acquired support prototype. To further improve the model, we propose a Query Rectification Module (QRM), which utilizes the prediction disparity in the query before and after support activation to identify challenging positive and negative samples from ambiguous regions for query self-rectification. Furthermore, we integrate the QFB, QAB, and QRM into a feedback and rectification layer and incorporate it into an iterative pipeline. This configuration enables the progressive enhancement of bidirectional reciprocative flow of category information between query and support, effectively providing query-adaptive support information and addressing the intra-class diversity problem. Extensive experiments conducted on both PASCAL-5i and COCO-20i datasets validate the effectiveness of our approach. The code is available at https://github.com/LIUYUANWEI98/IFRNet .
Yuanwei Liu, Junwei Han 0001, Xiwen Yao, Salman Khan 0001, Hisham Cholakkal, Rao Muhammad Anwer, Nian Liu 0002, Fahad Shahbaz Khan
ICML3
2024 Oriented R-CNN and Beyond
Xingxing Xie, Gong Cheng 0003, Jiabao Wang 0005, Ke Li 0005, Xiwen Yao, Junwei Han 0001
Int. J. Comput. Vis.5
2024 DIMA: Digging Into Multigranular Archetype for Fine-Grained Object Detection
abstract
Fine-grained remote sensing object detection aims at precisely locating objects and determining the fine-level categories. This task is exceptionally challenging due to the substantial interclass similarity, presenting difficulties in capturing discriminative features. We attribute this to the absence of essential information that can serve as supervision for the learning. This involves comprehensive visual patterns of objects and intrinsic relationships of multigranular features. In this article, we propose a novel scheme dubbed as digging into the multigranular archetype (DIMA) for fine-grained remote sensing object detection. In detail, we first design a simple yet effective frequency-aware representation supplement (FARS) mechanism learning from original images and their auxiliary frequency counterparts simultaneously. The FARS introduces high- and low-frequency representations to reinforce a range of visual cues, such as particular regions associated with the former and contours of objects related to the latter. Then, we further devise a module named hierarchical classification paradigm (HCP), which constructs the interhierarchy relationships between coarse and fine-level representations and then exploits them to guide fine-grained feature enhancement. HCP eventually selects and boosts samples that are hard to discriminate by keeping consistency in multilevels. Our method can be easily integrated into prevailing oriented object detectors and brings consistent performance improvements across these detectors. Notably, our method combined with oriented RCNN (ORCNN) achieves 44.44% (+3.62%) on the FAIR1M and 91.0% (+6.9%) on the MAR20. Moreover, thoughtful discussions about qualitative results and rich visualizations are provided to intuitively underscore the superiority of our approach. The source code is available athttps://github.com/chengjc2019/DIMA.
Jiacheng Cheng 0001, Xiwen Yao, Xuguang Yang, Xiaoxu Feng, Gong Cheng 0003, Xiankai Huang, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Complete and Invariant Instance Classifier Refinement for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images is used to detect high-value objects by utilizing image-level labels. However, the current models still have two problems. Firstly, the misclassification of neighboring instances is easily occurred because the one-hot label is assigned to all of seed instances and their neighboring instances. Secondly, the supervisory information of each instance classifier refinement (ICR) branch is generated from the predicted class score of upper ICR branch rather than the real label, thus the prediction mistake of each ICR branch will be accumulated with the propagation of supervisory information. To address the first problem, a complete definition of pseudo soft label (CPSL) of instances is proposed to directly train each ICR branch, where the CPSL of seed instances is defined according to the predicted class scores of upper ICR branch, and the CPSL of other instances are determined by the spatial distance weighted feature similarity between them and seed instances. To handle the second problem, an invariant multiple instance learning (IMIL) scheme is proposed to indirectly train each ICR branch by using the real image-level labels. Furthermore, the affine transformations of original image are incorporated into the baseline model to enhance the invariance of our model. The ablation studies verify the effectiveness of CPSL, IMIL and their combination. The quantitative comparisons with popular methods show that the 73.63% (31.08%) mAP and 79.88% (57.52%) CorLoc of our method is the best on the NWPU VHR-10.v2 (DIOR) dataset, and the qualitative comparisons intuitively demonstrate it again.
Xiaoliang Qian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.4
2024 Attention Erasing and Instance Sampling for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) trains detectors by only weak labels, aiming to save the burden of expensive bounding box-level annotations. Most previous efforts formulate WSOD as a multiple instance learning (MIL) problem, which is prone to detect discriminative object parts and miss object instances. This article proposes an attention erasing and instance sampling (AE-IS) approach to alleviate the above problems. Concretely, we first apply an attention erasing (AE) scheme to the WSOD model to hide the most discriminative region for capturing the integral extent of the object. Then, we employ an intersection-over-union (IoU)-balanced sampling component toward mining more object instances. Moreover, an instance reweighted loss (IRL) is designed to learn a larger portion of object instances, thereby further enhancing the performance of the object detector. Experimental results demonstrate that our method significantly improves the baseline approach by great margins and achieves competitive performance with the state-of-the-art algorithms on the NWPU VHR-10.v2 (72.0% mAP, 76.1% CorLoc) and DIOR (29.1% mAP, 55.9% CorLoc) datasets. The source code will be available athttps://github.com/XuanX/AE-IS.
Gong Cheng 0003, Xiaoxu Feng, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Learning Complementary Spatial-Temporal Transformer for Video Salient Object Detection
abstract
Besides combining appearance and motion information, another crucial factor for video salient object detection (VSOD) is to mine spatial-temporal (ST) knowledge, including complementary long-short temporal cues and global-local spatial context from neighboring frames. However, the existing methods only explored part of them and ignored their complementarity. In this article, we propose a novel complementary ST transformer (CoSTFormer) for VSOD, which has a short-global branch and a long-local branch to aggregate complementary ST contexts. The former integrates the global context from the neighboring two frames using dense pairwise attention, while the latter is designed to fuse long-term temporal information from more consecutive frames with local attention windows. In this way, we decompose the ST context into a short-global part and a long-local part and leverage the powerful transformer to model the context relationship and learn their complementarity. To solve the contradiction between local window attention and object motion, we propose a novel flow-guided window attention (FGWA) mechanism to align the attention windows with object and camera movements. Furthermore, we deploy CoSTFormer on fused appearance and motion features, thus enabling the effective combination of all three VSOD factors. Besides, we present a pseudo video generation method to synthesize sufficient video clips from static images for training ST saliency models. Extensive experiments have verified the effectiveness of our method and illustrated that we achieve new state-of-the-art results on several benchmark datasets.
Nian Liu 0002, Kepan Nan, Wangbo Zhao, Xiwen Yao, Junwei Han 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Multi-grained Temporal Prototype Learning for Few-shot Video Object Segmentation
abstract
Few-Shot Video Object Segmentation (FSVOS) aims to segment objects in a query video with the same category defined by a few annotated support images. However, this task was seldom explored. In this work, based on IPMT, a state-of-the-art few-shot image segmentation method that combines external support guidance information with adaptive query guidance cues, we propose to leverage multi-grained temporal guidance information for handling the temporal correlation nature of video data. We decompose the query video information into a clip prototype and a memory prototype for capturing local and long-term internal temporal guidance, respectively. Frame prototypes are further used for each frame independently to handle fine-grained adaptive guidance and enable bidirectional clip-frame prototype communication. To reduce the influence of noisy memory, we propose to leverage the structural similarity relation among different predicted regions and the support for selecting reliable memory frames. Furthermore, a new segmentation loss is also proposed to enhance the category discriminability of the learned prototypes. Experimental results demonstrate that our proposed video IPMT model significantly outperforms previous models on two benchmark datasets. Code is available at https://github.com/nankepan/VIPMT.
Nian Liu 0002, Kepan Nan, Wangbo Zhao, Yuanwei Liu, Xiwen Yao, Salman Khan 0001, Hisham Cholakkal, Rao Muhammad Anwer, Junwei Han 0001, Fahad Shahbaz Khan
ICCV5
2023 Towards Large-Scale Small Object Detection: Survey and Benchmarks
abstract
With the rise of deep convolutional neural networks, object detection has achieved prominent advances in past years. However, such prosperity could not camouflage the unsatisfactory situation of Small Object Detection (SOD), one of the notoriously challenging tasks in computer vision, owing to the poor visual appearance and noisy representation caused by the intrinsic structure of small targets. In addition, large-scale dataset for benchmarking small object detection methods remains a bottleneck. In this paper, we first conduct a thorough review of small object detection. Then, to catalyze the development of SOD, we construct two large-scale Small Object Detection dAtasets (SODA), SODA-D and SODA-A, which focus on the Driving and Aerial scenarios respectively. SODA-D includes 24828 high-quality traffic images and 278433 instances of nine categories. For SODA-A, we harvest 2513 high resolution aerial images and annotate 872069 instances over nine classes. The proposed datasets, as we know, are the first-ever attempt to large-scale benchmarks with a vast collection of exhaustively annotated instances tailored for multi-category SOD. Finally, we evaluate the performance of mainstream methods on SODA. We expect the released benchmarks could facilitate the development of SOD and spawn more breakthroughs in this field.
Gong Cheng 0003, Xiwen Yao, Kebing Yan, Xingxing Xie, Junwei Han 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Learning an Invariant and Equivariant Network for Weakly Supervised Object Detection
abstract
Weakly Supervised Object Detection (WSOD) is of increasing importance in the community of computer vision as its extensive applications and low manual cost. Most of the advanced WSOD approaches build upon an indefinite and quality-agnostic framework, leading to unstable and incomplete object detectors. This paper attributes these issues to the process of inconsistent learning for object variations and the unawareness of localization quality and constructs a novel end-to-end Invariant and Equivariant Network (IENet). It is implemented with a flexible multi-branch online refinement, to be naturally more comprehensive-perceptive against various objects. Specifically, IENet first performs label propagation from the predicted instances to their transformed ones in a progressive manner, achieving affine-invariant learning. Meanwhile, IENet also naturally utilizes rotation-equivariant learning as a pretext task and derives an instance-level rotation-equivariant branch to be aware of the localization quality. With affine-invariance learning and rotation-equivariant learning, IENet urges consistent and holistic feature learning for WSOD without additional annotations. On the challenging datasets of both natural scenes and aerial scenes, we substantially boost WSOD to new state-of-the-art performance. The codes have been released at: https://github.com/XiaoxFeng/IENet.
Xiaoxu Feng, Xiwen Yao, Hui Shen 0005, Gong Cheng 0003, Bin Xiao 0002, Junwei Han 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Uncertainty Exploration: Toward Explainable SAR Target Detection
abstract
Deep learning-based synthetic aperture radar (SAR) target detection has been developed for years, with many advanced methods proposed to achieve higher indicators of accuracy and speed. In spite of this, the current deep detectors cannot express the reliability and interpretation in trusting the predictions, which are crucial especially for those ordinary users without much expertise in understanding SAR images. To achieve explainable SAR target detection, it is necessary to answer the following questions: how much should we trust and why cannot we trust the results. With this purpose, we explore the uncertainty for SAR target detection in this article by quantifying the model uncertainty and explaining the ignorance of the detector. First, the Bayesian deep detectors (BDDs) are constructed for uncertainty quantification, answering how much to trust the classification and localization result. Second, an occlusion-based explanation method (U-RISE) for BDD is proposed to account for the SAR scattering features that cause uncertainty or promote trustworthiness. We introduce the probability-based detection quality (PDQ) and multielement decision space for evaluation besides the traditional metrics. The experimental results show that the proposed BDD outperforms the counterpart frequentist object detector, and the output probabilistic results successfully convey the model uncertainty and contribute to more comprehensive decision-making. Furthermore, the proposed U-RISE generates an attribution map with intuitive explanations to reveal the complex scattering phenomena about which BDD is uncertain. We deem our work will facilitate explainable and trustworthy modeling in the field of SAR image understanding and increase user comprehension of model decisions.
Zhongling Huang, Xiwen Yao, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Robust Few-Shot Aerial Image Object Detection via Unbiased Proposals Filtration
abstract
Few-shot aerial image object detection aims to rapidly detect object instances of novel category in aerial images by using few labeled samples. However, due to the complex background of aerial images, few labeled samples of novel categories, and the model trained with the few-shot learning paradigm is biased towards the base categories, it greatly increases the difficulty of identifying foreground objects of novel categories. In addition to this, tiny object detection is always a hot potato in aerial image object detection, and it is even more difficult for few-shot object detection. To this end, we propose a Few-shot aerial image object detection with Confidence-Iou collaborative proposal filtration and Tiny object constraint loss (FsCIT). Specifically, we first introduce a new confidence-iou collaborative proposal filtration scheme to RPN, which combines the unbiased IoU scores between the two bounding boxes with foreground-background confidence scores to filter redundant region proposals and rescue more foreground proposals for the novel categories in RPN. Then, we design a new tiny object loss constraint term to attempt at overcoming the challenge of tiny object detection in few-shot aerial image object detection. This term considers the central point distance, the size of ground-truth bounding boxes, and the distances between the four edges of the ground-truth bounding box and the predicted bounding box. Experiments on DIOR, AI-TOD and HRRSD datasets show that FsCIT is effective and can improve the performance of few-shot aerial image object detection.
Lingjun Li, Xiwen Yao, Dongpao Hong, Gong Cheng 0003, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Mining High-Quality Pseudoinstance Soft Labels for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection in remote sensing image (RSI) is still a challenge because of the lack of instance-level labels, and many existing methods have two problems. Firstly, most of the existing methods usually mine the pseudo ground truth (PGT) instances solely relying on proposal class scores (PCS). Actually, the reliability of PCS is not enough because of the bird’s eye view imaging and large-scale chaotic background of RSIs, and the instances with high PCS incline to cover the discriminative region rather than the whole object. Secondly, the existing methods assign a one-hot label to each instance, and the label of PGT instance is copied to its neighbor instances, which induces the misclassification problem to some extent. Actually, the probability that the neighbor instances contain the object with the same category is smaller than the PGT instance. For the first problem, the proposal quality score (PQS) is proposed for mining high-quality PGT instances, which contains PCS and dual-context projection score (DCPS). The DCPS is calculated through semantic segmentation, and is employed to measure the completeness that each proposal covers an object. For the second problem, a pseudo soft label assignment (PSLA) strategy is proposed to assign more precise soft label for each instance, where the soft label is determined by the spatial distance between each instance and its nearest PGT instance. The ablation study validates the effectiveness of the PQS and PSLA. The comprehensive comparisons with other WSOD methods on three popular benchmarks show the excellent performance of our method.
Xiaoliang Qian, Yu Huo 0001, Gong Cheng 0003, Chenyang Gao, Xiwen Yao, Wei Wang 0245
IEEE Trans. Geosci. Remote. Sens.5
2023 Building a Bridge of Bounding Box Regression Between Oriented and Horizontal Object Detection in Remote Sensing Images
abstract
Oriented object detection (OOD) aims to precisely detect the objects with arbitrary orientation in remote sensing images. Up to now, most of bounding box regression (BBR) losses for OOD are transferred from horizontal object detection (HOD) methods, however, the transferring requires lots of professional knowledge and experiences for designers, consequently, many excellent BBR losses for HOD have not been transferred to OOD. To accelerate the research progress of BBR loss for OOD, a unified transferring strategy (UTS) is proposed to facilitate the transferring of BBR loss from HOD to OOD. The UTS proposes that the BBR of oriented bounding box (OBB) can be converted into the joint BBR of its horizontal smallest enclosing rectangle (HSER) and two offsets, so the BBR loss in HOD can be easily transferred to OOD by using HSER as a bridge. Following the UTS, a BBR loss named Rotated-IoU (RIoU) loss is designed for OOD by transferring an advanced BBR loss in HOD, which can be considered as an example to show how to transfer. On the basis of RIoU loss, a focal rotated-IoU (FRIoU) loss is proposed to assign larger weights to hard samples in the BBR. The comparisons with other BBR losses show that the RIoU and FRIoU losses can give better performance. The ablation study shows that giving more attention to hard samples in BBR is effective. The comparisons with many advanced methods demonstrate that the combinations of baseline methods and FRIoU loss achieve state-of-the-art performance on the DOTA and DIOR-R datasets.
Xiaoliang Qian, Baokun Wu, Gong Cheng 0003, Xiwen Yao, Wei Wang 0245, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Learning to Assess Image Quality Like an Observer
abstract
Human observers are the ultimate receivers and evaluators of the image visual information and have powerful perception ability of visual quality with short-term global perception and long-term regional observation. Thus, it is natural to design an image quality assessment (IQA) computational model to act like an observer for accurately predicting the human perception of image quality. Inspired by this, here, we propose a novel observer-like network (OLN) to perform IQA by jointly considering the global glimpsing information and local scanning information. Specifically, the OLN consists of a global distortion perception (GDP) module and a local distortion observation (LDO) module. The GDP module is designed to mimic the observer's global perception of image quality through performing classification of images' distortion categories and levels. Simultaneously, to simulate the human local observation behavior, the LDO module attempts to gather the long-term regional observation information of the distorted images by continuously tracing the human scanpath in the observer-like scanning manner. By leveraging the bilinear pooling layer to collaborate the short-term global perception with the long-term regional observation, our network precisely predicts the quality scores of distorted images, such as human observers. Comprehensive experiments on the public datasets powerfully demonstrate that the proposed OLN achieves state-of-the-art performance.
Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 SCAN: Cross Domain Object Detection with Semantic Conditioned Adaptation
abstract
The domain gap severely limits the transferability and scalability of object detectors trained in a specific domain when applied to a novel one. Most existing works bridge the domain gap by minimizing the domain discrepancy in the category space and aligning category-agnostic global features. Though great success, these methods model domain discrepancy with prototypes within a batch, yielding a biased estimation of domain-level distribution. Besides, the category-agnostic alignment leads to the disagreement of class-specific distributions in the two domains, further causing inevitable classification errors. To overcome these two challenges, we propose a novel Semantic Conditioned AdaptatioN (SCAN) framework such that well-modeled unbiased semantics can support semantic conditioned adaptation for precise domain adaptive object detection. Specifically, class-specific semantics crossing different images in the source domain are graphically aggregated as the input to learn an unbiased semantic paradigm incrementally. The paradigm is then sent to a lightweight manifestation module to obtain conditional kernels to serve as the role of extracting semantics from the target domain for better adaptation. Subsequently, conditional kernels are integrated into global alignment to support the class-specific adaptation in a well-designed Conditional Kernel guided Alignment (CKA) module. Meanwhile, rich knowledge of the unbiased paradigm is transferred to the target domain with a novel Graph-based Semantic Transfer (GST) mechanism, yielding the adaptation in the category-based feature space. Comprehensive experiments conducted on three adaptation benchmarks demonstrate that SCAN outperforms existing works by a large margin.
Wuyang Li, Xinyu Liu 0001, Xiwen Yao, Yixuan Yuan
AAAI3
2022 Weakly Supervised Rotation-Invariant Aerial Object Detection Network
abstract
Object rotation is among longstanding, yet still unexplored, hard issues encountered in the task of weakly supervised object detection (WSOD) from aerial images. Existing predominant WSOD approaches built on regular CNNs which are not inherently designed to tackle object rotations without corresponding constraints, thereby leading to rotation-sensitive object detector. Meanwhile, current solutions have been prone to fall into the issue with unsTable detectors, as they ignore lower-scored instances and may regard them as backgrounds. To address these issues, in this paper, we construct a novel end-to-end weakly supervised Rotation-Invariant aerial object detection Network (RINet). It is implemented with a flexible multi-branch online detector refinement, to be naturally more rotation-perceptive against oriented objects. Specifically, RINet first performs label propagating from the predicted instances to their rotated ones in a progressive refinement manner. Meanwhile, we propose to couple the predicted in-stance labels among different rotation-perceptive branches for generating rotation-consistent supervision and mean-while pursuing all possible instances. With the rotation-consistent supervisions, RINet enforces and encourages consistent yet complementary feature learning for WSOD without additional annotations and hyper-parameters. On the challenging NWPU VHR-10.v2 and DIOR datasets, extensive experiments clearly demonstrate that we significantly boost existing WSOD methods to a new state-of-the-art performance. The code will be available at: https://github.com/XiaoxFeng/RINet.
Xiaoxu Feng, Xiwen Yao, Gong Cheng 0003, Junwei Han 0001
CVPR2
2022 Learning Non-target Knowledge for Few-shot Semantic Segmentation
abstract
Existing studies in few-shot semantic segmentation only focus on mining the target object information, however, often are hard to tell ambiguous regions, especially in non-target regions, which include background (BG) and Distracting Objects (DOs). To alleviate this problem, we propose a novel framework, namely Non-Target Region Eliminating (NTRE) network, to explicitly mine and eliminate BG and DO regions in the query. First, a BG Mining Module (BGMM) is proposed to extract the BG region via learning a general BG prototype. To this end, we design a BG loss to supervise the learning of BGMM only using the known target object segmentation ground truth. Then, a BG Eliminating Module and a DO Eliminating Module are proposed to successively filter out the BG and DO information from the query feature, based on which we can obtain a BG and DO-free target object segmentation result. Furthermore, we propose a prototypical contrastive learning algorithm to improve the model ability of distinguishing the target object from DOs. Extensive experiments on both PASCAL-5iand COCO-20idatasets show that our approach is effective despite its simplicity. Code is available at https://github.com/LIUYUANWEI98/NERTNet
Yuanwei Liu, Nian Liu 0002, Qinglong Cao, Xiwen Yao, Junwei Han 0001, Ling Shao 0001
CVPR4
2022 Intermediate Prototype Mining Transformer for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation aims to segment the target objects in query under the condition of a few annotated support images. Most previous works strive to mine more effective category information from the support to match with the corresponding objects in query. However, they all ignored the category information gap between query and support images. If the objects in them show large intra-class diversity, forcibly migrating the category information from the support to the query is ineffective. To solve this problem, we are the first to introduce an intermediate prototype for mining both deterministic category information from the support and adaptive category knowledge from the query. Specifically, we design an Intermediate Prototype Mining Transformer (IPMT) to learn the prototype in an iterative way. In each IPMT layer, we propagate the object information in both support and query features to the prototype and then use it to activate the query feature map. By conducting this process iteratively, both the intermediate prototype and the query feature can be progressively improved. At last, the final query feature is used to yield precise segmentation prediction. Extensive experiments on both PASCAL-5i and COCO-20i datasets clearly verify the effectiveness of our IPMT and show that it outperforms previous state-of-the-art methods by a large margin. Code is available at https://github.com/LIUYUANWEI98/IPMT
Yuanwei Liu, Nian Liu 0002, Xiwen Yao, Junwei Han 0001
NeurIPS3
2022 Guiding Clean Features for Object Detection in Remote Sensing Images
abstract
Recently, object detection has gained significant progress in remote sensing images. Nevertheless, we conclude two defects in remote sensing image object detection. At first, most methods rely on feature pyramid, but the features of different levels would influence each other when we use top-down operation. Second, the traditional label assignment strategy cannot assign suitable labels, as it adopts the fixed intersection over union (IoU) threshold to divide positive samples and negative samples during training. According to the problems we pointed out, a simple yet effective framework is employed to eliminate these two limitations. It integrates two novel components: aware feature pyramid network (AFPN) and group assignment strategy (GAS). AFPN is to mitigate the adverse effects caused by the first problem. Specifically, it learns a vector for the higher level features in the feature pyramid to obtain clean features. As for the second limitation, we recommend a new label assignment strategy named GAS to address this problem. Samples will be grouped according to their overlaps with ground truth, and then, they are assigned to positive or negative labels in each group. Extensive experiments are conducted on the large-scale object detection dataset DIOR and DOTA. With the newly introduced two key components, our model significantly improves the detection accuracy. Without bells and whistles, our proposed method achieves 2.0% and 1.9% higher mean average precision (mAP) than Faster R-CNN with FPN when using ResNet50 and ResNet101 as the backbones, respectively. Finally, we obtain 73.3% mAP on the DIOR dataset without any tricks. Our code is available athttps://github.com/hm-better/dior_detect.
Gong Cheng 0003, Hailong Hong, Xiwen Yao, Xiaoliang Qian, Lei Guo 0002
IEEE Geosci. Remote. Sens. Lett.4
2022 P-CNN: Part-Based Convolutional Neural Networks for Fine-Grained Visual Categorization
abstract
This paper proposes an end-to-end fine-grained visual categorization system, termed Part-based Convolutional Neural Network (P-CNN), which consists of three modules. The first module is a Squeeze-and-Excitation (SE) block, which learns to recalibrate channel-wise feature responses by emphasizing informative channels and suppressing less useful ones. The second module is a Part Localization Network (PLN) used to locate distinctive object parts, through which a bank of convolutional filters are learned as discriminative part detectors. Thus, a group of informative parts can be discovered by convolving the feature maps with each part detector. The third module is a Part Classification Network (PCN) that has two streams. The first stream classifies each individual object part into image-level categories. The second stream concatenates part features and global feature into a joint feature for the final classification. In order to learn powerful part features and boost the joint feature capability, we propose a Duplex Focal Loss used for metric learning and part classification, which focuses on training hard examples. We further merge PLN and PCN into a unified network for an end-to-end training process via a simple training technique. Comprehensive experiments and comparisons with state-of-the-art methods on three benchmark datasets demonstrate the effectiveness of our proposed method.
Junwei Han 0001, Xiwen Yao, Gong Cheng 0003, Xiaoxu Feng, Dong Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 SPNet: Siamese-Prototype Network for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot image classification has attracted extensive attention, which aims to recognize unseen classes given only a few labeled samples. Due to the large intraclass variances and interclass similarity of remote sensing scenes, the task under such circumstance is much more challenging than general few-shot image classification. Most existing prototype-based few-shot algorithms usually calculate prototypes directly from support samples and ignore the validity of prototypes, which results in a decline in the accuracy of subsequent inferences based on prototypes. To tackle this problem, we propose a Siamese-prototype network (SPNet) with prototype self-calibration (SC) and intercalibration (IC). First, to acquire more accurate prototypes, we utilize the supervision information from support labels to calibrate the prototypes generated from support features. This process is called SC. Second, we propose to consider the confidence scores of the query samples as another type of prototypes, which are then used to predict the support samples in the same way. Thus, the information interaction between support and query samples is implicitly a further calibration for prototypes (so-called IC). Our model is optimized with three losses, of which two additional losses help the model to learn more representative prototypes and make more accurate predictions. With no additional parameters to be learned, our model is very lightweight and convenient to employ. The experiments on three public remote sensing image datasets demonstrate competitive performance compared with other advanced few-shot image classification approaches. The source code is available athttps://github.com/zoraup/SPNet.
Gong Cheng 0003, Liming Cai, Chunbo Lang, Xiwen Yao, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Self-Guided Proposal Generation for Weakly Supervised Object Detection
abstract
Weakly Supervised Object Detection (WSOD) in remote sensing images remains a challenging task when learning object detectors with only image-level labels. As we know, object proposal generation plays a crucial role in WSOD. At present, the proposal generation of most existing WSOD methods mainly relies on heuristic strategies such as selective search and Edge Boxes. However, the proposals obtained by the above methods cannot well cover the entire objects, severely hindering the performance of WSOD. To address this issue, this paper proposes a Self-guided Proposal Generation approach, termed SPG. It can be easily implemented with most WSOD methods in a unified framework. To this end, we first introduce a confidence propagation approach to obtain the objectness confidence map for each image, which on the one hand highlights informative object locations, and on the other hand aggregates discriminative feature representation by combining the objectness confidence map with the deep features. Then, the proposal generation is implemented by mining informative regions as proposals on the objectness confidence map. Extensive evaluations on two challenging datasets demonstrate that our SPG significantly improves the baseline methods Online Instance Classifier Refinement (OICR) and Min-Entropy Latent Model (MELM) by large margins (for OICR: 15.86% mAP and 12.89% CorLoc gains on the NWPU VHR-10.v2 dataset, 3.65% mAP and 4.87% CorLoc gains on the DIOR dataset; for MELM: 20.51% mAP and 23.54% CorLoc gains on the NWPU VHR-10.v2 dataset, 7.11% mAP and 4.96% CorLoc gains on the DIOR dataset) and achieves the state-of-the-art results compared with existing methods.
Gong Cheng 0003, Weining Chen, Xiaoxu Feng, Xiwen Yao, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Dual-Aligned Oriented Detector
abstract
In the past few years, object detection in remote sensing images has achieved remarkable progress. However, the detection of oriented and densely packed objects are still unsatisfactory due to the following spatial and feature misalignments. 1) Most two-stage oriented detectors only introduce an orientation regression branch in the detection head, while still leverage horizontal proposals for classification and regression. This inevitably results in the spatial misalignment problem between horizontal proposals and oriented objects. 2) The features used for classification are in fact extracted from the region proposals which have shifted to the final predictions via the regression branch. This leads to the feature misalignment problem between the classification and the localization tasks. In this article, we present a two-stage oriented object detection method, termed dual-aligned oriented detector (DODet), toward evading the aforementioned problems of spatial and feature misalignments. In DODet, the first stage is an oriented proposal network (OPN), which generates high-quality oriented proposals via a novel representation scheme of oriented objects. The second stage is a localization-guided detection head (LDH) that aims at alleviating the feature misalignment between classification and localization. Comprehensive and extensive evaluations on three benchmarks, including DIOR-R, DOTA, and HRSC2016, indicate that our method could obtain consistent and substantial gains compared with the baseline method. The source code is publicly available athttps://github.com/yanqingyao1994/DODet.
Gong Cheng 0003, Shengyang Li, Ke Li 0005, Xingxing Xie, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Prototype-CNN for Few-Shot Object Detection in Remote Sensing Images
abstract
Recently, due to the excellent representation ability of convolutional neural networks (CNNs), object detection in remote sensing images has undergone remarkable development. However, when trained with a small number of samples, the performance of the object detectors drops sharply. In this article, we focus on the following three main challenges of few-shot object detection in remote sensing images: 1) since the sample number of novel classes is far less than base classes, object detectors would fail to quickly adapt to the features of novel classes, which would result in overfitting; 2) the scarcity of samples in novel classes leads to a sparse orientation space, while the objects in remote sensing images usually have arbitrary orientations; and 3) the distribution of object instances in remote sensing images is scattered and, therefore, it is hard to identify foreground objects from the complex background. To tackle these problems, we propose a simple yet effective method named prototype-CNN (P-CNN), which mainly consists of three parts: a prototype learning network (PLN) converting support images to class-aware prototypes, a prototype-guided region proposal network (P-G RPN) for better generation of region proposals, and a detector head extending the head of Faster region-based CNN (R-CNN) to further boost the performance. Comprehensive evaluations on the large-scale DIOR dataset demonstrate the effectiveness of our P-CNN. The source code is available athttps://github.com/Ybowei/P-CNN.
Gong Cheng 0003, Bowei Yan, Peizhen Shi, Ke Li 0005, Xiwen Yao, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 SAENet: Self-Supervised Adversarial and Equivariant Network for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images (RSIs) remains a challenge when learning a subtle object detection model with only image-level annotations. Most works tend to optimize the detection model via exploiting the most contributed region, thereby to be dominated by the most discriminative part of an object. Meanwhile, these methods ignore the consistency across different spatial transformations of the same image and always label them with different classes, which introduces potential ambiguities. To tackle these challenges, we propose a unique self-supervised adversarial and equivariant network (SAENet) and aim at learning complementary and consistent visual patterns for WSOD in RSIs. To this end, an adversarial dropout–activation block is first designed to facilitate the entire object detector via adaptively hiding the discriminative parts and highlighting the instance-related regions. Besides, we further introduce a flexible self-supervised transformation equivariance mechanism on each potential instance from multiple spatial transformations to obtain spatially consistent self-supervisions. Accordingly, the obtained supervisions can be leveraged to pursue a more robust and spatially consistent object detector. Comprehensive experiments on the challenging LEarning, VIsion and Remote sensing Laboratory (LEVIR), NorthWestern Polytechnical University (NWPU) VHR-10.v2, and detection in optical RSIs (DIOR) datasets validate that SAENet outperforms the previous state-of-the-art works and achieves 46.2%, 60.7%, and 27.1% mAP, respectively.
Xiaoxu Feng, Xiwen Yao, Gong Cheng 0003, Jungong Han, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 AIFS-DATASET for Few-Shot Aerial Image Scene Classification
abstract
Few-shot learning (FSL), which aims to rapidly recognize unseen categories with limited samples, has attracted wide attention in aerial image scene classification. However, the existing methods generally train and evaluate the model within a dataset, and changing the dataset requires retraining and evaluation, which only realizes the generalization of intra-dataset. Considering meta-learning, this brings in a natural assumption: FSL should learn meta-knowledge from cross-domain heterogeneous tasks and then can generalize to new data distributions (e.g., datasets) with few samples. To this end, we propose a new benchmark, dubbed aerial image few-shot dataset (AIFS-DATASET), which is composed of diverse datasets and can provide more realistic heterogeneous task distributions. On AIFS-DATASET, we use many heterogeneous tasks, across multi-domains without any aerial image category, to train the model, achieving “see more.” Then we transfer the learned knowledge to new tasks in aerial images to evaluate the generalization performance of the model, thus acquiring a “well-informed” few-shot aerial image scene classification model. Moreover, the challenges of inter-class similarity and intra-class discrepancy in aerial images still exist. We also develop a dual constrained distance metric learning (DC-DML) framework to deal with the variable learning tasks adaptively and to achieve compact data distribution within a class and clear distribution gaps between classes from the perspective of metric learning. DC-DML mainly uses a task-adapted feature extractor while devising a novel distance metric with a cross-class bias penalty. By conducting experiments on AIFS-DATASET, we observed that DC-DML outperforms the current prevailing FSL approaches by a large margin.
Lingjun Li, Xiwen Yao, Gong Cheng 0003, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Solo-to-Collaborative Dual-Attention Network for One-Shot Object Detection in Remote Sensing Images
abstract
In this article, we attempt to achieve one-shot object detection by mimicking the human ability to learn new concepts under limited reference, which aims at detecting all object instances of an unseen class in a target image when given a query image of the same unseen class. However, this one-shot learning ability of human benefits from the fact that human brain can quickly extract and process the associated information between the query–target images, which is an issue for the one-shot object detection framework to overcome. Moreover, the feature extraction of the query class in target images is intractable due to the complex and diversified background of remote sensing images. To solve these issues, we propose a solo-to-collaborative dual-attention network (SCoDANet) to hierarchically (image itself/pairs) enhance image feature representations. It consists of three components: 1) solo-attention head that strengthens the compactness of intraclass feature representations of an image and avoids background interference by selectively aggregating the similar features from the spatial and channel dimensions, respectively; 2) dual coattention module that guides RPN to generate an expected set of region proposals related to the query class by mining the coinformation of each query–target feature pair; and 3) nonlinear matching that provides a measure of similarity between the query feature and proposals of the target image to further learn a more robust detector. Our extensive experiments over two benchmarks demonstrate the effectiveness of our method under the one-shot scenario of detecting seen and unseen object categories.
Lingjun Li, Xiwen Yao, Gong Cheng 0003, Mingliang Xu 0001, Jungong Han, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Scale-Aware Detailed Matching for Few-Shot Aerial Image Semantic Segmentation
abstract
Few-shot semantic segmentation, aiming to segment query images with a few annotated support samples, has drawn increasing attention. Most existing few-shot methods leverage the single prototype obtained from global average pooling to represent all support information and further use the extracted prototype to segment the query images in a matching manner. Although promising results for natural images have been reported, these methods cannot be directly applied on aerial images. The main reason comes from that the extracted single support prototype can only provide a coarse guidance for matching between query and support images and could not handle the large variance of objects’ appearances and scales. To deal with these challenges on aerial images, we propose a scale-aware few-shot semantic segmentation network to perform detailed matching with multiple prototypes. More specifically, the detailed matching module is first constructed to compute the pixel-level similarity between the query features and the extracted multiple support prototypes for providing more accurate parsing guidance. Subsequently, to address the problem of scale imbalance, the scale-aware focal loss is designed to dynamically down-weight the loss assigned to large well-parsed objects and focus training on tiny hard-parsed objects. To facilitate the reproducible research on the task of few-shot semantic segmentation in aerial images, we further provide a few-shot segmentation benchmark iSAID-$5^{\mathrm {i}}$constructed from the large-scale iSAID dataset[1]. Comprehensive experiments and comparisons with the state-of-the-art few-shot segmentation methods on the iSAID-$5^{\mathrm {i}}$dataset clearly demonstrate the superiority of our proposed method. The code and dataset are available athttps://github.com/caoql98/SDM.
Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 R²IPoints: Pursuing Rotation-Insensitive Point Representation for Aerial Object Detection
abstract
Anchor-free aerial object detection methods have recently attracted much attention due to their simplicity and efficiency. However, the performance is still unsatisfactory due to the following two main limitations. On the one hand, the anchor-free detector employs ordinary convolution layers with axis-aligned receptive fields to extract object features, resulting in lacking internal mechanisms to handle the rotation variance. On the other hand, the detector sacrifices much semantic information to achieve faster detection, leading to the inability to deal with objects’ high inter-class similarity and intra-class diversity. To address these issues, in this paper, we present a unique anchor-free detector, termed Rotation-Insensitive Point Representation (R2IPoints), of which a set of category-aware points are employed to encode the spatial and semantic information of the arbitrary-oriented objects. Specifically, we first devise a Stacked Rotation convolution Module (SRM) to encourage the learning of rotation-insensitive point representation by adaptively modelling orientation-agnostic interdependencies over stochastically rotated features. Meanwhile, we further introduce a Class-specific Semantic enhancement Module (CSM). It performs category-aware semantic activation to recalibrate features, thus enabling the point representation to be aware of object categories. Through jointly optimizing the two proposed modules in an end-to-end manner, R2IPoints could simultaneously generate rotation-insensitive and category-aware point representation. Extensive experiments on the challenging DIOR and DOTA datasets demonstrate the superiority of the proposed method. We achieve 72.7% mAP on DIOR and 74.34% mAP on DOTA, surpassing the baseline method of +2.4% mAP and +2.49% mAP, respectively. The code is available at https://github.com/shnew/R2IPoints.
Xiwen Yao, Hui Shen 0005, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 DFENet for Domain Adaptation-Based Remote Sensing Scene Classification
abstract
Domain adaptation scene classification refers to the task of scene classification where the training set (called source domain) has different distributions from the test set (called target domain). Although remarkable results have been reported, the misalignment of source and target domain features still remains a big challenge when the large intraclass variances of remote sensing images encounter the insufficient exploration of discriminative feature representations for both domains. To address this challenge, a novel domain feature enhancement network (DFENet) is proposed to adaptively enhance the discriminative ability of the learned features for dealing with the domain variances of scene classification. Specifically, an adaptive context-aware feature refinement (CAFR) module is first designed to automatically recalibrate global and local features by explicitly modeling interdependencies between the channel and spatial for each domain. Then, a multilevel adversarial dropout (MAD) module is further designed to strengthen the generalization capability of our network by adaptively reconfiguring the sparsity of the feature level and decision level in the target domain. The cooperation of CAFR module and MAD module formulates a unique DFENet that can be learned in an end-to-end manner. Comprehensive experiments show that our proposed method is better than state-of-the-art methods on Merced$\to $RSSCN7, AID$\to $RSSCN7, NWPU$\to $RSSCN7, RSSCN7$\to $Merced, RSSCN7$\to $AID, and RSSCN7$\to $NWPU datasets.
Xiufei Zhang, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Oriented R-CNN for Object Detection
abstract
Current state-of-the-art two-stage detectors generate oriented proposals through time-consuming schemes. This diminishes the detectors’ speed, thereby becoming the computational bottleneck in advanced oriented object detection systems. This work proposes an effective and simple oriented object detection framework, termed Oriented R-CNN, which is a general two-stage oriented detector with promising accuracy and efficiency. To be specific, in the first stage, we propose an oriented Region Proposal Network (oriented RPN) that directly generates high-quality oriented proposals in a nearly cost-free manner. The second stage is oriented R-CNN head for refining oriented Regions of Interest (oriented RoIs) and recognizing them. Without tricks, oriented R-CNN with ResNet50 achieves state-of-the-art detection accuracy on two commonly-used datasets for oriented object detection including DOTA (75.87% mAP) and HRSC2016 (96.50% mAP), while having a speed of 15.1 FPS with the image size of 1024×1024 on a single RTX 2080Ti. We hope our work could inspire rethinking the design of oriented detectors and serve as a baseline for oriented object detection. Code is available at https://github.com/jbwang1997/OBBDetection.
Xingxing Xie, Gong Cheng 0003, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001
ICCV4
2021 Cross-Scale Feature Fusion for Object Detection in Optical Remote Sensing Images
abstract
For the time being, there are many groundbreaking object detection frameworks used in natural scene images. These algorithms have good detection performance on the data sets of open natural scenes. However, applying these frameworks to remote sensing images directly is not very effective. The existing deep-learning-based object detection algorithms still face some challenges when dealing with remote sensing images because these images usually contain a number of targets with large variations of object sizes as well as interclass similarity. Aiming at the challenges of object detection in optical remote sensing images, we propose an end-to-end cross-scale feature fusion (CSFF) framework, which can effectively improve the object detection accuracy. Specifically, we first use a feature pyramid network (FPN) to obtain multilevel feature maps and then insert a squeeze and excitation (SE) block into the top layer to model the relationship between different feature channels. Next, we use the CSFF module to obtain powerful and discriminative multilevel feature representations. Finally, we implement our work in the framework of Faster region-based CNN (R-CNN). In the experiment, we evaluate our method on a publicly available large-scale data set, named DIOR, and obtain an improvement of 3.0% measured in terms of mAP compared with Faster R-CNN with FPN.
Gong Cheng 0003, Yongjie Si, Hailong Hong, Xiwen Yao, Lei Guo 0002
IEEE Geosci. Remote. Sens. Lett.4
2021 Weighted bilateral K-means algorithm for fast co-clustering and fast spectral clustering
Kun Song 0001, Xiwen Yao, Feiping Nie 0001, Xuelong Li 0001, Mingliang Xu 0001
Pattern Recognit.2
2021 Two-Stream Encoder GAN With Progressive Training for Co-Saliency Detection
abstract
The recent end-to-end co-saliency models have good performance, however, they cannot express the semantic consistency among a group of images well and usually require many co-saliency labels. To this end, a two-stream encoder generative adversarial network (TSE-GAN) with progressive training is proposed in this paper. In the pre-training stage, the salient object detection generative adversarial networks (SOD-GAN) and classification network (CN) are separately trained by the salient object detection (SOD) datasets and co-saliency datasets with only category labels to learn the intra-saliency and preliminary inter-saliency cues and alleviate the problem of insufficient co-saliency labels. In the second training stage, the backbone of TSE-GAN is inherited from the trained SOD-GAN, the encoder of trained SOD-GAN (SOD-Encoder) is used to extract intra-saliency features, the group-wise semantic encoder (GS-Encoder) is constructed by the multi-level group-wise category features extracted from CN for extracting inter-saliency features with better semantic consistency, the TSE-GAN constructed by incorporating the GS-Encoder into SOD-GAN is trained on co-saliency datasets for co-saliency detection. The comprehensive comparisons with 13 state-of-the-art methods demonstrate the effectiveness of proposed method.
Xiaoliang Qian, Gong Cheng 0003, Xiwen Yao, Liying Jiang
IEEE Signal Process. Lett.4
2021 TCANet: Triple Context-Aware Network for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised object detection (WSOD) in remote sensing images (RSI) plays an essential role in RSI understanding applications. Currently, predominant works are inclined to first activate the most discriminative region and then pursue the whole object by analyzing the context information of the activated region. However, the most discriminative region usually only covers a small crucial part. Besides, many same-class instances often appear in adjacent locations. In such a case, treating proposals of large spatial overlap as the same-class instances not only introduces potential ambiguities but also misleads the detection model to recognize multiple adjacent instances as one object instance. To address these challenges, a novel triple context-aware network (TCANet) is proposed to learn complementary and discriminative visual patterns for WSOD in RSIs. Specifically, a global context-aware enhancement (GCAE) module is first designed to activate the features of the whole object by capturing the global visual scene context. Then, a dual-local context residual (DLCR) module is further developed to capture the instance-level discriminative cues by leveraging the semantic discrepancy of the local context. Furthermore, an effective adaptive-weighted refinement loss is integrated into the DLCR module to reduce the ambiguities in the label propagating process. The collaboration of GCAE and DLCR formulates a unique TCANet that can be learned in an end-to-end manner. Comprehensive experiments are carried out on the challenging NWPU VHR-10.v2 and DIOR data sets. We achieve a 58.8% mAP and a 25.8% mAP on the NWPU VHR-10.v2 and DIOR data sets, respectively, which both significantly outperform the state of the arts.
Xiaoxu Feng, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.3
2021 DLA-MatchNet for Few-Shot Remote Sensing Image Scene Classification
abstract
Few-shot scene classification aims to recognize unseen scene concepts from few labeled samples. However, most existing works are generally inclined to learn metalearners or transfer knowledge while ignoring the importance to learn discriminative representations and a proper metric for remote sensing images. To address these challenges, in this article, we propose an end-to-end network for boosting a few-shot remote sensing image scene classification, called discriminative learning of adaptive match network (DLA-MatchNet). Specifically, we first adopt the attention technique to delve into the interchannel and interspatial relationships to automatically discover discriminative regions. Then, the channel attention and spatial attention modules can be incorporated with the feature network by using different feature fusion schemes, achieving “discriminative learning.” Afterward, considering the issues of the large intraclass variances and interclass similarity of remote sensing images, instead of simply computing the distances between the support samples and query samples, we concatenate the support and query discriminative features in depth and utilize a matcher to “adaptively” select the semantically relevant sample pairs to assign similarity scores. Our method leverages an episode-based strategy to train the model. Once trained, our model can predict the category of query image without further fine-tuning. Experimental results on three public remote sensing image data sets demonstrate the effectiveness of our model in the few-shot scene classification task.
Lingjun Li, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.3
2021 Automatic Weakly Supervised Object Detection From High Spatial Resolution Remote Sensing Images via Dynamic Curriculum Learning
abstract
In this article, we focus on tackling the problem of weakly supervised object detection from high spatial resolution remote sensing images, which aims to learn detectors with only image-level annotations, i.e., without object location information during the training stage. Although promising results have been achieved, most approaches often fail to provide high-quality initial samples and thus are difficult to obtain optimal object detectors. To address this challenge, a dynamic curriculum learning strategy is proposed to progressively learn the object detectors by feeding training images with increasing difficulty that matches current detection ability. To this end, an entropy-based criterion is firstly designed to evaluate the difficulty for localizing objects in images. Then, an initial curriculum that ranks training images in ascending order of difficulty is generated, in which easy images are selected to provide reliable instances for learning object detectors. With the gained stronger detection ability, the subsequent order in the curriculum for retraining detectors is accordingly adjusted by promoting difficult images as easy ones. In such way, the detectors can be well prepared by training on easy images for learning from more difficult ones and thus gradually improve their detection ability more effectively. Moreover, an effective instance-aware focal loss function for detector learning is developed to alleviate the influence of positive instances of bad quality and meanwhile enhance the discriminative information of class-specific hard negative instances. Comprehensive experiments and comparisons with state-of-the-art methods on two publicly available data sets demonstrate the superiority of our proposed method.
Xiwen Yao, Xiaoxu Feng, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.1
2020 Progressive Contextual Instance Refinement for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly supervised learning has been attracting much attention due to its broad applications, which only requires image-level annotations to indicate whether there exist objects in the images. Currently, most of the existing weakly supervised object detection (WSOD) methods are inclined to seek only one top-scoring object instance per image from noisy proposals to train the corresponding object detector. However, more than one same-class instances often exist in the large-scale, cluttered remote sensing images. Thus, selecting only one top-scoring proposal usually results in highlighting the most representative part of an object rather than the whole object, which may cause learning a suboptimal object detector by losing much important information. To address this problem, a novel end-to-end progressive contextual instance refinement (PCIR) method is proposed to perform WSOD. Specifically, a dual-contextual instance refinement (DCIR) strategy is designed to divert the focus of the detection network from the local distinct part to the whole object and further to other potential instances by leveraging both local and global context information. Benefiting from DCIR, a progressive proposal self-pruning (PPSP) strategy is further developed to mitigate the influence of the complex background by dynamically rejecting the negative training proposals. Comprehensive experiments on the challenging NWPU VHR-10.v2 and DIOR data sets clearly demonstrate that the proposed method can significantly boost the detection accuracy compared with the state of the arts.
Xiaoxu Feng, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003
IEEE Trans. Geosci. Remote. Sens.3
2019 Performance Comparison of Two Pooling Strategies for Remote Sensing Image Scene Classification
abstract
With the advances of convolutional neural networks (CNNs), the accuracy of remote sensing image scene classification has been greatly boosted thanks to the powerful features extracted through CNNs. Although significant success has been achieved, most of existing methods are dominated by the use of fully-connected CNN features. This paper focuses on the performance comparison of two kinds of novel pooling strategies, including generalized max pooling (GMP) and taskdriven pooling (TDP), for remote sensing image scene classification. To this end, an off-the-shelf CNN model is used as backbone network to extract multi-scale convolutional features. Then, GMP and TDP are respectively adopted to obtain globally pooled features. Finally, scene classification is performed with support vector machine (SVM). In the experiment, we evaluate the performance of these two kinds of pooling schemes on a widely-used scene classification benchmark data set. The experimental results show that (i) using pooled CNN convolutional features can obtain better results than using fully-connected CNN features and (ii) TDP is slightly better than GMP.
Maoxiong Wu, Gong Cheng 0003, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001, Lei Guo 0002
IGARSS3
2019 Learning Region Response Ranking Features for Remote Sensing Image Scene Classification
abstract
Recently, deep learning especially convolutional neural networks (CNNs) has huge great success for remote sensing image scene classification. However, global CNN features still lack geometric invariance for addressing the problem of large intra-class variations and so are not optimal for scene classification. In this paper, we introduce a new feature representation for scene classification, named region response ranking (3R) feature representations by using off-the-shelf CNN models. Specifically, by considering each cube pixel of a certain convolutional feature map as one image region, we jointly train a class-specific support vector machine (SVM) base classifier and a decision function for each scene class. The base classifier is used to generate 3R feature by reordering the SVM responses of all image regions in descending order and the decision function is used for classification with 3R feature representations. Comprehensive evaluations on the publicly available NWPU-RESISC45 data set and comparisons with state-of-the-art methods demonstrate that the proposed 3R feature is effective for remote sensing image scene classification.1
Junyu Yang, Gong Cheng 0003, Xiwen Yao, Junwei Han 0001, Lei Guo 0002
IGARSS3
2019 Rotation-Invariant Latent Semantic Representation Learning for Object Detection in VHR Optical Remote Sensing Images
abstract
Object detection in very high resolution (VHR) optical remote sensing images is a fundamental yet challenging problem for the field of remote sensing image analysis. The detection performance is heavily dependent on the representation capability of the extracted features. Recently, convolutional neural networks (CNNs) have made a breakthrough for various applications in nature images. However, it is problematic to directly apply CNN to perform object detection in VHR optical remote sensing images due to the problem of object rotation variations. To address this issue, a novel rotation invariant probabilistic Latent Semantic Analysis (RI-pLSA) model is proposed to learn latent semantic representations for object detection. This is achieved by imposing a rotation-invariant regularization term on the objective function of pLSA to enforce the learned representation from all rotations of the same sample to be as consistent as possible. Additionally, the proposed RI-pLSA model takes the CNN features as input, which generates more powerful semantic representation for object detection. Comprehensive experiments on a publicly available ten-class object detection dataset demonstrate the superiority and effectiveness of our method compared with state-of-the-arts.
Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002
IGARSS1
2019 Scene Classification of High Resolution Remote Sensing Images Via Self-Paced Deep Learning
abstract
Scene classification of high resolution remote sensing (HRRS) images is a fundamental yet challenging problem for remote sensing image analysis. In this paper, we focus on tackling the problem of HRSS scene classification using a small pool of unlabeled images and only a few labeled images per category, namely, few-shot scene classification (FSSC), which is more challenging than common scene classification task. The key challenge arises from selecting trustworthy samples from the pool of unlabeled images that have high confidence. To address this challenge, a novel local manifold constrained self-paced deep learning method is proposed. Specifically, the model is learned by gradually selecting easy samples from the pool of unlabeled images, assigning them with pseudo-labels and further adopting them with labeled images as the new training set. In addition, a local manifold constraint is introduced to enforce that the pseudo-labels assigned by the initial model should be consistent with the local manifold of the labeled samples. In such way, the confidence of the selecting samples is increased and is beneficial to train more robust classifier. Experimental results on a publicly available large scale NWPU-RESISC45 data set demonstrated the effectiveness of our method in achieving competitive performance while significantly reducing manually labeled cost.
Xiwen Yao, Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002
IGARSS1
2018 Discriminative Joint-Feature Topic Model With Dual Constraints for WCE Classification
abstract
Wireless capsule endoscopy (WCE) enables clinicians to examine the digestive tract without any surgical operations, at the cost of a large amount of images to be analyzed. The main challenge for automatic computer-aided diagnosis arises from the difficulty of robust characterization of these images. To tackle this problem, a novel discriminative joint-feature topic model (DJTM) with dual constraints is proposed to classify multiple abnormalities in WCE images. We first propose a joint-feature probabilistic latent semantic analysis (PLSA) model, where color and texture descriptors extracted from same image patches are jointly modeled with their conditional distributions. Then the proposed dual constraints: visual words importance and local image manifold are embedded into the joint-feature PLSA model simultaneously to obtain discriminative latent semantic topics. The visual word importance is proposed in our DJTM to guarantee that visual words with similar importance come from close latent topics while the local image manifold constraint enforces that images within the same category share similar latent topics. Finally, each image is characterized by distribution of latent semantic topics instead of low level features. Our proposed DJTM showed an excellent overall recognition accuracy 90.78%. Comprehensive comparison results demonstrate that our method outperforms existing multiple abnormalities classification methods for WCE images.
Yixuan Yuan, Xiwen Yao, Junwei Han 0001, Lei Guo 0002, Max Q.-H. Meng
IEEE Trans. Cybern.2
2018 Exploring Hierarchical Convolutional Features for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification is an active and important research task driven by many practical applications. To leverage deep learning models especially convolutional neural networks (CNNs) for HSI classification, this paper proposes a simple yet effective method to extract hierarchical deep spatial feature for HSI classification by exploring the power of off-the-shelf CNN models, without any additional retraining or fine-tuning on the target data set. To obtain better classification accuracy, we further propose a unified metric learning-based framework to alternately learn discriminative spectral-spatial features, which have better representation capability and train support vector machine (SVM) classifiers. To this end, we design a new objective function that explicitly embeds a metric learning regularization term into SVM training. The metric learning regularization term is used to learn a powerful spectral-spatial feature representation by fusing spectral feature and deep spatial feature, which has small intraclass scatter but big between class separation. By transforming HSI data into new spectral-spatial feature space through CNN and metric learning, we can pull the pixels from the same class closer, while pushing the different class pixels farther away. In the experiments, we comprehensively evaluate the proposed method on three commonly used HSI benchmark data sets. State-of-the-art results are achieved when compared with the existing HSI classification methods.
Gong Cheng 0003, Junwei Han 0001, Xiwen Yao, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.4
2018 When Deep Learning Meets Metric Learning: Remote Sensing Image Scene Classification via Learning Discriminative CNNs
abstract
Remote sensing image scene classification is an active and challenging task driven by many applications. More recently, with the advances of deep learning models especially convolutional neural networks (CNNs), the performance of remote sensing image scene classification has been significantly improved due to the powerful feature representations learnt through CNNs. Although great success has been obtained so far, the problems of within-class diversity and between-class similarity are still two big challenges. To address these problems, in this paper, we propose a simple but effective method to learn discriminative CNNs (D-CNNs) to boost the performance of remote sensing image scene classification. Different from the traditional CNN models that minimize only the cross entropy loss, our proposed D-CNN models are trained by optimizing a new discriminative objective function. To this end, apart from minimizing the classification error, we also explicitly impose a metric learning regularization term on the CNN features. The metric learning regularization enforces the D-CNN models to be more discriminative so that, in the new D-CNN feature spaces, the images from the same scene class are mapped closely to each other and the images of different classes are mapped as farther apart as possible. In the experiments, we comprehensively evaluate the proposed method on three publicly available benchmark data sets using three off-the-shelf CNN models. Experimental results demonstrate that our proposed D-CNN methods outperform the existing baseline methods and achieve state-of-the-art results on all three data sets.
Gong Cheng 0003, Ceyuan Yang, Xiwen Yao, Lei Guo 0002, Junwei Han 0001
IEEE Trans. Geosci. Remote. Sens.3
2017 Remote Sensing Image Scene Classification Using Bag of Convolutional Features
abstract
More recently, remote sensing image classification has been moving from pixel-level interpretation to scene-level semantic understanding, which aims to label each scene image with a specific semantic class. While significant efforts have been made in developing various methods for remote sensing image scene classification, most of them rely on handcrafted features. In this letter, we propose a novel feature representation method for scene classification, named bag of convolutional features (BoCF). Different from the traditional bag of visual words-based methods in which the visual words are usually obtained by using handcrafted feature descriptors, the proposed BoCF generates visual words from deep convolutional features using off-the-shelf convolutional neural networks. Extensive evaluations on a publicly available remote sensing image scene classification benchmark and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed BoCF method for remote sensing image scene classification.
Gong Cheng 0003, Xiwen Yao, Lei Guo 0002, Zhongliang Wei
IEEE Geosci. Remote. Sens. Lett.3
2017 Revisiting Co-Saliency Detection: A Novel Approach Based on Two-Stage Multi-View Spectral Rotation Co-clustering
abstract
With the goal of discovering the common and salient objects from the given image group, co-saliency detection has received tremendous research interest in recent years. However, as most of the existing co-saliency detection methods are performed based on the assumption that all the images in the given image group should contain co-salient objects in only one category, they can hardly be applied in practice, particularly for the large-scale image set obtained from the Internet. To address this problem, this paper revisits the co-saliency detection task and advances its development into a new phase, where the problem setting is generalized to allow the image group to contain objects in arbitrary number of categories and the algorithms need to simultaneously detect multi-class co-salient objects from such complex data. To solve this new challenge, we decompose it into two sub-problems, i.e., how to identify subgroups of relevant images and how to discover relevant co-salient objects from each subgroup, and propose a novel co-saliency detection framework to correspondingly address the two sub-problems via two-stage multi-view spectral rotation co-clustering. Comprehensive experiments on two publically available benchmarks demonstrate the effectiveness of the proposed approach. Notably, it can even outperform the state-of-the-art co-saliency detection methods, which are performed based on the image subgroups carefully separated by the human labor.
Xiwen Yao, Junwei Han 0001, Dingwen Zhang, Feiping Nie 0001
IEEE Trans. Image Process.1
2016 Scene classification of high resolution remote sensing images using convolutional neural networks
abstract
Scene classification of high resolution remote sensing images plays an important role for a wide range of applications. While significant efforts have been made in developing various methods for scene classification, most of them are based on handcrafted or shallow learning-based features. In this paper, we investigate the use of deep convolutional neural network (CNN) for scene classification. To this end, we first adopt two simple and effective strategies to extract CNN features: (1) using pre-trained CNN models as universal feature extractors, and (2) domain-specifically fine-tuning pre-trained CNN models on our scene classification dataset. Then, scene classification is carried out by using simple classifiers such as linear support vector machine (SVM). In our work, three off-the-shelf CNN models including AlexNet [1], VGGNet [2], and GoogleNet [3] are investigated. Comprehensive evaluations on a publicly available 21 classes land use dataset and comparisons with several state-of-the-art approaches demonstrate that deep CNN features are effective for scene classification of high resolution remote sensing images.
Gong Cheng 0003, Chengcheng Ma, Peicheng Zhou, Xiwen Yao, Junwei Han 0001
IGARSS4
2016 Semantic annotation of satellite images via joint multi-feature learning with diversity constraint
abstract
Automatic semantic annotation of high-resolution optical satellite images is a task to assign one or several predefined semantic concepts to an image according to its content. The fundamental challenge arises from the difficulty of characterizing complex and ambiguous contents of the satellite images. To address this challenge, a diversity constrained joint multi-feature learning method is proposed to learn robust feature representations for annotating satellite images. The key motivation of our method is to make full use of the complementarity diversity information among the heterogeneous features in the learning process. Comprehensive experiments on an annotation dataset demonstrate the superiority and effectiveness of our method compared with baseline multi-feature learning method.
Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Peicheng Zhou, Lei Guo 0002
IGARSS1
2016 Semantic Annotation of High-Resolution Satellite Images via Weakly Supervised Learning
abstract
In this paper, we focus on tackling the problem of automatic semantic annotation of high resolution (HR) optical satellite images, which aims to assign one or several predefined semantic concepts to an image according to its content. The main challenges arise from the difficulty of characterizing complex and ambiguous contents of the satellite images and the high human labor cost caused by preparing a large amount of training examples with high-quality pixel-level labels in fully supervised annotation methods. To address these challenges, we propose a unified annotation framework by combining discriminative high-level feature learning and weakly supervised feature transferring. Specifically, an efficient stacked discriminative sparse autoencoder (SDSAE) is first proposed to learn high-level features on an auxiliary satellite image data set for the land-use classification task. Inspired by the motivation that the encoder of the prelearned SDSAE can be regarded as a generic high-level feature extractor for HR optical satellite images, we then transfer the learned high-level features to semantic annotation. To compensate the difference between the auxiliary data set and the annotation data set, the transferred high-level features are further fine-tuned in a weakly supervised scheme by using the tile-level annotated training data. Finally, the fine-tuning process is formulated as an ultimate optimization problem, which can be solved efficiently with our proposed alternate iterative optimization method. Comprehensive experiments on a publicly available land-use classification data set and an annotation data set demonstrate the superiority of our SDSAE-based high-level feature learning method and the effectiveness of our weakly supervised semantic annotation framework compared with state-of-the-art fully supervised annotation methods.
Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Xueming Qian, Lei Guo 0002
IEEE Trans. Geosci. Remote. Sens.1
2015 Semantic Segmentation based on Stacked Discriminative Autoencoders and Context-Constrained Weakly Supervised Learning
abstract
In this paper, we focus on tacking the problem of weakly supervised semantic segmentation. The aim is to predict the class label of image regions under weakly supervised settings, where training images are only provided with image-level labels indicating the classes they contain. The main difficulty of weakly supervised semantic segmentation arises from the complex diversity of visual classes and the lack of supervision information for learning a multi-classes classifier. To conquer the challenge, we propose a novel discriminative deep feature learning framework based on stacked autoencoders (SAE) by integrating pairwise constraints to serve as a discriminative term. Furthermore, to mine effective supervision information, global context about co-occurrence of visual classes as well as local context around each image region is exploited as constraints for training a multi-class classifier. Finally, the classifier training is formulated as an ultimate optimization problem, which can be solved efficiently by an alternate iterative optimization method. Comprehensive experiments on the MSRC 21 dataset demonstrate the superior performance compared with several state-of-the-art weakly supervised image segmentation methods.
Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002
ACM Multimedia1
2015 A coarse-to-fine model for airport detection from remote sensing images using target-oriented visual saliency and CRF
Xiwen Yao, Junwei Han 0001, Lei Guo 0002, Shuhui Bu, Zhenbao Liu
Neurocomputing1