EDBT 2026 Demo / reviewers in the wild / expert
Gong Cheng 0003
dblp:69/1215-3
· DBLP profile ↗
140ranked-venue papers
29as first author
106since 2021 · last 2026
0000-0001-5030-0683ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 69 · 20 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 6 first-author · 32 since 2021Artificial intelligence and machine learning · 41 · 6 first-author · 34 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UQ-ViT: Harmonizing Extreme Activations with Hardware-Friendly Uniform Quantization in Vision TransformersabstractPost-Training Quantization enables efficient Vision Transformer (ViTs) deployment with a small calibration data, and its prevalent use of uniform quantization harnesses AI accelerator matrix cores for high-speed inference. However, the application of uniform quantization is fundamentally challenged by the extreme non-uniformity of activation distributions.Specifically, the power-law nature of post-Softmax attention scores and the significant inter-channel variance in post-GELU activations create a dilemma for conventional quantization, as it struggles to preserve critical high-magnitude values without sacrificing overall precision. To resolve this core conflict, we introduce UQ-ViT (Uniform Quantization for Vision Transformers), a novel uniform quantization framework designed to reconcile high precision with hardware efficiency. Central to UQ-ViT are two operators: Dynamic Elimination of Maximum (DeMax) and Normalization Quantization (NormQuant). DeMax is a quantization operator for post-Softmax attention scores that utilizes uniform quantization. It dynamically eliminates and preserves dominant values, effectively mitigating quantization loss from the extreme values in the power-law distribution. NormQuant utilizes a per-channel quantization strategy during quantization and reverts to a per-tensor format for dequantization, achieving both high accuracy and computational efficiency. Crucially, it is applicable to any linear layer, enabling effective quantization of post-GELU activations in ViTs. Through extensive experiments on various ViTs and vision tasks, including image classification, object detection, and instance segmentation, we demonstrate that our proposed approach outperforms existing methods, achieving superior accuracy while ensuring hardware friendliness. Tao Jiang 0002, Yucheng Jiang, Xiwen Yao, Gong Cheng 0003, Junwei Han 0001 |
AAAI | 4 |
| 2026 | Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
Shilei Wang 0001, Pujian Lai, Jifeng Ning, Gong Cheng 0003 |
AAAI | 5 |
| 2026 | Unified Interaction Consistency Learning for Single-Source Domain-Generalized Object Detection in Urban SceneabstractDomain generalization remains a critical challenge for deploying neural networks, particularly in out-of-distribution object detection. The distributional discrepancy between training (e.g., daytime-sunny) and the realistic condition (e.g., night-rainy) inevitably produces imprecise localization and wrong classification. To address these issues, we propose a unified interaction consistency learning (UICL) framework, a novel single-source domain-generalized method designed to learn intra-class domain-invariant representations. Specifically, we put forth a cross-domain interaction mechanism to exchange region proposals between original and augmented pipelines, enriching the diversity of instance-level representations. Building upon this, we propose prediction-guided consistency learning to unify the interaction mechanism and harmonize the cross-domain representations, contributing to a discriminative prediction distribution under domain shift. In addition, we devise a cyclic interaction resilient detection strategy, which mitigates inaccurate predictions suffering from partial occlusion and ambiguous boundaries among different domains. Extensive experiments evidence that UICL significantly improves the robustness of detectors over several target domains, achieving state-of-the-art generalization performance on the diverse weather benchmark. Peng Zhang 0121, Gong Cheng 0003 |
AAAI | 3 |
| 2026 | Uncertainty-Aware and Decoupled Distillation for Semantic Segmentation
Gong Cheng 0003, Junwei Han 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | Evidential Robust Feature Learning for Generalized Few-Shot Segmentation
Weide Liu, Xiaoyang Zhong, Lu Wang 0001, Chunbo Lang, Yuming Fang 0001, Jun Cheng 0003, Xulei Yang, Gong Cheng 0003 |
Int. J. Comput. Vis. | 8 |
| 2026 | Relaxed Knowledge Distillation
Xiwen Yao, Xuguang Yang, Gong Cheng 0003, Junwei Han 0001 |
Int. J. Comput. Vis. | 6 |
| 2026 | Dual-perspective filter pruning via diversity and independence collaboration
Chenyang Gao, Qinglong Cao, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003 |
Pattern Recognit. | 5 |
| 2026 | Mining representative tokens via transformer-based multi-modal interaction for RGB-T tracking
Pujian Lai, Shilei Wang 0001, Gong Cheng 0003 |
Pattern Recognit. | 4 |
| 2026 | Generative attack in complex real-world scenarios
Hongyu Peng, Gong Cheng 0003, Xuxiang Sun 0001 |
Pattern Recognit. | 2 |
| 2026 | Cross-alignment for efficient visual object tracking
Shilei Wang 0001, Mingjiang Liang, Shaoli Huang, Jifeng Ning, Gong Cheng 0003 |
Pattern Recognit. | 5 |
| 2026 | Semantic contrastive learning via VLM for few-shot remote sensing object detection
Bowei Yan, Chunbo Lang, Gong Cheng 0003 |
Pattern Recognit. | 3 |
| 2026 | Domain adaptation for remote sensing image semantic segmentation with prototype-driven domain disentangle alignment
Xiufei Zhang, Yuanwei Liu, Xiaoliang Qian, Gong Cheng 0003, Xiwen Yao |
Pattern Recognit. | 4 |
| 2026 | Change Detection Mamba With Boundary-Specific SupervisionabstractEmphasis on modeling visual context underpins the current high-performance dense prediction models, including those for change detection. However, excessive context modeling tends to cause ambiguous feature representations around object boundaries and possibly overwhelms small or thin objects. This encourages maintaining large feature maps to provide sufficient information for an objective rescue. With the goal of efficiently harvesting global context from large feature maps while getting good sensitivity to change boundaries as well as small or thin changes, we propose a change detection model that features (i) an encoder-decoder architecture with state space model-based feature refinement, and (ii) boundary-specific supervision. Our encoder-decoder is equipped with novel modulated Mamba blocks capable of preserving local correlation and then achieving local-global context mixing. To bridge the gap between local and global information during modulation, we employ a spectral transform on the local features to holistically enhance the information encoded in key frequency components. Moreover, our custom-designed boundary-specific supervision explicitly induces the revision of change boundaries. With these improvements, our Mamba-grounded change detection model can efficiently garner boundary-sensitive large feature maps applicable to various change shapes and scales. Extensive experimental results on four public change detection datasets demonstrate that our method consistently outperforms state-of-the-art competitors in terms of key evaluation metrics. Our source code is available at https://github.com/xingronaldo/BSSMamba. Guangxing Wang 0001, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Weakly Supervised Object Detection for Aerial Images With Instance-Aware Label Assignment
Gong Cheng 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | TPTAF: Task-Prior Tripartite Attention for Infrared and Visible Image FusionabstractInfrared and visible image fusion aims to integrate complementary information to produce informative fused images for visual perception and downstream tasks. However, mainstream methods often focus on fusion reconstruction, while losses from different tasks may introduce competing optimization demands, making it difficult to balance visual quality and detection performance. To address this issue, we propose task-prior tripartite attention for infrared and visible image fusion (TPTAF), which introduces detection semantics as task-prior guidance for representation-level cross-modal interaction. TPTAF employs a decoupled encoder to organize infrared and visible features into structural and discriminative spaces, enabling cross-modal layout consistency and modality-specific detail preservation to be modeled with different roles. Meanwhile, detection semantics extracted from hybrid infrared-visible features are transformed into task priors and integrated with the decoupled representations through tripartite attention. In this interaction, structural cues stabilize the fusion layout, while task-prior guidance regulates the selection of discriminative details toward target-related evidence. For joint optimization, uncertainty-weighted learning further balances fusion and detection losses, reducing the dependence on manually assigned loss weights. Experiments on M3FD, RoadScene, AVMS, and MSRS demonstrate that TPTAF maintains stable fusion quality across different data distributions while improving the utility of fused images for downstream object detection and semantic segmentation evaluation. The code will be made available at https://github.com/Duxeno/TPTAF. Xuyang Du, Xiwen Yao, Ankang Zang, Gong Cheng 0003 |
IEEE Trans. Image Process. | 4 |
| 2026 | Unc-SOD: An Uncertainty Learning Framework for Small Object DetectionabstractSmall object detection (SOD) constitutes a notable yet immensely arduous task, stemming from the restricted informative regions inherent in size-limited instances, which further sparks off heightened uncertainty beyond the capacity of current two-stage detectors. Specifically, the intrinsic ambiguity in small objects undermines the prevailing sampling paradigms and may mislead the model to devote futile effort to those unrecognizable targets, while the inconsistency of features utilized for the detection at two stages further exposes the hierarchical uncertainty. In this paper, we develop an Uncertainty learning framework for Small Object Detection, dubbed as Unc-SOD. By incorporating an auxiliary uncertainty branch to conventional Region Proposal Network (RPN), we model the indeterminacy at instance-level which later on serves as a surrogate criterion for sampling, thereby unearthing adequate candidates dynamically based on the varying degrees of uncertainty and facilitating the learning of proposal networks. In parallel, a Perception-and-Interaction strategy is devised to capture rich and discriminative representations, through optimizing the intrinsic properties from the regional features at the original pyramid and the assigned one, in which the perceptual process unfolds in a mutual paradigm. As the seminal attempt to model uncertainty in SOD task, our Unc-SOD yields state-of-the-art performance on two large-scale small object detection benchmarks, SODA-D and SODA-A, and the results on several SOD-oriented datasets including COCO, VisDrone, and Tsinghua-Tencent 100K also exhibit the promotion to baseline detector. This underscores the efficacy of our approach and its superiority over prevailing detectors when dealing with small instances. Gong Cheng 0003, Jiacheng Cheng 0001, Ruixiang Yao, Junwei Han 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | $\Phi$-GAN: Physics-Inspired GAN for Generating SAR Images Under Limited Data
Xidan Zhang, Yihan Zhuang, Haodong Yang, Xuelin Qian, Gong Cheng 0003, Junwei Han 0001, Zhongling Huang |
ICCV | 6 |
| 2025 | Multi-State Tracker: Enhancing Efficient Object Tracking via Multi-State Specialization and InteractionabstractEfficient trackers achieve faster runtime by reducing computational complexity and model parameters. However, this efficiency often compromises the expense of weakened feature representation capacity, thus limiting their ability to accurately capture target states using single-layer features. To overcome this limitation, we propose Multi-State Tracker (MST), which utilizes highly lightweight state-specific enhancement (SSE) to perform specialized enhancement on multi-state features produced by multi-state generation (MSG) and aggregates them in an interactive and adaptive manner using cross-state interaction (CSI). This design greatly enhances feature representation while incurring minimal computational overhead, leading to improved tracking robustness in complex environments. Specifically, the MSG generates multiple state representations at multiple stages during feature extraction, while SSE refines them to highlight target-specific features. The CSI module facilitates information exchange between these states and ensures the integration of complementary features. Notably, the introduced SSE and CSI modules adopt a highly lightweight hidden state adaptation-based state space duality (HSA-SSD) design, incurring only 0.1 GFLOPs in computation and 0.66 M in parameters. Experimental results demonstrate that MST outperforms all previous efficient trackers across multiple datasets, significantly improving tracking accuracy and robustness. In particular, it shows excellent runtime performance, with an AO score improvement of 4.5% over the previous SOTA efficient tracker HCAT on the GOT-10K dataset. The code is available at https://github.com/wsumel/MST. Shilei Wang 0001, Gong Cheng 0003, Pujian Lai, Junwei Han 0001 |
ACM Multimedia | 2 |
| 2025 | Propagation rectified attack: on improving adversarial transferability
Xuxiang Sun 0001, Hongyu Peng, Gong Cheng 0003, Junwei Han 0001 |
Sci. China Inf. Sci. | 3 |
| 2025 | Learning Compact Discriminant Representation via Low-Rank Bilinear PoolingabstractIn this paper, we explain the mechanism of bilinear pooling as a module of hard sample generation, and find that bilinear pooling significantly expands variances of the first-order vectors when it produces discriminative bilinear features. In conjunction with the extremely high dimensionality of the obtained bilinear features, those variances lead to overfitting in subsequent learning models. To solve this issue, we construct a bi-level optimization problem, where the high-level problem is the supervised classification loss, and the low-level problem is the principal component analysis (PCA). Then, we find that PCA on bilinear features is equivalent to spectral clustering, which allows us to mathematically prove that the first $\log _{2}(C)$log2(C) principal components can support the discriminant information of $C$C classes. By removing the rest principal components, the dimensionality and variances are simultaneously reduced. To the best of our knowledge, this is the first work providing a lower bound for dimension reduction for bilinear pooling. However, the PCA projection matrix $\mathbf{L}$L is prone to overfitting due to having many parameters. To address this issue, we propose a rank-$k$k general bilinear projection (RK-GBP) that decomposes $\mathbf{L}$L into two small matrices $\mathbf{U}$U and $\mathbf{V}$V, whose learnable parameters are smaller. Different from traditional bilinear projections used in factorized bilinear pooling (FBiP), our RK-GBP can preserve the orthogonality of columns in $\mathbf{L}$L by constraining the orthogonality of columns in $\mathbf{U}$U and $\mathbf{V}$V. For computational efficiency, we relax the PCA in the low-level task into a dictionary learning problem, obtaining the rank-$k$k orthogonal factorization bilinear pooling (RK-OFBP). The RK-OFBP can be considered as a general form of current factorization bilinear pooling methods (e.g., Hadamard product-based ones). Finally, we evaluate our approach on fine-grained images and large-scale datasets, demonstrating that our proposed method not only produces extremely low-dimensional features but also outperforms other methods in classification tasks. For example, our RK-OFBP can employ 32-dimensional vectors to achieve comparable results to B-CNN (Lin, 2015) (dimension: 512*512) for the 200-class classification task. Kun Song 0001, Gong Cheng 0003, Junwei Han 0001, Feiping Nie 0001, Bin Gu 0001, Fakhri Karray |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | STDatav2: Accessing Efficient Black-Box Stealing for Adversarial AttacksabstractOn account of the extreme settings, stealing the black-box model without its training data is difficult in practice. On this topic, along the lines of data diversity, this paper substantially makes the following improvements based on our conference version (dubbed STDatav1, short for Surrogate Training Data). First, to mitigate the undesirable impacts of the potential mode collapse while training the generator, we propose the joint-data optimization scheme, which utilizes both the synthesized data and the proxy data to optimize the surrogate model. Second, we propose the self-conditional data synthesis framework, an interesting effort that builds the pseudo-class mapping framework via grouping class information extraction to hold the class-specific constraints while holding the diversity. Within this new framework, we inherit and integrate the class-specific constraints of STDatav1 and design a dual cross-entropy loss to fit this new framework. Finally, to facilitate comprehensive evaluations, we perform experiments on four commonly adopted datasets, and a total of eight kinds of models are employed. These assessments witness the considerable performance gains compared to our early work and demonstrate the competitive ability and promising potential of our approach. Xuxiang Sun 0001, Gong Cheng 0003, Chunbo Lang, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Physics-Guided Detector for SAR AirplanesabstractThe disperse structure distributions (discreteness) and variant scattering characteristics (variability) of SAR airplane targets lead to special challenges of object detection and recognition. The current deep learning-based detectors encounter challenges in distinguishing fine-grained SAR airplanes against complex backgrounds. To address it, we propose a novel physics-guided detector (PGD) learning paradigm for SAR airplanes that comprehensively investigate their discreteness and variability to improve the detection performance. It is a general learning paradigm that can be extended to different existing deep learning-based detectors with ”backbone-neck-head” architectures. The main contributions of PGD include the physics-guided self-supervised learning, feature enhancement, and instance perception, denoted as PGSSL, PGFE, and PGIP, respectively. PGSSL aims to construct a self-supervised learning task based on a wide range of SAR airplane targets that encodes the prior knowledge of various discrete structure distributions into the embedded space. Then, PGFE enhances the multi-scale feature representation of a detector, guided by the physics-aware information learned from PGSSL. PGIP is constructed at the detection head to learn the refined and dominant scattering point of each SAR airplane instance, thus alleviating the interference from the complex background. We propose two implementations, denoted as PGD and PGD-Lite, and apply them to various existing detectors with different backbones and detection heads. The experiments demonstrate the flexibility and effectiveness of the proposed PGD, which can improve existing detectors on SAR airplane detection with fine-grained classification task (an improvement of 3.1% mAP most), and achieve the state-of-the-art performance (90.7% mAP) on SAR-AIRcraft-1.0 dataset. The project is open-source at https://github.com/XAI4SAR/PGD. Zhongling Huang, Shuxin Yang, Zhirui Wang 0003, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Learning Discriminative Representation for Fine-Grained Object Detection in Remote Sensing ImagesabstractFine-grained object detection (FGOD) in remote sensing images is an emerging and challenging task in the field of image intelligent interpretation. It aims to localize objects while classifying them into different fine-grained categories. Modern FGOD methods are mainly derived from well-developed detectors and have made compelling progress. Despite this, these methods struggle to perform well in classifying objects at the subordinate level due to the limitations of their representation manners. In this paper, we propose a network capable of learning discriminative representation (DR) for fine-grained object detection in remote sensing images, named DRNet. First, a fine-grained branch that works in parallel with other task branches is introduced, where objects’ features are re-encoded with dual refinement to generate discriminative representation, enabling accurate fine-grained classification. Second, we design a confusion-minimized loss that automatically scales loss contributions according to the separability of samples to train the fine-grained branch, further boosting discriminative ability of the representation and better addressing hard-to-distinguish objects. Moreover, we devise an interaction verification strategy that empowers the network to fully utilize the results of fine-grained classification and coarse classification for achieving robust inference. On large-scale FAIR1M-1.0 and FAIR1M-2.0 datasets, our DRNet with ResNet50 and$1\times $training schedule obtains 40.87% mAP and 47.04% mAP, respectively, establishing new state-of-the-arts for fine-grained object detection in remote sensing images. The source code is available athttps://github.com//54wb//DRNet. Xingxing Xie, Gong Cheng 0003, Chunbo Lang, Peng Zhang 0121, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Centric Probability-Based Sample Selection for Oriented Object DetectionabstractIn object detection, particularly within remote sensing images, the quality of selected samples is crucial for the accuracy and robustness of detection models. However, current sampling strategies demonstrate inherent limitations. They empirically define positive sample sets using fixed thresholds or preset areas, ignoring the actual shapes of the objects and failing to distinguish the intrinsic value of each sample point. To address these critical issues, this article proposes a novel centric probability-based sample selection approach that includes centering probability mapping (CPM), Expectation-Maximization-based boundary optimization (EBO), and probabilistic random sampling (PRS) technologies. Specifically, the CPM is constructed to assign various confidence levels for all sample points based on their proximity to the center of bounding box, effectively discerning the value of individual samples. Then, the EBO is utilized to dynamically optimize the boundaries for positive and negative samples based on the EM algorithm, thus avoiding the sample imbalance problem associated with empirical thresholds. Finally, the PRS strategy is proposed to select training samples from the sample space constructed by CPM and EBO in a manner of random probability sampling, which could improve the diversity of samples while guaranteeing their quality. Experimental validation on three remote sensing image datasets, including DOTA-v1.0, DOTA-v2.0, and DIOR-R, demonstrates that our method achieves robust performance improvements over baseline and significantly surpasses the advanced sample selection methods. The source code will be available athttps://github.com/yanqingyao1994/CPSS. Gong Cheng 0003, Chunbo Lang, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Semantic Differentiation Aids Oriented Small Object DetectionabstractDetecting small, oriented objects in remote sensing images remains a bottleneck for prevailing detection paradigms. The discriminative cues essential for detecting small instances are often inaccessible owing to the restrained spatial extent and poor visual responses, which further compromises the model and necessitates reliance on low-level patterns for identification and localization, exacerbating vulnerability to structural distortions and intra-class confusion especially in complex scenarios. To address these desiderata, we devise a Semantic Differentiation (SemDiff) framework for oriented small object detection in remote sensing images. Starting with randomly initialized category-specific units, we deliver a differentiation pipeline where distinctive features steer the evolution of these embeddings via a tailored differentiation loss. Afterwards, these class-aligned vectors function as dynamic kernels, infusing hierarchical representations with semantic understanding. Moreover, an improved centerness metric that is more accommodating to size-constrained instances is introduced. Building upon this, we design an instance-level recalibration mechanism to regulate the training process, thereby ensuring adequate optimization even for exceptionally small instances. By integrating semantic in an explicit fashion, our SemDiff efficiently facilitates the discriminative capabilities of hierarchical features, thereby revitalizing foreground responses and alleviating semantic-level ambiguity. On the challenging small object detection benchmarks SODA-A and Tiny-DOTA, our approach outstrips prevailing single-stage paradigms by a substantial margin, and achieves competitive performance to its two-stage counterparts, but with an edge of speed. Codes will be available athttps://github.com/shaunyuan22/SemDiff. Gong Cheng 0003, Ruixiang Yao, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Multi-Scale Oriented Object Detection With Focus Error Ellipse LossabstractThe loss function and feature extraction framework are essential parts of the algorithm design and significantly affect the accuracy of oriented object detection in remote sensing images. Though considerable progress has been made, there are still challenges left to be explored, e.g., large variations in scales, arbitrary direction, and dense distribution of the objects, which may have some undesirable effects, such as inaccurate object position regression, high false alarm, and miss rate. To address the above problems, we propose a Focus Error Ellipse (FEE) loss function. This function bolsters the detection accuracy by narrowing the distance between the center points of the labeled and predicted bounding boxes based on the Error Ellipse. For the network part, we carefully crafted two unit modules: a Fine-grained and Context-augmented Module (FCM) and a Semantic Information Regrouping Module (SIRM). The FCM aligns fine-grained information with contextual information to establish dependencies between local and global features, which helps grasp the more holistic characteristics of objects. The SIRM reorganizes the acquired deep semantic features in the channel dimension, enhances the weight of task-beneficial semantic information, and further derives the optimal combination method of feature subsets for object detection. Based on the aforementioned work, we developed an oriented object detection framework, which further improves the detection accuracy of large aspect ratio objects and dense scenes. Experimental results show that the proposed method can produce competitive performance in oriented object detection compared to other state-of-the-art models. Xuanbei Lu, Ke Li 0005, Gong Cheng 0003, Xiong You |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Incorporating Multiscale Context and Task-Consistent Focal Loss into Oriented Object DetectionabstractOriented object detection (OOD) in remote sensing images (RSIs) aims to precisely localize and identify objects with arbitrary orientations. Two-stage OOD methods attract lots of interest due to their superior accuracy, however, they still face two major problems. First of all, the misclassification problem frequently occurs because the majority of classification strategies solely relies on the features of proposals. Secondly, most of loss functions cannot simultaneously concentrate on hard samples and boost the consistency between identification and localization, which restricts the further improvement of OOD models. To address the first problem, the multi-scale context (MSC) is incorporated into a two-stage OOD model in this paper. Specifically,Ncontextual branches are added to predict the class confidence score (CCS) of each proposal and itsNenlarged proposals which contain the MSC, and the final CCS of each proposal is determined by the mean value of aboveN+ 1 CCSs. To tackle the second problem, a task-consistent focal (TF) loss is proposed. The TF loss employs the difficulty of localization as the weight of classification loss, and the difficulty of identification is used as the weight of regression loss. Concentrating on hard samples and synchronous optimization of classification and regression can be achieved by minimizing the TF loss. The ablation studies show the validity of MSC, TF and their combination. The comparison with popular OOD models demonstrates the superior performance of our model on the DOTA and DIOR-R datasets. The source code can be obtained from https://github.com/qxlzengli/MSC-TF. Xiaoliang Qian, Qingqing Jian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | IPS-YOLO: Iterative Pseudo-Fully Supervised Training of YOLO for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised object detection (WSOD) in remote sensing images only requires image-level labels, greatly reducing the cost of manual annotations. Recently, the pseudo-fully supervised object detection (pseudo-FSOD) models give better performance than traditional WSOD models based on multiple instance learning. However, the existing pseudo-FSOD models still have two problems that need to be addressed. Firstly, existing models tend to focus on the salient parts of object rather than the whole object. Secondly, the pseudo ground truth (PGT) cannot be continuously improved during the training process. For the first problem, a filtering and weighted synthesis guided by category confidence score (FWSC) strategy is proposed to generate PGT. The FWSC strategy firstly removes the instances with low category confidence score (CCS), then, to cover the whole object as much as possible, any two instances are weighted and synthesized according to their CCSs if they have very high spatial overlap, otherwise, the non-maximum suppression (NMS) operation is conducted. For the second problem, an iterative refinement (IR) scheme of PGT is proposed. Specifically, the PGT instances produced by the FWSC strategy are firstly used to train a YOLO model, then a filtering and weighted synthesis guided by IoU (FWSI) strategy employs the detection results inferred from the trained YOLO model to refine the PGT instances, and above two steps can be repeated multiple times. Furthermore, three improvement strategies are proposed to enhance the traditional baseline WSOD model in proposals generation, the selection of augmented samples, and the definition of pseudo-labels, respectively. The ablation studies demonstrate the effectiveness of FWSC strategy, IR scheme of PGT, and the three improvement strategies of baseline model. The comparisons with popular WSOD models show that our model gives the best results on the NWPU VHR-10. v2 and DIOR datasets.The source codes have been released at https://github.com/qxlzengli/IPS-YOLO. Xiaoliang Qian, Baihui Zhang, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Global-Integrated and Drift-Rectified Imprinting for Few-Shot Remote Sensing Object DetectionabstractFew-shot object detection (FSOD) in remote sensing images is a marginally explored but highly challenging task that focuses on identifying unseen classes of objects with a limited number of annotations. Current FSOD approaches often fail to accurately localize the foreground and misalign targets with various orientations, resulting in poor detection performance. For this purpose, we develop a fresh and powerful meta-learning framework based on the idea of imprinting, which leverages tailored support information to model the regional correlation between query and support objects in different stages. Specifically, a global-integrated scheme is first proposed to guide the generation of high-quality proposals by increasing the activation of foreground features and integrating global support information. Considering the orientation discrepancy of objects in query and support sets, we introduce a drift-rectified technique to achieve adaptive alignment by implicitly capturing the positional correspondence between the instances in two sets. In stark contrast to conventional FSOD approaches, our method can extract key clues and establish directional relationships between objects from different training sets, leading to better generalization capability. Extensive experiments on two standard benchmarks (DIOR and NWPU VHR-10.V2) manifest the effectiveness, and our proposed method exhibits superior performance to other competitors with similar motivation. The source code is available athttps://github.com/Ybowei/GIDR Bowei Yan, Gong Cheng 0003, Chunbo Lang, Zhongling Huang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | NIRNet: Noise Incentive Robust Network in Remote Sensing Object Detection Under Cloud CorruptionabstractWithin remote sensing images, complex atmospheric environments commonly bring about distinct variations in imaging visibility and ambient occlusions, significantly transforming the appearance of objects. Nevertheless, modern detectors generally struggle to maintain promising accuracy when encountering realistic scenarios. Devoting to alleviating the issues, we develop a noise incentive robust network (NIRNet) for remote sensing object detection under cloud corruption without relying on hazy images for training. The proposed NIRNet preserves discriminative representations and calibrates them using an incentive mechanism. Firstly, we design a noise perception module (NPM) to deal with diverse cloud corruption types, which generates point-wise calibration weights dependent on the perceived discrepancy between objects and environmental noise. Secondly, aiming to detect difficult-to-discern objects thoroughly, a dual-path incentive calibration (DPIC) strategy is proposed to combine intensity and stability features weighted by NPM. Profiting from its universal design, the DPIC could be treated as a plug-and-play module for existing detectors, enhancing robustness against adverse weather. To evaluate the reliability of aerial detectors under intricate cloud corruptions, we present an elaborate Hazy-DIOR dataset, which contains numerous images with different cloud conditions and severity levels. Finally, extensive experiments on the Hazy-DIOR and DOTA-Cloud datasets simultaneously demonstrate the robustness of NIRNet, which especially achieves state-of-the-art accuracy and gets 2.16% mAP and 2.56% rPC improvements on the Hazy-DIOR compared to solid Oriented R-CNN detector. The code is available at https://github.com/zhangpeng2001/nirnet. Peng Zhang 0121, Gong Cheng 0003, Chunbo Lang, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Cross-Modality Domain Adaptation Based on Semantic Graph Learning: From Optical to SAR ImagesabstractSynthetic aperture radar (SAR) imaging provides a distinct advantage in scene understanding due to its capability for all-weather data acquisition. However, in comparison to easily annotated optical remote sensing images, the lower imaging quality of SAR images presents significant challenges in obtaining manually annotated training data, which poses substantial issues for SAR image analysis. In this paper, we employ the domain adaptation (DA) that leverages labeled optical images to better understand unlabeled SAR images. Global feature alignment as a method for DA has demonstrated effectiveness in transferring knowledge, yet it faces challenges in cross-modality adaptation from optical remote sensing to SAR images due to their differing imaging mechanisms. With distinct visual features between optical and SAR images, the semantic dependency is difficult to construct, which results in low-quality pseudo-label assignment for SAR images. To address the above issue, we propose a semantic graph learning framework to comprehensively align the global features of optical remote sensing and SAR images by modeling the cross-modality semantics and generating high-quality pseudo-labels. It can be applied for SAR scene classification and object detection when only optical remote sensing images are labeled. Specifically, a cross-modality semantic graph alignment (CSGA) module is constructed to model and align the second-order semantic dependencies by aggregating cross-modality visual semantic information. Then, an uncertainty-based robust pseudo-label generation (URPG) module is designed to generate pseudo-labels for effective semantic alignment and self-training by modeling the uncertainty of pseudo-labels for each SAR image. Comprehensive experiments show that our proposed method outperforms the state-of-the-art methods on scene classification (NWPU-RESISC45→WHU-SAR6, MLRSNet→NWPU-SAR6, MLRSNet→NWPU-SAR6, and NWPU-RESISC45→NWPU-SAR6) and object detection (MASATI-ship→SSDD, MVSRD→SARDet-vehicle, and DIOR-airplane→SAR-airplane) tasks. The code and datasets are publicly accessible at https://github.com/XZhang878/SGLF. Xiufei Zhang, Zhongling Huang, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | X-Fake: Juggling Utility Evaluation and Explanation of Simulated SAR ImagesabstractSynthetic aperture radar (SAR) image simulation has attracted much attention due to its great potential to supplement the scarce training data for deep learning algorithms. Consequently, evaluating the quality of the simulated SAR image is crucial for practical applications. The current literature primarily uses image quality assessment (IQA) techniques for evaluation that rely on human observers' perceptions. However, because of the unique imaging mechanism of SAR, these techniques may produce evaluation results that are not entirely valid. The distribution inconsistency between real and simulated data is the main obstacle that influences the utility of simulated SAR images. To this end, we propose a novel trustworthy utility evaluation framework with a counterfactual explanation for simulated SAR images for the first time, denoted as X-Fake. It unifies a probabilistic evaluator and a causal explainer to achieve a trustworthy utility assessment. We construct the evaluator using a probabilistic Bayesian deep model to learn the posterior distribution, conditioned on real data. Quantitatively, the predicted uncertainty of simulated data can reflect the distribution discrepancy. We build the causal explainer with an introspective variational auto-encoder (IntroVAE) to generate high-resolution counterfactuals. The latent code of IntroVAE is finally optimized with evaluation indicators and prior information to generate the counterfactual explanation, thus revealing the inauthentic details of simulated data explicitly. The proposed framework is validated on four simulated SAR image datasets obtained from electromagnetic models and generative artificial intelligence approaches. The results demonstrate the proposed X-Fake framework outperforms other IQA methods in terms of utility. Furthermore, the results illustrate that the generated counterfactual explanations are trustworthy, and can further improve the data utility in applications. Zhongling Huang, Yihan Zhuang, Zipei Zhong, Feng Xu 0001, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Fewer is more: efficient object detection in large aerial images
Xingxing Xie, Gong Cheng 0003, Qingyang Li 0001, Shicheng Miao, Ke Li 0005, Junwei Han 0001 |
Sci. China Inf. Sci. | 2 |
| 2024 | Few-Shot Segmentation via Divide-and-Conquer Proxies
Chunbo Lang, Gong Cheng 0003, Binfei Tu, Junwei Han 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Oriented R-CNN and Beyond
Xingxing Xie, Gong Cheng 0003, Jiabao Wang 0005, Ke Li 0005, Xiwen Yao, Junwei Han 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Boosting Knowledge Distillation via Intra-Class Logit Distribution SmoothingabstractPrevious arts built an intimate link between knowledge distillation (KD) and label smoothing (LS) that they both impose regularization on the model training. In this paper, we delve deeper into investigating the hidden reason rendering KD and LS to exert distinct effects on a model’s potential ability in sequential knowledge transferring. Specifically, we observe that the distilled model typically exhibits much higher intra-class variance than the regularized one, consequentially acting as the better teacher. Then we devise two exploratory experiments and identify that sufficient intra-class variance retained by a teacher model is an implicit distillation recipe for achieving competitive student performance. The observed properties allow us to further put forth a simple yet beneficial approach that promotes intra-class diversity at the optimizing process of the teacher models to accomplish the most promising performance of KD. Extensive experiments are conducted on various image classification tasks across three distillation paradigms, demonstrating our proposed method’s effectiveness and generalization. Additionally, we offer new interpretations to receive a more in-depth cognition of the gap issues,i,e., better teacher, worse student, and the success of multi-generation self-distillation, respectively. Code will be made available at https://github.com/swift1988. Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Task-Specific Importance-Awareness Matters: On Targeted Attacks Against Object DetectionabstractTargeted Attacks on Object Detection (TAOD) aim to deceive the victim detector into recognizing a specific instance as the predefined target category while minimizing the changes to the predicted bounding box of that instance. Yet, this kind of flexible attack paradigm, which is capable of manipulating the decision outcome of the victim detector, received limited attention, especially in the context of attacking object detection in optical remote sensing images, where relevant research remains a blank. To fill this gap, this paper concentrates on TAOD in optical remote sensing images, and pays attention to a fundamental question, how to deploy TAOD via the raw predictions (the predictions before non-maximum suppression) of a victim detector. In this regard, we depart from widely adopted task-independent importance measurements and hard-weighted ensemble optimization schemes present in existing methods. Instead, we first define the task-specific importance score, which considers both the qualities and the attack costs of predictions. Further, we propose the Task-Specific Importance-Aware Candidate Predictions Selection Scheme (TSIA-CPSS) alongside the Soft-Weighted Ensemble Optimization Scheme (SW-EOS). A total of eleven detectors on DIOR and DOTA, two commonly employed benchmarks, are included to comprehensively evaluate our approach. Furthermore, we indicate that the effectiveness of our approach is not only substantial for vanilla TAOD, but also can be better generalized to extended scenarios, which encompasses random TAOD, TAOD on oriented object detection, and targeted patch attacks, highlighting the noteworthy potential of our approach. Our codes will be released on Github. Xuxiang Sun 0001, Gong Cheng 0003, Hongyu Peng, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Cross-Level Attentive Feature Aggregation for Change DetectionabstractThis article studies change detection within pairs of optical images remotely sensed from overhead views. We consider that a high-performance solution to this task entails highly effective multi-level feature interaction. With that in mind, we propose a novel approach characterized by two attentive feature aggregation schemes that handle cross-level features in different processes. For the Siamese-based feature extraction of the bi-temporal image pair, we attach emphasis on constructing semantically strong and contextually rich pyramidal feature representations to enable comprehensive matching and differencing. To this end, we leverage a feature pyramid network and re-formulate its cross-level feature merging procedure as top-down modulation with multiplicative channel attention and additive gated attention. For the multi-level difference feature fusion, we progressively fuse the derived difference feature pyramid in an attend-then-filter manner. This makes the high-level fused features and the adjacent lower-level difference features constrain each other, and thus allows steady feature fusion for specifying change regions. In addition, we build an upsampling head as a replacement for the normal heads followed by static upsampling. Our implementation contains a stack of upsampling modules that allocate features for each pixel. Each has a learnable branch that produces attentive residuals for refining the statically upsampled results. We conduct extensive experiments on four public datasets and results show that our approach achieves state-of-the-art performance. Code is available at https://github.com/xingronaldo/CLAFA. Guangxing Wang 0001, Gong Cheng 0003, Peicheng Zhou, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Retentive Compensation and Personality Filtering for Few-Shot Remote Sensing Object DetectionabstractIn recent years, few-shot object detection (FSOD) in remote sensing images has attracted increasing attention. Numerous studies address the challenges posed by both intra-class and inter-class variance through strategies such as augmenting sample diversity and incorporating multi-scale features. However, these features still encompass a considerable amount of noise attributes due to the complex characteristic of satellite images, persistently and adversely affecting classification. In contrast, we advocate for the belief that a limited yet refined set of features surpasses a multitude of coarse features. Accordingly, we tackle above issues through the meticulous refinement of representative category features, enhancing performance by eliminating irrelevant attributes that interfere with classification. Specifically, two pivotal modules: retentive compensation module (RCM) and personality filtering module (PFM), are introduced. The former module RCM systematically scrutinizes features proximate to the category center, yielding prototypes that exhibit both intra-class compactness and inter-class distinctiveness. Furthermore, the latter module PFM utilizes previous obtained prototypes to supervise the filtering process, diminishing the intra-class variance by excluding personality features which could impede the classification task. The integration of the above two modules enables a holistic feature representation, capturing inherent similarities within individual classes while accentuating distinctions between classes. Experiments have been conducted on the DIOR and NWPU VHR-10.v2 datasets, and the results demonstrate that our proposed approach exceeds several state-of-the-art methods. Code is available at https://github.com/yomik-js/RP-FSOD. Jiashan Wu, Chunbo Lang, Gong Cheng 0003, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Understanding Negative Proposals in Generic Few-Shot Object DetectionabstractRecently, Few-Shot Object Detection (FSOD) has received considerable research attention as a strategy for reducing reliance on extensively labeled bounding boxes. However, current approaches encounter significant challenges due to the intrinsic issue of incomplete annotation while building the instance-level training benchmark. In such cases, the instances with missing annotations are regarded as background, resulting in erroneous training gradients back-propagated through the detector, thereby compromising the detection performance. To mitigate this challenge, we introduce a simple and highly efficient method that can be plugged into both meta-learning-based and transfer-learning-based methods. Our method incorporates two innovative components: Confusing Proposals Separation (CPS) and Affinity-Driven Gradient Relaxation (ADGR). Specifically, CPS effectively isolates confusing negatives while ensuring the contribution of hard negatives during model fine-tuning; ADGR then adjusts their gradients based on the affinity to different category prototypes. As a result, false-negative samples are assigned lower weights than other negatives, alleviating their harmful impacts on the few-shot detector without the requirement of additional learnable parameters. Extensive experiments conducted on the PASCAL VOC and MS-COCO datasets consistently demonstrate that our method significantly outperforms both the baseline and recent FSOD methods. Furthermore, its versatility and efficiency suggest the potential to become a stronger new baseline in the field of FSOD. Code is available at https://github.com/Ybowei/UNP. Bowei Yan, Chunbo Lang, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Hierarchical Mask Prompting and Robust Integrated Regression for Oriented Object DetectionabstractObject detection in remote sensing images has garnered significant attention due to its wide applications in real-world scenarios. However, most existing oriented object detectors still suffer from complex backgrounds and varying angles, limiting their performance to further improvement. In this paper, we propose a novel oriented detector withHierarchical mask prompting andRobust integrated regression, termed HRDet. Specifically, to cope with the first issue, we construct a hierarchical mask prompting module consisting of a semantic mask prediction branch and hierarchical Softmax technique. The former aims to isolate object instances from cluttered interferences guided by coarse box-wise masks, while the latter propagates differentiated features for adjacent layers using hierarchical attentive weights. To deal with the second issue, we strive for robust integrated regression and formulate an efficient oriented IoU loss, explicitly measuring the discrepancies of three geometric factors in oriented regression, i.e., the central point distance, side length, and angle. This innovative loss intends to overcome the problem that existing IoU-based losses are invariant during the regression of varying angles. We applied these two strategies to a simple one-stage detection pipeline, achieving a new level of trade-off between speed and accuracy. Extensive experiments on four large aerial imagery datasets, DOTA-v1.0, DOTA-v2.0, DIOR-R, and HRSC2016, demonstrate that our HRDet significantly improves the accuracy of the one-stage detector over refine-stage counterparts while maintaining the efficiency advantage. The source code will be available athttps://github.com/yanqingyao1994/HRDet. Gong Cheng 0003, Chunbo Lang, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | DIMA: Digging Into Multigranular Archetype for Fine-Grained Object DetectionabstractFine-grained remote sensing object detection aims at precisely locating objects and determining the fine-level categories. This task is exceptionally challenging due to the substantial interclass similarity, presenting difficulties in capturing discriminative features. We attribute this to the absence of essential information that can serve as supervision for the learning. This involves comprehensive visual patterns of objects and intrinsic relationships of multigranular features. In this article, we propose a novel scheme dubbed as digging into the multigranular archetype (DIMA) for fine-grained remote sensing object detection. In detail, we first design a simple yet effective frequency-aware representation supplement (FARS) mechanism learning from original images and their auxiliary frequency counterparts simultaneously. The FARS introduces high- and low-frequency representations to reinforce a range of visual cues, such as particular regions associated with the former and contours of objects related to the latter. Then, we further devise a module named hierarchical classification paradigm (HCP), which constructs the interhierarchy relationships between coarse and fine-level representations and then exploits them to guide fine-grained feature enhancement. HCP eventually selects and boosts samples that are hard to discriminate by keeping consistency in multilevels. Our method can be easily integrated into prevailing oriented object detectors and brings consistent performance improvements across these detectors. Notably, our method combined with oriented RCNN (ORCNN) achieves 44.44% (+3.62%) on the FAIR1M and 91.0% (+6.9%) on the MAR20. Moreover, thoughtful discussions about qualitative results and rich visualizations are provided to intuitively underscore the superiority of our approach. The source code is available athttps://github.com/chengjc2019/DIMA. Jiacheng Cheng 0001, Xiwen Yao, Xuguang Yang, Xiaoxu Feng, Gong Cheng 0003, Xiankai Huang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Target-Aware Transformer for Satellite Video Object TrackingabstractRecent years have witnessed the astonishing development of transformer-based paradigm in single object tracking (SOT) in generic videos. However, due to the fact that the targets of interest in satellite videos are small in size and weak in visual appearance, the advancements of transformer-based paradigm in satellite video object tracking are impeded. To alleviate this issue, a novel transformer-based recipe is proposed, which consists of a bi-direction propagation and fusion (Bi-PF) strategy and a target-aware enhancement (TAE) module. Concretely, we first adopt the Bi-PF strategy to make full use of multiscale information to generate discriminative representations of tracking targets. Then, the TAE module is employed to decouple an object query into content-aware embedding and spatial-aware embedding and produce a target prototype to help get high-quality content-aware embedding. It is worth mentioning that, different from the previous methods in satellite video tracking most of which evaluate their performance using only several videos, we conduct extensive experiments on the SatSOT dataset which consists of 105 videos. In particular, the proposed method achieves the success score of 45.6% and the precision score of 57.6%, surpassing the baseline method by 5.0% and 9.5%, respectively. The code will be released athttps://github.com/laybebe/TATrans_SVOT. Pujian Lai, Meili Zhang, Gong Cheng 0003, Shengyang Li, Xiankai Huang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Complete and Invariant Instance Classifier Refinement for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised object detection (WSOD) in remote sensing images is used to detect high-value objects by utilizing image-level labels. However, the current models still have two problems. Firstly, the misclassification of neighboring instances is easily occurred because the one-hot label is assigned to all of seed instances and their neighboring instances. Secondly, the supervisory information of each instance classifier refinement (ICR) branch is generated from the predicted class score of upper ICR branch rather than the real label, thus the prediction mistake of each ICR branch will be accumulated with the propagation of supervisory information. To address the first problem, a complete definition of pseudo soft label (CPSL) of instances is proposed to directly train each ICR branch, where the CPSL of seed instances is defined according to the predicted class scores of upper ICR branch, and the CPSL of other instances are determined by the spatial distance weighted feature similarity between them and seed instances. To handle the second problem, an invariant multiple instance learning (IMIL) scheme is proposed to indirectly train each ICR branch by using the real image-level labels. Furthermore, the affine transformations of original image are incorporated into the baseline model to enhance the invariance of our model. The ablation studies verify the effectiveness of CPSL, IMIL and their combination. The quantitative comparisons with popular methods show that the 73.63% (31.08%) mAP and 79.88% (57.52%) CorLoc of our method is the best on the NWPU VHR-10.v2 (DIOR) dataset, and the qualitative comparisons intuitively demonstrate it again. Xiaoliang Qian, Wei Wang 0245, Xiwen Yao, Gong Cheng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Spatial-Channel Attention Transformer With Pseudo Regions for Remote Sensing Image-Text RetrievalabstractRecently, remote sensing image-text retrieval (RSITR) has received significant attention due to its flexible query form and effective management of remote sensing images. However, prior work often relies on compact global features and ignores local features that can reflect salient objects in the images. Moreover, these methods primarily model interactions between features in the spatial domain, which is insufficient for mining the rich semantic information presented in remote sensing images. In this article, we propose a novel spatial-channel attention transformer (SCAT) with pseudo regions to address these issues. Concretely, in order to acquire the fine-grained perception of local objects, we introduce a pseudo region generation (PRG) module that adaptively aggregates grid features with similar semantic information into multiple clusters through a clustering algorithm. These generated cluster centers are able to flexibly and efficiently represent local objects in remote sensing images without relying on sophisticated object detectors. Furthermore, in order to achieve a comprehensive understanding of image semantics information, we carefully construct a novel SCAT. By exploiting spatial and channel attention to explore the dependencies between features at both spatial and channel domains, the proposed SCAT enhances the model’s ability to identify both “where to look” and “what it is,” thereby obtaining a more powerful representation. In addition, SCAT incorporates two novel designs that alleviate the high overhead caused by attention modeling. Extensive experiments on two benchmark datasets, RSICD and RSITMD, fully demonstrate the effectiveness and superiority of our proposed method. Dongqing Wu, Yinxuan Hou, Cuili Xu, Gong Cheng 0003, Lei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Attention Erasing and Instance Sampling for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) trains detectors by only weak labels, aiming to save the burden of expensive bounding box-level annotations. Most previous efforts formulate WSOD as a multiple instance learning (MIL) problem, which is prone to detect discriminative object parts and miss object instances. This article proposes an attention erasing and instance sampling (AE-IS) approach to alleviate the above problems. Concretely, we first apply an attention erasing (AE) scheme to the WSOD model to hide the most discriminative region for capturing the integral extent of the object. Then, we employ an intersection-over-union (IoU)-balanced sampling component toward mining more object instances. Moreover, an instance reweighted loss (IRL) is designed to learn a larger portion of object instances, thereby further enhancing the performance of the object detector. Experimental results demonstrate that our method significantly improves the baseline approach by great margins and achieves competitive performance with the state-of-the-art algorithms on the NWPU VHR-10.v2 (72.0% mAP, 76.1% CorLoc) and DIOR (29.1% mAP, 55.9% CorLoc) datasets. The source code will be available athttps://github.com/XuanX/AE-IS. Gong Cheng 0003, Xiaoxu Feng, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Oriented Object Detection via Contextual Dependence Mining and Penalty-Incentive AllocationabstractOriented object detection in aerial images has made significant advancements propelled by well-developed detection frameworks and diverse representation approaches to oriented bounding boxes. However, within modern oriented object detectors, the insufficient consideration given to certain factors, like contextual priors in aerial images and the sensitivity of the angle regression, hinder further improvement of detection performance. In this paper, we propose a dual-focused detector (DFDet), which simultaneously focuses on the exploration of contextual knowledge and the mitigation of angle sensitivity. Specifically, DFDet contains two novel designs: a contextual dependence mining network (CDMN) and a penalty-incentive allocation strategy (PIAS). CDMN constructs multiple features containing contexts across various ranges with low computational burden, and aggregates them into a compact yet informative representation that empowers the model for robust inference. PIAS dynamically calibrates the angle regression loss with a scalable penalty term determined by the angle regression sensitivity, incentivizing model to boost regression capacity for large aspect ratio objects challenging to be localized accurately. Extensive experiments on four widely-used benchmarks demonstrate the effectiveness of our approach, and new state-of-the-arts for one-stage object detection in aerial images are established. Without bells and whistles, DFDet with ResNet50 achieves 74.71% mAP running at 23.4 FPS on the most widely-used DOTA-v1.0 dataset. The source code is available at https://github.com/DDGRCF/DFDet. Xingxing Xie, Gong Cheng 0003, Chaofan Rao, Chunbo Lang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Modeling of Multiple Spatial-Temporal Relations for Robust Visual Object TrackingabstractRecently, one-stream trackers have achieved parallel feature extraction and relation modeling through the exploitation of Transformer-based architectures. This design greatly improves the performance of trackers. However, as one-stream trackers often overlook crucial tracking cues beyond the template, they prone to give unsatisfactory results against complex tracking scenarios. To tackle these challenges, we propose a multi-cue single-stream tracker, dubbed MCTrack here, which seamlessly integrates template information, historical trajectory, historical frame, and the search region for synchronized feature extraction and relation modeling. To achieve this, we employ two types of encoders to convert the template, historical frames, search region, and historical trajectory into tokens, which are then collectively fed into a Transformer architecture. To distill temporal and spatial cues, we introduce a novel adaptive update mechanism, which incorporates a thresholding component and a local multi-peak component to filter out less accurate and overly disturbed tracking cues. Empirically, MCTrack achieves leading performance on mainstream benchmark datasets, surpassing the most advanced SeqTrack by 2.0% in terms of the AO metric on GOT-10k. The code is available at https://github.com/wsumel/MCTrack. Shilei Wang 0001, Zhenhua Wang 0003, Qianqian Sun, Gong Cheng 0003, Jifeng Ning |
IEEE Trans. Image Process. | 4 |
| 2023 | Small Object Detection via Coarse-to-fine Proposal Generation and Imitation LearningabstractThe past few years have witnessed the immense success of object detection, while current excellent detectors struggle on tackling size-limited instances. Concretely, the well-known challenge of low overlaps between the priors and object regions leads to a constrained sample pool for optimization, and the paucity of discriminative information further aggravates the recognition. To alleviate the aforementioned issues, we propose CFINet, a two-stage framework tailored for small object detection based on the Coarse-to-fine pipeline and Feature Imitation learning. Firstly, we introduce Coarse-to-fine RPN (CRPN) to ensure sufficient and high-quality proposals for small objects through the dynamic anchor selection strategy and cascade regression. Then, we equip the conventional detection head with a Feature Imitation (FI) branch to facilitate the region representations of size-limited instances that perplex the model in an imitation manner. Moreover, an auxiliary imitation loss following supervised contrastive learning paradigm is devised to optimize this branch. When integrated with Faster RCNN, CFINet achieves state-of-the-art performance on the large-scale small object detection benchmarks, SODA-D and SODA-A, underscoring its superiority over baseline detector and other mainstream detection approaches. Code is available at https://github.com/shaunyuan22/CFINet. Gong Cheng 0003, Kebing Yan, Junwei Han 0001 |
ICCV | 2 |
| 2023 | A Shape-Based Quadrangle Detector for Aerial Images
Chaofan Rao, Xingxing Xie, Gong Cheng 0003 |
PRCV (4) | 4 |
| 2023 | Class attention network for image recognition
Gong Cheng 0003, Pujian Lai, Decheng Gao, Junwei Han 0001 |
Sci. China Inf. Sci. | 1 |
| 2023 | Holistic Prototype Activation for Few-Shot SegmentationabstractConventional deep CNN-based segmentation approaches have achieved satisfactory performance in recent years, however, they are essentially big data-driven technologies and are difficult to generalize to unseen categories. Few-shot segmentation is subsequently developed to perform pertinent operations in a low-data regime. Unfortunately, due to the training paradigm and network architecture factors, existing methods are prone to overfit the targets of base categories and yield inaccurate segmentation boundaries, which impedes the research progress to some extent. In this paper, we propose a Holistic Prototype Activation (HPA) network to alleviate these problems. Its novel designs can be summarized in three aspects: 1) A training-free scheme to derive the prior representations of base categories. 2) Prototype Activation Module (PAM) that generates reliable activation maps and well-matched query features by filtering the objects of irrelevant classes with high confidence. 3) Cross-Referenced Decoder (CRD) for interacted feature reweighting and multi-level feature aggregation. Extensive experiments on standard few-shot segmentation benchmarks (PASCAL-5$^{i}$and COCO-20$^{i}$) verify the effectiveness of our method. On top of that, the superior performance on multiple extended tasks, such as weak-label segmentation, zero-shot segmentation, and video object segmentation, also illustrates its flexibility and versatility. Our code is publicly available athttps://github.com/chunbolang/HPA. Gong Cheng 0003, Chunbo Lang, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Towards Large-Scale Small Object Detection: Survey and BenchmarksabstractWith the rise of deep convolutional neural networks, object detection has achieved prominent advances in past years. However, such prosperity could not camouflage the unsatisfactory situation of Small Object Detection (SOD), one of the notoriously challenging tasks in computer vision, owing to the poor visual appearance and noisy representation caused by the intrinsic structure of small targets. In addition, large-scale dataset for benchmarking small object detection methods remains a bottleneck. In this paper, we first conduct a thorough review of small object detection. Then, to catalyze the development of SOD, we construct two large-scale Small Object Detection dAtasets (SODA), SODA-D and SODA-A, which focus on the Driving and Aerial scenarios respectively. SODA-D includes 24828 high-quality traffic images and 278433 instances of nine categories. For SODA-A, we harvest 2513 high resolution aerial images and annotate 872069 instances over nine classes. The proposed datasets, as we know, are the first-ever attempt to large-scale benchmarks with a vast collection of exhaustively annotated instances tailored for multi-category SOD. Finally, we evaluate the performance of mainstream methods on SODA. We expect the released benchmarks could facilitate the development of SOD and spawn more breakthroughs in this field. Gong Cheng 0003, Xiwen Yao, Kebing Yan, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Learning an Invariant and Equivariant Network for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD) is of increasing importance in the community of computer vision as its extensive applications and low manual cost. Most of the advanced WSOD approaches build upon an indefinite and quality-agnostic framework, leading to unstable and incomplete object detectors. This paper attributes these issues to the process of inconsistent learning for object variations and the unawareness of localization quality and constructs a novel end-to-end Invariant and Equivariant Network (IENet). It is implemented with a flexible multi-branch online refinement, to be naturally more comprehensive-perceptive against various objects. Specifically, IENet first performs label propagation from the predicted instances to their transformed ones in a progressive manner, achieving affine-invariant learning. Meanwhile, IENet also naturally utilizes rotation-equivariant learning as a pretext task and derives an instance-level rotation-equivariant branch to be aware of the localization quality. With affine-invariance learning and rotation-equivariant learning, IENet urges consistent and holistic feature learning for WSOD without additional annotations. On the challenging datasets of both natural scenes and aerial scenes, we substantially boost WSOD to new state-of-the-art performance. The codes have been released at: https://github.com/XiaoxFeng/IENet. Xiaoxu Feng, Xiwen Yao, Hui Shen 0005, Gong Cheng 0003, Bin Xiao 0002, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Base and Meta: A New Perspective on Few-Shot SegmentationabstractDespite the progress made by few-shot segmentation (FSS) in low-data regimes, the generalization capability of most previous works could be fragile when countering hard query samples with seen-class objects. This paper proposes a fresh and powerful scheme to tackle such an intractable bias problem, dubbed base and meta (BAM). Concretely, we apply an auxiliary branch (base learner) to the conventional FSS framework (meta learner) to explicitly identify base-class objects, i.e., the regions that do not need to be segmented. Then, the coarse results output by these two learners in parallel are adaptively integrated to derive accurate segmentation predictions. Considering the sensitivity of meta learner, we further introduce adjustment factors to estimate the scene differences between support and query image pairs from both style and appearance perspectives, so as to facilitate the model ensemble forecasting. The remarkable performance gains on standard benchmarks (PASCAL-5$^{i}$, COCO-20$^{i}$, and FSS-1000) manifest the effectiveness, and surprisingly, our versatile scheme sets new state-of-the-arts even with two plain learners. Furthermore, in light of its unique nature, we also discuss several more practical but challenging extensions, including generalized FSS, 3D point cloud FSS, class-agnostic FSS, cross-domain FSS, weak-label FSS, and zero-shot segmentation. Our source code is available athttps://github.com/chunbolang/BAM. Chunbo Lang, Gong Cheng 0003, Binfei Tu, Chao Li 0028, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Mutual-Assistance Learning for Object DetectionabstractObject detection is a fundamental yet challenging task in computer vision. Despite the great strides made over recent years, modern detectors may still produce unsatisfactory performance due to certain factors, such as non-universal object features and single regression manner. In this paper, we draw on the idea of mutual-assistance (MA) learning and accordingly propose a robust one-stage detector, referred as MADet, to address these weaknesses. First, the spirit of MA is manifested in the head design of the detector. Decoupled classification and regression features are reintegrated to provide shared offsets, avoiding inconsistency between feature-prediction pairs induced by zero or erroneous offsets. Second, the spirit of MA is captured in the optimization paradigm of the detector. Both anchor-based and anchor-free regression fashions are utilized jointly to boost the capability to retrieve objects with various characteristics, especially for large aspect ratios, occlusion from similar-sized objects, etc. Furthermore, we meticulously devise a quality assessment mechanism to facilitate adaptive sample selection and loss term reweighting. Extensive experiments on standard benchmarks verify the effectiveness of our approach. On MS-COCO, MADet achieves 42.5% AP with vanilla ResNet50 backbone, dramatically surpassing multiple strong baselines and setting a new state of the art. Xingxing Xie, Chunbo Lang, Shicheng Miao, Gong Cheng 0003, Ke Li 0005, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | SFRNet: Fine-Grained Oriented Object Recognition via Separate Feature RefinementabstractFine-grained oriented object recognition (FGO2R) is a practical need for intellectually interpreting remote sensing images. It aims at realizing fine-grained classification and precise localization with oriented bounding boxes, simultaneously. Our considerations for the task are general but decisive: (i) the extraction of subtle differences carries a big weight in differentiating fine-grained classes, and (ii) oriented localization prefers rotation-sensitive features. In this article, we propose a network with separate feature refinement (SFRNet), in which two transformer-based branches are designed to perform function-specific feature refinement for fine-grained classification and oriented localization, separately. To highlight the discriminative information advantageous to fine-grained classification, we propose a spatial and channel transformer (SC-Former) to capture both the long-range spatial interactions and the key correlations hidden in the feature channels. Besides, we design a Multi-RoI loss (MRL) following the protocol of deep metric learning to enhance the separability of fine-grained classes further. For oriented localization, we integrate the oriented response convolution with the transformer structure (namely, OR-Former) to assist in encoding rotation information during regression. Extensive experimental results validate the effectiveness and robustness of our SFRNet. Without bells and whistles, our SFRNet achieves state-of-the-art performance on the large-scale FAIR1M datasets (FAIR1M-1.0 and FAIR1M-2.0). Code will be available at https://github.com/Ranchosky/SFRNet. Gong Cheng 0003, Qingyang Li 0001, Guangxing Wang 0001, Xingxing Xie, Lingtong Min, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Global Rectification and Decoupled Registration for Few-Shot Segmentation in Remote Sensing ImageryabstractFew-shot segmentation (FSS), which aims to determine specific objects in the query image given only a handful of densely labeled samples, has received extensive academic attention in recent years. However, most existing FSS methods are designed for natural images, and few works have been done to investigate more realistic and challenging applications,e.g., remote sensing image understanding. In such a setup, the complex nature of the raw images would undoubtedly further increase the difficulty of the segmentation task. To couple with potential inference failures, we propose a novel and powerful remote sensing FSS framework with global Rectification and decoupled Registration, termed R2Net. Specifically, a series of dynamically updated global prototypes are utilized to provide auxiliary non-target segmentation cues and to prevent inaccurate prototype activation resulting from the variability between query-support image pairs. The foreground and background information flows are then decoupled for more targeted and tailored object localization, avoiding unnecessary confusion from information redundancy. Furthermore, we impose additional constraints to promote the interclass separability and intraclass compactness. Extensive experiments on the standard benchmark iSAID-5idemonstrate the superiority of the proposed R2Net over state-of-the-art FSS models. The code will be made available. Chunbo Lang, Gong Cheng 0003, Binfei Tu, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Progressive Parsing and Commonality Distillation for Few-Shot Remote Sensing SegmentationabstractIn recent years, few-shot segmentation (FSS) has received widespread attention from scholars by virtue of its superiority in low-data regimes. Most existing research focuses on natural image processing, and very few studies are dedicated to the practical but challenging topic of remote sensing image understanding. Related experimental results show that directly transferring the previously proposed framework to the current domain is prone to produce unsatisfactory results withincomplete objectsandirrelevant distractors. Such phenomena can be attributed to the lack of modules specifically designed for the complex characteristics of remote sensing images,e.g., great intra-class diversity and low target-background contrast. In this paper, we propose a conceptually simple and easy-to-implement framework to tackle the aforementioned problems. Specifically, our innovative design embodies two main aspects: i) the support mask is progressively parsed into multiple valuable sub-regions that can be further exploited to compute local descriptors with segmentation cues about intractable parts; ii) the base-class memories stored in the meta-training phase are replayed and leveraged for the distillation of novel-class prototypes, where the commonalities between classes are adequately explored, more in line with the concept oflearning to learn. These two components, i.e., the progressive parsing module and commonality distillation module, contribute to each other and together constitute the proposed PCNet. We conduct extensive experiments on the standard benchmark to evaluate segmentation performance in few-shot settings. Quantitative and qualitative results illustrate that our PCNet distinctly outperforms previous FSS approaches and sets a new state-of-the-art. Chunbo Lang, Gong Cheng 0003, Binfei Tu, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Instance-Aware Distillation for Efficient Object Detection in Remote Sensing ImagesabstractPractical applications ask for object detection models that achieve high performance at low overhead. Knowledge distillation demonstrates favorable potential in this case by transferring knowledge from a cumbersome teacher model to a lightweight student model. However, previous distillation methods are plagued with massive misleading background information in remote sensing images and ignore investigating the relationships between different instances. In this article, we propose an instance-aware distillation (InsDist for short) method to derive efficient remote sensing object detectors. Our InsDist combines feature-based and relation-based knowledge distillation to make the most of instance-related information in the knowledge transfer from the teacher to the student. On one hand, we propose a parameter-free masking module to decouple instance-related foreground from instance-irrelevant background in multiscale features. On the other hand, we construct the relationships between different instances to enhance the learning of intraclass compactness and interclass dispersion. The student comprehensively imitates both features and relationships from the teacher, yielding considerable effectiveness in dealing with complex remote sensing images. In addition, our InsDist can be easily built on mainstream object detectors with negligible extra cost. Extensive experiments on two large-scale remote sensing object detection datasets, namely DIOR and DOTA, show that our InsDist obtains noticeable gains over other distillation methods for both one-stage and two-stage, as well as both anchor-based and anchor-free detectors. The source code will be publicly available athttps://github.com/swift1988/InsDist. Gong Cheng 0003, Guangxing Wang 0001, Peicheng Zhou, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Robust Few-Shot Aerial Image Object Detection via Unbiased Proposals FiltrationabstractFew-shot aerial image object detection aims to rapidly detect object instances of novel category in aerial images by using few labeled samples. However, due to the complex background of aerial images, few labeled samples of novel categories, and the model trained with the few-shot learning paradigm is biased towards the base categories, it greatly increases the difficulty of identifying foreground objects of novel categories. In addition to this, tiny object detection is always a hot potato in aerial image object detection, and it is even more difficult for few-shot object detection. To this end, we propose a Few-shot aerial image object detection with Confidence-Iou collaborative proposal filtration and Tiny object constraint loss (FsCIT). Specifically, we first introduce a new confidence-iou collaborative proposal filtration scheme to RPN, which combines the unbiased IoU scores between the two bounding boxes with foreground-background confidence scores to filter redundant region proposals and rescue more foreground proposals for the novel categories in RPN. Then, we design a new tiny object loss constraint term to attempt at overcoming the challenge of tiny object detection in few-shot aerial image object detection. This term considers the central point distance, the size of ground-truth bounding boxes, and the distances between the four edges of the ground-truth bounding box and the predicted bounding box. Experiments on DIOR, AI-TOD and HRRSD datasets show that FsCIT is effective and can improve the performance of few-shot aerial image object detection. Lingjun Li, Xiwen Yao, Dongpao Hong, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Mining High-Quality Pseudoinstance Soft Labels for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised object detection in remote sensing image (RSI) is still a challenge because of the lack of instance-level labels, and many existing methods have two problems. Firstly, most of the existing methods usually mine the pseudo ground truth (PGT) instances solely relying on proposal class scores (PCS). Actually, the reliability of PCS is not enough because of the bird’s eye view imaging and large-scale chaotic background of RSIs, and the instances with high PCS incline to cover the discriminative region rather than the whole object. Secondly, the existing methods assign a one-hot label to each instance, and the label of PGT instance is copied to its neighbor instances, which induces the misclassification problem to some extent. Actually, the probability that the neighbor instances contain the object with the same category is smaller than the PGT instance. For the first problem, the proposal quality score (PQS) is proposed for mining high-quality PGT instances, which contains PCS and dual-context projection score (DCPS). The DCPS is calculated through semantic segmentation, and is employed to measure the completeness that each proposal covers an object. For the second problem, a pseudo soft label assignment (PSLA) strategy is proposed to assign more precise soft label for each instance, where the soft label is determined by the spatial distance between each instance and its nearest PGT instance. The ablation study validates the effectiveness of the PQS and PSLA. The comprehensive comparisons with other WSOD methods on three popular benchmarks show the excellent performance of our method. Xiaoliang Qian, Yu Huo 0001, Gong Cheng 0003, Chenyang Gao, Xiwen Yao, Wei Wang 0245 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Building a Bridge of Bounding Box Regression Between Oriented and Horizontal Object Detection in Remote Sensing ImagesabstractOriented object detection (OOD) aims to precisely detect the objects with arbitrary orientation in remote sensing images. Up to now, most of bounding box regression (BBR) losses for OOD are transferred from horizontal object detection (HOD) methods, however, the transferring requires lots of professional knowledge and experiences for designers, consequently, many excellent BBR losses for HOD have not been transferred to OOD. To accelerate the research progress of BBR loss for OOD, a unified transferring strategy (UTS) is proposed to facilitate the transferring of BBR loss from HOD to OOD. The UTS proposes that the BBR of oriented bounding box (OBB) can be converted into the joint BBR of its horizontal smallest enclosing rectangle (HSER) and two offsets, so the BBR loss in HOD can be easily transferred to OOD by using HSER as a bridge. Following the UTS, a BBR loss named Rotated-IoU (RIoU) loss is designed for OOD by transferring an advanced BBR loss in HOD, which can be considered as an example to show how to transfer. On the basis of RIoU loss, a focal rotated-IoU (FRIoU) loss is proposed to assign larger weights to hard samples in the BBR. The comparisons with other BBR losses show that the RIoU and FRIoU losses can give better performance. The ablation study shows that giving more attention to hard samples in BBR is effective. The comparisons with many advanced methods demonstrate that the combinations of baseline methods and FRIoU loss achieve state-of-the-art performance on the DOTA and DIOR-R datasets. Xiaoliang Qian, Baokun Wu, Gong Cheng 0003, Xiwen Yao, Wei Wang 0245, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Learning Orientation-Aware Distances for Oriented Object DetectionabstractOriented object detectors have suffered severely from the discontinuous boundary problem for a long time. In this work, we ingeniously avoid this problem by relating regression outputs to regression target orientations. The core idea of our method is to build a contour function which imports orientations and outputs the corresponding distance predictions. Inspired by Fourier transformations, we assume this function can be represented as a linear combination of trigonometric functions and Fourier series. We replace the final 4D layer in the regression branch of fully convolutional one-stage object detector (FCOS) with a Fourier Series Transformation (FST) module and term this new network FCOSF. By this unique design, the regression outputs in FCOSF can adaptively vary according to the regression target orientations. Thus, the discontinuous boundary has no impact on our FCOSF. More importantly, FCOSF avoids building complicated oriented box representations, which usually cause extra computations and ambiguities. With only flipping augmentation and single-scale training and testing, FCOSF with ResNet-50 achieves 73.64% mAP on the DOTA-v1.0 dataset with up to 23.6 FPS speed, surpassing all one-stage oriented object detectors. On the more challenging DOTA-v2.0 dataset, FCOSF also achieves the highest results of 51.75% mAP among one-stage detectors. More experiments on DIOR-R and HRSC2016 are also conducted to verify the robustness of FCOSF. Code and models will be available at https://github.com/DDGRCF/FCOSF. Chaofan Rao, Jiabao Wang 0005, Gong Cheng 0003, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Threatening Patch Attacks on Object Detection in Optical Remote Sensing ImagesabstractAdvanced Patch Attacks (PAs) on object detection in natural images have pointed out the great safety vulnerability in methods based on deep neural networks. However, little attention has been paid to this topic in Optical Remote Sensing Images (O-RSIs). To this end, we focus on this research,i.e., PAs on object detection in O-RSIs, and propose a more Threatening PA without the scarification of the visual quality, dubbed TPA. Specifically, to address the problem of inconsistency between local and global landscapes in existing patch selection schemes, we propose leveraging the First-Order Difference (FOD) of the objective function before and after masking to select the sub-patches to be attacked. Further, considering the problem of gradient inundation when applying existing coordinate-based loss to PAs directly, we design an IoU-based objective function specific for PAs, dubbed Bounding box Drifting Loss (BDL), which pushes the detected bounding boxes far from the initial ones until there are no intersections between them. Finally, on two widely used benchmarks,i.e., DIOR and DOTA, comprehensive evaluations of our TPA with four typical detectors (Faster R-CNN, FCOS, RetinaNet, and YOLO-v4) witness its remarkable effectiveness. To the best of our knowledge, this is the first attempt to study the PAs on object detection in O-RSIs, and we hope this work can get our readers interested in studying this topic. Xuxiang Sun 0001, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | TransY-Net: Learning Fully Transformer Networks for Change Detection of Remote Sensing ImagesabstractIn the remote sensing field, Change Detection (CD) aims to identify and localize the changed regions from dual-phase images over the same places. Recently, it has achieved great progress with the advances of deep learning. However, current methods generally deliver incomplete CD regions and irregular CD boundaries due to the limited representation ability of the extracted visual features. To relieve these issues, in this work we propose a novel Transformer-based learning framework named TransY-Net for remote sensing image CD, which improves the feature extraction from a global view and combines multi-level visual features in a pyramid manner. More specifically, the proposed framework first utilizes the advantages of Transformers in long-range dependency modeling. It can help to learn more discriminative global-level features and obtain complete CD regions. Then, we introduce a novel pyramid structure to aggregate multi-level visual features from Transformers for feature enhancement. The pyramid structure grafted with a Progressive Attention Module (PAM) can improve the feature representation ability with additional inter-dependencies through spatial and channel attentions. Finally, to better train the whole framework, we utilize the deeply-supervised learning with multiple boundary-aware loss functions. Extensive experiments demonstrate that our proposed method achieves a new state-of-the-art performance on four optical and two SAR image CD benchmarks. The source code is released at https://github.com/Drchip61/TransYNet. Tianyu Yan, Zifu Wan, Gong Cheng 0003, Huchuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | On Improving Bounding Box Representations for Oriented Object DetectionabstractDetecting objects in remote sensing images (RSIs) using oriented bounding boxes (OBBs) is flourishing but challenging, wherein the design of OBB representations is the key to achieving accurate detection. In this article, we focus on two issues that hinder the performance of the two-stage oriented detectors: 1) the notorious boundary discontinuity problem, which would result in significant loss increases in boundary conditions, and 2) the inconsistency in regression schemes between the two stages. We propose a simple and effective bounding box representation by drawing inspiration from the polar coordinate system and integrate it into two detection stages to circumvent the two issues. The first stage specifically initializes four quadrant points as the starting points of the regression for producing high-quality oriented candidates without any postprocessing. In the second stage, the final localization results are refined using the proposed novel bounding box representation, which can fully release the capabilities of the oriented detectors. Such consistency brings a good trade-off between accuracy and speed. With only flipping augmentation and single-scale training and testing, our approach with ResNet-50-FPN harvests 76.25% mAP on the DOTA dataset with a speed of up to 16.5 frames/s, achieving the best accuracy and the fastest speed among the mainstream two-stage oriented detectors. Additional results on the DIOR-R and HRSC2016 datasets also demonstrate the effectiveness and robustness of our method. The source code is publicly available athttps://github.com/yanqingyao1994/QPDet. Gong Cheng 0003, Guangxing Wang 0001, Shengyang Li, Peicheng Zhou, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | On Single-Model Transferable Targeted Attacks: A Closer Look at Decision-Level OptimizationabstractKnown as a hard nut, the single-model transferable targeted attacks via decision-level optimization objectives have attracted much attention among scholars for a long time. On this topic, recent works devoted themselves to designing new optimization objectives. In contrast, we take a closer look at the intrinsic problems in three commonly adopted optimization objectives, and propose two simple yet effective methods in this paper to mitigate these intrinsic problems. Specifically, inspired by the basic idea of adversarial learning, we, for the first time, propose a unified Adversarial Optimization Scheme (AOS) to release both the problems of gradient vanishing in cross-entropy loss and gradient amplification in Po+Trip loss, and indicate that our AOS, a simple transformation on the output logits before passing them to the objective functions, can yield considerable improvements on the targeted transferability. Besides, we make a further clarification on the preliminary conjecture in Vanilla Logit Loss (VLL) and point out the problem of unbalanced optimization in VLL, in which the source logit may risk getting increased without the explicit suppression on it, leading to the low transferability. Then, the Balanced Logit Loss (BLL) is further proposed, where we take both the source logit and the target logit into account. Comprehensive validations witness the compatibility and the effectiveness of the proposed methods across most attack frameworks, and their effectiveness can also span two tough cases (i.e., the low-ranked transfer scenario and the transfer to defense methods) and three datasets (i.e., the ImageNet, CIFAR-10, and CIFAR-100). Our source code is available at https://github.com/xuxiangsun/DLLTTAA. Xuxiang Sun 0001, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | NCSiam: Reliable Matching via Neighborhood Consensus for Siamese-Based Object TrackingabstractAn essential need for accurate visual object tracking is to capture better correlations between the tracking target and the search region. However, the dominant Siamese-based trackers are limited to producing dense similarity maps at once via a cross-correlations operation, ignoring to remedy the contamination caused by erroneous or ambiguous matches. In this paper, we propose a novel tracker, termed neighborhood consensus constraint-based siamese tracker (NCSiam), which takes the idea of neighborhood consensus constraint to refine the produced correlation maps. The intuition behind our approach is that we can support the nearby erroneous or ambiguous matches by analyzing a larger context of the scene that contains a unique match. Specifically, we devise a 4D convolution-based multi-level similarity refinement (MLSR) strategy. Taking the primary similarity maps obtained from a cross-correlation as input, MLSR acquires reliable matches by analyzing neighborhood consensus patterns in 4D space, thus enhancing the discriminability between the tracking target and the distractors. Besides, traditional Siamese-based trackers directly perform classification and regression on similarity response maps which discard appearance or semantic information. Therefore, an appearance affinity decoder (AAD) is developed to take full advantage of the semantic information of the search region. To further improve performance, we design a task-specific disentanglement (TSD) module to decouple the learned representations into classification-specific and regression-specific embeddings. Extensive experiments are conducted on six challenging benchmarks, including GOT-10k, TrackingNet, LaSOT, UAV123, OTB2015, and VOT2020. The results demonstrate the effectiveness of our method. The code will be available at https://github.com/laybebe/NCSiam. Pujian Lai, Gong Cheng 0003, Meili Zhang, Jifeng Ning, Xiangtao Zheng, Junwei Han 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Retain and Recover: Delving Into Information Loss for Few-Shot SegmentationabstractBenefiting from advances in few-shot learning techniques, their application to dense prediction tasks (e.g., segmentation) has also made great strides in the past few years. However, most existing few-shot segmentation (FSS) approaches follow a similar pipeline to that of few-shot classification, where some core components are directly exploited regardless of various properties between tasks. We note that such an ill-conceived framework introduces unnecessary information loss, which is clearly unacceptable given the already very limited training sample. To this end, we delve into the typical types of information loss and provide a reasonably effective way, namely Retain And REcover (RARE). The main focus of this paper can be summarized as follows: (i) the loss of spatial information due to global pooling; (ii) the loss of boundary information due to mask interpolation; (iii) the degradation of representational power due to sample averaging. Accordingly, we propose a series of strategies to retain/recover the avoidable/unavoidable information, such as unidirectional pooling, error-prone region focusing, and adaptive integration. Extensive experiments on two popular benchmarks (i.e., PASCAL-5iand COCO-20i) demonstrate the effectiveness of our scheme, which is not restricted to a particular baseline approach. The ultimate goal of our work is to address different information loss problems within a unified framework, and it also exhibits superior performance compared to other methods with similar motivations. The source code will be made available at https://github.com/chunbolang/RARE. Chunbo Lang, Gong Cheng 0003, Binfei Tu, Chao Li 0028, Junwei Han 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Learning to Assess Image Quality Like an ObserverabstractHuman observers are the ultimate receivers and evaluators of the image visual information and have powerful perception ability of visual quality with short-term global perception and long-term regional observation. Thus, it is natural to design an image quality assessment (IQA) computational model to act like an observer for accurately predicting the human perception of image quality. Inspired by this, here, we propose a novel observer-like network (OLN) to perform IQA by jointly considering the global glimpsing information and local scanning information. Specifically, the OLN consists of a global distortion perception (GDP) module and a local distortion observation (LDO) module. The GDP module is designed to mimic the observer's global perception of image quality through performing classification of images' distortion categories and levels. Simultaneously, to simulate the human local observation behavior, the LDO module attempts to gather the long-term regional observation information of the distorted images by continuously tracing the human scanpath in the observer-like scanning manner. By leveraging the bilinear pooling layer to collaborate the short-term global perception with the long-term regional observation, our network precisely predicts the quality scores of distorted images, such as human observers. Comprehensive experiments on the public datasets powerfully demonstrate that the proposed OLN achieves state-of-the-art performance. Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Weakly Supervised Rotation-Invariant Aerial Object Detection NetworkabstractObject rotation is among longstanding, yet still unexplored, hard issues encountered in the task of weakly supervised object detection (WSOD) from aerial images. Existing predominant WSOD approaches built on regular CNNs which are not inherently designed to tackle object rotations without corresponding constraints, thereby leading to rotation-sensitive object detector. Meanwhile, current solutions have been prone to fall into the issue with unsTable detectors, as they ignore lower-scored instances and may regard them as backgrounds. To address these issues, in this paper, we construct a novel end-to-end weakly supervised Rotation-Invariant aerial object detection Network (RINet). It is implemented with a flexible multi-branch online detector refinement, to be naturally more rotation-perceptive against oriented objects. Specifically, RINet first performs label propagating from the predicted instances to their rotated ones in a progressive refinement manner. Meanwhile, we propose to couple the predicted in-stance labels among different rotation-perceptive branches for generating rotation-consistent supervision and mean-while pursuing all possible instances. With the rotation-consistent supervisions, RINet enforces and encourages consistent yet complementary feature learning for WSOD without additional annotations and hyper-parameters. On the challenging NWPU VHR-10.v2 and DIOR datasets, extensive experiments clearly demonstrate that we significantly boost existing WSOD methods to a new state-of-the-art performance. The code will be available at: https://github.com/XiaoxFeng/RINet. Xiaoxu Feng, Xiwen Yao, Gong Cheng 0003, Junwei Han 0001 |
CVPR | 3 |
| 2022 | Learning What Not to Segment: A New Perspective on Few-Shot SegmentationabstractRecently few-shot segmentation (FSS) has been extensively developed. Most previous works strive to achieve generalization through the meta-learning framework derived from classification tasks; however, the trained models are biased towards the seen classes instead of being ideally class-agnostic, thus hindering the recognition of new concepts. This paper proposes a fresh and straightforward insight to alleviate the problem. Specifically, we apply an additional branch (base learner) to the conventional FSS model (meta learner) to explicitly identify the targets of base classes, i.e., the regions that do not need to be segmented. Then, the coarse results output by these two learners in parallel are adaptively integrated to yield precise segmentation prediction. Considering the sensitivity of meta learner, we further introduce an adjustment factor to estimate the scene differences between the input image pairs for facilitating the model ensemble forecasting. The substantial performance gains on PASCAL-5iand COCO-20iverify the effectiveness, and surprisingly, our versatile scheme sets a new state-of-the-art even with two plain learners. Moreover, in light of the unique nature of the proposed approach, we also extend it to a more realistic but challenging setting, i.e., generalized FSS, where the pixels of both base and novel classes are required to be determined. The source code is available at github.com/chunbolang/BAM. Chunbo Lang, Gong Cheng 0003, Binfei Tu, Junwei Han 0001 |
CVPR | 2 |
| 2022 | Exploring Effective Data for Surrogate Training Towards Black-box AttackabstractWithout access to the training data where a black-box victim model is deployed, training a surrogate model for black-box adversarial attack is still a struggle. In terms of data, we mainly identify three key measures for effective surrogate training in this paper. First, we show that leveraging the loss introduced in this paper to enlarge the inter-class similarity makes more sense than enlarging the inter-class diversity like existing methods. Next, unlike the approaches that expand the intra-class diversity in an implicit model-agnostic fashion, we propose a loss function specific to the surrogate model for our generator to enhance the intra-class diversity. Finally, in accordance with the in-depth observations for the methods based on proxy data, we argue that leveraging the proxy data is still an effective way for surrogate training. To this end, we propose a triple-player framework by introducing a discriminator into the traditional data-free framework. In this way, our method can be competitive when there are few semantic overlaps between the scarce proxy data (with the size between 1 k and 5k) and the training data. We evaluate our method on a range of victim models and datasets. The extensive results witness the effectiveness of our method. Our source code is available at https://github.com/xuxiangsun/ST-Data. Xuxiang Sun 0001, Gong Cheng 0003, Junwei Han 0001 |
CVPR | 2 |
| 2022 | Dynamic Proposal Generation for Oriented Object Detection in Aerial ImagesabstractCurrent two-stage oriented object detectors for aerial images have achieved remarkable progress. However, they still suffer from some drawbacks. Firstly, most of them place redundant anchors or utilize complicated transformation to generate oriented proposals, which are inefficient. Secondly, the generation of proposals is static, which cannot adapt to the extremely nonuniform distribution of objects. To address these issues, we propose a Dynamic Proposal Generation Network (DPGN) which can generate high-quality oriented proposals directly and estimate the upper limit of proposals adaptively. To be specific, with Guided Anchor Regression (GAR), we obtain the coarse oriented anchors and utilize them to align the features. After this, we make further classification and regression to produce final oriented proposals. Meanwhile, we design Maximum Number Estimation (MNE) for predicting an approximate value to remain the proposals adaptively. Without tricks, our method can achieve competitive detection accuracy compared with other mainstream methods on DOTA dataset. Qingyang Li 0001, Gong Cheng 0003, Shicheng Miao |
IGARSS | 2 |
| 2022 | Precise Vertex Regression and Feature Decoupling for Oriented Object DetectionabstractOriented object detection is a key task in the field of remote sensing image interpretation. Although extensive efforts have been made over the past few years, accurate oriented object detection remains a big challenge due to the dense arrangement and diverse orientations of objects. In this paper, we propose an oriented object detector based on the Faster R-CNN, which mainly consists of a Precise Vertex Regression (PVR) module and a Feature Decoupling (FD) module. Specifically, the PVR module predicts the arbitrary quadrilaterals of oriented objects with the precise vertex regression manner, which discretizes the regression range of vertex into several bins and applies a classification network to predict which bin the vertex belongs to. The FD module decouples the RoI features for classification and regression tasks by lightweight affine transformation. Experimental results on DOTA and DIOR-R datasets validate the effectiveness of our proposed method. Code is available at https://github.com/ShichengMiao16/VRDet. Shicheng Miao, Gong Cheng 0003, Qingyang Li 0001 |
IGARSS | 2 |
| 2022 | Beyond the Prototype: Divide-and-conquer Proxies for Few-shot SegmentationabstractFew-shot segmentation, which aims to segment unseen-class objects given only a handful of densely labeled samples, has received widespread attention from the community. Existing approaches typically follow the prototype learning paradigm to perform meta-inference, which fails to fully exploit the underlying information from support image-mask pairs, resulting in various segmentation failures, e.g., incomplete objects, ambiguous boundaries, and distractor activation. To this end, we propose a simple yet versatile framework in the spirit of divide-and-conquer. Specifically, a novel self-reasoning scheme is first implemented on the annotated support image, and then the coarse segmentation mask is divided into multiple regions with different properties. Leveraging effective masked average pooling operations, a series of support-induced proxies are thus derived, each playing a specific role in conquering the above challenges. Moreover, we devise a unique parallel decoder structure that integrates proxies with similar attributes to boost the discrimination power. Our proposed approach, named divide-and-conquer proxies (DCP), allows for the development of appropriate and reliable information as a guide at the “episode” level, not just about the object cues themselves. Extensive experiments on PASCAL-5i and COCO-20i demonstrate the superiority of DCP over conventional prototype-based approaches (up to 5~10% on average), which also establishes a new state-of-the-art. Code is available at github.com/chunbolang/DCP. Chunbo Lang, Binfei Tu, Gong Cheng 0003, Junwei Han 0001 |
IJCAI | 3 |
| 2022 | Guiding Clean Features for Object Detection in Remote Sensing ImagesabstractRecently, object detection has gained significant progress in remote sensing images. Nevertheless, we conclude two defects in remote sensing image object detection. At first, most methods rely on feature pyramid, but the features of different levels would influence each other when we use top-down operation. Second, the traditional label assignment strategy cannot assign suitable labels, as it adopts the fixed intersection over union (IoU) threshold to divide positive samples and negative samples during training. According to the problems we pointed out, a simple yet effective framework is employed to eliminate these two limitations. It integrates two novel components: aware feature pyramid network (AFPN) and group assignment strategy (GAS). AFPN is to mitigate the adverse effects caused by the first problem. Specifically, it learns a vector for the higher level features in the feature pyramid to obtain clean features. As for the second limitation, we recommend a new label assignment strategy named GAS to address this problem. Samples will be grouped according to their overlaps with ground truth, and then, they are assigned to positive or negative labels in each group. Extensive experiments are conducted on the large-scale object detection dataset DIOR and DOTA. With the newly introduced two key components, our model significantly improves the detection accuracy. Without bells and whistles, our proposed method achieves 2.0% and 1.9% higher mean average precision (mAP) than Faster R-CNN with FPN when using ResNet50 and ResNet101 as the backbones, respectively. Finally, we obtain 73.3% mAP on the DIOR dataset without any tricks. Our code is available athttps://github.com/hm-better/dior_detect. Gong Cheng 0003, Hailong Hong, Xiwen Yao, Xiaoliang Qian, Lei Guo 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | P-CNN: Part-Based Convolutional Neural Networks for Fine-Grained Visual CategorizationabstractThis paper proposes an end-to-end fine-grained visual categorization system, termed Part-based Convolutional Neural Network (P-CNN), which consists of three modules. The first module is a Squeeze-and-Excitation (SE) block, which learns to recalibrate channel-wise feature responses by emphasizing informative channels and suppressing less useful ones. The second module is a Part Localization Network (PLN) used to locate distinctive object parts, through which a bank of convolutional filters are learned as discriminative part detectors. Thus, a group of informative parts can be discovered by convolving the feature maps with each part detector. The third module is a Part Classification Network (PCN) that has two streams. The first stream classifies each individual object part into image-level categories. The second stream concatenates part features and global feature into a joint feature for the final classification. In order to learn powerful part features and boost the joint feature capability, we propose a Duplex Focal Loss used for metric learning and part classification, which focuses on training hard examples. We further merge PLN and PCN into a unified network for an end-to-end training process via a simple training technique. Comprehensive experiments and comparisons with state-of-the-art methods on three benchmark datasets demonstrate the effectiveness of our proposed method. Junwei Han 0001, Xiwen Yao, Gong Cheng 0003, Xiaoxu Feng, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Adaptive Neighborhood Metric LearningabstractIn this paper, we reveal that metric learning would suffer from serious inseparable problem if without informative sample mining. Since the inseparable samples are often mixed with hard samples, current informative sample mining strategies used to deal with inseparable problem may bring up some side-effects, such as instability of objective function, etc. To alleviate this problem, we propose a novel distance metric learning algorithm, named adaptive neighborhood metric learning (ANML). In ANML, we design two thresholds to adaptively identify the inseparable similar and dissimilar samples in the training procedure, thus inseparable sample removing and metric parameter learning are implemented in the same procedure. Due to the non-continuity of the proposed ANML, we develop an ingenious function, named log-exp mean function to construct a continuous formulation to surrogate it, which can be efficiently solved by the gradient descent method. Similar to Triplet loss, ANML can be used to learn both the linear and deep embeddings. By analyzing the proposed method, we find it has some interesting properties. For example, when ANML is used to learn the linear embedding, current famous metric learning algorithms such as the large margin nearest neighbor (LMNN) and neighbourhood components analysis (NCA) are the special cases of the proposed ANML by setting the parameters different values. When it is used to learn deep features, the state-of-the-art deep metric learning algorithms such as Triplet loss, Lifted structure loss, and Multi-similarity loss become the special cases of ANML. Furthermore, the log-exp mean function proposed in our method gives a new perspective to review the deep metric learning methods such as Prox-NCA and N-pairs loss. At last, promising experimental results demonstrate the effectiveness of the proposed method. Kun Song 0001, Junwei Han 0001, Gong Cheng 0003, Jiwen Lu, Feiping Nie 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Weakly Supervised Object Localization and Detection: A SurveyabstractAs an emerging and challenging problem in the computer vision community, weakly supervised object localization and detection plays an important role for developing new generation computer vision systems and has received significant attention in the past decade. As methods have been proposed, a comprehensive survey of these topics is of great importance. In this work, we review (1) classic models, (2) approaches with feature representations from off-the-shelf deep networks, (3) approaches solely based on deep learning, and (4) publicly available datasets and standard evaluation metrics that are widely used in this field. We also discuss the key challenges in this field, development history of this field, advantages/disadvantages of the methods in each category, the relationships between methods in different categories, applications of the weakly supervised object localization and detection methods, and potential future directions to further promote the development of this research field. Dingwen Zhang, Junwei Han 0001, Gong Cheng 0003, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Query-efficient decision-based attack via sampling distribution reshaping
Xuxiang Sun 0001, Gong Cheng 0003, Junwei Han 0001 |
Pattern Recognit. | 2 |
| 2022 | Multi-criteria Selection of Rehearsal Samples for Continual Learning
Chen Zhuang, Shaoli Huang, Gong Cheng 0003, Jifeng Ning |
Pattern Recognit. | 3 |
| 2022 | SPNet: Siamese-Prototype Network for Few-Shot Remote Sensing Image Scene ClassificationabstractFew-shot image classification has attracted extensive attention, which aims to recognize unseen classes given only a few labeled samples. Due to the large intraclass variances and interclass similarity of remote sensing scenes, the task under such circumstance is much more challenging than general few-shot image classification. Most existing prototype-based few-shot algorithms usually calculate prototypes directly from support samples and ignore the validity of prototypes, which results in a decline in the accuracy of subsequent inferences based on prototypes. To tackle this problem, we propose a Siamese-prototype network (SPNet) with prototype self-calibration (SC) and intercalibration (IC). First, to acquire more accurate prototypes, we utilize the supervision information from support labels to calibrate the prototypes generated from support features. This process is called SC. Second, we propose to consider the confidence scores of the query samples as another type of prototypes, which are then used to predict the support samples in the same way. Thus, the information interaction between support and query samples is implicitly a further calibration for prototypes (so-called IC). Our model is optimized with three losses, of which two additional losses help the model to learn more representative prototypes and make more accurate predictions. With no additional parameters to be learned, our model is very lightweight and convenient to employ. The experiments on three public remote sensing image datasets demonstrate competitive performance compared with other advanced few-shot image classification approaches. The source code is available athttps://github.com/zoraup/SPNet. Gong Cheng 0003, Liming Cai, Chunbo Lang, Xiwen Yao, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Perturbation-Seeking Generative Adversarial Networks: A Defense Framework for Remote Sensing Image Scene ClassificationabstractThe methods for remote sensing image (RSI) scene classification based on deep convolutional neural networks (DCNNs) have achieved prominent success. However, confronted with adversarial examples obtained by adding imperceptible perturbations to clean images, the great vulnerability of DCNNs makes it worth exploring effective defense methods. To date, numerous countermeasures for adversarial examples have been proposed, but how to improve the defensive ability for unknown attacks still to be answered. To address this issue, in this article, we propose an effective defense framework specified for RSI scene classification, named perturbation-seeking generative adversarial networks (PSGANs). In brief, a new training framework is designed to train the classifier by introducing the examples generated during the image reconstruction process, in addition to clean examples and adversarial ones. These generated examples can be random kinds of unknown attacks during training and thus are utilized to eliminate the blind spots of a classifier. To assist the proposed training framework, a reconstruction method is developed. First, instead of modeling the distribution of clean examples, we model the distributions of the perturbations added in adversarial examples. Second, to make a tradeoff between the diversity of the reconstructed examples and the optimization of PSGAN, a scale factor named seeking radius is introduced to scale the generated perturbations before they are subtracted by the given adversarial examples. Comprehensive and extensive experimental results on three widely used benchmarks for RSI scene classification demonstrate the great effectiveness of PSGAN when faced with both known and unknown attacks. Our source code is available athttps://github.com/xuxiangsun/PSGAN. Gong Cheng 0003, Xuxiang Sun 0001, Ke Li 0005, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | ISNet: Towards Improving Separability for Remote Sensing Image Change DetectionabstractDeep learning has substantially pushed forward remote sensing image change detection through extracting discriminative hierarchical features. However, as the increasingly high resolution remote sensing images have abundant spatial details but limited spectral information, the use of conventional backbone networks would give rise to blurry boundaries between different semantics among hierarchical features. This explains why most false alarms in the final predictions distribute around change boundaries. To alleviate the problem, we pay attention to feature refinement and propose deep learning networks that deliver improved separability (ISNet). Our ISNet reaps the advantages from two strategies applied to refining bi-temporal feature hierarchies: (i) margin maximization that clarifies the gap between changed and unchanged semantics, and (ii) targeted arrangement of attention mechanisms that directs the use of channel attention and spatial attention for highlighting semantic and positional information, respectively. Specifically, we insert channel attention modules into share-weighted backbone networks to facilitate semantic-specific feature extraction. The semantic boundaries in the extracted bi-temporal hierarchical features are then clarified by margin maximization modules, followed by spatial attention modules to enhance positional change responses. A top-down fusion pathway makes the final refined features cover multi-scale representations and have strong separability for remote sensing image change detection. Extensive experimental evaluations demonstrate that our ISNet achieves state-of-the-art performance on the LEVIR-CD, SYSU-CD, and Season-Varying datasets, in terms of Overall Accuracy (OA), Intersection-of-Union (IoU), and F1 score. Code is available at https://github.com/xingronaldo/ISNet. Gong Cheng 0003, Guangxing Wang 0001, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Anchor-Free Oriented Proposal Generator for Object DetectionabstractOriented object detection is a practical and challenging task in remote sensing image interpretation. Nowadays, oriented detectors mostly use horizontal boxes as intermedium to derive oriented boxes from them. However, the horizontal boxes are inclined to get small Intersection-over-Unions (IoUs) with ground truths, which may have some undesirable effects, such as introducing redundant noise, mismatching with ground truths, detracting from the robustness of detectors, etc. In this paper, we propose a novel Anchor-free Oriented Proposal Generator (AOPG) that abandons horizontal box-related operations from the network architecture. AOPG first produces coarse oriented boxes by a Coarse Location Module (CLM) in an anchor-free manner and then refines them into high-quality oriented proposals. After AOPG, we apply a Fast R-CNN head to produce the final detection results. Furthermore, the shortage of large-scale datasets is also a hindrance to the development of oriented object detection. To alleviate the data insufficiency, we release a new dataset on the basis of our DIOR dataset and name it DIOR-R. Massive experiments demonstrate the effectiveness of AOPG. Particularly, without bells and whistles, we achieve the accuracy of 64.41%, 75.24% and 96.22% mAP on the DIOR-R, DOTA and HRSC2016 datasets respectively. Code and models are available at https://github.com/jbwang1997/AOPG. Gong Cheng 0003, Jiabao Wang 0005, Ke Li 0005, Xingxing Xie, Chunbo Lang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Self-Guided Proposal Generation for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD) in remote sensing images remains a challenging task when learning object detectors with only image-level labels. As we know, object proposal generation plays a crucial role in WSOD. At present, the proposal generation of most existing WSOD methods mainly relies on heuristic strategies such as selective search and Edge Boxes. However, the proposals obtained by the above methods cannot well cover the entire objects, severely hindering the performance of WSOD. To address this issue, this paper proposes a Self-guided Proposal Generation approach, termed SPG. It can be easily implemented with most WSOD methods in a unified framework. To this end, we first introduce a confidence propagation approach to obtain the objectness confidence map for each image, which on the one hand highlights informative object locations, and on the other hand aggregates discriminative feature representation by combining the objectness confidence map with the deep features. Then, the proposal generation is implemented by mining informative regions as proposals on the objectness confidence map. Extensive evaluations on two challenging datasets demonstrate that our SPG significantly improves the baseline methods Online Instance Classifier Refinement (OICR) and Min-Entropy Latent Model (MELM) by large margins (for OICR: 15.86% mAP and 12.89% CorLoc gains on the NWPU VHR-10.v2 dataset, 3.65% mAP and 4.87% CorLoc gains on the DIOR dataset; for MELM: 20.51% mAP and 23.54% CorLoc gains on the NWPU VHR-10.v2 dataset, 7.11% mAP and 4.96% CorLoc gains on the DIOR dataset) and achieves the state-of-the-art results compared with existing methods. Gong Cheng 0003, Weining Chen, Xiaoxu Feng, Xiwen Yao, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Dual-Aligned Oriented DetectorabstractIn the past few years, object detection in remote sensing images has achieved remarkable progress. However, the detection of oriented and densely packed objects are still unsatisfactory due to the following spatial and feature misalignments. 1) Most two-stage oriented detectors only introduce an orientation regression branch in the detection head, while still leverage horizontal proposals for classification and regression. This inevitably results in the spatial misalignment problem between horizontal proposals and oriented objects. 2) The features used for classification are in fact extracted from the region proposals which have shifted to the final predictions via the regression branch. This leads to the feature misalignment problem between the classification and the localization tasks. In this article, we present a two-stage oriented object detection method, termed dual-aligned oriented detector (DODet), toward evading the aforementioned problems of spatial and feature misalignments. In DODet, the first stage is an oriented proposal network (OPN), which generates high-quality oriented proposals via a novel representation scheme of oriented objects. The second stage is a localization-guided detection head (LDH) that aims at alleviating the feature misalignment between classification and localization. Comprehensive and extensive evaluations on three benchmarks, including DIOR-R, DOTA, and HRSC2016, indicate that our method could obtain consistent and substantial gains compared with the baseline method. The source code is publicly available athttps://github.com/yanqingyao1994/DODet. Gong Cheng 0003, Shengyang Li, Ke Li 0005, Xingxing Xie, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Prototype-CNN for Few-Shot Object Detection in Remote Sensing ImagesabstractRecently, due to the excellent representation ability of convolutional neural networks (CNNs), object detection in remote sensing images has undergone remarkable development. However, when trained with a small number of samples, the performance of the object detectors drops sharply. In this article, we focus on the following three main challenges of few-shot object detection in remote sensing images: 1) since the sample number of novel classes is far less than base classes, object detectors would fail to quickly adapt to the features of novel classes, which would result in overfitting; 2) the scarcity of samples in novel classes leads to a sparse orientation space, while the objects in remote sensing images usually have arbitrary orientations; and 3) the distribution of object instances in remote sensing images is scattered and, therefore, it is hard to identify foreground objects from the complex background. To tackle these problems, we propose a simple yet effective method named prototype-CNN (P-CNN), which mainly consists of three parts: a prototype learning network (PLN) converting support images to class-aware prototypes, a prototype-guided region proposal network (P-G RPN) for better generation of region proposals, and a detector head extending the head of Faster region-based CNN (R-CNN) to further boost the performance. Comprehensive evaluations on the large-scale DIOR dataset demonstrate the effectiveness of our P-CNN. The source code is available athttps://github.com/Ybowei/P-CNN. Gong Cheng 0003, Bowei Yan, Peizhen Shi, Ke Li 0005, Xiwen Yao, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | SAENet: Self-Supervised Adversarial and Equivariant Network for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised object detection (WSOD) in remote sensing images (RSIs) remains a challenge when learning a subtle object detection model with only image-level annotations. Most works tend to optimize the detection model via exploiting the most contributed region, thereby to be dominated by the most discriminative part of an object. Meanwhile, these methods ignore the consistency across different spatial transformations of the same image and always label them with different classes, which introduces potential ambiguities. To tackle these challenges, we propose a unique self-supervised adversarial and equivariant network (SAENet) and aim at learning complementary and consistent visual patterns for WSOD in RSIs. To this end, an adversarial dropout–activation block is first designed to facilitate the entire object detector via adaptively hiding the discriminative parts and highlighting the instance-related regions. Besides, we further introduce a flexible self-supervised transformation equivariance mechanism on each potential instance from multiple spatial transformations to obtain spatially consistent self-supervisions. Accordingly, the obtained supervisions can be leveraged to pursue a more robust and spatially consistent object detector. Comprehensive experiments on the challenging LEarning, VIsion and Remote sensing Laboratory (LEVIR), NorthWestern Polytechnical University (NWPU) VHR-10.v2, and detection in optical RSIs (DIOR) datasets validate that SAENet outperforms the previous state-of-the-art works and achieves 46.2%, 60.7%, and 27.1% mAP, respectively. Xiaoxu Feng, Xiwen Yao, Gong Cheng 0003, Jungong Han, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multiscale Generative Adversarial Network Based on Wavelet Feature Learning for SAR-to-Optical Image TranslationabstractThe synthetic aperture radar (SAR) system is a kind of active remote sensing, which can be carried on a variety of flight platforms and can observe the Earth under all-day and all-weather conditions, so it has a wide range of applications. However, the interpretation of SAR images is quite challenging and not suitable for nonexperts. In order to enhance the visual effect of SAR images, this article proposes a multiscale generative adversarial network based on wavelet feature learning (WFLM-GAN) to implement the translation from SAR images to optical images; the translated images not only retain the key content of SAR images but also have the style of optical images. The main advantages of this method over the previous SAR-to-optical image translation (S2OIT) methods are given as follows. First, the generator does not learn the mapping from SAR images to optical images directly but learns the mapping from SAR images to wavelet features and then reconstructs the gray-scale images to optimize the content, increasing the mapping relationships and helping to learn more effective features. Second, a multiscale coloring network based on detail learning and style learning is designed to further translate the gray-scale images into optical images, which makes the generated images have an excellent visual effect with details closer to real images. Extensive experiments on SAR image datasets in different regions and seasons demonstrate the superior performance of WFLM-GAN over the baseline algorithms in terms of structural similarity (SSIM), the peak signal-to-noise ratio (PSNR), the Frechet inception distance (FID), and the kernel inception distance (KID). Comprehensive ablation studies are also carried out to isolate the validity of each proposed component. Our codes will be available athttps://github.com/G2022G/WFLM-GAN. Cang Gu, Dongqing Wu, Gong Cheng 0003, Lei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | AIFS-DATASET for Few-Shot Aerial Image Scene ClassificationabstractFew-shot learning (FSL), which aims to rapidly recognize unseen categories with limited samples, has attracted wide attention in aerial image scene classification. However, the existing methods generally train and evaluate the model within a dataset, and changing the dataset requires retraining and evaluation, which only realizes the generalization of intra-dataset. Considering meta-learning, this brings in a natural assumption: FSL should learn meta-knowledge from cross-domain heterogeneous tasks and then can generalize to new data distributions (e.g., datasets) with few samples. To this end, we propose a new benchmark, dubbed aerial image few-shot dataset (AIFS-DATASET), which is composed of diverse datasets and can provide more realistic heterogeneous task distributions. On AIFS-DATASET, we use many heterogeneous tasks, across multi-domains without any aerial image category, to train the model, achieving “see more.” Then we transfer the learned knowledge to new tasks in aerial images to evaluate the generalization performance of the model, thus acquiring a “well-informed” few-shot aerial image scene classification model. Moreover, the challenges of inter-class similarity and intra-class discrepancy in aerial images still exist. We also develop a dual constrained distance metric learning (DC-DML) framework to deal with the variable learning tasks adaptively and to achieve compact data distribution within a class and clear distribution gaps between classes from the perspective of metric learning. DC-DML mainly uses a task-adapted feature extractor while devising a novel distance metric with a cross-class bias penalty. By conducting experiments on AIFS-DATASET, we observed that DC-DML outperforms the current prevailing FSL approaches by a large margin. Lingjun Li, Xiwen Yao, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Solo-to-Collaborative Dual-Attention Network for One-Shot Object Detection in Remote Sensing ImagesabstractIn this article, we attempt to achieve one-shot object detection by mimicking the human ability to learn new concepts under limited reference, which aims at detecting all object instances of an unseen class in a target image when given a query image of the same unseen class. However, this one-shot learning ability of human benefits from the fact that human brain can quickly extract and process the associated information between the query–target images, which is an issue for the one-shot object detection framework to overcome. Moreover, the feature extraction of the query class in target images is intractable due to the complex and diversified background of remote sensing images. To solve these issues, we propose a solo-to-collaborative dual-attention network (SCoDANet) to hierarchically (image itself/pairs) enhance image feature representations. It consists of three components: 1) solo-attention head that strengthens the compactness of intraclass feature representations of an image and avoids background interference by selectively aggregating the similar features from the spatial and channel dimensions, respectively; 2) dual coattention module that guides RPN to generate an expected set of region proposals related to the query class by mining the coinformation of each query–target feature pair; and 3) nonlinear matching that provides a measure of similarity between the query feature and proposals of the target image to further learn a more robust detector. Our extensive experiments over two benchmarks demonstrate the effectiveness of our method under the one-shot scenario of detecting seen and unseen object categories. Lingjun Li, Xiwen Yao, Gong Cheng 0003, Mingliang Xu 0001, Jungong Han, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Scale-Aware Detailed Matching for Few-Shot Aerial Image Semantic SegmentationabstractFew-shot semantic segmentation, aiming to segment query images with a few annotated support samples, has drawn increasing attention. Most existing few-shot methods leverage the single prototype obtained from global average pooling to represent all support information and further use the extracted prototype to segment the query images in a matching manner. Although promising results for natural images have been reported, these methods cannot be directly applied on aerial images. The main reason comes from that the extracted single support prototype can only provide a coarse guidance for matching between query and support images and could not handle the large variance of objects’ appearances and scales. To deal with these challenges on aerial images, we propose a scale-aware few-shot semantic segmentation network to perform detailed matching with multiple prototypes. More specifically, the detailed matching module is first constructed to compute the pixel-level similarity between the query features and the extracted multiple support prototypes for providing more accurate parsing guidance. Subsequently, to address the problem of scale imbalance, the scale-aware focal loss is designed to dynamically down-weight the loss assigned to large well-parsed objects and focus training on tiny hard-parsed objects. To facilitate the reproducible research on the task of few-shot semantic segmentation in aerial images, we further provide a few-shot segmentation benchmark iSAID-$5^{\mathrm {i}}$constructed from the large-scale iSAID dataset[1]. Comprehensive experiments and comparisons with the state-of-the-art few-shot segmentation methods on the iSAID-$5^{\mathrm {i}}$dataset clearly demonstrate the superiority of our proposed method. The code and dataset are available athttps://github.com/caoql98/SDM. Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | R²IPoints: Pursuing Rotation-Insensitive Point Representation for Aerial Object DetectionabstractAnchor-free aerial object detection methods have recently attracted much attention due to their simplicity and efficiency. However, the performance is still unsatisfactory due to the following two main limitations. On the one hand, the anchor-free detector employs ordinary convolution layers with axis-aligned receptive fields to extract object features, resulting in lacking internal mechanisms to handle the rotation variance. On the other hand, the detector sacrifices much semantic information to achieve faster detection, leading to the inability to deal with objects’ high inter-class similarity and intra-class diversity. To address these issues, in this paper, we present a unique anchor-free detector, termed Rotation-Insensitive Point Representation (R2IPoints), of which a set of category-aware points are employed to encode the spatial and semantic information of the arbitrary-oriented objects. Specifically, we first devise a Stacked Rotation convolution Module (SRM) to encourage the learning of rotation-insensitive point representation by adaptively modelling orientation-agnostic interdependencies over stochastically rotated features. Meanwhile, we further introduce a Class-specific Semantic enhancement Module (CSM). It performs category-aware semantic activation to recalibrate features, thus enabling the point representation to be aware of object categories. Through jointly optimizing the two proposed modules in an end-to-end manner, R2IPoints could simultaneously generate rotation-insensitive and category-aware point representation. Extensive experiments on the challenging DIOR and DOTA datasets demonstrate the superiority of the proposed method. We achieve 72.7% mAP on DIOR and 74.34% mAP on DOTA, surpassing the baseline method of +2.4% mAP and +2.49% mAP, respectively. The code is available at https://github.com/shnew/R2IPoints. Xiwen Yao, Hui Shen 0005, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | DFENet for Domain Adaptation-Based Remote Sensing Scene ClassificationabstractDomain adaptation scene classification refers to the task of scene classification where the training set (called source domain) has different distributions from the test set (called target domain). Although remarkable results have been reported, the misalignment of source and target domain features still remains a big challenge when the large intraclass variances of remote sensing images encounter the insufficient exploration of discriminative feature representations for both domains. To address this challenge, a novel domain feature enhancement network (DFENet) is proposed to adaptively enhance the discriminative ability of the learned features for dealing with the domain variances of scene classification. Specifically, an adaptive context-aware feature refinement (CAFR) module is first designed to automatically recalibrate global and local features by explicitly modeling interdependencies between the channel and spatial for each domain. Then, a multilevel adversarial dropout (MAD) module is further designed to strengthen the generalization capability of our network by adaptively reconfiguring the sparsity of the feature level and decision level in the target domain. The cooperation of CAFR module and MAD module formulates a unique DFENet that can be learned in an end-to-end manner. Comprehensive experiments show that our proposed method is better than state-of-the-art methods on Merced$\to $RSSCN7, AID$\to $RSSCN7, NWPU$\to $RSSCN7, RSSCN7$\to $Merced, RSSCN7$\to $AID, and RSSCN7$\to $NWPU datasets. Xiufei Zhang, Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Oriented R-CNN for Object DetectionabstractCurrent state-of-the-art two-stage detectors generate oriented proposals through time-consuming schemes. This diminishes the detectors’ speed, thereby becoming the computational bottleneck in advanced oriented object detection systems. This work proposes an effective and simple oriented object detection framework, termed Oriented R-CNN, which is a general two-stage oriented detector with promising accuracy and efficiency. To be specific, in the first stage, we propose an oriented Region Proposal Network (oriented RPN) that directly generates high-quality oriented proposals in a nearly cost-free manner. The second stage is oriented R-CNN head for refining oriented Regions of Interest (oriented RoIs) and recognizing them. Without tricks, oriented R-CNN with ResNet50 achieves state-of-the-art detection accuracy on two commonly-used datasets for oriented object detection including DOTA (75.87% mAP) and HRSC2016 (96.50% mAP), while having a speed of 15.1 FPS with the image size of 1024×1024 on a single RTX 2080Ti. We hope our work could inspire rethinking the design of oriented detectors and serve as a baseline for oriented object detection. Code is available at https://github.com/jbwang1997/OBBDetection. Xingxing Xie, Gong Cheng 0003, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001 |
ICCV | 2 |
| 2021 | Multi-Scale Bidirectional Feature Fusion for One-Stage Oriented Object Detection in Aerial ImagesabstractThis paper aims to address the problem of oriented object detection under the complex background of remote sensing images. To this end, we propose a one-stage object detection method with feature fusion structure, and modify the loss function to enhance the detection of small objects. More specifically, on the basis of the end-to-end one-stage object detection model RetinaNet, the method of gliding the vertices of the horizontal bounding box is used to describe an oriented object. In order to obtain multi-scale context information, we design a feature fusion module. Besides, we propose a novel area-weighted loss function to pay more attention to small objects. Experimental results conducted on the DOTA dataset demonstrate that the proposed framework outperforms several state-of-the-art baselines. Gong Cheng 0003, Xuxiang Sun 0001, Qingyang Li 0001, Meili Zhang, Shicheng Miao |
IGARSS | 2 |
| 2021 | Object Detection in Optical Remote Sensing Images Based on Positive Sample Reweighting and Feature DecouplingabstractThe object detection head of Faster R-CNN shared by classification and localization tasks can cause greater harm to training process due to the different features required by the tasks. Besides, Faster R-CNN assigns the same weight to positive proposals, which is unfair to the positive proposal with better regression performance. Aiming at these issues, this paper proposes a novel detection framework, which consists of two important parts, i.e., a Positive Sample Reweighting (PSR) module and a Feature Decoupling (FD) module. Specifically, PSR re-weights each positive proposal according to the positional relationship between each positive proposal and its ground truth. FD employs two parallel branches to extract the weights to decouple features for classification and localization tasks. Experimental results on the DIOR and DOTA datasets show the effectiveness of our proposed method. Wenqi Yu, Jiabao Wang 0005, Gong Cheng 0003 |
IGARSS | 3 |
| 2021 | Task-wise attention guided part complementary learning for few-shot image classification
Gong Cheng 0003, Chunbo Lang, Junwei Han 0001 |
Sci. China Inf. Sci. | 1 |
| 2021 | Cross-Scale Feature Fusion for Object Detection in Optical Remote Sensing ImagesabstractFor the time being, there are many groundbreaking object detection frameworks used in natural scene images. These algorithms have good detection performance on the data sets of open natural scenes. However, applying these frameworks to remote sensing images directly is not very effective. The existing deep-learning-based object detection algorithms still face some challenges when dealing with remote sensing images because these images usually contain a number of targets with large variations of object sizes as well as interclass similarity. Aiming at the challenges of object detection in optical remote sensing images, we propose an end-to-end cross-scale feature fusion (CSFF) framework, which can effectively improve the object detection accuracy. Specifically, we first use a feature pyramid network (FPN) to obtain multilevel feature maps and then insert a squeeze and excitation (SE) block into the top layer to model the relationship between different feature channels. Next, we use the CSFF module to obtain powerful and discriminative multilevel feature representations. Finally, we implement our work in the framework of Faster region-based CNN (R-CNN). In the experiment, we evaluate our method on a publicly available large-scale data set, named DIOR, and obtain an improvement of 3.0% measured in terms of mAP compared with Faster R-CNN with FPN. Gong Cheng 0003, Yongjie Si, Hailong Hong, Xiwen Yao, Lei Guo 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Two-Stream Encoder GAN With Progressive Training for Co-Saliency DetectionabstractThe recent end-to-end co-saliency models have good performance, however, they cannot express the semantic consistency among a group of images well and usually require many co-saliency labels. To this end, a two-stream encoder generative adversarial network (TSE-GAN) with progressive training is proposed in this paper. In the pre-training stage, the salient object detection generative adversarial networks (SOD-GAN) and classification network (CN) are separately trained by the salient object detection (SOD) datasets and co-saliency datasets with only category labels to learn the intra-saliency and preliminary inter-saliency cues and alleviate the problem of insufficient co-saliency labels. In the second training stage, the backbone of TSE-GAN is inherited from the trained SOD-GAN, the encoder of trained SOD-GAN (SOD-Encoder) is used to extract intra-saliency features, the group-wise semantic encoder (GS-Encoder) is constructed by the multi-level group-wise category features extracted from CN for extracting inter-saliency features with better semantic consistency, the TSE-GAN constructed by incorporating the GS-Encoder into SOD-GAN is trained on co-saliency datasets for co-saliency detection. The comprehensive comparisons with 13 state-of-the-art methods demonstrate the effectiveness of proposed method. Xiaoliang Qian, Gong Cheng 0003, Xiwen Yao, Liying Jiang |
IEEE Signal Process. Lett. | 3 |
| 2021 | TCANet: Triple Context-Aware Network for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised object detection (WSOD) in remote sensing images (RSI) plays an essential role in RSI understanding applications. Currently, predominant works are inclined to first activate the most discriminative region and then pursue the whole object by analyzing the context information of the activated region. However, the most discriminative region usually only covers a small crucial part. Besides, many same-class instances often appear in adjacent locations. In such a case, treating proposals of large spatial overlap as the same-class instances not only introduces potential ambiguities but also misleads the detection model to recognize multiple adjacent instances as one object instance. To address these challenges, a novel triple context-aware network (TCANet) is proposed to learn complementary and discriminative visual patterns for WSOD in RSIs. Specifically, a global context-aware enhancement (GCAE) module is first designed to activate the features of the whole object by capturing the global visual scene context. Then, a dual-local context residual (DLCR) module is further developed to capture the instance-level discriminative cues by leveraging the semantic discrepancy of the local context. Furthermore, an effective adaptive-weighted refinement loss is integrated into the DLCR module to reduce the ambiguities in the label propagating process. The collaboration of GCAE and DLCR formulates a unique TCANet that can be learned in an end-to-end manner. Comprehensive experiments are carried out on the challenging NWPU VHR-10.v2 and DIOR data sets. We achieve a 58.8% mAP and a 25.8% mAP on the NWPU VHR-10.v2 and DIOR data sets, respectively, which both significantly outperform the state of the arts. Xiaoxu Feng, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | DLA-MatchNet for Few-Shot Remote Sensing Image Scene ClassificationabstractFew-shot scene classification aims to recognize unseen scene concepts from few labeled samples. However, most existing works are generally inclined to learn metalearners or transfer knowledge while ignoring the importance to learn discriminative representations and a proper metric for remote sensing images. To address these challenges, in this article, we propose an end-to-end network for boosting a few-shot remote sensing image scene classification, called discriminative learning of adaptive match network (DLA-MatchNet). Specifically, we first adopt the attention technique to delve into the interchannel and interspatial relationships to automatically discover discriminative regions. Then, the channel attention and spatial attention modules can be incorporated with the feature network by using different feature fusion schemes, achieving “discriminative learning.” Afterward, considering the issues of the large intraclass variances and interclass similarity of remote sensing images, instead of simply computing the distances between the support samples and query samples, we concatenate the support and query discriminative features in depth and utilize a matcher to “adaptively” select the semantically relevant sample pairs to assign similarity scores. Our method leverages an episode-based strategy to train the model. Once trained, our model can predict the category of query image without further fine-tuning. Experimental results on three public remote sensing image data sets demonstrate the effectiveness of our model in the few-shot scene classification task. Lingjun Li, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003, Lei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Automatic Weakly Supervised Object Detection From High Spatial Resolution Remote Sensing Images via Dynamic Curriculum LearningabstractIn this article, we focus on tackling the problem of weakly supervised object detection from high spatial resolution remote sensing images, which aims to learn detectors with only image-level annotations, i.e., without object location information during the training stage. Although promising results have been achieved, most approaches often fail to provide high-quality initial samples and thus are difficult to obtain optimal object detectors. To address this challenge, a dynamic curriculum learning strategy is proposed to progressively learn the object detectors by feeding training images with increasing difficulty that matches current detection ability. To this end, an entropy-based criterion is firstly designed to evaluate the difficulty for localizing objects in images. Then, an initial curriculum that ranks training images in ascending order of difficulty is generated, in which easy images are selected to provide reliable instances for learning object detectors. With the gained stronger detection ability, the subsequent order in the curriculum for retraining detectors is accordingly adjusted by promoting difficult images as easy ones. In such way, the detectors can be well prepared by training on easy images for learning from more difficult ones and thus gradually improve their detection ability more effectively. Moreover, an effective instance-aware focal loss function for detector learning is developed to alleviate the influence of positive instances of bad quality and meanwhile enhance the discriminative information of class-specific hard negative instances. Comprehensive experiments and comparisons with state-of-the-art methods on two publicly available data sets demonstrate the superiority of our proposed method. Xiwen Yao, Xiaoxu Feng, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Progressive Contextual Instance Refinement for Weakly Supervised Object Detection in Remote Sensing ImagesabstractWeakly supervised learning has been attracting much attention due to its broad applications, which only requires image-level annotations to indicate whether there exist objects in the images. Currently, most of the existing weakly supervised object detection (WSOD) methods are inclined to seek only one top-scoring object instance per image from noisy proposals to train the corresponding object detector. However, more than one same-class instances often exist in the large-scale, cluttered remote sensing images. Thus, selecting only one top-scoring proposal usually results in highlighting the most representative part of an object rather than the whole object, which may cause learning a suboptimal object detector by losing much important information. To address this problem, a novel end-to-end progressive contextual instance refinement (PCIR) method is proposed to perform WSOD. Specifically, a dual-contextual instance refinement (DCIR) strategy is designed to divert the focus of the detection network from the local distinct part to the whole object and further to other potential instances by leveraging both local and global context information. Benefiting from DCIR, a progressive proposal self-pruning (PPSP) strategy is further developed to mitigate the influence of the complex background by dynamically rejecting the negative training proposals. Comprehensive experiments on the challenging NWPU VHR-10.v2 and DIOR data sets clearly demonstrate that the proposed method can significantly boost the detection accuracy compared with the state of the arts. Xiaoxu Feng, Junwei Han 0001, Xiwen Yao, Gong Cheng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | High-Quality Proposals for Weakly Supervised Object DetectionabstractDespite significant efforts made so far for Weakly Supervised Object Detection (WSOD), proposal generation and proposal selection are still two major challenges. In this paper, we focus on addressing the two challenges by generating and selecting high-quality proposals. To be specific, for proposal generation, we combine selective search and a Gradient-weighted Class Activation Mapping (Grad-CAM) based technique to generate more proposals having higher Intersection-Over-Union (IOU) with ground truth boxes than those obtained by greedy search approaches, which can better envelop the entire objects. As regards proposal selection, for each object class, we choose as many confident positive proposals as possible and meanwhile only select class-specific hard negatives to focus training on more discriminative negative proposals by up-weighting their losses, which can make training more effective. The proposed proposal generation and proposal selection approaches are generic and thus can be broadly applied to many WSOD methods. In this work, we unify them into the framework of Online Instance Classifier Refinement (OICR). Experimental results on the PASCAL VOC 2007 and 2012 datasets and MS COCO dataset demonstrate that our method significantly improves the baseline method OICR by large margins (13.4% mAP and 11.6% CorLoc gains on the VOC 2007 dataset, 15.0% mAP and 8.9% CorLoc gains on the VOC 2012 dataset, and 6.4% mAP and 5.0% CorLoc gains on the COCO dataset) and achieves the state-of-the-art results compared with existing methods. Gong Cheng 0003, Junyu Yang, Decheng Gao, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Performance Comparison of Two Pooling Strategies for Remote Sensing Image Scene ClassificationabstractWith the advances of convolutional neural networks (CNNs), the accuracy of remote sensing image scene classification has been greatly boosted thanks to the powerful features extracted through CNNs. Although significant success has been achieved, most of existing methods are dominated by the use of fully-connected CNN features. This paper focuses on the performance comparison of two kinds of novel pooling strategies, including generalized max pooling (GMP) and taskdriven pooling (TDP), for remote sensing image scene classification. To this end, an off-the-shelf CNN model is used as backbone network to extract multi-scale convolutional features. Then, GMP and TDP are respectively adopted to obtain globally pooled features. Finally, scene classification is performed with support vector machine (SVM). In the experiment, we evaluate the performance of these two kinds of pooling schemes on a widely-used scene classification benchmark data set. The experimental results show that (i) using pooled CNN convolutional features can obtain better results than using fully-connected CNN features and (ii) TDP is slightly better than GMP. Maoxiong Wu, Gong Cheng 0003, Xiwen Yao, Xiaoliang Qian, Junwei Han 0001, Lei Guo 0002 |
IGARSS | 2 |
| 2019 | Learning Region Response Ranking Features for Remote Sensing Image Scene ClassificationabstractRecently, deep learning especially convolutional neural networks (CNNs) has huge great success for remote sensing image scene classification. However, global CNN features still lack geometric invariance for addressing the problem of large intra-class variations and so are not optimal for scene classification. In this paper, we introduce a new feature representation for scene classification, named region response ranking (3R) feature representations by using off-the-shelf CNN models. Specifically, by considering each cube pixel of a certain convolutional feature map as one image region, we jointly train a class-specific support vector machine (SVM) base classifier and a decision function for each scene class. The base classifier is used to generate 3R feature by reordering the SVM responses of all image regions in descending order and the decision function is used for classification with 3R feature representations. Comprehensive evaluations on the publicly available NWPU-RESISC45 data set and comparisons with state-of-the-art methods demonstrate that the proposed 3R feature is effective for remote sensing image scene classification.1 Junyu Yang, Gong Cheng 0003, Xiwen Yao, Junwei Han 0001, Lei Guo 0002 |
IGARSS | 2 |
| 2019 | Rotation-Invariant Latent Semantic Representation Learning for Object Detection in VHR Optical Remote Sensing ImagesabstractObject detection in very high resolution (VHR) optical remote sensing images is a fundamental yet challenging problem for the field of remote sensing image analysis. The detection performance is heavily dependent on the representation capability of the extracted features. Recently, convolutional neural networks (CNNs) have made a breakthrough for various applications in nature images. However, it is problematic to directly apply CNN to perform object detection in VHR optical remote sensing images due to the problem of object rotation variations. To address this issue, a novel rotation invariant probabilistic Latent Semantic Analysis (RI-pLSA) model is proposed to learn latent semantic representations for object detection. This is achieved by imposing a rotation-invariant regularization term on the objective function of pLSA to enforce the learned representation from all rotations of the same sample to be as consistent as possible. Additionally, the proposed RI-pLSA model takes the CNN features as input, which generates more powerful semantic representation for object detection. Comprehensive experiments on a publicly available ten-class object detection dataset demonstrate the superiority and effectiveness of our method compared with state-of-the-arts. Xiwen Yao, Xiaoxu Feng, Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002 |
IGARSS | 3 |
| 2019 | Scene Classification of High Resolution Remote Sensing Images Via Self-Paced Deep LearningabstractScene classification of high resolution remote sensing (HRRS) images is a fundamental yet challenging problem for remote sensing image analysis. In this paper, we focus on tackling the problem of HRSS scene classification using a small pool of unlabeled images and only a few labeled images per category, namely, few-shot scene classification (FSSC), which is more challenging than common scene classification task. The key challenge arises from selecting trustworthy samples from the pool of unlabeled images that have high confidence. To address this challenge, a novel local manifold constrained self-paced deep learning method is proposed. Specifically, the model is learned by gradually selecting easy samples from the pool of unlabeled images, assigning them with pseudo-labels and further adopting them with labeled images as the new training set. In addition, a local manifold constraint is introduced to enforce that the pseudo-labels assigned by the initial model should be consistent with the local manifold of the labeled samples. In such way, the confidence of the selecting samples is increased and is beneficial to train more robust classifier. Experimental results on a publicly available large scale NWPU-RESISC45 data set demonstrated the effectiveness of our method in achieving competitive performance while significantly reducing manually labeled cost. Xiwen Yao, Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002 |
IGARSS | 3 |
| 2019 | Learning Compact and Discriminative Stacked Autoencoder for Hyperspectral Image ClassificationabstractAs one of the fundamental research topics in remote sensing image analysis, hyperspectral image (HSI) classification has been extensively studied so far. However, how to discriminatively learn a low-dimensional feature space, in which the mapped features have small within-class scatter and big between-class separation, is still a challenging problem. To address this issue, this paper proposes an effective framework, named compact and discriminative stacked autoencoder (CDSAE), for HSI classification. The proposed CDSAE framework comprises two stages with different optimization objectives, which can learn discriminative low-dimensional feature mappings and train an effective classifier progressively. First, we impose a local Fisher discriminant regularization on each hidden layer of stacked autoencoder (SAE) to train discriminative SAE (DSAE) by minimizing reconstruction error. This stage can learn feature mappings, in which the pixels from the same land-cover class are mapped as nearly as possible and the pixels from different land-cover categories are separated by a large margin. Second, we learn an effective classifier and meanwhile update DSAE with a local Fisher discriminant regularization being embedded on the top of feature representations. Moreover, to learn a compact DSAE with as small number of hidden neurons as possible, we impose a diversity regularization on the hidden neurons of DSAE to balance the feature dimensionality and the feature representation capability. The experimental results on three widely-used HSI data sets and comprehensive comparisons with existing methods demonstrate that our proposed method is effective. Peicheng Zhou, Junwei Han 0001, Gong Cheng 0003, Baochang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Learning Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object DetectionabstractThe performance of object detection has recently been significantly improved due to the powerful features learnt through convolutional neural networks (CNNs). Despite the remarkable success, there are still several major challenges in object detection, including object rotation, within-class diversity, and between-class similarity, which generally degenerate object detection performance. To address these issues, we build up the existing state-of-the-art object detection systems and propose a simple but effective method to train rotation-invariant and Fisher discriminative CNN models to further boost object detection performance. This is achieved by optimizing a new objective function that explicitly imposes a rotation-invariant regularizer and a Fisher discrimination regularizer on the CNN features. Specifically, the first regularizer enforces the CNN feature representations of the training samples before and after rotation to be mapped closely to each other in order to achieve rotation-invariance. The second regularizer constrains the CNN features to have small within-class scatter but large between-class separation. We implement our proposed method under four popular object detection frameworks, including region-CNN (R-CNN), Fast R- CNN, Faster R- CNN, and R- FCN. In the experiments, we comprehensively evaluate the proposed method on the PASCAL VOC 2007 and 2012 data sets and a publicly available aerial image data set. Our proposed methods outperform the existing baseline methods and achieve the state-of-the-art results. Gong Cheng 0003, Junwei Han 0001, Peicheng Zhou, Dong Xu 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Multi-scale and Discriminative Part Detectors Based Features for Multi-label Image ClassificationabstractConvolutional neural networks (CNNs) have shown their promise for image classification task. However, global CNN features still lack geometric invariance for addressing the problem of intra-class variations and so are not optimal for multi-label image classification. This paper proposes a new and effective framework built upon CNNs to learn Multi-scale and Discriminative Part Detectors (MsDPD)-based feature representations for multi-label image classification. Specifically, at each scale level, we (i) first present an entropy-rank based scheme to generate and select a set of discriminative part detectors (DPD), and then (ii) obtain a number of DPD-based convolutional feature maps with each feature map representing the occurrence probability of a particular part detector and learn DPD-based features by using a task-driven pooling scheme. The two steps are formulated into a unified framework by developing a new objective function, which jointly trains part detectors incrementally and integrates the learning of feature representations into the classification task. Finally, the multi-scale features are fused to produce the predictions. Experimental results on PASCAL VOC 2007 and VOC 2012 datasets demonstrate that the proposed method achieves better accuracy when compared with the existing state-of-the-art multi-label classification methods. Gong Cheng 0003, Decheng Gao, Yang Liu 0007, Junwei Han 0001 |
IJCAI | 1 |
| 2018 | Identifying affective levels on music video via completing the missing modality
Gong Cheng 0003, Lei Guo 0002 |
Multim. Tools Appl. | 2 |
| 2018 | A Unified Metric Learning-Based Framework for Co-Saliency DetectionabstractCo-saliency detection, which focuses on extracting commonly salient objects in a group of relevant images, has been attracting research interest because of its broad applications. In practice, the relevant images in a group may have a wide range of variations, and the salient objects may also have large appearance changes. Such wide variations usually bring about large intra-co-salient objects (intra-COs) diversity and high similarity between COs and background, which makes the co-saliency detection task more difficult. To address these problems, we make the earliest effort to introduce metric learning to co-saliency detection. Specifically, we propose a unified metric learning-based framework to jointly learn discriminative feature representation and co-salient object detector. This is achieved by optimizing a new objective function that explicitly embeds a metric learning regularization term into support vector machine (SVM) training. Here, the metric learning regularization term is used to learn a powerful feature representation that has small intra-COs scatter, but big separation between background and COs and the SVM classifier is used for subsequent co-saliency detection. In the experiments, we comprehensively evaluate the proposed method on two commonly used benchmark data sets. The state-of-the-art results are achieved in comparison with the existing co-saliency detection methods. Junwei Han 0001, Gong Cheng 0003, Dingwen Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Exploring Hierarchical Convolutional Features for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is an active and important research task driven by many practical applications. To leverage deep learning models especially convolutional neural networks (CNNs) for HSI classification, this paper proposes a simple yet effective method to extract hierarchical deep spatial feature for HSI classification by exploring the power of off-the-shelf CNN models, without any additional retraining or fine-tuning on the target data set. To obtain better classification accuracy, we further propose a unified metric learning-based framework to alternately learn discriminative spectral-spatial features, which have better representation capability and train support vector machine (SVM) classifiers. To this end, we design a new objective function that explicitly embeds a metric learning regularization term into SVM training. The metric learning regularization term is used to learn a powerful spectral-spatial feature representation by fusing spectral feature and deep spatial feature, which has small intraclass scatter but big between class separation. By transforming HSI data into new spectral-spatial feature space through CNN and metric learning, we can pull the pixels from the same class closer, while pushing the different class pixels farther away. In the experiments, we comprehensively evaluate the proposed method on three commonly used HSI benchmark data sets. State-of-the-art results are achieved when compared with the existing HSI classification methods. Gong Cheng 0003, Junwei Han 0001, Xiwen Yao, Lei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | When Deep Learning Meets Metric Learning: Remote Sensing Image Scene Classification via Learning Discriminative CNNsabstractRemote sensing image scene classification is an active and challenging task driven by many applications. More recently, with the advances of deep learning models especially convolutional neural networks (CNNs), the performance of remote sensing image scene classification has been significantly improved due to the powerful feature representations learnt through CNNs. Although great success has been obtained so far, the problems of within-class diversity and between-class similarity are still two big challenges. To address these problems, in this paper, we propose a simple but effective method to learn discriminative CNNs (D-CNNs) to boost the performance of remote sensing image scene classification. Different from the traditional CNN models that minimize only the cross entropy loss, our proposed D-CNN models are trained by optimizing a new discriminative objective function. To this end, apart from minimizing the classification error, we also explicitly impose a metric learning regularization term on the CNN features. The metric learning regularization enforces the D-CNN models to be more discriminative so that, in the new D-CNN feature spaces, the images from the same scene class are mapped closely to each other and the images of different classes are mapped as farther apart as possible. In the experiments, we comprehensively evaluate the proposed method on three publicly available benchmark data sets using three off-the-shelf CNN models. Experimental results demonstrate that our proposed D-CNN methods outperform the existing baseline methods and achieve state-of-the-art results on all three data sets. Gong Cheng 0003, Ceyuan Yang, Xiwen Yao, Lei Guo 0002, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Rotation-Insensitive and Context-Augmented Object Detection in Remote Sensing ImagesabstractMost of the existing deep-learning-based methods are difficult to effectively deal with the challenges faced for geospatial object detection such as rotation variations and appearance ambiguity. To address these problems, this paper proposes a novel deep-learning-based object detection framework including region proposal network (RPN) and local-contextual feature fusion network designed for remote sensing images. Specifically, the RPN includes additional multiangle anchors besides the conventional multiscale and multiaspect-ratio ones, and thus can deal with the multiangle and multiscale characteristics of geospatial objects. To address the appearance ambiguity problem, we propose a double-channel feature fusion network that can learn local and contextual properties along two independent pathways. The two kinds of features are later combined in the final layers of processing in order to form a powerful joint representation. Comprehensive evaluations on a publicly available ten-class object detection data set demonstrate the effectiveness of the proposed method. Ke Li 0005, Gong Cheng 0003, Shuhui Bu, Xiong You |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Duplex Metric Learning for Image Set ClassificationabstractImage set classification has attracted much attention because of its broad applications. Despite the success made so far, the problems of intra-class diversity and inter-class similarity still remain two major challenges. To explore a possible solution to these challenges, this paper proposes a novel approach, termed duplex metric learning (DML), for image set classification. The proposed DML consists of two progressive metric learning stages with different objectives used for feature learning and image classification, respectively. The metric learning regularization is not only used to learn powerful feature representations but also well explored to train an effective classifier. At the first stage, we first train a discriminative stacked autoencoder (DSAE) by layer-wisely imposing a metric learning regularization term on the neurons in the hidden layers and meanwhile minimizing the reconstruction error to obtain new feature mappings in which similar samples are mapped closely to each other and dissimilar samples are mapped farther apart. At the second stage, we discriminatively train a classifier and simultaneously fine-tune the DSAE by optimizing a new objective function, which consists of a classification error term and a metric learning regularization term. Finally, two simple voting strategies are devised for image set classification based on the learnt classifier. In the experiments, we extensively evaluate the proposed framework for the tasks of face recognition, object recognition, and face verification on several commonly-used data sets and state-of-the-art results are achieved in comparison with existing methods. Gong Cheng 0003, Peicheng Zhou, Junwei Han 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Remote Sensing Image Scene Classification Using Bag of Convolutional FeaturesabstractMore recently, remote sensing image classification has been moving from pixel-level interpretation to scene-level semantic understanding, which aims to label each scene image with a specific semantic class. While significant efforts have been made in developing various methods for remote sensing image scene classification, most of them rely on handcrafted features. In this letter, we propose a novel feature representation method for scene classification, named bag of convolutional features (BoCF). Different from the traditional bag of visual words-based methods in which the visual words are usually obtained by using handcrafted feature descriptors, the proposed BoCF generates visual words from deep convolutional features using off-the-shelf convolutional neural networks. Extensive evaluations on a publicly available remote sensing image scene classification benchmark and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed BoCF method for remote sensing image scene classification. Gong Cheng 0003, Xiwen Yao, Lei Guo 0002, Zhongliang Wei |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Semi-direct tracking and mapping with RGB-D camera for MAV
Shuhui Bu, Ke Li 0005, Gong Cheng 0003, Zhenbao Liu |
Multim. Tools Appl. | 5 |
| 2017 | Remote Sensing Image Scene Classification: Benchmark and State of the ArtabstractRemote sensing image scene classification plays an important role in a wide range of applications and hence has been receiving remarkable attention. During the past years, significant efforts have been made to develop various data sets or present a variety of approaches for scene classification from remote sensing images. However, a systematic review of the literature concerning data sets and methods for scene classification is still lacking. In addition, almost all existing data sets have a number of limitations, including the small scale of scene classes and the image numbers, the lack of image variations and diversity, and the saturation of accuracy. These limitations severely limit the development of new approaches especially deep learning-based methods. This paper first provides a comprehensive review of the recent progress. Then, we propose a large-scale data set, termed “NWPU-RESISC45,” which is a publicly available benchmark for REmote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This data set contains 31 500 images, covering 45 scene classes with 700 images in each class. The proposed NWPU-RESISC45 1) is large-scale on the scene classes and the total image number; 2) holds big variations in translation, spatial resolution, viewpoint, object pose, illumination, background, and occlusion; and 3) has high within-class diversity and between-class similarity. The creation of this data set will enable the community to develop and evaluate various data-driven algorithms. Finally, several representative methods are evaluated using the proposed data set, and the results are reported as a useful baseline for future research. Gong Cheng 0003, Junwei Han 0001, Xiaoqiang Lu |
Proc. IEEE | 1 |
| 2016 | RIFD-CNN: Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object DetectionabstractThanks to the powerful feature representations obtained through deep convolutional neural network (CNN), the performance of object detection has recently been substantially boosted. Despite the remarkable success, the problems of object rotation, within-class variability, and between-class similarity remain several major challenges. To address these problems, this paper proposes a novel and effective method to learn a rotation-invariant and Fisher discriminative CNN (RIFD-CNN) model. This is achieved by introducing and learning a rotation-invariant layer and a Fisher discriminative layer, respectively, on the basis of the existing high-capacity CNN architectures. Specifically, the rotation-invariant layer is trained by imposing an explicit regularization constraint on the objective function that enforces invariance on the CNN features before and after rotating. The Fisher discriminative layer is trained by imposing the Fisher discrimination criterion on the CNN features so that they have small within-class scatter but large between-class separation. In the experiments, we comprehensively evaluate the proposed method for object detection task on a public available aerial image dataset and the PASCAL VOC 2007 dataset. State-of-the-art results are achieved compared with the existing baseline methods. Gong Cheng 0003, Peicheng Zhou, Junwei Han 0001 |
CVPR | 1 |
| 2016 | Scene classification of high resolution remote sensing images using convolutional neural networksabstractScene classification of high resolution remote sensing images plays an important role for a wide range of applications. While significant efforts have been made in developing various methods for scene classification, most of them are based on handcrafted or shallow learning-based features. In this paper, we investigate the use of deep convolutional neural network (CNN) for scene classification. To this end, we first adopt two simple and effective strategies to extract CNN features: (1) using pre-trained CNN models as universal feature extractors, and (2) domain-specifically fine-tuning pre-trained CNN models on our scene classification dataset. Then, scene classification is carried out by using simple classifiers such as linear support vector machine (SVM). In our work, three off-the-shelf CNN models including AlexNet [1], VGGNet [2], and GoogleNet [3] are investigated. Comprehensive evaluations on a publicly available 21 classes land use dataset and comparisons with several state-of-the-art approaches demonstrate that deep CNN features are effective for scene classification of high resolution remote sensing images. Gong Cheng 0003, Chengcheng Ma, Peicheng Zhou, Xiwen Yao, Junwei Han 0001 |
IGARSS | 1 |
| 2016 | Semantic annotation of satellite images via joint multi-feature learning with diversity constraintabstractAutomatic semantic annotation of high-resolution optical satellite images is a task to assign one or several predefined semantic concepts to an image according to its content. The fundamental challenge arises from the difficulty of characterizing complex and ambiguous contents of the satellite images. To address this challenge, a diversity constrained joint multi-feature learning method is proposed to learn robust feature representations for annotating satellite images. The key motivation of our method is to make full use of the complementarity diversity information among the heterogeneous features in the learning process. Comprehensive experiments on an annotation dataset demonstrate the superiority and effectiveness of our method compared with baseline multi-feature learning method. Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Peicheng Zhou, Lei Guo 0002 |
IGARSS | 3 |
| 2016 | Approximative Bayes optimality linear discriminant analysis for Chinese handwriting character recognition
Gong Cheng 0003 |
Neurocomputing | 2 |
| 2016 | Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing ImagesabstractObject detection in very high resolution optical remote sensing images is a fundamental problem faced for remote sensing image analysis. Due to the advances of powerful feature representations, machine-learning-based object detection is receiving increasing attention. Although numerous feature representations exist, most of them are handcrafted or shallow-learning-based features. As the object detection task becomes more challenging, their description capability becomes limited or even impoverished. More recently, deep learning algorithms, especially convolutional neural networks (CNNs), have shown their much stronger feature representation power in computer vision. Despite the progress made in nature scene images, it is problematic to directly use the CNN feature for object detection in optical remote sensing images because it is difficult to effectively deal with the problem of object rotation variations. To address this problem, this paper proposes a novel and effective approach to learn a rotation-invariant CNN (RICNN) model for advancing the performance of object detection, which is achieved by introducing and learning a new rotation-invariant layer on the basis of the existing CNN architectures. However, different from the training of traditional CNN models that only optimizes the multinomial logistic regression objective, our RICNN model is trained by optimizing a new objective function via imposing a regularization constraint, which explicitly enforces the feature representations of the training samples before and after rotating to be mapped close to each other, hence achieving rotation invariance. To facilitate training, we first train the rotation-invariant layer and then domain-specifically fine-tune the whole RICNN network to further boost the performance. Comprehensive evaluations on a publicly available ten-class object detection data set demonstrate the effectiveness of the proposed method. Gong Cheng 0003, Peicheng Zhou, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Semantic Annotation of High-Resolution Satellite Images via Weakly Supervised LearningabstractIn this paper, we focus on tackling the problem of automatic semantic annotation of high resolution (HR) optical satellite images, which aims to assign one or several predefined semantic concepts to an image according to its content. The main challenges arise from the difficulty of characterizing complex and ambiguous contents of the satellite images and the high human labor cost caused by preparing a large amount of training examples with high-quality pixel-level labels in fully supervised annotation methods. To address these challenges, we propose a unified annotation framework by combining discriminative high-level feature learning and weakly supervised feature transferring. Specifically, an efficient stacked discriminative sparse autoencoder (SDSAE) is first proposed to learn high-level features on an auxiliary satellite image data set for the land-use classification task. Inspired by the motivation that the encoder of the prelearned SDSAE can be regarded as a generic high-level feature extractor for HR optical satellite images, we then transfer the learned high-level features to semantic annotation. To compensate the difference between the auxiliary data set and the annotation data set, the transferred high-level features are further fine-tuned in a weakly supervised scheme by using the tile-level annotated training data. Finally, the fine-tuning process is formulated as an ultimate optimization problem, which can be solved efficiently with our proposed alternate iterative optimization method. Comprehensive experiments on a publicly available land-use classification data set and an annotation data set demonstrate the superiority of our SDSAE-based high-level feature learning method and the effectiveness of our weakly supervised semantic annotation framework compared with state-of-the-art fully supervised annotation methods. Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Xueming Qian, Lei Guo 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2015 | Learning coarse-to-fine sparselets for efficient object detection and scene classificationabstractPart model-based methods have been successfully applied to object detection and scene classification and have achieved state-of-the-art results. More recently the “sparselets” work [1-3] were introduced to serve as a universal set of shared basis learned from a large number of part detectors, resulting in notable speedup. Inspired by this framework, in this paper, we propose a novel scheme to train more effective sparselets with a coarse-to-fine framework. Specifically, we first train coarse sparselets to exploit the redundancy existing among part detectors by using an unsupervised single-hidden-layer auto-encoder. Then, we simultaneously train fine sparselets and activation vectors using a supervised single-hidden-layer neural network, in which sparselets training and discriminative activation vectors learning are jointly embedded into a unified framework. In order to adequately explore the discriminative information hidden in the part detectors and to achieve sparsity, we propose to optimize a new discriminative objective function by imposing L0-norm sparsity constraint on the activation vectors. By using the proposed framework, promising results for multi-class object detection and scene classification are achieved on PASCAL VOC 2007, MIT Scene-67, and UC Merced Land Use datasets, compared with the existing sparselets baseline methods. Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002, Tianming Liu 0001 |
CVPR | 1 |
| 2015 | Semantic Segmentation based on Stacked Discriminative Autoencoders and Context-Constrained Weakly Supervised LearningabstractIn this paper, we focus on tacking the problem of weakly supervised semantic segmentation. The aim is to predict the class label of image regions under weakly supervised settings, where training images are only provided with image-level labels indicating the classes they contain. The main difficulty of weakly supervised semantic segmentation arises from the complex diversity of visual classes and the lack of supervision information for learning a multi-classes classifier. To conquer the challenge, we propose a novel discriminative deep feature learning framework based on stacked autoencoders (SAE) by integrating pairwise constraints to serve as a discriminative term. Furthermore, to mine effective supervision information, global context about co-occurrence of visual classes as well as local context around each image region is exploited as constraints for training a multi-class classifier. Finally, the classifier training is formulated as an ultimate optimization problem, which can be solved efficiently by an alternate iterative optimization method. Comprehensive experiments on the MSRC 21 dataset demonstrate the superior performance compared with several state-of-the-art weakly supervised image segmentation methods. Xiwen Yao, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002 |
ACM Multimedia | 3 |
| 2015 | Auto-encoder-based shared mid-level visual dictionary learning for scene classification using very high resolution remote sensing imagesabstractEffective representation and classification of scenes using very high resolution (VHR) remote sensing images cover a wide range of applications. Although robust low‐level image features have been proven to be effective for scene classification, they are not semantically meaningful and thus have difficulty to deal with challenging visual recognition tasks. In this study, the authors propose a new and effective auto‐encoder‐based method to learn a shared mid‐level visual dictionary. This dictionary serves as a shared and universal basis to discover mid‐level visual elements. On the one hand, the mid‐level visual dictionary learnt using machine learning technique is more discriminative and contains rich semantic information, compared with the traditional low‐level visual words. On the other hand, the mid‐level visual dictionary is more robust to occlusions and image clutters. In the authors' scene‐classification scheme, they use discriminative mid‐level visual elements, rather than individual pixels or low‐level image features, to represent images. This new image representation is able to capture much of the high‐level meaning and contents of the image, facilitating challenging remote sensing image scene‐classification tasks. Comprehensive evaluations on a challenging VHR remote sensing images data set and comparisons with state‐of‐the‐art approaches demonstrate the effectiveness and superiority of their study. Gong Cheng 0003, Peicheng Zhou, Junwei Han 0001, Lei Guo 0002, Jungong Han |
IET Comput. Vis. | 1 |
| 2015 | Weakly Supervised Learning for Target Detection in Remote Sensing ImagesabstractIn this letter, we develop a novel framework of leveraging weakly supervised learning techniques to efficiently detect targets from remote sensing images, which enables us to reduce the tedious manual annotation for collecting training data while maintaining the detection accuracy to large extent. The proposed framework consists of a weakly supervised training procedure to yield the detectors and an effective scheme to detect targets from testing images. Comprehensive evaluations on three benchmarks which have different spatial resolutions and contain different types of targets as well as the comparisons with traditional supervised learning schemes demonstrate the efficiency and effectiveness of the proposed framework. Dingwen Zhang, Junwei Han 0001, Gong Cheng 0003, Zhenbao Liu, Shuhui Bu, Lei Guo 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Effective and Efficient Midlevel Visual Elements-Oriented Land-Use Classification Using VHR Remote Sensing ImagesabstractLand-use classification using remote sensing images covers a wide range of applications. With more detailed spatial and textural information provided in very high resolution (VHR) remote sensing images, a greater range of objects and spatial patterns can be observed than ever before. This offers us a new opportunity for advancing the performance of land-use classification. In this paper, we first introduce an effective midlevel visual elementsoriented land-use classification method based on “partlets,” which are a library of pretrained part detectors used for midlevel visual elements discovery. Taking advantage of midlevel visual elements rather than low-level image features, a partlets-based method represents images by computing their responses to a large number of part detectors. As the number of part detectors grows, a main obstacle to the broader application of this method is its computational cost. To address this problem, we next propose a novel framework to train coarse-to-fine shared intermediate representations, which are termed “sparselets,” from a large number of pretrained part detectors. This is achieved by building a single-hidden-layer autoencoder and a single-hidden-layer neural network with an L0-norm sparsity constraint, respectively. Comprehensive evaluations on a publicly available 21-class VHR landuse data set and comparisons with state-of-the-art approaches demonstrate the effectiveness and superiority of this paper. Gong Cheng 0003, Junwei Han 0001, Lei Guo 0002, Zhenbao Liu, Shuhui Bu, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Object Detection in Optical Remote Sensing Images Based on Weakly Supervised Learning and High-Level Feature LearningabstractThe abundant spatial and contextual information provided by the advanced remote sensing technology has facilitated subsequent automatic interpretation of the optical remote sensing images (RSIs). In this paper, a novel and effective geospatial object detection framework is proposed by combining the weakly supervised learning (WSL) and high-level feature learning. First, deep Boltzmann machine is adopted to infer the spatial and structural information encoded in the low-level and middle-level features to effectively describe objects in optical RSIs. Then, a novel WSL approach is presented to object detection where the training sets require only binary labels indicating whether an image contains the target object or not. Based on the learnt high-level features, it jointly integrates saliency, intraclass compactness, and interclass separability in a Bayesian framework to initialize a set of training examples from weakly labeled images and start iterative learning of the object detector. A novel evaluation criterion is also developed to detect model drift and cease the iterative learning. Comprehensive experiments on three optical RSI data sets have demonstrated the efficacy of the proposed approach in benchmarking with several state-of-the-art supervised-learning-based object detection approaches. Junwei Han 0001, Dingwen Zhang, Gong Cheng 0003, Lei Guo 0002, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Visual attention computation in video of driving environmentabstractWe here study the problem of visual attention computation in video of driving environment via the learning from eye movements. We collect a large-scale database of eye movements from 28 subjects on 30 videos of road scenes, which simulate the driving environment. The analysis on this eye movement database reveals that visual attention in driving environment is directed by high-level cognitive factors such as objects. We then present a new high-level representation called Traffic Object Bank (TOB), which is comprised of many individual road object detectors trained comprehensively in semantic space as well as viewpoint space. TOB provides semantically rich object-level features. Finally, we develop a computational model to predict where drivers look via the mapping from TOB-based representation and to gaze data. Experimental results on our traffic scene video benchmark indicate high accordance with human eye movement and show great promise for further applications. Junwei Han 0001, Liye Sun, Dingwen Zhang, Xintao Hu, Gong Cheng 0003, Lei Guo 0002 |
ICME | 5 |
| 2014 | Scalable multi-class geospatial object detection in high-spatial-resolution remote sensing imagesabstractIn this paper we present a conceptually simple but surprisingly effective multi-class geospatial object detection method based on Collection of Part Detectors (COPD), which can be easily scaled to a larger number of object classes. The presented COPD is composed of a set of representative and discriminative part detectors, where each part detector is a linear support vector machine (SVM) classifier trained using a weakly supervised learning method that only requires image labels indicating the presence of objects for the training data. Here, each part detector corresponds to a particular viewpoint of an object class, so the collection of them provides a feasible solution for rotation-invariant and simultaneous detection of multi-class geospatial objects. Comprehensive evaluations on high-spatial-resolution remote sensing images and comparisons with a number of state-of-the-art approaches demonstrate the effectiveness and superiority of the presented method. Gong Cheng 0003, Junwei Han 0001, Peicheng Zhou, Lei Guo 0002 |
IGARSS | 1 |
| 2014 | Image visual attention computation and application via the learning of object attributes
Junwei Han 0001, Ling Shao 0001, Xiaoliang Qian, Gong Cheng 0003, Jungong Han |
Mach. Vis. Appl. | 5 |
| 2013 | Optimal contrast based saliency detection
Xiaoliang Qian, Junwei Han 0001, Gong Cheng 0003, Lei Guo 0002 |
Pattern Recognit. Lett. | 3 |