Xianpeng Guo

dblp:229/6967 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-3733-2570ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs
abstract
Multimodal Large Language Models typically assume linguistic context invariably enhances visual understanding.We study this assumption in semantic adversarial scenarios, specifically magic tricks, where narration deliberately diverges from physical reality.We introduce MagicBench, a diagnostic benchmark of 402 videos for evaluating MLLMs under hierarchical linguistic interference, together with a Physical Constraint Set (PCS) protocol for assessing adherence to physical laws.Evaluation uncovers a Semantic Dependency Paradox: (1) Semantic anchoring: Entity nouns act as anchors aiding localization, paradoxically boosting performance despite false predicates.(2) Visual Agency Loss: In semantic vacuums, multimodal performance collapses 12.4% (p < 0.01) below the vision-only capability probe.This gap persists under symmetric prompting, suggesting a form of functional perception suppression in which autonomous visual search may be under-utilized in multimodal settings without linguistic triggers.Causal interventions via spatial prompting and signal magnification provide evidence that internal reasoning remains functional, supporting the interpretation of a perceptual access bottleneck.Our findings suggest MLLMs function as language-guided passive observers, advocating for perceptuallyindependent architectures that decouple sensory agency from linguistic dominance.
Tang Da Huang, Weidong Tang, Wen Qi Xu, Xianpeng Guo
ACL (1)4
2025 Interpretable Fine-Grained Aircraft Classification Network for Remote Sensing Image With Image Pair Interaction and Neural Tree
abstract
This paper proposes a novel interpretable framework for fine-grained aircraft classification in high-stakes remote sensing applications. Our approach addresses three key challenges: small inter-class variance, large intra-class variance, and the need for model interpretability. Specifically, our framework is built on the Swin Transformer (SwinT) backbone and includes three main modules. First, we present the Dynamic Attention Fusion Module (DAFM), which adaptively fuses multi-stage attention maps from the SwinT backbone. By leveraging a dispersion-based weighting mechanism, DAFM balances the contributions of coarse and fine-grained features, capturing both global structures and localized details. Second, we propose the Adaptive Image Pair Interaction Module (AIPI), which dynamically adjusts feature interaction strategies based on intra-class and inter-class similarity, effectively enhancing informative regions and improving robustness. To further optimize discriminative power, we incorporate an AIPI loss function that enforces intra-class consistency and inter-class separability. Finally, we develop a Binary Neural Tree Module (BNTM) to hierarchically select and propagate informative image patches, enhancing both feature refinement and interpretability through explicit path-based decision-making. Extensive experiments on benchmark datasets demonstrate that our framework significantly improves classification accuracy and interpretability, making it well-suited for applications requiring transparent and reliable decision-making. The Code can be found at https://github.com/StarmanGzx/BNTM.
Zhengxi Guo, Biao Hou, Xianpeng Guo, Chen Yang 0027, Zitong Wu, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Masked Angle-Aware Autoencoder for Remote Sensing Images
Zhihao Li 0005, Biao Hou, Siteng Ma, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Licheng Jiao
ECCV (8)5
2024 Weakly supervised object localization via knowledge distillation based on foreground-background contrast
Siteng Ma, Biao Hou, Zhihao Li 0005, Zitong Wu, Xianpeng Guo, Chen Yang 0027, Licheng Jiao
Neurocomputing5
2024 Unreliable Pixel Contrast Based on von Mises-Fisher Distribution for Semi-Supervised SAR Segmentation
abstract
The unique visual properties and the huge size of synthetic aperture radar (SAR) images pose challenges in labeling data. The scarcity of labeled data limits the training of SAR image segmentation networks. To address this issue, semi-supervised methods are used to train the network. However, traditional semi-supervised algorithms like Mean Teacher often generate erroneous pseudo-labels that carry over to subsequent training epochs, leading to overfitting and affecting network performance. This overfitting stems from an overreliance on unreliable pixels in Mean Teacher. In this study, an enhanced approach is proposed called unreliable pixel contrast (UPCo), where a von Mises-Fisher distribution is applied to constrain unreliable pixels in feature space. We augment the segmentation network with a feature output header for pixel-level contrastive learning in UPCo. Moreover, to minimize the computational effort during the training phase, hard sample selection and negative object non-uniform selection strategies are designed to facilitate contrastive learning. The proposed UPCo was evaluated on two large-scene SAR images and demonstrated its superiority over other comparative algorithms, achieving more optimal performance in semi-supervised segmentation of SAR images.
Zitong Wu, Biao Hou, Xianpeng Guo, Bo Ren 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 MSRIP-Net: Addressing Interpretability and Accuracy Challenges in Aircraft Fine-Grained Recognition of Remote Sensing Images
abstract
The task of fine-grained aircraft recognition is crucial in the field of remote sensing. Despite some progress achieved by traditional deep learning methods in addressing this challenge, they are often perceived as a “black box,” lacking transparent explanations for model decisions. Current interpretable methods based on attention mechanisms, although providing some interpretability, do not align with human thought logic. Therefore, we propose a multiscale rotation-invariant prototype network (MSRIP-Net). Our approach simulates the intuitive reasoning process of humans in identifying objects by segmenting them into multiple components. Importantly, MSRIP-Net has the capability to automatically recognize rigid components on aircraft targets without relying on additional part annotations, using only image-level class labels. In addition, our approach effectively addresses challenges presented by noise, deformations, and multiscale variations in remote sensing targets and has been comprehensively evaluated on datasets FAIR1M1.0 and Rareplane. Our results demonstrate that MSRIP-Net achieves higher accuracy compared with existing fine-grained recognition methods. Furthermore, we provide insights into the model’s decision-making process to illustrate the interpretability of our approach.
Zhengxi Guo, Biao Hou, Xianpeng Guo, Zitong Wu, Chen Yang 0027, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2024 Rebalancing Gaussian Location Loss for High-Precision Detection on Remote Sensing Images
abstract
Aerial image objects are usually orientated arbitrarily, with a large scale range, and densely distributed. Traditional horizontal bounding box (HBB) detectors tend to filter out densely distributed objects leading to missed detections, such as ship (SH) and vehicle. Therefore, oriented object detection has become a mainstream solution in recent years. The 2-D Gaussian distribution representation of the oriented bounding boxes (OBBs) solves the problem of angular discontinuity and boundary discontinuity and thus gets more attention. However, as the aspect ratio of the object gradually decreases, its predicted angular performance continues to decrease. We find that the angular gradient of an object decreases sharply as the aspect ratio decreases, resulting in a large gradient gap between a small aspect ratio object (SARO) and a large aspect ratio object (LARO). It makes the detector prefer to ignore SARO during training, which weakens the high precision performance of SARO. We call this phenomenon shape imbalance. To solve the problem, we proposed a simple gradient rebalancing strategy named shape balance. Since the shape imbalance is only related to the aspect ratio of the object, we designed a modulation function with an inverse aspect ratio to calculate the balance coefficient. The principle of the function is that the larger the aspect ratio, the smaller the balance coefficient; the smaller the aspect ratio, the larger the balance coefficient. We aim to get the balance coefficients for objects with different aspect ratios. Location loss multiplied by a balance coefficient can directly adjust the gradient gap between objects with different aspect ratios to achieve a rebalancing effect. Extensive experiments conducted on DOTA-v1.0 dataset and DIOR-R dataset verify the effectiveness of our proposed method. Our method improves the detection performance of Gaussian location loss by an average of 2.08%/1.01%(AP75/mAP) metrics on the DOTA-v1.0 dataset and 1.17%/0.82%(AP75/mAP) improvements for DIOR-R dataset.
Biao Hou, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Zhongle Ren, Chen Yang 0027, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Contrastive Learning Based on Multiscale Hard Features for Remote-Sensing Image Scene Classification
abstract
The overwhelming majority of models for remote sensing image (RSI) scene classification generally require the weights pre-trained on natural images for initialization before formal training. However, differences in imaging mechanisms lead to huge discrepancies between natural images and RSIs, and the strong visual representation learned from massive natural images limits the performance of models when inferencing RSIs. To address this issue, the well-established self-supervised contrastive learning paradigm in the natural image field is introduced to the RSI field. We propose a contrastive learning method based on multi-scale hard features, MHCL, which aims to use finite RSIs to learn sufficient visual representations in an unsupervised contrastive manner, thus provide a powerful upstream pre-trained model for fine-tuning downstream scene classification task. Multi-level features extracted by intermediate layers of each encoder’s backbone are first gathered, and then a hard features transformation method is proposed to create hard positive features and diverse queues that save hard negatives, thereby enriching the finite scene information in small-scale RSIs. Furthermore, we redesign the multi-scale hard features joint contrastive loss to boost the model to explore sufficient invariant representations by additionally pulling hard positive pairs closer and pushing hard negative pairs farther away in the embedding space. Extensive experiments demonstrate that the upstream pre-training model generated by MHCL achieves competitive transferred performance on three popular scene classification datasets, outperforming the traditional model pre-trained on ImageNet and models pre-trained by other state-of-the-art contrastive learning methods. Our code will be released at: https://github.com/benesakitam/MHCL.
Zhihao Li 0005, Biao Hou, Xianpeng Guo, Siteng Ma, Yanyu Cui, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2023 Complete Rotated Localization Loss Based on Super-Gaussian Distribution for Remote Sensing Images
abstract
Localization regression in oriented object detection tasks has long faced boundary discontinuity and angular discontinuity problems induced by periodic angles. These problems were successfully resolved by using a 2d Gaussian distribution to modelling the oriented bounding box (OBB). However, the angular information of square-like objects will be lost when they are converted to 2d Gaussian distribution, forming a systematic problem. Its fundamental reason is that when the aspect ratio of the object tends to 1, the equiprobability curve of 2d Gaussian distribution degenerates from an ellipse to a circle, thus losing the orientation information of the rotated object. This results in the bounding boxes of such square-like objects not being learned effectively. To resolve this problem, we used the Lamé curve (or superellipse) to modify the existing 2d Gaussian function and designed a super-Gaussian distribution. This distribution can maintain anisotropy at arbitrary aspect ratios, thus preserving the angular information of the oriented object. We used the Kullback-Leibler (KL) divergence to measure the distance between two super-Gaussian distributions and convert it into a localization loss (SGKLD) by a function. SGKLD is an improved version of KLD loss. By modifying the form of the probability distribution, we elegantly fix the angle missing problem of the traditional Gaussian distribution. We validated the effectiveness of the proposed algorithm on several datasets and obtained the performance of SOTA. Our algorithm achieves a mean average precision (mAP) of 80.07, 76.59, 62.27, and 90.55/98.13 on the DOTA-v1.0, DOTA-v1.5, DOTA-v2.0, and HRSC2016 datasets, respectively.
Biao Hou, Zitong Wu, Zhengxi Guo, Bo Ren 0001, Xianpeng Guo, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2023 Unsupervised Prototype-Wise Contrastive Learning for Domain Adaptive Semantic Segmentation in Remote Sensing Image
abstract
Labeling data in the field of remote sensing is time-consuming and labor-intensive, making domain adaptation between different domains an urgently needed solution. To address the domain gap between diverse datasets in the remote sensing domain, numerous methods tailored for domain adaptation in high-resolution remote sensing imagery have emerged. Some of the existing methods focus on reducing the domain gap at either the feature level or the pixel level, often overlooking their underlying connection. To tackle this issue, we introduce a prototype-wise contrastive feature alignment paradigm (PCFA) aimed at bridging the representations between the feature and pixel levels. By dynamically updating, we acquire prototype information encompassed by different mini-batches and employ an optimal transport mechanism to reasonably apply the prototype feature distribution in guiding the learning of target domain features. We conduct extensive domain adaptation semantic segmentation (DASS) experiments on the ISPRS Vaihingen and Potsdam datasets, achieving an improvement about 4%~5% in mIoU (mean Intersection over Union) compared to previous methods using the DeepLabV2 framework.
Siteng Ma, Biao Hou, Xianpeng Guo, Zitong Wu, Zhihao Li 0005, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2023 Automatic Aug-Aware Contrastive Proposal Encoding for Few-Shot Object Detection of Remote Sensing Images
abstract
In the annotation of remote sensing images (RSIs), the effectiveness of common object detection methods trained on only a few samples decreases instantly, which has prompted increasing research on the few-shot problem in remote sensing. RSIs often exhibit suboptimal performance in few-shot scenarios due to the intricate nature of scene information interference and the high degree of cosine similarity, both of which present significant challenges to their effectiveness. In this paper, a two-stage detection framework based on fine-tuning is selected to deal with the common problems in the few-shot task of remote sensing domain. Considering the excessive scale variation of instances in remote sensing datasets, we introduce an automatically learned aug-aware search module to provide an intelligent data augmentation solution for Faster R-CNN using different optimal augmentation policies searched by the network to fit the current dataset. We introduce a contrastive RoI branch to better classify novel class proposal features that are easily confused by the base class. We named our work AACE and conducted extensive experiments on two common object detection datasets in remote sensing, NWPU VHR 10 and DIOR, on which AACE achieved about 2.30% and 2.61% improvement, respectively, in the number of shots listed in the paper, compared to other algorithms.
Siteng Ma, Biao Hou, Zitong Wu, Zhihao Li 0005, Xianpeng Guo, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Network Pruning for Remote Sensing Images Classification Based on Interpretable CNNs
abstract
Convolutional neural network (CNN)-based research has been successfully applied in remote sensing image classification due to its powerful feature representation ability. However, these high-capacity networks bring heavy inference costs and are easily overparameterized, especially for the deep CNNs pretrained on natural image datasets. Network pruning is regarded as a prevalent approach for compressing networks, but most existing research ignores model interpretability while formulating pruning criterion. To address these issues, a network pruning method for remote sensing image classification based on interpretable CNNs is proposed. More specifically, an original interpretable CNN with a predefined pruning ratio is trained at first. The filters, namely channels in the high convolutional layer, are able to learn specific semantic meanings in proportion to the predefined pruning ratio. The filters without interpretability are supposed to be removed. As for other convolutional layers, a sensitivity function is designed to assess the risk of pruning channels for each layer, and furthermore, the pruning ratio for each layer is corrected adaptively. The pruning method based on the proposed sensitivity function is effective and requires little computational costs to search abandoned channels without damaging classification performance. To demonstrate the effectiveness, the proposed method is implemented on different scales of modern CNN models, including VGG-VD and AlexNet. The experimental results, obtained on the UC Merced dataset and NWPU-RESISC45 dataset, prove that our method significantly reduces the inference costs and improves the interpretability of networks.
Xianpeng Guo, Biao Hou, Bo Ren 0001, Zhongle Ren, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 A Neural Network Based on Consistency Learning and Adversarial Learning for Semisupervised Synthetic Aperture Radar Ship Detection
abstract
Ship detection in synthetic aperture radar (SAR) images has important application value. Sea clutter, complex scenes, a large size change in ships, and the arbitrary directionality of ships make ship detection challenging. With the development of deep learning, many deep learning algorithms have been applied to SAR images. These algorithms need a lot of labeled data for training. It is time-consuming to label SAR data, and the unlabeled data are easy to obtain. It is necessary to use the unlabeled data effectively to improve the performance of the algorithm. In this study, a semisupervised SAR ship detection network, named the semisupervised consistency learning adversarial network (SCLANet), is presented. SCLANet is a two-stage detection network. The local features around the ship can be extracted by the SCLANet, and the features generated from unlabeled data become closer to those generated from labeled data by using adversarial learning. There are two consistency learning modules in SCLANet: noise robustness consistency learning and output encoding consistency learning. Noise robustness consistency learning can increase the robustness of the SCLANet. Maintaining consistency between the noisy results and the original results can train the unlabeled data. In output encoding consistency learning, outputs are mapped to a picture that is fed into an encoder to obtain the intermediate representation embedding. Another embedding is a layer in the main network of the SCLANet. Reducing the error between two embeddings can train the SCLANet with unlabeled data. Two types of consistency learning can be used as pretext tasks for semisupervised learning. Experiments were conducted on two SAR ship datasets. Compared with other algorithms, the SCLANet achieved the highest detection accuracy, indicating that it is more advantageous to use in ship detection.
Biao Hou, Zitong Wu, Bo Ren 0001, Xianpeng Guo, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Multicrop Fusion Strategy Based on Prototype Assignment for Remote Sensing Image Scene Classification
abstract
The gap between self-supervised visual representation learning and supervised learning is gradually closing. Self-supervised learning does not rely on a large amount of labeled data and reduces the loss of human labeled information. Compared with natural images, remote sensing images require rich samples and human annotation by experts. Moreover, many algorithms have poor interpretability and unconvincing results. Therefore, this paper proposes a self-supervised method based on prototype assignment by designing a pretext task so that the network maps features to prototypes in the process of learning, swaps the code corresponding to the obtained features, combines them with another data-enhancing feature, and then optimizes the network. The prototype is introduced to explain the clustering idea embodied in the whole process. Considering the existence of the scene information-rich characteristic of remote sensing images, we introduce multiple views with different resolutions to capture more detailed information on the images. Finally, if the data enhancement method is not powerful enough, the network can easily fall into an overfitting state, which prevents the network from learning subtle differences and detailed information. To address this shortcoming, we propose a fusion strategy to flatten the decision boundary of the framework so that the model can also learn the soft similarity between sample pairs. We name the whole framework MFPC. In extensive experiments conducted on three common remote sensing image datasets (i.e., UCMerced, AID, and NWPU45), MFPC achieves a maximum improvement of 4.3% over some existing self-supervised algorithms, indicating that it can achieve good results.
Siteng Ma, Biao Hou, Xianpeng Guo, Zhihao Li 0005, Zitong Wu, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2022 Reconstruction Error-Based Decomposition Feature Selection for PolSAR Image
abstract
Target decomposition features are the cornerstone of subsequent analyses for PolSAR images. Generally, adopting single or several decomposition algorithms limits the representation ability for original terrain characteristics. Using all the existing decomposition features, however, will definitely increase computational complexity. Besides, some features even have a negative effect on the following tasks. To address these problems, a sparse variational autoencoder feature selection framework (SVAE-FS) is proposed in this article. In detail, the encoder transforms the original feature set into latent space and then decoder reconstructs the corresponding pseudo set on this latent space. Similarly, a pseudo subset is subsequently obtained by the SVAE. The discrepancy, namely reconstruction error, between the pseudo set and the pseudo subset is taken as an evaluation criterion which reflects the feature representation ability of pseudo subset. Sparse constraint in the encoder makes the representative features stand out. Meanwhile, the linear feature transformation layer of the encoder enables the SVAE to evaluate different scale subsets without repeated training. Finally, a greedy selection approach with search scale$K$is proposed to find the suboptimal subset. This procedure not only reduces time consumption, but also ensures the performance of the subset. The selected features are analyzed on four real PolSAR datasets according to the terrain scattering characteristics. Furthermore, these features have achieved competitive performance on three PolSAR image tasks.
Chen Yang 0027, Biao Hou, Xianpeng Guo, Bo Ren 0001, Jocelyn Chanussot, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2018 PolSAR Image Classification Based on DBN and Tensor Dimensionality Reduction
abstract
This paper proposes a new semi-supervised PolSAR image classification method using deep belief network (DBN) and tensor dimensionality reduction, which uses multilinear principle component analysis (MPCA) to reduce the dimension of tensor form PolSAR data, and regards the multiple features of PolSAR data as the input of DBN. In order to take full advantage of neighborhood information of each pixel of PolSAR data, we take each pixel and its neighborhood as tensor form. For PolSAR data, simple feature has been proven not to be able to effectively classify complex terrains. Therefore, we combine multiple features of PolSAR data to obtain more abundant information, which can reflect some spatial structure of PolSAR data. The experimental results show that the overall classification accuracy based on the proposed method outperforms the traditional classification strategies.
Biao Hou, Xianpeng Guo, Weidan Hou, Shuang Wang 0001, Xiangrong Zhang, Licheng Jiao
IGARSS2