Haixia Xu 0003

dblp:26/5544-3 · also Hai-Xia Xu 0003 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0003-3721-665XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ACDR-CRAFF Net: A Multi-Scale Network Based on Adaptive Channel and Coordinate Relational Attention Network for Remote Sensing Scene Classification
abstract
ABSTRACT Accurate classification of remote sensing scene images is crucial for diverse applications, from environmental monitoring to urban planning. While convolutional neural networks (CNNs) have dramatically improved classification accuracy, challenges remain due to the complex distribution of small objects, varied spatial configurations, and intra‐class multimodality in remote sensing images. In this work, we make three key contributions to address these challenges. (1) We propose the adaptive channel and coordinate relational attention network (ACDR‐CRAFF), a novel multi‐scale feature fusion framework designed to enhance feature representation across scales. (2) We introduce two innovative modules: the adaptive channel dimensionality reduction (ACDR) module, which dynamically adjusts channel representations to retain essential low‐dimensional features, and the coordinate relational attention multi‐scale feature fusion (CRAFF) module, which effectively captures and transfers spatial information between feature levels. (3) By integrating ACDR and CRAFF, our model achieves a progressive fusion of local to global features, ensuring robust feature expressiveness at multiple scales. Experimental results on four widely used benchmark datasets demonstrate that ACDR‐CRAFF consistently outperforms several state‐of‐the‐art methods, achieving significant improvements in classification accuracy and setting a new benchmark for complex remote sensing scene classification tasks. These results underscore the effectiveness of our approach in addressing the limitations of existing methods and advancing the state of the art in remote sensing image analysis.
Haixia Xu 0003, Furong Shi, Xianbin Wen
IET Image Process.2
2024 LPNet: A remote sensing scene classification method based on large kernel convolution and parameter fusion
abstract
Abstract Remote sensing scene images contain numerous feature targets with unrelated semantic information, so how to extract to the local key information and semantic features of the image becomes the key to achieving accurate classification. Existing Convolutional Neural Networks (CNNs) mostly concentrate on the global representation of an image and lose the shallow features. To overcome these issues, this paper proposes LPNet for remote sensing scene image classification. First, LPNet employs LKConv to extract the semantic features in the image, while using standard convolution to extract local key information in the image. Additionally, the LPNet applies a shortcut residual concatenation branch to reuse features. Then, parameter fusion combines parameters from previous branches, improving the capacity of the model to obtain a more comprehensive and rich feature representation of the image. Finally, considering the relationship between the classification ability of the model and the depth of feature extraction, the Feature Mixture (FM) Block is used to deepen the model for feature extraction. Comparative experiments on four publicly available datasets show that LPNet provides comparable results to other state‐of‐the‐art methods. The effectiveness of LPNet is further demonstrated by visualizing the effective receptive fields (ERFs).
Furong Shi, Haixia Xu 0003, Xianbin Wen
IET Image Process.4
2023 A lightweight and stochastic depth residual attention network for remote sensing scene classification
abstract
Abstract Due to the rapid development of satellite technology, high‐spatial‐resolution remote sensing (HRRS) images have highly complex spatial distributions and multiscale features, making the classification of such images a challenging task. The key to scene classification is to accurately understand the main semantic information contained in images. Convolutional neural networks (CNNs) have outstanding advantages in this field. Deep CNNs (D‐CNNs) with better performance tend to have more parameters and higher complexity. However, shallow CNNs have difficulty extracting the key features of complex remote sensing images. In this paper, we propose a lightweight network with a random depth strategy for remote sensing scene classification (LRSCM). We construct a convolutional feature extraction module, DCAB, which incorporates depthwise separable convolutional and inverted residual structures, effectively reducing the numbers of required parameters and computations, and retains and utilizes low‐level features. In addition, coordinate attention (CA) is integrated into the module, thereby further improving the network's ability to extract key local information. To further reduce the complexity of model training, the residual module adopts a stochastic depth strategy, providing the network with a random depth. Comparative experiments on five public datasets show that the LRSCM network can achieve results comparable to those of other state‐of‐the‐art methods.
Haixia Xu 0003, Xianbin Wen
IET Image Process.2
2022 Global Correlative Network for Person re-identification
Gengsheng Xie, Xianbin Wen, Haixia Xu 0003, Zhanlu Liu
Neurocomputing4
2022 DASFTOT: Dual attention spatiotemporal fused transformer for object tracking
Ruixu Wu, Xianbin Wen, Haixia Xu 0003
Knowl. Based Syst.4
2022 STASiamRPN: visual tracking based on spatiotemporal and attention
Ruixu Wu, Xianbin Wen, Zhanlu Liu, Haixia Xu 0003
Multim. Syst.5
2022 Correction: STASiamRPN: visual tracking based on spatiotemporal and attention
Ruixu Wu, Xianbin Wen, Zhanlu Liu, Haixia Xu 0003
Multim. Syst.5
2022 Multiview clustering via consistent and specific nonnegative matrix factorization with graph regularization
Haixia Xu 0003, Limin Gong, Hai-Zhen Xuan, Xusheng Zheng, Zan Gao 0001, Xianbin Wen
Multim. Syst.1
2021 Channel-exchanged feature representations for person re-identification
Jianchen Wang, Haixia Xu 0003, Gengsheng Xie, Xianbin Wen
Inf. Sci.3
2021 Dense capsule networks with fewer parameters
Xianbin Wen, Haixia Xu 0003
Soft Comput.4
2020 Learnable Bag Similarity Based Deep Multi-Instance Network for Breast Cancer Diagnosis
abstract
Computer-aided diagnosis for breast cancer is an important and challenging research problem in medical image analysis. The main difficulty lies in that only image-level labels rather than fine-grained patch-level labels can be gotten for medical images in general. This situation fits well with the settings of Multi-Instance Learning (MIL). Following this line of research, Bag Similarity Network (BSN) uses the inter-bag similarities to learn the relationship between bags (images) and achieves a good performance in automatic diagnosis of breast cancer. Nevertheless, the inter-bag similarities are pre-defined rather than learnable. In this paper, we propose a Learnable Bag Similarity Network for deep MIL, called LBSN, to aid breast cancer diagnosis. To implement automatic similarity learning, LBSN first extracts fixed numbers of global representations of each bag using the attention mechanism, and then employs channel and spatial attention on similarity matrices for learning the relation of bags and that of instances, respectively. The experimental results on a publicly available breast cancer data set demonstrate that the proposed LBSN outperforms BSN by a large margin in terms of classification accuracy.
Rui Cheng 0008, Haixia Xu 0003, Zhenliang Li, Xianbin Wen
BIBM3
2020 Deep Multi-Instance Learning with Induced Self-Attention for Medical Image Classification
abstract
Existing Multi-Instance learning (MIL) methods for medical image classification typically segment an image (bag) into small patches (instances) and learn a classifier to predict the label of an unknown bag. Most of such methods assume that instances within a bag are independently and identically distributed. However, instances in the same bag often interact with each other. In this paper, we propose an Induced SelfAttention based deep MIL method that uses the self-attention mechanism for learning the global structure information within a bag. To alleviate the computational complexity of the naive implementation of self-attention, we introduce an inducing point based scheme into the self-attention block. We show empirically that the proposed method is superior to other deep MIL methods in terms of performance and interpretability on three medical image data sets. We also employ a synthetic MIL data set to provide an intensive analysis of the effectiveness of our method. The experimental results reveal that the induced self-attention mechanism can learn very discriminative and different features for target and non-target instances within a bag, and thus fits more generalized MIL problems.
Zhenliang Li, Haixia Xu 0003, Rui Cheng 0008, Xianbin Wen
BIBM3
2020 Local-Variance-Based Attention For Visual Tracking
abstract
The RoIAlign module incorporated into the deep tracking-by-detection framework, which can thus receive the entire image as the input to the convolutional layer, alleviating high computational complexity induced by multiple proposals. Nevertheless, this would also produce an ambiguous feature discriminative boundary between the target and background in the feature map, which makes the following target identification and localization very difficult. To solve this problem, we apply a novel local-variance-based regularization for optimizing the convolutional layer, the local variance calculated from the attention map, i.e., the average pooling of the convolutional feature map. Therefore, the binary classification loss function integrated with local-variance-based regularization item can explicitly make the response of the target and background very distinguishable, specifically strengthening the response of target and weakening that of background. Extensive experiments on large-scale benchmark data sets demonstrate that the proposed algorithm is highly comparable to other state-of-the-art methods.
Changlun Guo, Xianbin Wen, Haixia Xu 0003
ICME4
2018 Multiple- Instance Learning with Empirical Estimation Guided Instance Selection
abstract
The embedding based framework handles the multiple-instance learning (MIL) via the instance selection and embedding. It is how to select instance prototypes that becomes the main difference between various algorithms. Most current studies depend on single criteria for selecting instance prototypes. In this paper, we adopt two kinds of instance-selection criteria from two different views. For the combination of the two-view criteria, we also present an empirical estimator under which the two criteria compete for the instance selection. Experimental results validate the effectiveness of the proposed empirical estimator based instance-selection method for MIL.
Xianbin Wen, Haixia Xu 0003
ICPR3
2018 An Iterative Instance Selection Based Framework for Multiple-Instance Learning
abstract
The instance selection based model is an effective multiple-instance learning (MIL) framework, which solves the MIL problems by embedding examples (bags of instances) into a new feature space formed by some concepts (represented by some selected instances). Most previous studies use single-point concepts for the instance selection, where every possible concept is represented by only a single instance. In this paper, we apply multiple-point concepts for choosing instances, in which each possible concept is jointly represented by a group of similar instances. Furthermore, we establish an iterative instance selection based MIL framework based on multiple-point concepts, which is guaranteed to automatically converge to the needed number of concepts for a given problem. The experimental results demonstrate that the proposed framework can better handle not only common MIL problems but also hybrid ones compared to state-of-the-art MIL algorithms.
Xianbin Wen, Haixia Xu 0003
ICTAI4
2017 A robust approach of watermarking in contourlet domain based on probabilistic neural network
Jiaxing Liu 0001, Xianbin Wen, Haixia Xu 0003
Multim. Tools Appl.4
2015 Multi-instance learning via instance-based and bag-based representation transformations
abstract
Recent studies show that multi-instance learning can be cast to the standard supervised learning through representation transformation in that every bag is embedded into a feature space defined by bags or instances in the training set. However, all instances from the same bag are considered to be of equal importance in the bag-based representation transformation. In this paper, we propose a new multi-instance learning algorithm by jointly considering both instance-based and bag-based representation transformations. It can be roughly divided into two steps. In the first step, the instance-based transformation is used to evaluate the importance of every instance in a bag. In the second step, the importance information is exploited to compute the weighted distances from the bag to all training bags in order to achieve the bag-based transformation. We have performed extensive experiments on several multi-instance data sets. The experimental results demonstrate the effectiveness of the proposed algorithm.
Haixia Xu 0003
ICIP3
2010 Adaptive kernel principal component analysis
Mingtao Ding, Zheng Tian 0001, Haixia Xu 0003
Signal Process.3