Shufang Xu

dblp:92/2729 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Complementary information-guided interactive fusion network for HSI and LiDAR data joint classification
Shufang Xu, Qiyuan Xue, Zhonghao Chen, Shuyu Fei, Hongmin Gao 0001
Expert Syst. Appl.1
2025 MKGFA: Multimodal Knowledge Graph Construction and Fact-Assisted Reasoning for VQA
abstract
Knowledge-based visual question answering relies on open-ended external knowledge and a fine-grained comprehension of both the visual content of images and semantic information. Existing methods for utilizing knowledge have the following limitations: (1) Language pre-training methods output answers in the form of plain text, which only understand shallow visual content; (2) The knowledge retrieved by image objects as labels is represented as first-order logic, making it difficult to infer complex questions. To address the above problems, this paper integrates visual-textual multimodal information, accumulates domain-specific and external multi-modal knowledge, introduces and supplements external objective facts, and proposes a multimodal knowledge graph construction and fact-assisted reasoning network (MKGFA). The network consists of three parts: the multimodal knowledge graph construction module (MKGC), the objective fact-assisted reasoning module (FAR), and the answer inference module. The MKGC engages in the coarse-to-fine-grained learning of triplet representations for multimodal knowledge units. The FAR establishes deep cross-modal relations between visual objects and factual words for correlating real answers. The answer inference module makes the final decision based on the results of both. Among them, the former two modules employ a pre-training and fine-tuning strategy, systematically accumulating foundational and domain-specific knowledge. Compared with the state-of-the-arts, MKGFA achieves 1.09% and 0.7% higher accuracy on the two challenging OKVQA and KRVQA datasets, respectively. The experimental results demonstrate the complementary advantages of the integration of the two modules.
Longbao Wang, Libing Zhang, Shufang Xu, Hongmin Gao 0001
Int. J. Comput. Intell. Appl.5
2025 Dual-Feature Attention Hybrid GCN Mamba Network for Joint Hyperspectral and LiDAR Classification
abstract
Hyperspectral images (HSIs) and light detection and ranging (LiDAR) data provide complementary spectral-spatial and elevation information, respectively, whose fusion can significantly improve classification accuracy. However, their inherent heterogeneity challenges effective spectral-geospatial integration. Although convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformer models have advanced multimodal remote sensing classification, each shows distinct limitations. CNNs excel in spatial feature aggregation but lack global context, whereas RNNs and Transformers, despite capturing long-range spectral features, face issues such as computational inefficiency. To address these limitations, we propose a dual-feature attention hybrid graph convolutional network (GCN) Mamba network (DAHGMN) for joint HSI and LiDAR classification. Specifically, multimodal image cubes are first extracted by a CNN to obtain initial features. Subsequently, a dual-feature attention (DA) module is introduced to adaptively recalibrate spectral and spatial feature weights, enhancing discriminability. Furthermore, we propose a hybrid GCN Mamba (HGM) module with both low parameter complexity and time complexity, which combines the local geometric modeling capability of GCNs with the global long-range dependency modeling of Mamba’s state-space model (SSM). A probability-based decision fusion strategy is employed to integrate multi-level classification results, achieving an efficient combination of the spatial-spectral contextual features. Extensive experiments on three benchmark HSI-LiDAR datasets demonstrate that DAHGMN achieves superior classification accuracy while significantly reducing parameter complexity compared to state-of-the-art methods. The implementation code is publicly available at https://github.com/RogsXie/DAHGMN.
Zhenyang Xie, Hongmin Gao 0001, Shufang Xu, Haihua Xie
IEEE Trans. Geosci. Remote. Sens.4
2025 Multiscale Segmentation-Guided Fusion Network for Hyperspectral Image Classification
abstract
Convolution Neural Networks (CNNs) have demonstrated strong feature extraction capabilities in Euclidean spaces, achieving remarkable success in hyperspectral image (HSI) classification tasks. Meanwhile, Graph convolution networks (GCNs) effectively capture spatial-contextual characteristics by leveraging correlations in non-Euclidean spaces, uncovering hidden relationships to enhance the performance of HSI classification (HSIC). Methods combining GCNs with CNNs have achieved excellent results. However, existing GCN methods primarily rely on single-scale graph structures, limiting their ability to extract features across different spatial ranges. To address this issue, this paper proposes a multiscale segmentation-guided fusion network (MS2FN) for HSIC. This method constructs pixel-level graph structures based on multiscale segmentation data, enabling the GCN to extract features across various spatial ranges. Moreover, effectively utilizing features extracted from different spatial scales is crucial for improving classification performance. This paper adopts distinct processing strategies for different feature types to enhance feature representation. Comparative experiments demonstrate that the proposed method outperforms several state-of-the-art (SOTA) approaches in accuracy. The source code will be released at https://github.com/shengrunhua/MS2FN.
Hongmin Gao 0001, Runhua Sheng, Yuanchao Su, Zhonghao Chen, Shufang Xu, Lianru Gao
IEEE Trans. Image Process.5
2024 Dual Adaptive Compression for Efficient Communication in Heterogeneous Federated Learning
abstract
In federated learning, multiple rounds of communication are involved between clients and the server to train a global model. The extensive model updates transmitted during the training lead to significant communication costs. Previous methods usually employ quantization or sparsification to compress model updates. However, the lossy compression leads to a decline in accuracy, it is challenging to strike a balance between communication efficiency and model accuracy. Meanwhile, due to the data heterogeneity, local updates among different clients are biased towards each other. Employing the same compression ratios for each local updates will further degrade the model accuracy. To achieve the trade-off between communication efficiency and model accuracy, we propose FedDAC, a Dual Adaptive Compression method in heterogeneous federated learning. In the local computation phase, the loss queue is adopted to detect the convergence trends within each client. FedDAC can then dynamically quantify model updates and allow for various compression ratios among heterogeneous clients. In the global aggregation phase, FedDAC can determine the fluctuations in training based on the similarity between clients and the server, thereby adjusting the sparsity ratio flexibly. To alleviate the reduction in model accuracy caused by lossy compression, we introduce residual updates in the local computation and global aggregation phases to maintain model accuracy. Experiment results show that compared with one-way compression methods NAGC and AdaQuantFL, FedDAC can maintain comparable accuracy while the accumulated communication volume is reduced by about 29.6 times, and 22.8 times, respectively. Moreover, the global model accuracy of FedDAC surpasses the two-way compression method T-FedAvg by about 2.4%, and the accumulated communication volume is about 2.5 times lower than T-FedAvg.
Yingchi Mao, Chenxin Li, Jiakai Zhang, Shufang Xu, Jie Wu 0001
CCGrid5
2024 A cross-modal feature aggregation and enhancement network for hyperspectral and LiDAR joint classification
Hongmin Gao 0001, Jun Zhou 0001, Pedram Ghamisi, Shufang Xu, Bing Zhang 0001
Expert Syst. Appl.6
2024 A dual-branch siamese spatial-spectral transformer attention network for Hyperspectral Image Change Detection
Shufang Xu, Hongmin Gao 0001
Expert Syst. Appl.4
2024 Interactive Enhanced Network Based on Multihead Self-Attention and Graph Convolution for Classification of Hyperspectral and LiDAR Data
abstract
The fusion of multimodal data plays a crucial role in classification tasks. However, existing research typically mines and analyzes the individual features of each data source separately before considering how to fuse them. In contrast, our approach first constructs interactive enhanced fusion features (IEFFs) for initial fusion while considering the extraction of individual features and, finally, integrates them effectively to utilize the information from each data source more comprehensively. To this end, we propose a novel interactive enhanced network based on multihead self-attention (MSA) and graph convolution. Specifically, we extract individual features from hyperspectral image (HSI) and light detection and ranging (LiDAR) data and then construct IEFFs based on the row and column features of the central pixel. Individual features focus on the local characteristics of a single data source, while IEFFs strengthen the feature expression of the central pixel through matrix operations, integrating the complementary information of multimodal data. Subsequently, we use graph convolutional networks (GCNs) to construct graph structures for four types of features (interactive enhanced HSI features, interactive enhanced LiDAR features, HSI individual features, and LiDAR individual features), modeling the pixels as nodes and capturing spatial relationships. On this basis, we apply an MSA mechanism to mine spectral dependencies, further extracting global spectral features. Finally, we design a multimodal gated fusion module (MGFM) that effectively integrates these features through its weighting mechanism. The weight allocation is adjusted dynamically according to the characteristics of the feature, achieving optimal fusion of multimodal data. Extensive experiments on three popular HSI and LiDAR datasets verify the superior performance of our method. Our code will be available athttps://github.com/haofeng0003/MSA-GCN.
Hongmin Gao 0001, Shuyu Fei, Runhua Sheng, Shufang Xu, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Multiscale Random-Shape Convolution and Adaptive Graph Convolution Fusion Network for Hyperspectral Image Classification
abstract
Convolution neural networks (CNNs) are extensively utilized in hyperspectral image (HSI) classification due to their remarkable capability to extract features from patterns with fixed shapes. These networks have been shown to effectively capture features at the pixel level. However, the fixed shape of convolution kernels poses a challenge for CNNs to adapt to the diverse shapes found in HSIs. Graph neural networks (GNNs), particularly graph convolution networks (GCNs), possess robust feature extraction capabilities on graph structures and are extensively applied in HSI classification. However, one significant challenge in using GNNs is the selection of appropriate neighboring nodes for information aggregation. To address the existing challenges of GCN and CNN and leverage their respective advantages, this paper introduces a novel patch-based CNN-GCN fusion classification network, named multi-scale random-shape convolution and adaptive graph convolution fusion network (MRCAGCFN). It consists of a spectral transformation module and three main modules we proposed: a multi-scale random-shape convolution module for extracting convolution features, where the shape of the convolution kernel is randomized and a multi-scale approach is applied to enhance adaptability to data with diverse shapes; an adaptive feature-fusion graph convolution module for extracting graph convolution features, where the weights for neighborhood aggregation are learned adaptively to reduce feature fusion from dissimilar nodes and strengthen feature fusion from similar nodes; and an adaptive local feature processing module for processing features, where two different methods are employed to convert patch-level features to pixel-level features, thereby improving feature representation. MRCAGCFN combines the strengths of CNN and GCN while introducing enhancements to better accommodate diverse feature shapes. Experimental results on three HSI classification datasets demonstrate that our proposed MRCAGCFN outperforms some existing methods. The codes of our MRCAGCFN will be available at https://github.com/shengrunhua/MRCAGCFN.
Hongmin Gao 0001, Runhua Sheng, Zhonghao Chen, Haiyun Liu, Shufang Xu, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Airborne Small Target Detection Method Based on Multimodal and Adaptive Feature Fusion
abstract
The detection of airborne small targets amidst cluttered environments poses significant challenges. Factors such as the susceptibility of a single RGB image to interference from the environment in target detection and the difficulty of retaining small target information in detection necessitate the development of a new method to improve the accuracy and robustness of airborne small target detection. This article proposes a novel approach to achieve this goal by fusing RGB and infrared (IR) images, which is based on the existing fusion strategy with the addition of an attention mechanism. The proposed method employs the YOLO-SA network, which integrates a YOLO model optimized for the downsampling step with an enhanced image set. The fusion strategy employs an early fusion method to retain as much target information as possible for small target detection. To refine the feature extraction process, we introduce the self-adaptive characteristic aggregation fusion (SACAF) module, leveraging spatial and channel attention mechanisms synergistically to focus on crucial feature information. Adaptive weighting ensures effective enhancement of valid features while suppressing irrelevant ones. Experimental results indicate 1.8% and 3.5% improvements in mean average precision (mAP) over the LRAF-Net model and Infusion-Net detection network, respectively. Additionally, ablation studies validate the efficacy of the proposed algorithm’s network structure.
Shufang Xu, Tianci Liu 0007, Zhonghao Chen, Hongmin Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Dual-Feature Attention-Based Contrastive Prototypical Clustering for Multimodal Remote Sensing Data
abstract
The integrated use of multisource remote sensing (RS) data in Earth observation missions has garnered considerable attention. Hyperspectral images (HSIs) offer extensive spatial and spectral detail, whereas light detection and ranging (LiDAR) data provide elevation information. Therefore, the fusion of HSI and LiDAR data can enhance the accuracy (ACC) of image classification. However, contemporary supervised multimodal deep learning techniques depend heavily on extensive human-annotated training datasets. To address this challenge, we propose a contrastive prototypical clustering network enhanced with a dual-feature attention module. Specifically, two sets of enhanced modal views are constructed from the multimodal RS images for the subsequent contrastive learning. The proposed dual-feature attention module emphasizes channel and spatial attention separately for each modality, integrating both to adjust the feature representation across different channels and positions. By learning the importance weights of each channel and position, this module highlights the hierarchical structure and enhances the discriminative quality of the features. The learned features are utilized through an online clustering mechanism and a self-supervised training strategy that combines contrastive loss and cluster loss to achieve efficient and effective land cover classification. Extensive experiments on three widely used HSI and LiDAR datasets demonstrate that the proposed method outperforms current state-of-the-art approaches. The code for this method is openly available at:https://github.com/RogsDing/DFCPC.
Shufang Xu, Xinchen Ding, Zhen Zhang 0019, Hongmin Gao 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Cognitive Fusion of Graph Neural Network and Convolutional Neural Network for Enhanced Hyperspectral Target Detection
abstract
In recent years, deep learning has emerged as a prominent technique in hyperspectral target detection (HTD). Extensive research has highlighted the potential of Graph Neural Network (GNN) as a promising framework for exploring non-Euclidean dependencies within hyperspectral imagery. However, GNN has not been introduced to HTD. Additionally, achieving a balanced training set while effectively suppressing background remains a challenge. Therefore, we propose the cognitive fusion of GNN and Convolutional Neural Network (CNN) for enhanced HTD (named as CFGC), which marks the first integration of GNN and CNN in HTD. Initially, using sparse subspace clustering and a similarity measurement strategy, we select the most representative background samples for HTD. Subsequently, linear interpolation combines the prior target with the Laplacian-weighted prior target, yielding abundant targets with meaningful transformations. Finally, a fused network of CNN and GNN is utilized for training both the prior target and the constructed training set. Significantly, the incorporation of attention mechanism in both the CNN and GNN branches stands out as a noteworthy advantage, augmenting the models’ ability to selectively prioritize crucial information. Four benchmark hyperspectral images have been used in extensive experiments, and the results demonstrate that CFGC exhibits superior performance in HTD.
Shufang Xu, Sijie Geng, Zhonghao Chen, Hongmin Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Strengthened Residual Graph and Multiscale Gated Guided Convolutional Fusion Network for Hyperspectral Change Detection
abstract
Hyperspectral image (HSI) change detection (CD) focuses on identifying changes in the internal components of land cover and land use. Convolutional neural networks (CNNs) have made significant progress in HSI-CD. Concurrently, graph convolutional networks (GCNs) have gained considerable attention for their ability to utilize unlabeled data and explicitly exploit correlations between adjacent parcels. However, CNNs are constrained by fixed, small-size convolutional kernels, which severely limit their receptive field. On the other hand, GCNs use superpixels to reduce the number of nodes, which will lead to losing pixel-level features, resulting in partial feature representations from both networks. To leverage the strengths of both CNNs and GCNs, a model was proposed that incorporates two subnetworks: decomposed multiscale gated guided CNNs and strengthened residual graph convolution. The decomposed multiscale gated guided CNNs are designed to capture pixel-level features at various scales using different kernel sizes. A gated change information fusion (GCF) unit integrates these multiscale pixel-level features. Meanwhile, the strengthened residual graph convolution was used to aggregate change information, which can prevent node information from becoming homogeneous. Additionally, a feature fusion module (FFM) is employed to combine features from the two subnetworks. The proposed model effectively utilizes both multiscale convolution and graph features, facilitating the learning of multilevel contextual semantic features. The experimental results on three HSI datasets demonstrate that this model outperforms several state-of-the-art methods. The code is available athttps://github.com/zhangyiyan001/srgmgn.
Shufang Xu, Xiangfei Xia, Runhua Sheng, Hongmin Gao 0001, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Two-way Delayed Updates with Model Similarity in Communication-Efficient Federated Learning
abstract
The great achievement of IoT and the wide use of edge devices have brought explosive growth in data. The quality and scale of data determine the performances of machine learning models. Federated learning has attracted widespread attention for its ability to use isolated data and protect data privacy. Models can represent excellent generalization capabilities through federated training. However, the large number of devices and complex models involved in federated training exacerbate the communication costs and degrade the performance of the global model. Although existing approaches can reduce communication costs, they ignore the degradation of global model accuracy in a heterogeneous environment. To alleviate the huge communication costs in federated learning, this paper focuses on reducing upstream and downstream communication frequency while ensuring global model accuracy. We propose a Two-way Delayed Updates method with Model Similarity in Communication-Efficient Federated Learning (FedTDMS). FedTDMS employs personalized local computation to improve global model accuracy on heterogeneous data. Combining 10-cal update relevance check and global model compensation, FedTDMS reduces the communication frequency in Federated Learning. We conduct experiments on the MNIST-FL and CFAR-10-FL datasets. Results show that FedTDMS can greatly optimize communication efficiency while maintaining good global model accuracy.
Yingchi Mao, Jun Wu 0001, Lijuan Shen, Shufang Xu, Jie Wu 0001
MSN5
2023 FRDet: Few-shot object detection via feature reconstruction
abstract
Abstract State‐of‐the‐art object detection models rely on large‐scale datasets for training to achieve good precision. Without sufficient samples, the model can suffer from severe overfitting. Current explorations in few‐shot object detection are mainly divided into meta‐learning‐based methods and fine‐tuning‐based methods. However, existing models do not focus on how feature maps should be processed to present more accurate regions of interest (RoIs), leading to many non‐supporting RoIs. These non‐supporting RoIs can increase the burden of subsequent classification and even lead to misclassification. Additionally, catastrophic forgetting is inevitable in both few‐shot object detection models. Many models classify directly in low‐dimensional spaces due to insufficient resources, but this transformation of the data space can confuse some categories and lead to misclassification. To address these problems, the Feature Reconstruction Detector (FRDet) is proposed, a simple yet effective fine‐tune‐based approach for few‐shot object detection. FRDet includes a region proposal network (RPN) based on channel attention and space attention called Multi‐Attention RPN (MARPN) and a head based on feature reconstruction called Feature Reconstruction Head (FRHead). MARPN utilizes channel attention to suppress non‐supporting classes and spatial attention to enhance support classes based on Attention RPN, resulting in fewer but more accurate RoIs. Meanwhile, FRHead utilizes support features to reconstruct query RoI features through a closed‐form solution, allowing for a comprehensive and fine‐grained comparison. The model was validated on the PASCAL VOC, MS COCO, FSOD, and CUB200 datasets and achieved better results.
Yingchi Mao, Yong Qian, Zhenxiang Pan, Shufang Xu
IET Image Process.5
2023 Depthwise Separable Convolutional Autoencoders for Hyperspectral Image Change Detection
abstract
Hyperspectral image change detection (HSI-CD) has recently become a research hotspot. Current methods rely heavily on a huge amount of training samples to perform the change detection tasks. While acquiring data from the same region of bi-temporal HSIs is extraordinarily time-consuming and laborious. Therefore, this letter proposes an unsupervised method based on three dimensional (3D) depthwise separable convolutional autoencoders (DSConvAE). First, the dual-branch symmetrical 3D DSConvAE is pre-trained with limited samples to obtain the optimal weights, which facilitates extracting discriminative spatial and spectral features subsequently. Second, we adopt the temporal-specific feature concatenation strategy to acquire comprehensive characteristics from bi-temporal HSIs. Third, the general autoencoders are employed at the end of the model to further explore the high-level and abstract feature vectors. Finally, we compare the mean square loss calculated from the spatial-spectral branches and apply threshold judgement to generate the ultimate detection maps. Experimental results on three public HSI datasets demonstrate that the proposed framework outperforms other comparative methods by significant improvements.
Yongfeng Zhou, Shufang Xu, Danfeng Hong, Hongmin Gao 0001, Qiqiang Zhong, Bing Zhang 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 A Multidepth and Multibranch Network for Hyperspectral Target Detection Based on Band Selection
abstract
Deep learning (DL) has recently risen to prominence in hyperspectral target detection (HTD). Nevertheless, how to tackle the extreme training sample imbalance together with achieving target highlighting and background suppression is challenging. Additionally, due to the spectral redundancy of hyperspectral imagery (HSI), it is a new course for HTD through band selection (BS) to retain crucial bands thereupon improving the subsequent detection performance. Accordingly, we propose a DL-based BS-HTD (DLBSTD) algorithm, incorporating DL-based BS with DL-based HTD for the first time. Most significantly, a multi-depth and multi-branch network (MDBN) for HTD based on a novel BS method is proposed. First of all, the BS method including an alternating local-global reconstruction network (ALGRN) and a correlation measurement strategy provides representative bands containing key target information for MDBN. For the training sample imbalance of MDBN, we develop a BS-based method to select multifarious representative background training samples and propose a target band random substitution (TBRS) strategy to augment an ample target training set. Lastly, the MDBN composed of a multi-depth feature extraction (MDFE) module, three fusion strategies, and the parallel local convolution and gated recurrent unit (Conv-GRU) fully taps the spectral feature relationships to highlight targets and suppress backgrounds. Compared with nine competitive HTD algorithms, we carry out plentiful experiments on four classical datasets exhibiting that the proposed DLBSTD has strong generalization and salient detection performance of target highlighting and background suppression.
Hongmin Gao 0001, Zhonghao Chen, Shufang Xu, Danfeng Hong, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 AMSSE-Net: Adaptive Multiscale Spatial-Spectral Enhancement Network for Classification of Hyperspectral and LiDAR Data
abstract
With the abundant emergence of remote sensing data sources, multimodal remote sensing observation has become an active field. Extracting valuable information from multi-modal data has the potential to make a significant contribution to applications such as urban planning and monitoring. However, existing studies are deficient in extracting spectral and spatial features from hyperspectral remote sensing data. Meanwhile, the method of fusing multimodal features has limitations and poses a challenge to the convergence of the model loss function, which increases the complexity of the network model optimisation process. Therefore, this paper proposes an Adaptive Multi-scale Spatial–Spectral Enhancement Network for Classification of Hyperspectral and LiDAR Data called AMSSE-Net. First, we perform deep mining of spectral features in hyperspectral images by the involution operator. The main idea is to take full advantage of the involution operator in characterising spectral features by using the property that the convolution kernel shares the feature channels within the group. Furthermore, the multi-branching approach is used to extract the multi-scale information, and then the spectral-spatial features are formed with the strategy of hierarchical fusion. Meanwhile, we employ three-layer convolution for extracting shallow features from LiDAR data, offering supplementary information. Finally, we propose the ”Adaptive Feature Fusion Module,” an innovative and comprehensive mechanism designed for the fusion of features from diverse sources in multi-source data fusion. These dynamically assigned weights guide the selection of the optimal model, which is determined by the joint loss across the three methods, ultimately leading to the generation of an accurate prediction map. This approach not only helps to deeply explore the spectral spatial information in the hyperspectral data, but also effectively fuses the hyperspectral information with the elevation information from the LiDAR data. The expression ability of model features is rapidly improved by adaptive weighting, which in turn enhances the performance and generalisation ability of the model. Compared with some existing methods, extensive experiments on three popular HSI and LiDAR datasets show that our proposed AMSSE-Net can achieve better classification performance. The codes will be available at https://github.com/haofeng0003/AMSSE-Net, contributing to the RS community.
Hongmin Gao 0001, Shufang Xu, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Multimodal Transformer Network for Hyperspectral and LiDAR Classification
abstract
The land cover classification of single-modal remote sensing (RS) data has recently reached a bottleneck. The joint use of multi-modal RS data to improve classification performances has received much attention. Convolutional Neural Networks are powerful tools in feature extraction and contextual modeling. While they have attendant drawbacks to capture the sequence attributes of spectral signatures and struggle to acquire discriminative spectral-spatial features from a global perspective due to limitations inherent in their network backbones. The transformer backbone is a promising approach for addressing these challenges and generating novel insights in multi-modal RS image classification. In this article, we present a new model called Multi-modal Transformer Network (MTNet) that leverages transformer advantages to capture both the specific and shared characteristics of hyperspectral (HS) and light detection and ranging (LiDAR) data. HS images contain a wide range of bands with rich spectral information and LiDAR data provide accurate elevation information without affecting by environmental factors. The well-designed module Hyperspectral Spectral Transformer can learn spectrally local sequence information from neighbouring bands of HS images, yielding group-wise spectral embeddings comprising rich diagnostic information about land covers. Furthermore, the HS and LiDAR spatial transformers aim to mine the pixel-wise feature embedding relationships in a global manner, capturing spatial and elevation information of HS and LiDAR, respectively. Finally, the feature embedding tokens of two modalities are integrated jointly and a new transformer encoder is redesigned to explore the shared spatial characteristics between the two modalities. We evaluate the classification performances of the proposed MTNet on three public HS-LiDAR datasets by conducting extensive experiments, exhibiting superiority over conventional classifiers and state-of-the-art networks.
Shufang Xu, Danfeng Hong, Hongmin Gao 0001, Meiqiao Bi
IEEE Trans. Geosci. Remote. Sens.2
2014 Acoustic investigation of /th/ lenition in brunei Mandarin
abstract
This study investigates the acoustic characteristics of /th/ lenition in conversational speech of Brunei Mandarin, a variety of Mandarin Chinese. Based on data from 20 Chinese Bruneians, /th / lenition was found in the third-person pronoun tā /tha/, which is frequently pronounced as hā [ha]. Perceptual judgments, spectrographic analysis and acoustic measurements were conducted to examine the features of this sound change. In comparison with the perceptual judgments, it was found that the spectrographic inspection yielded 83.6 % correct classification of [th] and [h] for female speakers and 77.2 % for male speakers, indicating there is reasonably high reliability in identification in terms of spectral properties. Results of the acoustic measurements showed that there is an increase in high frequency intensity after the release of the closure for [th] while there is little change in intensity during the frication for [h]. The results showed that the lack of burst and little increase in intensity are reasonably reliable cues for stop lenition. Index Terms: Brunei Mandarin, /th / lenition, auditory judgment, spectrographic observation, intensity 1.
Shufang Xu
INTERSPEECH1
2007 A New Directional Weighted Median Filter for Removal of Random-Valued Impulse Noise
abstract
The known median-based denoising methods tend to work well for restoring the images corrupted by random-valued impulse noise with low noise level but poorly for highly corrupted images. This letter proposes a new impulse detector, which is based on the differences between the current pixel and its neighbors aligned with four main directions. Then, we combine it with the weighted median filter to get a new directional weighted median (DWM) filter. Extensive simulations show that the proposed filter not only can provide better performance of suppressing impulse with high noise level but can preserve more detail features, even thin lines. As extended to restoring corrupted color images, this filter also performs very well
Yiqiu Dong, Shufang Xu
IEEE Signal Process. Lett.2
2007 A Detection Statistic for Random-Valued Impulse Noise
abstract
This paper proposes an image statistic for detecting random-valued impulse noise. By this statistic, we can identify most of the noisy pixels in the corrupted images. Combining it with an edge-preserving regularization, we obtain a powerful two-stage method for denoising random-valued impulse noise, even for noise levels as high as 60%. Simulation results show that our method is significantly better than a number of existing techniques in terms of image restoration and noise detection.
Yiqiu Dong, Raymond Chan 0001, Shufang Xu
IEEE Trans. Image Process.3
2006 Matrix factorizations for reversible integer implementation of orthonormal M-band wavelet transforms
Tony Lin 0001, Pengwei Hao, Shufang Xu
Signal Process.3
2005 Factoring M-band wavelet transforms into reversible integer mappings and lifting steps
abstract
In this paper, a matrix factorization method is presented for reversible integer M-band wavelet transforms. Based on an algebraic construction of orthonormal M-band wavelets with perfect reconstruction, the polyphase matrix can be factorized into a finite sequence of elementary reversible matrices that map integers to integers reversibly. We show that the reversible integer mapping is essentially equivalent to the lifting scheme, thus we extend the classical lifting scheme to a more flexible framework.
Tony Lin 0001, Pengwei Hao, Shufang Xu
ICASSP (4)3