Miaozhong Xu

dblp:153/8779 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-8434-8009ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Identity Model Transformation for boosting performance and efficiency in object detection network
Zhongyuan Lu, Jin Liu 0029, Miaozhong Xu
Neural Networks3
2025 Enhancing Object Detection With Fourier Series
abstract
Traditional object detection models often lose the detailed outline information of the object. To address this problem, we propose the Fourier Series Object Detection (FSD). It encodes the object's outline closed curve into two one-dimensional periodic Fourier series. The Fourier Series Model (FSM) is constructed to regress the Fourier series for each object in the image. Thus, during inference, the detailed outline information of each object can be retrieved. We introduce Rolling Optimization Matching for Fourier loss to ensure that the model's learning process is not affected by the sequence of the starting points of the labeled contour points, speeding up the training process. The FSM demonstrates improved feature extraction and descriptive capabilities for non-rectangular or elongated object regions. The model achieves AP50 = 73.3% on the DOTA 1.5 dataset, which surpasses the state-of-the-art (SOTA) method by 6.44% at 66.86%. On the UCAS dataset, the model achieves AP50 = 97.25%, also surpassing the performance indicators of the SOTA methods. Furthermore, we introduce the object's Fourier power spectrum to describe outline features and the Fourier vector to indicate its direction. This enhances the scene semantic representation of the object detection model and paves a new pathway for the evolution of object detection methodologies.
Jin Liu 0029, Zhongyuan Lu, Yaorong Cen, Yong Hong, Miaozhong Xu
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 CFPNet: Coarse-to-Fine Progressive Network for Cloud Detection in Remote Sensing Images
abstract
Accurate cloud detection in remote sensing images constitutes a critical preprocessing requirement for ensuring the validity of subsequent analytical applications. Existing cloud detection methods effectively identify primary cloud regions but struggle with precise boundary delineation and distinguishing spectrally similar surfaces, particularly between clouds and terrain with comparable reflectance properties. To address these challenges, we propose the coarse-to-fine progressive network (CFPNet), a progressive detection framework integrating across channel and spatial dimensions. The framework incorporates two innovative components: a channel spatial attention residual fusion module (CSARF) and a multi-mask adaptive attention module (MMAA). The CSARF module achieves the network’s initial focus on clouds, while the MMAA module enables fine feature extraction of clouds. Specifically, The CSARF module employs channel attention (CAM) and spatial attention (SAM) mechanisms to suppress background noise while enhancing discriminative feature representation. MMAA employs muti-mask self-attention (MMSA) to compute channel self-attention and capture long-range dependencies, while using a multi-mask strategy to filter important channels. Deformable contextual feed-forward network (DCFN) then adaptively extracts cloud boundary features through deformable convolution, minimizing non-cloud pixel interference. Hence, the network enables coarse-to-fine feature extraction of clouds across both channel and spatial dimensions. Experimental results on the GF1-WFV, AIR-CD and Sentinel-2 datasets demonstrate that our method achieves superior performance in boundary detail accuracy and inter-class feature classification compared to other methods.
Hao Deng 0012, Mingjun Deng, Yonghua Jiang 0001, Miaozhong Xu, Yuexi Peng
IEEE Trans. Geosci. Remote. Sens.5
2023 Crisscross-Global Vision Transformers Model for Very High Resolution Aerial Image Semantic Segmentation
abstract
Semantic segmentation is a key means for understanding very-high resolution (VHR) aerial imagery. With the explosive development of deep learning, deep learning methods are being applied to the segmentation of VHR images, with convolutional neural networks (CNNs) as the basic framework. However, owing to the highly complex details present in VHR images and the high spatial dependence of geographical objects, CNN-based methods are inadequate. This is because the inherent locality of CNNs limits the size of the receptive field, thus limiting the ability to obtain long-range context information. To solve this problem, in this paper, we propose a transformer-based novel deep learning model called crisscross-global vision transformers (CGVT). CGVT exploits the transformer’s inherent ability to obtain long-range context information to solve the restricted receptive field problem. Specifically, we redesign the self-attention mechanism in the transformer and call it crisscross-global attention. It consists of two parts: crisscross transformer encoder block (CC-TEB) and global squeeze transformer encoder block (GS-TEB). CC-TEB overcomes the limitation of the traditional self-attention design (specifically, difficulty applying it to VHR aerial image segmentation) and further increases the local feature representation ability of the model. GS-TEB increases the global feature representation ability of the model. The results of experiments conducted on the popular ISPRS Vaihingen, IEEE GRSS Data Fusion Contest Zeebrugge, and LoveDA Semantic Segmentation Challenge datasets verify the effectiveness and superiority of our proposed method. Specifically, it achieved state-of-the-art performance on both Zeebrugge and LoveDA datasets, and is currently ranked second in Vaihingen dataset.
Guohui Deng, Zhaocong Wu, Miaozhong Xu, Zhiye Wang, Zhongyuan Lu
IEEE Trans. Geosci. Remote. Sens.3
2022 Cross-Modality Image Matching Network With Modality-Invariant Feature Representation for Airborne-Ground Thermal Infrared and Visible Datasets
abstract
Thermal infrared (TIR) remote-sensing imagery can allow objects to be imaged clearly at night through the long-wave infrared, so that the fusion of thermal infrared and visible (VIS) imagery is a way to improve the remote-sensing interpretation ability. However, due to the large radiation difference between the two kinds of images, it is very difficult to match them. One of the most important issues is the lack of comprehensive consideration of the modality-specific information and modality-shared information, which makes it difficult for the existing methods to obtain a modality-invariant feature representation. In this article, a cross-modality image matching network, which we refer to as CMM-Net, is proposed to realize thermal infrared and visible image matching by learning a modality-invariant feature representation. First, in order to extract the modality-specific features of the imagery, the framework constructs a shallow two-branch network to make full use of the modality-specific information, without sharing parameters. Second, in order to extract the high-level semantic information between the different modalities, modality-shared layers are embedded into the deep layers of the network. In addition, three novel loss functions are designed and combined to learn the modality-invariant feature representation, that is, the discriminative loss of the non-corresponding features in the same modality, the cross-modality loss of the corresponding features between different modalities, and the cross-modality triplet (CMT) loss. The multimodal matching experiments conducted with ground- and airborne-based thermal infrared images and visible images showed that the proposed method outperforms the existing image matching methods by about 2% and 6% for the ground and airborne images, respectively.
Ailong Ma, Yuting Wan, Yanfei Zhong, Bin Luo 0005, Miaozhong Xu
IEEE Trans. Geosci. Remote. Sens.6
2022 MAP-Net: SAR and Optical Image Matching via Image-Based Convolutional Network With Attention Mechanism and Spatial Pyramid Aggregated Pooling
abstract
The complementarity of synthetic aperture radar (SAR) and optical images allows remote sensing observations to “see” unprecedented discoveries. Image matching plays a fundamental role in the fusion and application of SAR and optical images. However, both the geometric imaging pattern and the physical radiation mechanism of these two sensors are significantly different, so that the images show complex geometric distortion and nonlinear radiation differences. This phenomenon brings great challenges to image matching, which neither the handcrafted descriptors nor the deep learning-based methods have adequately addressed. In this article, a novel image-based matching method for SAR to optical images via an image-based convolutional network with spatial pyramid aggregated pooling (SPAP) and an attention mechanism is proposed, namely MAP-Net. The original image is embedded through the convolutional neural network to generate the feature map. Through the information extraction and abstraction of the original imagery, the embedded features containing the high-level semantic information are more robust to the geometric distortion and radiation variation among the different modal images, which is beneficial to the matching of cross-modal images. The adoption of the SPAP module makes the network more capable of integrating global and local contextual information. The attention block weights the dense features generated from the network to extract the key features that are invariant, distinguishable, repeatable, and suitable for the image matching task. In the experiments, five sets of multisource and multiresolution SAR and optical images with wide and varied ground coverage were used to evaluate the accuracy of MAP-Net, compared to both handcrafted and deep learning-based methods. The experimental results show that the MAP-Net method is superior to the current state-of-the-art image matching methods for SAR to optical images.
Ailong Ma, Liangpei Zhang 0001, Miaozhong Xu, Yanfei Zhong
IEEE Trans. Geosci. Remote. Sens.4
2022 CCANet: Class-Constraint Coarse-to-Fine Attentional Deep Network for Subdecimeter Aerial Image Semantic Segmentation
abstract
Semantic segmentation is important for the understanding of subdecimeter aerial images. In recent years, deep convolutional neural networks (DCNNs) have been used widely for semantic segmentation in the field of remote sensing. However, because of the highly complex subdecimeter resolution of aerial images, inseparability often occurs among some geographic entities of interest in the spectral domain. In addition, the semantic segmentation methods based on DCNNs mostly obtain context information using extra information within the added receptive field. However, the context information obtained this way is not explicit. We propose a novel class-constraint coarse-to-fine attentional (CCA) deep network, which enables the formation of class information constraints to obtain explicit long-range context information. Further, the performance of subdecimeter aerial image semantic segmentation can be improved, particularly for fine-structured geographic entities. Based on coarse-to-fine technology, we obtained a coarse segmentation result and constructed an image class feature library. We propose the use of the attention mechanism to obtain strong class-constrained features. Consequently, pixels of different geographic entities can adaptively match the corresponding categories in the class feature library. Additionally, we employed a novel loss function, CCA-loss to realize end-to-end training. The experimental results obtained using two popular open benchmarks, International Society for Photogrammetry and Remote Sensing (ISPRS) 2-D semantic labeling Vaihingen data set and Institute of Electrical and Electronics Engineers (IEEE) Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest Zeebrugge data set, validated the effectiveness and superiority of our proposed model. The proposed method achieved state-of-the-art performance on the IEEE GRSS Data Fusion Contest Zeebrugge data set.
Guohui Deng, Zhaocong Wu, Miaozhong Xu, Yanfei Zhong
IEEE Trans. Geosci. Remote. Sens.4
2022 Translution-SNet: A Semisupervised Hyperspectral Image Stripe Noise Removal Based on Transformer and CNN
abstract
Hyperspectral remote sensing images (HSIs) have been applied in urban planning, environmental monitoring, and other fields. However, they are susceptible to noise interference, such as Gaussian noise, stripe, and mixed noises, from various factors in the imaging process, which greatly limits their applications. Although previous efforts to improve HSI quality have achieved remarkable results, there are still many challenges to be solved. To avoid the poor generalization ability and improve the stripe removal performance of the network in real scenarios. In this paper, we proposed a novel deep learning model (Translution-SNet) for HSI stripe noise removal based on a semi-supervised training strategy that applies a convolution and transformer for feature extraction. Moreover, we used an unbiased estimation method to calculate the loss function of the unsupervised part from noisy data without a clean image. The semi-supervised method improved the ability of Translution-SNet to deal with various complex stripe noises during stripe removal and strengthened its robustness and generalization ability. Our experimental results showed that Translution-SNet could robustly handle stripe noise of images with different loads and achieve satisfactory results, proving its feasibility and effectiveness. In addition, Translution-SNet showed good generalization ability.
Miaozhong Xu, Yonghua Jiang 0001, Guo Zhang 0001, Hao Cui 0002, Litao Li
IEEE Trans. Geosci. Remote. Sens.2
2022 Hyperspectral Image Stripe Removal Network With Cross-Frequency Feature Interaction
abstract
Remote sensing images, especially hyperspectral images (HSIs), are extremely vulnerable to random noise and stripe noise. As a key aspect of HSI data quality improvement, stripe noise removal has always been a pervasive issue in remote sensing image processing. Convolutional neural networks have been applied for HSI data destriping. However, the existing methods lose the stripe-free component of the original image to a certain extent. These models also ignore the global spatial context of images and the correlation between spatial information and spectral information. Therefore, we propose a novel destriping convolutional network to overcome the problems with the existing methods. Octave convolution is used to extract cross-frequency features, and separate and compress the low-frequency information of the images, while dilation convolution (Dila-Conv) is used to reduce the amount of required calculation and also preserve the key image information. In addition, Dila-Conv can expand the receptive field to obtain multiscale features. Finally, a cross-channel enhanced spatial–spectral feature fusion module is used to acquire and integrate spatial context information and interchannel dependencies on a global scale as auxiliary information so that the network model can learn and pay attention to key feature information, specifically, “what to look for” and “where to look at,” which can facilitate the distinction between stripe and stripe-free components. Experimental results obtained using multiple datasets demonstrated that the proposed method can outperform the existing comparable methods and can produce satisfactory results in terms of visual effects and quantitative evaluation.
Miaozhong Xu, Yonghua Jiang 0001, Guohui Deng, Zhongyuan Lu, Guo Zhang 0001, Hao Cui 0002
IEEE Trans. Geosci. Remote. Sens.2
2019 A Multi-Satellite Regional Imaging Mission Planning Method Based on Moom for Emergency Surveying and Mapping
abstract
Aiming at the imaging problem of regional target in emergency surveying and mapping, a multi-objective optimization model(MOOM) is proposed, which takes the imaging lateral swing angles of satellite as decision variables and takes the maximum coverage rate of regional target and the minimum number of holes as objective functions. Aiming at the two key problems of evaluation function calculation and multi-objective model solving, Vatti algorithm and NSGAII algorithm are used to solve them respectively. Finally, STK simulation data are used to verify the feasibility of the optimization method.
Yaxin Chen, Xin Shen 0001, Shixue Li, Guo Zhang 0001, Miaozhong Xu, Junfei Xu
IGARSS5
2017 Unsupervised-Restricted Deconvolutional Neural Network for Very High Resolution Remote-Sensing Image Classification
abstract
As the acquisition of very high resolution (VHR) satellite images becomes easier owing to technological advancements, ever more stringent requirements are being imposed on automatic image interpretation. Moreover, per-pixel classification has become the focus of research interests in this regard. However, the efficient and effective processing and the interpretation of VHR satellite images remain a critical task. Convolutional neural networks (CNNs) have recently been applied to VHR satellite images with considerable success. However, the prevalent CNN models accept input data of fixed sizes and train the classifier using features extracted directly from the convolutional stages or the fully connected layers, which cannot yield pixel-to-pixel classifications. Moreover, training a CNN model requires large amounts of labeled reference data. These are challenging to obtain because per-pixel labeled VHR satellite images are not open access. In this paper, we propose a framework called the unsupervised-restricted deconvolutional neural network (URDNN). It can solve these problems by learning an end-to-end and pixel-to-pixel classification and handling a VHR classification using a fully convolutional network and a small number of labeled pixels. In URDNN, supervised learning is always under the restriction of unsupervised learning, which serves to constrain and aid supervised training in learning more generalized and abstract feature. To some degree, it will try to reduce the problems of overfitting and undertraining, which arise from the scarcity of labeled training data, and to gain better classification results using fewer training samples. It improves the generality of the classification model. We tested the proposed URDNN on images from the Geoeye and Quickbird sensors and obtained satisfactory results with the highest overall accuracy (OA) achieved as 0.977 and 0.989, respectively. Experiments showed that the combined effects of additional kernels and stages may have produced better results, and two-stage URDNN consistently produced a more stable result. We compared URDNN with four methods and found that with a small ratio of selected labeled data items, it yielded the highest and most stable results, whereas the accuracy values of the other methods quickly decreased. For some categories with fewer training pixels, accuracy for categories from other methods was considerably worse than that in URDNN, with the largest difference reaching almost 10%. Hence, the proposed URDNN can successfully handle the VHR image classification using a small number of labeled pixels. Furthermore, it is more effective than state-of-the-art methods.
Yiting Tao, Miaozhong Xu, Fan Zhang 0006, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Weakly Supervised Learning Based on Coupled Convolutional Neural Networks for Aircraft Detection
abstract
Aircraft detection from very high resolution (VHR) remote sensing images has been drawing increasing interest in recent years due to the successful civil and military applications. However, several challenges still exist: 1) extracting the high-level features and the hierarchical feature representations of the objects is difficult; 2) manual annotation of the objects in large image sets is generally expensive and sometimes unreliable; and 3) locating objects within such a large image is difficult and time consuming. In this paper, we propose a weakly supervised learning framework based on coupled convolutional neural networks (CNNs) for aircraft detection, which can simultaneously solve these problems. We first develop a CNN-based method to extract the high-level features and the hierarchical feature representations of the objects. We then employ an iterative weakly supervised learning framework to automatically mine and augment the training data set from the original image. We propose a coupled CNN method, which combines a candidate region proposal network and a localization network to extract the proposals and simultaneously locate the aircraft, which is more efficient and accurate, even in large-scale VHR images. In the experiments, the proposed method was applied to three challenging high-resolution data sets: the Sydney International Airport data set, the Tokyo Haneda Airport data set, and the Berlin Tegel Airport data set. The extensive experimental results confirm that the proposed method can achieve a higher detection accuracy than the other methods.
Fan Zhang 0006, Bo Du 0001, Liangpei Zhang 0001, Miaozhong Xu
IEEE Trans. Geosci. Remote. Sens.4
2014 Estimation of CODMn in Tai lake basin using Landsat-8 satellite
abstract
CODMnis employed as a kind of water quality parameter indicating eutrophication. This paper focuses on the estimation of CODMnin Tai lake basin, China, and has proposed an Advanced CODMnForecast Index (ACFI) with 3 sub-indices X1, X2and X3, based on empirical and semi-empirical methods. Landsat-8 has been involved in the research for the purpose of exploring its band features in water quality monitoring. Comparing in-situ data and Landsat-8 image on July 19th 2013, we selected band 8 for X1, combination of band 1 and band 7 for X2and the ratio of band 4 to band3 for X3. ACFI model is proved to work well on the estimation of CODMnwith the smallest average absolute relative error 7.8% and the biggest average absolute relative error 17.1% from June and March respectively using the fitting coefficient of July. What is more, the research also shows that the ACFI model is suitable when CODMnis in the range of 3-5mg/L.
Yiting Tao, Miaozhong Xu
IGARSS2