VLDB 2026 Research / reviewers in the wild / expert
Xiaoliang Meng
dblp:31/7844
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Cross-Temporal Difference: Style-Aligned and Fusion-Difference Learning for Change DetectionabstractAccurate change detection in bi-temporal remote sensing images remains challenging due to style differences and multi-scale target differences. Existing methods relying on cross-temporal feature comparison often suffer from style-induced false alarms. Our research proposes a Style-Aligned and Fusion-Difference Network (SAFDNet) which can distinguish style differences from genuine changes. In the encoding stage, we designed a style alignment module based on the idea of histogram equalization to normalize feature distributions of bi-temporal images after feature extraction. When calculating the difference between bi-temporal images, we optimized the existing methods from two perspectives: the object and the scale. For objects, we make full use of layer-exchange fusion mechanism: features are exchanged between temporal streams, fused within each stream, and then compared to their original versions to capture changes while mitigating style interference. For scale, we use a multi-scale Structural Similarity (SSIM) metric replaces conventional distance measures to better quantify structural differences across varying object sizes. With this design, we effectively avoid the noise effect caused by directly using the cross-temporal feature maps to calculate the changes. Our model was tested on five remote sensing datasets: WHUCD, LEVIR-CD, SYSU-CD, CDD and PX-CLCD, and all showed excellent performance. The code and pre-trained model are available at https://github.com/Vgrant0/SAFDNet.git. Siming Fu, Sijun Dong, Xiaoliang Meng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multi-Scale Adversarial Cross-Domain Vehicle Detection from UAV to Satellite ImageryabstractVehicle detection based on satellite imagery has extensive applications in the field of intelligent transportation. However, current algorithms generally rely on expensive data annotation, leading to increased cost burdens. To address this challenge, our paper proposes a domain-adaptive multi-scale adversarial cross-domain vehicle detection network, enabling the transfer of knowledge learned from annotated and easily accessible UAV imagery to the unlabeled satellite imagery domain. Specifically, our network, built upon the two-stage object detection framework, achieves domain alignment at the image-level by embedding a multi-scale information aggregation module between the source domain (UAV imagery) and the target domain (satellite imagery) to handle data differences at various scales effectively. To tackle the sample misalignment issue, an adversarial gradient reversal layer is designed to adversarially mine samples that are challenging to align. The integration of domain adaptation methods further enhances the performance of unsupervised cross-domain vehicle detection. Experimental results demonstrate a significant improvement in the VisDrone → UCAS-AOD cross-domain vehicle detection scenario with substantial domain scale differences. Compared to non-domain adaptation methods, our proposed approach achieves a notable 15.15% increase in AP. Xiaoliang Meng |
IGARSS | 5 |
| 2024 | Towards Robust Semantic Segmentation for Remote Sensing Images Under Adversarial Noise ConditionabstractRemote sensing image analysis tasks perform exceptionally well on clear, noise-free remote sensing images. However, their accuracy significantly declines when faced with adversarial conditions, such as noise and interference. To address this challenge, we propose the MDWT-Transformer, which is a transformer-based autoencoder that is enhanced with multi-level discrete wavelet transformations and nonlinear activation free transformer blocks. Our approach uses multi-level discrete wavelet transformations to gradually decrease the size of the image. Then, the image is gradually restored through its inverse transformation. This structure is beneficial for capturing multi-scale information in images. Additionally, the nonlinear activation free transformer block offers a dual advantage by combining lower computational costs with preserving a significant portion of details in the image. Experimental results on the ISPRS Vaihingen and Potsdam datasets demonstrate the MDWT-Transformer’s superiority in enhancing image quality and mitigating adversarial interference compared to advanced autoencoder benchmarks. Yuechi Yang, Xiaoliang Meng |
IGARSS | 6 |
| 2024 | Label-Guided Cross-Modal Attention Network for Multi-Label Aerial Image ClassificationabstractMulti-label aerial image classification is a fundamental yet complex task in remote sensing interpretation that aims to identify multiple labels in a single image. In this letter, we propose a Label-Guided Cross-Modal Attention (L-GCMA) network, which first introduces a novel approach to enriching the semantic information of labels and utilizes the multi-head attention module to extract diverse features. The proposed method consists of two components before the cross-modal attention. Firstly, the visual features of the image are obtained using a transformer encoder. Additionally, to capture the rich semantic relationship of the scene, we design a Label-Sentence Mapping Attention (L-SMA) module. This module performs word embedding encoding on the labels and applies BERT encoding on the sentence prompts, followed by multi-head attention to extract comprehensive inter- and intra-class relationships for the labels, specifically obtaining label-scene text features. Subsequently, by treating the text features as a query, the visual features and text features are combined using cross-modal attention. This progressive integration narrows the semantic gap between vision and text, facilitating accurate label recognition. Experiments on the UCM and AID multi-label datasets demonstrate the superior performance of our L-GCMA, surpassing state-of-the-art methods with mAP scores of 99.10% (UCM) and 85.96% (AID). Ying Chen 0035, Xiaoliang Meng, Mianxin Gao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | EfficientCD: A New Strategy for Change Detection Based With Bi-Temporal Layers ExchangedabstractWith the widespread application of remote sensing (RS) technology in environmental monitoring, the demand for efficient and accurate RS image change detection (CD) for natural environments is growing. We propose a novel deep learning framework named EfficientCD, specifically designed for RS image CD. The framework employs EfficientNet as its backbone network for feature extraction. To enhance the information exchange between bi-temporal image feature maps, we have designed a new feature pyramid network module targeted at RS CD, named ChangeFPN. In addition, to make full use of the multilevel feature maps in the decoding stage, we have developed a layer-by-layer feature upsampling module combined with Euclidean distance to improve feature fusion and reconstruction during the decoding stage. The EfficientCD has been experimentally validated on four RS datasets: LEVIR-CD, SYSU-CD, CLCD, and WHUCD. The experimental results demonstrate that EfficientCD exhibits outstanding performance in CD accuracy. The code and pretrained models will be released athttps://github.com/dyzy41/mmrscd. Sijun Dong, Yuwei Zhu, Xiaoliang Meng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MetaSegNet: Metadata-Collaborative Vision-Language Representation Learning for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images plays a vital role in a wide range of Earth Observation applications, such as land-use land-cover (LULC) mapping, environment monitoring, and sustainable development. Driven by rapid developments in artificial intelligence, deep learning (DL) has emerged as the mainstream for semantic segmentation and has achieved many breakthroughs in the field of remote sensing. However, most DL-based methods focus on unimodal visual data while ignoring rich multimodal information involved in the real world. Nonvisual data, such as text, can gather extra knowledge from the real world, which can strengthen the interpretability, reliability, and generalization of visual models. Inspired by this, we propose a novel metadata-collaborative segmentation network (MetaSegNet) that applies vision-language representation learning for the semantic segmentation of remote sensing images. Unlike the common model structure that only uses unimodal visual data, we extract the key characteristic (e.g., the climate zone) from freely available remote sensing image metadata and transfer it into geographic text prompts via the generic ChatGPT. Then, we construct an image encoder, a text encoder, and a crossmodal attention fusion subnetwork to extract the image and text feature and apply image-text interaction. Benefiting from such a design, the proposed MetaSegNet not only demonstrates superior generalization in zero-shot testing but also achieves competitive accuracy with the state-of-the-art semantic segmentation methods on the large-scale OpenEarthMap dataset [70.4% mean intersection over union (mIoU)] and the Potsdam dataset (93.3% mean${F}1$score) as well as the LoveDA dataset (52.0% mIoU). Sijun Dong, Xiaoliang Meng, Shenghui Fang, Songlin Fei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Speed Monitoring of Heavy Vehicles on Construction Plants by Fusing Camera Visual Image with UAV LiDAR Point CloudabstractDriving heavy vehicles at high speeds through construction sections can pose significant dangers. Implementing a ground speed radar system during construction progress is challenging due to the high costs associated with installation and maintenance. This paper proposes an innovative and cost-effective method for monitoring vehicle speed on construction sites by combining 2D/3D data. This method involves collecting 3D point cloud data using UAV LiDAR and capturing 2D visual image data using IoT cameras placed around construction sections. The EPnP algorithm is employed to calculate the matching results of pixel coordinates and world coordinates using the 2D and 3D data. Subsequently, the YOLOv5 object detection algorithm is combined with the DeepSORT tracking algorithm to enable real-time detection and tracking of heavy vehicles. Lastly, vehicle speed monitoring is conducted by analyzing the 2D/3D coordinate matching and the results of vehicle detection tracking. Pertinent experiments conducted in real scenarios validate that the method achieves a maximum error of 7.24% and an average relative error of 4.16%, satisfying the requirements for real-time monitoring of heavy vehicle speeds on construction sites. Zehua Wu, Xiaoliang Meng, Jieyan Sun, Tinghao Li |
IGARSS | 3 |
| 2022 | Class-Guided Swin Transformer for Semantic Segmentation of Remote Sensing ImageryabstractSemantic segmentation of remote sensing images plays a crucial role in a wide variety of practical applications, including land cover mapping, environmental protection, and economic assessment. In the last decade, convolutional neural network (CNN) is the mainstream deep learning-based method of semantic segmentation. Compared with conventional methods, CNN-based methods learn semantic features automatically, thereby achieving strong representation capability. However, the local receptive field of the convolution operation limits CNN-based methods from capturing long-range dependencies. In contrast, Vision Transformer (ViT) demonstrates its great potential in modeling long-range dependencies and obtains superior results in semantic segmentation. Inspired by this, in this letter, we propose a class-guided Swin Transformer (CG-Swin) for semantic segmentation of remote sensing images. Specifically, we adopt a Transformer-based encoder–decoder structure, which introduces the Swin Transformer backbone as the encoder and designs a class-guided Transformer block to construct the decoder. The experimental results on ISPRS Vaihingen and Potsdam datasets demonstrate the significant breakthrough of the proposed method over ten benchmarks, outperforming both advanced CNN-based and recent Transformer-based approaches. Xiaoliang Meng, Yuechi Yang, Rui Li 0036, Ce Zhang 0005 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | A Novel Transformer Based Semantic Segmentation Scheme for Fine-Resolution Remote Sensing ImagesabstractThe fully convolutional network (FCN) with an encoder-decoder architecture has been the standard paradigm for semantic segmentation. The encoder-decoder architecture utilizes an encoder to capture multilevel feature maps, which are incorporated into the final prediction by a decoder. As the context is crucial for precise segmentation, tremendous effort has been made to extract such information in an intelligent fashion, including employing dilated/atrous convolutions or inserting attention modules. However, these endeavors are all based on the FCN architecture with ResNet or other backbones, which cannot fully exploit the context from the theoretical concept. By contrast, we introduce the Swin Transformer as the backbone to extract the context information and design a novel decoder of densely connected feature aggregation module (DCFAM) to restore the resolution and produce the segmentation map. The experimental results on two remotely sensed semantic segmentation datasets demonstrate the effectiveness of the proposed scheme. Rui Li 0036, Chenxi Duan, Ce Zhang 0005, Xiaoliang Meng, Shenghui Fang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Building Extraction With Vision TransformerabstractAs an important carrier of human productive activities, the extraction of buildings is not only essential for urban dynamic monitoring but also necessary for suburban construction inspection. Nowadays, accurate building extraction from remote sensing images remains a challenge due to the complex background and diverse appearances of buildings. The convolutional neural network (CNN) based building extraction methods, although increased the accuracy significantly, are criticized for their inability for modelling global dependencies. Thus, this paper applies the Vision Transformer for building extraction. However, the actual utilization of the Vision Transformer often comes with two limitations. First, the Vision Transformer requires more GPU memory and computational costs compared to CNNs. This limitation is further magnified when encountering large-sized inputs like fine-resolution remote sensing images. Second, spatial details are not sufficiently preserved during the feature extraction of the Vision Transformer, resulting in the inability for fine-grained building segmentation. To handle these issues, we propose a novel Vision Transformer (BuildFormer), with a dual-path structure. Specifically, we design a spatial-detailed context path to encode rich spatial details and a global context path to capture global dependencies. Besides, we develop a window-based linear multi-head self-attention to make the complexity of the multi-head self-attention linear with the window size, which strengthens the global context extraction by using large windows and greatly improves the potential of the Vision Transformer in processing large-sized remote sensing images. The proposed method yields state-of-the-art performance (75.74% IoU) on the Massachusetts building dataset. Code will be available at https://github.com/WangLibo1995/BuildFormer. Shenghui Fang, Xiaoliang Meng, Rui Li 0036 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Linear Regression Algorithm against Device Diversity for the WLAN Indoor Localization SystemabstractRecent years have witnessed a growing interest in using WLAN fingerprint‐based methods for the indoor localization system because of their cost‐effectiveness and availability compared to other localization systems. In this system, the received signal strength (RSS) values are measured as the fingerprint from the access points (AP) at each reference point (RP) in the offline phase. However, signal strength variations across diverse devices become a major problem in this system, especially in the crowdsourcing‐based localization system. In this paper, the device diversity problem and the adverse effects caused by this problem are analyzed firstly. Then, the intrinsic relationship between different RSS values collected by different devices is mined by the linear regression (LR) algorithm. Based on the analysis, the LR algorithm is proposed to create a unique radio map in the offline phase and precisely estimate the user’s location in the online phase. After applying the LR algorithm in the crowdsourcing systems, the device diversity problem is solved effectively. Finally, we verify the LR algorithm using the theoretical study of the probability of error detection. Experimental results in a typical office building show that the proposed method results in a higher reliability and localization accuracy. Liye Zhang 0001, Xiaoliang Meng, Chao Fang 0003 |
Wirel. Commun. Mob. Comput. | 2 |
| 2016 | A grid-based reliable routing protocol for wireless sensor networks with randomly distributed clusters
Xiaoliang Meng, Xiaochuan Shi |
Ad Hoc Networks | 1 |