VLDB 2026 Research / reviewers in the wild / expert
Nan Xu 0008
dblp:26/4304-8
· DBLP profile ↗
16ranked-venue papers
2as first author
16since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | GeoMoE: Geometry-Driven prompts with adaptive Mixture-of-Experts for remote sensing image semantic segmentation
Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You |
Expert Syst. Appl. | 5 |
| 2026 | Cos-UMamba: Optimizing salient object detection with cosine scanning and bias-corrected feature fusion in optical remote sensing images
Zhen Wang 0020, Fu-Lin He, Nan Xu 0008, Zhu-Hong You |
Expert Syst. Appl. | 3 |
| 2025 | Fine-Tuning SAM for Forward-Looking Sonar With Collaborative Prompts and EmbeddingabstractThe Segment Anything Model (SAM) represents a significant advancement in semantic segmentation, particularly for natural images, but encounters notable limitations when applied to forward-looking sonar (FLS) images. The primary challenges lie in the inherent boundary ambiguity of FLS images, which complicates the use of prompt strategies for accurate boundary delineation, and the lack of effective interaction between prompts and image features. In this letter, we introduce a collaborative prompting strategy to address these issues by generating dense prompt embeddings and sonar tokens that focus on contour and boundary features, thereby replacing the original dense prompt embedding and IoU token. To further enhance segmentation, we employ embedding compensation techniques based on Mamba and KAN, which increase boundary information to image embedings and improve the fusion of prompts within image embeddings. We conducted comprehensive experiments, including comparative analyses and ablation studies, to validate the superiority of our proposed approach. Results show that our method significantly improves segmentation performance for FLS images, effectively addressing boundary ambiguity and optimizing prompt utilization. The source code and dataset will be available on https://github.com/darkseid-arch/FLSSAM. Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Preliminary Analysis of Jitter Detection and Causes on the Ziyuan3 (ZY3) Satellite Series PlatformsabstractJitter is a critical factor affecting the geometric accuracy of Earth observation satellites. After a satellite is launched, it is essential to detect and mitigate platform jitter to minimize its impact. The Ziyuan-3 (ZY3) series, China’s first civilian stereo mapping satellite constellation, meets the 1:50000 scale mapping accuracy requirement without ground control points. Detecting and suppressing jitter during the satellite’s orbit is a key technology for ensuring mapping accuracy. This letter investigates the causes of platform jitter across three satellites in the ZY3 series by analyzing gyroscope data using fast Fourier transform (FFT) at different stages of the satellite lifecycle. It compares the frequency and amplitude of jitter across different platforms and time periods. The results show that ZY3-01 and ZY3-02 exhibit consistent jitter patterns. They have significant amplitude in the 0.6 Hz region early in orbit and less later on, correlating with the satellites’ lifespans. In contrast, the 0.2 Hz region remains relatively stable. The ZY3-03 satellite, equipped with a highly stable platform, shows no significant jitter, evidencing the effectiveness of such platforms in jitter suppression. Fan Mo 0001, Xinming Tang, Yingbao Yang, Junfeng Xie 0001, Nan Xu 0008, Haoran Xia |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Ranging Bias Correction of Fully Saturated Data Over Waters for ICESat-2 Photon-Counting LidarabstractThe recent capabilities of the Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) photon counting lidar in monitoring water levels have been demonstrated through its precise elevation measurement and small footprint. The accuracy in water level measurements is, however, significantly impacted by the first photon bias, especially when photon-counting detectors are fully saturated due to particular reflections from calm water surfaces. In this study, we propose an analytical model to correct the first photon bias in scenarios where the ICESat-2/Advanced Topographic Altimeter System (ATLAS) is fully saturated. Notably, the model innovatively recovers and estimates the required signal level using after-pulses, which are typically considered as noise. These after-pulses can be used to effectively estimate the signal level when the detector is fully saturated. The experiment analysis, conducted on eight ICESat-2 ground tracks over calm water surfaces near the Great Lakes and the Tibetan Plateau, indicates that the actual received signal photons can surpass 200 and in some cases, reach up to 600 counts for strong beams, introducing a first photon bias exceeding 15 cm. The findings prove that 1) after-pulses can be used to retrieve water surface elevation and reflectance when the primary surface return is distorted by detector saturation and 2) calm waters reflect 5–40 times more than ice and snow surfaces, where first photon bias is a predominant error in water level measurements. The method holds great significance for the accurate monitoring of water levels in small inland water bodies using ICESat-2 and may also inform the design of lidar systems for inland water observations. Yuanfei Gu, Jian Yang 0033, Yue Ma 0002, Yao Li 0027, Nan Xu 0008, Xiaohua Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Vision Foundation Model-Driven Multiscale Expert Tuning for Multimodal Remote Sensing Semantic SegmentationabstractMultimodal remote sensing semantic segmentation based on Optical and Digital Surface Model (Opt-DSM) data is pivotal for comprehensive scene interpretation. However, prevailing methodologies often lack a unified vision foundation model and encounter significant challenges in bridging modality gaps and achieving effective feature fusion. Conventional models, such as the Segment Anything Model (SAM), exhibit inherent limitations when addressing the unique complexities of multimodal remote sensing, particularly in managing cross-modal discrepancies and intricate surface structures. In this study, we present VF-MET (Vision Foundation Model-Driven Multi-Scale Expert Tuning), an innovative framework meticulously tailored for Opt-DSM semantic segmentation tasks. VF-MET incorporates an adaptive Multi-Scale Expert Tuning (AMET) strategy, which substantially enhances the feature extraction capabilities of vision foundation models. This enables the robust capture of cross-scale and morphologically irregular objects, while simultaneously preserving superior generalization ability. To further address the segmentation of densely distributed and weakly correlated regions, we propose a collaborative Box-Point Prompt Mechanism (CBPM), which significantly improves spatial localization and contextual discrimination. Moreover, we introduce a Two-Stage Mask Decoder (TSMD) that facilitates efficient multimodal feature fusion and augments contextual understanding. Extensive experiments conducted on public Opt-DSM benchmark datasets unequivocally demonstrate that VF-MET achieves state-of-the-art performance. Comprehensive ablation studies further substantiate the indispensable contributions of each constituent module within the proposed architecture. The source code and datasets are publicly accessible at https://github.com/NWPUFranklee/VF-MET.git. Zhen Wang 0020, Nan Xu 0008, Zhu-Hong You, De-Shuang Huang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | A Denoising Algorithm for ICESat-2 Bathymetric Photons Based on Point Cloud GriddingabstractIce, Cloud, and Land Elevation Satellite-2 (ICESat-2), utilizing a 532 nm green laser, provides critical data for high-precision depth measurements in shallow water areas. However, due to the high sensitivity of photon detection to solar radiation noise and complex underwater topography, the data contain significant noise, making precise extraction of underwater terrain a major challenge. Vertical segmentation methods are ineffective in accurately capturing steep terrain features. This article proposes a denoising algorithm based on point cloud gridding, which converts discrete photon data into a raster grid, emphasizing the spatial distribution of the water surface and underwater terrain. Through multiscale density-adaptive processing, the algorithm effectively separates photons above, at, and below the water surface, accurately locating underwater terrain points, making it particularly suitable for photon data analysis in complex underwater environments. Experimental validation across multiple regions shows that the proposed algorithm significantly outperforms density-based spatial clustering of applications with noise (DBSCAN) and ordering points to identify the clustering structure (OPTICS) in terms of accuracy and robustness. Experimental results show that the proposed algorithm achieves a coefficient of determination ($R^{2}$) of 0.99 across different regions, with root mean square error (RMSE) ranging from 0.27 to 0.38 m, and mean absolute error (MAE) ranging from 0.21 to 0.33 m. Compared to traditional methods, the proposed algorithm effectively removes noise points in sparse photon regions while preserving the continuity and details of underwater terrain in photon-dense areas. The experimental results demonstrate that the proposed algorithm accurately extracts the underwater contours of complex terrain and shows strong adaptability and stability in various aquatic environments. This method demonstrates outstanding precision and denoising capabilities in underwater terrain extraction in shallow water areas, providing a reliable and efficient approach for high-precision depth measurement of complex underwater topography. Yu Li 0037, Lei Zhou 0015, Dongzhen Jia, Nan Xu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Combining Airborne LiDAR Data and Optical Imagery for Improved National-Scale Beach Topography Estimation: A Case Study in New ZealandabstractAccurate beach topography mapping is crucial for understanding coastal dynamics and mitigating climate change impacts. However, traditional methods such as airborne LiDAR have limitations, leading to substantial gaps in national-scale elevation data. This study presents an innovative framework to reconstruct missing elevation data along New Zealand’s coastline by integrating airborne LiDAR, Sentinel-2 optical imagery, and geometric features (distance) using machine learning methods. Our results show that Artificial Neural Network (ANN) emerged as the best model (test set: R²=0.79, RMSE=0.91 m; validation set: 0.79, RMSE=0.93 m), outperforming other models in accuracy. The produced 10-m DEM for national-scale sandy beaches expands area coverage by 286.6% (114.15 km²), filling gaps in 1249 beaches, including remote areas such as Stewart Island. This novel framework offers a scalable solution for improving the comprehensiveness and accuracy of beach topography. It provides essential support for inundation prediction, habitat management, and the development of climate adaptation strategies, thereby facilitating more informed decision-making in coastal zone management and climate change mitigation efforts. Conghong Huang, Yue Ma 0002, Xin Ma 0007, Yifu Ou, Chunpeng Chen, Shaoguang Zhou, Dongzhen Jia, Zhen Wang 0020, Qingquan Li 0001, Nan Xu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 12 |
| 2025 | FDMamba: Frequency-Driven Dual-Branch Mamba Network for Road Extraction From Remote Sensing ImagesabstractRoad extraction from remote sensing imagery is crucial for a variety of applications, including transportation monitoring, disaster response, and urban planning. However, existing methods often fail to accurately delineate sparse, curvilinear, and boundary-blurred road structures in high-resolution images, leading to incomplete detail preservation and inadequate contextual understanding. To address these challenges, we propose a novel Frequency-Driven Dual-Branch Mamba Network (FDMamba) for precise road extraction from remote sensing imagery. The proposed FDMamba integrates frequency-aware modeling with a dual-branch architecture, enabling collaborative learning of fine-grained edge details and global spatial dependencies. Specifically, FDMamba comprises three key modules: a Fourier Reconstruction Attention Mechanism (FRAM) to enhance high-frequency boundary information and low-frequency structural representation; a Rotation-aware Mamba Module (RAMamba) that leverages multi-path state space modeling for robust directional perception of road structures; and a Phase-guided Feature Fusion Module (PFFM) for effective cross-scale alignment and fusion of high- and low-frequency features. Furthermore, to mitigate the issue of blurred or ambiguous boundaries, we introduce a hybrid loss function that combines binary cross-entropy, focal loss, and frequency-aware loss, explicitly guiding the model to focus on edge structure and multi-frequency complementary information. Extensive experiments on three benchmark datasets, CHN6-CUG, DeepGlobe, and Massachusetts, demonstrate that FDMamba consistently outperforms state-of-the-art methods in terms of F1-score and IoU, achieving superior boundary clarity and structural continuity while preserving overall geometric integrity. The code is available at https://github.com/darkseid-arch/RE-FDMamba. Zhen Wang 0020, Shen-Ao Yuan, Nan Xu 0008, Zhu-Hong You, De-Shuang Huang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | UAVSeg: Dual-Encoder Cross-Scale Attention Network for UAV Images' Semantic SegmentationabstractBenefiting from the powerful feature extraction and feature correlation modeling capabilities of convolutional neural networks (CNNs) and Transformer models, these techniques have been widely used in unmanned aerial vehicle (UAV) aerial image semantic segmentation tasks. However, the ground objects in aerial images contain feature information with different scales, and existing methods directly cascade low-level visual features and high-level semantic features without processing, resulting in low semantic segmentation precision. To address these challenges, we propose a dual-encoder cross-scale attention network, which efficiently extracts local and global context information from aerial images and performs fine-grained fusion of multiscale features to improve semantic segmentation performance. First, we introduce the dual-CNN-Transformer encoder, which embeds the scan-focus window Transformer (SFWT) into CNNs as an auxiliary encoder to supplement the local feature information lost in the global context information extraction process. Second, the cross-scale lightweight integration (CSLI) module is designed, which uses a light dot-product attention mechanism (DPAM) to fusion multiscale features and reduce model calculation parameters. Finally, the linear multilayer perceptron (LMLP) is used to restore the feature map resolution while expanding the deconvolution receptive field. To validate the effectiveness of the proposed method, we conducted extensive experiments on real aerial scene datasets, including UAVid, Urban Drone, and AeroScapes. The experimental results show that our method achieves state-of-the-art performance while maintaining superior real-time efficiency. Implementation codes will be available athttps://github.com/darkseid-arch/UAVSeg. Zhen Wang 0020, Zhu-Hong You, Nan Xu 0008, Chuanlei Zhang, De-Shuang Huang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Locality-Enhanced Transformer for Semantic Segmentation of High-Resolution Remote Sensing ImagesabstractTransformers have emerged as a transformative tool in various computer vision tasks, excelling at capturing long-range dependencies. Their potential applicability and scalability in the interpretation of high-resolution remote sensing images (HRRSIs) have thus garnered substantial interest. However, unlike natural images, HRRSIs present intricate scenes characterized by scale variations and diverse appearances. These challenges underscore the importance of enabling networks to effectively assimilate both local intricacies and global context. In this letter, we introduce LETFormer, a semantic segmentation transformer. LETFormer balances capturing longrange dependencies with preserving local details through its unique LETFormer block, featuring an anchor token. This token aggregates localized contextual information within a designated window and promotes meaningful interactions among anchor tokens. With a mask transformer decoder, LETFormer gains ample contextual cues for precise semantic mask prediction. Empirical findings based on evaluations using the ISPRS Potsdam and LoveDA benchmarks unequivocally establish LETFormer’s superiority over state-of-the-art models. Additionally, we analyze the parameter size and floating-point operations per second (FLOPs) of LETFormer. Xin Li 0090, Feng Xu 0008, Runliang Xia, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Qian Huang 0008, Xin Lyu 0001 |
ICASSP | 4 |
| 2024 | AAFormer: Attention-Attended Transformer for Semantic Segmentation of Remote Sensing ImagesabstractThe rapid advancements in remote sensing technology have enabled the widespread availability of fine-resolution remote sensing images (RSIs), offering rich spatial details and semantics. Despite the applicability and scalability of transformers in semantic segmentation of RSIs by learning pairwise contextual affinity, they inevitably introduce irrelevant context, hindering accurate inference of patch semantics. To address this, we propose a novel multi-head attention-attended module (AAM) that refines the multi-head self-attention mechanism. The AAM filters out irrelevant context while highlighting informative ones by considering the relevance between self-attention maps and the query vector. The AAM generates an attention gate to complement contextual affinity and emphasize the useful ones with a higher weight simultaneously. Leveraging multi-head AAM as the core unit, we construct a lightweight attention-attended transformer block (ATB). Subsequently, we devise AAFormer, a pure transformer with a mask transformer decoder, for achieving semantic segmentation of RSIs. We extensively evaluate our approach on the ISPRS Potsdam and LoveDA datasets, demonstrating compelling performance compared to mainstream methods. Additionally, we conduct evaluations to analyze the effects of AAM. Xin Li 0090, Feng Xu 0008, Linyang Li, Nan Xu 0008, Fan Liu 0003, Chi Yuan, Xin Lyu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Coastal Bathymetry Determined From Water Waves Observed by Airborne Lidars: A Case Study Near Ganquan Island, South China SeaabstractThe passive multispectral imaging and active bathymetric lidar make a great achievement for bathymetry in optically shallow waters. However, due to the attenuation of the water column precludes deep penetration of the light, accurately obtaining the underwater topography in turbid waters through remote sensing techniques, both passive and active, is still a challenging task. Airborne lidars can obtain water surface topography with high accuracy and resolution, which can further be used to derive the water depth based on wave theory. In this study, an ‘indirect’ method to determine water depth is proposed using airborne lidar measured water surface points. As the wavelength and wave direction can be accurately tracked from the water surface topography, a 20m×20m underwater topography near Ganquan Island, South China Sea, is generated with an RMSE of 0.91 m and a MAPE of 7.1%. The basic theory of deriving water depths is totally different from airborne lidar bathymetry, i.e., this method is independent of water clarity and can be used in turbid waters or even works with near infrared airborne lidar that can only obtain water surface points. Jian Yang 0033, Yue Ma 0002, Nan Xu 0008, Hui Zhou 0013, Xiaohua Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | A Method to Derive Bathymetry for Dynamic Water Bodies Using ICESat-2 and GSWD Data SetsabstractDetailed information on lake bathymetry is essential for both hydrology-related studies and water resource management. Conventionally, lake bathymetry was mapped using high-cost approaches (e.g., ship/boat-based multibeam echosounders or airborne bathymetric lidars). With only satellite remotely sensed data sets, a method for deriving high-resolution bathymetry for dynamic areas was proposed by combining the new Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) lidar data and the Landsat-based Global Surface Water Data Set (GSWD). First, ICESat-2 can provide accurate along-track topographic points after the point cloud processing and bathymetric error correction, and the GSWD can supply water occurrence information within the lake dynamic area between 1984 and 2018. Second, using the derived relationship between the elevation and water occurrence, the bathymetric map of Lake Mead, USA, was produced with the dynamic area exceeding 235 km2, elevation ranging nearly 37 m, and a resolution of 30 m. The local reference data (i.e., the airborne topographic lidar data and ship/boat-based bathymetric data) in six areas around Lake Mead were used for the validations. In general, the produced lake bathymetry achieved an accuracy of approximately 2 m in elevation with$R^{2}$of 0.97. The proposed method is promising to obtain global bathymetry for inland water bodies (e.g., the lake and reservoir) and coastal areas (e.g., the tidal zone) where water level fluctuations are strong and the water clarity is sufficient. Nan Xu 0008, Yue Ma 0002, Hui Zhou 0013, Zhiyu Zhang 0006, Xiaohua Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | SCFNet: Semantic Condition Constraint Guided Feature Aware Network for Aircraft Detection in SAR ImagesabstractAircraft detection in synthetic aperture radar (SAR) images plays an essential role in satellite observation and military decisions. Due to discrete scattering properties, speckle noise interference, and various aircraft types, many existing methods struggle to achieve the desired detection performance. In this article, we propose an innovative semantic condition constraint guided feature aware network (SCFNet) for detecting different aircraft categories in SAR images. First, considering the discrete scattering properties of aircraft, we design a local-global feature aware module (LGA-M) and morphological-semantic feature aware module (MSF-M), which can effectively extract the fine-grained feature information contained in SAR images. Second, to effectively fuse different feature information, we construct a feature fusion pyramid (FFP), which uses different branches and paths to reasonably merge multiple feature information types and suppresses background information interference. Third, according to the structure characteristics of aircraft, the global coordinate attention mechanism (G-CAT) is presented to highlight foreground target features and suppress speckle noise interference. Finally, we construct semantic condition constraints, including constraint condition setting, semantic information calculation, and template matching, to improve aircraft localization and recognition accuracy. Extensive experiments demonstrate that the proposed SCFNet can obtain state-of-the-art performance on the SAR aircraft detection dataset, which achieves AP and F1 Score of 94.83% and 95.58%, respectively. The related implementation codes will be made publicly available at https://github.com/darkseid-arch/AirDetection. Zhen Wang 0020, Nan Xu 0008, Jianxin Guo, Chuanlei Zhang, Buhong Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Surface-Water-Level Changes During 2003-2019 in Australia Revealed by ICESat/ICESat-2 Altimetry and Landsat ImageryabstractSurface-water-level changes reflect Earth's water resource variations (e.g., trends and fluctuations) and they are helpful to understand potential drivers (e.g., climate change and human activities). Currently, Australia is facing serious water crisis owing to rainfall shortage and climate change, and national-scale data set on surface-water-level changes is required for supporting sustainable water resource management. Here, we used all Landsat Thematic Mapper (TM)/Enhanced Thematic Mapper (ETM)+/Operational Land Imager (OLI) data available on Google Earth Engine to obtain annual surface water during 2003-2019 and produced 1506 boundaries of water bodies with areas greater than 1 km2across Australia. The produced surface water map in Australia is more accurate than the existing global lake databases (e.g., the Global Lakes and Wetlands Database), which can be downloaded for free. Then, 52 water bodies (lakes and reservoirs) with areas larger than 1 km2and available Ice, Cloud, and land Elevation Satellite (ICESat/ICESat-2) data for more than 5 years were combined to estimate trends in surface water levels in Australia. Across Australia, from 2003 to 2019, the area-weighted mean of water level change rates is -0.046 m/year with 17 lakes (32.7%) with increasing water levels and 35 lakes (67.3%) decreasing with water levels. In detail, the largest lakes (>100 km2) dominate the total change trend and most of the largest lakes underwent decreasing trend (-0.046 m/year), whereas the mean water levels of small lakes (2) increased in the past 17 years. In situ water levels of three typical lakes/reservoirs were used to validate our estimation results, which exhibited a very good agreement (R2= 0.98). Nan Xu 0008, Yue Ma 0002, Xiaohua Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |