EDBT 2026 Demo / reviewers in the wild / expert
Xiongwu Xiao
dblp:136/5536
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-3035-7727ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AstroHSP: A hybrid supervision framework for robust monocular astronaut pose estimation
Haohang Jian, Xiongwu Xiao, Jianguo Yan, Zhigang Tu 0001 |
Neural Networks | 3 |
| 2025 | RTO-LLI: Robust Real-Time Image Orientation Method With Rapid Multilevel Matching and Third-Times Optimizations for Low-Overlap Large-Format UAV ImagesabstractUAV real-time photogrammetry is important to promote the rapid generation of photogrammetry 4D product, intelligent information extraction and rapid remote sensing mapping, and efficient large-scale 3D modeling. However, for real-time processing of low-overlap large-format image sequence, there remains two challenges: (1) Large-format images result in greater data volume and computational load, posing challenges for real-time online processing on regular-performance computing units, requiring more efficient algorithms; (2) Low-overlap images make it difficult for matching correspondences to cover the entire overlapping area at real-time, leading to significant challenges for real-time and robust relative orientation. Therefore, this paper proposes a robust Real-Time Orientation method for Low-overlap Large-format UAV Images (RTO-LLI), which can robustly handle these kind of data in real-time. Firstly, robust initialization method for real-time processing of low-overlap large-format images was designed to ensure a high-success-rate of SLAM initialization. Secondly, constant velocity hypothesis tracking enables fast orientation during constant-speed flight. Thirdly, when the second step false, using real-time pose estimation method based on multilevel matching and coarse-to-fine optimization to robustly solve the precision image pose. Fourthly, final (third-level) pose optimization method based on the IRLS algorithm with suitable search area, which can compute higher-precision image pose in real-time. Finally, real-time mapping based on parallel processing for low-overlap images can generate high-precision 3D point maps and complete feature extraction for the next frame in real-time. Experiments conducted on several different types of scenes show that: (1) the processing speed of RTO-LLI significantly surpasses traditional offline methods: PhotoScan, OpenMVG, Colmap. RTO-LLI can handle large-format UAV image sequence (single-imagery has 20-million-pixels) at a speed of 1.5 frames-per-second, meeting the demands of real-time UAV photogrammetry tasks; (2) RTO-LLI is the only method that has successfully completed real-time tasks in all 50-times repeated experiments for four different types of scenes, demonstrating robustness far superior to other classical SLAM solutions; (3) the-displacement-error of the estimated Pose by RTO-LLI is less than 1/2000 of the-trajectory-length, and the average-reprojection-error is less than 1.5 pixels, almost as well as traditional offline methods. RTO-LLI method meets the efficiency, robustness and accuracy requirements of real-time photogrammetry for low-overlap large-format UAV images. Xiongwu Xiao, Gui-Song Xia, Jianya Gong, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | SLCGC: A lightweight Self-supervised Low-Pass Contrastive Graph Clustering Network for Hyperspectral ImagesabstractSelf-supervised hyperspectral image (HSI) clustering remains a fundamental yet challenging task due to the absence of labeled data and the inherent complexity of spatial-spectral interactions. While recent advancements have explored innovative approaches, existing methods face critical limitations in clustering accuracy, feature discriminability, computational efficiency, and robustness to noise, hindering their practical deployment. In this paper, a self-supervised efficient low-pass contrastive graph clustering (SLCGC) is introduced for HSIs. Our approach begins with homogeneous region generation, which aggregates pixels into spectrally consistent regions to preserve local spatial-spectral coherence while drastically reducing graph complexity. We then construct a structural graph using an adjacency matrix A and introduce a low-pass graph denoising mechanism to suppress high-frequency noise in the graph topology, ensuring stable feature propagation. A dual-branch graph contrastive learning module is developed, where Gaussian noise perturbations generate augmented views through two multilayer perceptrons (MLPs), and a cross-view contrastive loss enforces structural consistency between views to learn noise-invariant representations. Finally, latent embeddings optimized by this process are clustered via K-means. Extensive experiments and repeated comparative analysis have verified that our SLCGC contains high clustering accuracy, low computational complexity, and strong robustness. The code source will be available athttps://github.com/DY-HYX. Yao Ding 0010, Aitao Yang, Yaoming Cai, Xiongwu Xiao, Danfeng Hong, Junsong Yuan 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | LUMNet: Land Use Knowledge Guided Multiscale Network for Height Estimation From Single Remote Sensing ImagesabstractFor estimating ground object height from single remote sensing images, this study proposes a land use knowledge guided multiscale height estimation network (LUMNet), which takes single image and land use data as inputs and produces an estimated height map as output. First, in the encoder part, the visual geometry group network (VGGNet) is used to extract multilevel deep semantic features from images. Second, the features of the encoder part are fused with those of the decoder part through skip connection with land use knowledge attention weight and Dense Atrous Spatial Pyramid Pooling (DenseASPP). Third, in the decoder part, the features are decoded using upsampling and convolution, and a joint loss function is constructed to supervise network training. Experiments show that the proposed method achieves the best visualizations and quantitative evaluation results among all the tested methods. For the Vaihingen dataset, theRMSE,MAE, andZNCCof LUMNet are 1.523, 0.969, and 0.908, respectively. For the Potsdam dataset, these are 2.158, 1.222, and 0.871, respectively. The source code of LUMNet has been made public at the following link: https://figshare.com/s/eaf206879ade35a61f88. Shouhang Du, Jianghe Xing, Xiongwu Xiao, Jun Li 0021, Hao Liu 0105 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Patch Similarity Self-Knowledge Distillation for Cross-View Geo-LocalizationabstractCross-view geo-localization is an extremely challenging task due to drastic discrepancies in scene context and object scale between different views. Existing works normally concentrate on aligning the global appearance between two views but underestimate these two discrepancies. In practice, only a small region in the retrieved aerial image can be matched to the whole query ground image (i.e. scene context change). On the other hand, the retrieved aerial images are only able to describe the coarse-grained information but the query ground images can capture the fine-grained details (i.e. object scale change). In this paper, we propose a novel self-distillation framework called Patch Similarity Self-Knowledge Distillation (PaSS-KD), which provides the local and multi-scale knowledge as fine-grained location-related supervision to guide cross-view image feature extraction and representation in a self-enhanced manner. Specifically, we develop an auxiliary image-to-patch retrieval task to explore the scene context change and devise a multi-scale patch partition strategy to sense the object scale change across views. Additionally, our self-distilling framework can be removed to avoid additional computation cost at the inference stage. Extensive experiments show that our method not only achieves the state-of-the-art image retrieval performance on the CVUSA and CVACT benchmarks, but also significantly boosts the fine-grained localization accuracy on the VIGOR dataset. Remarkably, for 10 meter-level localization, we improve the relative accuracy by a factor of 0.8× and 1.6× on the VIGOR dataset under same-area and cross-area evaluation, respectively. Songlian Li, Xiongwu Xiao, Zhigang Tu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | TL2GH²T: Triple-Path Local-to-Global Network With Hybrid Head Transformer for Hyperspectral Change DetectionabstractWith the aid of transformers, significant progress has been achieved in hyperspectral image change detection (HSI-CD) in recent times. Nonetheless, most contemporary detection methods fail to incorporate diverse diagnostic features extracted from hyperspectral (HS) images. In addition, relying solely on algebraic-based techniques to extract information of difference is insufficient for achieving satisfactory detection performance. In this regard, we propose an innovative triple-path local-to-global network (TL2GN), complemented by a hybrid head transformer (HybridHT), called TL2GH2T, tailored for HSI-CD tasks. To be specific, TL2GH2T first investigates spatial, spectral, and spatial–spectral features from a local-to-global perspective. Then, a novel spatial and spectral token fusion (SSTF) module is developed to integrate the above three tokenized features, producing discriminative features from two HS images separately. Moreover, drawing inspiration from chromosomal crossover mechanisms, we propose a HybridHT. Its goal is to simultaneously learn cross correlation and self-correlation information of bitemporal features from a global perspective, producing highly discriminative distinctions. Our approach, validated through extensive experimentation on four varied HS benchmarks, exhibits exceptional performance in HSI-CD, outperforming contemporary methods in both visual and quantitative evaluations. Zhonghao Chen, Swalpa Kumar Roy, Hongmin Gao 0001, Yao Ding 0010, Xiongwu Xiao, Bing Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | A Novel LOD Rendering Method With Multilevel Structure-Keeping Mesh Simplification and Fast Texture Alignment for Realistic 3-D ModelsabstractFast, high-precision texture maps and high-frame-rate level of detail (LOD) generation for realistic 3-D models are foundational data infrastructures for smart cities. However, LOD generation faces three main issues: reduced model accuracy from mesh simplification, inefficient texture memory utilization, and browsing lag with detail loss. This article proposed a novel LOD rendering method with multilevel structure-keeping mesh simplification and fast texture alignment for realistic 3-D models. First, a multilevel structure-keeping mesh simplification method with mesh segmentation and vertex classification was used to generate a simplified mesh with high-precision structure preservation. Second, a fast texture alignment method was proposed that uses segmentation information and least-squares conformal map (LSCM) parameterization to acquire texture blocks. The method integrated integral images and a precise multitemplate strategy to align texture blocks, to obtain texture maps with high completeness and high occupancy. Finally, by integrating these methods, an LOD generation method with fast multilevel pyramid construction and adaptive tree organization is proposed. This method achieved high-precision multilevel structure keeping, along with a high-occupancy rate of texture maps, facilitating a high frame rate for LOD model construction. Compared with quadratic error function (QEF), quadratic error metrics (QEMs), low-poly, computational geometry algorithms library (CGAL), and Nvdiffrec, the proposed mesh simplification algorithm demonstrated average accuracy improvements of 12.1%, 24.2%, 57.1%, 17.9%, and 3.2%, respectively. Compared with open multi-view environment (OpenMVE), ContextCapture, and Xatlas, the proposed texture alignment algorithm achieved average occupancy improvements of 27.74%, 11.89%, and 4.80%, respectively. Compared with the state-of-the-art ContextCapture and Smart3D, the proposed method for browsing large-scale realistic 3-D models increased the frame rate by 17.3% and 14.4%, respectively. Yingwei Ge, Xiongwu Xiao, Bingxuan Guo, Jianya Gong, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | LiST-Net: Enhanced Flood Mapping With Lightweight SAR Transformer Network and Dimension-Wise AttentionabstractDetecting flood-induced changes using synthetic aperture radar (SAR) is crucial for crisis management and damage assessment. Nevertheless, current methodologies predominantly focus on changes in buildings within optical images, struggling with the complex structures of floods. These structures are marked by widespread speckle noise and are accompanied by an increase in computational cost. These challenges hinder their success in real-world applications, necessitating a novel approach. This paper proposes LiST-Net, a lightweight SAR transformer network with dimension-wise attention to improve flood detection accuracy. LiST-Net offers three key advantages. Firstly, the graph neighbor module (GNM) is designed to enhance both detailed information of neighboring pixels and multi-date features within the encoder. Secondly, the dimension-wise interactive attention (DIA) module is proposed to effectively reduce computational complexity while enhancing feature representation. Thirdly, an attentive supervised learning module (ASLM) is incorporated to mitigate noise through a pixel mask gate, allowing change water information to pass through and improving the accuracy of water edge delineation. The effectiveness of LiST-Net is evaluated on two flood detection datasets, S1GFloods and ETCI-2021. Experimental results demonstrate that LiST-Net outperforms existing methods, showcasing a 94.7% improvement in F1 and an 88.7% enhancement in IoU on the S1GFloods datasets, with lower computational costs (11.78G) and fewer parameters (7.34M). This underscores LiST-Net as a promising strategy for precise and effective mapping of floods within SAR images in real-world applications. A public release of the demo code will be available at https://github.com/Tamer-Saleh. Tamer Saleh, Shimaa Holail, Xiongwu Xiao, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | An Efficient and Lightweight Spectral-Spatial Feature Graph Contrastive Learning Framework for Hyperspectral Image ClusteringabstractDue to the scarcity of prior information and the high complexity of spectral data, hyperspectral image (HSI) clustering presents a significant challenge. Although recent deep clustering methods have demonstrated remarkable performance, their intricate network structures and poor robustness hinder their practical application. To address this issue, we propose an efficient and lightweight spectral-spatial feature graph contrastive learning (S2GCL) framework for robust HSI clustering. Specifically, we have designed a novel spectral-spatial feature encoder that fully leverages the information in HSI by incorporating both spatial structure and spectral similarity matrices. To establish a lightweight model, we implement several effective designs: First, S2GCL eliminates the commonly used data augmentation and discriminator in GCL during the generation of positive embeddings. Second, we use a multilayer perceptron (MLP) to produce low-dimensional embeddings instead of relying on graph convolutional networks (GCNs). Third, negative embeddings are generated through row-shuffling, avoiding the use of neural networks. Finally, we propose a multiple boundary loss function to extract complementary information from spatial structures and neighboring nodes, while also constraining the interclass differences between positive and negative examples. We conducted extensive experiments on four publicly available datasets and compared S2GCL with state-of-the-art clustering methods. The results indicate that S2GCL achieves satisfactory performance. The code for S2GCL will be released athttps://github.com/ahappyyang/S2GCL. Aitao Yang, Min Li 0030, Yao Ding 0010, Xiongwu Xiao, Yujie He 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | KT-Net: Knowledge Transfer for Unpaired 3D Shape CompletionabstractUnpaired 3D object completion aims to predict a complete 3D shape from an incomplete input without knowing the correspondence between the complete and incomplete shapes. In this paper, we propose the novel KTNet to solve this task from the new perspective of knowledge transfer. KTNet elaborates a teacher-assistant-student network to establish multiple knowledge transfer processes. Specifically, the teacher network takes complete shape as input and learns the knowledge of complete shape. The student network takes the incomplete one as input and restores the corresponding complete shape. And the assistant modules not only help to transfer the knowledge of complete shape from the teacher to the student, but also judge the learning effect of the student network. As a result, KTNet makes use of a more comprehensive understanding to establish the geometric correspondence between complete and incomplete shapes in a perspective of knowledge transfer, which enables more detailed geometric inference for generating high-quality complete shapes. We conduct comprehensive experiments on several datasets, and the results show that our method outperforms previous methods of unpaired point cloud completion by a large margin. Code is available at https://github.com/a4152684/KT-Net. Xin Wen 0003, Zhen Dong 0005, Yu-Shen Liu, Xiongwu Xiao, Bisheng Yang |
AAAI | 6 |
| 2023 | AFDE-Net: Building Change Detection Using Attention-Based Feature Differential Enhancement for Satellite ImageryabstractBuilding Change Detection (BCD) from satellite imagery is critical for monitoring urbanization, managing agricultural land, and updating geospatial databases. However, complex variations in building roofs that resemble the background of their surroundings pose challenges for deep learning-based change detection methods due to their focus on color and texture. Additionally, downsampling can result in the loss of spatial information, leading to incomplete buildings and irregular output boundaries. To address these challenges, a novel Siamese network called AFDE-Net is proposed, which combines differential image features and attention modules using a learnable parameter. The AFDE-Net employs an ensemble spatial-channel attention fusion (ESCAF) module, along with a deep supervision module, to mitigate the loss of spatial information and refine deep features in high-dimensional inputs. Besides, we have created a new dataset (EGY-BCD) comprising high-resolution and multi-temporal satellite images captured in four urban and coastal areas in Egypt to detect building changes. The EGY-BCD dataset includes images with complex types of change, such as tall and dense buildings with roofs that resemble the background of their surroundings, which is a challenge for deep learning algorithms. The proposed method outperforms other methods on the EGY-BCD dataset with an overall accuracy of 94.3%, an F1-score of 88.8%, and an mIoU of 86.6%. The datasets and codes will be released at https://github.com/oshholail/EGY-BCD. Shimaa Holail, Tamer Saleh, Xiongwu Xiao, DeRen Li |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2013 | Stereo Matching with Weighted Feature Constraints for Aerial ImagesabstractThe stereo correspondence problem represents a great challenge in Photogrammetry. In recent years, the Semi-Global Matching (SGM) method proposed by Hirschmüller et. al. is a promising approach for solving stereo matching problems in different applications and data sets, including aerial images. In this paper, a Ground Control Points based Semi-Global stereo Matching (GCP-SGM) method is proposed for aerial images. Ground Control Points are incorporated into the global energy function model in SGM to lessen matching ambiguities. Significant accuracy improvement is obtained by utilizing the prior information in the ground control points as soft constraints. Experimental results in real scenes show that when such prior information exists, the GCP-SGM method can effectively improve accuracy. Xiongwu Xiao, Bingxuan Guo, Yueru Shi |
ICIG | 1 |