Shaobo Xia

dblp:178/8198 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0003-2890-838XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 NoKSR: Kernel-Free Neural Surface Reconstruction via Point Cloud Serialization
abstract
We present a novel approach to large-scale point cloud surface reconstruction by developing an efficient framework that converts an irregular point cloud into a signed distance field (SDF). Our backbone builds upon recent transformer-based architectures (i.e. PointTransformerV3), that serializes the point cloud into a locality-preserving sequence of tokens. We efficiently predict the SDF value at a point by aggregating nearby tokens, where fast approximate neighbors can be retrieved thanks to the serialization. We serialize the point cloud at different levels/scales, and non-linearly aggregate a feature to predict the SDF value. We show that aggregating across multiple scales is critical to overcome the approximations introduced by the serialization (i.e. false negatives in the neighborhood). Our frameworks sets the new state-of-the-art in terms of accuracy and efficiency (better or similar performance with half the latency of the best prior method, coupled with a simpler implementation), particularly on outdoor datasets where sparse-grid methods have shown limited performance.
Zhen Li 0052, Shrisudhan Govindarajan, Shaobo Xia, Daniel Rebain, Kwang Moo Yi, Andrea Tagliasacchi
3DV4
2025 HyperEDL: Spectral-Spatial Evidence Deep Learning for Cross-Scene Hyperspectral Image Classification
abstract
Cross-scene hyperspectral image (HSI) classification presents significant challenges due to domain shifts, which amplify epistemic uncertainty and lead to substantial performance drops in unseen scenes. While evidence deep learning (EDL) has shown promise in modeling uncertainty, existing methods fall short, as they do not explicitly account for the epistemic uncertainty arising from spatial-spectral feature interactions. To address these challenges, we propose the spectral-spatial evidence deep learning for cross-scene hyperspectral image classification (HyperEDL) framework, which introduces the spatial-spectral multiorder aggregation module (SS-Moga). This module effectively captures and adaptively encodes multiorder contextual interactions from both spatial and spectral perspectives. By combining multiorder contextual encoding with spatial-spectral confidence, our approach fully aggregates multiorder evidence to mitigate epistemic uncertainty arising from knowledge gaps between seen and unseen scenes. Specifically, it uses Dirichlet distribution to capture correlation between spatial-spectral knowledge about different scenes, which can be generalized to unseen scenes. Extensive experiments on three benchmark datasets demonstrate that HyperEDL outperforms state-of-the-art methods, showcasing its effectiveness and strong generalization ability.
Yangbo Feng, Shuhe Wang, Jun Yue 0004, Shaobo Xia, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.5
2025 Exemplar-Free Lifelong Hyperspectral Image Classification With Spectral Consistency
abstract
Hyperspectral image (HSI) classification suffers from severe catastrophic forgetting in exemplar-free lifelong learning, where models must continuously learn new land cover categories without accessing historical training samples. This challenge persists due to high-dimensional spectral-volumetric complexity and cross-task spectral drift, which current methods inadequately address. We propose HyperSC, a novel framework that synergizes spectral-consistent auxiliary samples synthesis with stability-plasticity fused learning. The framework consists of three key components: a Spectral Consistency Model Inversion (SCMI) module, a Spectral Progressive Enhancement (SPE) module, and a Fusion Distillation Learning (FDL) module. The SCMI module synthesizes class-conditional auxiliary samples through spectral moment matching, in which the mean and variance of each spectral band are constrained to match class-specific real historical data distributions, thereby achieving spectral consistency. The SPE module injects class-specific Gaussian noise and applies momentum-based updating to enhance sample diversity while preserving spectral fidelity. The FDL module jointly trains on fused real and auxiliary samples by coordinating cross-entropy classification, output-layer knowledge distillation, and intermediate-layer feature alignment, thereby enabling plasticity for new class learning while maintaining stability against catastrophic forgetting for previous tasks. Extensive experiments on three HSI datasets (Indian Pines, Houston, Salinas) demonstrate HyperSC’s superiority compared to previous exemplar-free lifelong learning methods. The code is available at https://github.com/lzlsxs/hypersc.
Zhenlin Li, Shaobo Xia, Shuhe Wang, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.2
2025 HyperKD: Lifelong Hyperspectral Image Classification With Cross-Spectral-Spatial Knowledge Distillation
abstract
Hyperspectral image (HSI) classification models suffer from a phenomenon known as catastrophic forgetting, which refers to the sharp decline in performance on previously learned tasks after learning a new one when continuously acquiring new knowledge from a sequence of tasks. In recent years, some lifelong learning approaches have been proposed for HSI classification. Despite some progress, the challenge of catastrophic forgetting in lifelong learning remains significant and unresolved. In this article, we propose a novel lifelong learning framework for HSI classification, which is based on exemplar replay and cross-spectral–spatial feature knowledge distillation (KD), termed HyperKD. Specifically, the proposed framework incorporates a min-max cross-selection (MMCS) module tailored to HSI characteristics with a cross-spectral-spatial knowledge distillation (CSSKD) module. The MMCS module selects the most representative or diverse samples as exemplars from previous tasks for replay. Additionally, the CSSKD module not only transfers the prediction logit distribution from the previous network to the current network but also transfers the spectral-spatial feature distribution via cross-network KD, without directly assessing the similarity of feature distributions, thereby retaining more knowledge and mitigating forgetting. Through experiments conducted on a series of tasks, including the Pavia, Indian Pines, Salinas, and Houston datasets, our approach demonstrates superior performance compared to previous lifelong learning methods for HSI classification, effectively mitigating catastrophic forgetting. The code implementation of our approach will be publicly available athttps://github.com/lzlsxs/hyperkd.
Zhenlin Li, Shaobo Xia, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.2
2025 Benchmarking ULS-TLS Point Cloud Registration Algorithms in Forest Environments
abstract
Integrating Unmanned aerial vehicle Laser Scanning (ULS) and Terrestrial Laser Scanning (TLS) data in complex forest environments remains a significant challenge. Despite the availability of numerous registration algorithms, robust comparative studies are limited by the lack of reliable multi-platform benchmark datasets. In this study, we introduce the first multiplatform benchmark dataset for ULS-TLS point cloud registration in forests, encompassing 17 plots from seven diverse regions with about 1.56 billion points. The dataset is categorized into three difficulty levels based on overlap ratio and rigid overlap.We evaluated the performance of five registration algorithms against this benchmark. Chen2022 achieved the highest accuracy with a 100% success rate across all difficulty levels. While Wu2024 demonstrated robust performance in lower difficulties but faced challenges in more complex scenarios. We also found that terrain variations and rigid overlap significantly impacted registration accuracy, particularly for algorithms reliant on individual tree positions such as Hyypp¨a2021 and Feng2024. These findings underscore the need for improved data collection strategy, ground filtering techniques, and feature matching algorithms to enhance performance in challenging environments. We present the first openly accessible multi-platform benchmark dataset for forested regions and anticipate that future research will expand this work to additional areas. The dataset can be downloaded from: DatasetDownloadLink.
Wangjun Liu, Sheng Nie, Shaobo Xia, Cheng Wang 0016, Jinliang Wang 0002, Xiaohuan Xi
IEEE Trans. Geosci. Remote. Sens.3
2024 Spectral Query Spatial: Revisiting the Role of Center Pixel in Transformer for Hyperspectral Image Classification
abstract
Recently, there have been significant advancements in Hyperspectral Image (HSI) classification methods employing Transformer architectures. However, these methods, while extracting spectral-spatial features, may introduce irrelevant spatial information that interferes with HSI classification. To address this issue, this paper proposes a Spectral Query Spatial Transformer (SQSFormer) framework. The proposed framework utilizes the center pixel (i.e., pixel to be classified) to adaptively query relevant spatial information from neighboring pixels, thereby preserving spectral features while reducing the introduction of irrelevant spatial information. Specifically, this paper introduces a Rotation-Invariant Position Embedding module to integrate random central rotation and center relative position embedding, mitigating the interference of absolute position and orientation information on spatial feature extraction. Moreover, a Spectral-Spatial Center Attention module is designed to enable the network to focus on the center pixel by adaptively extracting spatial features from neighboring pixels at multiple scales. The pivotal characteristic of the proposed framework achieves adaptive spectral-spatial information fusion using the Spectral Query Spatial paradigm, reducing the introduction of irrelevant information and effectively improving classification performance. Experimental results on multiple public datasets demonstrate that our framework outperforms previous state-of-the-art methods. For the sake of reproducibility, the source code of SQSFormer will be publicly available at https://github.com/chenning0115/SQSFormer.
Leyuan Fang, Shaobo Xia, Hui Liu 0041, Jun Yue 0004
IEEE Trans. Geosci. Remote. Sens.4
2024 Semantic Segmentation of Airborne LiDAR Point Clouds With Noisy Labels
abstract
High-quality point cloud annotation is labor-intensive and time-consuming, but it serves as a critical factor driving the success of LiDAR point cloud semantic segmentation. Leveraging low-quality labels in LiDAR point cloud processing is overlooked, despite the fact that noisy annotation has low labeling costs and abundant cross-modal resources (e.g., labels from images). To this end, we thoroughly investigate the performance of airborne LiDAR point cloud semantic segmentation models using noisy labels for the first time and find that it is closely related to object categories and learning stages. Then we propose a new semantic segmentation framework for LiDAR point cloud noisy learning called adaptive dynamic noise label correction (ADNLC), which consists of weak category priority, dynamic monitoring (DM), and historical choice (HC). With these methods, we can adaptively correct the noise labels of different categories according to their specific learning situations. Finally, we provide a comprehensive process for noise simulation, accuracy evaluation, and comparisons in airborne LiDAR point cloud learning from noisy labels. We conduct experiments on the ISPRS 3-D Labeling Vaihingen and Large-scale ALS data for Semantic Labeling in Dense Urban Areas (LASDU) datasets, and the results show that our ADNLC outperforms baseline methods by 30% and 16%, respectively, verifying the superiority of ADNLC and demonstrating the potential of noise labels in LiDAR data processing.
Yuan Gao 0058, Shaobo Xia, Cheng Wang 0016, Xiaohuan Xi, Bisheng Yang, Chou Xie
IEEE Trans. Geosci. Remote. Sens.2
2024 HyperMamba: A Spectral-Spatial Adaptive Mamba for Hyperspectral Image Classification
abstract
Transformers have significantly advanced hyperspectral image (HSI) classification through their proficiency in modeling long sequences. However, the high dimensionality of HSIs poses a particular challenge for Transformers due to their quadratic computational complexity. In natural language processing, state-space models (SSMs) such as Mamba hold great promise for handling long sequence tasks with significantly reduced computational overhead. However, the original Mamba lacks consideration for the spectral and spatial information inherent in HSIs. Inspired by this, we propose the HyperMamba, a novel spectral-spatial adaptive Mamba for HSI classification. The core idea of HyperMamba involves adaptively scanning spatial neighborhood pixels and dynamically enhancing spectral bands for spectral scanning based on acquired spatial neighborhood information. Specifically, HyperMamba consists of two core modules: the spatial neighborhood adaptive scanning (SNAS) module and the spectral adaptive enhancement scanning (SAES) module. Initially, the SNAS module analyzes the spectral characteristics of classified pixels, adaptively selecting the optimal neighborhood for spatial scanning by balancing spatial neighborhood information and local spatial structure. Subsequently, the SAES module dynamically enhances the spectral features of classified pixels using neighborhood spectral information and conducts spectral scanning. Finally, the spectral features of the target pixels are fed into a single fully connected layer classifier, achieving high-precision HSI classification. Extensive experiments demonstrate the effectiveness of HyperMamba, surpassing state-of-the-art methods across three widely used HSI datasets. The code will be available athttps://github.com/chiangliu/HyperMamba.
Jun Yue 0004, Shaobo Xia, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.4
2024 Diffusion Models Meet Remote Sensing: Principles, Methods, and Perspectives
abstract
As a newly emerging advance in deep generative models, diffusion models have achieved state-of-the-art results in many fields, including computer vision, natural language processing, and molecule design. The remote sensing (RS) community has also noticed the powerful ability of diffusion models and quickly applied them to a variety of tasks for image processing. Given the rapid increase in research on diffusion models in the field of RS, it is necessary to conduct a comprehensive review of existing diffusion model-based RS papers, to help researchers recognize the potential of diffusion models and provide some directions for further exploration. Specifically, this article first introduces the theoretical background of diffusion models, and then systematically reviews the applications of diffusion models in RS, including image generation, enhancement, and interpretation. Finally, the limitations of existing RS diffusion models and worthy research directions for further exploration are discussed and summarized.
Yidan Liu, Jun Yue 0004, Shaobo Xia, Pedram Ghamisi, Weiying Xie, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.3
2024 HyperMLL: Toward Robust Hyperspectral Image Classification With Multisource Label Learning
Xia Yue, Anfeng Liu, Shaobo Xia, Jun Yue 0004, Leyuan Fang
IEEE Trans. Geosci. Remote. Sens.4
2023 SpectralDiff: A Generative Framework for Hyperspectral Image Classification With Diffusion Models
abstract
Hyperspectral Image (HSI) classification is an important issue in remote sensing field with extensive applications in earth science. In recent years, a large number of deep learning-based HSI classification methods have been proposed. However, existing methods have limited ability to handle high-dimensional, highly redundant, and complex data, making it challenging to capture the spectral-spatial distributions of data and relationships between samples. To address this issue, we propose a generative framework for HSI classification with diffusion models (SpectralDiff) that effectively mines the distribution information of high-dimensional and highly redundant data by iteratively denoising and explicitly constructing the data generation process, thus better reflecting the relationships between samples. The framework consists of a spectral-spatial diffusion module, and an attention-based classification module. The spectral-spatial diffusion module adopts forward and reverse spectral-spatial diffusion processes to achieve adaptive construction of sample relationships without requiring prior knowledge of graphical structure or neighborhood information. It captures spectral-spatial distribution and contextual information of objects in HSI and mines unsupervised spectral-spatial diffusion features within the reverse diffusion process. Finally, these features are fed into the attention-based classification module for per-pixel classification. The diffusion features can facilitate cross-sample perception via reconstruction distribution, leading to improved classification performance. Experiments on three public HSI datasets demonstrate that the proposed method can achieve better performance than state-of-the-art methods. For the sake of reproducibility, the source code of SpectralDiff will be publicly available at https://github.com/chenning0115/SpectralDiff.
Jun Yue 0004, Leyuan Fang, Shaobo Xia
IEEE Trans. Geosci. Remote. Sens.4
2023 Dif-Fusion: Toward High Color Fidelity in Infrared and Visible Image Fusion With Diffusion Models
abstract
Color plays an important role in human visual perception, reflecting the spectrum of objects. However, the existing infrared and visible image fusion methods rarely explore how to handle multi-spectral/channel data directly and achieve high color fidelity. This paper addresses the above issue by proposing a novel method with diffusion models, termed as Dif-Fusion, to generate the distribution of the multi-channel input data, which increases the ability of multi-source information aggregation and the fidelity of colors. In specific, instead of converting multi-channel images into single-channel data in existing fusion methods, we create the multi-channel data distribution with a denoising network in a latent space with forward and reverse diffusion process. Then, we use the the denoising network to extract the multi-channel diffusion features with both visible and infrared information. Finally, we feed the multi-channel diffusion features to the multi-channel fusion module to directly generate the three-channel fused image. To retain the texture and intensity information, we propose multi-channel gradient loss and intensity loss. Along with the current evaluation metrics for measuring texture and intensity fidelity, we introduce Delta E as a new evaluation metric to quantify color fidelity. Extensive experiments indicate that our method is more effective than other state-of-the-art image fusion methods, especially in color fidelity. The source code is available at https://github.com/GeoVectorMatrix/Dif-Fusion.
Jun Yue 0004, Leyuan Fang, Shaobo Xia, Yue Deng 0001, Jiayi Ma 0001
IEEE Trans. Image Process.3
2022 A Gap-Based Method for LiDAR Point Cloud Division
abstract
As many LiDAR point cloud processing steps, such as reconstruction, are often time- and memory-consuming, dividing LiDAR point clouds into subregions is common and necessary during preprocessing. However, the existing data dividing methods rely on tedious manual work or regular grids and result in oversegmentation around cutting lines. In this letter, we propose a new gap-based data dividing method for various LiDAR point clouds that can minimize the intersections between cutting lines and objects. The basic idea is to find a set of optimal paths that consist of gaps between objects as potential cutting lines. The experiments and comparisons in three data sets demonstrate that the proposed method is much better than the baseline method in terms visual inspection and cutting line quality.
Shaobo Xia, Sheng Nie, Dong Chen 0009, Sheng Xu 0003, Cheng Wang 0016
IEEE Geosci. Remote. Sens. Lett.1
2022 Uncertainty Quantification of Hyperspectral Image Denoising Frameworks Based on Sliding-Window Low-Rank Matrix Approximation
abstract
Sliding-window-based low-rank matrix approximation (LRMA) is a technique widely used in hyperspectral images (HSIs) denoising or completion. However, the uncertainty quantification of the restored HSI has not been addressed to date. Accurate uncertainty quantification of the denoised HSI facilitates applications such as multisource or multiscale data fusion, data assimilation, and product uncertainty quantification since these applications require an accurate approach to describe the statistical distributions of the input data. Therefore, we propose a prior-free closed-form element-wise uncertainty quantification method for LRMA-based HSI restoration. Our closed-form algorithm overcomes the difficulty of handling uncertainty in HSI patch mixing caused by the sliding-window strategy used in the conventional LRMA process. The proposed approach only requires the uncertainty of the observed HSI and provides the uncertainty result relatively rapidly and with similar computational complexity as the LRMA technique. We conduct extensive experiments to validate the estimation accuracy of the proposed closed-form uncertainty approach. The method is robust to at least 10% random impulse noise at the cost of 10%–20% of additional processing time compared to the LRMA. The experiments indicate that the proposed closed-form uncertainty quantification method is more applicable to real-world applications than the baseline Monte Carlo test, which is computationally expensive.
Jingwei Song, Shaobo Xia, Jun Wang 0129, Dong Chen 0009
IEEE Trans. Geosci. Remote. Sens.2
2022 Building Instance Mapping From ALS Point Clouds Aided by Polygonal Maps
abstract
Building region extraction from ALS point clouds has been widely studied, whereas instance-level building mapping has been overlooked and remains unsolved. In this study, we present a method to extract individual buildings from ALS point clouds with the help of widely accessible polygonal footprints. The key idea is to merge roof segments to a set of building candidates, from which correct instances are selected by finding optimal matches between polygonal footprints and building candidates. The method has three steps: roof segmentation, building candidate generation, and instance-polygon matching. The method is tested on two large-scale scenes of different building types and can generally achieve high instance-level building mapping accuracy (around 90%) when there are large positioning errors (6.0 m) among polygons. Future work will focus on classification errors in preprocessing, shape inconsistency between point clouds and polygons, and building footprint delineation and updating in postprocessing.
Shaobo Xia, Sheng Xu 0003, Ruisheng Wang 0001, Jonathan Li 0001, Guanghui Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Tunnel Reconstruction With Block Level Precision by Combining Data-Driven Segmentation and Model-Driven Assembly
abstract
Metro subway systems with underground tunnels form the backbone of urban transportations and therefore, accurate monitoring and maintenance of such subway systems are extremely necessary for a hassle-free daily commutation of billions of people. Though 3-D models of tunnels are widely used for the deformation monitoring of such subway tunnels, existing model-based tunnel monitoring systems rely on coarse geometric models and hence fail to capture complete tunnel health information. We present a two-stage algorithm to create high-fidelity geometric models of tunnel lining from Terrestrial Laser Scanning (TLS) point clouds. Tunnel geometry, defined at the detailed block entity level, is constructed through a data-driven block segmentation algorithm and a model-driven assembly technique. In our approach, the 3-D tunnel block segmentation problem has been translated into a bolt and lining joint recognition problem from 2-D images unfolded from the 3-D scans. The segmented 3-D blocks are matched with a set of predefined 3-D templates from a primitive library via a constraint total least squares matching method and the matched 3-D templates are assembled to create the final watertight tunnel model. The proposed tunnel modeling method has been comprehensively evaluated on Changzhou, Nanjing, and Wuhan tunnel data sets in terms of outliers, missing data, point density, topological representation, robustness, and geometric accuracy. The experiments on Nanjing and Changzhou metro tunnels show that the geometric model fitting incurs an error of only 7 mm, which is almost consistent with a mean density of 6 mm of these two data sets. Experimental results validate the advantages and potentials of the proposed tunnel modeling method.
Dong Chen 0009, Jiju Poovvancheri, Zhenxin Zhang, Shaobo Xia, Liqiang Zhang 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 Curved Buildings Reconstruction From Airborne LiDAR Data by Matching and Deforming Geometric Primitives
abstract
Airborne light detection and ranging (LiDAR) data are widely applied in building reconstruction, with studies reporting success in typical buildings. However, the reconstruction of curved buildings remains an open research problem. To this end, we propose a new framework for curved building reconstruction via assembling and deforming geometric primitives. The input LiDAR point clouds are first converted into contours where individual buildings are identified. After recognizing geometric units (primitives) from building contours, we get initial models by matching the basic geometric primitives to these primitives. To polish assembly models, we employ a warping field for model refinements. Specifically, an embedded deformation (ED) graph is constructed via downsampling the initial model. Then, the point to model displacements is minimized by adjusting node parameters in the ED graph based on our objective function. The presented framework is validated on several highly curved buildings collected by various LiDAR in different cities. The experimental results, as well as accuracy comparison, demonstrate the advantage and effectiveness of our method. The new insight attributes to an efficient reconstruction manner. Moreover, we prove that the primitive-based framework significantly reduces the data storage to 10%-20% of classical mesh models.
Jingwei Song, Shaobo Xia, Jun Wang 0129, Dong Chen 0009
IEEE Trans. Geosci. Remote. Sens.2
2019 Semiautomatic Construction of 2-D Façade Footprints From Mobile LiDAR Data
abstract
Although mobile light detection and ranging (LiDAR) technology has excellent potential in mapping street scenes, there is little research in constructing façade footprints from unorganized, uneven, and incomplete mobile LiDAR point clouds. In fact, façade footprint vectorization from mobile LiDAR data still involves a lot of manual work, especially in complex street scenes with various types of buildings. In this paper, we present a new and effective framework for extracting 2-D façade footprints from mobile LiDAR point clouds. The proposed framework consists of three steps: 1) line segment extraction from projected point clouds based on a hypotheses and selection strategy; 2) completion of missing parts between adjacent walls using line intersections; and 3) delineation of footprints through finding the least cost path in the graph of the line segments. We compare our method with several existing ones and discuss its robustness against data missing and noise such as nonwall structures and vegetation. Our proposed method is also tested in two large-scale data sets, a residential data set, and an urban data set. The coverage ratio, i.e., the percentage of outer wall points covered by the generated outlines in the residential data set is 93.4% and 91.7% in the urban data set is achieved. The mean distance between points of ground truth and constructed footprints for the residential data set and urban data set is 0.019 and 0.028 m, respectively. The experimental results demonstrate that the proposed framework is effective in modeling various façade footprints from mobile LiDAR point clouds.
Shaobo Xia, Ruisheng Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 A Fast Edge Extraction Method for Mobile Lidar Point Clouds
abstract
Edges in mobile light detection and ranging (lidar) point clouds are important for many applications but usually overlooked. In this letter, we propose a fast edge extraction method for mobile lidar. First, an edge index based on geometric center is introduced and then gradients in unorganized 3-D point clouds are defined. By analyzing the ratio between eigenvalues, edge candidates can be detected. Finally, an edge linking algorithm named graph snapping is proposed. The method is tested extensively and the experimental results demonstrate that the proposed method is able to quickly extract most of 3-D edges with higher accuracy than the existing methods.
Shaobo Xia, Ruisheng Wang 0001
IEEE Geosci. Remote. Sens. Lett.1