Guozheng Xu

dblp:67/3018 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 A Coarse-to-Fine Boundary Relabeling Approach for Roof Plane Segmentation
abstract
Building roof plane segmentation is important for three-dimensional (3D) building model reconstruction from airborne light detection and ranging (LiDAR) point data. During the roof plane segmentation, challenges such as pseudo planes, over- and under-segmentation often arise, particularly evident in the boundary regions. To improve the accuracy of plane segmentation, various energy optimization-based methods have been proposed to refine the roof planes. However, the existing methods optimize the energy function at the point level, which may lead to getting stuck in local optima. To address these problems, we propose a coarse-to-fine boundary relabeling approach for roof plane segmentation. Starting from an initial plane segmentation result, the proposed method iteratively refines the planes by adjusting the boundaries from the voxel level to the point level. In addition, we also design a new energy function that considers accurateness, smoothness and compactness to guide the optimization. The experimental results constructed on two datasets demonstrate that the proposed method outperforms the existing roof plane segmentation methods, achieving high accuracy and smooth boundary extraction. The source code of the proposed approach will be publicly available at https://github.com/Li-Li-Whu/Coarse2FineRoofPlane.
Guozheng Xu, Siyuan You, Li Li 0047, Jian Yao 0002
IEEE Geosci. Remote. Sens. Lett.1
2025 A Parallelizable Global Color Consistency Optimization Algorithm for Multiple Images
abstract
The global optimization-based color correction approach aims to minimize the color differences of multiple images by optimizing the correction model for each image. The color differences in multisource and multitemporal remote sensing images are difficult to express using a simple correction model with few parameters. When employing a more flexible correction model, the number of correction parameters and optimization equations grows rapidly with the increase in the number and resolution of input images. In addition, the correction parameters of all images are coupled together and need to be solved simultaneously. An excessive number of parameters results in solving slowly or potential failure. To solve this problem, we propose a parallelizable color correction approach that decouples the correlation of correction parameters in the optimization equations and optimizes each image separately. First, we introduce auxiliary variables that replace values related to other images in the cost function. Second, we construct optimization equations for each image and parallelly solve the correction parameters. Finally, we correct the input images through a weighted correction model to better eliminate correction artifacts. Our approach iteratively optimizes auxiliary variables and correction parameters until the correction results converge. The experimental results on several challenging datasets show that our approach significantly improves execution efficiency and obtains the global optimal solution using the flexible correction model.
Hongche Yin, Pengwei Zhou, Guozheng Xu, Gaoming He, Li Li 0047, Jian Yao 0002
IEEE Geosci. Remote. Sens. Lett.3
2025 A Transformer-Based Roof Plane Segmentation Approach for Airborne LiDAR Point Clouds
abstract
In the fields of photogrammetry and computer vision, three-dimensional (3D) urban building model reconstruction from airborne Light Detection and Ranging (LiDAR) point clouds has attracted significant attention in recent years. Accurately and automatically extracting local geometric structures, such as planar patches, from 3D point cloud data directly determines the quality of subsequent 3D model reconstruction. Considering that the roof is a crucial component of a real building, roof plane segmentation is a critical procedure in building 3D reconstruction. In this paper, a novel dual-branch transformer-based network is designed to accurately segment roof planes from airborne LiDAR point clouds. We first use PointNet++ followed with a transformer encoder to extract point-wise feature embeddings. Then, in the first branch, a transformer decoder module is applied to directly learn the instance centers of planar patches by giving a set of learned queries. Because the transformer can effectively model the relations of the queries and the global context information, the instance center positions of all planes included in the input point clouds can be accurately predicted. In this way, the number and center positions of roof planes are known before performing roof plane segmentation. In the second branch, we predict the offsets for each point using its point-wise feature to shift it towards the corresponding instance center. After that, the plane parameters for each plane instance can be estimated using the shifted points around the predicted centers, and the rest of points are assigned to its nearest plane to generate the final roof planes. The experimental results illustrate that our approach can successfully address the plane segmentation challenge for diverse building roof structures while achieving performance superior to the current state-of-the-art techniques. We will make the source code of our approach publicly available at https://github.com/Li-Li-Whu/PlaneTransformer.
Siyuan You, Guozheng Xu, Pengwei Zhou, Yubing Wei, Jian Yao 0002, Li Li 0047
IEEE Trans. Geosci. Remote. Sens.2
2024 A Novel ITU-Net for GPR Image Clutter Removing
abstract
GPR clutter removal significantly benefits subsequent target recognition, detection, and imaging, enhancing the subsequent processing quality. Traditional clutter removal approaches can only remove noise in simple environments. To solve this problem, We proposed a novel improved triplet attention u-net(ITU-Net) which focuses on the hyperbolic feature w e need while disregarding irrelevant ground clutter and other background noise. The ITU-Net network enhances image reconstruction capability and facilitates rapid image processing. The improved triplet attention module captures cross-domain interaction between any two domains between H, W, and C and considers long-distance dependencies separately in H, W, and C. The experimental results demonstrate that we can effectively retain the information of the hyperbolas while eliminating noise in complex environments.
Mengyang Shi, Guozheng Xu, Yesheng Gao, Xingzhao Liu
IGARSS3
2024 Speckle-Based Residual Optronic Convolutional Neural Network for SAR Target Recognition in Scattering Imaging Scenarios
abstract
Scattering imaging is a pervasive scenario in many areas, especially challenging the performance of remote sensing and automatic target recognition (ATR). Recently, deep learning was utilized for synthetic aperture radar (SAR) ATR in scattering scenarios by extracting the feature of speckle patterns. However, huge computational costs and power consumption challenge its development. Here, we develop a speckle-based residual optronic convolutional neural network (S-ROPCNN) for SAR target recognition. Specifically, we model the light scattering scenarios and build the optical imaging system to produce the speckle patterns for network training. The S-ROPCNN performs SAR target recognition in optical platforms with the speed of light, low computational cost, and low energy consumption. Experiments on the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset demonstrate the feasibility of S-ROPCNN for SAR target recognition in scattering imaging scenarios.
Fengyuan Hu, Guozheng Xu, Mengyang Shi, Yesheng Gao
IGARSS4
2024 Transformer-Based Incomplete Multi-Modal Learning for Land Cover Classification
abstract
Land cover (LC) classification via remote sensing is crucial for ecosystem monitoring and urban planning but faces the challenge of inconsistent multimodal data availability. Current techniques often falter with incomplete modalities, resulting in reduced performance and adaptability. Addressing these issues, this study propose the Transformer-based Incomplete Multi-Modal Learning (TIMML) framework. TIMML incorporates a Bernoulli indicator module during training to facilitate adaptation to missing modalities. This module, in tandem with a fusion token, is instrumental in enabling the model to handle the random omission of modalities by selectively nullifying data streams and effectively aggregates information from the remaining available modalities. Moreover, TIMML integrates a modality-aware regularization module designed to enhance the stability of the feature extraction process, especially when perturbed by the Bernoulli indicator during training. Our comprehensive experiments demonstrate that TIMML not only proficiently manages the challenge of missing modalities but also outperforms existing methods in LC classification tasks, marking a significant advancement in the field.
Guozheng Xu, Xue Jiang 0001, Yue Zhou 0005, Xingzhao Liu
IGARSS1
2024 A fusiform network of indoor scene classification with the stylized semantic description for service-robot applications
Bo Zhu 0009, Junzhe Xie, Guozheng Xu
Expert Syst. Appl.4
2024 Semi-Supervised Scene Classification for Optical Remote Sensing Images via Label and Embedding Consistency
abstract
The utilization of unlabeled samples has contributed significantly to the achievements of semi-supervised methods in optical remote sensing image (ORSI) scene classification. However, existing methods face the challenge of effectively integrating labeled and unlabeled data during model training. To mitigate these challenges, a semi-supervised label and embedding consistency network (SS-LEC) is proposed for OSRI scene classification. Specifically, given an image, SS-LEC enables the high-confidence prediction from a weak-augmentation view consistent with the prediction from a strong-augmentation view, while also ensuring consistency in embeddings derived from middle-augmentation views. Moreover, a soft learning schedule is proposed to strategically focus on varied consistency tasks at different stages of training. Our experiments on two ORSI datasets showcase SS-LEC’s superior classification performance over existing semi-supervised methods. Notably, under label-scarce scenarios with only four labeled images per category, SS-LEC achieves classification accuracies of 92.04% on the EuroSAT dataset and 70.19% on the NWPU-RESISC45 dataset. These results set new benchmarks and demonstrating superior classification performance in challenging conditions with limited labeled data.
Guozheng Xu, Xue Jiang 0001, Yue Zhou 0005, Xingzhao Liu
IEEE Geosci. Remote. Sens. Lett.1
2024 Robust Land Cover Classification With Multimodal Knowledge Distillation
abstract
In recent years, enormous studies have been conducted to improve the land cover (LC) classification performance of multimodal remote sensing (RS) data, which outperforms single-modal-based methods by a large margin due to information diversity. To go a step further, we develop a two-branch patch-based convolutional neural network (CNN) with an encoder–decoder (ED) module to fuse multimodal RS data information. A knowledge distillation in model (DIM) module is proposed to guild per-modality encoder learning with the final fused information to enable multimodal data fusion more effectively. Moreover, utilizing multimodal information to guide single-modal learning still remains to be explored. To this end, a knowledge distillation cross-model (DCM) module is designed to improve single-modal LC classification with multimodal knowledge distillation, which bridges the gap between single-modal-based and multimodal-based methods. In particular, the multimodal-based method is taken as a teacher to transfer knowledge to single-modal-based methods. Extensive experiments are carried out on two multimodal RS datasets, including hyperspectral (HS) and light detection and ranging (LiDAR) data, i.e., the Houston2013 dataset, and HS and synthetic aperture radar (SAR) data, i.e., the Berlin dataset. The results demonstrate the effectiveness and superiority of the proposed multimodal fusion strategy in comparison with several state-of-the-art multimodal RS data classification methods. Also, the proposed DCM module improves the LC classification performance of single-modal methods by a large margin.
Guozheng Xu, Xue Jiang 0001, Yue Zhou 0005, Shutao Li 0001, Xingzhao Liu, Peiwen Lin
IEEE Trans. Geosci. Remote. Sens.1
2024 DGA: Direction-Guided Attack Against Optical Aerial Detection in Camera Shooting Direction-Agnostic Scenarios
abstract
Patch-based adversarial attacks have increasingly aroused concerns due to their application potential in military and civilian fields. In aerial imagery, numerous targets exhibit inherent directionality, such as vehicles and ships, giving rise to the emergence of oriented object detection tasks; similarly, adversarial patches also exhibit intrinsic orientation due to their lack of perfect symmetry. Existing methods presuppose a static alignment between the adversarial patch’s orientation and the camera’s coordinate system – an assumption that is frequently violated in aerial images, whose effectiveness degrades in real-world scenarios. In this paper, we investigate the often-neglected aspect of patch orientation in adversarial attacks and its impact on camouflage effectiveness, particularly when the orientation is not congruent with the target. A new Directional Guided Attack (DGA) framework is proposed for deceiving real-world aerial detectors, which shows robust and adaptable attack performance in camera shooting direction agnostic (CSDA) scenarios. The core idea of DGA is to utilize affine transformations to constrain the relative orientation of the patch to the target and introduce three types of loss to reduce target detection confidence, make the color printable, and smooth the patch color. We introduce a direction-guided evaluation methodology to bridge the gap between patch performance in the digital domain and its actual real-world efficacy. Moreover, we establish a drone-based vehicle detection dataset (SJTU-4K), which labels the orientation of the target, to assess the robustness of patches under various shooting altitudes and views. Extensive proportionally scaled and 1:1 experiments are performed in physical scenarios, demonstrating the superiority and potential of the proposed framework for real-world attacks.
Yue Zhou 0005, Shu-Qi Sun, Xue Jiang 0001, Guozheng Xu, Fengyuan Hu, Xingzhao Liu
IEEE Trans. Geosci. Remote. Sens.4
2023 Color-Aware Self-Supervised Learning for Scene Classification and Segmentation of Remote Sensing Images
abstract
Recently, fully supervised deep learning has achieved excellent success in remote sensing (RS) scene classification and segmentation. However, supervised learning requires tremendous labels, which are difficult to obtain in the field of RS. Self-supervised contrastive methods alleviate this problem by learning impressive transferable representations invariant to different data augmentations, e.g. color jittering. Such invariance could be harmful to RS scene classification and segmentation, which is sensitive to color changes. Therefore, we introduce a color-aware self-supervised learning framework (ColorSelf) for RS scene classification and segmentation. Our model encourages to preserve color-aware information in learned representation to improve their transferability. Extensive experiments on two challenging RS datasets demonstrate the proposed ColorSelf brings a significant performance improvement in both RS scene classification and segmentation task.
Guozheng Xu, Xue Jiang 0001, Xingzhao Liu
IGARSS1
2021 Self-Supervised Disentangled Embedding For Robust Image Classification
Lanqing Liu, Zhenyu Duan, Guozheng Xu, Yi Xu 0001
ICIP3
2021 Reinventing 2D Convolutions for 3D Images
abstract
There have been considerable debates over 2D and 3D representation learning on 3D medical images. 2D approaches could benefit from large-scale 2D pretraining, whereas they are generally weak in capturing large 3D contexts. 3D approaches are natively strong in 3D contexts, however few publicly available 3D medical dataset is large and diverse enough for universal 3D pretraining. Even for hybrid (2D + 3D) approaches, the intrinsic disadvantages within the 2D/3D parts still exist. In this study, we bridge the gap between 2D and 3D convolutions by reinventing the 2D convolutions. We propose ACS (axial-coronal-sagittal) convolutions to perform natively 3D representation learning, while utilizing the pretrained weights on 2D datasets. In ACS convolutions, 2D convolution kernels are split by channel into three parts, and convoluted separately on the three views (axial, coronal and sagittal) of 3D representations. Theoretically, ANY 2D CNN (ResNet, DenseNet, or DeepLab) is able to be converted into a 3D ACS CNN, with pretrained weight of a same parameter size. Extensive experiments validate the consistent superiority of the pretrained ACS CNNs, over the 2D/3D CNN counterparts with/without pretraining. Even without pretraining, the ACS convolution can be used as a plug-and-play replacement of standard 3D convolution, with smaller model size and less computation.
Jiancheng Yang, Jingwei Xu 0005, Canqian Yang, Guozheng Xu, Bingbing Ni
IEEE J. Biomed. Health Informatics6
2018 Continuous Shared Control for Robotic Arm Reaching Driven by a Hybrid Gaze-Brain Machine Interface
abstract
The brain-machine interface (BMI) has been reported to offer the potential for controlling the assistive robot for the motor impaired people, using the non-invasively obtained electroencephalogram (EEG) signals. However, the EEG based BMI may not be sufficient and stable to drive the robot moving freely in its 2D or 3D workspace. The robot autonomy may provide assistance for the BMI users with the shared control paradigm. Nevertheless, users suffers from several limitations of the current shared control paradigms applied on BMI, e.g., loss of sense of control, high mental workload due to unintuitive control with the human-robot interface and fixed level of assistance. To overcome these drawbacks, we propose a new control paradigm for the robotic arm reaching task where the robot autonomy is dynamically blended with the gaze-BMI control from a user. In this paradigm, the hybrid gaze-BMI constitutes an intuitive and effective input to continuously control the robotic arm end-effector moving freely in its 2D workspace, with an adjustable speed proportional to the motion intention strength. Furthermore, the adjustable level of assistance by our paradigm allows the system to balance the user's capabilities and feelings of control while compensating for the reaching task's difficulty. The proposed paradigm is verified in the task where a healthy subject utilizes the hybrid gaze-BMI to control the robotic arm end-effector reaching for a target object while avoiding the obstacle in the path. The experimental results demonstrate that the movements with our shared control paradigm are safer, more efficient and less difficult than those without shared control.
Guozheng Xu, Aiguo Song, Baoguo Xu, Hong Zeng 0001
IROS2
2017 Integration of semantic and visual hashing for image retrieval
Songhao Zhu, Dongliang Jin, Yajie Sun, Guozheng Xu
J. Vis. Commun. Image Represent.6