Tong Wang 0017

dblp:51/6856-17 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 FACTS: Training-free zero-shot diffusion framework for facade texture restoration in 3D urban models
Juexiao Cheng, Xiangru Huang, Guanzhou Chen 0001, Tong Wang 0017, Jiaqi Wang 0015, Xiaoliang Tan, Aiyi Jiang, Xiaodong Zhang 0027
Adv. Eng. Informatics4
2026 From roof structure to 3D building models: A RoofStructNet-Based framework for roof extraction towards multi-level 3D reconstruction
Jiaqi Wang 0015, Xiaodong Zhang 0027, Guanzhou Chen 0001, Tong Wang 0017, Xiaoliang Tan, Wenchao Guo, Kun Zhu 0003
Expert Syst. Appl.4
2025 BFA-YOLO: A balanced multiscale object detection network for building façade elements detection
Yangguang Chen, Tong Wang 0017, Guanzhou Chen 0001, Kun Zhu 0003, Xiaoliang Tan, Jiaqi Wang 0015, Wenchao Guo, Qing Wang 0054, Xiaolong Luo, Xiaodong Zhang 0027
Adv. Eng. Informatics2
2025 INVITATION: A Framework for Enhancing UAV Image Semantic Segmentation Accuracy Through Depth Information Fusion
abstract
With the increasing use of uncrewed aerial vehicles (UAVs), improving the accuracy of semantic segmentation is becoming critical. Depth information preserves geometric structure, serving as an invaluable supplement to color-rich UAV imagery. Inspired by this, we proposed a novel framework named INVITATION, which exclusively takes original UAV imagery as input, yet is capable of obtaining complemented depth information and fusing into RGB semantic segmentation models effectively, thereby enhancing UAV semantic segmentation accuracy. Concretely, this framework supports two distinct depth generation approaches: high-precision multiview stereo (MVS) depth reconstruction using multiple views or video sequences via structure from motion (SfM) and monocular depth estimation using individual images. Our empirical evaluations conducted on the UAVid dataset showed that mIoU metric of INVITATION used precise reconstructed depth maps via MVS improved from 66.02% to 70.57%, while used depth predictions from pretrained models reached 69.69%, which supports the effectiveness of extracting and fusing depth information from original imagery in enhancing UAV semantic segmentation. This study explores a novel approach to acquire UAV multimodal information at low data cost, highlights the advantages of incorporating depth information into UAV semantic analysis, and paves the way for further studies on the integration of multimodal UAV information. Our code is available athttps://github.com/CVEO/INVITATION.
Xiaodong Zhang 0027, Wenlin Zhou, Guanzhou Chen 0001, Jiaqi Wang 0015, Xiaoliang Tan, Tong Wang 0017
IEEE Geosci. Remote. Sens. Lett.7
2025 LMFNet: Lightweight Multimodal Fusion Network for high-resolution remote sensing image segmentation
Tong Wang 0017, Guanzhou Chen 0001, Xiaodong Zhang 0027, Jiaqi Wang 0015, Xiaoliang Tan, Wenlin Zhou, Chanjuan He
Pattern Recognit.1
2024 Segment Change Model (SCM) for Unsupervised Change Detection in VHR Remote Sensing Images: A Case Study of Buildings
abstract
The field of Remote Sensing (RS) widely employs Change Detection (CD) on very-high-resolution (VHR) images. A majority of extant deep-learning-based methods hinge on annotated samples to complete the CD process. Recently, the emergence of Vision Foundation Model (VFM) enables zeroshot predictions in particular vision tasks. In this work, we propose an unsupervised CD method named Segment Change Model (SCM), built upon the Segment Anything Model (SAM) and Contrastive Language-Image Pre-training (CLIP). Our method recalibrates features extracted at different scales and integrates them in a top-down manner to enhance discriminative change edges. We further design an innovative Piecewise Semantic Attention (PSA) scheme, which can offer semantic representation without training, thereby minimize pseudo change phenomenon. Through conducting experiments on two public datasets, the proposed SCM increases the mIoU from 46.09% to 53.67% on the LEVIR-CD dataset, and from 47.56% to 52.14% on the WHU-CD dataset. Our codes are available at: https://github.com/StephenApX/UCDSCM.
Xiaoliang Tan, Guanzhou Chen 0001, Tong Wang 0017, Jiaqi Wang 0015, Xiaodong Zhang 0027
IGARSS3
2024 S3Net: Innovating Stereo Matching and Semantic Segmentation with a Single-Branch Semantic Stereo Network in Satellite Epipolar Imagery
abstract
Stereo matching and semantic segmentation are significant tasks in binocular satellite 3D reconstruction. However, previous studies primarily view these as independent parallel tasks, lacking an integrated multitask learning framework. This work introduces a solution, the Single-branch Semantic Stereo Network (S3Net), which innovatively combines semantic segmentation and stereo matching using Self-Fuse and Mutual-Fuse modules. Unlike preceding methods that utilize semantic or disparity information independently, our method identifies and leverages the intrinsic link between these two tasks, leading to a more accurate understanding of semantic information and disparity estimation. Comparative testing on the US3D dataset proves the effectiveness of our S3Net. Our model improves the mIoU in semantic segmentation from 61.38 to 67.39, and reduces the D1-Error and average endpoint error (EPE) in disparity estimation from 10.051 to 9.579 and 1.439 to 1.403 respectively, surpassing existing competitive methods. Our codes are available at: https://github.com/CVEO/S3Net.
Guanzhou Chen 0001, Xiaoliang Tan, Tong Wang 0017, Jiaqi Wang 0015, Xiaodong Zhang 0027
IGARSS4
2024 A Deep Transfer Learning Framework Using Teacher-Student Structure for Land Cover Classification of Remote-Sensing Imagery
abstract
Deep learning techniques are widely used for land cover classification in remote sensing, primarily because they can effectively extract complex features from imagery data, which is essential for accurate land cover classification. However, the heterogeneity of data can pose a challenge to the generalizability of deep models. To address this issue, we propose a novel transfer learning framework for land cover classification using teacher-student structure. The proposed framework utilizes the knowledge acquired by the teacher model from large datasets to facilitate fine-tuning of the student model on small datasets, which prevents problems such as overfitting resulting from training large models on small datasets and inferior performance arising from the limited learning capacity of small models. To achieve this goal, we design a loss function base on central moment discrepancy and high-temperature softmax. In addition, we conducted experiments on three distinct pairs of datasets and found that our proposed framework outperforms both training the student model from scratch on the target domain and simply fine-tuning the teacher and student models from the source domain to the target domain. Specifically, we observed an average increase in mIoU of 9.9%, 2.1%, and 4.3%. These results demonstrate the effectiveness and generalizability of our proposed framework for remote sensing land cover classification.
Xiaodong Zhang 0027, Xianwei Li 0003, Guanzhou Chen 0001, Puyun Liao, Tong Wang 0017, Haobo Yang, Chanjuan He, Wenlin Zhou
IEEE Geosci. Remote. Sens. Lett.5
2024 S2Net: A Multitask Learning Network for Semantic Stereo of Satellite Image Pairs
abstract
Stereo matching and semantic segmentation are two significant tasks in remote sensing. Recently, deep learning approaches have been applied to these tasks separately. However, the lack of semantic supervision makes the training of stereo matching model susceptible to data disturbance, resulting in inferior generalization ability; foreground objects are sometimes confused with background pixels in RGB images, limiting the classification accuracy. By exploring the relationship between these two tasks, semantic stereo solves these problems simultaneously with multitask learning. Previous methods took semantic stereo as two parallel processing tasks, so they did not take full advantages of the additional information from both tasks and only obtained slight improvement. In this work, we designed a multitask learning framework semantic stereo network (S2Net). The proposed network generates cost volumes with feature maps supervised by semantic information to estimate disparity maps and fuses RGB-D feature maps to predict classification maps, therefore gathering multitask learning information. To enhance the performance of trained model, we also considered the continuity of disparity values and the duality of stereo image pairs in data augmentation. When applied in datasets without training, S2Net obtained 2.937% D1-Error in the WHU dataset, lower than 4.297% of the previous best method, depicting the generalization ability improvement from semantic supervision. In terms of semantic segmentation, the introduction of disparity maps increases the mean intersection over union (mIoU) from 61.375% to 69.096% in the US3D datasets. The experiments on the KITTI semantics benchmark show that our proposed method obtains 60.76% mIoU, achieving state-of-the-art among multitask learning methods.
Puyun Liao, Xiaodong Zhang 0027, Guanzhou Chen 0001, Tong Wang 0017, Xianwei Li 0003, Haobo Yang, Wenlin Zhou, Chanjuan He, Qing Wang 0054
IEEE Trans. Geosci. Remote. Sens.4
2022 A Superpixel-Guided Unsupervised Fast Semantic Segmentation Method of Remote Sensing Images
abstract
Semantic segmentation is one of the fundamental tasks of pixel-level remote sensing image analysis. Currently, most high-performance semantic segmentation methods are trained in a supervised learning manner. These methods require a large number of image labels as support, but manual annotations are difficult to obtain. To address the problem, we propose an efficient unsupervised remote sensing image segmentation method based on superpixel segmentation and fully convolutional networks (FCNs) in this letter. Our method can achieve pixel-level images segmentation of various scales rapidly without any manual labels or prior knowledge. We use the superpixel segmentation results as synthetic ground truth to guide the gradient descent direction during FCN training. In experiments, our method achieved high performance compared to current unsupervised image segmentation methods on three public datasets. Specifically, our method achieves an Adjusted Mutual Information (AMI) score of 0.2955 on the Gaofen Image Dataset (GID) dataset, while processing each image of size 7200 × 6800 pixels in just 30 seconds.
Guanzhou Chen 0001, Chanjuan He, Tong Wang 0017, Kun Zhu 0003, Puyun Liao, Xiaodong Zhang 0027
IEEE Geosci. Remote. Sens. Lett.3
2022 Object-Based Classification Framework of Remote Sensing Images With Graph Convolutional Networks
abstract
Object-based image classification (OBIC) on very-high-resolution (VHR) remote sensing (RS) images is utilized in a wide range of applications. Nowadays, many existing OBIC methods only focus on features of each object itself, neglecting the contextual information among adjacent objects and resulting in low classification accuracy. Inspired by a spectral graph theory, we construct a graph structure from objects generated from VHR RS images and propose an OBIC framework based on truncated sparse singular value decomposition and graph convolutional network (GCN) model, aiming to make full use of relativities among objects and produce an accurate classification. Through conducting experiments on two annotated RS image data sets, our framework obtained 97.2% and 66.9% overall accuracy, respectively, in automatic and manual object segmentation circumstances, within a processing time of about 1/100 of convolutional neural network (CNN)-based methods’ training time.
Xiaodong Zhang 0027, Xiaoliang Tan, Guanzhou Chen 0001, Kun Zhu 0003, Puyun Liao, Tong Wang 0017
IEEE Geosci. Remote. Sens. Lett.6
2022 Multi-Oriented Rotation-Equivariant Network for Object Detection on Remote Sensing Images
abstract
Object detection has attracted a lot of attention in the field of image automatic interpretation. Detectors based on convolution neural networks (CNNs) applied in natural scene images encode detection results with horizontal bounding boxes (HBBs), which can not accurately calibrate the position and shape of the arbitrary-orientation objects on remote sensing images (RSIs). To solve these issues, we propose an object detection framework named multi-oriented rotation-equivariant network (MORE-Net) in this letter. The MORE-Net consists of a multi-oriented rotating filter (MORF) and a multi-oriented ground objects detector. The MORF generates rotation-equivariant features by rotating canonical filter according to the predefined discrete directions. Equipped with MORF, the multi-oriented ground objects detector extracts rotation-invariant semantic representations for each decoded rotated region of interest (RRoI). We also propose area smooth (AS) L1 loss to impose tighter shape and position constraints to the RROIs. Extensive experiments and comprehensive evaluations on the large-scale DOTA dataset demonstrate the effectiveness of the proposed framework, which achieves a mean average precision (mAP) value of 81.27 on the DOTA-v1.0 dataset and a mAP value of 78.03 on the DOTA-v1.5 dataset.
Kun Zhu 0003, Xiaodong Zhang 0027, Guanzhou Chen 0001, Xianwei Li 0003, Peihua Cai, Puyun Liao, Tong Wang 0017
IEEE Geosci. Remote. Sens. Lett.7
2020 Convective Clouds Extraction From Himawari-8 Satellite Images Based on Double-Stream Fully Convolutional Networks
abstract
Auto-extraction of convective clouds is of great significance. Convective clouds often bring heavy rain, strong winds, and other disastrous weather. Early warning of convection can effectively reduce loss. Using remote sensing images, we can get large-scale cloud information, which provides many effective methods for convective clouds detection. In this letter, we proposed a novel method to extract convective clouds. We introduce a novel deep network using only$1 \times 1$convolution (3ONet) to extract the spectral characteristics. We then combine a 3ONet with the symmetrical dense-shortcut deep fully convolutional networks (SDFCNs) with a double-stream fully convolutional network to extract convective clouds. In the experiment, we used 12 000 Himawari–8 satellite image patches to verify the proposed framework. Experimental results with 0.5882 mean intersection over union (mIOU) pointed out the proposed method can extract convective clouds effectively.
Xiaodong Zhang 0027, Tong Wang 0017, Guanzhou Chen 0001, Xiaoliang Tan, Kun Zhu 0003
IEEE Geosci. Remote. Sens. Lett.2