Wei Huang 0068

dblp:81/6685-68 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0002-5862-3126ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Semi-Supervised Building Footprint Extraction Using Debiased Pseudo-Labels
abstract
Accurate extraction of building footprints from satellite imagery is of high value. Currently, deep learning methods are predominant in this field due to their powerful representation capabilities. However, they generally require extensive pixel-wise annotations, which constrains their practical application. Semi-supervised learning (SSL) significantly mitigates this requirement by leveraging large volumes of unlabeled data for model self-training (ST), thus enhancing the viability of building footprint extraction. Despite its advantages, SSL faces a critical challenge: the imbalanced distribution between the majority background class and the minority building class, which often results in model bias toward the background during training. To address this issue, this article introduces a novel method called DeBiased matching (DBMatch) for semi-supervised building footprint extraction. DBMatch comprises three main components: 1) a basic supervised learning module (SUP) that uses labeled data for initial model training; 2) a classical weak-to-strong ST module that generates pseudo-labels from unlabeled data for further model ST; and 3) a novel logit debiasing (LDB) module that calculates a global logit bias between building and background, allowing for dynamic pseudo-label calibration. To verify the effectiveness of the proposed DBMatch, extensive experiments are performed on three public building footprint extraction datasets covering six global cities in SSL setting. The experimental results demonstrate that our method significantly outperforms some advanced SSL methods in semi-supervised building footprint extraction. Our codes will be publicly provided athttps://github.com/zhu-xlab/SSL_Buildings.
Wei Huang 0068, Ziqi Gu, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Height-Assisted Semi-Supervised Building Footprint Extraction From Optical Remote Sensing Images
abstract
Automatic building footprint extraction from optical remote sensing (RS) images is popular and crucial for various downstream applications. Current building footprint extraction methods are mainly based on deep learning, which requires large amounts of manually labeled data for model training, limiting their practical deployment. Semi-supervised semantic segmentation (SSS), which leverages limited labeled data for supervised learning and abundant unlabeled data for unsupervised self-training, offers a promising solution to reduce this reliance. Nonetheless, directly applying existing SSS methods to building footprint extraction with limited labels fails to fully exploit the geometric structural features of buildings—key characteristics that distinguish them from background. To tackle this challenge, we propose a semi-supervised learning framework, HeightMatch, which integrates real or synthetic height information with RS images to extract more comprehensive and discriminative feature representations of buildings, particularly in limited-label scenarios. During training, these height maps effectively enhance the model’s ability to capture geometric structures, leading to more accurate pseudo-labels for unlabeled data and thereby enabling more effective self-training. At inference, building predictions rely solely on RS images, ensuring the practicality of the proposed method. Extensive experimental results on five widely-used building footprint extraction datasets demonstrate the effectiveness and superiority of our method in comparison with multiple state-of-the-art SSS methods. Our code is available at https://github.com/zhu-xlab/HeightMatch.
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Embedding Generalized Semantic Knowledge Into Few-Shot Remote Sensing Segmentation
abstract
Few-shot segmentation (FSS) for remote sensing (RS) imagery leverages supporting information from limited annotated samples to achieve query segmentation of novel classes. Previous efforts are dedicated to mining segmentation-guiding visual cues from a constrained set of support samples. However, they still struggle to address the pronounced intra-class differences in RS images, as sparse visual cues make it challenging to establish robust class-specific representations. In this article, we propose a holistic semantic embedding (HSE) approach that effectively harnesses general semantic knowledge, i.e., class description (CD) embeddings. Instead of the naive combination of CD embeddings and visual features for segmentation decoding, we investigate embedding the general semantic knowledge during the feature extraction stage. Specifically, in HSE, a spatial dense interaction (SDI) module allows the interaction of visual support features with CD embeddings along the spatial dimension via self-attention. Furthermore, a global content modulation (GCM) module efficiently augments the global information of the target category in both support and query features, thanks to the transformative fusion of visual features and CD embeddings. These two components holistically synergize CD embeddings and visual cues, constructing a robust class-specific representation. Through extensive experiments on the standard FSS benchmark, the proposed HSE approach demonstrates superior performance compared to peer work, setting a new state-of-the-art.
Qi Wang 0009, Yuyu Jia, Wei Huang 0068, Junyu Gao 0001, Qiang Li 0042
IEEE Trans. Geosci. Remote. Sens.3
2024 Representation Enhancement-Stabilization: Reducing Bias-Variance of Domain Generalization
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
ECCV (36)1
2024 Disentangling Semi-Supervised Semantic Segmentation of Remote Sensing Images
abstract
In Earth observation, semantic understanding of Remote Sensing (RS) images holds significant importance, yet it is hindered in practice by the need for extensive manual pixel-level labeling. Semi-supervised semantic segmentation (SSS) of RS images would be a promising solution, which fully utilizes unlabeled data for model self-training under the guidance of limited labeled data. The mainstream SSS methods use pseudo-labels of the unlabeled data for model training, however, their performance is bottlenecked because of confirmation bias, i.e., stubborn incorrect pseudo-labels. To counter this, our study introduces a novel disentanglement learning (DL) method tailored for RS-SSS. It separates the predictions of the labeled and unlabeled data by two individual prediction heads during current training, and then integrates them during follow-up training. The experimental results verify its effectiveness on two widely-used RS semantic segmentation datasets in semi-supervised setting.
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IGARSS1
2023 GLCM: Global-Local Captioning Model for Remote Sensing Image Captioning
abstract
Remote sensing image captioning (RSIC), which describes a remote sensing image with a semantically related sentence, has been a cross-modal challenge between computer vision and natural language processing. For visual features extracted from remote sensing images, global features provide the complete and comprehensive visual relevance of all the words of a sentence simultaneously, while local features can emphasize the discrimination of these words individually. Therefore, not only global features are important for caption generation but also local features are meaningful for making the words more discriminative. In order to make full use of the advantages of both global and local features, in this article, we propose an attention-based global-local captioning model (GLCM) to obtain global-local visual feature representation for RSIC. Based on the proposed GLCM, the correlation of all the generated words and the relation of each separate word and the most related local visual features can be visualized in a similarity-based manner, which provides more interpretability for RSIC. In the extensive experiments, our method achieves comparable results in UCM-captions and superior results in Sydney-captions and RSICD which is the largest RSIC dataset.
Qi Wang 0009, Wei Huang 0068, Xuelong Li 0001
IEEE Trans. Cybern.2
2023 AdaptMatch: Adaptive Matching for Semisupervised Binary Segmentation of Remote Sensing Images
abstract
There are various binary semantic segmentation tasks in remote sensing (RS) that aim to extract the foreground areas of interest, such as buildings and roads, from the background in satellite images. In particular, semi-supervised learning, which can use limited labeled data to guide a large amount of unlabeled data for model training, can significantly promote the fast applications of these tasks in practice. However, due to the predominance of the background in RS images, the foreground only accounts for a small proportion of the pixels. It poses a challenge: models are biased toward the majority class of the background, leading to poor performance on the minority class of the foreground. To address this issue, this paper proposes a novel and effective semi-supervised learning framework, Adaptive Matching (AdaptMatch), for RS binary segmentation. AdaptMatch calculates individual and adaptive thresholds of the foreground and background based on their convergence difficulty in an online manner at the training stage; the adaptive thresholds are then used to select the high-confidence pseudo-labeled data of the two classes for model self-training in turn. Extensive experiments are conducted on two widely-studied RS binary segmentation tasks, building footprint extraction and road extraction, to demonstrate the effectiveness and generalizability of the proposed method. The results show that the proposed AdaptMatch achieves superior performance compared with some state-of-the-art semi-supervised methods in RS binary segmentation tasks. The codes will be publicly available at https://github.com/zhu-xlab/AdaptMatch.
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Exploring Hard Samples in Multiview for Few-Shot Remote Sensing Scene Classification
abstract
Few-shot remote sensing scene classification is of high practical value in real situations where data are scarce and annotated costly. The few-shot learner needs to identify new categories with limited examples, and the core issue of this assignment is how to prompt the model to learn transferable knowledge from a large-scale base dataset. Although current approaches based on transfer learning or meta-learning have achieved significant performance on this task, there are still two problems to be addressed: (i) as an essential characteristic of remote sensing images, spatial rotation insensitivity surprisingly remains largely unexplored; (ii) the high distribution uncertainty of hard samples reduces the discriminative power of the model decision boundary. Stimulated by these, we propose a corresponding end-to-end framework termed a Hard Sample Learning (HSL) and Multi-view Integration (MI) Network (HSL-MINet). First, the MI module contains a pretext task introduced to guide the knowledge transfer, and a multiview-attention mechanism used to extract correlational information across different rotation views of images. Second, aiming at increasing the discrimination of the model decision boundary, the HSL module is designed to evaluate and select hard samples via a class-wise adaptive threshold strategy, and then decrease the uncertainty of their feature distributions by a devised triplet loss. Extensive evaluations on NWPU-RESISC45, WHU-RS19, and UCM datasets show that the effectiveness of our HSL-MINet surpasses the former state-of-the-art approaches.
Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.3
2023 Holistic Mutual Representation Enhancement for Few-Shot Remote Sensing Segmentation
abstract
Few-shot segmentation endeavors to utilize a minimal amount of annotated samples (support) to guide the segmentation of unseen objects (query). Previous techniques primarily employ asupport-to-queryparadigm, neglecting to sufficiently leverage the mutual representation between query and support images, which leaves models suffering from intra-class variations and background interference in remote sensing images. This paper proposes a Holistic Mutual Representation Enhancement (HMRE) method to bridge these gaps. First, a Dual Activation (DA) module is devised to establish information symmetry between the two branches and forms the foundation for mutual representation enhancement. Subsequently, the holistic mutual enhancement is jointly constructed by the Global Semantic (GS) and Spatial Dense (SD) mutual enhancement modules. In the prediction stage for segmentation, we integrate the enhanced mutual representation into the Mutual-Fusion Decoder to activate the homologous object regions bidirectionally. To expedite the replication of investigation in this task, we further create a corresponding benchmark Flood-3i. The whole dataset is attainable at https://drive.google.com/drive/folders/1FMAKf2sszoFKjq0UrUmSLnJDbwQSpfxR. Extensive experiments on two benchmarks iSAID-5i and Flood-3i demonstrate the superiority of our proposed method, which also sets a new state-of-the-art.
Yuyu Jia, Junyu Gao 0001, Wei Huang 0068, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.3
2023 THE Benchmark: Transferable Representation Learning for Monocular Height Estimation
abstract
Generating 3D city models rapidly is crucial for many applications. Monocular height estimation is one of the most efficient and timely ways to obtain large-scale geometric information. However, existing works focus primarily on training and testing models using unbiased datasets, which does not align well with real-world applications. Therefore, we propose a new benchmark dataset to study the transferability of height estimation models in a cross-dataset setting. To this end, we first design and construct a large-scale benchmark dataset for cross-dataset transfer learning on the height estimation task. This benchmark dataset includes a newly proposed large-scale synthetic dataset, a newly collected real-world dataset, and four existing datasets from different cities. Next, a new experimental protocol,few-shot cross-dataset transfer, is designed. Furthermore, in this paper, we propose a scale-deformable convolution module to enhance the window-based Transformer for handling the scale-variation problem in the height estimation task. Experimental results have demonstrated the effectiveness of the proposed methods in traditional and cross-dataset transfer settings. The datasets and codes are publicly available at https://mediatum.ub.tum.de/1662763 and https://thebenchmarkh.github.io/.
Zhitong Xiong, Wei Huang 0068, Jingtao Hu, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Pseudo Features-Guided Self-Training for Domain Adaptive Semantic Segmentation of Satellite Images
abstract
Semantic segmentation is a fundamental and crucial task that is of great importance to real-world satellite image-based applications. Yet a widely acknowledged issue that occurs when applying the semantic segmentation models to unseen scenery is that the model will perform much poorer than when it was applied to scenery similar to the training data. This phenomenon is usually termed as the domain shift problem. To tackle it, this article presents a self-training-based unsupervised domain adaptation (UDA) method. Different from the previous self-training approaches which focus on rectifying and improving the quality of the pseudo labels, we instead seek to exploit feature-level relation among neighboring pixels to structure and regularize the prediction of the adapted model. Based on the assumption that spatial topological relation is maintained despite the impact of the domain shift, we propose a novel self-training mechanism to perform DA by exploiting local relation in the feature space spanned by the teacher model, from which the pseudo labels are generated. Quantitative experiments on four different public benchmarks demonstrate that the proposed method can outperform the other UDA methods. Besides, analytical experiments also intuitively verify the proposed assumption. Codes will be publicly available athttps://github.com/zhu-xlab/PFST.
Fahong Zhang 0001, Yilei Shi, Zhitong Xiong, Wei Huang 0068, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Looking Closer at the Scene: Multiscale Representation Learning for Remote Sensing Image Scene Classification
abstract
Remote sensing image scene classification has attracted great attention because of its wide applications. Although convolutional neural network (CNN)-based methods for scene classification have achieved excellent results, the large-scale variation of the features and objects in remote sensing images limits the further improvement of the classification performance. To address this issue, we present multiscale representation for scene classification, which is realized by a global-local two-stream architecture. This architecture has two branches of the global stream and local stream, which can individually extract the global features and local features from the whole image and the most important area. In order to locate the most important area in the whole image using only image-level labels, a weakly supervised key area detection strategy of structured key area localization (SKAL) is specially designed to connect the above two streams. To verify the effectiveness of the proposed SKAL-based two-stream architecture, we conduct comparative experiments based on three widely used CNN models, including AlexNet, GoogleNet, and ResNet18, on four public remote sensing image scene classification data sets, and achieve the state-of-the-art results on all the four data sets. Our codes are provided in https://github.com/hw2hwei/SKAL.
Qi Wang 0009, Wei Huang 0068, Zhitong Xiong, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 Denoising-Based Multiscale Feature Fusion for Remote Sensing Image Captioning
abstract
With the benefits from deep learning technology, generating captions for remote sensing images has become achievable, and great progress has been made in this field in the recent years. However, a large-scale variation of remote sensing images, which would lead to errors or omissions in feature extraction, still limits the further improvement of caption quality. To address this problem, we propose a denoising-based multi-scale feature fusion (DMSFF) mechanism for remote sensing image captioning in this letter. The proposed DMSFF mechanism aggregates multiscale features with the denoising operation at the stage of visual feature extraction. It can help the encoder-decoder framework, which is widely used in image captioning, to obtain the denoising multiscale feature representation. In experiments, we apply the proposed DMSFF in the encoder-decoder framework and perform the comparative experiments on two public remote sensing image captioning data sets including UC Merced (UCM)-captions and Sydney-captions. The experimental results demonstrate the effectiveness of our method.
Wei Huang 0068, Qi Wang 0009, Xuelong Li 0001
IEEE Geosci. Remote. Sens. Lett.1
2021 Truncation Cross Entropy Loss for Remote Sensing Image Captioning
abstract
Recently, remote sensing image captioning (RSIC) has drawn an increasing attention. In this field, the encoder-decoder-based methods have become the mainstream due to their excellent performance. In the encoder-decoder framework, the convolutional neural network (CNN) is used to encode a remote sensing image into a semantic feature vector, and a sequence model such as long short-term memory (LSTM) is subsequently adopted to generate a content-related caption based on the feature vector. During the traditional training stage, the probability of the target word at each time step is forcibly optimized to 1 by the cross entropy (CE) loss. However, because of the variability and ambiguity of possible image captions, the target word could be replaced by other words like its synonyms, and therefore, such an optimization strategy would result in the overfitting of the network. In this article, we explore the overfitting phenomenon in the RSIC caused by CE loss and correspondingly propose a new truncation cross entropy (TCE) loss, aiming to alleviate the overfitting problem. In order to verify the effectiveness of the proposed approach, extensive comparison experiments are performed on three public RSIC data sets, including UCM-captions, Sydney-captions, and RSICD. The state-of-the-art result of Sydney-captions and RSICD and the competitive results of UCM-captions achieved by TCE loss demonstrate that the proposed method is beneficial to RSIC.
Xuelong Li 0001, Wei Huang 0068, Qi Wang 0009
IEEE Trans. Geosci. Remote. Sens.3
2021 Word-Sentence Framework for Remote Sensing Image Captioning
abstract
Remote sensing image captioning (RSIC), which aims at generating a well-formed sentence for a remote sensing image, has attracted more attention in recent years. The general framework for RSIC is the encoder–decoder architecture containing two submodels of encoder and decoder. Although the significant performance is obtained, the encoder–decoder architecture is a black-box model with a lack of explainability. To overcome this drawback, in this article, we propose a new explainable word–sentence framework for RSIC. The proposed word–sentence framework consists of two parts: word extractor and sentence generator, where the former extracts the valuable words in the given remote sensing image, while the latter organizes these words into a well-formed sentence. The proposed framework decomposes RSIC into a word classification task and a word sorting task, which is more in line with human intuitive understanding. On the basis of the word–sentence framework, some ablation experiments are conducted on the three public RSIC data sets of Sydney-captions, UCM-captions, and RSICD to explore the specific and effective network structures. In order to evaluate the proposed word–sentence framework objectively, we further conduct some comparative experiments on these three data sets and achieve comparable results in comparison with the encoder–decoder-based methods.
Qi Wang 0009, Wei Huang 0068, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 SSR-NET: Spatial-Spectral Reconstruction Network for Hyperspectral and Multispectral Image Fusion
abstract
The fusion of a low-spatial-resolution hyperspectral image (HSI) (LR-HSI) with its corresponding high-spatial-resolution multispectral image (MSI) (HR-MSI) to reconstruct a high-spatial-resolution HSI (HR-HSI) has been a significant subject in recent years. Nevertheless, it is still difficult to achieve the cross-mode information fusion of spatial mode and spectral mode when reconstructing HR-HSI for the existing methods. In this article, based on a convolutional neural network (CNN), an interpretable spatial-spectral reconstruction network (SSR-NET) is proposed for more efficient HSI and MSI fusion. More specifically, the proposed SSR-NET is a physical straightforward model that consists of three components: 1) cross-mode message inserting (CMMI); this operation can produce the preliminary fused HR-HSI, preserving the most valuable information of LR-HSI and HR-MSI; 2) spatial reconstruction network (SpatRN); the SpatRN concentrates on reconstructing the lost spatial information of LR-HSI with the guidance of spatial edge loss (Lspat); and 3) spectral reconstruction network (SpecRN); the SpecRN pays attention to reconstruct the lost spectral information of HR-MSI under the constraint of spatial edge loss (Lspec). Comparative experiments are conducted on six HSI data sets of Urban, Pavia University (PU), Pavia Center (PC), Botswana, Indian Pines (IP), and Washington DC Mall (WDCM), and the proposed SSR-NET achieves the superior or competitive results in comparison with seven state-of-the-art methods. The code of SSR-NET is available at https://github.com/hw2hwei/SSRNET.
Wei Huang 0068, Qi Wang 0009, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 KALM: Key Area Localization Mechanism for Abnormality Detection in Musculoskeletal Radiographs
abstract
Recently abnormality detection in musculoskeletal radio-graphs has attracted many attentions. For abnormality detection, it is crucial to locate the most important area in the musculoskeletal radiographs. To achieve this goal, we propose a key area localization mechanism (KALM) for abnormality detection for the first time in this paper. The proposed KALM explicitly defines the process of selecting the most important area from the whole image with using only image-level label. Based on KALM, we further present a joint global and local feature representation strategy for abnormality detection which takes as input both the entire image and the selected local area. The experimental results based on several classical convolutional neural network (CNN) architectures of MURA, the largest abnormality detection dataset of musculoskeletal radiographs, demonstrate the effectiveness of our KALM.
Wei Huang 0068, Zhitong Xiong, Qi Wang 0009, Xuelong Li 0001
ICASSP1
2019 Feature Sparsity in Convolutional Neural Networks for Scene Classification of Remote Sensing Image
abstract
Recently, the analysis of remote sensing images has attracted a lot of attention. In the domain of scene classification, deep learning methods, especially convolutional networks (CNNs), currently achieve the best results. Although the classification performance has reached a high level, there are still some factors limiting the improvement of classification accuracy. Based on obeservation of remote sensing scene images, we fing that some scenes are quite similar though they belong to different classes. To improve the classification performance between different scenes with similar characteristics, we propose a significant Feature Sparsity Layer that can be esaily embedded into various convolutional network architectures. The proposed layer can inhibit the confusing features meanwhile stress the discriminative features, and it is used to sparse the multi-layer feature map, which is extracted by the convolutional layers. The proposed method achieves the state-of-the-art results on three datasets UC Merced Land Use, Aerial Image Data and OPTIMAL-31, and competitive result on dataset WHU-RS19.
Wei Huang 0068, Qi Wang 0009, Xuelong Li 0001
IGARSS1