Wenzhen Wang

dblp:275/8881 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Topological Information Aggregation Network for Few-Shot Cross-Domain Hyperspectral Image Classification
abstract
In recent advancements, hyperspectral image (HSI) classification through few-shot learning (FSL) has significantly progressed. Domain adaptation, integrated with FSL, effectively utilizes transferable knowledge from a source domain (SD) with abundant labeled data to excel in classification tasks within a target domain (TD) with scarce labels. However, most existing methods usually use traditional convolutional neural networks (CNNs) to extract local spatial information to characterize and mine feature and distribution information while ignoring the underlying topological relationships among feature classes. Therefore, we propose a topology graph perception cross-domain FSL (TGP-CFSL) framework that leverages graph information aggregation. Specifically, to construct the extended topological relationships of the target, we have designed a topological graph-based multiscale fusion (TGMF) feature extraction module, which is adept at fully mining the topological spatial neighborhood information of the target. Meanwhile, a dual-graph information perception (DGIP) module is designed, which is able to characterize and aggregate intradomain topological relationships in terms of both feature representations and interdomain distribution similarities and to extract higher order domain distribution information for realizing domain alignment. Experimental results on three public HSI datasets demonstrate that the proposed method outperforms existing methods.
Kai Shi 0001, Wenzhen Wang, Qichao Liu, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Text-Driven Adaptive Semantic Alignment Network for Cross-Scene Hyperspectral Image Classification
abstract
Land cover in different scenes generally exhibits scene-invariant category semantic, typically represented and described consistently in a textual modality. Traditional cross-scene classification methods often treat categories as discrete class labels, neglecting their semantic information, or use category names merely as auxiliary textual modalities to enhance the discriminative representations of land cover. However, the cross-scene consistency of category semantic for land cover remains underexplored and underutilized. To address this issue, the text-driven adaptive semantic alignment network (TASA-Net) is proposed in this article for cross-scene hyperspectral image classification (HSIC). TASA-Net employs hand-crafted template prompts for stable category descriptions and vision-guided fine semantic prompts (VG-FSPs) for dynamic scene adaptation. Through a dual-gated adaptive mechanism, TASA-Net optimally weights coarse- and fine-grained semantics in a shared space, ensuring stable yet discriminative semantic representation. Additionally, cross-modal semantic alignment projects visual features into the shared semantic space, while a soft alignment strategy dynamically adjusts category correlations to enhance intraclass consistency and mitigate domain shifts. Ultimately, by leveraging text-driven semantic consistency representation, TASA-Net achieves zero-shot cross-scene transfer for unsupervised classification. Experiments demonstrate superior performance across multiple hyperspectral datasets, validating the critical role of textual modality in enhancing model robustness and cross-scene generalization ability.
Wenzhen Wang, Fang Liu 0034, Hongyuan Zhu 0002, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 I2MEP-Net: Inter- and Intra-Modality Enhancing Prototypical Network for Few-Shot Hyperspectral Image Classification
abstract
Prototypical network, celebrated for its flexible network structure and metric computation capability, has become a prevailing strategy in addressing challenges associated with few-shot hyperspectral image classification. However, factors such as insufficient training samples, spectral mixing, and noise interference in complex scenarios severely impact the stability of its prototypes, ultimately leading to a degradation in classification performance. Therefore, this paper proposes a novel method called I2MEP-Net, which incorporates two auxiliary modalities to facilitate both inter- and intra-modality enhancing prototype learning with the base modality. Specifically, I2MEP-Net employs the auxiliary LiDAR modality with base HS modality from the same scene for inter-modality enhancing prototype learning, which offsets the sparsity of few-shot features through a cross-modal approach. In addition, it utilizes the target few-shot labeled data as the auxiliary HS modality for intra-modality enhancing prototype learning on the enhanced prototypes, in a way that adaptively generates diversity features, thereby further enriching the prototype embedding space and achieving more fine-grained and stable prototypes. Comprehensive experiments are conducted on the publicly available hyperspectral image datasets. These experiments indicate that the proposed I2MEP-Net outshines the existing state-of-the-art deep learning techniques and few-shot classification methodologies.
Wenzhen Wang, Fang Liu 0034, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 A Multi-Scale Deep Feature Learning and Semantic Enhancement Approach for Remote Sensing Scene Classification
abstract
Deep learning has made great success in remote sensing scene classification since the powerful feature representation and complex nonlinear relationship learning. However, existing methods ignore the information redundancy and semantic ambiguity among them. To cope with this problem, we propose a multi-scale deep feature learning and semantic enhancement approach (MDFL-SE). First, we employ Pyramid Convolution PyConvResNet as the backbone to extract multilayer convolutional features. Then, a progressive deep feature aggregation module (PDFA) is designed to use high-level features to guide the low-level ones to choose the discriminative features. Finally, a global multiscale semantic extraction module (GMSE) and a grouped semantic extraction module (GSE) are combined to extract the channel and spatial information of multilayer fusion features. Experiments performed on AID and NWPURESISC45 RSSC datasets demonstrate that the proposed framework can obtain outstanding performance compared with state-of-the-art approaches.
Hengyi Huang, Wenzhen Wang, Wenzi Liao, Liang Xiao 0001
IGARSS2
2023 Cross-Domain Few-Shot Hyperspectral Image Classification With Class-Wise Attention
abstract
Few-shot learning (FSL) is an effective method to solve the problem of hyperspectral image (HSI) classification with few labeled samples. It learns transferable knowledge from sufficient labeled auxiliary data to classify unseen classes with limited labeled samples for training. However, the distribution difference between auxiliary data and unseen classes results in the learned transferable knowledge not being well applied to the new task. Therefore, a class-wise attentive cross-domain FSL (CA-CFSL) framework is proposed in this article, in which a feature extractor is learned to extract data features with discriminability and domain invariance. The class-wise attention metric module (CAMM) introduces class-wise attention on the FSL framework to learn more discriminative features, which improves the interclass decision boundaries. Furthermore, an asymmetric domain adversarial module (ADAM) is designed to enhance the ability of extracting domain-invariant representations, which combines asymmetric adversarial training with embedded domain-specific information. Experimental results on four public HSI datasets demonstrate that the proposed method outperforms the existing methods.
Wenzhen Wang, Fang Liu 0001, Jia Liu 0020, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Cross-Modal Graph Knowledge Representation and Distillation Learning for Land Cover Classification
abstract
Complementary multimodal remote sensing (RS) data often leads to more robust and accurate classification performance. However, not all modal data can be available at the time of inference due to imaging conditions. To mitigate this issue, cross-modal knowledge distillation becomes an effective method, as it can leverage the complementary characteristics of multimodal data to guide cross-modal classification in cases with missing data. Therefore, this paper examines the shortcomings of traditional CNN cross-modal distillation methods in land cover classification: 1) insufficient knowledge representation; and 2) unstable knowledge transfer. Moreover, a novel cross-modal graph knowledge representation and distillation learning (CGKR-DL) framework is proposed to enhance land cover classification performance. The proposed CGKR-DL designs a single-stream joint feature learning network with convolutional neural network and graph convolutional network (CNN-GCN) to effectively construct the remote topology of data based on the strong correlation between land objects, thus enhancing the knowledge representation ability of the network. In addition, a multi-granularity graph distillation method is proposed to compensate for the inability of traditional CNN distillation in handling graph-structured information, where a feature distillation module based on graph discrimination (FD-GDM) is designed for stable graph feature distillation. We evaluate CGKR-DL on three publicly available multimodal RS datasets (HS-LiDAR, HS-SAR and HS-SAR-DSM) and achieve a significant improvement in comparison with several state-of-the-art methods.
Wenzhen Wang, Fang Liu 0034, Wenzi Liao, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.1
2020 Novel segmentation algorithm for jacquard patterns based on multi-view image fusion
abstract
Pattern regeneration is one of the applications of reverse engineering technology in the textile field, which realises the process of textile‐pattern regeneration‐textile, and fundamentally provides an intelligent design means of textile. At present, the method of pattern regeneration in the jacquard fabric is to use the image segmentation algorithm to segment the image digitalised by unidirectional imaging, and then the segmented pattern could be identified to regeneration for the design of new fabrics. However, due to the concave and convex pattern textures on the surface of jacquard fabric, the traditional unidirectional imaging method cannot be used for the full characterisation of its structural information, resulting in unsatisfactory pattern segmentation effect. To solve this problem, a novel segmentation algorithm for jacquard patterns based on multi‐view image fusion was proposed in this study. Based on multi‐view image acquisition and fusion, the pattern image of jacquard fabric could be cluster‐segmented by extracting the complete texture information of the fused image and the actual colour information of the calibrated image. Compared with the traditional unidirectional imaging method, the experimental results show that the enhanced texture information of the fused image is more workable for the pattern segmentation, it validates the effectiveness of the proposed method.
Wenzhen Wang, Na Deng, Binjie Xin, Yiliang Wang, Shuaigang Lu
IET Image Process.1