Haoyu Wang 0008

dblp:50/8499-8 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0002-8905-822XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ULDGN: Uncertainty-aware language-guided domain generalization network for cross-scene hyperspectral image classification
Tianyang Duan, Haoyu Wang 0008
Pattern Recognit.3
2026 Test-Time Training: Bayesian Meta-Hessian Network for Single-Source Domain Generalization
abstract
Existing domain generalization methods for hyper-spectral image (HSI) classification face challenges related to negative transfer risks and a lack of dynamic adaptation mechanisms. To address these issues, this paper proposes the Bayesian Meta-Hessian Network (BHMN), introducing a novel test-time training paradigm for single-source domain generalization. First, to mitigate negative transfer, BMHN moves beyond traditional feature alignment by employing an innovative Hessian matrix-based gradient regularization strategy. This approach captures domain-invariant knowledge from second-order gradient information, guiding the model toward flatter solution regions and enhancing robustness. Second, to overcome the limitations of static models, we develop a test-time training mechanism based on Bayesian posterior inference, enabling the model to dynamically adapt its parameters to shifts in the target domain. These components are unified within a meta-learning framework that simulates test-time scenarios during source domain training, teaching the model how to adapt. Experimental results on multiple HSI datasets demonstrate that the proposed BMHN achieves state-of-the-art performance.
Haoyu Wang 0008
IEEE Trans. Circuits Syst. Video Technol.1
2026 HMAMRL: Multicriterion Flexible Coordinated Control for Coal-Fired Power Generation Systems under Wide Load Operation
abstract
Flexible and efficient wide-load tracking in coal-fired power generation systems (CPGSs) is crucial for integrating renewable energy. To address the challenges arising from the dynamic characteristics and task distribution differences during the wide-load operation of thermal power units, this article proposes a novel hierarchical model-agnostic meta reinforcement learning (HMAMRL) framework. This framework combines inner meta-learning for quick adaptation within task categories and outer meta-learning for sharing general task knowledge, ensuring robust generalization under different load conditions. Meanwhile, an adaptive multicriterion reward function design method is proposed to dynamically balance load tracking costs, coal consumption costs, and input fluctuation costs. Moreover, a truncated proximal policy optimization (TPPO) algorithm ensures precise load control within physical constraints. Experimental results on the 160 and 1000 MW CPGSs demonstrate the effectiveness and superiority of the proposed algorithm.
Mengjun Yu, Chunyu Yang 0001, Haoyu Wang 0008, Linna Zhou, Huaichun Zhou
IEEE Trans. Cybern.4
2026 GIDDM: Generating Labels With Diffusion Model to Promote Cross-Domain Open-Set Image Recognition
abstract
Due to the lack of prior knowledge about unknown classes during training, existing methods for cross-domain open-set image recognition typically rely on threshold-based solutions. However, such approaches often struggle to capture the complex boundary relationships between known and unknown classes, which can lead to negative transfer effects caused by feature confusion between the two. To address this issue, this paper proposes a graph isomorphic distillation diffusion model (GIDDM) that aims to learn the boundary relationships between known and unknown classes from a closed-set classifier that models predictive uncertainty. First, a diffusion classifier is designed to quantify model predictive uncertainty through a Monte Carlo sampling strategy performed on the noise distribution during the reverse denoising process. The uncertainty distribution is modeled, and the cumulative distribution function is used to compute the probability of a sample belonging to an unknown class. Second, an open-set recognition framework is constructed, treating the closed-set diffusion classifier as a teacher classifier, and guiding the student classifier to learn the complex boundary relationships between known and unknown classes through knowledge distillation. Third, the knowledge distillation process is further formalized as a graph isomorphic optimization problem, where the predictive manifolds of the student and teacher classifiers are constrained to be consistent, thereby enhancing knowledge transfer between the classifiers. Finally, the entire process is integrated into a unified open-set adversarial domain adaptation framework, reconstructing the traditional optimization objectives of closed-set adversarial domain adaptation to ensure sufficient separation between known and unknown classes while aligning the distributions of known classes in both the source and target domains. Experiments conducted on multiple hyperspectral image (HSI) datasets demonstrate that the proposed method achieves state-of-the-art performance on cross-domain open-set image recognition tasks. The code demo can be accessed on the following website: https://github.com/wzr78998/GIDDM.
Haoyu Wang 0008, Yuhu Cheng 0001, Wei Zhang 0382, Xuesong Wang 0001
IEEE Trans. Image Process.1
2026 VDBAN: Suppressing Intermodal Information Interference Faced by Multimodal Object Detection
abstract
Objects in complex environmental conditions such as low light and smoke occlusion are difficult to be detected accurately, and multimodal fusion of visible and infrared images with complementary physical properties provides a solution. However, different modalities often exhibit significant modal heterogeneity due to differences in imaging mechanisms, which can easily induce intermodal information interference. We decomposed inter modal information interference into task coupling interference, distribution difference interference, and noise superposition interference. To address the above issues, the variational decoupled bottleneck adaptation network (VDBAN) has been proposed. First, we decoupled the class distribution and spatial distribution of the two modalities, which in turn is targeted to captured task-related distribution information from multimodal data. In addition, we also customized differentiated fusion parameters and strategies for object recognition and localization tasks, adapting to the feature learning preferences of different tasks. Second, we designed the global mutual information adaptation mechanism and the structural mutual information adaptation mechanism, which promoted cross-modal knowledge sharing by enhancing the shared information between modalities. Finally, we constructed a variational fusion information bottleneck to compress the information flow in the multimodal feature fusion process to make the model focus on the task related information, and filter out the task-independent noise information. Numerous experimental results have demonstrated that VDBAN shows state-of-the-art detection performance in multimodal object detection task.
Haoyu Wang 0008, Yuhu Cheng 0001, Xuesong Wang 0001
IEEE Trans. Multim.1
2025 Open Set Cross-Domain Hyperspectral Image Classification Based on Critical Reflective Learning Network
abstract
Limited by the lack of supervisory information of unknown classes, existing open set cross-domain hyperspectral image (HSI) classification methods often rely on threshold-based methods when identifying unknown classes, and are difficult to adapt to the complex inter-class variability between unknown and known classes in HSI. To solve this problem, this paper proposes a critical-reflective learning network (CRLN) based on teacher-student network, which obtains the unknown classes probability from the prediction of the closed set classifier (teacher network) and provides supervisory information for the open set classifier (student network). Specifically, first, the teacher network is trained with source domain data, and the unknown class probability predictions of the teacher network are obtained by quantifying the uncertainty of the network predictions. Second, the student network learns the complex boundary relationship between known and unknown classes based on the output of the teacher network to achieve accurate recognition of unknown classes. Furthermore, considering the credibility of the teacher network, a critical-reflective learning mechanism is proposed to allow the student network to reflect on the erroneous experiences of the teacher network, thus alleviating the performance damage caused by this potentially erroneous knowledge. Finally, the class alignment-separation module is proposed, which uses contrastive learning to promote the separation of known and unknown classes in the feature space, so as to reduce the risk of negative transfer induced by the confusion of the two types of features during cross-domain distribution adaptation. Experiments on three datasets show that the proposed method achieves state-of-the-art performance in the open set cross-domain HSI classification task.
Haoyu Wang 0008, Zhenzhuang Qiao, Wei Zhang 0382
IEEE Trans. Geosci. Remote. Sens.1
2024 Inducing Causal Meta-Knowledge From Virtual Domain: Causal Meta-Generalization for Hyperspectral Domain Generalization
abstract
Cross-domain hyperspectral image (HSI) classification can improve the model’s classification performance in the target domain by utilizing the rich knowledge from the source domain. However, existing cross-domain HSI classification methods mostly belong to transductive learning, which is difficult to apply to domain generalization tasks where the target domain is unseen during model learning. Inspired by human causal reasoning and knowledge induction mechanisms, this article develops an inductive learning-based framework for hyperspectral domain generalization: Causal meta-generalization. By simulating domain generalization scenarios, the framework helps the model induct domain-invariant causal meta-knowledge, thereby ensuring its strong generalization ability to unseen target domains. Specifically, we first propose a bottleneck variational auto-encoder (B-VAE) based on a forward–reverse information bottleneck, decoupling the domain distribution and class distribution of HSIs. By perturbing the domain distribution to generate virtual domains, we simulate potential domain distribution changes in the real world, providing a data basis for the induction of causal meta-knowledge. Second, in the process of simulating domain generalization scenarios, we establish a dual-layer optimization mechanism (DLOM) based on invariant-generalization risk minimization. In the inner layer optimization, by minimizing the model’s invariant causal effect loss (ICEL) in the virtual source domains, we guide the model to learn domain-invariant causal meta-knowledge. In the outer optimization, by minimizing the model’s generalization risk in the unseen virtual target domain, we enhance the applicability of causal meta-knowledge in domain generalization tasks. This proposed method has potential applications in remote sensing signal-processing tasks, such as the recognition of crop pests and diseases and the identification of minerals. The code can be accessed athttps://github.com/wzr78998/CMG.
Haoyu Wang 0008, Zhenzhuang Qiao, Hanqing Tao
IEEE Trans. Geosci. Remote. Sens.1
2024 Multimodal Remote Sensing Data Classification Based on Gaussian Mixture Variational Dynamic Fusion Network
abstract
With the development of sensor technology, the rational use of multimodal data has become a research hotspot in the field of remote sensing. The multimodal fusion method can effectively improve the accuracy of remote sensing data classification by using the complementary information of different modalities. However, the existing multimodal fusion methods face many challenges, including difficulties in suppressing spectral noise, fully mining contextual information, and learning the strong adaptive fusion pattern. To address the above challenges, a Gaussian mixture variational dynamic fusion network (GM-VDFN) is proposed. First, a multimodal multiscale spatial graph is constructed, and the graph convolution is used to learn the multiscale features. In this process, a spatial topology constraint based on GM (STC-GM) is proposed, which suppresses spectral noise by constraining the topological consistency of the two modalities. Second, a multiscale dynamic graph aggregation module (MDGAM) is constructed, which can capture the shareable class identification information from multiscale features and mine personalized fusion patterns suitable for each sample. Finally, the evidence lower bound for the multimodal joint distribution is derived, and a multimodal variational autoencoder (M-VAE) is designed. Optimizing the evidence lower bound to model multimodal joint distributions, thereby learning the strong adaptive fusion pattern between modalities. Experimental results on four fusion datasets (Houston 2013, Trento, MUUFL, and Houston 2018) show that GM-VDFN achieved state-of-the-art performance in multimodal remote sensing data classification tasks.
Haoyu Wang 0008, Zhenzhuang Qiao, Guoqing Wang 0003
IEEE Trans. Geosci. Remote. Sens.1
2024 Value Distribution DDPG With Dual-Prioritized Experience Replay for Coordinated Control of Coal-Fired Power Generation Systems
abstract
The grid connection of renewable energy poses challenges to the coordinated control of coal-fired power generation systems. Model uncertainty makes model-driven methods less effective due to the lack of adaptive capability. Large inertia of thermal process leads to local aggregation of state information, and the direct grafting reinforcement learning methods will affect the learning efficiency due to insufficient data utilization. To this end, this article proposes dual-prioritized experience replay value distribution deep deterministic policy gradient (DPER-VDP3G) algorithm. Value distribution is introduced to reflect the influence of model uncertainty on the evaluation of coordinated control policy, thus improving the accuracy of prediction cost function. The DPER is designed to reduce the nonuniform sampling bias and remove redundant data to enhance sample diversity. Comparative experiments demonstrate the advantages of the proposed method for improving network training efficiency, ameliorating load tracking accuracy and speed, and reducing energy consumption.
Mengjun Yu, Chunyu Yang 0001, Linna Zhou, Haoyu Wang 0008, Huaichun Zhou
IEEE Trans. Ind. Informatics5
2024 Reinforcement Learning Based Markov Edge Decoupled Fusion Network for Fusion Classification of Hyperspectral and LiDAR
abstract
Hyperspectral images (HSIs) and light detection and ranging (LiDAR) are two critical and frequently used types of remote sensing data, each containing rich spectral and elevation information. Fusing HSI and LiDAR can exploit the complementary properties of the two modalities for ground object classification. The performance of existing fusion classification methods is often limited by the difficulty of adapting feature extraction operators to complex spatial distributions, and the correlation and specificity between different modalities are not reasonably exploited. Therefore, the reinforcement learningbased markov edge decoupled fusion network (MEDFN) is proposed. This network can intelligently compose graphs based on different modal characteristics and tasks to adapt to complex spatial distributions; it can also suppress noise to complete fusion classification while fully utilizing complementary information of different modalities. First, a reinforcement learning-based graph construction subnetwork (RLGN) is proposed to learn a twomodal graph construction strategy suitable for classification tasks by transforming regular multimodal data into irregular graph data. Second, a multimodal edge attention module (MEAM) is proposed to extract edge features between spatial neighboring nodes and model the importance of each node, thereby capturing the spatial topology information encompassed in the multimodal data. Finally, the decoupled multimodal fusion module (DMFM) is proposed to decouple multimodal features into shared and unshared parts and enhance the model's ability to distinguish features by targeting the modal-shared feature between modalities and modal-specific feature. The experimental results based on three well-known HSI and LiDAR datasets demonstrate the effectiveness of the proposed MEDFN in fusion classification tasks.
Haoyu Wang 0008, Yuhu Cheng 0001, Xuesong Wang 0001
IEEE Trans. Multim.1
2023 Bi-Classifier Adversarial Network for Cross-Scene Hyperspectral Image Classification
abstract
Labeling hyperspectral images (HSIs) is time-consuming and labor-intensive for researchers, so the deficiency of adequate labeling samples is a giant obstacle to conducting HSI classification. Especially, such issue is exacerbated when there are no available labeled samples in the target scene. For the sake of resolving aforesaid issue, we put forward a novel cross-scene HSI classification method namely bi-classifier adversarial augmentation network (BCAN) so as to transfer knowledge from a similar but different source domain to an unlabeled target domain. First, the source and target domain distributions are aligned by maximizing and minimizing the decision discrepancy between two classifiers, respectively. Then, more accurate samples corresponding to pseudo-labels are selected as reliable samples and added to the training set. Finally, the spectral band random zeroing (SBRZ) method is proposed to expand the training samples for reliable samples, which handles the problem of insufficient network training resulted from insufficient samples in the source domain. By using multi-classifiers for domain adaptation and data augmentation, the accuracy of the network for cross-scene HSI classification tasks are improved. BCAN can extract the source domain’s helpful information to complete the target domain classification task. Experiments conducted on ten HSI data pairs show that BCAN outperforms many state-of-the-art baselines.
Haoyu Wang 0008, Yuhu Cheng 0001, Yi Kong 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 Causal Meta-Transfer Learning for Cross-Domain Few-Shot Hyperspectral Image Classification
abstract
Few-shot hyperspectral image (HSI) classification poses challenges due to sample selection bias in few-shot scenarios, potentially leading to incorrect statistical associations between noncausal factors and category semantics. To address these challenges, an original HSI is treated as a mixture comprising causal and noncausal factors. By integrating the causal learning, meta-learning, and transfer learning, a cross domain few-shot HSI classification method based on causal meta-transfer learning (CMTL) is developed. First, a Mask Transformer is implemented to identify noncausal factors unrelated to categories. Second, an independent causal constraint is applied to separate the causal and noncausal factors, and enhancing the inclusion of pure and independent causal factors in the features. Finally, the meta-transfer learning is leveraged to enable the classification model to extract causal factors highly correlated with category semantics from data, facilitating the cross-domain knowledge transfer. Meanwhile, a causal association module is employed to maximize the mutual information between causal factors and category predictions, thereby ensuring a strong causal association between causal factors and classification tasks. Experimental results show that CMTL achieves competitive performance in cross-domain few-shot HSI classification tasks.
Yuhu Cheng 0001, Wei Zhang 0382, Haoyu Wang 0008, Xuesong Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Graph Meta Transfer Network for Heterogeneous Few-Shot Hyperspectral Image Classification
abstract
Since obtaining labeled hyperspectral images (HSIs) is difficult and time-consuming, the shortage of training samples has always been a challenge for HSI classification. In practical applications, only a few labeled samples are available in the task domain (target domain), while sufficient labeled samples are available in another domain (source domain). At the same time, these two domains are heterogeneous and contain different categories. This scenario makes it difficult to effectively transfer knowledge from the source domain to the target domain. To address this challenge, we propose a novel heterogeneous few-shot learning (FSL) method, namely graph meta transfer network (GMTN). Specifically, the graph sample and aggregate network (GraphSAGE) and meta-learning, which are both inductive learning, are integrated into a unified framework. In this way, the aggregation function is generalized from abundant few-shot tasks for feature extraction on the source and target domains. The spatial importance strategy (SIS) is designed to guide the feature propagation and alleviate the information interference caused by different categories. The neighborhood receptive field spectral attention (RFSA) mechanism is proposed to model the importance of spectral band using the information of the neighborhood pixels, which enables GMTN to pay more attention to bands with discriminative features in both domains. In addition, the node spatial information reset method is proposed to augment samples based on the spatial position relationship of nodes. Furthermore, to alleviate the domain shift in heterogeneous scenarios, the conditional domain adversarial strategy is used to achieve effective meta-knowledge transfer. Experiments show that GMTN outperforms the compared state-of-the-art methods.
Haoyu Wang 0008, Xuesong Wang 0001, Yuhu Cheng 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Hyperspectral Image Classification Based on Domain Adversarial Broad Adaptation Network
abstract
For hyperspectral image (HSI) classification tasks, obtaining sufficient labeled samples is usually difficult, time-consuming, and expensive. To address the aforementioned issue, by transferring the labeled sample information of a relevant source domain to the unlabeled target domain, an HSI classification method based on the domain adversarial broad adaptation network (DABAN) is proposed. First, the bottleneck adaptation module composed of a bottleneck layer and a domain adaptation layer is constructed and introduced to the domain adversarial neural network; thus, the domain adversarial adaptation network (DAAN) is designed. By simultaneously performing domain adversarial learning, reducing both the marginal distribution difference and second-order statistic difference between two domains, the distributions of the source and target domains are aligned. Then, the conditional distribution adaptation regularization term based on the maximum mean discrepancy is embedded into a broad learning system to obtain the conditional adaptation broad network (CABN). On the one hand, CABN can perform the class-level distribution adaptation on the domain-invariant features extracted by DAAN. On the other hand, the representation ability of the domain-invariant features expanded by CABN can be further enhanced. Experimental results on ten real hyperspectral data pairs show that, compared with the existing mainstream methods, DABAN can effectively utilize relevant source-domain information to assist in improving the classification accuracy of the target domain.
Haoyu Wang 0008, Yuhu Cheng 0001, C. L. Philip Chen, Xuesong Wang 0001
IEEE Trans. Geosci. Remote. Sens.1