Shuying Huang

dblp:04/925 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0003-2771-8461ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 11 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UMNet: Uncertainty-guided Memory Network for Hyperspectral Pansharpening
abstract
At present, most hyperspectral (HS) sharpening methods have not fully utilized the feature correlation between adjacent bands in HS images, nor have they explored the problem of feature uncertainty generated by the model during the fusion process. This may lead to inaccurate fusion features generated by the model, resulting in spatial and spectral distortions in the fusion results. To address these issues, we propose an uncertainty-guided memory network (UMNet) for HS pansharpening. A spatial-spectral recurrent fusion unit (SRFU) is designed based on the concept of temporal data modeling, which utilizes the correlation between adjacent bands to fuse spectral and spatial features from PAN and LRHS images. In SRFU, a state memory interaction unit (SMIU) is constructed based on non-negative matrix factorization (NMF) to learn the global spatial-spectral dependency of PAN and HS images in the recurrent state space. Moreover, based on uncertainty theory, we define two spatial-spectral uncertainty-guided loss functions for the HS pansharpening task to train the model step by step, ensuring that the network can reconstruct more accurate spectral and spatial features. Extensive experiments on three widely used datasets demonstrate that, compared with some state-of-the-art (SOTA) methods, the proposed UMNet has achieved significant improvements in both spatial and spectral quality metrics.
Yong Yang 0001, Shuying Huang, Nayu Liu
AAAI3
2026 Learning from not-all-negative N-tuples and unlabeled data
Shuying Huang, Changchun Hua, Yana Yang
Pattern Recognit.1
2026 PCFFusion: Progressive cross-modal feature fusion network for infrared and visible images
abstract
Infrared and visible image fusion (IVIF) aims to fuse thermal target information in infrared images and spatial texture information in visible images, improving the observability and comprehensibility of the fused images. Currently, most IVIF methods suffer from the loss of salient target information and texture details in fused images. To alleviate this problem, a progressive cross-modal feature fusion network (PCFFusion) for IVIF is proposed, which comprises two stages: feature extraction and feature fusion. In the feature extraction stage, to enhance the network’s feature representation capability, a feature decomposition module (FDM) is constructed to extract two modal features of different scales by defining a feature decomposition operation (FDO). In addition, by establishing correlations between the high- frequency and low-frequency components of two modal features, a cross-modal feature enhancement module (CMFEM) is built to realize correction and enhancement of the two features at each scale. The feature fusion stage achieves the fusion of two modal features at each scale and the supplementation of adjacent scale features by constructing three cross-domain fusion module (CDFMs). To constrain the fused results preserve more salient targets and richer texture details, a dual-feature fidelity loss function is defined by constructing a salient weight map to balance the two loss terms. Extensive experiments demonstrate that fusion results of the proposed method highlight prominent targets from infrared images while retaining rich background details from visible images, and the performance of PCFFusion is superior to some advanced methods. Specifically, compared to the optimal results obtained by other comparison methods, the proposed network achieves an average increase of 30.35 % and 10.9 % in metrics Mutual Information (MI) and Standard deviation (SD) on the TNO dataset, respectively.
Shuying Huang, Yong Yang 0001, Weiguo Wan
Pattern Recognit.1
2026 EHEN: Eigendecomposition-based hyperchannel enhancement network for hyperspectral image super-resolution
Yong Yang 0001, Aoqi Zhao, Shuying Huang, Weiguo Wan
Pattern Recognit.3
2026 A mask protection-based and multi-probabilistic prior dictionary-guided unfolding model for pansharpening
Shengna Wei, Yong Yang 0001, Shuying Huang, Weiguo Wan, Changjie Chen 0002
Signal Process.3
2026 Learning From Mixed-Class N-Tuples and Unlabeled Data
abstract
To alleviate the annotation burden in supervised learning, various weakly-supervised learning (WSL) configurations have emerged-such as pairwise, triplet, and N-tuple based supervision-that leverage instance-level relationships to reduce labeling costs. While these methods have demonstrated strong performance across diverse tasks, most existing methods are limited to homogeneous tuples and do not support learning from mixed-class tuples, particularly when the tuple size exceeds three. To address this gap, this paper introduces a framework called Mixed-class Triplet and Unlabeled Data (MTU) learning, which enables classification from triplets containing instances from different classes, combined with unlabeled data. We further extend this framework to a more general setting, Mixed-class N-tuple and Unlabeled (MNU) learning, which handles arbitrary-length tuples ( N$\geq$3 ) with mixed-class composition. Both MTU and MNU leverage unlabeled instances to enhance supervisory signals and improve classification performance. We formulate tailored empirical risk minimization (ERM) objectives for both learning settings and derive generalization error bounds to provide theoretical guarantees. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed methods under various weak supervision scenarios.
Shuying Huang, Changchun Hua, Yana Yang
IEEE Trans. Big Data1
2025 FMPM-DNet: Hyperspectral Pansharpening Dynamic Network Based on Feature Modulation and Probability Mask
abstract
Currently, most Hyperspectral (HS) pansharpening methods have two problems, namely the lack of consideration the spatial variations of HS images and inaccurate feature reconstruction in multi-channel complex mapping relationships, leading to spectral and spatial distortions in the fusion results. To address these issues, we propose a dynamic network based on feature modulation and probability mask (FMPM-DNet) for HS pansharpening, including two stages of spectral-spatial feature modulation and feature reconstruction. In the first stage, to increase the feature representation ability of the model, a wave function is defined based on complex transformation to convert spatial features into wave-like features. On this basis, considering the spatial variations of HS images, a dynamic feature modulation unit (DFMU) is constructed to achieve adaptive modulation and coarse fusion of features by dynamically generating spectral-spatial correction matrix. In the second stage, a feature probability mask unit (FPMU) is designed to realize global feature embedding at different depths and local feature embedding at the same depth to obtain refined fused features. Extensive experiments on three widely used datasets demonstrate that the proposed FMPM-Net achieves significant improvements in both spatial and spectral quality metrics compared to some state-of-the-art (SOTA) methods.
Yong Yang 0001, Shuying Huang, Hangyuan Lu, Weiguo Wan, Aoqi Zhao
AAAI3
2025 ID-RWKV: Image Deraining RWKV
abstract
Although Transformer has achieved considerable results in image deraining tasks, the quadratic complexity of self-attention in this structure limits its ability to process high-resolution rainy images. The Receptance Weighted Key Value (RWKV) in the field of natural language processing (NLP) solves the expensive computational cost of self-attention in Transformers by using linear computational complexity to learn long-range dependencies of features. Based on advanced design of RWKV, we propose an image deraining RWKV (ID-RWKV) that gradually obtains fine rainless images by constructing multiple deraining stages. At each stage, a U-shaped rain removal network (U-RRNet) is constructed to encode and decode image features, and output the rain removal results for the current stage. Each layer of U-RRNet consists of Fourier enhancement module (FFM) and local-global RWKV Block (LGRB). FEM is constructed to extract and enhance the features in the frequency domain. LGRB is designed based on RWKV structure to improve the performance and efficiency of deraining models. To better learn local and global contextual information, we propose a LG-WKV attention mechanism to enhance local and global dependencies. To reduce the loss of background information and increase the stability of the network, we construct a deep-shallow feature fusion module (DSFFM) to supplement shallow features. Following extensive experimentation, we demonstrate that our method not only outperforms the current state-of-the-art methods, but also requires fewer parameters and less computation than the Transformer-based method.
Yong Yang 0001, Jiaxuan Yang, Shuying Huang
ICASSP3
2025 Binary classification from N-Tuple Comparisons data
Shuying Huang, Changchun Hua, Yana Yang
Neural Networks2
2025 Learning from not-all-negative pairwise data and unlabeled data
Shuying Huang, Changchun Hua, Yana Yang
Pattern Recognit.1
2025 MSAN: Multiscale self-attention network for pansharpening
Hangyuan Lu, Yong Yang 0001, Shuying Huang, Rixian Liu, Huimin Guo
Pattern Recognit.3
2025 DFCFN: Dual-Stage Feature Correction Fusion Network for Hyperspectral Pansharpening
abstract
Hyperspectral (HS) pansharpening aims to fuse high-spatial-resolution panchromatic (PAN) images with low-spatial-resolution hyperspectral (LRHS) images to generate high-spatial-resolution hyperspectral (HRHS) images. Due to the lack of consideration for the modal feature difference between PAN and LRHS images, most deep leaning-based methods suffer from spectral and spatial distortions in the fusion results. In addition, most methods use upsampled LRHS images as network input, resulting in spectral distortion. To address these issues, we propose a dual-stage feature correction fusion network (DFCFN) that achieves accurate fusion of PAN and LRHS images by constructing two fusion sub-networks: a feature correction compensation fusion network (FCCFN) and a multi-scale spectral correction fusion network (MSCFN). Based on the lattice filter structure, FCCFN is designed to obtain the initial fusion result by mutually correcting and supplementing the modal features from PAN and LRHS images. To suppress spectral distortion and obtain fine HRHS results, MSCFN based on 2D discrete wavelet transform (2D-DWT) is constructed to gradually correct the spectral features of the initial fusion result by designing a conditional entropy transformer (CE-Transformer). Extensive experiments on three widely used simulated datasets and one real dataset demonstrate that the proposed DFCFN achieves significant improvements in both spatial and spectral quality metrics over other state-of-the-art (SOTA) methods. Specifically, the proposed method improves the SAM metric by 6.4%, 6.2%, and 5.3% compared to the second-best comparison approach on Pavia center, Botswana, and Chikusei datasets, respectively. The codes are made available at: https://github.com/EchoPhD/DFCFN.
Yong Yang 0001, Shuying Huang, Weiguo Wan, Long Zhang 0009, Aoqi Zhao
IEEE Trans. Geosci. Remote. Sens.3
2025 MFJLN: Multi-Frequency Feature Joint Learning Network for Rain Removal
Yong Yang 0001, Jiaxuan Yang, Shuying Huang, Weiguo Wan
IEEE Trans. Multim.3
2024 MFTN: Multi-Level Feature Transfer Network Based on MRI-Transformer for MR Image Super-resolution
abstract
Due to the unique environment and inherent properties of magnetic resonance imaging (MRI) instruments, MR images typically have lower resolution. Therefore, improving the resolution of MR images is beneficial for assisting doctors in diagnosing the condition. Currently, the existing MR image super-resolution (SR) methods still have the problem of insufficient detail reconstruction. To overcome this issue, this paper proposes a multi-level feature transfer network (MFTN) based on MRI-Transformer to realize SR of low-resolution MRI data. MFTN consists of a multi-scale feature reconstruction network (MFRN) and a multi-level feature extraction branch (MFEB). MFRN is constructed as a pyramid structure to gradually reconstruct image features at different scales by integrating the features obtained from MFEB, and MFEB is constructed to provide detail information at different scales for low resolution MR image SR reconstruction by constructing multiple MRI-Transformer modules. Each MRI-Transformer module is designed to learn the transfer features from the reference image by establishing feature correlations between the reference image and low-resolution MR image. In addition, a contrast learning constraint item is added to the loss function to enhance the texture details of the SR image. A large number of experiments show that our network can effectively reconstruct high-quality MR Images and achieves better performance compared to some state-of-the-art methods. The source code of this work will be released on GitHub.
Shuying Huang, Yong Yang 0001, Chenbin Liang
AAAI1
2024 MFTN: A Multi-scale Feature Transfer Network Based on IMatchFormer for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HISR) aims to fuse a low-resolution hyperspectral image (LR-HSI) with a high-resolution multispectral image (HR-MSI) to obtain a high-resolution hyperspectral image (HR-HSI). Due to some existing HISR methods ignoring the significant feature difference between LR-HSI and HR-MSI, the reconstructed HR-HSI typically exhibits spectral distortion and blurring of spatial texture. To solve this issue, we propose a multi-scale feature transfer network (MFTN) for HISR. Firstly, three multi-scale feature extractors are constructed to extract features of different scales from the input images. Then, a multi-scale feature transfer module (MFTM) consisting of three improved feature matching Transformers (IMatchFormers) is designed to learn the detail features of different scales from HR-MSI by establishing the cross-model feature correlation between LR-HSI and degraded HR-MSI. Finally, a multiscale dynamic aggregation module (MDAM) containing three spectral aware aggregation modules (SAAMs) is constructed to reconstruct the final HR-HSI by gradually aggregating features of different scales. Extensive experimental results on three commonly used datasets demonstrate that the proposed model achieves better performance compared to state- of-the-art (SOTA) methods.
Shuying Huang, Mingyang Ren, Yong Yang 0001, Yingzhi Wei
ICML1
2024 SCPSN: Spectral Clustering-based Pyramid Super-resolution Network for Hyperspectral Images
abstract
Single hyperspectral image super-resolution aims to reconstruct a high-resolution hyperspectral image (HRHSI) from an observed low resolution hyperspectral image (LRHSI). Most current methods combine CNN and Transformer structures to directly extract features of all channels in LRHSI for image reconstruction, but they do not consider the interference of redundant information in adjacent bands, resulting in spectral and spatial distortions in the reconstruction results and an increase in model computational complexity. To address this issue, this paper proposes a spectral clustering-based pyramid super-resolution network (SCPSN) to progressively reconstruct HRHSI at different scales. In each image reconstruction layer, a clustering super-resolution block (CSRB) consisting of spectral clustering block (SCB), patch non local attention block (PNAB), and dynamic fusion block (DFB) is designed to achieve the reconstruction of detail features. Specifically, for the high correlation between adjacent spectral bands in LRHSI, a SCB is first constructed to achieve clustering of spectral channels and filtering of hyperchannels. This can reduce the interference of redundant spectral information and the computational complexity of the model. Then, by utilizing the non-local similarity of features within the channel, a patch non-local attention block (PNAB) is constructed to enhance the features of hyperchannels. Next, a dynamic fusion block (DFB) is designed to reconstruct the features of all channels in LRHSI by establishing correlations between enhanced hyperchannels and other channels. Finally, the reconstructed channels are upsampled and added to the corresponding channels to obtain the reconstructed HRHSI. Extensive experiments validate that the performance of SCPSN is superior to that of some other state-of-the-art (SOTA) HSSR methods in terms of visual effects and quantitative metrics. In addition, our model does not require training on large-scale datasets compared to other methods. The dataset and code will be released on GitHub.
Yong Yang 0001, Aoqi Zhao, Shuying Huang, Yajing Fan
ACM Multimedia3
2024 ECLB: Efficient contrastive learning on bi-level for noisy labels
Juwei Guan, Shuying Huang, Yong Yang 0001
Knowl. Based Syst.3
2024 MFITN: A Multilevel Feature Interaction Transformer Network for Pansharpening
abstract
In this letter, to better supplement the advantages of features at different levels and improve the feature extraction ability of the network, a novel multi-level feature interaction transformer network (MFITN) is proposed for pansharpening, aiming to fuse multispectral (MS) and panchromatic (PAN) images. In MFITN, a multi-level feature interaction transformer encoding module is designed to extract and correct global multi-level features by considering the modality difference between source images. These features are then fused using the proposed multi-level feature mixing (MFM) operation, which enables features to fuse interactively to obtain richer information. Furthermore, the global features are fed into a CNN-based local decoding module to better reconstruct high-spatial-resolution multispectral (HRMS) images. Additionally, based on the spatial consistency between MS and PAN images, a band compression loss is defined to improve the fidelity of fused images. Numerous simulated and real experiments demonstrate that the proposed method has the optimal performance compared to state-of-the-art methods. Specifically, the proposed method improves the SAM metric by 7.89% and 6.41% compared to the second-best comparison approach on Pléiades and WorldView-3, respectively.
Changjie Chen 0002, Yong Yang 0001, Shuying Huang, Hangyuan Lu, Weiguo Wan, Shengna Wei, Wenying Wen
IEEE Geosci. Remote. Sens. Lett.3
2024 Pansharpening Based on Fuzzy Logic and Edge Activity
abstract
Pansharpening technology aims to extract spatial information from a panchromatic (PAN) image and integrate it into a multispectral (MS) image to generate a high-spatial resolution MS image. To overcome the problem of spatial and spectral distortion of traditional methods, this letter presents a pansharpening method based on fuzzy logic and edge activity. To obtain more accurate spatial details of source images, details extracted using the component substitution and multiresolution analysis methods are fused via the proposed fuzzy logic algorithm. To preserve the edges of the fusion result, the edge maps of the source images are detected and fused based on the edge activity. The optimized detail maps are obtained by multiplying the fused details and edge maps, which are then injected into the upsampled MS image to obtain the final pansharpened image. Reduced- and full-scale experimental results on the Pléiades and IKONOS datasets demonstrate the effectiveness of our method compared with state-of-the-art pansharpening algorithms. Specifically, the proposed method improves the ERGAS metrics by 9.0% and 11.1% compared to the second-best comparison approach on Pléiades and IKONOS, respectively.
Yong Yang 0001, Shuying Huang, Weiguo Wan, Hangyuan Lu
IEEE Geosci. Remote. Sens. Lett.3
2024 LSRN-AED: lightweight super-resolution network based on asymmetric encoder-decoder
Shuying Huang, Yong Yang 0001, Weiguo Wan, Houzeng Lai
Soft Comput.1
2024 VSDM: Variable-Scale Diffusion Model Based on Dynamic Condition Guidance for Pansharpening
abstract
Pansharpening aims to obtain a high-spatial-resolution multispectral (MS) image by fusing a lower-spatial resolution MS image with a high-spatial-resolution panchromatic (PAN) image. Currently, the results obtained by most pansharpening methods still suffer from spatial and spectral distortion issues. The diffusion model has shown outstanding performance in various image-processing tasks. However, maintaining the full image size throughout the diffusion process imposes a large computational burden, and the simultaneous use of PAN and MS images acquired by different sensors as a condition for guiding noise prediction leads to spatial and spectral distortions. To solve these problems, a variable-scale diffusion model (VSDM) based on dynamic condition guidance for pansharpening is proposed, which achieves better fusion performance by improving the diffusion manner of the diffusion model and injecting dynamic conditions to guide the reverse process. In VSDM, a variable-scale diffusion manner (VSDMN) is designed to reduce the computational complexity of the model by reducing the size of the image in the diffusion process. A condition generator (CG) is constructed to generate dynamic conditions using the features learned from the PAN and upsampled MS images. In CG, a cross-attention dynamic convolution is built to extract features from the PAN image by designing a spatial and spectral attention mechanism, which can improve the spatial and spectral consistency in the dynamic condition. Extensive experiments validate the effectiveness of the proposed VSDM against other state-of-the-art (SOTA) pansharpening methods in both quantitative and qualitative assessments. The source code will be released athttps://github.com/MELiMZ/VSDM.
Yong Yang 0001, Shuying Huang, Weiguo Wan, Hangyuan Lu, Wei Tu 0002
IEEE Trans. Geosci. Remote. Sens.3
2023 Low-Light Image Enhancement Network Based on Multi-Scale Feature Complementation
abstract
Images captured in low-light environments have problems of insufficient brightness and low contrast, which will affect subsequent image processing tasks. Although most current enhancement methods can obtain high-contrast images, they still suffer from noise amplification and color distortion. To address these issues, this paper proposes a low-light image enhancement network based on multi-scale feature complementation (LIEN-MFC), which is a U-shaped encoder-decoder network supervised by multiple images of different scales. In the encoder, four feature extraction branches are constructed to extract features of low-light images at different scales. In the decoder, to ensure the integrity of the learned features at each scale, a feature supplementary fusion module (FSFM) is proposed to complement and integrate features from different branches of the encoder and decoder. In addition, a feature restoration module (FRM) and an image reconstruction module (IRM) are built in each branch to reconstruct the restored features and output enhanced images. To better train the network, a joint loss function is defined, in which a discriminative loss term is designed to ensure that the enhanced results better meet the visual properties of the human eye. Extensive experiments on benchmark datasets show that the proposed method outperforms some state-of-the-art methods subjectively and objectively.
Yong Yang 0001, Wenzhi Xu, Shuying Huang, Weiguo Wan
AAAI3
2023 MMPN: Multi-supervised Mask Protection Network for Pansharpening
abstract
Pansharpening is to fuse a panchromatic (PAN) image with a multispectral (MS) image to obtain a high-spatial-resolution multispectral (HRMS) image. The deep learning-based pansharpening methods usually apply the convolution operation to extract features and only consider the similarity of gradient information between PAN and HRMS images, resulting in the problems of edge blur and spectral distortion in the fusion results. To solve this problem, a multi-supervised mask protection network (MMPN) is proposed to prevent spatial information from being damaged and overcome spectral distortion in the learning process. Firstly, by analyzing the relationships between high-resolution images and corresponding degraded images, a mask protection strategy (MPS) for edge protection is designed to guide the recovery of fused images. Then, based on the MPS, an MMPN containing four branches is constructed to generate the fusion and mask protection images. In MMPN, each branch employs a dual-stream multi-scale feature fusion module (DMFFM), which is built to extract and fuse the features of two input images. Finally, different loss terms are defined for the four branches, and combined into a joint loss function to realize network training. Experiments on simulated and real satellite datasets show that our method is superior to state-of-the-art methods both subjectively and objectively.
Changjie Chen 0002, Yong Yang 0001, Shuying Huang, Wei Tu 0002, Weiguo Wan, Shengna Wei
IJCAI3
2023 CTCP: Cross Transformer and CNN for Pansharpening
abstract
Pansharpening is to fuse a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to obtain an enhanced LRMS image with high spectral and spatial resolution. The current Transformer-based pansharpening methods neglect the interaction between the extracted long- and short-range features, resulting in spectral and spatial distortion in the fusion results. To address this issue, a novel cross Transformer and convolutional neural network (CNN) for pansharpening (CTCP) is proposed to achieve better fusion results by designing a cross mechanism, which can enhance the interaction between long- and short-range features. First, a dual branch feature extraction module (DBFEM) is constructed to extract the features from the LRMS and PAN images, respectively, reducing the aliasing of the two image features. In the DBFEM, to improve the feature representation ability of the network, a cross long-short-range feature module (CLSFM) is designed by combining the feature learning capabilities of Transformer and CNN via the cross mechanism, which achieves the integration of long-short-range features. Then, to improve the ability of spectral feature representation, a spectral feature enhancement fusion module (SFEFM) based on a frequency channel attention is constructed to realize feature fusion. Finally, the shallow features from the PAN image are reused to provide detail features, which are integrated with the fused features to obtain the final pansharpened results. To the best of our knowledge, this is the first attempt to introduce the cross mechanism between Transformer and CNN in pansharpening field. Numerous experiments show that our CTCP outperforms some state-of-the-art (SOTA) approaches both subjectively and objectively. The source code will be released at https://github.com/zhsu99/CTCP.
Zhao Su, Yong Yang 0001, Shuying Huang, Weiguo Wan, Wei Tu 0002, Hangyuan Lu, Changjie Chen 0002
ACM Multimedia3
2023 Multi-scale Spatial-Spectral Attention Guided Fusion Network for Pansharpening
abstract
Pansharpening is to fuse high-resolution panchromatic (PAN) images with low-resolution multispectral (LR-MS) images to generate high-resolution multispectral (HR-MS) images. Most of the deep learning-based pansharpening methods did not consider the inconsistency of the PAN and LR-MS images and used simple concatenation to fuse the source images, which may cause spectral and spatial distortion in the fused results. To address this problem, a multi-scale spatial-spectral attention guided fusion network for pansharpening is proposed. First, the spatial features from the PAN image and spectral features from the LR-MS image are independently extracted to obtain the shallow features. Then, a spatial-spectral attention feature fusion module (SAFFM) is constructed to guide the reconstruction of spatial-spectral features by generating a guidance map to achieve the fusion of reconstructed features at different scales. In SAFFM, the guidance map is designed to ensure the spatial-spectral consistency of the reconstructed features. Finally, considering the difference between multiply scale features, a multi-level feature integration scheme is proposed to progressively achieve fusion of multi-scale features from different SAFFMs. Extensive experiments validate the effectiveness of the proposed network against other state-of-the-art (SOTA) pansharpening methods in both quantitative and qualitative assessments. The source code will be released at https://github.com/MELiMZ/ssaff.
Yong Yang 0001, Shuying Huang, Hangyuan Lu, Wei Tu 0002, Weiguo Wan
ACM Multimedia3
2023 FRAN: feature-filtered residual attention network for realistic face sketch-to-photo transformation
Weiguo Wan, Yong Yang 0001, Shuying Huang, Lixin Gan
Appl. Intell.3
2023 Intensity mixture and band-adaptive detail fusion for pansharpening
Hangyuan Lu, Yong Yang 0001, Shuying Huang, Hongfu Su, Wei Tu 0002
Pattern Recognit.3
2023 AWFLN: An Adaptive Weighted Feature Learning Network for Pansharpening
abstract
Deep learning (DL)-based pansharpening methods have shown great advantages in extracting spectral–spatial features from multispectral (MS) and panchromatic (PAN) images compared with traditional methods. However, most DL-based methods ignore the local inner connection between the source images and the high-resolution MS (HRMS) image, which cannot fully extract spectral–spatial information and attempt to improve the quality of fusion by increasing the complexity of the network. To solve these problems, a lightweight network based on adaptive weighted feature learning network (AWFLN) is proposed for pansharpening. Specifically, a novel detail extraction model is first built by exploring the local relationship between HRMS and source images, thereby improving the accuracy of details and the interpretability of the network. Guided by this model, we then design a residual multiple receptive-field structure to fully extract spectral–spatial features of source images. In this structure, an adaptive feature learning block based on spectral–spatial interleaving attention is proposed to adaptively learn the weights of features and improve the accuracy of the extracted details. Finally, the pansharpened result is obtained by a detail injection model in AWFLN. Numerous experiments are carried out to validate the effectiveness of the proposed method. Compared to traditional and state-of-the-art methods, AWFLN performs the best both subjectively and objectively, with high efficiency. The code is available athttps://github.com/yotick/AWFLN.
Hangyuan Lu, Yong Yang 0001, Shuying Huang, Biwei Chi, Aizhu Liu, Wei Tu 0002
IEEE Trans. Geosci. Remote. Sens.3
2023 STCP: Synergistic Transformer and Convolutional Neural Network for Pansharpening
abstract
Pansharpening is a process of fusing a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to obtain a high-resolution multispectral (HRMS) image. Convolutional neural networks (CNNs) have been commonly utilized in this field because of their remarkable learning capabilities. However, their convolutional operators limit the long-range feature extraction ability of CNN. Meanwhile, the Transformer models have exhibited strong capabilities in modeling long-range representations, but there are shortcomings in modeling local-range feature dependencies. To this end, we propose a novel synergistic transformer and CNN for pansharpening (STCP). First, a parallel U-shaped feature extraction module (PUFEM) is constructed for extracting the features of the LRMS and PAN images, which improves the feature representation ability for the two source images. In the PUFEM, combining the different feature learning capabilities of the CNN and transformer, we design a long-short-range feature integration block (LSFIB) to extract the short-range features and long-range features at different scales in parallel. Then, a channel attention module (CAM)-based feature fusion module (CFFM) is constructed to integrate the features extracted by the PUFEM. Finally, the shallow features from the PAN image are reused to provide detailed features, which are integrated with the fused features from the CFFM to achieve the final pansharpened results. Numerous experiments show that our STCP outperforms some state-of-the-art approaches both subjectively and objectively.
Zhao Su, Yong Yang 0001, Shuying Huang, Weiguo Wan, Jiancheng Sun, Wei Tu 0002, Changjie Chen 0002
IEEE Trans. Geosci. Remote. Sens.3
2022 An Efficient Pansharpening Approach Based on Texture Correction and Detail Refinement
abstract
Pansharpening aims at fusing a multispectral (MS) image and panchromatic (PAN) image to obtain a high spatial resolution multispectral (HRMS) image. To obtain accurate details and reduce spectral distortion, this letter proposes an efficient pansharpening approach based on texture correction (TC) and detail refinement. First, a TC model is constructed based on spatial and spectral fidelity constraints to obtain a texture image that is highly correlated with the MS image. Second, a detail acquisition model is proposed by the consecutive parameter regression to adaptively refine the extracted details. Finally, the extracted details are injected into the up-sampled MS (UPMS) image to obtain the fused HRMS image. Experimental results demonstrate that the proposed method can obtain the high-quality results with high efficiency.
Hangyuan Lu, Yong Yang 0001, Shuying Huang, Wei Tu 0002
IEEE Geosci. Remote. Sens. Lett.3
2022 DCNP: Dual-Information Compensation Network for Pansharpening
abstract
To reduce the loss of detail and spectral information during the network propagation and better extract detail and spectral features at different scales, a novel dual-information compensation network for pansharpening (DCNP) is proposed for fusing multispectral (MS) and panchromatic (PAN) images. In the network, the domain-specific knowledge is considered to design our DCNP architecture by focusing on the two aims of the pansharpening: spatial and spectral preservation. Specifically, a cascaded U-shaped structure is constructed to improve the feature representation ability of the network. To preserve more spatial details in the pansharpened image, the details of the PAN image are extracted based on Laplacian operator and then compensated into the network. Furthermore, for spectral preservation, the MS image is conducted by the transposed convolution as the compensation information of the network. Experiments on the full- and reduced-scale data indicate that the proposed DCNP achieves significant improvement over state-of-the-art methods in terms of subjective and objective evaluation. Specifically, DCNP improves the PSNR and ERGAS metrics by 12.9% and 44.5% respectively compared to the deep learning-based approach with the best average values on Pléiades.
Yong Yang 0001, Zhao Su, Shuying Huang, Weiguo Wan, Wei Tu 0002, Changjie Chen 0002
IEEE Geosci. Remote. Sens. Lett.3
2022 MMDN: Multi-Scale and Multi-Distillation Dilated Network for Pansharpening
abstract
Pansharpening is a technology involving information integration and processing in remote sensing imagery. It is applied to generate a high-resolution multispectral (HRMS) image through an effective fusion of a low spatial resolution multispectral image and a panchromatic (PAN) image. In this paper, we propose an end-to-end multi-scale and multi-distillation dilated network (MMDN) for pansharpening. In MMDN, to extract more abundant spatial details from source images, a clique structure-based multi-scale dilated block (CSMDB) is presented. The clique structure in CSMDB can fully transfer the information between feature maps obtained by the multi-scale dilated convolutional filters. Then, a multi-distillation residual information block (MRIB) is constructed to help the network capture the spatial structure of different scales in MS and PAN images. Finally, to reuse and supplement the feature information, a feature embedding strategy is designed by feeding the sum result of the output of cascaded CSMDBs and the shallow features to each MRIB. Experimental results verify that the proposed MMDN outperforms other compared state-of-the-art approaches in terms of objective and subjective evaluations.
Wei Tu 0002, Yong Yang 0001, Shuying Huang, Weiguo Wan, Lixin Gan, Hangyuan Lu
IEEE Trans. Geosci. Remote. Sens.3
2022 Dual-Stream Convolutional Neural Network With Residual Information Enhancement for Pansharpening
abstract
Deep-learning-based pansharpening methods have achieved remarkable results due to their powerful feature representation ability. However, the existing deep-learning-based pansharpening methods not only lack information exchange and sharing between features of different resolutions but also cannot effectively use the residual information at different levels. These disadvantages may lead to the loss of spatial information and spectral information in the pansharpened image. To address the above problems, we propose a novel dual-stream convolutional neural network with residual information enhancement (DSCNN-RIE) for pansharpening. The proposed network is mainly composed of a set of dual-stream information complementation blocks (DSICBs), which can extract various spatial details at two different resolutions using convolutional filters of various sizes simultaneously, and can transfer complementary information effectively between two different resolutions. Furthermore, to improve the learning ability of the network and enhance the feature extraction, an RIE strategy is presented to stack different levels of residuals into the outputs of cascaded DSICBs. The final pansharpened image is obtained by integrating the extracted features using the shallow feature information of the source images. Experimental results on three datasets demonstrate that DSCNN-RIE outperforms ten other state-of-the-art pansharpening methods in both subjective and objective image-quality evaluations.
Yong Yang 0001, Wei Tu 0002, Shuying Huang, Hangyuan Lu, Weiguo Wan, Lixin Gan
IEEE Trans. Geosci. Remote. Sens.3
2022 A Unified Pansharpening Model Based on Band-Adaptive Gradient and Detail Correction
abstract
Pansharpening is used to fuse a panchromatic (PAN) image with a multispectral (MS) image to obtain a high-spatial-resolution multispectral (HRMS) image. Traditional pansharpening methods face difficulties in obtaining accurate details and have low computational efficiency. In this study, a unified pansharpening model based on the band-adaptive gradient and detail correction is proposed. First, a spectral fidelity constraint is designed by keeping each band of the HRMS image consistent with that of the MS image. Then, a band-adaptive gradient correction model is constructed by exploring the gradient relationship between a PAN image and each band of the MS image, so as to adaptively obtain an accurate spatial structure for the estimated HRMS image. To refine the spatial details, a detail correction constraint is defined based on the parameter transfer by designing a reduced-scale parameter acquisition model. Finally, a unified model is constructed based on the gradient and detail corrections, which is then solved by an alternating direction multiplier method. Both reduced-scale and full-scale experiments are conducted on several datasets. Compared with state-of-the-art pansharpening methods, the proposed method can achieve the best results in terms of fusion quality and has high efficiency. Specifically, our method improves the SAM and ERGAS metrics by 17.6% and 21.2% respectively compared to the traditional approach with the best average values, and improves these two metrics by 4.3% and 10.3% respectively compared to the learning-based approach with the best average values.
Hangyuan Lu, Yong Yang 0001, Shuying Huang, Wei Tu 0002, Weiguo Wan
IEEE Trans. Image Process.3
2022 End-to-End Rain Removal Network Based on Progressive Residual Detail Supplement
abstract
Methods of rain removal based on deep learning have rapidly developed, and the image quality after rain removal is continuously improving. However, the results of most methods have some common problems, including a loss of details, a blurring of edges, and the existence of artifacts. To remove rain-related information more thoroughly and retain more edge details, this paper proposes an end-to-end rain removal network based on the progressive residual detail supplement (ERRN-PRDS) approach. The entire network structure is designed in an iterative manner to obtain higher-quality rain removal images from coarse to fine. In the network, a diamond residual block is constructed as the main module of iteration to learn the feature information of the background layer. Meanwhile, to keep more texture details in the background layer, a detail supplement mechanism is designed between the iterative layers to transfer more information to the next iterative operation. Experimental results show that this method can remove the rain information more completely and better retain the image edges compared with previous state-of-the-art methods. In addition, because of the sparsity of the detail injection, our network also achieves high-quality results for image denoising tasks.
Yong Yang 0001, Juwei Guan, Shuying Huang, Weiguo Wan, Yating Xu
IEEE Trans. Multim.3
2021 Infrared and Visible Image Fusion Based on Modal Feature Fusion Network and Dual Visual Decision
abstract
Infrared and visible image fusion can integrate the complementary information of infrared and visible images to realize a more accurate scene interpretation. In this paper, a novel infrared and visible image fusion framework is proposed, which is based on a modal feature fusion network (MFFN) and a dual visual decision fusion module. Firstly, the infrared and visible images are decomposed by side window filtering to obtain the modal feature components, which can emphasize the target information of the source images. Secondly, MFFN is designed to merge the modal feature components to get a modal fused image. Then, a dual visual decision fusion module is built to obtain a supplement image which can provide more visual supplementary information for the modal fused image. Lastly, the final fusion result is achieved by combining the supplement image and the modal fused image. Experimental results show that the proposed method can generate fusion results with clearer targets and richer texture details, compared with other state-of-the-art fusion methods.
Yong Yang 0001, Shuying Huang, Weiguo Wan, Xiangkai Kong, Wang Zhang 0004
ICME3
2021 Infrared and Visible Image Fusion Based on Multiscale Network with Dual-channel Information Cross Fusion Block
abstract
The purpose of infrared and visible image fusion is to combine the complementary information of an infrared image and a visible image into a single image. In this paper, we propose an infrared and visible image fusion method based on dual-channel information cross fusion block (DICFB), which is developed to crossly extract and preliminarily fuse the multi-scale features of the source images. With the cascaded DICFB, we can obtain a series of fusion feature maps of the source images at different scales. Then, a progressive feature reconstruction module (PFRM) is designed to reconstruct the multi-scale fusion features to obtain the final fused image. Moreover, to better train the network, we design a joint loss function, in which a saliency map-based loss term is proposed to enhance the saliency targets in the fused images. Experimental results show that the proposed method has better performance than other state-of-the-art image fusion methods both objectively and subjectively.
Yong Yang 0001, Xiangkai Kong, Shuying Huang, Weiguo Wan, Wang Zhang 0004
IJCNN3
2021 An efficient and high-quality pansharpening model based on conditional random fields
Yong Yang 0001, Hangyuan Lu, Shuying Huang, Yuming Fang 0001, Wei Tu 0002
Inf. Sci.3
2021 Infrared and Visible Image Fusion via Texture Conditional Generative Adversarial Network
abstract
This paper proposes an effective infrared and visible image fusion method based on a texture conditional generative adversarial network (TC-GAN). The constructed TC-GAN generates a combined texture map for capturing gradient changes in image fusion. The generator in the TC-GAN is designed as a codec structure for extracting more details, and a squeeze-and-excitation module is applied to this codec structure to increase the weight of significant texture information in the combined texture map. The generator loss function is designed by combing the gradient loss and adversarial loss to retain the texture information of the source images. The discriminator brings the texture of the generated image closer to the visible image. To obtain significant texture information from the source images, a multiple decision map-based fusion strategy is proposed using a combined texture map and an adaptive guided filter. Extensive experiments on the public TNO and RoadScene datasets demonstrate that the proposed method is superior to other state-of-the-art algorithms in terms of a subjective evaluation and quantitative indicators.
Yong Yang 0001, Shuying Huang, Weiguo Wan, Wenying Wen, Juwei Guan
IEEE Trans. Circuits Syst. Video Technol.3
2020 An Efficient Pansharpening Method Based On Conditional Random Fields
abstract
Pansharpening is to fuse the existing low spatial resolution multi-spectral (MS) image with high spatial resolution panchromatic (PAN) image, so as to obtain high spatial resolution MS (HRMS) image. An efficient pansharpening model based on conditional random fields (CRFs) is proposed in this paper. In the model, a state feature function is designed to force the blurred HRMS image in accordance with the upsampled MS (UPMS) image to keep the spectral fidelity. Meanwhile, a transition feature function is defined to keep the sharpness of fused image. Besides, a new Gaussian filter acquisition algorithm is proposed to effectively satisfy the blur function in the model. To improve the efficiency of algorithm, a new initialization method based on fitting normal distribution is presented. Experiments are conducted on both reduced-scale images and full-scale images. Compared with some classical and state-of-art pansharpening methods, the proposed method achieves the best results in terms of fusion quality and efficiency.
Yong Yang 0001, Hangyuan Lu, Shuying Huang, Wei Tu 0002
ICME3
2020 Recognition of Tea Diseases under Natural Background Based on Particle Swarm Optimization Algorithm Optimized Support Vector Machine
abstract
In order to detect tea diseases quickly and meet the needs of practical production, a method was proposed for the classification and identification of tea diseases under natural conditions. According to the characteristic factors, the foreground main blade region was segmented, and then an improved image segmentation algorithm is used to segment each color component. The$\pmb{H}, \pmb{a}^{*},$and$\pmb{b}^{\star}$components were superimposed to obtain an ideal lesion area. A total of 34 features were extracted from three aspects of color, shape, and texture of lesion area. The first 9 principal components were selected as the modeling variables by principal component analysis. Diseases were classified by back propagation (BP) neural network, random forest, and support vector machine (SVM), respectively. The results show that the recognition rate of SVM classification algorithm is 87.3%, which is superior to BP neural network and random forest, and its operating speed is faster. In order to improve the recognition rate, particle swarm optimization (PSO) algorithm to optimize SVM was proposed, which improved the recognition rate of SVM to 91.7%. The experimental results show that PSO optimized SVM algorithm can quickly realize the classification and identification of tea diseases under natural background, and provide a valuable reference for the intelligent agricultural informatization of tea garden agriculture.
Xiuguo Zou, Shuying Huang, Jiaping Wang, Heyang Yao, Yuanyuan Song
INDIN3
2020 Remote Sensing Image Fusion Based on Fuzzy Logic and Salience Measure
abstract
Remote sensing image fusion is to fuse low spatial resolution multispectral (MS) images with high spatial resolution panchromatic (PAN) images to get high spatial resolution multispectral images. The component substitute (CS)-based methods are popular approaches for their high efficiency and high spatial resolution. However, they may produce spectral distortion, especially when there are large radiometric differences between PAN images and MS images. For tackling this problem, a new framework based on the CS model with fuzzy logic and salience measure is proposed. In order to get details that are highly relevant to MS image, a novel fuzzy logic rule based on the global salience measure is designed to fuse the details extracted from both PAN image and MS image. Furthermore, to better preserve the edges of the fused image, a new edge-preserving algorithm is defined to fuse the edges from the PAN image and MS image according to the local salience measure. A series of experiments are conducted and analyzed on both simulated images and real images from the data sets of four satellites to illustrate the effectiveness of the proposed method. Compared with some state-of-the-art methods, our method performs the best in both objective and subjective evaluations. Besides, our method has a low computational cost and is suitable for practical application.
Yong Yang 0001, Hangyuan Lu, Shuying Huang, Wei Tu 0002
IEEE Geosci. Remote. Sens. Lett.3
2019 Multilevel and Multiscale Network for Single-Image Super-Resolution
abstract
In recent years, deep convolutional neural networks (CNNs) have achieved great success in Single Image Super-Resolution (SISR). Most existing networks for Super-Resolution (SR) concentrate on wider or deeper network designs, leading to neglect of the feature correlations of intermediate layers. In this letter, a novel Multilevel and Multiscale Network for SISR (M2SR) is presented. The proposed network framework consists of four parts, including the feature extraction network, the cascade residual U-shaped blocks, the channel-wise attention U-shaped block and the fusion reconstruction network. Specially, the residual U-shaped blocks are designed to extract different scales of features, which are stacked to better refine the multifeatures. Then, to fully exploit the different levels of features, a channel-wise attention U-shaped block (At-U) is proposed to adjust the feature weights, which can adaptively enhance the feature expression and correlation learning. Finally, a fusion reconstruction network is constructed to fuse the different scales of the enhanced features to achieve the reconstructed result. Quantitative and qualitative evaluations of four public datasets show that the proposed method can achieve better performance compared with the state-of-the-art SR methods.
Yong Yang 0001, Shuying Huang, Jiajun Wu 0005
IEEE Signal Process. Lett.3
2019 Multimodal Medical Image Fusion Based on Fuzzy Discrimination With Structural Patch Decomposition
abstract
Multimodal medical image fusion, emerging as a hot topic, aims to fuse images with complementary multi-source information. In this paper, we propose a novel multimodal medical image fusion method based on structural patch decomposition (SPD) and fuzzy logic technology. First, the SPD method is employed to extract two salient features for fusion discrimination. Next, two novel fusion decision maps called an incomplete fusion map and supplemental fusion map are constructed from salient features. In this step, the supplemental map is constructed by our defined two different fuzzy logic systems. The supplemental and incomplete maps are then combined to construct an initial fusion map. The final fusion map is obtained by processing the initial fusion map with a Gaussian filter. Finally, a weighted average approach is adopted to create the final fused image. Additionally, an effective color medical image fusion scheme that can effectively prevent color distortion and obtain superior diagnostic effects is also proposed to enhance fused images. Experimental results clearly demonstrate that the proposed method outperforms state-of-the-art methods in terms of subjective visual and quantitative evaluations.
Yong Yang 0001, Jiahua Wu 0004, Shuying Huang, Yuming Fang 0001, Pan Lin, Yue Que 0001
IEEE J. Biomed. Health Informatics3
2018 Face Deduplication in Video Surveillance
abstract
The video surveillance system based on face analysis has played an increasingly important role in the security industry. Compared with identification methods of other physical characteristics, face verification method is easy to be accepted by people. In the video surveillance scene, it is common to capture multiple faces belonging to a same person. We cannot get a good result of face recognition if we use all the images without considering image quality. In order to solve this problem, we propose a face deduplication system which is combined with face detection and face quality evaluation to obtain the highest quality face image of a person. The experimental results in this paper also show that our method can effectively detect the faces and select the high-quality face images, so as to improve the accuracy of face recognition.
Ye Shen, Shuying Huang
Int. J. Pattern Recognit. Artif. Intell.5
2018 Compensation Details-Based Injection Model for Remote Sensing Image Fusion
abstract
Remote sensing image fusion has a potential spectral distortion problem due to the global/local spectral and spatial correlations between panchromatic (PAN) and multispectral (MS) images. To overcome this problem, in this letter, a compensation details-based injection (CDI) fusion model is presented from a new perspective of compensatory learning. In contrast to the traditional method, the two categories of details, namely, the PAN details and the CD, are considered to compensate for the spatial and spectral differences between low-resolution MS (LRMS) and high-resolution MS images. To obtain the CD, a robust sparse representation was employed to calculate the difference between the PAN and MS images during the fusion. The CD combined with the PAN details extracted by a multiscale-guided filter are then injected into the upsampled LRMS image to achieve a fused image. Extensive experiments were undertaken on several image data sets, and the results demonstrate the effectiveness of the proposed CDI method.
Yong Yang 0001, Shuying Huang, Jiancheng Sun, Weiguo Wan, Jiahua Wu 0004
IEEE Geosci. Remote. Sens. Lett.3