EDBT 2026 Demo / reviewers in the wild / expert
Weiguo Wan
dblp:185/7245
· DBLP profile ↗
37ranked-venue papers
5as first author
34since 2021 · last 2026
0000-0002-3537-979XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FSGNet: A frequency-aware and semantic guidance network for infrared small target detection
Yingmei Zhang 0001, Wangtao Bao, Yong Yang 0001, Weiguo Wan, Xueting Zou |
Expert Syst. Appl. | 4 |
| 2026 | GIW-MEF: Multi-exposure fusion via gated interaction and wavelet attention
Weiguo Wan |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Transformer embedded X-shaped encoding-decoding GAN for NIR-VIS face synthesis
Yue Que 0001, Jiyu Sun, Weiguo Wan, Tijian Cai, Yuejin Zhang |
Multim. Syst. | 3 |
| 2026 | PCFFusion: Progressive cross-modal feature fusion network for infrared and visible imagesabstractInfrared and visible image fusion (IVIF) aims to fuse thermal target information in infrared images and spatial texture information in visible images, improving the observability and comprehensibility of the fused images. Currently, most IVIF methods suffer from the loss of salient target information and texture details in fused images. To alleviate this problem, a progressive cross-modal feature fusion network (PCFFusion) for IVIF is proposed, which comprises two stages: feature extraction and feature fusion. In the feature extraction stage, to enhance the network’s feature representation capability, a feature decomposition module (FDM) is constructed to extract two modal features of different scales by defining a feature decomposition operation (FDO). In addition, by establishing correlations between the high- frequency and low-frequency components of two modal features, a cross-modal feature enhancement module (CMFEM) is built to realize correction and enhancement of the two features at each scale. The feature fusion stage achieves the fusion of two modal features at each scale and the supplementation of adjacent scale features by constructing three cross-domain fusion module (CDFMs). To constrain the fused results preserve more salient targets and richer texture details, a dual-feature fidelity loss function is defined by constructing a salient weight map to balance the two loss terms. Extensive experiments demonstrate that fusion results of the proposed method highlight prominent targets from infrared images while retaining rich background details from visible images, and the performance of PCFFusion is superior to some advanced methods. Specifically, compared to the optimal results obtained by other comparison methods, the proposed network achieves an average increase of 30.35 % and 10.9 % in metrics Mutual Information (MI) and Standard deviation (SD) on the TNO dataset, respectively. Shuying Huang, Yong Yang 0001, Weiguo Wan |
Pattern Recognit. | 4 |
| 2026 | EHEN: Eigendecomposition-based hyperchannel enhancement network for hyperspectral image super-resolution
Yong Yang 0001, Aoqi Zhao, Shuying Huang, Weiguo Wan |
Pattern Recognit. | 4 |
| 2026 | A mask protection-based and multi-probabilistic prior dictionary-guided unfolding model for pansharpening
Shengna Wei, Yong Yang 0001, Shuying Huang, Weiguo Wan, Changjie Chen 0002 |
Signal Process. | 4 |
| 2025 | FMPM-DNet: Hyperspectral Pansharpening Dynamic Network Based on Feature Modulation and Probability MaskabstractCurrently, most Hyperspectral (HS) pansharpening methods have two problems, namely the lack of consideration the spatial variations of HS images and inaccurate feature reconstruction in multi-channel complex mapping relationships, leading to spectral and spatial distortions in the fusion results. To address these issues, we propose a dynamic network based on feature modulation and probability mask (FMPM-DNet) for HS pansharpening, including two stages of spectral-spatial feature modulation and feature reconstruction. In the first stage, to increase the feature representation ability of the model, a wave function is defined based on complex transformation to convert spatial features into wave-like features. On this basis, considering the spatial variations of HS images, a dynamic feature modulation unit (DFMU) is constructed to achieve adaptive modulation and coarse fusion of features by dynamically generating spectral-spatial correction matrix. In the second stage, a feature probability mask unit (FPMU) is designed to realize global feature embedding at different depths and local feature embedding at the same depth to obtain refined fused features. Extensive experiments on three widely used datasets demonstrate that the proposed FMPM-Net achieves significant improvements in both spatial and spectral quality metrics compared to some state-of-the-art (SOTA) methods. Yong Yang 0001, Shuying Huang, Hangyuan Lu, Weiguo Wan, Aoqi Zhao |
AAAI | 5 |
| 2025 | Face Sketch Synthesis via Sparse Spatial Channel Generators and Multi-scale Discriminators
Weiguo Wan, Yingmei Zhang 0001, Benting Wan, Mingzhang Liu |
CGI (1) | 2 |
| 2025 | Multispectral-Hyperspectral Image Fusion via Similarity-Guided Graph Attention and VAE-TransformerabstractFusion of a high-spatial-resolution multispectral image (MSI) and a low-spatial-resolution hyperspectral image (HSI) aims to generate a high-spatial-resolution HSI (HR-HSI). Most fusion methods use simple upsampling techniques to increase the resolution of HSI without guidance, which can introduce unwanted artifacts and lead to spectral distortion. Additionally, they face challenges in generalization and robustness. To address these challenges, this paper introduces a cross-modal fusion network for MSI and HSI, named CSGAV, which is built on similarity-guided graph attention (SGA) and a variational autoencoder-Transformer (VAET). Specifically, we develop a similarity measure algorithm to compute the similarity degree between the source images and construct an SGA module to mitigate modal differences, producing precise upsampled outputs. Moreover, we present an adaptive weighted Transformer, with the weights guided by a variational autoencoder, thereby enhancing the generalization and robustness of the model. The SGA and VAET are integrated in the cross-modal interactive architecture to achieve the final HR-HSI image. Experimental results conducted on four public datasets show that CSGAV is superior compared to existing state-of-the-art fusion methods both in fusion performance and generalization. The code of this work is available at https://github.com/yotick/CSGAV. Biwei Chi, Hangyuan Lu, Rixian Liu, Yong Yang 0001, Lingrong Xu, Weiguo Wan |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | DFCFN: Dual-Stage Feature Correction Fusion Network for Hyperspectral PansharpeningabstractHyperspectral (HS) pansharpening aims to fuse high-spatial-resolution panchromatic (PAN) images with low-spatial-resolution hyperspectral (LRHS) images to generate high-spatial-resolution hyperspectral (HRHS) images. Due to the lack of consideration for the modal feature difference between PAN and LRHS images, most deep leaning-based methods suffer from spectral and spatial distortions in the fusion results. In addition, most methods use upsampled LRHS images as network input, resulting in spectral distortion. To address these issues, we propose a dual-stage feature correction fusion network (DFCFN) that achieves accurate fusion of PAN and LRHS images by constructing two fusion sub-networks: a feature correction compensation fusion network (FCCFN) and a multi-scale spectral correction fusion network (MSCFN). Based on the lattice filter structure, FCCFN is designed to obtain the initial fusion result by mutually correcting and supplementing the modal features from PAN and LRHS images. To suppress spectral distortion and obtain fine HRHS results, MSCFN based on 2D discrete wavelet transform (2D-DWT) is constructed to gradually correct the spectral features of the initial fusion result by designing a conditional entropy transformer (CE-Transformer). Extensive experiments on three widely used simulated datasets and one real dataset demonstrate that the proposed DFCFN achieves significant improvements in both spatial and spectral quality metrics over other state-of-the-art (SOTA) methods. Specifically, the proposed method improves the SAM metric by 6.4%, 6.2%, and 5.3% compared to the second-best comparison approach on Pavia center, Botswana, and Chikusei datasets, respectively. The codes are made available at: https://github.com/EchoPhD/DFCFN. Yong Yang 0001, Shuying Huang, Weiguo Wan, Long Zhang 0009, Aoqi Zhao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | HAFNet: Hierarchical Attention Fusion Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) involves identifying targets that are typically small in spatial extent, have low signal-to-clutter ratios, and are often embedded in dynamic and complex backgrounds, making the task particularly challenging. Benefiting from the powerful feature extraction and multiscale feature fusion characteristics, U-Net performs well in the IRSTD task. However, existing U-Net methods often focus solely on optimizing backbone feature extraction or skip connections, which limits their performance in complex scenes and makes it difficult to recognize small targets effectively. To address this limitation, we propose a novel hierarchical attention fusion network based on the U-Net architecture, namely HAFNet. Specifically, a dual-branch semantic perception module (DSPM) is designed as the feature extraction backbone to enhance contextual semantic interactions. This module integrates dual-branch feature extraction using standard and dilated convolutions while utilizing spatial and channel attention modules (CAMs) to effectively separate small targets from background noise. In addition, we extend the skip connection by merging a hierarchical feature fusion encoder (HFFE) and a hierarchical feature fusion decoder (HFFD). These modules utilize hierarchical attention-guided and encoded feature injection skip connections (ESCs) to achieve effective fusion of multiscale and multilevel semantic features between the encoder and decoder. Extensive experiments on three public datasets (NUAA-SIRST, IRSTD-1K, and NUDT-SIRST) demonstrate that the proposed HAFNet outperforms the existing IRSTD methods and achieves state-of-the-art (SOTA) detection performance. The code will be released onhttps://github.com/Wangtao-Bao/HAFNet Yingmei Zhang 0001, Wangtao Bao, Yong Yang 0001, Weiguo Wan, Xueting Zou |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | MFJLN: Multi-Frequency Feature Joint Learning Network for Rain Removal
Yong Yang 0001, Jiaxuan Yang, Shuying Huang, Weiguo Wan |
IEEE Trans. Multim. | 4 |
| 2024 | Transformer-Based adversarial network for semi-supervised face sketch synthesis
Zhihua Shi, Weiguo Wan |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | MFITN: A Multilevel Feature Interaction Transformer Network for PansharpeningabstractIn this letter, to better supplement the advantages of features at different levels and improve the feature extraction ability of the network, a novel multi-level feature interaction transformer network (MFITN) is proposed for pansharpening, aiming to fuse multispectral (MS) and panchromatic (PAN) images. In MFITN, a multi-level feature interaction transformer encoding module is designed to extract and correct global multi-level features by considering the modality difference between source images. These features are then fused using the proposed multi-level feature mixing (MFM) operation, which enables features to fuse interactively to obtain richer information. Furthermore, the global features are fed into a CNN-based local decoding module to better reconstruct high-spatial-resolution multispectral (HRMS) images. Additionally, based on the spatial consistency between MS and PAN images, a band compression loss is defined to improve the fidelity of fused images. Numerous simulated and real experiments demonstrate that the proposed method has the optimal performance compared to state-of-the-art methods. Specifically, the proposed method improves the SAM metric by 7.89% and 6.41% compared to the second-best comparison approach on Pléiades and WorldView-3, respectively. Changjie Chen 0002, Yong Yang 0001, Shuying Huang, Hangyuan Lu, Weiguo Wan, Shengna Wei, Wenying Wen |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Pansharpening Based on Fuzzy Logic and Edge ActivityabstractPansharpening technology aims to extract spatial information from a panchromatic (PAN) image and integrate it into a multispectral (MS) image to generate a high-spatial resolution MS image. To overcome the problem of spatial and spectral distortion of traditional methods, this letter presents a pansharpening method based on fuzzy logic and edge activity. To obtain more accurate spatial details of source images, details extracted using the component substitution and multiresolution analysis methods are fused via the proposed fuzzy logic algorithm. To preserve the edges of the fusion result, the edge maps of the source images are detected and fused based on the edge activity. The optimized detail maps are obtained by multiplying the fused details and edge maps, which are then injected into the upsampled MS image to obtain the final pansharpened image. Reduced- and full-scale experimental results on the Pléiades and IKONOS datasets demonstrate the effectiveness of our method compared with state-of-the-art pansharpening algorithms. Specifically, the proposed method improves the ERGAS metrics by 9.0% and 11.1% compared to the second-best comparison approach on Pléiades and IKONOS, respectively. Yong Yang 0001, Shuying Huang, Weiguo Wan, Hangyuan Lu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Unsupervised masked face inpainting based on contrastive learning and attention mechanism
Weiguo Wan, Shunming Chen, Yingmei Zhang 0001 |
Multim. Syst. | 1 |
| 2024 | LSRN-AED: lightweight super-resolution network based on asymmetric encoder-decoder
Shuying Huang, Yong Yang 0001, Weiguo Wan, Houzeng Lai |
Soft Comput. | 4 |
| 2024 | Denoising Diffusion Probabilistic Model for Face Sketch-to-Photo SynthesisabstractThe field of face sketch-to-photo synthesis involves generating photographic facial images with enhanced details and a heightened sense of style realism. In recent years, the advancement of deep learning techniques has significantly contributed to the development of methods for synthesizing photographic face images from sketches. Nevertheless, challenges remain in synthesizing facial photographs with richer details and more accurate structural representation. This paper introduces a novel architecture for face sketch-to-photo synthesis, using denoising diffusion probabilistic models (DDPM). Our approach simplifies the complex transformation process into sequential forward and backward denoising steps. We incorporate a pretrained coarse generator to effectively encode sketch information, integrating it into each backward step to guide the generative process toward accurate photo space representation. Furthermore, we design a detail diffusion branch to refine the coarse photo face generated from the coarse generator. By deeply fusing multiscale detail features from this branch with a sophisticated conditional noise predictor, our model effectively captures the correlation between detail and stylistic elements both in sketches and in photographic faces. Extensive experimental evaluations on three datasets show the effectiveness of our model, emphasizing its ability to synthesize facial photographs with remarkable realism and rich detail. The synthesized facial images consistently demonstrate superior face recognition accuracy, surpassing that of state-of-the-art methods. Yue Que 0001, Li Xiong 0018, Weiguo Wan, Xue Xia 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | VSDM: Variable-Scale Diffusion Model Based on Dynamic Condition Guidance for PansharpeningabstractPansharpening aims to obtain a high-spatial-resolution multispectral (MS) image by fusing a lower-spatial resolution MS image with a high-spatial-resolution panchromatic (PAN) image. Currently, the results obtained by most pansharpening methods still suffer from spatial and spectral distortion issues. The diffusion model has shown outstanding performance in various image-processing tasks. However, maintaining the full image size throughout the diffusion process imposes a large computational burden, and the simultaneous use of PAN and MS images acquired by different sensors as a condition for guiding noise prediction leads to spatial and spectral distortions. To solve these problems, a variable-scale diffusion model (VSDM) based on dynamic condition guidance for pansharpening is proposed, which achieves better fusion performance by improving the diffusion manner of the diffusion model and injecting dynamic conditions to guide the reverse process. In VSDM, a variable-scale diffusion manner (VSDMN) is designed to reduce the computational complexity of the model by reducing the size of the image in the diffusion process. A condition generator (CG) is constructed to generate dynamic conditions using the features learned from the PAN and upsampled MS images. In CG, a cross-attention dynamic convolution is built to extract features from the PAN image by designing a spatial and spectral attention mechanism, which can improve the spatial and spectral consistency in the dynamic condition. Extensive experiments validate the effectiveness of the proposed VSDM against other state-of-the-art (SOTA) pansharpening methods in both quantitative and qualitative assessments. The source code will be released athttps://github.com/MELiMZ/VSDM. Yong Yang 0001, Shuying Huang, Weiguo Wan, Hangyuan Lu, Wei Tu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Low-Light Image Enhancement Network Based on Multi-Scale Feature ComplementationabstractImages captured in low-light environments have problems of insufficient brightness and low contrast, which will affect subsequent image processing tasks. Although most current enhancement methods can obtain high-contrast images, they still suffer from noise amplification and color distortion. To address these issues, this paper proposes a low-light image enhancement network based on multi-scale feature complementation (LIEN-MFC), which is a U-shaped encoder-decoder network supervised by multiple images of different scales. In the encoder, four feature extraction branches are constructed to extract features of low-light images at different scales. In the decoder, to ensure the integrity of the learned features at each scale, a feature supplementary fusion module (FSFM) is proposed to complement and integrate features from different branches of the encoder and decoder. In addition, a feature restoration module (FRM) and an image reconstruction module (IRM) are built in each branch to reconstruct the restored features and output enhanced images. To better train the network, a joint loss function is defined, in which a discriminative loss term is designed to ensure that the enhanced results better meet the visual properties of the human eye. Extensive experiments on benchmark datasets show that the proposed method outperforms some state-of-the-art methods subjectively and objectively. Yong Yang 0001, Wenzhi Xu, Shuying Huang, Weiguo Wan |
AAAI | 4 |
| 2023 | MMPN: Multi-supervised Mask Protection Network for PansharpeningabstractPansharpening is to fuse a panchromatic (PAN) image with a multispectral (MS) image to obtain a high-spatial-resolution multispectral (HRMS) image. The deep learning-based pansharpening methods usually apply the convolution operation to extract features and only consider the similarity of gradient information between PAN and HRMS images, resulting in the problems of edge blur and spectral distortion in the fusion results. To solve this problem, a multi-supervised mask protection network (MMPN) is proposed to prevent spatial information from being damaged and overcome spectral distortion in the learning process. Firstly, by analyzing the relationships between high-resolution images and corresponding degraded images, a mask protection strategy (MPS) for edge protection is designed to guide the recovery of fused images. Then, based on the MPS, an MMPN containing four branches is constructed to generate the fusion and mask protection images. In MMPN, each branch employs a dual-stream multi-scale feature fusion module (DMFFM), which is built to extract and fuse the features of two input images. Finally, different loss terms are defined for the four branches, and combined into a joint loss function to realize network training. Experiments on simulated and real satellite datasets show that our method is superior to state-of-the-art methods both subjectively and objectively. Changjie Chen 0002, Yong Yang 0001, Shuying Huang, Wei Tu 0002, Weiguo Wan, Shengna Wei |
IJCAI | 5 |
| 2023 | CTCP: Cross Transformer and CNN for PansharpeningabstractPansharpening is to fuse a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to obtain an enhanced LRMS image with high spectral and spatial resolution. The current Transformer-based pansharpening methods neglect the interaction between the extracted long- and short-range features, resulting in spectral and spatial distortion in the fusion results. To address this issue, a novel cross Transformer and convolutional neural network (CNN) for pansharpening (CTCP) is proposed to achieve better fusion results by designing a cross mechanism, which can enhance the interaction between long- and short-range features. First, a dual branch feature extraction module (DBFEM) is constructed to extract the features from the LRMS and PAN images, respectively, reducing the aliasing of the two image features. In the DBFEM, to improve the feature representation ability of the network, a cross long-short-range feature module (CLSFM) is designed by combining the feature learning capabilities of Transformer and CNN via the cross mechanism, which achieves the integration of long-short-range features. Then, to improve the ability of spectral feature representation, a spectral feature enhancement fusion module (SFEFM) based on a frequency channel attention is constructed to realize feature fusion. Finally, the shallow features from the PAN image are reused to provide detail features, which are integrated with the fused features to obtain the final pansharpened results. To the best of our knowledge, this is the first attempt to introduce the cross mechanism between Transformer and CNN in pansharpening field. Numerous experiments show that our CTCP outperforms some state-of-the-art (SOTA) approaches both subjectively and objectively. The source code will be released at https://github.com/zhsu99/CTCP. Zhao Su, Yong Yang 0001, Shuying Huang, Weiguo Wan, Wei Tu 0002, Hangyuan Lu, Changjie Chen 0002 |
ACM Multimedia | 4 |
| 2023 | Multi-scale Spatial-Spectral Attention Guided Fusion Network for PansharpeningabstractPansharpening is to fuse high-resolution panchromatic (PAN) images with low-resolution multispectral (LR-MS) images to generate high-resolution multispectral (HR-MS) images. Most of the deep learning-based pansharpening methods did not consider the inconsistency of the PAN and LR-MS images and used simple concatenation to fuse the source images, which may cause spectral and spatial distortion in the fused results. To address this problem, a multi-scale spatial-spectral attention guided fusion network for pansharpening is proposed. First, the spatial features from the PAN image and spectral features from the LR-MS image are independently extracted to obtain the shallow features. Then, a spatial-spectral attention feature fusion module (SAFFM) is constructed to guide the reconstruction of spatial-spectral features by generating a guidance map to achieve the fusion of reconstructed features at different scales. In SAFFM, the guidance map is designed to ensure the spatial-spectral consistency of the reconstructed features. Finally, considering the difference between multiply scale features, a multi-level feature integration scheme is proposed to progressively achieve fusion of multi-scale features from different SAFFMs. Extensive experiments validate the effectiveness of the proposed network against other state-of-the-art (SOTA) pansharpening methods in both quantitative and qualitative assessments. The source code will be released at https://github.com/MELiMZ/ssaff. Yong Yang 0001, Shuying Huang, Hangyuan Lu, Wei Tu 0002, Weiguo Wan |
ACM Multimedia | 6 |
| 2023 | FRAN: feature-filtered residual attention network for realistic face sketch-to-photo transformation
Weiguo Wan, Yong Yang 0001, Shuying Huang, Lixin Gan |
Appl. Intell. | 1 |
| 2023 | STCP: Synergistic Transformer and Convolutional Neural Network for PansharpeningabstractPansharpening is a process of fusing a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to obtain a high-resolution multispectral (HRMS) image. Convolutional neural networks (CNNs) have been commonly utilized in this field because of their remarkable learning capabilities. However, their convolutional operators limit the long-range feature extraction ability of CNN. Meanwhile, the Transformer models have exhibited strong capabilities in modeling long-range representations, but there are shortcomings in modeling local-range feature dependencies. To this end, we propose a novel synergistic transformer and CNN for pansharpening (STCP). First, a parallel U-shaped feature extraction module (PUFEM) is constructed for extracting the features of the LRMS and PAN images, which improves the feature representation ability for the two source images. In the PUFEM, combining the different feature learning capabilities of the CNN and transformer, we design a long-short-range feature integration block (LSFIB) to extract the short-range features and long-range features at different scales in parallel. Then, a channel attention module (CAM)-based feature fusion module (CFFM) is constructed to integrate the features extracted by the PUFEM. Finally, the shallow features from the PAN image are reused to provide detailed features, which are integrated with the fused features from the CFFM to achieve the final pansharpened results. Numerous experiments show that our STCP outperforms some state-of-the-art approaches both subjectively and objectively. Zhao Su, Yong Yang 0001, Shuying Huang, Weiguo Wan, Jiancheng Sun, Wei Tu 0002, Changjie Chen 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | DCNP: Dual-Information Compensation Network for PansharpeningabstractTo reduce the loss of detail and spectral information during the network propagation and better extract detail and spectral features at different scales, a novel dual-information compensation network for pansharpening (DCNP) is proposed for fusing multispectral (MS) and panchromatic (PAN) images. In the network, the domain-specific knowledge is considered to design our DCNP architecture by focusing on the two aims of the pansharpening: spatial and spectral preservation. Specifically, a cascaded U-shaped structure is constructed to improve the feature representation ability of the network. To preserve more spatial details in the pansharpened image, the details of the PAN image are extracted based on Laplacian operator and then compensated into the network. Furthermore, for spectral preservation, the MS image is conducted by the transposed convolution as the compensation information of the network. Experiments on the full- and reduced-scale data indicate that the proposed DCNP achieves significant improvement over state-of-the-art methods in terms of subjective and objective evaluation. Specifically, DCNP improves the PSNR and ERGAS metrics by 12.9% and 44.5% respectively compared to the deep learning-based approach with the best average values on Pléiades. Yong Yang 0001, Zhao Su, Shuying Huang, Weiguo Wan, Wei Tu 0002, Changjie Chen 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | MMDN: Multi-Scale and Multi-Distillation Dilated Network for PansharpeningabstractPansharpening is a technology involving information integration and processing in remote sensing imagery. It is applied to generate a high-resolution multispectral (HRMS) image through an effective fusion of a low spatial resolution multispectral image and a panchromatic (PAN) image. In this paper, we propose an end-to-end multi-scale and multi-distillation dilated network (MMDN) for pansharpening. In MMDN, to extract more abundant spatial details from source images, a clique structure-based multi-scale dilated block (CSMDB) is presented. The clique structure in CSMDB can fully transfer the information between feature maps obtained by the multi-scale dilated convolutional filters. Then, a multi-distillation residual information block (MRIB) is constructed to help the network capture the spatial structure of different scales in MS and PAN images. Finally, to reuse and supplement the feature information, a feature embedding strategy is designed by feeding the sum result of the output of cascaded CSMDBs and the shallow features to each MRIB. Experimental results verify that the proposed MMDN outperforms other compared state-of-the-art approaches in terms of objective and subjective evaluations. Wei Tu 0002, Yong Yang 0001, Shuying Huang, Weiguo Wan, Lixin Gan, Hangyuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dual-Stream Convolutional Neural Network With Residual Information Enhancement for PansharpeningabstractDeep-learning-based pansharpening methods have achieved remarkable results due to their powerful feature representation ability. However, the existing deep-learning-based pansharpening methods not only lack information exchange and sharing between features of different resolutions but also cannot effectively use the residual information at different levels. These disadvantages may lead to the loss of spatial information and spectral information in the pansharpened image. To address the above problems, we propose a novel dual-stream convolutional neural network with residual information enhancement (DSCNN-RIE) for pansharpening. The proposed network is mainly composed of a set of dual-stream information complementation blocks (DSICBs), which can extract various spatial details at two different resolutions using convolutional filters of various sizes simultaneously, and can transfer complementary information effectively between two different resolutions. Furthermore, to improve the learning ability of the network and enhance the feature extraction, an RIE strategy is presented to stack different levels of residuals into the outputs of cascaded DSICBs. The final pansharpened image is obtained by integrating the extracted features using the shallow feature information of the source images. Experimental results on three datasets demonstrate that DSCNN-RIE outperforms ten other state-of-the-art pansharpening methods in both subjective and objective image-quality evaluations. Yong Yang 0001, Wei Tu 0002, Shuying Huang, Hangyuan Lu, Weiguo Wan, Lixin Gan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Unified Pansharpening Model Based on Band-Adaptive Gradient and Detail CorrectionabstractPansharpening is used to fuse a panchromatic (PAN) image with a multispectral (MS) image to obtain a high-spatial-resolution multispectral (HRMS) image. Traditional pansharpening methods face difficulties in obtaining accurate details and have low computational efficiency. In this study, a unified pansharpening model based on the band-adaptive gradient and detail correction is proposed. First, a spectral fidelity constraint is designed by keeping each band of the HRMS image consistent with that of the MS image. Then, a band-adaptive gradient correction model is constructed by exploring the gradient relationship between a PAN image and each band of the MS image, so as to adaptively obtain an accurate spatial structure for the estimated HRMS image. To refine the spatial details, a detail correction constraint is defined based on the parameter transfer by designing a reduced-scale parameter acquisition model. Finally, a unified model is constructed based on the gradient and detail corrections, which is then solved by an alternating direction multiplier method. Both reduced-scale and full-scale experiments are conducted on several datasets. Compared with state-of-the-art pansharpening methods, the proposed method can achieve the best results in terms of fusion quality and has high efficiency. Specifically, our method improves the SAM and ERGAS metrics by 17.6% and 21.2% respectively compared to the traditional approach with the best average values, and improves these two metrics by 4.3% and 10.3% respectively compared to the learning-based approach with the best average values. Hangyuan Lu, Yong Yang 0001, Shuying Huang, Wei Tu 0002, Weiguo Wan |
IEEE Trans. Image Process. | 5 |
| 2022 | End-to-End Rain Removal Network Based on Progressive Residual Detail SupplementabstractMethods of rain removal based on deep learning have rapidly developed, and the image quality after rain removal is continuously improving. However, the results of most methods have some common problems, including a loss of details, a blurring of edges, and the existence of artifacts. To remove rain-related information more thoroughly and retain more edge details, this paper proposes an end-to-end rain removal network based on the progressive residual detail supplement (ERRN-PRDS) approach. The entire network structure is designed in an iterative manner to obtain higher-quality rain removal images from coarse to fine. In the network, a diamond residual block is constructed as the main module of iteration to learn the feature information of the background layer. Meanwhile, to keep more texture details in the background layer, a detail supplement mechanism is designed between the iterative layers to transfer more information to the next iterative operation. Experimental results show that this method can remove the rain information more completely and better retain the image edges compared with previous state-of-the-art methods. In addition, because of the sparsity of the detail injection, our network also achieves high-quality results for image denoising tasks. Yong Yang 0001, Juwei Guan, Shuying Huang, Weiguo Wan, Yating Xu |
IEEE Trans. Multim. | 4 |
| 2021 | Infrared and Visible Image Fusion Based on Modal Feature Fusion Network and Dual Visual DecisionabstractInfrared and visible image fusion can integrate the complementary information of infrared and visible images to realize a more accurate scene interpretation. In this paper, a novel infrared and visible image fusion framework is proposed, which is based on a modal feature fusion network (MFFN) and a dual visual decision fusion module. Firstly, the infrared and visible images are decomposed by side window filtering to obtain the modal feature components, which can emphasize the target information of the source images. Secondly, MFFN is designed to merge the modal feature components to get a modal fused image. Then, a dual visual decision fusion module is built to obtain a supplement image which can provide more visual supplementary information for the modal fused image. Lastly, the final fusion result is achieved by combining the supplement image and the modal fused image. Experimental results show that the proposed method can generate fusion results with clearer targets and richer texture details, compared with other state-of-the-art fusion methods. Yong Yang 0001, Shuying Huang, Weiguo Wan, Xiangkai Kong, Wang Zhang 0004 |
ICME | 4 |
| 2021 | Infrared and Visible Image Fusion Based on Multiscale Network with Dual-channel Information Cross Fusion BlockabstractThe purpose of infrared and visible image fusion is to combine the complementary information of an infrared image and a visible image into a single image. In this paper, we propose an infrared and visible image fusion method based on dual-channel information cross fusion block (DICFB), which is developed to crossly extract and preliminarily fuse the multi-scale features of the source images. With the cascaded DICFB, we can obtain a series of fusion feature maps of the source images at different scales. Then, a progressive feature reconstruction module (PFRM) is designed to reconstruct the multi-scale fusion features to obtain the final fused image. Moreover, to better train the network, we design a joint loss function, in which a saliency map-based loss term is proposed to enhance the saliency targets in the fused images. Experimental results show that the proposed method has better performance than other state-of-the-art image fusion methods both objectively and subjectively. Yong Yang 0001, Xiangkai Kong, Shuying Huang, Weiguo Wan, Wang Zhang 0004 |
IJCNN | 4 |
| 2021 | Generative adversarial learning for detail-preserving face sketch synthesisabstractFace sketch synthesis aims to generate a face sketch image from a corresponding photo image and has wide applications in law enforcement and digital entertainment. Despite the remarkable achievements that have been made in face sketch synthesis, most existing works pay main attention to the facial content transfer, at the expense of facial detail information. In this paper, we present a new generative adversarial learning framework to focus on detail preservation for realistic face sketch synthesis. Specifically, the high-resolution network is modified as generator to transform a face image from photograph to sketch domain. Except for the common adversarial loss, we design a detail loss to force the synthesized face sketch images have proximate details to its corresponding photo images. In addition, the style loss is adopted to restrain the synthesized face sketch images have vivid sketch style as the hand-drawn sketch images. Experimental results demonstrate that the proposed approach achieves superior performance, compared to state-of-the-art approaches, both on visual perception and objective evaluation. Specifically, this study indicated the higher FSIM values (0.7345 and 0.7080) and Scoot values (0.5317 and 0.5091) than most comparison methods on the CUFS and CUFSF datasets, respectively. Weiguo Wan, Yong Yang 0001, Hyo Jong Lee |
Neurocomputing | 1 |
| 2021 | Infrared and Visible Image Fusion via Texture Conditional Generative Adversarial NetworkabstractThis paper proposes an effective infrared and visible image fusion method based on a texture conditional generative adversarial network (TC-GAN). The constructed TC-GAN generates a combined texture map for capturing gradient changes in image fusion. The generator in the TC-GAN is designed as a codec structure for extracting more details, and a squeeze-and-excitation module is applied to this codec structure to increase the weight of significant texture information in the combined texture map. The generator loss function is designed by combing the gradient loss and adversarial loss to retain the texture information of the source images. The discriminator brings the texture of the generated image closer to the visible image. To obtain significant texture information from the source images, a multiple decision map-based fusion strategy is proposed using a combined texture map and an adaptive guided filter. Extensive experiments on the public TNO and RoadScene datasets demonstrate that the proposed method is superior to other state-of-the-art algorithms in terms of a subjective evaluation and quantitative indicators. Yong Yang 0001, Shuying Huang, Weiguo Wan, Wenying Wen, Juwei Guan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Generative Adversarial Multi-Task Learning for Face Sketch Synthesis and RecognitionabstractFace sketch synthesis and recognition have wide range of applications in law enforcement. Despite the impressive progresses have been made in faces sketch and recognition, most existing researches regard them as two separate tasks. In this paper, we propose a generative adversarial multitask learning method in order to deal with face sketch synthesis and recognition simultaneously. Our framework is based on generative adversarial networks (GAN), in which an improved deep network named residual dense U-Net is used as generator to synthesize face sketch image and a multi-task discriminator is designed to not only guide the generator to produce more realistic sketch image, but also extract discriminative face feature. In addition, except the common adversarial loss, the perceptual loss and triplet loss are adopted for the learning of generator and discriminator, respectively. Compared with the state-of-the-art methods, the proposed method obtains better results in terms of face sketch synthesis and recognition. Weiguo Wan, Hyo Jong Lee |
ICIP | 1 |
| 2019 | Transfer deep feature learning for face sketch recognition
Weiguo Wan, Yongbin Gao, Hyo Jong Lee |
Neural Comput. Appl. | 1 |
| 2018 | Compensation Details-Based Injection Model for Remote Sensing Image FusionabstractRemote sensing image fusion has a potential spectral distortion problem due to the global/local spectral and spatial correlations between panchromatic (PAN) and multispectral (MS) images. To overcome this problem, in this letter, a compensation details-based injection (CDI) fusion model is presented from a new perspective of compensatory learning. In contrast to the traditional method, the two categories of details, namely, the PAN details and the CD, are considered to compensate for the spatial and spectral differences between low-resolution MS (LRMS) and high-resolution MS images. To obtain the CD, a robust sparse representation was employed to calculate the difference between the PAN and MS images during the fusion. The CD combined with the PAN details extracted by a multiscale-guided filter are then injected into the upsampled LRMS image to achieve a fused image. Extensive experiments were undertaken on several image data sets, and the results demonstrate the effectiveness of the proposed CDI method. Yong Yang 0001, Shuying Huang, Jiancheng Sun, Weiguo Wan, Jiahua Wu 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |