Yingying Wang 0005

dblp:87/6339-5 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
24since 2021 · last 2026
0009-0004-3748-1191ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
abstract
Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks.
Yingying Wang 0005, Xuanhua He, Jialing Huang, Suiyun Zhang, Xinghao Ding, Haoxuan Che
AAAI1
2026 Self-supervised Multiplex Consensus Mamba for General Image Fusion
abstract
Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks.
Yingying Wang 0005, Rongjin Zhuang, Hui Zheng 0003, Xuanhua He, Ke Cao 0001, Xiaotong Tu, Xinghao Ding
AAAI1
2025 AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image Enhancement
abstract
Existing low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often impractical and sometimes even unavailable, and 2) accurately modeling complex degradations presents a non-trivial problem. To overcome them, we propose the Attribute Guidance Diffusion framework (AGLLDiff), a training-free method for effective real-world LIE. Instead of specifically defining the degradation process, AGLLDiff shifts the paradigm and models the desired attributes, such as image exposure, structure and color of normal-light images. These attributes are readily available and impose no assumptions about the degradation process, which guides the diffusion sampling process to a reliable high-quality solution space. Extensive experiments demonstrate that our approach outperforms the current leading unsupervised LIE methods across benchmarks in terms of distortion-based and perceptual-based metrics, and it performs well even in sophisticated wild degradation.
Yunlong Lin, Tian Ye 0001, Sixiang Chen, Zhenqi Fu, Yingying Wang 0005, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003, Xinghao Ding
AAAI5
2025 DPLUT: Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion Priors
abstract
Low-light image enhancement (LIE) aims at precisely and efficiently recovering an image degraded in poor illumination environments. Recent advanced LIE techniques are using deep neural networks, which require lots of low-normal light image pairs, network parameters, and computational resources. As a result, their practicality is limited. In this work, we devise a novel unsupervised LIE framework based on diffusion priors and lookup tables (DPLUT) to achieve efficient low-light image recovery. The proposed approach comprises two critical components: a light adjustment lookup table (LLUT) and a noise suppression lookup table (NLUT). LLUT is optimized with a set of unsupervised losses. It aims at predicting pixel-wise curve parameters for the dynamic range adjustment of a specific image. NLUT is designed to remove the amplified noise after the light brightens. As diffusion models are sensitive to noise, diffusion priors are introduced to achieve high-performance noise suppression. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of visual quality and efficiency.
Yunlong Lin, Zhenqi Fu, Kairun Wen, Tian Ye 0001, Sixiang Chen, Ge Meng, Yingying Wang 0005, Chui Kong, Yue Huang 0001, Xiaotong Tu, Xinghao Ding
AAAI7
2025 Accelerated Diffusion via High-Low Frequency Decomposition for Pan-Sharpening
abstract
Pan-sharpening aims to preserve the spectral information of the multi-spectral (MS) image while leveraging the high-frequency details from the guided high-resolution panchromatic (PAN) image to enhance its spatial resolution. The key challenge is how to preserve the spectral information from the MS image and the spatial details from the PAN image as much as possible. Diffusion models have achieved favorable results in image restoration and synthesis tasks but suffer from excessive computational resource and time consumption. In this paper, we design a novel and computationally efficient diffusion-based pan-sharpening network that achieves accelerated diffusion while reducing task complexity by decoupling the high and low-frequency components of the fused image. Specifically, leveraging the information-preserving characteristic of the wavelet transformation, we introduce a Wavelet-based Low-frequency Diffusion Model (WLDM). WLDM generates the low-frequency coefficient of high-resolution MS (HRMS) image from the low-resolution MS (LRMS) image. This approach significantly reduces computational resources and complexity compared to the direct restoration of the HRMS image. Furthermore, we have devised a High-frequency Information Restoration Module (HIRM) to restore the high-frequency information in the HRMS image through the interaction of high-frequency coefficients from the PAN image in three directions. Extensive experiments on three different datasets demonstrate that our method outperforms existing approaches in both quantitative metrics, qualitative metrics, and inference efficiency.
Ge Meng, Jingjia Huang, Jingyan Tu, Yingying Wang 0005, Yunlong Lin, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
AAAI4
2025 Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image Reconstruction
abstract
Hyperspectral image (HSI) reconstruction aims to restore the original 3D HSIs from the 2D hyperspectral snapshot compressive images (SCIs). The key to high-fidelity HSI reconstruction lies in designing refined spatial and spectral attention mechanisms, which are crucial for generating fine-grained representations of HSI based on the limited spatial and spectral information available in SCI. Recently, Mamba has demonstrated remarkable performance and efficiency in modeling spatial correlations. Its implicit attention mechanism generates three orders of magnitude more attention matrices than transformers, significantly raising the performance ceiling for HSI reconstruction. In this paper, we propose a novel joint SSM network named Sp3ctralMamba for HSI reconstruction. Sp3ctralMamba integrates frequency domain knowledge and physical priors to enhance reconstruction quality. Specifically, we first perform hierarchical decomposition of the 3D HSI embedding to mitigate the negative impact of distant bands on reconstruction. Next, we design a joint SSM block S3Mamba (S3MAB) to perform parallel scans of the embeddings from different bands. In addition to the conventional vanilla scan, S3MAB introduces a local scanning scheme to address the reconstruction challenges posed by the spatial sparsity of spectral information. Furthermore, a spiral scanning scheme in the frequency domain is incorporated to enhance the order correlation between different frequency signals. Finally, we introduce energy priors and structural priors to constrain the generation of spectral and spatial representations during the training process. Extensive experiments on both simulated and real datasets demonstrate that Sp3ctralMamba significantly elevates HSI reconstruction performance to a new level, surpassing SOTA methods in both quantitative and qualitative metrics.
Ge Meng, Jingyan Tu, Jingjia Huang, Yunlong Lin, Yingying Wang 0005, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
AAAI5
2025 Freq-RWKV: Granularity-Aware Spatial-Frequency Synergy via Dual-Domain Recurrent Scanning for Pan-sharpening
Xueheng Li, Xuanhua He, Tao Hu 0027, Jie Zhang 0033, Man Zhou 0003, Chengjun Xie, Yingying Wang 0005, Bo Huang 0001
ACM Multimedia7
2025 Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables
abstract
Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicability of these methods in real-world scenarios, particularly in the absence of dedicated computing devices such as GPUs and TPUs. To address these challenges, we propose Pan-LUT, a novel learnable look-up table (LUT) framework for pan-sharpening that strikes a balance between performance and computational efficiency for large remote sensing images. Our method makes it possible to process 15K$\times$15K remote sensing images on a 24GB GPU. To finely control the spectral transformation, we devise the PAN-guided look-up table (PGLUT) for channel-wise spectral mapping. To effectively capture fine-grained spatial details, we introduce the spatial details look-up table (SDLUT). Furthermore, to adaptively aggregate channel information for generating high-resolution multispectral images, we design an adaptive output look-up table (AOLUT). Our model contains fewer than 700K parameters and processes a 9K$\times$9K image in under 1 ms using one RTX 2080 Ti GPU, demonstrating significantly faster performance compared to other methods. Experiments reveal that Pan-LUT efficiently processes large remote sensing images in a lightweight manner, bridging the gap to real-world applications. Furthermore, our model surpasses SOTA methods in full-resolution scenes under real-world conditions, highlighting its effectiveness and efficiency. We also extend our method to general image fusion tasks.
Zhongnan Cai, Yingying Wang 0005, Hui Zheng 0003, Panwang Pan, Zixu Lin, Ge Meng, Chenxin Li, Chunming He, Jiaxin Xie, Yunlong Lin, Junbin Lu, Yue Huang 0001, Xinghao Ding
NeurIPS2
2025 FRN: Fractal-Based Recursive Spectral Reconstruction Network
abstract
Generating hyperspectral images (HSIs) from RGB images through spectral reconstruction can significantly reduce the cost of HSI acquisition. In this paper, we propose a Fractal-Based Recursive Spectral Reconstruction Network (FRN), which differs from existing paradigms that attempt to directly integrate the full-spectrum information from the R, G, and B channels in a one-shot manner. Instead, it treats spectral reconstruction as a progressive process, predicting from broad to narrow bands or employing a coarse-to-fine approach for predicting the next wavelength. Inspired by fractals in mathematics, FRN establishes a novel spectral reconstruction paradigm by recursively invoking an atomic reconstruction module. In each invocation, only the spectral information from neighboring bands is used to provide clues for the generation of the image at the next wavelength, which follows the low-rank property of spectral data. Moreover, we design a band-aware state space model that employs a pixel-differentiated scanning strategy at different stages of the generation process, further suppressing interference from low-correlation regions caused by reflectance differences. Through extensive experimentation across different datasets, FRN achieves superior reconstruction performance compared to state-of-the-art methods. Code is available at https://github.com/mongko007/frn.
Ge Meng, Zhongnan Cai, Ruizhe Chen, Jingyan Tu, Yingying Wang 0005, Yue Huang 0001, Xinghao Ding
NeurIPS5
2025 SOMA: A semantic-guided Order-aware Mamba Architecture for multivariate time series forecasting
Jinkai Zhang, Yingying Wang 0005, Shengbin Ma, Xinghao Ding, Xiaotong Tu
Adv. Eng. Informatics2
2025 Spatial-frequency dual-domain Kolmogorov-Arnold networks for multimodal medical image fusion
Lewu Lin, Jiaxin Xie, Yingying Wang 0005, Jialing Huang, Rongjin Zhuang, Xiaotong Tu, Xinghao Ding, Na Shen
Neurocomputing3
2025 Vision-Language Model Priors-Driven State Space Model for Infrared-Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) aims to effectively integrate complementary information from both infrared and visible modalities, enabling a more comprehensive understanding of the scene and improving downstream semantic tasks. Recent advancements in Mamba have shown remarkable performance in image fusion, owing to its linear complexity and global receptive fields. However, leveraging Vision-Language Model (VLM) priors to drive Mamba for modality-specific feature extraction and using them as constraints to enhance fusion results has not been fully explored. To address this gap, we introduce VLMPD-Mamba, a Vision-Language Model Priors-Driven Mamba framework for IVIF. Initially, we employ the VLM to adaptively generate modality-specific textual descriptions, which enhance image quality and highlight critical target information. Next, we present Text-Controlled Mamba (TCM), which integrates textual priors from the VLM to facilitate effective modality-specific feature extraction. Furthermore, we design the Cross-modality Fusion Mamba (CFM) to fuse features from different modalities, utilizing VLM priors as constraints to enhance fusion outcomes while preserving salient targets with rich details. In addition, to promote effective cross modality feature interactions, we introduce a novel bi-modal interaction scanning strategy within the CFM. Extensive experiments on various datasets for IVIF, as well as downstream visual tasks, demonstrate the superiority of our approach over state-of-the-art (SOTA) image fusion algorithms.
Rongjin Zhuang, Yingying Wang 0005, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
IEEE Signal Process. Lett.2
2025 Fusion2Void: Unsupervised Multi-Focus Image Fusion Based on Image Inpainting
abstract
Multi-focus image fusion aims to integrate clear segments from different partially focused images, creating an ‘all-in-focus’ composite. Due to the lack of ground-truth for multi-focus image fusion, supervised deep learning methods are deemed inappropriate for this task. In this paper, we present an unsupervised approach for multi-focus image fusion, named Fusion2Void. Fusion2Void ingeniously tackles the challenge of missing ground-truth by framing image inpainting as an auxiliary task. Specifically, Fusion2Void utilizes a fusion network to merge focused regions from multiple source images. Following the fusion process, image patches in the source images are randomly dropped to construct an additional image inpainting task. Subsequently, an image inpainting network uses the fused image as a guide to restore the missing content in the source images. The missing content in the source images includes both focused and defocused regions. Restoring focused image patches is significantly more challenging than restoring their defocused counterparts due to their inclusion of more high-frequency details. If the focused image patches are effectively restored, the repair of the defocused image patches becomes notably easier. Therefore, the image inpainting network implicitly compels the fused image to incorporate all focused content from the source images, as these can be utilized to restore the missing focused regions in the source images perfectly. Based on image inpainting, the fusion network generates ‘all-in-focus’ images in an unsupervised manner. Experiments on several synthetic and real-world datasets highlight Fusion2Void’s state-of-the-art performance relative to other methods.
Huangxing Lin, Yunlong Lin, Jingyuan Xia, Linyu Fan, Yingying Wang 0005, Xinghao Ding
IEEE Trans. Circuits Syst. Video Technol.6
2025 Frequency Decoupled Domain-Irrelevant Feature Learning for Pan-Sharpening
abstract
Pan-sharpening aims to generate high-detail multi-spectral images (HRMS) through the fusion of panchromatic (PAN) and multi-spectral (MS) images. However, existing pan-sharpening methods often suffer from significant performance degradation when dealing with out-of-distribution data, as they assume the training and test datasets are independent and identically distributed. To overcome this challenge, we propose a novel frequency domain-irrelevant feature learning framework that exhibits exceptional generalization capabilities. Our approach involves parallel extraction and processing of domain-irrelevant information from the amplitude and phase components of the input images. Specifically, we design a frequency information separation module to extract the amplitude and phase components of the paired images. The learnable high-pass filter is then employed to eliminate domain-specific information from the amplitude spectrums. After that, we devised two specialized sub-networks (AFL-Net and PFL-Net) to perform targeted learning of the frequency domain-irrelevant information. This allows our method to effectively capture the complementary domain-irrelevant information contained in the amplitude and phase spectra of the images. Finally, the information fusion and restoration module dynamically adjusts the feature channel weights, enabling the network to output high-quality HRMS images. Through this frequency domain-irrelevant feature learning framework, our method balances generalization capability and network performance on the distribution of training dataset. Extensive experiments conducted on various satellite datasets demonstrate the effectiveness of our method for generalized pan-sharpening. Our proposed network outperforms state-of-the-art methods in terms of both quantitative metrics and visual quality, showcasing its superior ability to handle diverse, out-of-distribution data.
Jie Zhang 0033, Ke Cao 0001, Yunlong Lin, Xuanhua He, Yingying Wang 0005, Rui Li 0027, Chengjun Xie, Jun Zhang 0034, Man Zhou 0003
IEEE Trans. Circuits Syst. Video Technol.6
2025 Learning Diffusion High-Quality Priors for Pan-Sharpening: A Two-Stage Approach With Time-Aware Adapter Fine-Tuning
abstract
Pan-sharpening aims to enhance the spatial resolution of the low-resolution multispectral (LRMS) image by incorporating high-frequency details from the panchromatic (PAN) image, while maintaining the spectral qualities of the LRMS image. Recent advancements in diffusion models have shown remarkable capabilities in image restoration and generation. However, simply applying diffusion models in pan-sharpening yields suboptimal outcomes in terms of fine-grained details and spectral fidelity. To this end, we introduce TA-DiffHQP, a two-stage approach that integrates the diffusion high-quality priors model (DiffHQP) and the time-aware adapter (TA-Adapter). Initially, we perform self-reconstruction pretraining DiffHQP with a fixed sampling strategy on approximately 24K high-resolution remote sensing datasets to explicitly model the high-quality texture details and spectral fidelity, after which we freeze most of DiffHQP’s parameters. In stage two, we integrate time-aware fusion adapters with the DiffHQP, enabling rapid adaptation to the pan-sharpening task. The TA-Adapters prioritize low-frequency main scenes during the early phases of the denoising process and refine high-frequency details in the later phases, achieving cross-modal information fusion from coarse to fine. Extensive experiments conducted on three satellite datasets demonstrate that our approach attains state-of-the-art (SOTA) performance over existing methods, revealing superior fusion outcomes in pan-sharpening.
Yingying Wang 0005, Yunlong Lin, Xuanhua He, Hui Zheng 0003, Linyu Fan, Yue Huang 0001, Xinghao Ding
IEEE Trans. Geosci. Remote. Sens.1
2025 Toward Generalizable Pansharpening: Conditional Flow-Based Learning Guided by Implicit High-Frequency Priors
abstract
The goal of pansharpening is to restore the missing high-frequency details in the low-resolution multispectral (LRMS) image to generate its high-resolution multispectral (HRMS) counterpart by exploiting the high-resolution panchromatic (PAN) image as guidance. Previous research has predominantly focused on improving pansharpening performance for single satellites, often neglecting the challenge of generalization. Moreover, pansharpening is inherently an ill-posed problem. Precise and generalizable prior guidance is crucial for effectively addressing this issue. To this end, we propose conditional flow-based learning guided by implicit high-frequency priors (CFLIHPs) toward generalizable pansharpening. Specifically, we utilize implicit neural representation (INR) to precisely align implicit high-frequency texture priors from LRMS and PAN images within Fourier and gradient domains. The flow-based restoration module then leverages these priors as the guiding condition to restore domain-irrelevant high-frequency details, thereby facilitating effective cross-satellite generalization. Furthermore, to tackle the complex degradation process in real-world scenarios, we introduce noise perturbation to the high-frequency learning part, enhancing generalizability across diverse spatial resolutions and improving the robustness of our framework. Extensive experiments conducted on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art (SOTA) methods, achieving superior performance and excellent generalization results in both cross-satellite scenarios and full-resolution scenes.
Yingying Wang 0005, Hui Zheng 0003, Yunlong Lin, Linyu Fan, Xuanhua He, Yue Huang 0001, Xinghao Ding
IEEE Trans. Geosci. Remote. Sens.1
2024 Progressive High-Frequency Reconstruction for Pan-Sharpening with Implicit Neural Representation
abstract
Pan-sharpening aims to leverage the high-frequency signal of the panchromatic (PAN) image to enhance the resolution of its corresponding multi-spectral (MS) image. However, deep neural networks (DNNs) tend to prioritize learning the low-frequency components during the training process, which limits the restoration of high-frequency edge details in MS images. To overcome this limitation, we treat pan-sharpening as a coarse-to-fine high-frequency restoration problem and propose a novel method for achieving high-quality restoration of edge information in MS images. Specifically, to effectively obtain fine-grained multi-scale contextual features, we design a Band-limited Multi-scale High-frequency Generator (BMHG) that generates high-frequency signals from the PAN image within different bandwidths. During training, higher-frequency signals are progressively injected into the MS image, and corresponding residual blocks are introduced into the network simultaneously. This design enables gradients to flow from later to earlier blocks smoothly, encouraging intermediate blocks to concentrate on missing details. Furthermore, to address the issue of pixel position misalignment arising from multi-scale features fusion, we propose a Spatial-spectral Implicit Image Function (SIIF) that employs implicit neural representation to effectively represent and fuse spatial and spectral features in the continuous domain. Extensive experiments on different datasets demonstrate that our method outperforms existing approaches in terms of quantitative and visual measurements for high-frequency detail recovery.
Ge Meng, Jingjia Huang, Yingying Wang 0005, Zhenqi Fu, Xinghao Ding, Yue Huang 0001
AAAI3
2024 Efficient Perceiving Local Details via Adaptive Spatial-Frequency Information Integration for Multi-focus Image Fusion
abstract
Multi-focus image fusion (MFIF) aims to combine multiple images with different focused regions into a single all-in-focus image. Existing unsupervised deep learning-based methods only fuse structural information of images in the spatial domain, neglecting potential solutions from the frequency domain exploration. In this paper, we make the first attempt to integrate spatial-frequency information to achieve high-quality MFIF. We propose a novel unsupervised spatial-frequency interaction MFIF network named SFIMFN, which consists of three key components: Adaptive Frequency Domain Information Interaction Module (AFIM), Ret-Attention-Based Spatial Information Extraction Module (RASEM), and Invertible Dual-domain Feature Fusion Module (IDFM). Specifically, in AFIM, we interactively explore global contextual information by combining the amplitude and phase information of multiple images separately. In RASEM, we design a customized transformer to encourage the network to capture important local high-frequency information by redesigning the self-attention mechanism with a bidirectional, two-dimensional form of explicit decay. Finally, we employ IDFM to fuse spatial-frequency information without information loss to generate the desired all-in-focus image. Extensive experiments on different datasets demonstrate that our method significantly outperforms state-of-the-art unsupervised methods in terms of qualitative and quantitative metrics as well as the generalization ability.
Jingjia Huang, Jingyan Tu, Ge Meng, Yingying Wang 0005, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ACM Multimedia4
2024 Diffusion-Based Continuous Feature Representation for Infrared Small-Dim Target Detection
abstract
Infrared small-dim target detection plays a pivotal role in missions involving rescue, surveillance, and early warning systems. Despite remarkable strides made by existing methods, certain limitations still hinder the detection accuracy, including deficiency in high-resolution representation, inadequacy in addressing dim targets, and difficulty in tackling low-contrast targets against complex backgrounds. To overcome these limitations, we propose a diffusion-based continuous feature representation network (DCFR-Net), comprising two crucial branches: diffusion-based continuous high-resolution feature representation (DCHFR) and infrared small-dim target detection (ISDTD). Specifically, to precisely capture extremely small target contours, DCHFR integrates implicit neural representation (INR) into a conditional denoising diffusion model, super-resolving infrared targets in a self-supervised strategy. ISDTD leverages the shared encoder from DCHFR to construct high-resolution feature representation, which is fed into multi-scale implicit feature alignment (MIFA) and spatial-frequency feature interaction (SFFI). To alleviate the impact of dim and vulnerable targets, MIFA delicately aggregates different-layer features in a resolution-free manner. Furthermore, to enhance the contrast between infrared targets and intricate backgrounds, SFFI achieves profound spatial-frequency feature interaction and global-local receptive field mixture. Extensive experiments conducted on three challenging datasets of NUAA-SIRST, IRSTD-1k and NUDT-SIRST reveal that our DCFR-Net outperforms the state-of-the-art (SOTA) methods, demonstrating the superiority and robustness of our approach in infrared small-dim target detection. Code will be available at https://github.com/flyannie/DCFR-Net.
Linyu Fan, Yingying Wang 0005, Guoliang Hu, Hui Zheng 0003, Yue Huang 0001, Xinghao Ding
IEEE Trans. Geosci. Remote. Sens.2
2024 Cross-Modality Interaction Network for Pan-Sharpening
abstract
Pan-sharpening seeks to generate a high-resolution multispectral (HRMS) image by merging the high-resolution panchromatic (PAN) image and its low-resolution multispectral (LRMS) counterpart. The main challenge lies in enhancing modality-aware features and efficiently integrating complementary information between PAN and MS pairs. To achieve desired fusion results, it is crucial to fully utilize both intramodality characteristics and intermodality relationships. Current research often overlooks the exploration of cross-modality relationships and neglects the enhancement of modality-aware features in pan-sharpening. In this work, we introduce an innovative pan-sharpening framework, named cross-modality interaction network (CMINet), which comprises three core designs: a modality-aware feature enhancement (MAFE) module to enhance the feature representation of both modalities, a cross-modality attention (CMA) module that effectively extracts the intramodality features and fully leverages the intermodality complementary information, and a modality alignment (MA) module to address modality-aware misalignment issue during fusion. Extensive experiments are conducted to verify the effectiveness of our proposed network and showcase its superior performance in comparison to other state-of-the-art approaches.
Yingying Wang 0005, Xuanhua He, Yunlong Lin, Yue Huang 0001, Xinghao Ding
IEEE Trans. Geosci. Remote. Sens.1
2023 Domain-irrelevant Feature Learning for Generalizable Pan-sharpening
abstract
Pan-sharpening aims to spatially enhance the low-resolution multispectral image (LRMS) by transferring high-frequency details from a panchromatic image (PAN) while preserving the spectral characteristics of LRMS. Previous arts mainly focus on how to learn a high-resolution multispectral image (HRMS) on the i.i.d. assumption. However, the distribution of training and testing data often encounters significant shifts in different satellites. To this end, this paper proposes a generalizable pan-sharpening network via domain-irrelevant feature learning. On the one hand, a structural preservation module (STP) is designed to fuse high-frequency information of PAN and LRMS. Our STP is performed on the gradient domain because it consists of structure and texture details that can generalize well on different satellites. On the other hand, to avoid spectral distortion while promoting the generalization ability, a spectral preservation module (SPP) is developed. The key design of SPP is to learn a phase fusion network of PAN and LRMS. The amplitude of LRMS, which contains 'satellite style' information is directly injected in different fusion stages. Extensive experiments have demonstrated the effectiveness of our method against state-of-the-art methods in both single-satellite and cross-satellite scenarios. Code is available at: https://github.com/LYL1015/DIRFL.
Yunlong Lin, Zhenqi Fu, Ge Meng, Yingying Wang 0005, Linyu Fan, Hedeng Yu, Xinghao Ding
ACM Multimedia4
2023 Learning High-frequency Feature Enhancement and Alignment for Pan-sharpening
abstract
Pan-sharpening aims to utilize the high-resolution panchromatic (PAN) image as a guidance to super-resolve the spatial resolution of the low-resolution multispectral (MS) image. The key challenge in pan-sharpening is how to effectively and precisely inject high-frequency edges and textures from the PAN image into the low-resolution MS image. To address this issue, we propose a High-frequency Feature Enhancement and Alignment Network (HFEAN) for effectively encouraging the high-frequency learning. To implement it, three core designs are customized: a Fourier convolution based efficient feature enhancement module (FEM), an implicit neural alignment module (INA), and a preliminary alignment module (Pre-align). To be specific, FEM employs the fast Fourier convolution with attention mechanism to achieve the mixed global-local receptive field on each scale of the high-frequency domain, thus yielding the informative latent codes. INA leverages implicit neural function to precisely align the latent codes from different scales in the continuous domain. In this way, the high frequency signals at different scales are represented as functions of continuous coordinates, enabling a precise feature alignment in a resolution-free manner. Pre-align is developed to further address the inherent misalignment between PAN and MS pairs. Extensive experiments over multiple satellite datasets validate the effectiveness of the proposed network and demonstrate its favorable performance against the existing state-of-the-art methods both visually and quantitatively. Code is available at: https://github.com/Gracewangyy/HFEAN.
Yingying Wang 0005, Yunlong Lin, Ge Meng, Zhenqi Fu, Linyu Fan, Hedeng Yu, Xinghao Ding, Yue Huang 0001
ACM Multimedia1
2023 High-Resolution Feature Representation Driven Infrared Small-Dim Object Detection
Yingying Wang 0005, Linyu Fan, Xinghao Ding, Yue Huang 0001
PRCV (12)2
2023 DP-INNet: Dual-Path Implicit Neural Network for Spatial and Spectral Features Fusion in Pan-Sharpening
Jingjia Huang, Ge Meng, Yingying Wang 0005, Yunlong Lin, Yue Huang 0001, Xinghao Ding
PRCV (8)3