VLDB 2026 Research / reviewers in the wild / expert
Xuanhua He
dblp:348/6039
· DBLP profile ↗
25ranked-venue papers
5as first author
25since 2021 · last 2026
0009-0004-6637-853XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RelaCtrl: Relevance-Guided Efficient Control for Diffusion TransformersabstractThe Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource allocation due to their failure to account for the varying relevance of control information across different transformer layers. To address this, we propose the Relevance-Guided Efficient Controllable Generation framework, RelaCtrl, enabling efficient and resource-optimized integration of control signals into the Diffusion Transformer. First, we evaluate the relevance of each layer in the Diffusion Transformer to the control information by assessing the ControlNet Relevance Score, which measures the impact of skipping each control layer on both the quality of generation and the control effectiveness during inference. Based on the strength of the relevance, we then tailor the positioning, parameter scale, and modeling capacity of the control layers to reduce unnecessary parameters and redundant computations. Additionally, to further improve efficiency, we replace the self-attention and FFN in the commonly used copy block with the carefully designed Two-Dimensional Shuffle Mixer (TDSM), enabling efficient implementation of both the token mixer and channel mixer. Both qualitative and quantitative experimental results demonstrate that our approach achieves superior performance with only 15% of the parameters and computational complexity compared to PixArt-delta. Ke Cao 0001, Jing Wang 0021, Ao Ma 0005, Jiasong Feng, Xuanhua He, Run Ling, Haozhe Wang 0002, Hongjuan Pei, Yihua Shao, Zhanjie Zhang, Jie Zhang 0033 |
AAAI | 5 |
| 2026 | ContextFlow: Training-Free Video Object Editing via Adaptive Context EnrichmentabstractTraining-free video object editing aims to achieve precise object-level manipulation, including object insertion, swapping, and deletion. However, it faces significant challenges in maintaining fidelity and temporal consistency. Existing methods, often designed for U-Net architectures, suffer from two primary limitations: inaccurate inversion due to first-order solvers, and contextual conflicts caused by crude "hard" feature replacement. These issues are more challenging in Diffusion Transformers (DiTs), where the unsuitability of prior layer-selection heuristics makes effective guidance challenging. To address these limitations, we introduce ContextFlow, a novel training-free framework for DiT-based video object editing. In detail, we first employ a high-order Rectified Flow solver to establish a robust editing foundation. The core of our framework is Adaptive Context Enrichment (for specifying what to edit), a mechanism that addresses contextual conflicts. Instead of replacing features, it enriches the self-attention context by concatenating Key-Value pairs from parallel reconstruction and editing paths, empowering the model to dynamically fuse information. Additionally, to determine where to apply this enrichment (for specifying where to edit), we propose a systematic, data-driven analysis to identify task-specific vital layers. Based on a novel Guidance Responsiveness Metric, our method pinpoints the most influential DiT blocks for different tasks (e.g., insertion, swapping), enabling targeted and highly effective guidance. Extensive experiments show that ContextFlow significantly outperforms existing training-free methods and even surpasses several state-of-the-art training-based approaches, delivering temporally coherent, high-fidelity results. Xuanhua He, Xiujun Ma, Jack Ma |
AAAI | 2 |
| 2026 | MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image EnhancementabstractPan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks. Yingying Wang 0005, Xuanhua He, Jialing Huang, Suiyun Zhang, Xinghao Ding, Haoxuan Che |
AAAI | 2 |
| 2026 | Self-supervised Multiplex Consensus Mamba for General Image FusionabstractImage fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks. Yingying Wang 0005, Rongjin Zhuang, Hui Zheng 0003, Xuanhua He, Ke Cao 0001, Xiaotong Tu, Xinghao Ding |
AAAI | 4 |
| 2026 | Shuffle Mamba: State Space Models With Random Shuffle for Multi-Modal Image FusionabstractMulti-modal image fusion integrates complementary information from different modalities to produce enhanced and informative images. Although State-Space Models, such as Mamba, are proficient in long-range modeling with linear complexity, most Mamba-based approaches use fixed scanning strategies, which can introduce biased prior information. To mitigate this issue, we propose a novel Bayesian-inspired scanning strategy called Random Shuffle, supplemented by a theoretically feasible inverse shuffle to maintain information coordination invariance, aiming to eliminate biases associated with fixed sequence scanning. Based on this transformation pair, we customized the Shuffle Mamba Framework, penetrating modality-aware information representation and cross-modality information interaction across spatial and channel axes to ensure robust interaction and an unbiased global receptive field for multi-modal image fusion. Furthermore, we develop a testing methodology based on Monte-Carlo averaging to ensure the model’s output aligns more closely with expected results. Extensive experiments across multiple multi-modal image fusion tasks demonstrate the effectiveness of our proposed method, yielding excellent fusion quality compared to state-of-the-art alternatives. The code is available at https://github.com/caoke-963/Shuffle-Mamba. Ke Cao 0001, Xuanhua He, Tao Hu 0027, Chengjun Xie, Man Zhou 0003, Jie Zhang 0033 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Distilling Textual Priors From LLM to Efficient Image FusionabstractMulti-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances in text-guided methods leverage large model priors to overcome these limitations, but at the cost of significant computational overhead, both in memory and inference time. To address this challenge, we propose a novel framework for distilling large model priors, eliminating the need for text guidance during inference while dramatically reducing model size. Our framework utilizes a teacher-student architecture, where the teacher network incorporates large model priors and transfers this knowledge to a smaller student network via a tailored distillation process. Crucially, our experiments demonstrate that this knowledge transfer is the primary driver of performance gains, rather than mere architectural optimization. Additionally, we introduce a spatial-channel cross-fusion module to enhance the model’s ability to leverage textual priors across both spatial and channel dimensions. Our method achieves a favorable trade-off between computational efficiency and fusion quality. The distilled network, requiring only 10% of the parameters and inference time of the teacher network, retains 90% of its performance and outperforms existing SOTA methods. Extensive experiments demonstrate the effectiveness of our approach. Codes are available at https://github.com/Zirconium233/DTPF. Xuanhua He, Ke Cao 0001, Liu Liu 0012, Li Zhang 0104, Man Zhou 0003, Jie Zhang 0033, Dan Guo 0001, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | GameGen-X: Interactive Open-world Game Video GenerationabstractWe introduce GameGen-$\mathbb{X}$, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos.
This model facilitates high-quality, open-domain generation by approximating various game elements, such as innovative characters, dynamic environments, complex actions, and diverse events.
Additionally, it provides interactive controllability, predicting and altering future content based on the current clip, thus allowing for gameplay simulation.
To realize this vision, we first collected and built an Open-World Video Game Dataset (OGameData) from scratch.
It is the first and largest dataset for open-world game video generation and control, which comprises over one million diverse gameplay video clips with informative captions.
GameGen-$\mathbb{X}$ undergoes a two-stage training process, consisting of pre-training and instruction tuning.
Firstly, the model was pre-trained via text-to-video generation and video continuation, enabling long-sequence open-domain game video generation with improved fidelity and coherence.
Further, to achieve interactive controllability, we designed InstructNet to incorporate game-related multi-modal control signal experts.
This allows the model to adjust latent representations based on user inputs, advancing the integration of character interaction and scene content control in video generation.
During instruction tuning, only the InstructNet is updated while the pre-trained foundation model is frozen, enabling the integration of interactive controllability without loss of diversity and quality of generated content.
GameGen-$\mathbb{X}$ contributes to advancements in open-world game design using generative models.
It demonstrates the potential of generative models to serve as auxiliary tools to traditional rendering techniques, demonstrating the potential for merging creative generation with interactive capabilities.
The project will be available at https://github.com/GameGen-X/GameGen-X. Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin 0003, Hao Chen 0011 |
ICLR | 2 |
| 2025 | Freq-RWKV: Granularity-Aware Spatial-Frequency Synergy via Dual-Domain Recurrent Scanning for Pan-sharpening
Xueheng Li, Xuanhua He, Tao Hu 0027, Jie Zhang 0033, Man Zhou 0003, Chengjun Xie, Yingying Wang 0005, Bo Huang 0001 |
ACM Multimedia | 2 |
| 2025 | WKV-sharing embraced random shuffle RWKV high-order modeling for pan-sharpeningabstractPan-sharpening aims to generate a spatially and spectrally enriched multi-spectral image by integrating complementary
cross-modality information from low-resolution multi-spectral image and texture-rich panchromatic counterpart. In this work, we propose a
WKV-sharing embraced random shuffle RWKV high-order modeling paradigm for pan-sharpening from Bayesian perspective, coupled with random weight manifold distribution training strategy derived from Functional theory to regularize the solution space adhering to the
following principles: 1) Random-shuffle RWKV. Recently, the Vision RWKV model, with its inherent linear complexity in global modeling,
has inspired us to explore its untapped potential in pan-sharpening tasks. However, its attention mechanism, relying on a recurrent
bidirectional scanning strategy, suffers from biased effects and demands significant processing time. To address this, we propose a novel
Bayesian-inspired scanning strategy called Random Shuffle, complemented by a theoretically-sound inverse shuffle to preserve
information coordination invariance, effectively eliminating biases associated with fixed sequence scanning. The Random Shuffle
approach mitigates preconceptions in global 2D dependencies in mathematical expectation, providing the model with an unbiased prior.
In line with similar spirit of Dropout, we introduce a testing methodology based on Monte Carlo averaging to ensure the model’s output
aligns more closely with expected results. 2) WKV-sharing high-order. Regarding KV’s attention score calculation in spatial mixer of RWKV, we leverage WKV-sharing mechanism to transfer KV activations across RWKV layers, achieving lower latency and improved trainability, and revisit the channel mixer in RWKV, originally a first-order weighting function, and redevelop its high-order potential by sharing the gate mechanism across RWKV layer. Comprehensive experiments across pan-sharpening benchmarks demonstrate our model’s effectiveness, consistently outperforming state-of-the-art alternatives Man Zhou 0003, Xuanhua He, Danfeng Hong, Bo Huang 0001 |
NeurIPS | 2 |
| 2025 | Toward Resolution Mismatching: Modality-Aware Feature-Aligned Network for Pan-SharpeningabstractPanchromatic (PAN) and multi-spectral (MS) remote satellite image fusion, known as pan-sharpening, aims to produce high-resolution MS images by combining the complementary information from the high-resolution, texture-rich PAN and the low-resolution but high spectral-resolution MS counterparts. Despite notable advancements in this field, the current state-of-the-art pan-sharpening techniques do not explicitly address the spatial resolution mismatching problem between the two modalities of PAN and MS images. This mismatching issue can lead to misalignment in feature representation and the creation of blurry artifacts in the model output, ultimately hindering the generation of high-frequency textures and impeding the performance improvement of such methods. To address the aforementioned spatial resolution mismatching problem in pan-sharpening, we propose a novel modality-aware feature-aligned pan-sharpening framework in this paper. The framework comprises three primary stages: modality-aware feature extraction, modality-aware feature aligning, and context integrated image reconstruction. First, we introduce the half-instance normalization strategy as the backbone to filter out the inconsistent features and promote the learning of consistent features between the PAN and MS modalities. Second, a learnable modality-aware feature interpolation is devised to effectively address the misalignment issue. Specifically, the extracted features from the backbone are integrated to predict the transformation offsets of each pixel, which allows for the adaptive selection of custom contextual information and enables the modality-aware features to be more aligned. Finally, within the context of the interactive offset correction, multi-stage information is aggregated to generate the feasible pan-sharpened model output. Extensive experimental results over multiple satellite datasets demonstrate that the proposed algorithm outperforms other state-of-the-art methods both qualitatively and quantitatively, exhibiting great generalization ability to real-world scenes. Man Zhou 0003, Xuanhua He, Danfeng Hong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Probing Synergistic High-Order Interaction for Multi-Modal Image FusionabstractMulti-modal image fusion aims to generate a fused image by integrating and distinguishing the cross-modality complementary information from multiple source images. While the cross-attention mechanism with global spatial interactions appears promising, it only captures second-order spatial interactions, neglecting higher-order interactions in both spatial and channel dimensions. This limitation hampers the exploitation of synergies between multi-modalities. To bridge this gap, we introduce a Synergistic High-order Interaction Paradigm (SHIP), designed to systematically investigate spatial fine-grained and global statistics collaborations between the multi-modal images across two fundamental dimensions: 1) Spatial dimension: we construct spatial fine-grained interactions through element-wise multiplication, mathematically equivalent to global interactions, and then foster high-order formats by iteratively aggregating and evolving complementary information, enhancing both efficiency and flexibility. 2) Channel dimension: expanding on channel interactions with first-order statistics (mean), we devise high-order channel interactions to facilitate the discernment of inter-dependencies between source images based on global statistics. We further introduce an enhanced version of the SHIP model, called SHIP++ that enhances the cross-modality information interaction representation by the cross-order attention evolving mechanism, cross-order information integration, and residual information memorizing mechanism. Harnessing high-order interactions significantly enhances our model's ability to exploit multi-modal synergies, leading in superior performance over state-of-the-art alternatives, as shown through comprehensive experiments across various benchmarks in two significant multi-modal image fusion tasks: pan-sharpening, and infrared and visible image fusion. Man Zhou 0003, Naishan Zheng, Xuanhua He, Danfeng Hong, Jocelyn Chanussot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | High-Order State Space Model for Multi-Modal Accelerated MRI ReconstructionabstractIn this work, we propose a novel high-order state space model embedded within an interpretable unfolding algorithm for MRI reconstruction, marking the first adaptation of the Mamba framework to this domain. Traditional methods often rely on stacking second-order interactions through transformer architectures to capture high-order global information. However, these approaches exhibit inherent quadratic complexity, resulting in significant computational overhead and limiting their efficiency in practical environments. Through a reexamination of the intrinsic properties of Mamba, we observe that it is fundamentally a first-order function and leveraging this insight, we introduce a high-order variant of Mamba, wherein the linear complexity attention mechanism within the initial Mamba is shared across successive Mamba calculations, achieving the same functionality while mitigating the computational cost of stacked methods. This adaptation allows our model to retain the benefits of cascaded modeling while significantly improving computational efficiency. By capturing high-order interactions more efficiently, our approach strikes a balance between performance and resource usage, providing a practical solution for accelerated MRI reconstruction. Furthermore, extensive experiments on multiple MRI benchmark datasets demonstrate that our approach delivers competitive performance compared to other state-of-the-arts. Bei Tang, Xuanhua He, Yuhan Zhan |
IEEE Signal Process. Lett. | 2 |
| 2025 | Frequency Decoupled Domain-Irrelevant Feature Learning for Pan-SharpeningabstractPan-sharpening aims to generate high-detail multi-spectral images (HRMS) through the fusion of panchromatic (PAN) and multi-spectral (MS) images. However, existing pan-sharpening methods often suffer from significant performance degradation when dealing with out-of-distribution data, as they assume the training and test datasets are independent and identically distributed. To overcome this challenge, we propose a novel frequency domain-irrelevant feature learning framework that exhibits exceptional generalization capabilities. Our approach involves parallel extraction and processing of domain-irrelevant information from the amplitude and phase components of the input images. Specifically, we design a frequency information separation module to extract the amplitude and phase components of the paired images. The learnable high-pass filter is then employed to eliminate domain-specific information from the amplitude spectrums. After that, we devised two specialized sub-networks (AFL-Net and PFL-Net) to perform targeted learning of the frequency domain-irrelevant information. This allows our method to effectively capture the complementary domain-irrelevant information contained in the amplitude and phase spectra of the images. Finally, the information fusion and restoration module dynamically adjusts the feature channel weights, enabling the network to output high-quality HRMS images. Through this frequency domain-irrelevant feature learning framework, our method balances generalization capability and network performance on the distribution of training dataset. Extensive experiments conducted on various satellite datasets demonstrate the effectiveness of our method for generalized pan-sharpening. Our proposed network outperforms state-of-the-art methods in terms of both quantitative metrics and visual quality, showcasing its superior ability to handle diverse, out-of-distribution data. Jie Zhang 0033, Ke Cao 0001, Yunlong Lin, Xuanhua He, Yingying Wang 0005, Rui Li 0027, Chengjun Xie, Jun Zhang 0034, Man Zhou 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | PanDiT: A Few-Step Diffusion Transformer for High-Fidelity and Efficient PansharpeningabstractPansharpening plays a crucial role in remote sensing by fusing low-resolution multispectral (LRMS) images and high-resolution panchromatic (PAN) images to generate high-resolution multispectral (HRMS) images. While denoising diffusion models offer potential for high-fidelity image generation, their practical application in pansharpening has been severely hindered by huge computational costs from iterative sampling and naive conditioning strategies that struggle to fuse multi-modal information effectively. In this paper, we introduce PanDiT, a novel Diffusion Transformer framework designed to address these challenges. PanDiT is built on the core principle of a Decoupled Conditioning Mechanism, which explicitly disentangles and injects spatial and spectral guidance, and is engineered for practical, few-step inference. Our framework leverages a powerful Diffusion Transformer (DiT) backbone, where conditioning is achieved through two specialized injection blocks that capture spatial and time-frequency features. Crucially, by integrating an implicit sampling strategy, we accelerate the inference process to as few as two steps. Extensive experiments on multiple benchmark datasets demonstrate that PanDiT not only establishes a new state-of-the-art in fusion quality and achieves a good quality-efficiency trade-off. Code is available at https://github.com/para133/PanDIT. Jiabin Fang, Ke Cao 0001, Xuanhua He, Jie Zhang 0033, Man Zhou 0003, Liu Liu 0012 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Exploring Text-Guided Information Fusion Through Chain-of-Reasoning for PansharpeningabstractPan-sharpening aims to enhance the spatial resolution of low-resolution multispectral (LRMS) images by integrating high-frequency information from a corresponding texture-rich panchromatic (PAN) image, while maintaining the spectral integrity of the LRMS image. Although text-guided multi-modal learning has made considerable strides in the natural image domain, its potential to pan-sharpening remains underexplored, primarily due to the limited availability of multi-modal remote sensing datasets. To this end, we construct an entirely new pan-sharpening framework by making efforts from three key aspects: (1) text-equipped multi-modal data collection through chain-of-reasoning, (2) large model prior-driven multi-modal information fusion, and (3) visual information interaction through prompt engineering, leveraging textual information to guide the pan-sharpening process within a multi-modal fusion framework. We initially utilize the generic large language model priors to generate descriptive captions for MS images, forming a multi-modal pan-sharpening dataset. By integrating super-resolved imagery and segmentation maps generated by segment anything, we apply Chain-of-Thought (CoT) prompting to generate spatially focused captions across diverse satellite datasets. These captions enhance visual features and provide high-level contextual information, improving semantic understanding for pan-sharpening. Building on the aforementioned multi-modal data, we tailor two text-guided information fusion modules: Textual Enhancement Block (TEB) standing on large model prior and Textual Modulated Block (TMB) utilizing text information to effectively guide and refine the pan-sharpening fusion process. Extensive experiments on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art methods, highlighting its effectiveness and superior performance in pan-sharpening. Xueheng Li, Xuanhua He, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong, Bo Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Learning Diffusion High-Quality Priors for Pan-Sharpening: A Two-Stage Approach With Time-Aware Adapter Fine-TuningabstractPan-sharpening aims to enhance the spatial resolution of the low-resolution multispectral (LRMS) image by incorporating high-frequency details from the panchromatic (PAN) image, while maintaining the spectral qualities of the LRMS image. Recent advancements in diffusion models have shown remarkable capabilities in image restoration and generation. However, simply applying diffusion models in pan-sharpening yields suboptimal outcomes in terms of fine-grained details and spectral fidelity. To this end, we introduce TA-DiffHQP, a two-stage approach that integrates the diffusion high-quality priors model (DiffHQP) and the time-aware adapter (TA-Adapter). Initially, we perform self-reconstruction pretraining DiffHQP with a fixed sampling strategy on approximately 24K high-resolution remote sensing datasets to explicitly model the high-quality texture details and spectral fidelity, after which we freeze most of DiffHQP’s parameters. In stage two, we integrate time-aware fusion adapters with the DiffHQP, enabling rapid adaptation to the pan-sharpening task. The TA-Adapters prioritize low-frequency main scenes during the early phases of the denoising process and refine high-frequency details in the later phases, achieving cross-modal information fusion from coarse to fine. Extensive experiments conducted on three satellite datasets demonstrate that our approach attains state-of-the-art (SOTA) performance over existing methods, revealing superior fusion outcomes in pan-sharpening. Yingying Wang 0005, Yunlong Lin, Xuanhua He, Hui Zheng 0003, Linyu Fan, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Toward Generalizable Pansharpening: Conditional Flow-Based Learning Guided by Implicit High-Frequency PriorsabstractThe goal of pansharpening is to restore the missing high-frequency details in the low-resolution multispectral (LRMS) image to generate its high-resolution multispectral (HRMS) counterpart by exploiting the high-resolution panchromatic (PAN) image as guidance. Previous research has predominantly focused on improving pansharpening performance for single satellites, often neglecting the challenge of generalization. Moreover, pansharpening is inherently an ill-posed problem. Precise and generalizable prior guidance is crucial for effectively addressing this issue. To this end, we propose conditional flow-based learning guided by implicit high-frequency priors (CFLIHPs) toward generalizable pansharpening. Specifically, we utilize implicit neural representation (INR) to precisely align implicit high-frequency texture priors from LRMS and PAN images within Fourier and gradient domains. The flow-based restoration module then leverages these priors as the guiding condition to restore domain-irrelevant high-frequency details, thereby facilitating effective cross-satellite generalization. Furthermore, to tackle the complex degradation process in real-world scenarios, we introduce noise perturbation to the high-frequency learning part, enhancing generalizability across diverse spatial resolutions and improving the robustness of our framework. Extensive experiments conducted on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art (SOTA) methods, achieving superior performance and excellent generalization results in both cross-satellite scenarios and full-resolution scenes. Yingying Wang 0005, Hui Zheng 0003, Yunlong Lin, Linyu Fan, Xuanhua He, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier DomainabstractRAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB images, a difference that goes beyond the color matrix and extends to spatial structure due to resolution variations. Recent methods directly rebuild color mapping and spatial structure via shared deep representation, limiting optimal performance. Inspired by Image Signal Processing (ISP) pipeline, which distinguishes image restoration and enhancement, we present a novel Neural ISP framework, named FourierISP. This approach breaks the image down into style and structure within the frequency domain, allowing for independent optimization. FourierISP is comprised of three subnetworks: Phase Enhance Subnet for structural refinement, Amplitude Refine Subnet for color learning, and Color Adaptation Subnet for blending them in a smooth manner. This approach sharpens both color and structure, and extensive evaluations across varied datasets confirm that our approach realizes state-of-the-art results. Code will be available at https://github.com/alexhe101/FourierISP. Xuanhua He, Tao Hu 0027, Guoli Wang 0004, Zejin Wang, Qian Zhang 0009, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
AAAI | 1 |
| 2024 | Frequency-Adaptive Pan-Sharpening with Mixture of ExpertsabstractPan-sharpening involves reconstructing missing high-frequency information in multi-spectral images with low spatial resolution, using a higher-resolution panchromatic image as guidance. Although the inborn connection with frequency domain, existing pan-sharpening research has not almost investigated the potential solution upon frequency domain. To this end, we propose a novel Frequency Adaptive Mixture of Experts (FAME) learning framework for pan-sharpening, which consists of three key components: the Adaptive Frequency Separation Prediction Module, the Sub-Frequency Learning Expert Module, and the Expert Mixture Module. In detail, the first leverages the discrete cosine transform to perform frequency separation by predicting the frequency mask. On the basis of generated mask, the second with low-frequency MOE and high-frequency MOE takes account for enabling the effective low-frequency and high-frequency information reconstruction. Followed by, the final fusion module dynamically weights high frequency and low-frequency MOE knowledge to adapt to remote sensing images with significant content variations. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes. Code will be made publicly at https://github.com/alexhe101/FAME-Net. Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
AAAI | 1 |
| 2024 | Frequency Decomposition-Driven Network for JPEG Artifacts RemovalabstractJPEG compression, a widely adopted image format, often introduces visual artifacts and quality degradation in image quality. Removal of these JPEG artifacts, especially under high compression rates, proves challenging and typically results in overly smoothed images. This issue primarily arises due to the prevalence of low-frequency regions in natural images. This distribution leads models towards capturing low-frequency information and generating over-smoothing results. To address this issue, we propose the Frequency Decomposition-Driven Network (FDDNet) for JPEG artifact removal. FDDNet incorporates three core modules: the Decomposition Module (DM), inspired by wavelet lifting schemes, extracts both low-frequency and high-frequency components by considering feature channel relationships. The lightweight Low-Frequency Restoration Module and the High-Frequency Refinement Module, are each adept at handling distinct frequency components effectively. By emphasizing high-frequency components, our method surpasses existing approaches in terms of both quantitative metrics and visual quality across various datasets. Ke Cao 0001, Xuanhua He, Tao Hu 0027, Rui Li 0027, Chengjun Xie, Jie Zhang 0033 |
ICME | 2 |
| 2024 | Spatially-Adaptive Large-Kernel Network for Efficient Image Super-ResolutionabstractIn the realm of image super-resolution (SR), efficiency on low-power devices remains a significant challenge due to the high computational demands of current methods. In response to this issue, we propose a novel Spatially-Adaptive Large-Kernel Network(SALKN) tailored for efficient image super-resolution. Inspired by the spectrum convolution theorem, we introduce a spatially-adaptive large kernel convolution unit integrated into a vision transformer architecture. This approach implements dynamic large kernel convolution through an input-adaptive frequency domain multiplication and multi-head mechanism, significantly reducing computational overhead. Our implementation approach, which involves dynamically generating spatially-adaptive frequency filters, enables the dynamic large kernel convolution to be performed with reduced computational costs, facilitating the realization of a global receptive field and promoting scale diversity within the features. Extensive experiments validate that SALKN outperforms existing efficient super-resolution methods with reduced complexity, achieving state-of-the-art performance. Xuanhua He, Ke Cao 0001, Tao Hu 0027, Jie Zhang 0033, Rui Li 0027 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Cross-Modality Interaction Network for Pan-SharpeningabstractPan-sharpening seeks to generate a high-resolution multispectral (HRMS) image by merging the high-resolution panchromatic (PAN) image and its low-resolution multispectral (LRMS) counterpart. The main challenge lies in enhancing modality-aware features and efficiently integrating complementary information between PAN and MS pairs. To achieve desired fusion results, it is crucial to fully utilize both intramodality characteristics and intermodality relationships. Current research often overlooks the exploration of cross-modality relationships and neglects the enhancement of modality-aware features in pan-sharpening. In this work, we introduce an innovative pan-sharpening framework, named cross-modality interaction network (CMINet), which comprises three core designs: a modality-aware feature enhancement (MAFE) module to enhance the feature representation of both modalities, a cross-modality attention (CMA) module that effectively extracts the intramodality features and fully leverages the intermodality complementary information, and a modality alignment (MA) module to address modality-aware misalignment issue during fusion. Extensive experiments are conducted to verify the effectiveness of our proposed network and showcase its superior performance in comparison to other state-of-the-art approaches. Yingying Wang 0005, Xuanhua He, Yunlong Lin, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Pan-Sharpening With Wavelet-Enhanced High-Frequency InformationabstractPan-sharpening is essentially a panchromatic (PAN)-guided super-resolution process, primarily focused on enhancing multi-spectral image quality. This methodology intricately incorporates the high-frequency derived from texture-rich PAN images into the lower-resolution multi-spectral (LRMS) counterparts. However, current spatial domain techniques frequently face challenges in accurately restoring texture details, while frequency domain methods lack efficient interaction with spatial domains, thus restricting the overall model performance. In response to these challenges, we introduce a novel High-frequency Wavelet Network that capitalizes on the spatial-frequency interaction and frequency division capabilities inherent in wavelet transform. In particular, our approach consists of two fundamental modules: the Wavelet-Inspired Fusion Block and the High-Frequency Enhancement Block. The former is inspired by wavelet lifting schemes, enabling the fusion of frequencies and facilitating information exchange across various subbands. The latter harnesses wavelet’s frequency division attributes to enhance high-frequency information learning. Comprehensive experiments over multiple satellite datasets demonstrate that our approach outperforms state-of-the-art techniques in both quantitative and qualitative assessments. Moreover, our model showcases exceptional generalization capabilities in real-world scenarios. Code is available at https://github.com/alexhe101/WINet. Jie Zhang 0033, Xuanhua He, Ke Cao 0001, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Pyramid Dual Domain Injection Network for Pan-sharpeningabstractPan-sharpening, a panchromatic image guided low-spatial-resolution multi-spectral super-resolution task, aims to reconstruct the missing high-frequency information of high-resolution multi-spectral counterpart. Although the inborn connection with frequency domain, existing pan-sharpening research has almost investigated the potential solution upon frequency domain, thus limiting the model performance improvement. To this end, we first revisit the degradation process of pan-sharpening in Fourier space, and then devise a Pyramid Dual Domain Injection pan-sharpening Network upon the above observation by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the proposed network is organized with multi-scale U-shape manner and composed by two core parts: a spatial guidance pyramid sub-network for fusing local spatial information and a frequency guidance pyramid sub-network for fusing global frequency domain information, thus encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to enable generating high-quality pan-sharpening results. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes. Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
ICCV | 1 |
| 2023 | Multiscale Dual-Domain Guidance Network for Pan-SharpeningabstractThe goal of pan-sharpening is to produce a high-spatial-resolution multi-spectral (HRMS) image from a low-spatial-resolution multi-spectral (LRMS) counterpart by super-resolving the LRMS one under the guidance of a texture-rich panchromatic (PAN) image. Existing research has concentrated on using spatial information to generate HRMS images, but has neglected to investigate the frequency domain, which severely restricts the performance improvement. In this work, we propose a novel pan-sharpening approach, named Multi-Scale Dual-Domain Guidance Network (MSDDN) by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the network is inborn with multi-scale U-shape manner and composed by two core parts: a spatial guidance sub-network for fusing local spatial information and a frequency guidance sub-network for fusing global frequency domain information and encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to help it generate high-quality pan-sharpening results. Employing the proposed model on different datasets, the quantitative and qualitative results demonstrate that our method performs appreciatively against other state-of-the-art approaches and comprises a strong generalization ability for real-world scenes. The source code is available at https://github.com/alexhe101/MSDDN. Xuanhua He, Jie Zhang 0033, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 1 |