VLDB 2026 Research / reviewers in the wild / expert
Man Zhou 0003
dblp:165/8236-3
· DBLP profile ↗
77ranked-venue papers
21as first author
77since 2021 · last 2026
0000-0003-2872-605XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 15 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 7 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shuffle Mamba: State Space Models With Random Shuffle for Multi-Modal Image FusionabstractMulti-modal image fusion integrates complementary information from different modalities to produce enhanced and informative images. Although State-Space Models, such as Mamba, are proficient in long-range modeling with linear complexity, most Mamba-based approaches use fixed scanning strategies, which can introduce biased prior information. To mitigate this issue, we propose a novel Bayesian-inspired scanning strategy called Random Shuffle, supplemented by a theoretically feasible inverse shuffle to maintain information coordination invariance, aiming to eliminate biases associated with fixed sequence scanning. Based on this transformation pair, we customized the Shuffle Mamba Framework, penetrating modality-aware information representation and cross-modality information interaction across spatial and channel axes to ensure robust interaction and an unbiased global receptive field for multi-modal image fusion. Furthermore, we develop a testing methodology based on Monte-Carlo averaging to ensure the model’s output aligns more closely with expected results. Extensive experiments across multiple multi-modal image fusion tasks demonstrate the effectiveness of our proposed method, yielding excellent fusion quality compared to state-of-the-art alternatives. The code is available at https://github.com/caoke-963/Shuffle-Mamba. Ke Cao 0001, Xuanhua He, Tao Hu 0027, Chengjun Xie, Man Zhou 0003, Jie Zhang 0033 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Distilling Textual Priors From LLM to Efficient Image FusionabstractMulti-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances in text-guided methods leverage large model priors to overcome these limitations, but at the cost of significant computational overhead, both in memory and inference time. To address this challenge, we propose a novel framework for distilling large model priors, eliminating the need for text guidance during inference while dramatically reducing model size. Our framework utilizes a teacher-student architecture, where the teacher network incorporates large model priors and transfers this knowledge to a smaller student network via a tailored distillation process. Crucially, our experiments demonstrate that this knowledge transfer is the primary driver of performance gains, rather than mere architectural optimization. Additionally, we introduce a spatial-channel cross-fusion module to enhance the model’s ability to leverage textual priors across both spatial and channel dimensions. Our method achieves a favorable trade-off between computational efficiency and fusion quality. The distilled network, requiring only 10% of the parameters and inference time of the teacher network, retains 90% of its performance and outperforms existing SOTA methods. Extensive experiments demonstrate the effectiveness of our approach. Codes are available at https://github.com/Zirconium233/DTPF. Xuanhua He, Ke Cao 0001, Liu Liu 0012, Li Zhang 0104, Man Zhou 0003, Jie Zhang 0033, Dan Guo 0001, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Freq-RWKV: Granularity-Aware Spatial-Frequency Synergy via Dual-Domain Recurrent Scanning for Pan-sharpening
Xueheng Li, Xuanhua He, Tao Hu 0027, Jie Zhang 0033, Man Zhou 0003, Chengjun Xie, Yingying Wang 0005, Bo Huang 0001 |
ACM Multimedia | 5 |
| 2025 | WKV-sharing embraced random shuffle RWKV high-order modeling for pan-sharpeningabstractPan-sharpening aims to generate a spatially and spectrally enriched multi-spectral image by integrating complementary
cross-modality information from low-resolution multi-spectral image and texture-rich panchromatic counterpart. In this work, we propose a
WKV-sharing embraced random shuffle RWKV high-order modeling paradigm for pan-sharpening from Bayesian perspective, coupled with random weight manifold distribution training strategy derived from Functional theory to regularize the solution space adhering to the
following principles: 1) Random-shuffle RWKV. Recently, the Vision RWKV model, with its inherent linear complexity in global modeling,
has inspired us to explore its untapped potential in pan-sharpening tasks. However, its attention mechanism, relying on a recurrent
bidirectional scanning strategy, suffers from biased effects and demands significant processing time. To address this, we propose a novel
Bayesian-inspired scanning strategy called Random Shuffle, complemented by a theoretically-sound inverse shuffle to preserve
information coordination invariance, effectively eliminating biases associated with fixed sequence scanning. The Random Shuffle
approach mitigates preconceptions in global 2D dependencies in mathematical expectation, providing the model with an unbiased prior.
In line with similar spirit of Dropout, we introduce a testing methodology based on Monte Carlo averaging to ensure the model’s output
aligns more closely with expected results. 2) WKV-sharing high-order. Regarding KV’s attention score calculation in spatial mixer of RWKV, we leverage WKV-sharing mechanism to transfer KV activations across RWKV layers, achieving lower latency and improved trainability, and revisit the channel mixer in RWKV, originally a first-order weighting function, and redevelop its high-order potential by sharing the gate mechanism across RWKV layer. Comprehensive experiments across pan-sharpening benchmarks demonstrate our model’s effectiveness, consistently outperforming state-of-the-art alternatives Man Zhou 0003, Xuanhua He, Danfeng Hong, Bo Huang 0001 |
NeurIPS | 1 |
| 2025 | Toward Resolution Mismatching: Modality-Aware Feature-Aligned Network for Pan-SharpeningabstractPanchromatic (PAN) and multi-spectral (MS) remote satellite image fusion, known as pan-sharpening, aims to produce high-resolution MS images by combining the complementary information from the high-resolution, texture-rich PAN and the low-resolution but high spectral-resolution MS counterparts. Despite notable advancements in this field, the current state-of-the-art pan-sharpening techniques do not explicitly address the spatial resolution mismatching problem between the two modalities of PAN and MS images. This mismatching issue can lead to misalignment in feature representation and the creation of blurry artifacts in the model output, ultimately hindering the generation of high-frequency textures and impeding the performance improvement of such methods. To address the aforementioned spatial resolution mismatching problem in pan-sharpening, we propose a novel modality-aware feature-aligned pan-sharpening framework in this paper. The framework comprises three primary stages: modality-aware feature extraction, modality-aware feature aligning, and context integrated image reconstruction. First, we introduce the half-instance normalization strategy as the backbone to filter out the inconsistent features and promote the learning of consistent features between the PAN and MS modalities. Second, a learnable modality-aware feature interpolation is devised to effectively address the misalignment issue. Specifically, the extracted features from the backbone are integrated to predict the transformation offsets of each pixel, which allows for the adaptive selection of custom contextual information and enables the modality-aware features to be more aligned. Finally, within the context of the interactive offset correction, multi-stage information is aggregated to generate the feasible pan-sharpened model output. Extensive experimental results over multiple satellite datasets demonstrate that the proposed algorithm outperforms other state-of-the-art methods both qualitatively and quantitatively, exhibiting great generalization ability to real-world scenes. Man Zhou 0003, Xuanhua He, Danfeng Hong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | A General Spatial-Frequency Learning Framework for Multimodal Image FusionabstractMultimodal image fusion involves tasks like pan-sharpening and depth super-resolution. Both tasks aim to generate high-resolution target images by fusing the complementary information from the texture-rich guidance and low-resolution target counterparts. They are inborn with reconstructing high-frequency information. Despite their inherent frequency domain connection, most existing methods only operate solely in the spatial domain and rarely explore the solutions in the frequency domain. This study addresses this limitation by proposing solutions in both the spatial and frequency domains. We devise a Spatial-Frequency Information Integration Network, abbreviated as SFINet for this purpose. The SFINet includes a core module tailored for image fusion. This module consists of three key components: a spatial-domain information branch, a frequency-domain information branch, and a dual-domain interaction. The spatial-domain information branch employs the spatial convolution-equipped invertible neural operators to integrate local information from different modalities in the spatial domain. Meanwhile, the frequency-domain information branch adopts a modality-aware deep Fourier transformation to capture the image-wide receptive field for exploring global contextual information. In addition, the dual-domain interaction facilitates information flow and the learning of complementary representations. We further present an improved version of SFINet, SFINet++, that enhances the representation of spatial information by replacing the basic convolution unit in the original spatial domain branch with the information-lossless invertible neural operator. We conduct extensive experiments to validate the effectiveness of the proposed networks and demonstrate their outstanding performance against state-of-the-art methods in two representative multimodal image fusion tasks: pan-sharpening and depth super-resolution. Man Zhou 0003, Jie Huang 0017, Danfeng Hong, Xiuping Jia, Jocelyn Chanussot, Chongyi Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Probing Synergistic High-Order Interaction for Multi-Modal Image FusionabstractMulti-modal image fusion aims to generate a fused image by integrating and distinguishing the cross-modality complementary information from multiple source images. While the cross-attention mechanism with global spatial interactions appears promising, it only captures second-order spatial interactions, neglecting higher-order interactions in both spatial and channel dimensions. This limitation hampers the exploitation of synergies between multi-modalities. To bridge this gap, we introduce a Synergistic High-order Interaction Paradigm (SHIP), designed to systematically investigate spatial fine-grained and global statistics collaborations between the multi-modal images across two fundamental dimensions: 1) Spatial dimension: we construct spatial fine-grained interactions through element-wise multiplication, mathematically equivalent to global interactions, and then foster high-order formats by iteratively aggregating and evolving complementary information, enhancing both efficiency and flexibility. 2) Channel dimension: expanding on channel interactions with first-order statistics (mean), we devise high-order channel interactions to facilitate the discernment of inter-dependencies between source images based on global statistics. We further introduce an enhanced version of the SHIP model, called SHIP++ that enhances the cross-modality information interaction representation by the cross-order attention evolving mechanism, cross-order information integration, and residual information memorizing mechanism. Harnessing high-order interactions significantly enhances our model's ability to exploit multi-modal synergies, leading in superior performance over state-of-the-art alternatives, as shown through comprehensive experiments across various benchmarks in two significant multi-modal image fusion tasks: pan-sharpening, and infrared and visible image fusion. Man Zhou 0003, Naishan Zheng, Xuanhua He, Danfeng Hong, Jocelyn Chanussot |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Frequency Decoupled Domain-Irrelevant Feature Learning for Pan-SharpeningabstractPan-sharpening aims to generate high-detail multi-spectral images (HRMS) through the fusion of panchromatic (PAN) and multi-spectral (MS) images. However, existing pan-sharpening methods often suffer from significant performance degradation when dealing with out-of-distribution data, as they assume the training and test datasets are independent and identically distributed. To overcome this challenge, we propose a novel frequency domain-irrelevant feature learning framework that exhibits exceptional generalization capabilities. Our approach involves parallel extraction and processing of domain-irrelevant information from the amplitude and phase components of the input images. Specifically, we design a frequency information separation module to extract the amplitude and phase components of the paired images. The learnable high-pass filter is then employed to eliminate domain-specific information from the amplitude spectrums. After that, we devised two specialized sub-networks (AFL-Net and PFL-Net) to perform targeted learning of the frequency domain-irrelevant information. This allows our method to effectively capture the complementary domain-irrelevant information contained in the amplitude and phase spectra of the images. Finally, the information fusion and restoration module dynamically adjusts the feature channel weights, enabling the network to output high-quality HRMS images. Through this frequency domain-irrelevant feature learning framework, our method balances generalization capability and network performance on the distribution of training dataset. Extensive experiments conducted on various satellite datasets demonstrate the effectiveness of our method for generalized pan-sharpening. Our proposed network outperforms state-of-the-art methods in terms of both quantitative metrics and visual quality, showcasing its superior ability to handle diverse, out-of-distribution data. Jie Zhang 0033, Ke Cao 0001, Yunlong Lin, Xuanhua He, Yingying Wang 0005, Rui Li 0027, Chengjun Xie, Jun Zhang 0034, Man Zhou 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | PanDiT: A Few-Step Diffusion Transformer for High-Fidelity and Efficient PansharpeningabstractPansharpening plays a crucial role in remote sensing by fusing low-resolution multispectral (LRMS) images and high-resolution panchromatic (PAN) images to generate high-resolution multispectral (HRMS) images. While denoising diffusion models offer potential for high-fidelity image generation, their practical application in pansharpening has been severely hindered by huge computational costs from iterative sampling and naive conditioning strategies that struggle to fuse multi-modal information effectively. In this paper, we introduce PanDiT, a novel Diffusion Transformer framework designed to address these challenges. PanDiT is built on the core principle of a Decoupled Conditioning Mechanism, which explicitly disentangles and injects spatial and spectral guidance, and is engineered for practical, few-step inference. Our framework leverages a powerful Diffusion Transformer (DiT) backbone, where conditioning is achieved through two specialized injection blocks that capture spatial and time-frequency features. Crucially, by integrating an implicit sampling strategy, we accelerate the inference process to as few as two steps. Extensive experiments on multiple benchmark datasets demonstrate that PanDiT not only establishes a new state-of-the-art in fusion quality and achieves a good quality-efficiency trade-off. Code is available at https://github.com/para133/PanDIT. Jiabin Fang, Ke Cao 0001, Xuanhua He, Jie Zhang 0033, Man Zhou 0003, Liu Liu 0012 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Bilateral Adaptive Evolution Transformer for Multispectral Image FusionabstractPansharpening is the critical technology for generating high-resolution (HR) multispectral (MS) images by learning the cross-modality complementary representations between the panchromatic (PAN) images and low-resolution (LR) MS images. Though methods based on convolutional neural networks (CNNs) have dominated the pansharpening community, they still suffer from the limited global modeling capability due to the inherent property of the convolutional operator. To remedy this common limitation, the transformer family has recently gained great popularity in this field. However, existing cascaded transformer designs inevitably introduce a heavy memory footprint and computational cost due to the dense dot-product self-attention (SA) computation. More importantly, these paradigms simply ignore the innate sparsity of remote sensing images, leading to information redundancy and a challenging optimization process. To alleviate these issues, we propose the bilateral adaptive evolution transformer (BAEFormer), which is built upon two core mechanisms: bilateral attention computation and adaptive attention evolution. Specifically, we first decompose the conventional quadratic complexity SA into linear-degree height and width computing at the first stage, respectively, which significantly reduces the computational complexity. Given the data-specific properties, furthermore, we devise a novel yet effective neighboring layer-dependent strategy to adaptively update the attention map of two spatial dimensions, thereby avoiding the repetitive SA computation while taking into account the dynamics toward the evolution of attention weights. Our model, called BAEFormer, outperforms other state-of-the-art pansharpening methods on various remote sensing datasets while showing fewer network parameters and computational requirements. The code is available athttps://github.com/coder-JMHou/BAEFormer. Junming Hou, Chenxu Wu, Man Zhou 0003, Junling Li, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | A General Cooperative Optimization Driven High-Frequency Enhancement Framework for Multispectral Image FusionabstractPan-sharpening essentially to boost the spatial resolution of a multispectral (MS) image guided by its paired panchromatic (PAN) image. In other words, this process intricately integrates the high-frequency components extracted from texture-rich PAN images into the low-resolution (LR) MS images, resulting in texture-rich MS images. Though existing deep learning (DL)-based techniques have made impressive performance compared with traditional algorithms, they still face challenges in accurately restoring high-frequency details in MS images, thus limiting overall pan-sharpening performance. In addition, reference high-resolution (HR) MS images are often underutilized, typically serving only as training labels. In this work, we present a general high-frequency enhancement framework for pan-sharpening, which is implemented through a cooperative optimization strategy using mutual information (MI) maximization and contrastive learning. Specifically, our model comprises two fundamental modules: the high-frequency feature alignment (HFFA) module and the high-frequency detail calibration (HFDC) module. The first employs MI maximization to align the high-frequency semantic statistical distribution between PAN images and reference HRMS images. The latter is designed to calibrate the high-frequency components of MS modality under the guidance of the PAN counterparts through the contrastive learning constraint, thereby producing more accurate high-frequency information on MS modality. By integrating the calibrated high-frequency features of MS modality and those of PAN modality, we can obtain a more comprehensive and precise high-frequency feature representation of these two modalities, facilitating the reconstruction of LRMS images. Our model, incorporating the aforementioned key elements, significantly surpasses other state-of-the-art (SOTA) techniques across multiple satellite datasets in both quantitative and qualitative experiments. Moreover, the real-world full-resolution and cross-sensor assessments testify to its exceptional generalization capabilities. The code is available athttps://github.com/Vcocoi/CONet. Chentong Huang, Junming Hou, Chenxu Wu, Xiaofeng Cong, Man Zhou 0003, Junling Li, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Pan-Sharpening via Causal-Aware Feature Distribution CalibrationabstractIn this work, we reveal an interesting observation within the multi-spectral modality: high-frequency components exhibit a long-tailed distribution, in contrast to the Gaussian distribution of dominant low-frequency components. This dual-distribution characteristic presents a challenge for network optimization, leading to overfitting on low-frequency information while neglecting essential high-frequency details. Addressing this issue from a causal inference perspective, we identify optimizer momentum as a confounding factor that biases models towards focusing on the head part of the high-frequency distribution during training. To counteract this effect, we propose a novel optimization strategy and supplement the global-modeling network architecture to balance the frequencies learning. In the training stage, we employ the Recurrent Weighted Key-Value (RWKV) architecture, which features a global receptive field, to effectively learn the long-tailed distribution of high-frequency components and quantify the cumulative direction of feature bias. During the testing stage, we apply counterfactual reasoning to adjust feature distributions based on the quantified bias. To our knowledge, this is the first time to investigate the imbalance of frequency learning within pan-sharpening from the causal inference perspective. Extensive experiments on three benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, showcasing its effectiveness and robustness in pan-sharpening tasks. Xueheng Li, Tao Hu 0027, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Exploring Text-Guided Information Fusion Through Chain-of-Reasoning for PansharpeningabstractPan-sharpening aims to enhance the spatial resolution of low-resolution multispectral (LRMS) images by integrating high-frequency information from a corresponding texture-rich panchromatic (PAN) image, while maintaining the spectral integrity of the LRMS image. Although text-guided multi-modal learning has made considerable strides in the natural image domain, its potential to pan-sharpening remains underexplored, primarily due to the limited availability of multi-modal remote sensing datasets. To this end, we construct an entirely new pan-sharpening framework by making efforts from three key aspects: (1) text-equipped multi-modal data collection through chain-of-reasoning, (2) large model prior-driven multi-modal information fusion, and (3) visual information interaction through prompt engineering, leveraging textual information to guide the pan-sharpening process within a multi-modal fusion framework. We initially utilize the generic large language model priors to generate descriptive captions for MS images, forming a multi-modal pan-sharpening dataset. By integrating super-resolved imagery and segmentation maps generated by segment anything, we apply Chain-of-Thought (CoT) prompting to generate spatially focused captions across diverse satellite datasets. These captions enhance visual features and provide high-level contextual information, improving semantic understanding for pan-sharpening. Building on the aforementioned multi-modal data, we tailor two text-guided information fusion modules: Textual Enhancement Block (TEB) standing on large model prior and Textual Modulated Block (TMB) utilizing text information to effectively guide and refine the pan-sharpening fusion process. Extensive experiments on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art methods, highlighting its effectiveness and superior performance in pan-sharpening. Xueheng Li, Xuanhua He, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong, Bo Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier DomainabstractRAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB images, a difference that goes beyond the color matrix and extends to spatial structure due to resolution variations. Recent methods directly rebuild color mapping and spatial structure via shared deep representation, limiting optimal performance. Inspired by Image Signal Processing (ISP) pipeline, which distinguishes image restoration and enhancement, we present a novel Neural ISP framework, named FourierISP. This approach breaks the image down into style and structure within the frequency domain, allowing for independent optimization. FourierISP is comprised of three subnetworks: Phase Enhance Subnet for structural refinement, Amplitude Refine Subnet for color learning, and Color Adaptation Subnet for blending them in a smooth manner. This approach sharpens both color and structure, and extensive evaluations across varied datasets confirm that our approach realizes state-of-the-art results. Code will be available at https://github.com/alexhe101/FourierISP. Xuanhua He, Tao Hu 0027, Guoli Wang 0004, Zejin Wang, Qian Zhang 0009, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
AAAI | 12 |
| 2024 | Frequency-Adaptive Pan-Sharpening with Mixture of ExpertsabstractPan-sharpening involves reconstructing missing high-frequency information in multi-spectral images with low spatial resolution, using a higher-resolution panchromatic image as guidance. Although the inborn connection with frequency domain, existing pan-sharpening research has not almost investigated the potential solution upon frequency domain. To this end, we propose a novel Frequency Adaptive Mixture of Experts (FAME) learning framework for pan-sharpening, which consists of three key components: the Adaptive Frequency Separation Prediction Module, the Sub-Frequency Learning Expert Module, and the Expert Mixture Module. In detail, the first leverages the discrete cosine transform to perform frequency separation by predicting the frequency mask. On the basis of generated mask, the second with low-frequency MOE and high-frequency MOE takes account for enabling the effective low-frequency and high-frequency information reconstruction. Followed by, the final fusion module dynamically weights high frequency and low-frequency MOE knowledge to adapt to remote sensing images with significant content variations. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes. Code will be made publicly at https://github.com/alexhe101/FAME-Net. Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
AAAI | 6 |
| 2024 | Revisiting Spatial-Frequency Information Integration from a Hierarchical Perspective for Panchromatic and Multi-Spectral Image FusionabstractPan-sharpening is a super-resolution problem that essentially relies on spectra fusion of panchromatic (PAN) images and low-resolution multi-spectral (LRMS) images. The previous methods have validated the effectiveness of information fusion in the Fourier space of the whole image. However, they haven't fully explored the Fourier relationships at different hierarchies between PAN and LRMS images. To this end, we propose a Hierarchical Frequency Integration Network (HFIN) to facilitate hierarchical Fourier information integration for pan-sharpening. Specifically, our network consists of two designs: information stratification and information integration. For information stratification, we hierarchically decompose PAN and LRMS information into spatial, global Fourier and local Fourier information, and fuse them independently. For information integration, the above hierarchical fused information is processed to further enhance their relationships and undergo comprehensive integration. Our method extend a new space for exploring the relationships of PAN and LRMS images, enhancing the integration of spatial-frequency information. Extensive experiments robustly validate the effectiveness of the proposed network, showcasing its superior performance compared to other state-of-the-art methods and generalization in real-world scenes and other fusion tasks as a general image fusion framework. Code is available at https://github.com/JosephTiTan/HFIN. Jiangtong Tan, Jie Huang 0017, Naishan Zheng, Man Zhou 0003, Danfeng Hong, Feng Zhao 0004 |
CVPR | 4 |
| 2024 | Empowering Resampling Operation for Ultra-High-Definition Image Enhancement with Model-Aware GuidanceabstractImage enhancement algorithms have made remarkable advancements in recent years, but directly applying them to Ultra-high-definition (UHD) images presents intractable computational overheads. Therefore, previous straightforward solutions employ resampling techniques to reduce the resolution by adopting a “Downsampling-Enhancement-Upsampling” processing paradigm. However, this paradigm disentangles the resampling operators and inner enhancement algorithms, which results in the loss of information that is favored by the model, further leading to sub-optimal outcomes. In this paper, we propose a novel method of Learning Model-Aware Resampling (LMAR), which learns to customize resampling by extracting model-aware information from the UHD input image, under the guidance of model knowledge. Specifically, our method consists of two core designs, namely compensatory kernel estimation and steganographic resampling. At the first stage, we dynamically predict compensatory kernels tailored to the specific input and resampling scales. At the second stage, the image-wise compensatory information is derived with the compensatory kernels and embedded into the rescaled input images. This promotes the representation of the newly derived downscaled inputs to be more consistent with the full-resolution UHD inputs, as perceived by the model. Our LMAR enables model-aware and model-favored resampling while maintaining compatibility with existing resampling operators. Extensive experiments on multiple UHD image enhancement datasets and different backbones have shown consistent performance gains after correlating resizer and enhancer; e.g., up to 1.2dB PSNR gain for ×1.8 resampling scale on UHD-LOL4K. The code is available at https://github.com/YPatrickW/LMAR. Jie Huang 0017, Bing Li 0024, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004 |
CVPR | 6 |
| 2024 | Probing Synergistic High-Order Interaction in Infrared and Visible Image FusionabstractInfrared and visible image fusion aims to generate a fused image by integrating and distinguishing complementary information from multiple sources. While the cross-attention mechanism with global spatial interactions appears promising, it only capture second-order spatial inter-actions, neglecting higher-order interactions in both spatial and channel dimensions. This limitation hampers the ex-ploitation of synergies between multi-modalities. To bridge this gap, we introduce a Synergistic High-order Interaction Paradigm (SHIP), designed to systematically investigate the spatial fine-grained and global statistics collaborations between infrared and visible images across two fundamental dimensions: 1) Spatial dimension: we construct spatial fine-grained interactions through element-wise multiplication, mathematically equivalent to global interactions, and then foster high-order formats by iteratively aggregating and evolving complementary information, enhancing both efficiency andflexibility; 2) Channel dimension: expanding on channel interactions with first-order statistics (mean), we devise high-order channel interactions to facilitate the discernment of inter-dependencies between source images based on global statistics. Harnessing high-order interactions significantly enhances our model's ability to exploit multi-modal synergies, leading to superior performance over state-of-the-art alternatives, as shown through comprehensive experiments across various benchmarks. Code is available at https://github.com/zheng980629/SHIP. Naishan Zheng, Man Zhou 0003, Jie Huang 0017, Junming Hou, Haoying Li, Feng Zhao 0004 |
CVPR | 2 |
| 2024 | Linearly-evolved Transformer for Pan-sharpening
Junming Hou, Zihan Cao, Naishan Zheng, Xuan Li 0012, Xiaofeng Cong, Danfeng Hong, Man Zhou 0003 |
ACM Multimedia | 9 |
| 2024 | HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image EnhancementabstractUnderwater image enhancement presents a significant challenge due to the complex and diverse underwater environments that result in severe degradation phenomena such as light absorption, scattering, and color distortion. More importantly, obtaining paired training data for these scenarios is a challenging task, which further hinders the generalization performance of enhancement models. To address these issues, we propose a novel approach, the Hybrid Contrastive Learning Regularization (HCLR-Net). Our method is built upon a distinctive hybrid contrastive learning regularization strategy that incorporates a unique methodology for constructing negative samples. This approach enables the network to develop a more robust sample distribution. Notably, we utilize non-paired data for both positive and negative samples, with negative samples are innovatively reconstructed using local patch perturbations. This strategy overcomes the constraints of relying solely on paired data, boosting the model’s potential for generalization. The HCLR-Net also incorporates an Adaptive Hybrid Attention module and a Detail Repair Branch for effective feature extraction and texture detail restoration, respectively. Comprehensive experiments demonstrate the superiority of our method, which shows substantial improvements over several state-of-the-art methods in terms of quantitative metrics, significantly enhances the visual quality of underwater images, establishing its innovative and practical applicability. Our code is available at: https://github.com/zhoujingchun03/HCLR-Net . Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 5 |
| 2024 | Correction: HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image Enhancement
Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 5 |
| 2024 | Learning Spatio-Temporal Sharpness Map for Video DeblurringabstractVideo deblurring is a challenging task because only input blurry sequences are available. To further constrain the optimization process, existing methods explore various additional information,e.g., events, depth and sharpness prior. However, they consume large computing costs or generate unpleasant visual results due to the insufficient exploitation of spatio-temporal information. In this work, we develop a novel spatio-temporal sharpness map learned by a prior-based generation network implicitly. The proposed generation network blends both spatial and temporal sharpness priors in a blurry sequence, while few extra parameters are added. We show that the proposed map has better spatial continuity and guidance for video deblurring than the previous method. Furthermore, different from the simply concatenation in the previous work, we allow the sharpness map to accommodate to more effective video deblurring via a dual-stream network. Specifically, the network is decomposed by two branches, namely inter-frame and intra-frame reconstructions. The inter-frame reconstruction obtains the sharp patches of cecutive frames from the sharpness map to restore textures well. Meanwhile, the other intra-frame branch is responsible for recovering structures of the latent frame, where a novel histogram statistical method is developed to quantify and count textures in the feature under the modulation of the sharpness map. Quantitative and qualitative experiments successfully validate the effectiveness of our proposed method. Qi Zhu 0010, Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Pan-Sharpening With Wavelet-Enhanced High-Frequency InformationabstractPan-sharpening is essentially a panchromatic (PAN)-guided super-resolution process, primarily focused on enhancing multi-spectral image quality. This methodology intricately incorporates the high-frequency derived from texture-rich PAN images into the lower-resolution multi-spectral (LRMS) counterparts. However, current spatial domain techniques frequently face challenges in accurately restoring texture details, while frequency domain methods lack efficient interaction with spatial domains, thus restricting the overall model performance. In response to these challenges, we introduce a novel High-frequency Wavelet Network that capitalizes on the spatial-frequency interaction and frequency division capabilities inherent in wavelet transform. In particular, our approach consists of two fundamental modules: the Wavelet-Inspired Fusion Block and the High-Frequency Enhancement Block. The former is inspired by wavelet lifting schemes, enabling the fusion of frequencies and facilitating information exchange across various subbands. The latter harnesses wavelet’s frequency division attributes to enhance high-frequency information learning. Comprehensive experiments over multiple satellite datasets demonstrate that our approach outperforms state-of-the-art techniques in both quantitative and qualitative assessments. Moreover, our model showcases exceptional generalization capabilities in real-world scenarios. Code is available at https://github.com/alexhe101/WINet. Jie Zhang 0033, Xuanhua He, Ke Cao 0001, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | IRVR: A General Image Restoration Framework for Visual RecognitionabstractImages corrupted with degradations often result in a performance drop in downstream image recognition models trained on clean images. Previous image restoration (IR) methods either restore the images without delicately considering the semantic recovery, or the training objectives cannot meet unseen recognition models, leading to poor and non-generalizable performance for various downstream recognition tasks. In this paper, we propose a general image restoration framework for visual recognition, IRVR, which is addressed for generalized and effective semantic recovery in image restoration for a range of high-level tasks. Concretely, for better generalization, we train the IR models with semantic recovery as the primary objective, and image regression as a regularization term, respectively, where the primary objective gradient is calibrated with the regularization gradient to ensure the generalization of IR to unseen recognition models. For effectiveness, we introduce an intrinsic semantic consistency constraint to match the semantic statistical distribution between restored and clean image pairs. Our IRVR is recognition-agnostic and orthogonal to IR, making it a plug-and-play component that can be incorporated into existing IR methods without adding any computation cost during inference. Extensive experiments demonstrate the effectiveness and generalization of our IRVR for improving the performance of IR in diverse downstream high-level tasks. The IRVR's ability to accurately recover intrinsic semantics in images is instrumental in high-level machine analysis, which ensures the integrity and authenticity of multimedia content. Zizheng Yang, Jie Huang 0017, Man Zhou 0003, Naishan Zheng, Feng Zhao 0004 |
IEEE Trans. Multim. | 3 |
| 2024 | Rethinking Pan-Sharpening in Closed-Loop RegularizationabstractIt is generally known that pan-sharpening is fundamentally a PAN-guided multispectral (MS) image super-resolution problem that involves learning the nonlinear mapping from low-resolution (LR) to high-resolution (HR) MS images. Since an infinite number of HR-MS images can be downsampled to produce the same corresponding LR-MS image, learning the mapping from LR-MS to HR-MS image is typically ill-posed and the space of the possible pan-sharpening functions can be extremely large, making it difficult to estimate the optimal mapping solution. To address the above issue, we propose a closed-loop scheme that learns the two opposite mapping including the pan-sharpening and its corresponding degradation process simultaneously to regularize the solution space in a single pipeline. More specifically, an invertible neural network (INN) is introduced to perform a bidirectional closed-loop: the forward operation for LR-MS pan-sharpening and the backward operation for learning the corresponding HR-MS image degradation process. In addition, given the vital importance of high-frequency textures for the Pan-sharpened MS images, we further strengthen the INN by designing a specified multiscale high-frequency texture extraction module. Extensive experimental results demonstrate that the proposed algorithm performs favorably against state-of-the-art methods qualitatively and quantitatively with fewer parameters. Ablation studies also verify the effectiveness of the closed-loop mechanism in pan-sharpening. The source code is made publicly available at https://github.com/manman1995/pan-sharpening-Team-zhouman/. Man Zhou 0003, Jie Huang 0017, Danfeng Hong, Feng Zhao 0004, Chongyi Li, Jocelyn Chanussot |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Learning Semantic Degradation-Aware Guidance for Recognition-Driven Unsupervised Low-Light Image EnhancementabstractLow-light images suffer severe degradation of low lightness and noise corruption, causing unsatisfactory visual quality and visual recognition performance. To solve this problem while meeting the unavailability of paired datasets in wide-range scenarios, unsupervised low-light image enhancement (ULLIE) techniques have been developed. However, these methods are primarily guided to alleviate the degradation effect on visual quality rather than semantic levels, hence limiting their performance in visual recognition tasks. To this end, we propose to learn a Semantic Degradation-Aware Guidance (SDAG) that perceives the low-light degradation effect on semantic levels in a self-supervised manner, which is further utilized to guide the ULLIE methods. The proposed SDAG utilizes the low-light degradation factors as augmented signals to degrade the low-light images, and then capture their degradation effect on semantic levels. Specifically, our SDAG employs the subsequent pre-trained recognition model extractor to extract semantic representations, and then learns to self-reconstruct the enhanced low-light image and its augmented degraded images. By constraining the relative reconstruction effect between the original enhanced image and the augmented formats, our SDAG learns to be aware of the degradation effect on semantic levels in a relative comparison manner. Moreover, our SDAG is general and can be plugged into the training paradigm of the existing ULLIE methods. Extensive experiments demonstrate its effectiveness for improving the ULLIE approaches on the downstream recognition tasks while maintaining a competitive visual quality. Code will be available at https://github.com/zheng980629/SDAG. Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Zizheng Yang, Qi Zhu 0010, Feng Zhao 0004 |
AAAI | 3 |
| 2023 | Learning Sample Relationship for Exposure CorrectionabstractExposure correction task aims to correct the underexposure and its adverse overexposure images to the normal exposure in a single network. As well recognized, the optimization flow is the opposite. Despite great advancement, existing exposure correction methods are usually trained with a mini-batch of both underexposure and overexposure mixed samples and have not explored the relationship between them to solve the optimization inconsistency. In this paper, we introduce a new perspective to conjunct their optimization processes by correlating and constraining the relationship of correction procedure in a mini-batch. The core designs of our framework consist of two steps: 1) formulating the exposure relationship of samples across the batch dimension via a context-irrelevant pretext task. 2) delivering the above sample relationship design as the regularization term within the loss function to promote optimization consistency. The proposed sample relationship design as a general term can be easily integrated into existing exposure correction methods without any computational burden in inference time. Extensive experiments over multiple representative exposure correction benchmarks demonstrate consistent performance gains by introducing our sample relationship design. Jie Huang 0017, Feng Zhao 0004, Man Zhou 0003, Jie Xiao 0002, Naishan Zheng, Zhiwei Xiong |
CVPR | 3 |
| 2023 | Visual Recognition-Driven Image Restoration for Multiple Degradation with Intrinsic Semantics RecoveryabstractDeep image recognition models suffer a significant performance drop when applied to low-quality images since they are trained on high-quality images. Although many studies have investigated to solve the issue through image restoration or domain adaptation, the former focuses on visual quality rather than recognition quality, while the latter requires semantic annotations for task-specific training. In this paper, to address more practical scenarios, we propose a Visual Recognition-Driven Image Restoration network for multiple degradation, dubbed VRD-IR, to recover high-quality images from various unknown corruption types from the perspective of visual recognition within one model. Concretely, we harmonize the semantic representations of diverse degraded images into a unified space in a dynamic manner, and then optimize them towards intrinsic semantics recovery. Moreover, a prior-ascribing optimization strategy is introduced to encourage VRD-IR to couple with various downstream recognition tasks better. Our VRD-IR is corruption- and recognition-agnostic, and can be inserted into various recognition tasks directly as an image enhancement module. Extensive experiments on multiple image distortions demonstrate that our VRD-IR surpasses existing image restoration methods and show superior performance on diverse high-level tasks, including classification, detection, and person re-identification. Zizheng Yang, Jie Huang 0017, Man Zhou 0003, Hu Yu 0001, Feng Zhao 0004 |
CVPR | 4 |
| 2023 | Ingredient-oriented Multi-Degradation Learning for Image RestorationabstractLearning to leverage the relationship among diverse image restoration tasks is quite beneficial for unraveling the intrinsicingredients behind the degradation. Recent years have witnessed the flourish of various All-in-one methods, which handle multiple image degradations within a single model. In practice, however, few attempts have been made to excavate task correlations in that exploring the underlying fundamentalingredients of various image degradations, resulting in poor scalability as more tasks are involved. In this paper, we propose a novel perspective to delve into the degradation via aningredients-oriented rather than previous task-oriented manner for scalable learning. Specifically, our method, named Ingredients-oriented Degradation Reformulation framework (IDR), consists of two stages, namely task-oriented knowledge collection and ingredients-oriented knowledge integration. In the first stage, we conduct ad hoc operations on different degradations according to the underlying physics principles, and establish the corresponding prior hubs for each type of degradation. While the second stage progressively reformulates the preceding task-oriented hubs into single ingredients-oriented hub via learnable Principal Component Analysis (PCA), and employs a dynamic routing mechanism for probabilistic unknown degradation removal. Extensive experiments on various image restoration tasks demonstrate the effectiveness and scalability of our method. More importantly, our IDR exhibits the favorable generalization ability to unknown downstream tasks. Jie Huang 0017, Mingde Yao, Zizheng Yang, Hu Yu 0001, Man Zhou 0003, Feng Zhao 0004 |
CVPR | 6 |
| 2023 | Probability-based Global Cross-modal Upsampling for PansharpeningabstractPansharpening is an essential preprocessing step for remote sensing image processing. Although deep learning (DL) approaches performed well on this task, current upsampling methods used in these approaches only utilize the local information of each pixel in the low-resolution multispectral (LRMS) image while neglecting to exploit its global information as well as the cross-modal information of the guiding panchromatic (PAN) image, which limits their performance improvement. To address this issue, this paper develops a novel probability-based global cross-modal upsampling (PGCU) method for pan-sharpening. Precisely, we first formulate the PGCU method from a probabilistic perspective and then design an efficient network module to implement it by fully utilizing the information mentioned above while simultaneously considering the channel specificity. The PGCU module consists of three blocks, i.e., information extraction (IE), distribution and expectation estimation (DEE), and fine adjustment (FA). Extensive experiments verify the superiority of the PGCU method compared with other popular upsampling methods. Additionally, experiments also show that the PGCU module can help improve the performance of existing SOTA deep learning pansharpening methods. The codes are available at https://github.com/Zeyu-Zhu/PGCU. Xiangyong Cao, Man Zhou 0003, Deyu Meng |
CVPR | 3 |
| 2023 | Pyramid Dual Domain Injection Network for Pan-sharpeningabstractPan-sharpening, a panchromatic image guided low-spatial-resolution multi-spectral super-resolution task, aims to reconstruct the missing high-frequency information of high-resolution multi-spectral counterpart. Although the inborn connection with frequency domain, existing pan-sharpening research has almost investigated the potential solution upon frequency domain, thus limiting the model performance improvement. To this end, we first revisit the degradation process of pan-sharpening in Fourier space, and then devise a Pyramid Dual Domain Injection pan-sharpening Network upon the above observation by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the proposed network is organized with multi-scale U-shape manner and composed by two core parts: a spatial guidance pyramid sub-network for fusing local spatial information and a frequency guidance pyramid sub-network for fusing global frequency domain information, thus encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to enable generating high-quality pan-sharpening results. Quantitative and qualitative experiments over multiple datasets demonstrate that our method performs the best against other state-of-the-art ones and comprises a strong generalization ability for real-world scenes. Xuanhua He, Rui Li 0027, Chengjun Xie, Jie Zhang 0033, Man Zhou 0003 |
ICCV | 6 |
| 2023 | PanFlowNet: A Flow-Based Deep Network for Pan-sharpeningabstractPan-sharpening aims to generate a high-resolution multispectral (HRMS) image by integrating the spectral information of a low-resolution multispectral (LRMS) image with the texture details of a high-resolution panchromatic (PAN) image. It essentially inherits the ill-posed nature of the super-resolution (SR) task that diverse HRMS images can degrade into an LRMS image. However, existing deep learning-based methods recover only one HRMS image from the LRMS image and PAN image using a deterministic mapping, thus ignoring the diversity of the HRMS image. In this paper, to alleviate this ill-posed issue, we propose a flow-based pan-sharpening network (PanFlowNet) to directly learn the conditional distribution of HRMS image given LRMS image and PAN image instead of learning a deterministic mapping. Specifically, we first transform this unknown conditional distribution into a given Gaussian distribution by an invertible network, and the conditional distribution can thus be explicitly defined. Then, we design an invertible Conditional Affine Coupling Block (CACB) and further build the architecture of PanFlowNet by stacking a series of CACBs. Finally, the PanFlowNet is trained by maximizing the log-likelihood of the conditional distribution given a training set and can then be used to predict diverse HRMS images. The experimental results verify that the proposed PanFlowNet can generate various HRMS images given an LRMS image and a PAN image. Additionally, the experimental results on different kinds of satellite datasets also demonstrate the superiority of our PanFlowNet compared with other state-of-the-art methods both visually and quantitatively. Code is available at Github. Xiangyong Cao, Wenzhe Xiao, Man Zhou 0003, Aiping Liu, Xun Chen 0001, Deyu Meng |
ICCV | 4 |
| 2023 | Generalized Lightness Adaptation with Channel Selective NormalizationabstractLightness adaptation is vital to the success of image processing to avoid unexpected visual deterioration, which covers multiple aspects, e.g., low-light image enhancement, image retouching, and inverse tone mapping. Existing methods typically work well on their trained lightness conditions but perform poorly in unknown ones due to their limited generalization ability. To address this limitation, we propose a novel generalized lightness adaptation algorithm that extends conventional normalization techniques through a channel filtering design, dubbed Channel Selective Normalization (CSNorm). The proposed CSNorm purposely normalizes the statistics of lightness-relevant channels and keeps other channels unchanged, so as to improve feature generalization and discrimination. To optimize CSNorm, we propose an alternating training strategy that effectively identifies lightness-relevant channels. The model equipped with our CSNorm only needs to be trained on one lightness condition and can be well generalized to unknown lightness conditions. Experimental results on multiple benchmark datasets demonstrate the effectiveness of CSNorm in enhancing the generalization ability for the existing lightness adaptation methods. Code is available at https://github.com/mdyao/CSNorm. Mingde Yao, Jie Huang 0017, Ruikang Xu, Shenglong Zhou 0002, Man Zhou 0003, Zhiwei Xiong |
ICCV | 6 |
| 2023 | Empowering Low-Light Image Enhancer through Customized Learnable PriorsabstractDeep neural networks have achieved remarkable progress in enhancing low-light images by improving their brightness and eliminating noise. However, most existing methods construct end-to-end mapping networks heuristically, neglecting the intrinsic prior of image enhancement task and lacking transparency and interpretability. Although some unfolding solutions have been proposed to relieve these issues, they rely on proximal operator networks that deliver ambiguous and implicit priors. In this work, we propose a paradigm for low-light image enhancement that explores the potential of customized learnable priors to improve the transparency of the deep unfolding paradigm. Motivated by the powerful feature representation capability of Masked Autoencoder (MAE), we customize MAE-based illumination and noise priors and redevelop them from two perspectives: 1) structure flow: we train the MAE from a normal-light image to its illumination properties and then embed it into the proximal operator design of the unfolding architecture; and 2) optimization flow: we train MAE from a normal-light image to its gradient representation and then employ it as a regularization term to constrain noise in the model output. These designs improve the interpretability and representation capability of the model. Extensive experiments on multiple low-light image enhancement datasets demonstrate the superiority of our proposed paradigm over state-of-the-art methods. Code is available at https://github.com/zheng980629/CUE. Naishan Zheng, Man Zhou 0003, Yanmeng Dong, Xiangyu Rui, Jie Huang 0017, Chongyi Li, Feng Zhao 0004 |
ICCV | 2 |
| 2023 | Learned Image Reasoning Prior Penetrates Deep Unfolding Network for Panchromatic and Multi-Spectral Image FusionabstractThe success of deep neural networks for pan-sharpening is commonly in a form of black box, lacking transparency and interpretability. To alleviate this issue, we propose a novel model-driven deep unfolding framework with image reasoning prior tailored for the pan-sharpening task. Different from existing unfolding solutions that deliver the proximal operator networks as the uncertain and vague priors, our framework is motivated by the content reasoning ability of masked autoencoders (MAE) with insightful designs. Specifically, the pre-trained MAE with spatial masking strategy, acting as intrinsic reasoning prior, is embedded into unfolding architecture. Meanwhile, the pre-trained MAE with spatial-spectral masking strategy is treated as the regularization term within loss function to constrain the spatial-spectral consistency. Such designs penetrate the image reasoning prior into deep unfolding networks while improving its interpretability and representation capability. The uniqueness of our framework is that the holistic learning process is explicitly integrated with the inherent physical mechanism underlying the pan-sharpening task. Extensive experiments on multiple satellite datasets demonstrate the superiority of our method over the existing state-of-the-art approaches. Code will be released at https://manman1995.github.io/. Man Zhou 0003, Jie Huang 0017, Naishan Zheng, Chongyi Li |
ICCV | 1 |
| 2023 | Exploring Temporal Frequency Spectrum in Deep Video DeblurringabstractVideo deblurring aims to restore the latent video frames from their blurred counterparts. Despite the remarkable progress, most promising video deblurring methods only investigate the temporal priors in the spatial domain and rarely explore their its potential in the frequency domain. In this paper, we revisit the blurred sequence in the Fourier space and figure out some intrinsic frequency-temporal priors that imply the temporal blur degradation can be accessibly decoupled in the potential frequency domain. Based on these priors, we propose a novel Fourier-based frequency-temporal video deblurring solution, where the core design accommodates the temporal spectrum to a popular video deblurring pipeline of feature extraction, alignment, aggregation, and optimization. Specifically, we design a Spectrum Prior-guided Alignment module by leveraging enlarged blur information in the potential spectrum to mitigate the blur effects on the alignment. Then, Temporal Energy prior-driven Aggregation is implemented to replenish the original local features by estimating the temporal spectrum energy as the global sharpness guidance. In addition, the customized frequency loss is devised to optimize the proposed method for decent spectral distribution. Extensive experiments demonstrate that our model performs favorably against other state-of-the-art methods, thus confirming the effectiveness of frequency-temporal prior modeling. Qi Zhu 0010, Man Zhou 0003, Naishan Zheng, Chongyi Li, Jie Huang 0017, Feng Zhao 0004 |
ICCV | 2 |
| 2023 | Embedding Fourier for Ultra-High-Definition Low-Light Image Enhancement
Chongyi Li, Chunle Guo, Man Zhou 0003, Zhexin Liang, Shangchen Zhou, Ruicheng Feng, Chen Change Loy |
ICLR | 3 |
| 2023 | Random Shuffle Transformer for Image RestorationabstractNon-local interactions play a vital role in boosting performance for image restoration. However, local window Transformer has been preferred due to its efficiency for processing high-resolution images. The superiority in efficiency comes at the cost of sacrificing the ability to model non-local interactions. In this paper, we present that local window Transformer can also function as modeling non-local interactions. The counterintuitive function is based on the permutation-equivariance of self-attention. The basic principle is quite simple: by *randomly shuffling* the input, local self-attention also has the potential to model non-local interactions without introducing extra parameters. Our random shuffle strategy enjoys elegant theoretical guarantees in extending the local scope. The resulting Transformer dubbed *ShuffleFormer* is capable of processing high-resolution images efficiently while modeling non-local interactions. Extensive experiments demonstrate the effectiveness of ShuffleFormer across a variety of image restoration tasks, including image denoising, deraining, and deblurring. Code is available at https://github.com/jiexiaou/ShuffleFormer. Jie Xiao 0002, Xueyang Fu, Man Zhou 0003, Hongjian Liu, Zhengjun Zha |
ICML | 3 |
| 2023 | Fourmer: An Efficient Global Modeling Paradigm for Image RestorationabstractGlobal modeling-based image restoration frameworks have become popular. However, they often require a high memory footprint and do not consider task-specific degradation. Our work presents an alternative approach to global modeling that is more efficient for image restoration. The key insights which motivate our study are two-fold: 1) Fourier transform is capable of disentangling image degradation and content component to a certain extent, serving as the image degradation prior, and 2) Fourier domain innately embraces global properties, where each pixel in the Fourier space is involved with all spatial pixels. While adhering to the ``spatial interaction + channel evolution'' rule of previous studies, we customize the core designs with Fourier spatial interaction modeling and Fourier channel evolution. Our paradigm, Fourmer, achieves competitive performance on common image restoration tasks such as image de-raining, image enhancement, image dehazing, and guided image super-resolution, while requiring fewer computational resources. The code for Fourmer will be made publicly available. Man Zhou 0003, Jie Huang 0017, Chunle Guo, Chongyi Li |
ICML | 1 |
| 2023 | Learning Non-Uniform-Sampling for Ultra-High-Definition Image EnhancementabstractUltra-high-definition (UHD) image enhancement is a challenging problem that aims to effectively and efficiently recover clean UHD images. To maintain efficiency, the straightforward approach is to downsample and perform most computations on low-resolution images. However, previous studies typically rely on the uniform and content-agnostic downsampling method that equally treats various regions regardless of their complexities, thus limiting the detail reconstruction in UHD image enhancement. To alleviate this issue, we propose a novel spatial-variant and invertible non-uniform downsampler that adaptively adjusts the sampling rate according to the richness of details. It magnifies important regions to preserve more information (e.g., sparse sampling points for sky, dense sampling points for buildings). Therefore, we propose a novel Non-uniform-Sampling Enhancement Network (NSEN) consisting of two core designs: 1) content-guided downsampling that extracts texture representation to guide the sampler to perform content-aware downsampling for producing detail-preserved low-resolution images; 2) invertible pixel-alignment which remaps the forward sampling process in an iterative manner to eliminate the deformations caused by the non-uniform downsampling, thus producing detail-rich clean UHD images. To demonstrate the superiority of our proposed model, we conduct extensive experiments on various UHD enhancement tasks. The results show that the proposed NSEN yields better performance against other state-of-the-art methods both visually and quantitatively. Qi Zhu 0010, Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004 |
ACM Multimedia | 5 |
| 2023 | Transition-constant Normalization for Image EnhancementabstractNormalization techniques that capture image style by statistical representation have become a popular component in deep neural networks.
Although image enhancement can be considered as a form of style transformation, there has been little exploration of how normalization affect the enhancement performance.
To fully leverage the potential of normalization, we present a novel Transition-Constant Normalization (TCN) for various image enhancement tasks.
Specifically, it consists of two streams of normalization operations arranged under an invertible constraint, along with a feature sub-sampling operation that satisfies the normalization constraint.
TCN enjoys several merits, including being parameter-free, plug-and-play, and incurring no additional computational costs.
We provide various formats to utilize TCN for image enhancement, including seamless integration with enhancement networks, incorporation into encoder-decoder architectures for downsampling, and implementation of efficient architectures.
Through extensive experiments on multiple image enhancement tasks, like low-light enhancement, exposure correction, SDR2HDR translation, and image dehazing, our TCN consistently demonstrates performance improvements.
Besides, it showcases extensive ability in other tasks including pan-sharpening and medical segmentation.
The code is available at \textit{\textcolor{blue}{https://github.com/huangkevinj/TCNorm}}. Jie Huang 0017, Man Zhou 0003, Mingde Yao, Chongyi Li, Zhiwei Xiong, Feng Zhao 0004 |
NeurIPS | 2 |
| 2023 | Deep Fractional Fourier TransformabstractExisting deep learning-based computer vision methods usually operate in the spatial and frequency domains, which are two orthogonal \textbf{individual} perspectives for image processing.
In this paper, we introduce a new spatial-frequency analysis tool, Fractional Fourier Transform (FRFT), to provide comprehensive \textbf{unified} spatial-frequency perspectives.
The FRFT is a unified continuous spatial-frequency transform that simultaneously reflects an image's spatial and frequency representations, making it optimal for processing non-stationary image signals.
We explore the properties of the FRFT for image processing and present a fast implementation of the 2D FRFT, which facilitates its widespread use.
Based on these explorations, we introduce a simple yet effective operator, Multi-order FRactional Fourier Convolution (MFRFC), which exhibits the remarkable merits of processing images from more perspectives in the spatial-frequency plane. Our proposed MFRFC is a general and basic operator that can be easily integrated into various tasks for performance improvement.
We experimentally evaluate the MFRFC on various computer vision tasks, including object detection, image classification, guided super-resolution, denoising, dehazing, deraining, and low-light enhancement. Our proposed MFRFC consistently outperforms baseline methods by significant margins across all tasks. Hu Yu 0001, Jie Huang 0017, Lingzhi Li 0002, Man Zhou 0003, Feng Zhao 0004 |
NeurIPS | 4 |
| 2023 | Rubik's Cube: High-Order Channel Interactions with a Hierarchical Receptive FieldabstractImage restoration techniques, spanning from the convolution to the transformer paradigm, have demonstrated robust spatial representation capabilities to deliver high-quality performance.Yet, many of these methods, such as convolution and the Feed Forward Network (FFN) structure of transformers, primarily leverage the basic first-order channel interactions and have not maximized the potential benefits of higher-order modeling. To address this limitation, our research dives into understanding relationships within the channel dimension and introduces a simple yet efficient, high-order channel-wise operator tailored for image restoration. Instead of merely mimicking high-order spatial interaction, our approach offers several added benefits: Efficiency: It adheres to the zero-FLOP and zero-parameter principle, using a spatial-shifting mechanism across channel-wise groups. Simplicity: It turns the favorable channel interaction and aggregation capabilities into element-wise multiplications and convolution units with $1 \times 1$ kernel. Our new formulation expands the first-order channel-wise interactions seen in previous works to arbitrary high orders, generating a hierarchical receptive field akin to a Rubik's cube through the combined action of shifting and interactions. Furthermore, our proposed Rubik's cube convolution is a flexible operator that can be incorporated into existing image restoration networks, serving as a drop-in replacement for the standard convolution unit with fewer parameters overhead. We conducted experiments across various low-level vision tasks, including image denoising, low-light image enhancement, guided image super-resolution, and image de-blurring. The results consistently demonstrate that our Rubik's cube operator enhances performance across all tasks. Code is publicly available at https://github.com/zheng980629/RubikCube. Naishan Zheng, Man Zhou 0003, Chong Zhou, Chen Change Loy |
NeurIPS | 2 |
| 2023 | Training Your Image Restoration Network Better with Random Weight Network as Optimization FunctionabstractThe blooming progress made in deep learning-based image restoration has been largely attributed to the availability of high-quality, large-scale datasets and advanced network structures. However, optimization functions such as L_1 and L_2 are still de facto. In this study, we propose to investigate new optimization functions to improve image restoration performance. Our key insight is that ``random weight network can be acted as a constraint for training better image restoration networks''. However, not all random weight networks are suitable as constraints. We draw inspiration from Functional theory and show that alternative random weight networks should be represented in the form of a strict mathematical manifold. We explore the potential of our random weight network prototypes that satisfy this requirement: Taylor's unfolding network, invertible neural network, central difference convolution, and zero-order filtering. We investigate these prototypes from four aspects: 1) random weight strategies, 2) network architectures, 3) network depths, and 4) combinations of random weight networks. Furthermore, we devise the random weight in two variants: the weights are randomly initialized only once during the entire training procedure, and the weights are randomly initialized in each training epoch. Our approach can be directly integrated into existing networks without incurring additional training and testing computational costs. We perform extensive experiments across multiple image restoration tasks, including image denoising, low-light image enhancement, and guided image super-resolution to demonstrate the consistent performance gains achieved by our method. Upon acceptance of this paper, we will release the code. Man Zhou 0003, Naishan Zheng, Chunle Guo, Chongyi Li |
NeurIPS | 1 |
| 2023 | FouriDown: Factoring Down-Sampling into Shuffling and SuperposingabstractSpatial down-sampling techniques, such as strided convolution, Gaussian, and Nearest down-sampling, are essential in deep neural networks. In this study, we revisit the working mechanism of the spatial down-sampling family and analyze the biased effects caused by the static weighting strategy employed in previous approaches. To overcome this limitation, we propose a novel down-sampling paradigm in the Fourier domain, abbreviated as FouriDown, which unifies existing down-sampling techniques. Drawing inspiration from the signal sampling theorem, we parameterize the non-parameter static weighting down-sampling operator as a learnable and context-adaptive operator within a unified Fourier function. Specifically, we organize the corresponding frequency positions of the 2D plane in a physically-closed manner within a single channel dimension. We then perform point-wise channel shuffling based on an indicator that determines whether a channel's signal frequency bin is susceptible to aliasing, ensuring the consistency of the weighting parameter learning. FouriDown, as a generic operator, comprises four key components: 2D discrete Fourier transform, context shuffling rules, Fourier weighting-adaptively superposing rules, and 2D inverse Fourier transform. These components can be easily integrated into existing image restoration networks. To demonstrate the efficacy of FouriDown, we conduct extensive experiments on image de-blurring and low-light image enhancement. The results consistently show that FouriDown can provide significant performance improvements. We will make the code publicly available to facilitate further exploration and application of FouriDown. Qi Zhu 0010, Man Zhou 0003, Jie Huang 0017, Naishan Zheng, Hongzhi Gao, Chongyi Li, Feng Zhao 0004 |
NeurIPS | 2 |
| 2023 | Memory-Augmented Deep Unfolding Network for Guided Image Super-resolution
Man Zhou 0003, Jinshan Pan, Wenqi Ren, Qi Xie 0002, Xiangyong Cao |
Int. J. Comput. Vis. | 1 |
| 2023 | Multiscale Dual-Domain Guidance Network for Pan-SharpeningabstractThe goal of pan-sharpening is to produce a high-spatial-resolution multi-spectral (HRMS) image from a low-spatial-resolution multi-spectral (LRMS) counterpart by super-resolving the LRMS one under the guidance of a texture-rich panchromatic (PAN) image. Existing research has concentrated on using spatial information to generate HRMS images, but has neglected to investigate the frequency domain, which severely restricts the performance improvement. In this work, we propose a novel pan-sharpening approach, named Multi-Scale Dual-Domain Guidance Network (MSDDN) by fully exploring and exploiting the distinguished information in both the spatial and frequency domains. Specifically, the network is inborn with multi-scale U-shape manner and composed by two core parts: a spatial guidance sub-network for fusing local spatial information and a frequency guidance sub-network for fusing global frequency domain information and encouraging dual-domain complementary learning. In this way, the model can capture multi-scale dual-domain information to help it generate high-quality pan-sharpening results. Employing the proposed model on different datasets, the quantitative and qualitative results demonstrate that our method performs appreciatively against other state-of-the-art approaches and comprises a strong generalization ability for real-world scenes. The source code is available at https://github.com/alexhe101/MSDDN. Xuanhua He, Jie Zhang 0033, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Learning Deep Multiscale Local Dissimilarity Prior for PansharpeningabstractVarious deep neural networks (DNNs) have been constructed to inject the spatial information of the panchromatic (PAN) image into the low spatial resolution multispectral (LR MS) image. However, most of them ignore the local dissimilarity (LD) prior between MS and PAN images, which has a negative influence on the fused image. Considering the above-mentioned issues, we propose a deep multiscale local dissimilarity network (DMLD-Net) to learn the LD prior at different scales and enhance the spatial and spectral information in the fused image better. Specifically, we first synthesize a downsampled PAN image from the original PAN image to match the scale of the LR MS image. Then, a LD metric is designed to calculate the dissimilarity map between the two images in feature space. According to the learned dissimilarity map, we utilize a LD-guided attention block (LDGAB) to suppress the impact of LD, which filters out the dissimilar information in the features of the PAN image. To learn the LD prior between MS and PAN images sufficiently, the multiscale architecture is considered and we infer the dissimilar maps hierarchically and inject filtered features into the LR MS image progressively. Finally, the fused image is generated by a reconstruction block. Through the LD learning at different scales, reasonable spatial information is extracted from the PAN image, by which the distortions in the fused image caused by LD can be reduced efficiently. Extensive experiments are conducted on GeoEye-1 and WorldView-2 datasets and the results demonstrate the effectiveness of the proposed DMLD-Net in terms of spatial and spectral preservation. The code is available at https://github.com/RSMagneto/DMLD-Net. Kai Zhang 0010, Guishuo Yang, Feng Zhang 0028, Wenbo Wan, Man Zhou 0003, Jiande Sun 0001, Huaxiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Deep Adaptive Pansharpening via Uncertainty-Aware Image FusionabstractPansharpening is a procedure that fuses high-resolution panchromatic (PAN) images and low-resolution multispectral (LMS) images to derive high-resolution multispectral (HMS) images. Despite its rapid development, most existing pansharpening techniques integrate the information of PAN and LMS invariantly in the spatial dimension, ignoring the uneven spatial dependence of restoring HMS with the aid of PAN information and resulting in ineffective fusion results. In this work, we propose an Uncertainty-aware Adaptive Pansharpening Network (UAPN) that integrates PAN information spatial-variantly to restore LMS information with an uncertainty mechanism. Specifically, we first estimate the epistemic and aleatoric uncertainties together, which model the spatial-variant distributions of restoring the LMS image to the HMS image. Then, we introduce Uncertainty-conditioned Adaptive Convolution (UAC) to adaptively integrate LMS and PAN information, where its parameters are spatially variable by conditioning on the uncertainty estimations. Furthermore, we propose a multi-stage uncertainty-driven loss function to explicitly force the network to concentrate on restoring challenging areas of the LMS image. Extensive experimental results demonstrate the superiority of our UAPN with fewer parameters and flops, outperforming other state-of-the-art methods both qualitatively and quantitatively on multiple satellite datasets. The code is available at https://github.com/keviner1/UAPN.. Jie Huang 0017, Man Zhou 0003, Danfeng Hong, Feng Zhao 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Modality-Aware Feature Integration for Pan-SharpeningabstractPan-sharpening aims to super-solve low-spatial resolution multiple spectral (MS) images with the guidance of high-resolution (HR) texture-rich panchromatic (PAN) images. Recently, deep-learning-based pan-sharpening approaches have dominated this field and achieved remarkable advancement. However, most promising algorithms are devised in one-way mapping and have not fully explored the mutual dependencies between PAN and MS modalities, thus impacting the model performance. To address this issue, we propose a novel information compensation and integration network for pan-sharpening by effective cross-modality joint learning in this work. First, the cross-central difference convolution is employed to explicitly extract the texture details of the PAN images. Second, we implement the compensation process by imitating the classical back-projection (BP) technique where the extracted PAN textures are employed to guide the intrinsic information learning of MS images iteratively. Subsequently, we devise the hierarchical transformer to integrate the comprehensive relations of stage-iteration information from spatial and temporal contexts. Extensive experiments over multiple satellite datasets demonstrate the superiority of our method to the existing state-of-the-art methods. The source code is available athttps://github.com/manman1995/pansharpening. Man Zhou 0003, Jie Huang 0017, Feng Zhao 0004, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Exposure Normalization and Compensation for Multiple-Exposure CorrectionabstractImages captured with improper exposures usually bring unsatisfactory visual effects. Previous works mainly focus on either underexposure or overexposure correction, resulting in poor generalization to various exposures. An alternative solution is to mix the multiple exposure data for training a single network. However, the procedures of correcting underexposure and overexposure to normal exposures are much different from each other, leading to large discrepancies for the network in correcting multiple-exposures, thus resulting in poor performance. The key point to address this issue lies in bridging different exposure representations. To achieve this goal, we design a multiple exposure correction framework based on an Exposure Normalization and Compensation (ENC) module. Specifically, the ENC module consists of an exposure normalization part for mapping different exposure features to the exposure-invariant feature space, and a compensation part for integrating the initial features unprocessed by the exposure normalization part to ensure the completeness of information. Besides, to further alleviate the imbalanced performance caused by variations in the optimization process, we introduce a parameter regularization fine-tuning strategy to improve the performance of the worst-performed exposure without degrading other exposures. Our model empowered by ENC outperforms the existing methods by more than 2dB and is robust to multiple image enhancement tasks, demonstrating its effectiveness and generalization capability for real-world applications. Code: https://github.com/KevinJ-Huang/ExposureNorm-Compensation. Jie Huang 0017, Xueyang Fu, Man Zhou 0003, Yang Wang 0015, Feng Zhao 0004, Zhiwei Xiong |
CVPR | 4 |
| 2022 | Memory-augmented Deep Conditional Unfolding Network for PansharpeningabstractPansharpening aims to obtain high-resolution multispectral (MS) images for remote sensing systems and deep learning-based methods have achieved remarkable success. However, most existing methods are designed in a black-box principle, lacking sufficient interpretability. Additionally, they ignore the different characteristics of each band of MS images and directly concatenate them with panchromatic (PAN) images, leading to severe copy artifacts [9]. To address the above issues, we propose an interpretable deep neural network, namely Memory-augmented Deep Conditional Unfolding Network with two specified core designs. Firstly, considering the degradation process, it formulates the Pansharpening problem as the minimization of a variational model with denoising-based prior and non-local auto-regression prior which is capable of searching the similarities between long-range patches, benefiting the texture enhancement. A novel iteration algorithm with built-in CNNs is exploited for transparent model design. Secondly, to fully explore the potentials of different bands of MS images, the PAN image is combined with each band of MS images, selectively providing the high-frequency details and alleviating the copy artifacts. Extensive experimental results validate the superiority of the proposed algorithm against other state-of-the-art methods. Man Zhou 0003, Aiping Liu, Xueyang Fu, Fan Wang 0005 |
CVPR | 2 |
| 2022 | Mutual Information-driven Pan-sharpeningabstractPan-sharpening aims to integrate the complementary information of texture-rich PAN images and multi-spectral (MS) images to produce the texture-rich MS images. Despite the remarkable progress, existing state-of-the-art Pansharpening methods don't explicitly enforce the complementary information learning between two modalities of PAN and MS images. This leads to information redundancy not being handled well, which further limits the performance of these methods. To address the above issue, we propose a novel mutual information-driven Pan-sharpening framework in this paper. To be specific, we first project the PAN and MS image into modality-aware feature space independently, and then impose the mutual information minimization over them to explicitly encourage the complementary information learning. Such operation is capable of reducing the information redundancy and improving the model performance. Extensive experimental results over multiple satellite datasets demonstrate that the proposed algorithm outperforms other state-of-the-art methods qualitatively and quantitatively with great generalization ability to real-world scenes. Man Zhou 0003, Jie Huang 0017, Zihe Yang, Xueyang Fu, Feng Zhao 0004 |
CVPR | 1 |
| 2022 | Deep Fourier-Based Exposure Correction Network with Spatial-Frequency Interaction
Jie Huang 0017, Feng Zhao 0004, Man Zhou 0003, Zhiwei Xiong |
ECCV (19) | 7 |
| 2022 | Memory-Augmented Model-Driven Network for Pansharpening
Man Zhou 0003, Li Zhang 0104, Chengjun Xie |
ECCV (19) | 2 |
| 2022 | Frequency and Spatial Dual Guidance for Image Dehazing
Hu Yu 0001, Naishan Zheng, Man Zhou 0003, Jie Huang 0017, Zeyu Xiao 0002, Feng Zhao 0004 |
ECCV (19) | 3 |
| 2022 | Spatial-Frequency Domain Information Integration for Pan-Sharpening
Man Zhou 0003, Jie Huang 0017, Hu Yu 0001, Xueyang Fu, Aiping Liu, Xian Wei, Feng Zhao 0004 |
ECCV (18) | 1 |
| 2022 | Exploring Fourier Prior for Single Image Rain RemovalabstractDeep convolutional neural networks (CNNs) have become dominant in the task of single image rain removal. Most of current CNN methods, however, suffer from the problem of overfitting on one single synthetic dataset as they neglect the intrinsic prior of the physical properties of rain streaks. To address this issue, we propose a simple but effective prior - Fourier prior to improve the generalization ability of an image rain removal model. The Fourier prior is a kind of property of rainy images. It is based on a key observation of us - replacing the Fourier amplitude of rainy images with that of clean images greatly suppresses the synthetic and real-world rain streaks. This means the amplitude contains most of the rain streak information and the phase keeps the similar structures of the background. So it is natural for single image rain removal to process the amplitude and phase information of the rainy images separately. In this paper, we develop a two-stage model where the first stage restores the amplitude of rainy images to clean rain streaks, and the second stage restores the phase information to refine fine-grained background structures. Extensive experiments on synthetic rainy data demonstrate the power of Fourier prior. Moreover, when trained on synthetic data, a robust generalization ability to real-world images can also be obtained. The code will be publicly available at https://github.com/willinglucky/ExploringFourier-Prior-for-Single-Image-Rain-Removal. Xin Guo 0018, Xueyang Fu, Man Zhou 0003, Zhen Huang 0007, Jialun Peng, Zhengjun Zha |
IJCAI | 3 |
| 2022 | Exposure-Consistency Representation Learning for Exposure CorrectionabstractImages captured under improper exposures including underexposure and overexposure often suffer from unsatisfactory visual effects. Since their correction procedures are quite different, it is challenging for a single network to correct various exposures. The key to addressing this issue is consistently learning underexposure and overexposure corrections. To achieve this goal, we propose an Exposure-Consistency Processing (ECP) module to consistently learn the representation of both underexposure and overexposure in the feature space. Specifically, the ECP module employs the bilateral activation mechanism that derives both underexposure and overexposure property features for exposure-consistency representation modeling, which is followed by two shared-weight branches to process these features. Based on the ECP module, we build the whole network by utilizing it as the basic unit. Additionally, to further assist the exposure-consistency learning, we develop an Exposure-Consistency Constraining (ECC) strategy that augments the various local region exposures and then constrains the feature representation change between the exposure augmented image and the original one. Our proposed network is lightweight and outperforms existing methods remarkably, while the ECP module can also be extended to other baselines, demonstrating its superiority and scalability. code: https://github.com/KevinJ-Huang/ECLNet. Jie Huang 0017, Man Zhou 0003, Mingde Yao, Feng Zhao 0004, Zhiwei Xiong |
ACM Multimedia | 2 |
| 2022 | SIR-Former: Stereo Image Restoration Using TransformerabstractStereo image pairs record the scene from two different views and introduce cross-view information for image restoration. However, there are two challenges in utilizing the cross-view information for stereo image restoration: cross-view alignment and information fusion. Most existing methods adopt convolutional neural networks to align the views and fuse the information locally, which has difficulty in capturing the global correspondence across stereo images for view alignment and makes it hard to integrate the long-term information across views. In this paper, we propose to address the stereo image restoration with transformer by leveraging its powerful capability of modeling long-range context dependencies. Specifically, we construct a stereo image restoration transformer (SIR-Former) to effectively exploit the cross-view correlations. First, to explore the global correspondence for view alignment effectively, we devise a stereo alignment transformer (SAT) module across stereo images, enabling robust alignment under the epipolar constraint. Then, we design a stereo fusion transformer (SFT) module for aggregating the cross-view information in a small horizontal neighborhood, aiming to enhance important features for succeeding restoration. Extensive experiments show that SIR-Former can remarkably boost quantitative and qualitative quality on various image restoration tasks (e.g., super-resolution, deblurring, deraining, and low-light enhancement), which demonstrate the effectiveness of the proposed framework. Zizheng Yang, Mingde Yao, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004 |
ACM Multimedia | 4 |
| 2022 | Model-Guided Multi-Contrast Deep Unfolding Network for MRI Super-resolution ReconstructionabstractMagnetic resonance imaging (MRI) with high resolution (HR) provides more detailed information for accurate diagnosis and quantitative image analysis. Despite the significant advances, most existing super-resolution (SR) reconstruction network for medical images has two flaws: 1) All of them are designed in a black-box principle, thus lacking sufficient interpretability and further limiting their practical applications. Interpretable neural network models are of significant interest since they enhance the trustworthiness required in clinical practice when dealing with medical images. 2) most existing SR reconstruction approaches only use a single contrast or use a simple multi-contrast fusion mechanism, neglecting the complex relationships between different contrasts that are critical for SR improvement. To deal with these issues, in this paper, a novel Model-Guided interpretable Deep Unfolding Network (MGDUN) for medical image SR reconstruction is proposed. The Model-Guided image SR reconstruction approach solves manually designed objective functions to reconstruct HR MRI. We show how to unfold an iterative MGDUN algorithm into a novel model-guided deep unfolding network by taking the MRI observation matrix and explicit multi-contrast relationship matrix into account during the end-to-end optimization. Extensive experiments on the multi-contrast IXI dataset and BraTs 2019 dataset demonstrate the superiority of our proposed model. Li Zhang 0104, Man Zhou 0003, Aiping Liu, Xun Chen 0001, Zhiwei Xiong, Feng Wu 0001 |
ACM Multimedia | 3 |
| 2022 | Source-Free Domain Adaptation for Real-World Image DehazingabstractDeep learning-based source dehazing methods trained on synthetic datasets have achieved remarkable performance but suffer from dramatic performance degradation on real hazy images due to domain shift. Although certain Domain Adaptation (DA) dehazing methods have been presented, they inevitably require access to the source dataset to reduce the gap between the source synthetic and target real domains. To address these issues, we present a novel Source-Free Unsupervised Domain Adaptation (SFUDA) image dehazing paradigm, in which only a well-trained source model and an unlabeled target real hazy dataset are available. Specifically, we devise the Domain Representation Normalization (DRN) module to make the representation of real hazy domain features match that of the synthetic domain to bridge the gaps. With our plug-and-play DRN module, unlabeled real hazy images can adapt existing well-trained source networks. Besides, the unsupervised losses are applied to guide the learning of the DRN module, which consists of frequency losses and physical prior losses. Frequency losses provide structure and style constraints, while the prior loss explores the inherent statistic property of haze-free images. Equipped with our DRN module and unsupervised loss, existing source dehazing models are able to dehaze unlabeled real hazy images. Extensive experiments on multiple baselines demonstrate the validity and superiority of our method visually and quantitatively. Hu Yu 0001, Jie Huang 0017, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004 |
ACM Multimedia | 5 |
| 2022 | Structure- and Texture-Aware Learning for Low-Light Image EnhancementabstractStructure and texture information is critically important for low-light image enhancement, in terms of stable global adjustment and fine details recovery. However, most existing methods tend to learn the structure and texture of low-light images in a coupled manner, without well considering the heterogeneity between them, which challenges the capability of the model to learn both adequately. In this paper, we tackle this problem in a divide and conquer strategy, based on the observation that the structure and texture representations are highly separated in the frequency spectrum. Specifically, we propose a Structure and Texture Aware Network (STAN) for low-light image enhancement, which consists of a structure sub-network and a texture sub-network. The former exploits the low-pass characteristic of the transformer to capture low-frequency-related structural representation. While the latter builds upon central difference convolution to capture high-frequency-related texture representation. We establish the Multi-Spectrum Interaction (MSI) module between two sub-networks to bidirectionally provide complementary information. In addition, to further elevate the capability of the model, we introduce a dual distillation scheme that assists the learning process of two sub-networks via counterparts' normal-light structure and texture representations. Comprehensive experiments show that the proposed STAN outperforms the state-of-the-art methods qualitatively and quantitatively. Jie Huang 0017, Mingde Yao, Man Zhou 0003, Feng Zhao 0004 |
ACM Multimedia | 4 |
| 2022 | Enhancement by Your Aesthetic: An Intelligible Unsupervised Personalized Enhancer for Low-Light ImagesabstractLow-light image enhancement is an inherently subjective process whose targets vary with the user's aesthetic. Motivated by this, several personalized enhancement methods have been investigated. However, the enhancement process based on user preferences in these techniques is invisible, i.e., a "black box". In this work, we propose an intelligible unsupervised personalized enhancer (iUP-Enhancer) for low-light images, which establishes the correlations between the low-light and the unpaired reference images with regard to three user-friendly attributions (brightness, chromaticity, and noise). The proposed iUP-Enhancer is trained with the guidance of these correlations and the corresponding unsupervised loss functions. Rather than a "black box" process, our iUP-Enhancer presents an intelligible enhancement process with the above attributions. Extensive experiments demonstrate that the proposed algorithm produces competitive qualitative and quantitative results while maintaining excellent flexibility and scalability. This can be validated by personalization with single/multiple references, cross-attribution references, or merely adjusting parameters. Naishan Zheng, Jie Huang 0017, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004, Zhengjun Zha |
ACM Multimedia | 4 |
| 2022 | Adaptively Learning Low-high Frequency Information Integration for Pan-sharpeningabstractPan-sharpening aims to generate high-spatial resolution multi-spectral (MS) image by fusing high-spatial resolution panchromatic (PAN) image and its corresponding low-spatial resolution MS image. Despite the remarkable progress, most existing pan-sharpening methods only work in the spatial domain and rarely explore the potential solutions in the frequency domain. In this paper, we propose a novel pan-sharpening framework by adaptively learning low-high frequency information integration in the spatial and frequency dual domains. It consists of three key designs: mask prediction sub-network, low-frequency learning sub-network and high-frequency learning sub-network. Specifically, the first is responsible for measuring the modality-aware frequency information difference of PAN and MS images and further predicting the low-high frequency boundary in the form of a two-dimensional mask. In view of the mask, the second adaptively picks out the corresponding low-frequency components of different modalities and then restores the expected low-frequency one by spatial and frequency dual domains information integration while the third combines the above refined low-frequency and the original high-frequency for the latent high-frequency reconstruction. In this way, the low-high frequency information is adaptively learned, thus leading to the pleasing results. Extensive experiments validate the effectiveness of the proposed network and demonstrate the favorable performance against other state-of-the-art methods. The source code will be released at https://github.com/manman1995/pansharpening. Man Zhou 0003, Jie Huang 0017, Chongyi Li, Hu Yu 0001, Naishan Zheng, Feng Zhao 0004 |
ACM Multimedia | 1 |
| 2022 | Normalization-based Feature Selection and Restitution for Pan-sharpeningabstractPan-sharpening is essentially a panchromatic (PAN) image-guided low-spatial resolution MS image super-resolution problem. The commonly challenging issue of pan-sharpening is how to correctly select consistent features and propagate them, and properly handle inconsistent ones between PAN and MS modalities. To solve this issue, we propose a Normalization-based Feature Selection and Restitution mechanism, which is capable of filtering out the inconsistent features and promoting to learn the consistent ones. Specifically, we first modulate the PAN feature as the MS style in feature space by AdaIN operation \citeAdaIN. However, such operation inevitably removes the favorable features. We thus propose to distill the effective information from the removed part and restitute it back to the modulated part. To better distillation, we enforce a contrastive learning constraint to close the distance between the restituted feature and the ground truth, and push the removed part away from the ground truth. In this way, the consistent features of PAN images are correctly selected and the inconsistent ones are filtered out, thus relieving the over-transferred artifacts in the process of PAN-guided MS super-resolution. Extensive experiments validate the effectiveness of the proposed network and demonstrate its favorable performance against other state-of-the-art methods. The source code will be released at https://github.com/manman1995/pansharpening. Man Zhou 0003, Jie Huang 0017, Aiping Liu, Chongyi Li, Feng Zhao 0004 |
ACM Multimedia | 1 |
| 2022 | Panchromatic and Multispectral Image Fusion via Alternating Reverse Filtering NetworkabstractPanchromatic (PAN) and multi-spectral (MS) image fusion, named Pan-sharpening, refers to super-resolve the low-resolution (LR) multi-spectral (MS) images in the spatial domain to generate the expected high-resolution (HR) MS images, conditioning on the corresponding high-resolution PAN images. In this paper, we present a simple yet effective alternating reverse filtering network for pan-sharpening. Inspired by the classical reverse filtering that reverses images to the status before filtering, we formulate pan-sharpening as an alternately iterative reverse filtering process, which fuses LR MS and HR MS in an interpretable manner. Different from existing model-driven methods that require well-designed priors and degradation assumptions, the reverse filtering process avoids the dependency on pre-defined exact priors. To guarantee the stability and convergence of the iterative process via contraction mapping on a metric space, we develop the learnable multi-scale Gaussian kernel module, instead of using specific filters. We demonstrate the theoretical feasibility of such formulations. Extensive experiments on diverse scenes to thoroughly verify the performance of our method, significantly outperforming the state of the arts. Man Zhou 0003, Jie Huang 0017, Feng Zhao 0004, Chengjun Xie, Chongyi Li, Danfeng Hong |
NeurIPS | 2 |
| 2022 | Deep Fourier Up-SamplingabstractExisting convolutional neural networks widely adopt spatial down-/up-sampling for multi-scale modeling. However, spatial up-sampling operators (e.g., interpolation, transposed convolution, and un-pooling) heavily depend on local pixel attention, incapably exploring the global dependency. In contrast, the Fourier domain is in accordance with the nature of global modeling according to the spectral convolution theorem. Unlike the spatial domain that easily performs up-sampling with the property of local similarity, up-sampling in the Fourier domain is more challenging as it does not follow such a local property. In this study, we propose a theoretically feasible Deep Fourier Up-Sampling (FourierUp) to solve these issues. We revisit the relationships between spatial and Fourier domains and reveal the transform rules on the features of different resolutions in the Fourier domain, which provide key insights for FourierUp's designs. FourierUp as a generic operator consists of three key components: 2D discrete Fourier transform, Fourier dimension increase rules, and 2D inverse Fourier transform, which can be directly integrated with existing networks. Extensive experiments across multiple computer vision tasks, including object detection, image segmentation, image de-raining, image dehazing, and guided image super-resolution, demonstrate the consistent performance gains obtained by introducing our FourierUp. Code will be publicly available. Man Zhou 0003, Hu Yu 0001, Jie Huang 0017, Feng Zhao 0004, Jinwei Gu, Chen Change Loy, Deyu Meng, Chongyi Li |
NeurIPS | 1 |
| 2022 | When Pansharpening Meets Graph Convolution Network and Knowledge DistillationabstractIn this article, we propose a novel graph convolutional network (GCN) for pansharpening, defined as GCPNet, which consists of three main modules: the spatial GCN module (SGCN), the spectral band GCN module (BGCN), and the atrous spatial pyramid module (ASPM). Specifically, due to the nature of GCN, the proposed SGCN and BGCN are capable of exploring the long-range relationship between the object and the global state in the spatial and spectral aspects, which benefits pansharpened results and has not been fully investigated before. In addition, the designed ASPM is equipped with multiscale atrous convolutions and learns richer local feature information, so as to cover the objects of different sizes in satellite images. To further enhance the representation of our proposed GCPNet, asynchronous knowledge distillation is introduced to provide compact features by heterogeneous task imitation in a teacher–student paradigm. In the paradigm, the teacher network acts as a variational autoencoder to extract compact features of the ground-truth MS images. The student network, devised for pansharpening, is trained with the assistance of the teacher network to transfer the important information of the expected ground-truth MS images. Extensive experimental results on different satellite datasets demonstrate that our proposed network outperforms the state-of-the-art methods both visually and quantitatively. The source code is released athttps://github.com/Keyu-Yan/GCPNet. Man Zhou 0003, Liu Liu 0012, Chengjun Xie, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Effective Pan-Sharpening With Transformer and Invertible Neural NetworkabstractIn remote sensing imaging systems, pan-sharpening is an important technique to obtain high-resolution multispectral images from a high-resolution panchromatic image and its corresponding low-resolution multispectral image. Due to the powerful learning capability of convolution neural networks (CNNs), CNN-based methods have dominated this field. However, due to the limitation of the convolution operator, long-range spatial features are often not accurately obtained, thus limiting the overall performance. To this end, we propose a novel and effective method by exploiting a customized transformer architecture and information-lossless invertible neural module for long-range dependencies modeling and effective feature fusion in this article. Specifically, the customized transformer formulates the panchromatic (PAN) and multispectral (MS) features as queries and keys to encourage joint feature learning across two modalities, while the designed invertible neural module enables effective feature fusion to generate the expected pan-sharpened results. To the best of our knowledge, this is the first attempt to introduce a transformer and a invertible neural network into the pan-sharpening field. Extensive experiments over different kinds of satellite datasets demonstrate that our method outperforms state-of-the-art algorithms both visually and quantitatively with fewer parameters and flops. Furthermore, the ablation experiments also prove the effectiveness of the proposed customized long-range transformer and effective invertible neural feature fusion module for pan-sharpening. Man Zhou 0003, Xueyang Fu, Jie Huang 0017, Feng Zhao 0004, Aiping Liu, Rujing Wang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Effective Pan-Sharpening by Multiscale Invertible Neural Network and Heterogeneous Task DistillingabstractAs recognized, the ground truth multi-spectral (MS) images possess the complementary information (e.g., high-frequency component) of low-resolution (LR) MS images, which can be considered as privileged information to alleviate the spectral distortion and insufficient spatial texture enhancement. Since existing supervised pan-sharpening methods only utilize the ground truth MS image to supervise the network training, its potential value has not been fully explored. To accomplish this, we propose a heterogeneous knowledge-distilling pan-sharpening framework that distills pan-sharpening by imitating the ground truth reconstruction task in both the feature space and network output. In our work, the teacher network performs as a variational auto-encoder to extract effective features of the ground truth MS. The student network, acting as pan-sharpening, is trained by the assistance of the teacher network with the process-oriented feature imitation learning. Moreover, we design a customized information-lossless multi-scale invertible neural module to effectively fuse LR-MS and panchromatic (PAN) images, producing expected pan-sharpened results. To reduce the artifacts generated by the knowledge distillation process, a knowledge-driven refinement sub-network is further devised according to the pan-sharpening imaging model. Extensive experimental results on different satellite datasets validate that the proposed network outperforms the state-of-the-art methods both visually and quantitatively. The source code will be released at https://github.com/manman1995/pansharpening. Man Zhou 0003, Jie Huang 0017, Xueyang Fu, Feng Zhao 0004, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Image De-Raining via Continual LearningabstractWhile deep convolutional neural networks (CNNs) have achieved great success on image de-raining task, most existing methods can only learn fixed mapping rules between paired rainy/clean images on a single dataset. This limits their applications in practical situations with multiple and incremental datasets where the mapping rules may change for different types of rain streaks. However, the catastrophic forgetting of traditional deep CNN model challenges the design of generalized framework for multiple and incremental datasets. A strategy of sharing the network structure but in-dependently updating and storing the network parameters on each dataset has been developed as a potential solution. Nevertheless, this strategy is not applicable to compact systems as it dramatically increases the overall training time and parameter space. To alleviate such limitation, in this study, we propose a parameter importance guided weights modification approach, named PIGWM. Specifically, with new dataset (e.g. new rain dataset), the well-trained network weights are updated according to their importance evaluated on previous training dataset. With extensive experimental validation, we demonstrate that a single network with a single parameter set of our proposed method can process multiple rain datasets almost without performance degradation. The proposed model is capable of achieving superior performance on both inhomogeneous and incremental datasets, and is promising for highly compact systems to gradually learn myriad regularities of the different types of rain streaks. The results indicate that our proposed method has great potential for other computer vision tasks with dynamic learning environments. Man Zhou 0003, Jie Xiao 0002, Yifan Chang, Xueyang Fu, Aiping Liu, Jinshan Pan, Zhengjun Zha |
CVPR | 1 |
| 2021 | Improving De-raining Generalization via Neural ReorganizationabstractMost existing image de-raining networks could only learn fixed mapping rules between paired rainy/clean images on single synthetic dataset and then stay static for lifetime. However, since single synthetic dataset merely provides a partial view for the distribution of rain streaks, deep models well trained on an individual synthetic dataset tend to overfit on this biased distribution. This leads to the inability of these methods to well generalize to complex and changeable real-world rainy scenes, thus limiting their practical applications. In this paper, we try for the first time to accumulate the de-raining knowledge from multiple synthetic datasets on a single network parameter set to improve the de-raining generalization of deep networks. To achieve this goal, we explore Neural Reorganization (NR) to allow the de-raining network to keep a subtle stability-plasticity trade-off rather than naive stabilization after training phase. Specifically, we design our NR algorithm by borrowing the synaptic consolidation mechanism in the biological brain and knowledge distillation. Equipped with our NR algorithm, the deep model can be trained on a list of synthetic rainy datasets by overcoming catastrophic forgetting, making it a general-version de-raining network. Extensive experimental validation shows that due to the successful accumulation of de-raining knowledge, our proposed method can not only process multiple synthetic datasets consistently, but also achieve state-of-the-art results when dealing with real-world rainy images. Jie Xiao 0002, Man Zhou 0003, Xueyang Fu, Aiping Liu, Zhengjun Zha |
ICCV | 2 |
| 2021 | Reinforcedet: Object Detection By Integrating Reinforcement Learning With Decoupled PipelineabstractRecent object detection methods largely rely on numerous pre-defined anchors that suffer from huge computational cost and resource consumption. To solve this issue, we propose a low-memory deep reinforcement learning based anchor-free object detection approach, namely ReinforceDet, which computes few but accurate region proposals for detection. Specifically, the extracted feature maps are fed into a reinforcement learning network to localize objects as initial region proposals with our re-designed reward function and then adopt another neural network to refine them. To speed up this process in test phase, we decouple the two-branch CNN networks as light-head cascaded subnetworks, named IoU-net and bounding box net. Experimental results show that ReinforceDet could obtain the state-of-the-art performance with much lower compitational and memory cost. Man Zhou 0003, Liu Liu 0012, Rujing Wang |
ICIP | 1 |
| 2021 | Unfolding Taylor's Approximations for Image RestorationabstractDeep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing methods empirically construct encapsulated end-to-end mapping networks without deepening into the rationality, and neglect the intrinsic prior knowledge of restoration task. To solve the above problems, inspired by Taylor’s Approximations, we unfold Taylor’s Formula to construct a novel framework for image restoration. We find the main part and the derivative part of Taylor’s Approximations take the same effect as the two competing goals of high-level contextualized information and spatial details of image restoration respectively. Specifically, our framework consists of two steps, which are correspondingly responsible for the mapping and derivative functions. The former first learns the high-level contextualized information and the later combines it with the degraded input to progressively recover local high-order spatial details. Our proposed framework is orthogonal to existing methods and thus can be easily integrated with them for further improvement, and extensive experiments demonstrate the effectiveness and scalability of our proposed framework. Man Zhou 0003, Xueyang Fu, Zeyu Xiao 0002, Aiping Liu, Zhiwei Xiong |
NeurIPS | 1 |
| 2021 | ReinforceNet: A reinforcement learning embedded object detection framework with region selection network
Man Zhou 0003, Rujing Wang, Chengjun Xie, Liu Liu 0012, Rui Li 0027, Fangyuan Wang 0001, Dengshan Li |
Neurocomputing | 1 |
| 2021 | Learning region-guided scale-aware feature selection for object detection
Liu Liu 0012, Rujing Wang, Chengjun Xie, Rui Li 0027, Fangyuan Wang 0001, Man Zhou 0003 |
Neural Comput. Appl. | 6 |