EDBT 2026 Demo / reviewers in the wild / expert
Xueyang Fu
dblp:136/9389
· DBLP profile ↗
127ranked-venue papers
20as first author
97since 2021 · last 2026
0000-0001-8036-4071ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 89 · 13 first-author · 62 since 2021Artificial intelligence and machine learning · 82 · 12 first-author · 73 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CompEvent: Complex-valued Event-RGB Fusion for Low-light Video Enhancement and DeblurringabstractLow-light video deblurring poses significant challenges in applications like nighttime surveillance and autonomous driving due to dim lighting and long exposures. While event cameras offer potential solutions with superior low-light sensitivity and high temporal resolution, existing fusion methods typically employ staged strategies, limiting their effectiveness against combined low-light and motion blur degradations. To overcome this, we propose CompEvent, a complex neural network framework enabling holistic full-process fusion of event data and RGB frames for enhanced joint restoration. CompEvent features two core components: 1) Complex Temporal Alignment GRU, which utilizes complex-valued convolutions and processes video and event streams iteratively via GRU to achieve temporal alignment and continuous fusion; and 2) Complex Space-Frequency Learning module, which performs unified complex-valued signal processing in both spatial and frequency domains, facilitating deep fusion through spatial structures and system-level characteristics. By leveraging the holistic representation capability of complex-valued neural networks, CompEvent achieves full-process spatiotemporal fusion, maximizes complementary learning between modalities, and significantly strengthens low-light video deblurring capability. Extensive experiments demonstrate that CompEvent outperforms SOTA methods in addressing this challenging task. Mingchen Zhong, Xin Lu 0008, Dong Liu 0002, Senyan Xu, Ruixuan Jiang, Xueyang Fu |
AAAI | 6 |
| 2026 | Learning Robust Event-Guided Representations for Person Re-Identification
Chengzhi Cao, Xueyang Fu, Senyan Xu, Chengjie Ge, Zhengjun Zha |
Int. J. Comput. Vis. | 2 |
| 2026 | Efficient Real-World Image Super-Resolution Via Adaptive Directional Gradient Convolution
Long Peng 0003, Zhanfeng Feng, Renjing Pei, Wenbo Li 0002, Jiaming Guo, Xueyang Fu, Yang Wang 0015, Yang Cao 0010, Zhengjun Zha |
Int. J. Comput. Vis. | 6 |
| 2026 | SkyFind: A Large-Scale Benchmark Unveiling Referring Expression Comprehension for UAVabstractUncrewed aerial vehicles (UAV) are increasingly deployed to assist humans in diverse tasks, where understanding human intentions is critical to effective collaboration. Referring expression comprehension (REC) links language to visual targets, allowing UAV to recognize human-intended targets of interest, thereby supporting subsequent actions. However, existing REC research is almost exclusively confined to ground-based scenarios, leaving aerial scenarios largely unexplored. In this paper, we formally define UAV-based REC as a new research problem and highlight its unique challenges, including abundant background interference, small target size, and complex referring relations. To enable systematic study, we introduce SkyFind, a large-scale dataset with one million high-quality target-expression pairs, providing a solid foundation. In addition, we propose AerialREC, a baseline framework that reduces background interference in UAV imagery by searching for a potential target region before localization. We establish benchmark results on SkyFind using ten representative REC methods and validate the effectiveness of the AerialREC framework. Guanbo Wu, Xueyang Fu, Kean Liu, Xin Lu 0008, Chengjie Ge, Wei Zhai, Zhengjun Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Bayesian Window Transformer for Image RestorationabstractTransformers have excelled in image restoration due to their advanced representational abilities. However, their reliance on a fixed local window for attention often undermines translation invariance and local relationship preservation. This limitation can reduce network stability, especially when dealing with positional changes in degradation scenarios. In this research, we present a new Bayesian Window Transformer, which innovates by employing a probability distribution for window shifts, overcoming the limitations of fixed window configurations in traditional transformers. This approach allows for more flexible coverage beyond a predetermined region. During the evaluation procedure, we further develop two approximate inference algorithms: Layer Expectation Propagation and Monte Carlo Average. These two algorithms calculate expectations derived from the introduced distribution to effectively approximate the marginalization results of the probabilistic variables. Hence, our Bayesian Window Transformer not only inherits the powerful representation ability but also maintains essential properties like translation invariance and local relationship preservation for image restoration. We also provide a theoretical guarantee, demonstrating that our method is aligned with the classic sliding window technique in terms of receptive field sizes and sliding behavior. Comprehensive experiments validate the exceptional effectiveness of our Bayesian Window Transformer across multiple image restoration tasks, including image deraining, denoising, and deblurring. Jie Xiao 0002, Xueyang Fu, Yurui Zhu, Zhengjun Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Toward Better De-Raining Generalization via Rainy Characteristics Memorization and ReplayabstractCurrent image de-raining methods primarily learn from a limited dataset, leading to inadequate performance in varied real-world rainy conditions. To tackle this, we introduce a new framework that enables networks to progressively expand their de-raining knowledge base by tapping into a growing pool of datasets, significantly boosting their adaptability. Drawing inspiration from the human brain's ability to continually absorb and generalize from ongoing experiences, our approach borrows the mechanism of the complementary learning system. Specifically, we first deploy generative adversarial networks (GANs) to capture and retain the unique features of new data, mirroring the hippocampus's role in learning and memory. Then, the de-raining network is trained with both existing and GAN-synthesized data, mimicking the process of hippocampal replay and interleaved learning. Furthermore, we employ knowledge distillation with the replayed data to replicate the synergy between the neocortex's activity patterns triggered by hippocampal replays and the preexisting neocortical knowledge. This comprehensive framework empowers the de-raining network to accumulate knowledge from various datasets, continually enhancing its performance on previously unseen rainy scenes. Our testing on three benchmark de-raining networks confirms the framework's effectiveness. It not only facilitates continual knowledge accumulation across six datasets but also surpasses state-of-the-art methods in generalizing to new real-world scenarios. Our code is available at https://github.com/wangkunyu241/CLGID. Xueyang Fu, Chengzhi Cao, Chengjie Ge, Wei Zhai, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video ReconstructionabstractLeveraging its robust linear global modeling capability, Mamba has notably excelled in computer vision. Despite its success, existing Mamba-based vision models have overlooked the nuances of event-driven tasks, especially in video reconstruction. Event-based video reconstruction (EBVR) demands spatial translation invariance and close attention to local event relationships in the spatio-temporal domain. Unfortunately, conventional Mamba algorithms apply static window partitions and standard reshape scanning methods, leading to significant losses in local connectivity. To overcome these limitations, we introduce EventMamba—a specialized model designed for EBVR task. EventMamba innovates by incorporating random window offset (RWO) in the spatial domain, moving away from the restrictive fixed partitioning. Additionally, it features a new consistent traversal serialization approach in the spatio-temporal domain, which maintains the proximity of adjacent events both spatially and temporally. These enhancements enable EventMamba to retain Mamba’s robust modeling capabilities while significantly preserving the spatio-temporal locality of event data. Comprehensive testing on multiple datasets shows that EventMamba markedly enhances video reconstruction, drastically improving computation speed while delivering superior visual quality compared to Transformer-based methods. Chengjie Ge, Xueyang Fu, Peng He 0004, Chengzhi Cao, Zhengjun Zha |
AAAI | 2 |
| 2025 | DreamUHD: Frequency Enhanced Variational Autoencoder for Ultra-High-Definition Image RestorationabstractExisting ultra-high-definition (UHD) image restoration methods often struggle with consistency due to downsampling. We aim to address these challenges by leveraging the powerful latent space representation and reconstruction capabilities of Variational Autoencoders (VAE). However, applying VAE to UHD image restoration presents challenges: 1) High-performing VAEs have large parameter sizes, leading to significant carbon footprints; 2) The self-reconstruction property of VAE hinders bridging the domain gap between clean and degraded images; 3) Latent encoding in VAE can lose high-frequency information, compromising image detail. To overcome these challenges, we propose a frequency enhanced VAE UHD image restoration framework by integrating frequency priors. First, we design the Fourier-based lightweight frequency learning within the VAE to improve parameter efficiency. Then, we introduce a wavelet-based adapter that extracts multi-scale image information and employs frequency-aware adaptive modulation to bridge the domain gap by integrating degraded image data into the pre-trained VAE. Additionally, the adapter injects high-frequency information into the VAE decoder, enhancing detail in the restored images. In this way, our method effectively combines the powerful latent space representation with frequency priors to enhance UHD image restoration. Extensive experiments on various UHD image restoration tasks show that our method surpasses state-of-the-art methods both qualitatively and quantitatively. Yidi Liu, Dong Li 0055, Jie Xiao 0002, Yuanfei Bao, Senyan Xu, Xueyang Fu |
AAAI | 6 |
| 2025 | SCott: Accelerating Diffusion Models with Stochastic Consistency DistillationabstractThe iterative sampling procedure employed by diffusion models (DMs) often leads to significant latency. To address this, we propose Stochastic Consistency Distillation (SCott) to enable accelerated text-to-image generation, where high-quality generations can be achieved with just 2-4 sampling steps or even1 step, and further improvements can be obtained by additional cost, e.g., 4 steps. In contrast to vanilla consistency distillation (CD) which distills the ordinary differential equation solvers-based sampling process of a pre-trained teacher model into a student, SCott explores the possibility and validates the efficacy of integrating stochastic differential equation (SDE) solvers into CD to fully unleash the potential of the teacher. SCott is augmented with elaborate strategies to control the noise strength and sampling process of the SDE solver. An adversarial loss is further incorporated to strengthen the sample quality with rare sampling steps. Empirically, on the MSCOCO-2017 5K dataset with a Stable Diffusion-V1.5 teacher, SCott achieves an FID of 21.9, surpassing that of the 1-step InstaFlow (23.4) and the 4-step UFOGen (22.1). Moreover, SCott can yield more diverse samples than other consistency models for high-resolution image generation, with up to 16% improvement in a qualified metric. Hongjian Liu, Qingsong Xie, Tianxiang Ye, Zhijie Deng, Chen Chen 0015, Shixiang Tang, Xueyang Fu, Haonan Lu, Zhengjun Zha |
AAAI | 7 |
| 2025 | Boosting Image De-Raining via Central-Surrounding Synergistic ConvolutionabstractRainy images suffer from quality degradation due to the synergistic effect of rain streaks and accumulation. The rain streaks are anisotropic and show a specific directional arrangement, while the rain accumulation is isotropic and shows a consistent concentration distribution in local regions. This distribution difference makes unified representation learning for rain streaks and accumulation challenging, which may lead to structure distortion and contrast degradation in the deraining results. To address this problem, a central-surrounding mechanism inspired Synergistic Convolution (SC) is proposed to extract rain streaks and accumulation features simultaneously. Specifically, the SC consists of two parallel novel convolutions: Central-Surrounding Difference Convolution (CSD) and Central-Surrounding Addition Convolution (CSA). In CSD, the difference operation between central and surrounding pixels is injected into the feature extraction process of convolution to perceive the direction distribution of rain streaks. In CSA, the addition operation between central and surrounding pixels is injected into the feature extraction process of convolution to facilitate the modeling of rain accumulation properties. The SC can be used as a general unit to substitute Vanilla Convolution (VC) in current de-raining networks to boost performance. To reduce computational costs, CSA and CSD in SC are merged into a single VC kernel by our parameter equivalent transformation before inferencing. Evaluations of twelve de-raining methods on nine public datasets demonstrate that our proposed SC can comprehensively improve the performance of twelve de-raining networks under various rainy conditions without changing the original network structure or introducing extra computational costs. Even for the current SOTA methods, SC can further achieve SOTA++ performance. The source codes will be publicly available. Long Peng 0003, Yang Wang 0015, Xin Di, Peizhe Xia, Xueyang Fu, Yang Cao 0010, Zhengjun Zha |
AAAI | 5 |
| 2025 | DCTMamba: Advancing JPEG Image Restoration Through Long-Sequence Modeling and Adaptive Frequency StrategyabstractDespite the advanced long-sequence modeling of Mamba, which has expanded its applications in image restoration, there remains a lack of exploration combining its strengths with the specific characteristics of JPEG image restoration, where high-frequency components are lost after the Discrete Cosine Transform (DCT). To address this, we introduce DCTMamba, a new framework designed to apply Mamba more effectively to JPEG image restoration. Specifically, our method integrates the Discrete Cosine Transform (DCT) into the Mamba to establish the sequential scanning from lower to higher frequencies, enabling the network to initially reconstruct coarse structures and progressively refine the image with more intricate details. Furthermore, recognizing the variable frequency distributions that arise from DCT transformations across different image sizes, we have developed Scale-Adaptive Normalization to manage these variations adeptly. Comprehensive experiments confirm that DCTMamba significantly outperforms existing solutions, achieving high fidelity in both coarse structures and fine details.CTMamba significantly outperforms existing solutions, achieving high fidelity in both coarse structures and fine details. Xi Wang 0018, Xueyang Fu, Liang Li 0003, Zhengjun Zha |
AAAI | 2 |
| 2025 | Motion-adaptive Transformer for Event-based Image DeblurringabstractEvent cameras, which capture pixel-level brightness changes asynchronously, provide rich motion information that is often missed during traditional frame-based camera exposures, thereby offering fresh perspectives for motion deblurring. Although current approaches incorporate event intensity, they neglect essential spatial motion information. Unlike their CNN architectures, Transformers excel in modeling long-range dependencies but struggle with establishing relevant non-local connections in sparse events and fail to highlight significant interactions in dense images. To address these limitations, we introduce a Motion-Adaptive Transformer network (MAT) that utilizes spatial motion information to forge robust global connections. The core design is an Adaptive Motion Mask Predictor (AMMP) that identifies key motion regions, guiding the Motion-Sparse Attention (MSA) to eliminate irrelevant event tokens and enabling the Motion-Aware Attention (MAA) to focus on relevant ones, thereby enhancing long-range dependency modeling. Additionally, we elaborately design a Cross-Modal Intensity Gating mechanism that efficiently merges intensity data across modalities while minimizing parameter use. The learnable Expansion-Controlled Spatial Gating further optimizes the transmission of event features. Comprehensive testing confirms that our approach sets a new benchmark in image deblurring, surpassing previous methods by up to 0.60dB on the GoPro dataset, 1.04dB on the HS-ERGB dataset, and achieving an average improvement of 0.52dB across two real-world datasets. Senyan Xu, Zhijing Sun, Mingchen Zhong, Chengzhi Cao, Yidi Liu, Xueyang Fu, Yan Chen 0007 |
AAAI | 6 |
| 2025 | A Lottery Ticket Hypothesis Approach with Sparse Fine-tuning and MAE for Image Forgery Detection and LocalizationabstractThe rise in sophisticated image forgery techniques, driven by advancements in image editing and generation, has posed new security challenges. Traditional methods, designed for specific tampering artifacts, struggle with out-of-distribution image forgery detection. In this paper, we propose a shift in paradigm, placing greater emphasis on the universal characteristics of authentic images, as opposed to solely focusing on specific forgery signals. We introduce an enhancement to the Masked Autoencoder (MAE), aptly termed the Forgery MAE (FMAE). This modification retains the inherent characteristics of natural images while integrating multi-source forgery information. Our implementation involves applying the lottery ticket hypothesis during pre-training to identify forgery-sensitive parameters, followed by their sparse fine-tuning to target the forgery detection and localization task. Concurrently, we develop a ``mixture of experts'' noise extractor to compile multi-source forgery data. Our FMAE effectively extracts forgery features and shows strong resilience against unseen forgeries. Extensive experiments across multiple datasets confirm our method's superior accuracy and generalization capability over existing techniques. Jiaying Zhu, Dong Li 0055, Xueyang Fu, Gege Shi, Jie Xiao 0002, Aiping Liu, Zhengjun Zha |
AAAI | 3 |
| 2025 | Continuous Adverse Weather Removal via Degradation-Aware DistillationabstractAll-in-one models for adverse weather removal aim to process various degraded images using a single set of parameters, making them ideal for real-world scenarios. However, they encounter two main challenges: catastrophic forgetting and limited degradation awareness. The former causes the model to lose knowledge of previously learned scenarios, reducing its overall effectiveness. While the later hampers the model’s ability to accurately identify and respond to specific types of degradation, limiting its performance across diverse adverse weather conditions. To address these issues, we introduce the Incremental Learning Adverse Weather Removal (ILAWR) framework, which uses a novel degradation-aware distillation strategy for continuous weather removal. Specifically, we first design a degradation-aware module that utilizes Fourier priors to capture a broad range of degradation features, effectively mitigating catastrophic forgetting in low-level visual tasks. Then, we implement multilateral distillation, which combines knowledge from multiple teacher models using an importance-guided aggregation approach. This enables the model to balance adaptation to new degradation types with the preservation of background details. Extensive experiments confirm that ILAWR outperforms existing models across multiple benchmarks, proving its effectiveness in continuous adverse weather removal. Xin Lu 0008, Jie Xiao 0002, Yurui Zhu, Xueyang Fu |
CVPR | 4 |
| 2025 | UHD-processer: Unified UHD Image Restoration with Progressive Frequency Learning and Degradation-aware PromptsabstractWe introduce UHD-Processor, a unified and robust framework for all-in-one image restoration, which is particularly resource-efficient for Ultra-High-Definition (UHD) images. To address the limitations of traditional all-in-one methods that rely on complex restoration backbones, our strategy employs a frequency domain decoupling progressive learning technique, motivated by curriculum learning, to incrementally learn restoration mappings from low to high frequencies. This approach incorporates specialized sub-network modules to effectively tackle different frequency bands in a divide-and-conquer manner, significantly enhancing the learning capability of simpler networks. Moreover, to accommodate the high-resolution characteristics of UHD images, we developed a variational autoencoder (VAE)-based framework that reduces computational complexity by modeling a concise latent space. It integrates task-specific degradation awareness in the encoder and frequency selection in the decoder, enhancing task comprehension and generalization. Our unified model is able to handle various degradations such as denoising, deblurring, dehazing, low-lighting, etc. Experimental evaluations extensively showcase the effectiveness of our dual-strategy approach, significantly improving UHD image restoration and achieving cutting-edge performance across diverse conditions. The code will be available at https://github.com/lyd-2022/UHD-processer Yidi Liu, Dong Li 0055, Xueyang Fu, Xin Lu 0008, Jie Huang 0017, Zhengjun Zha |
CVPR | 3 |
| 2025 | Efficient Test-time Adaptive Object Detection via Sensitivity-Guided PruningabstractContinual test-time adaptive object detection (CTTA-OD) aims to online adapt a source pre-trained detector to everchanging environments during inference under continuous domain shifts. Most existing CTTA-OD methods prioritize effectiveness while overlooking computational efficiency, which is crucial for resource-constrained scenarios. In this paper, we propose an efficient CTTA-OD method via pruning. Our motivation stems from the observation that not all learned source features are beneficial; certain domain-sensitive feature channels can adversely affect target domain performance. Inspired by this, we introduce a sensitivity-guided channel pruning strategy that quantifies each channel based on its sensitivity to domain discrepancies at both image and instance levels. We apply weighted sparsity regularization to selectively suppress and prune these sensitive channels, focusing adaptation efforts on invariant ones. Additionally, we introduce a stochastic channel reactivation mechanism to restore pruned channels, enabling recovery of potentially useful features and mitigating the risks of early pruning. Extensive experiments on three benchmarks show that our method achieves superior adaptation performance while reducing computational overhead by 12% in FLOPs compared to the recent SOTA method. Xueyang Fu, Xin Lu 0008, Chengjie Ge, Chengzhi Cao, Wei Zhai, Zhengjun Zha |
CVPR | 2 |
| 2025 | Enhanced Pansharpening Via Quaternion Spatial-Spectral Interactions
Dong Liu 0002, Chunhui Luo, Yuanfei Bao, Jie Xiao 0002, Xueyang Fu, Zhengjun Zha |
ICCV | 6 |
| 2025 | Decouple to Reconstruct: High Quality UHD Restoration Via Active Feature Disentanglement and Reversible Fusion
Yidi Liu, Dong Liu 0002, Jie Huang 0017, Xueyang Fu, Zhengjun Zha |
ICCV | 6 |
| 2025 | EVDM: Event-based Real-World Video Deblurring with Mamba
Zhijing Sun, Senyan Xu, Kean Liu, Runze Tian, Xueyang Fu, Zhengjun Zha |
ICCV | 5 |
| 2025 | Event Denoising Based on Iterative Tree-Structured Information AggregationabstractEvent cameras play a crucial role in the visual field; however, they are susceptible to noise. Traditional denoising algorithms for event cameras often struggle to balance accuracy and speed. To address this issue, this paper proposes an algorithm named Event Denoising Based on Iterative Tree-Structured Information Aggregation (EDIST). Specifically, the proposed method first establishes connections in event streams using a spatiotemporal window to extract Relation Tree. Then, a pruning algorithm is employed to streamline the subsequent information aggregation process, followed by a multi-stage convolution module designed to process the Relation Tree and obtain aggregated features. Finally, these aggregated features are fed into a classification module to determine whether an event is noise. During the inference stage, an information storage reuse module is designed to enable iterative execution, thereby enhancing inference speed. Experimental results demonstrate that the proposed algorithm outperforms existing methods in both denoising accuracy and inference speed on the DVSNOISE20 dataset. Yueyang Xu, Chengjie Ge, Xueyang Fu, Zhengjun Zha |
ICIP | 3 |
| 2025 | FourierMamba: Fourier Learning Integration with State Space Models for Image DerainingabstractImage deraining aims to remove rain streaks from rainy images and restore clear backgrounds. Currently, some research that employs the Fourier transform has proved to be effective for image deraining, due to it acting as an effective frequency prior for capturing rain streaks. However, despite there exists dependency of low frequency and high frequency in images, these Fourier-based methods rarely exploit the correlation of different frequencies for conjuncting their learning procedures, limiting the full utilization of frequency information for image deraining. Alternatively, the recently emerged Mamba technique depicts its effectiveness and efficiency for modeling correlation in various domains (e.g., spatial, temporal), and we argue that introducing Mamba into its unexplored Fourier spaces to correlate different frequencies would help improve image deraining. This motivates us to propose a new framework termed FourierMamba, which performs image deraining with Mamba in the Fourier space. Owing to the unique arrangement of frequency orders in Fourier space, the core of FourierMamba lies in the scanning encoding of different frequencies, where the low-high frequency order formats exhibit differently in the spatial dimension (unarranged in axis) and channel dimension (arranged in axis). Therefore, we design FourierMamba that correlates Fourier space information in the spatial and channel dimensions with distinct designs. Specifically, in the spatial dimension Fourier space, we introduce the zigzag coding to scan the frequencies to rearrange the orders from low to high frequencies, thereby orderly correlating the connections between frequencies; in the channel dimension Fourier space with arranged orders of frequencies in axis, we can directly use Mamba to perform frequency correlation and improve the channel information representation. Extensive experiments reveal that our method outperforms state-of-the-art methods both qualitatively and quantitatively. Dong Li 0055, Yidi Liu, Xueyang Fu, Jie Huang 0017, Senyan Xu, Qi Zhu 0010, Zhengjun Zha |
ICML | 3 |
| 2025 | Directing Mamba to Complex Textures: An Efficient Texture-Aware State Space Model for Image RestorationabstractImage restoration aims to recover details and enhance contrast in degraded images. With the growing demand for high-quality imaging (e.g., 4K and 8K), achieving a balance between restoration quality and computational efficiency has become increasingly critical. Existing methods, primarily based on CNNs, Transformers, or their hybrid approaches, apply uniform deep representation extraction across the image. However, these methods often struggle to effectively model long-range dependencies and largely overlook the spatial characteristics of image degradation (regions with richer textures tend to suffer more severe damage), making it hard to achieve the best trade-off between restoration quality and efficiency. To address these issues, we propose a novel texture-aware image restoration method, TAMambaIR, which simultaneously perceives image textures and achieves a trade-off between performance and efficiency. Specifically, we introduce a novel Texture-Aware State Space Model, which enhances texture awareness and improves efficiency by modulating the transition matrix of the state-space equation and focusing on regions with complex textures. Additionally, we design a Multi-Directional Perception Block to improve multi-directional receptive fields while maintaining low computational overhead. Extensive experiments on benchmarks for image super-resolution, deraining, and low-light image enhancement demonstrate that TAMambaIR achieves state-of-the-art performance with significantly improved efficiency, establishing it as a robust and efficient framework for image restoration. Long Peng 0003, Xin Di, Zhanfeng Feng, Wenbo Li 0002, Renjing Pei, Yang Wang 0015, Xueyang Fu, Yang Cao 0010, Zhengjun Zha |
IJCAI | 7 |
| 2025 | Learnable Frequency Decomposition for Image Forgery Detection and LocalizationabstractConcern for image authenticity spurs research in image forgery detection and localization (IFDL). Most deep learning-based methods focus primarily on spatial domain modeling and have not fully explored frequency domain strategies. In this paper, we observe and analyze the frequency characteristic changes caused by image tampering. Observations indicate that manipulation traces are especially prominent in phase components and span both low and high-frequency bands. Based on these findings, we propose a forensic frequency decomposition network (F2D-Net), which incorporates deep Fourier transforms and leverages both phase information and high and low-frequency components to enhance IFDL. Specifically, F2D-Net consists of the Spectral Decomposition Subnetwork (SDSN) and the Frequency Separation Subnetwork (FSSN). The former decomposes the image into amplitude and phase, focusing on learning the semantic content in the phase spectrum to identify forged objects, thus improving forgery detection accuracy. The latter further adaptively decomposes the output of the SDSN to obtain corresponding high and low frequencies, and applies a divide-and-conquer strategy to refine each frequency band, mitigating the optimization difficulties caused by coupled forgery traces across different frequencies, thereby better capturing the pixels belonging to the forged object to improve localization accuracy. Experiments on multiple datasets demonstrate that our method outperforms state-of-the-art image forgery detection and localization techniques both qualitatively and quantitatively. Dong Li 0055, Jiaying Zhu, Yidi Liu, Xin Lu 0008, Xueyang Fu, Jiawei Liu 0001, Aiping Liu, Zhengjun Zha |
IJCAI | 5 |
| 2025 | PanComplex: Leveraging Complex-Valued Neural Networks for Enhanced PansharpeningabstractPansharpening combines panchromatic and low-resolution multispectral images to generate high-resolution multispectral images. Previous studies have explored the connection between pansharpening and the frequency domain, but mostly in the real-valued domain, leaving the complex domain relatively unexplored. To redefine the pansharpening task, we propose a complex-valued spatial-frequency dual-domain framework, PanComplex. To achieve this, we first establish complex representations and introduce basic complex operators tailored to pansharpening, enabling the transformation of multispectral real-valued signals into the complex domain for learning. We then model both spatial and frequency branches to capture global frequency features and local spatial features comprehensively. Finally, we employ a complex-based interaction module to fuse the spatial and frequency features, achieving complementary information across both domains. By using the representation power of the complex domain, PanComplex effectively extracts complementary features from PAN and MS images, thereby enhancing pansharpening performance. Experiments on multiple datasets demonstrate that our method achieves optimal performance with the fewest parameters and exhibits strong generalization ability to other tasks. The source code for this work is publicly available at https://github.com/lch-ustc/PanComplex. Chunhui Luo, Dong Li 0055, Xin Lu 0008, Jiangtong Tan, Xueyang Fu |
IJCAI | 7 |
| 2025 | Efficient Prompt-based Multimodal Interaction for Audio-Visual Event LocalizationabstractAudio-Visual Event Localization (AVEL) requires localizing an event by jointly processing audio and visual information. Most existing AVEL methods commonly utilize two distinct models independently trained on image and audio datasets to encode features and take the extracted features as the input of the model. However, features extracted from unimodal pre-trained models lack cross-modal interaction and may also contain noise irrelevant to AVEL, which leads to sub-optimal performance. To address this issue, we propose an efficient prompt-based multimodal interaction approach for audio-visual event localization. Specifically, our method freezes a pre-trained transformer model and designs query and global prompt to facilitate information exchange and fusion across modalities. Combined with end-to-end training from raw data to event localization, our method can obtain more task-relevant features. Since only the parameters of prompts are updated, our method avoids the significant computational resource consumption associated with fine-tuning the entire transformer model. Additionally, our method enables the adaptation of the visual pre-trained model to downstream audio-visual tasks and facilitates information exchange and fusion between the video and audio modalities. Experimental results on the public AVE dataset demonstrate that our method, when compared to state-of-the-art approaches, achieves competitive performance while significantly reducing the number of trainable parameters. Longzhuo Huang, Liang Li 0003, Xueyang Fu, Zhengjun Zha |
ICMR | 3 |
| 2025 | Neural Fractional Attention Differential EquationsabstractThe integration of differential equations with neural networks has created powerful tools for modeling complex dynamics effectively across diverse machine learning applications. While standard integer-order neural ordinary differential equations (ODEs) have shown considerable success, they are limited in their capacity to model systems with memory effects and historical dependencies. Fractional calculus offers a mathematical framework capable of addressing this limitation, yet most current fractional neural networks use static memory weightings that cannot adapt to input-specific contextual requirements. This paper proposes a generalized neural Fractional Attention Differential Equation (FADE), which combines the memory-retention capabilities of fractional calculus with contextual learnable attention mechanisms. Our approach replaces fixed kernel functions in fractional operators with neural attention kernels that adaptively weight historical states based on their contextual relevance to current predictions. This allows our framework to selectively emphasize important temporal dependencies while filtering less relevant historical information. Our theoretical analysis establishes solution boundedness, problem well-posedness, and numerical equation solver convergence properties of the proposed model. Furthermore, through extensive evaluation on tasks such as fluid flow, graph learning problems and spatio-temporal traffic flow forecasting, we demonstrate that our adaptive attention-based fractional framework outperforms both integer-order neural ODE models and existing fractional approaches. The results confirm that our framework provides superior modeling capacity for complex dynamics with varying temporal dependencies. The code is available at \url{https://github.com/cuiwjTech/NeurIPS2025_FADE}. Qiyu Kang, Wenjun Cui, Xuhao Li, Xueyang Fu, Wee-Peng Tay, Yidong Li, Zhengjun Zha |
NeurIPS | 5 |
| 2025 | Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable RefinementabstractUltra-High Definition (UHD) image restoration struggles to balance computational efficiency and detail retention.
While Variational Autoencoders (VAEs) offer improved efficiency by operating in the latent space, with the Gaussian variational constraint, this compression preserves semantics but sacrifices critical high-frequency attributes specific to degradation and thus compromises reconstruction fidelity.
% This compromises reconstruction fidelity, even when global semantics are preserved.
Consequently, a VAE redesign is imperative to foster a robust semantic representation conducive to generalization and perceptual quality, while simultaneously enabling effective high-frequency information processing crucial for reconstruction fidelity.
To address this, we propose \textit{Latent Harmony}, a two-stage framework that reinvigorates VAEs for UHD restoration by concurrently regularizing the latent space and enforcing high-frequency-aware reconstruction constraints.
Specifically, Stage One introduces the LH-VAE, which fortifies its latent representation through visual semantic constraints and progressive degradation perturbation for enhanced semantics robustness; meanwhile, it incorporates latent equivariance to bolster its high-frequency reconstruction capabilities.
Then, Stage Two facilitates joint training of this refined VAE with a dedicated restoration model.
This stage integrates High-Frequency Low-Rank Adaptation (HF-LoRA), featuring two distinct modules: an encoder LoRA, guided by a fidelity-oriented high-frequency alignment loss, tailored for the precise extraction of authentic details from degradation-sensitive high-frequency components; and a decoder LoRA, driven by a perception-oriented loss, designed to synthesize perceptually superior textures. These LoRA modules are meticulously trained via alternating optimization with selective gradient propagation to preserve the integrity of the pre-trained latent structure. This methodology culminates in a flexible fidelity-perception trade-off at inference, managed by an adjustable parameter
$\alpha$.
Extensive experiments demonstrate that \textit{Latent Harmony} effectively balances perceptual and reconstructive objectives with efficiency, achieving superior restoration performance across diverse UHD and standard-resolution scenarios. Yidi Liu, Xueyang Fu, Jie Huang 0017, Jie Xiao 0002, Dong Li 0055, Lei Bai 0001, Zhengjun Zha |
NeurIPS | 2 |
| 2025 | PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time AdaptationabstractContinual Test-Time Adaptation (CTTA) aims to online adapt a pre-trained model to changing environments during inference. Most existing methods focus on exploiting target data, while overlooking another crucial source of information, the pre-trained weights, which encode underutilized domain-invariant priors. This paper takes the geometric attributes of pre-trained weights as a starting point, systematically analyzing three key components: magnitude, absolute angle, and pairwise angular structure. We find that the pairwise angular structure remains stable across diverse corrupted domains and encodes domain-invariant semantic information, suggesting it should be preserved during adaptation. Based on this insight, we propose PAID (Pairwise Angular Invariant Decomposition), a prior-driven CTTA method that decomposes weight into magnitude and direction, and introduces a learnable orthogonal matrix via Householder reflections to globally rotate direction while preserving the pairwise angular structure. During adaptation, only the magnitudes and the orthogonal matrices are updated. PAID achieves consistent improvements over recent SOTA methods on four widely used CTTA benchmarks, demonstrating that preserving pairwise angular structure offers a simple yet effective principle for CTTA. Our code is available at https://github.com/wangkunyu241/PAID. Xueyang Fu, Yuanfei Bao, Chengjie Ge, Chengzhi Cao, Wei Zhai, Zhengjun Zha |
NeurIPS | 2 |
| 2025 | Exploring Local Sparse Structure Prior for Image Deraining and DesnowingabstractExisting image deraining and desnowing methods are typically trained under specific weather conditions, which limits their effectiveness in locating rain streaks and snowflakes in diverse, open scenes. This restriction often leads to suboptimal restoration performance. To address these limitations, we propose a novel local sparse structure prior for rain and snow, characterized by high pixel intensity and the locally sparse spatial distribution of rain streaks and snowflakes. Leveraging this prior, we developed an algorithm that extracts rain and snow structure masks, enabling precise localization of rain streaks and snowflake regions across open scenes. In addition, we introduce a refinement and compensation process to remove irrelevant information from the masks and correct mask estimation errors. We further construct a Mask-Guided Restoration Network (MGNet) that utilizes the rain and snow structure masks effectively and includes a mask-conditioned attention module to focus restoration efforts on degraded areas affected by rain streaks and snowflakes. Extensive experimental results demonstrate that our method significantly outperforms current state-of-the-art techniques in open scenes, effectively restoring various types of rain streaks and snowflakes with a single model parameter configuration. Xin Guo 0018, Xueyang Fu, Zhengjun Zha |
IEEE Signal Process. Lett. | 2 |
| 2025 | Deep Unfolding Network for Image Desnowing With Snow Shape PriorabstractEffectively leveraging snow image formulation, which accounts for atmospheric light and snow masks, is crucial for enhancing image desnowing performance and improving interpretability. However, current direct-learning approaches often neglect this formulation, while model-based methods use it in overly simplistic ways. To address this, we propose a novel unfolding network that iteratively refines the desnowing process for more thorough optimization. Additionally, model-based techniques usually rely on real-world snow masks for supervision, a requirement that is impractical in many real-world applications. To overcome this limitation, we introduce a snow shape prior as a surrogate supervision signal. We further integrate the physical properties of atmospheric light and heavy snow by decomposing the optimization task into manageable sub-problems within our unfolding network. Extensive evaluations on multiple benchmark datasets confirm that our method outperforms current state-of-the-art techniques. Xin Guo 0018, Xi Wang 0018, Xueyang Fu, Zhengjun Zha |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Event-Based Video Reconstruction With Deep Spatial-Frequency Unfolding NetworkabstractCurrent event-based video reconstruction methods, limited to the spatial domain, face challenges in decoupling brightness and structural information, leading to exposure distortion, and in efficiently acquiring non-local information without relying on computationally expensive Transformer models. To address these issues, we propose the Deep Spatial-Frequency Unfolding Reconstruction Network (DSFURNet), which explores and utilizes knowledge in the frequency domain for event-based video reconstruction. Specifically, we construct a variational model and propose three regularization terms: a brightness regularization term approximated by Fourier amplitudes, a structural regularization term approximated by Fourier phases, and an initialization regularization term that converts event representations into initial video frames. Then, we design corresponding spatial-frequency domain approximation operators for each regularization term. Benefiting from the global nature of computations in the frequency domain, the designed approximation operators can integrate local spatial and global frequency information at a lower computational cost. Furthermore, we combine the learned knowledge of the three regularization terms and unfold the optimization algorithm into an iterative deep network. Through this approach, the pixel-level initialization regularization constraint and the frequency domain brightness and structural regularization constraints can continuously play a role during the testing process, achieving a gradual improvement in the quality of the reconstructed video frames. Compared to existing methods, our network significantly reduces the number of network parameters while improving evaluation metrics. Chengjie Ge, Xueyang Fu, Zhengjun Zha |
IEEE Trans. Image Process. | 2 |
| 2025 | Event-Driven Video Restoration With Spiking-Convolutional ArchitectureabstractWith high temporal resolution, high dynamic range, and low latency, event cameras have made great progress in numerous low-level vision tasks. To help restore low-quality (LQ) video sequences, most existing event-based methods usually employ convolutional neural networks (CNNs) to extract sparse event features without considering the spatial sparse distribution or the temporal relation in neighboring events. It brings about insufficient use of spatial and temporal information from events. To address this problem, we propose a new spiking-convolutional network (SC-Net) architecture to facilitate event-driven video restoration. Specifically, to properly extract the rich temporal information contained in the event data, we utilize a spiking neural network (SNN) to suit the sparse characteristics of events and capture temporal correlation in neighboring regions; to make full use of spatial consistency between events and frames, we adopt CNNs to transform sparse events as an extra brightness prior to being aware of detailed textures in video sequences. In this way, both the temporal correlation in neighboring events and the mutual spatial information between the two types of features are fully explored and exploited to accurately restore detailed textures and sharp edges. The effectiveness of the proposed network is validated in three representative video restoration tasks: deblurring, super-resolution, and deraining. Extensive experiments on synthetic and real-world benchmarks have illuminated that our method performs better than existing competing methods. Chengzhi Cao, Xueyang Fu, Yurui Zhu, Zhijing Sun, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | DDCNet: Advanced Decoupling of Degradation and Content for Adverse Weather Image RestorationabstractAdverse weather image restoration aims to recover clear images from those affected by weather conditions such as rain, haze, and snow. Different weather types affect images in distinct ways, necessitating specific degradation removal strategies, while content reconstruction generally benefits from a consistent approach since the underlying image structure remains largely consistent. Previous methods, despite their ability to handle multiple weather conditions within a single framework, often failed to adequately separate these two critical processes, thereby adversely affecting image restoration quality. In this article, we present DDCNet, a novel framework designed to explicitly decouple degradation removal and content reconstruction when processing various adverse weather conditions within a unified network. We achieve this by separating tailored degradation removal from uniform content reconstruction at the feature level, based on channel statistics. Additionally, we utilize the Fourier transform to enhance both processes. Furthermore, to address the differing optimization directions required by different adverse weather types, we propose a novel degradation mapping (DM) loss function to constrain their respective optimization paths. Extensive experiments show that DDCNet establishes new performance standards across multiple adverse weather scenarios. Xi Wang 0018, Xueyang Fu, Yurui Zhu, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Neuromorphic Event Signal-Driven Network for Video De-rainingabstractConvolutional neural networks-based video de-raining methods commonly rely on dense intensity frames captured by CMOS sensors. However, the limited temporal resolution of these sensors hinders the capture of dynamic rainfall information, limiting further improvement in de-raining performance. This study aims to overcome this issue by incorporating the neuromorphic event signal into the video de-raining to enhance the dynamic information perception. Specifically, we first utilize the dynamic information from the event signal as prior knowledge, and integrate it into existing de-raining objectives to better constrain the solution space. We then design an optimization algorithm to solve the objective, and construct a de-raining network with CNNs as the backbone architecture using a modular strategy to mimic the optimization process. To further explore the temporal correlation of the event signal, we incorporate a spiking self-attention module into our network. By leveraging the low latency and high temporal resolution of the event signal, along with the spatial and temporal representation capabilities of convolutional and spiking neural networks, our model captures more accurate dynamic information and significantly improves de-raining performance. For example, our network achieves a 1.24dB improvement on the SynHeavy25 dataset compared to the previous state-of-the-art method, while utilizing only 39% of the parameters. Chengjie Ge, Xueyang Fu, Peng He 0004, Chengzhi Cao, Zhengjun Zha |
AAAI | 2 |
| 2024 | Learning Discriminative Noise Guidance for Image Forgery Detection and LocalizationabstractThis study introduces a new method for detecting and localizing image forgery by focusing on manipulation traces within the noise domain. We posit that nearly invisible noise in RGB images carries tampering traces, useful for distinguishing and locating forgeries. However, the advancement of tampering technology complicates the direct application of noise for forgery detection, as the noise inconsistency between forged and authentic regions is not fully exploited. To tackle this, we develop a two-step discriminative noise-guided approach to explicitly enhance the representation and use of noise inconsistencies, thereby fully exploiting noise information to improve the accuracy and robustness of forgery detection. Specifically, we first enhance the noise discriminability of forged regions compared to authentic ones using a de-noising network and a statistics-based constraint. Then, we merge a model-driven guided filtering mechanism with a data-driven attention mechanism to create a learnable and differentiable noise-guided filter. This sophisticated filter allows us to maintain the edges of forged regions learned from the noise. Comprehensive experiments on multiple datasets demonstrate that our method can reliably detect and localize forgeries, surpassing existing state-of-the-art methods. Jiaying Zhu, Dong Li 0055, Xueyang Fu, Jie Huang 0017, Aiping Liu, Zhengjun Zha |
AAAI | 3 |
| 2024 | HomoFormer: Homogenized Transformer for Image Shadow RemovalabstractThe spatial non-uniformity and diverse patterns of shadow degradation conflict with the weight sharing manner of dominant models, which may lead to an unsatisfactory compromise. To tackle with this issue, we present a novel strategy from the view of shadow transformation in this paper: directly homogenizing the spatial distribution of shadow degradation. Our key design is the random shuffle operation and its corresponding inverse operation. Specifically, random shuffle operation stochastically rearranges the pixels across spatial space and the inverse operation recovers the original order. After randomly shuffling, the shadow diffuses in the whole image and the degradation appears in a homogenized way, which can be effectively processed by the local self-attention layer. Moreover, we further devise a new feed forward network with position modeling to exploit image structural information. Based on these elements, we construct the final local window based transformer named HomoFormer for image shadow removal. Our HomoFormer can enjoy the linear complexity of local transformers while bypassing challenges of non-uniformity and diversity of shadow. Extensive experiments are conducted to verify the superiority of our HomoFormer across public datasets. Code is available at https://github.com/jiexiaou/HomoFormer. Jie Xiao 0002, Xueyang Fu, Yurui Zhu, Dong Li 0055, Jie Huang 0017, Kai Zhu 0004, Zhengjun Zha |
CVPR | 2 |
| 2024 | Revisiting Single Image Reflection Removal in the WildabstractThis research focuses on the issue of single-image reflection removal (SIRR) in real-world conditions, examining it from two angles: the collection pipeline of real reflection pairs and the perception of real reflection locations. We devise an advanced reflection collection pipeline that is highly adaptable to a wide range of real-world reflection scenarios and incurs reduced costs in collecting large-scale aligned reflection pairs. In the process, we develop a large-scale, high-quality reflection dataset named Reflection Removal in the Wild (RRW). RRW contains over 14,950 high-resolution real-world reflection pairs, a dataset forty-five times larger than its predecessors. Regarding perception of reflection locations, we identify that numerous virtual reflection objects visible in reflection images are not present in the corresponding ground-truth images. This observation, drawn from the aligned pairs, leads us to conceive the Maximum Reflection Filter (MaxRF). The MaxRF could accurately and explicitly characterize reflection locations from pairs of images. Building upon this, we design a reflection location-aware cascaded framework, specifically tailored for SIRR. Powered by these innovative techniques, our solution achieves superior performance than current leading methods across multiple real-world benchmarks. Codes and datasets are available at here. Yurui Zhu, Xueyang Fu, Peng-Tao Jiang, Hao Zhang 0063, Qibin Sun, Jinwei Chen 0003, Zhengjun Zha, Bo Li 0130 |
CVPR | 2 |
| 2024 | Noise-Assisted Prompt Learning for Image Forgery Detection and Localization
Dong Li 0055, Jiaying Zhu, Xueyang Fu, Xun Guo 0001, Yidi Liu, Jiawei Liu 0001, Zhengjun Zha |
ECCV (11) | 3 |
| 2024 | Motion Aware Event Representation-Driven Image Deblurring
Zhijing Sun, Xueyang Fu, Longzhuo Huang, Aiping Liu, Zhengjun Zha |
ECCV (46) | 2 |
| 2024 | DreamClean: Restoring Clean Image Using Deep Diffusion PriorabstractImage restoration poses a garners substantial interest due to the exponential surge in demands for recovering high-quality images from diverse mobile camera devices, adverse lighting conditions, suboptimal shooting environments, and frequent image compression for efficient transmission purposes. Yet this problem gathers significant challenges as people are blind to the type of restoration the images suffer, which, is usually the case in real-day scenarios and is most urgent to solve for this field. Current research, however, heavily relies on prior knowledge of the restoration type, either explicitly through rules or implicitly through the availability of degraded-clean image pairs to define the restoration process, and consumes considerable effort to collect image pairs of vast degradation types. This paper introduces DreamClean, a training-free method that needs no degradation prior knowledge but yields high-fidelity and generality towards various types of image degradation. DreamClean embeds the degraded image back to the latent of pre-trained diffusion models and re-sample it through a carefully designed diffusion process that mimics those generating clean images. Thanks to the rich image prior in diffusion models and our novel Variance Preservation Sampling (VPS) technique, DreamClean manages to handle various different degradation types at one time and reaches far more satisfied final quality than previous competitors. DreamClean relies on elegant theoretical supports to assure its convergence to clean image when VPS has appropriate parameters, and also enjoys superior experimental performance over various challenging tasks that could be overwhelming for previous methods when degradation prior is unavailable. Jie Xiao 0002, Ruili Feng, Han Zhang 0010, Zhantao Yang, Yurui Zhu, Xueyang Fu, Kai Zhu 0004, Yu Liu 0063, Zhengjun Zha |
ICLR | 7 |
| 2024 | CCM: Real-Time Controllable Visual Content Creation Using Text-to-Image Consistency ModelsabstractConsistency Models (CMs) have showed a promise in creating high-quality images with few steps. However, the way to add new conditional controls to the pre-trained CMs has not been explored. In this paper, we explore the pivotal subject of leveraging the generative capacity and efficiency of consistency models to facilitate controllable visual content creation via ControlNet. First, it is observed that ControlNet trained for diffusion models (DMs) can be directly applied to CMs for high-level semantic controls but sacrifice image low-level details and realism. To tackle with this issue, we develop a CMs-tailored training strategy for ControlNet using the consistency training. It is substantiated that ControlNet can be successfully established through the consistency training technique. Besides, a unified adapter can be trained utilizing the consistency training, which enhances the adaptation of DM’s ControlNet. We quantitatively and qualitatively evaluate all strategies across various conditional controls, including sketch, hed, canny, depth, human pose, low-resolution image and masked image, with the pre-trained text-to-image latent consistency models. Jie Xiao 0002, Kai Zhu 0004, Han Zhang 0010, Yujun Shen, Zhantao Yang, Ruili Feng, Yu Liu 0063, Xueyang Fu, Zhengjun Zha |
ICML | 9 |
| 2024 | TSA2: Temporal Segment Adaptation and Aggregation for Video HarmonizationabstractVideo composition merges the foreground and background of different videos, presenting challenges due to variations in capture conditions (e.g., saturation, brightness, and contrast). Video harmonization is a vital process in achieving a realistic composite by seamlessly adjusting the foreground’s appearance to match the background. In this paper, we propose TSA2, a novel method for video harmonization that incorporates temporal segment adaptation and aggregation. TSA2divides the inharmonious input sequence into temporal segments, each corresponding to a different frame rate, allowing effective utilization of complementary information within each segment. The method includes the Temporal Segment Adaptation module, which learns and remaps the distribution difference between background and foreground regions, and the Temporal Segment Aggregation module, which emphasizes and aggregates cross-segment information through element-wise correlations. Experimental results demonstrate that TSA2outperforms advanced image and video harmonization methods quantitatively and qualitatively. Zeyu Xiao 0002, Yurui Zhu, Xueyang Fu, Zhiwei Xiong |
WACV | 3 |
| 2024 | Event-Driven Heterogeneous Network for Video Deraining
Xueyang Fu, Chengzhi Cao, Senyan Xu, Fanrui Zhang, Zhengjun Zha |
Int. J. Comput. Vis. | 1 |
| 2024 | Towards Generalized UAV Object Detection: A Novel Perspective from Frequency Domain Disentanglement
Xueyang Fu, Chengjie Ge, Chengzhi Cao, Zhengjun Zha |
Int. J. Comput. Vis. | 2 |
| 2024 | Vision-and-Language Navigation via Latent Semantic Alignment LearningabstractVision-and-Language Navigation (VLN) requires that an agent can comprehensively understand the given instructions and the immediate visual information obtained from the environment, so as to make correct actions to achieve the navigation goal. Therefore, semantic alignment across modalities is crucial for the agent understanding its own state during the navigation process. However, the potential of semantic alignment has not been systematically explored in current studies, which limits the further improvement of navigation performance. To address this issue, we propose a new Latent Semantic Alignment Learning method to develop the semantically aligned relationships contained in the environment. Specifically, we introduce three novel pre-training tasks: Trajectory-conditioned Masked Fragment Modeling, Action Prediction of Masked Observation, and Hierarchical Triple Contrastive Learning. The first two tasks are used to reason about cross-modal dependencies, while the third one is able to learn semantically consistent representations across modalities. In this way, the Latent Semantic Alignment Learning method establishes a consistent perception of the environment and makes the agent's actions easier to explain. Experiments on common benchmarks verify the effectiveness of our proposed methods. For example, we improve the Success Rate by 1.6% on the R2R validation unseen set and 4.3% on the R4R validation unseen set over the baseline model. Siying Wu, Xueyang Fu, Feng Wu 0005, Zhengjun Zha |
IEEE Trans. Multim. | 2 |
| 2024 | Hue Guidance Network for Single Image Reflection RemovalabstractReflection from glasses is ubiquitous in daily life, but it is usually undesirable in photographs. To remove these unwanted noises, existing methods utilize either correlative auxiliary information or handcrafted priors to constrain this ill-posed problem. However, due to their limited capability to describe the properties of reflections, these methods are unable to handle strong and complex reflection scenes. In this article, we propose a hue guidance network (HGNet) with two branches for single image reflection removal (SIRR) by integrating image information and corresponding hue information. The complementarity between image information and hue information has not been noticed. The key to this idea is that we found that hue information can describe reflections well and thus can be used as a superior constraint for the specific SIRR task. Accordingly, the first branch extracts the salient reflection features by directly estimating the hue map. The second branch leverages these effective features, which can help locate salient reflection regions to obtain a high-quality restored image. Furthermore, we design a new cyclic hue loss to provide a more accurate optimization direction for the network training. Experiments substantiate the superiority of our network, especially its excellent generalization ability to various reflection scenes, as compared with state-of-the-arts both qualitatively and quantitatively. Source codes are available at https://github.com/zhuyr97/HGRR. Yurui Zhu, Xueyang Fu, Zheyu Zhang 0002, Aiping Liu, Zhiwei Xiong, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Event-Guided Person Re-Identification via Sparse-Dense Complementary LearningabstractVideo-based person reidentification (Re-ID) is a prominent computer vision topic due to its wide range of video surveillance applications. Most existing methods utilize spatial and temporal correlations in frame sequences to obtain discriminative person features. However, inevitable degradation, e.g., motion blur contained in frames, leading to the loss of identity-discriminating cues. Recently, a new bio-inspired sensor called event camera, which can asynchronously record intensity changes, brings new vitality to the Re-ID task. With the microsecond resolution and low latency, it can accurately capture the movements of pedestrians even in the degraded environments. In this work, we propose a Sparse-Dense Complementary Learning (SDCL) Framework, which effectively extracts identity features by fully exploiting the complementary information of dense frames and sparse events. Specifically, for frames, we build a CNN-based module to aggregate the dense features of pedestrian appearance step by step, while for event streams, we design a bio-inspired spiking neural network (SNN) backbone, which encodes event signals into sparse feature maps in a spiking form, to extract the dynamic motion cues of pedestrians. Finally, a cross feature alignment module is constructed to fuse motion information from events and appearance cues from frames to enhance identity representation learning. Experiments on several benchmarks show that by employing events and SNN into Re-ID, our method significantly outperforms competitive methods. The code is available at https://github.com/Chengzhi-Cao/SDCL. Chengzhi Cao, Xueyang Fu, Hongjian Liu, Jiebo Luo 0001, Zhengjun Zha |
CVPR | 2 |
| 2023 | Edge-aware Regional Message Passing Controller for Image Forgery LocalizationabstractDigital image authenticity has promoted research on image forgery localization. Although deep learning-based methods achieve remarkable progress, most of them usually suffer from severe feature coupling between the forged and authentic regions. In this work, we propose a two-step Edge-aware Regional Message Passing Controlling strategy to address the above issue. Specifically, the first step is to account for fully exploiting the edge information. It consists of two core designs: context-enhanced graph construction and threshold-adaptive differentiable binarization edge algorithm. The former assembles the global semantic information to distinguish the features between the forged and authentic regions, while the latter stands on the output of the former to provide the learnable edges. In the second step, guided by the learnable edges, a region message passing controller is devised to weaken the message passing between the forged and authentic regions. In this way, our ERMPC is capable of explicitly modeling the inconsistency between the forged and authentic regions and enabling it to perform well on refined forged images. Extensive experiments on several challenging benchmarks show that our method is superior to state-of-the-art image forgery localization methods qualitatively and quantitatively. Dong Li 0055, Jiaying Zhu, Menglu Wang 0003, Jiawei Liu 0001, Xueyang Fu, Zhengjun Zha |
CVPR | 5 |
| 2023 | Generalized UAV Object Detection via Frequency Domain DisentanglementabstractWhen deploying the Unmanned Aerial Vehicles object detection (UAV-OD) network to complex and unseen real-world scenarios, the generalization ability is usually reduced due to the domain shift. To address this issue, this paper proposes a novel frequency domain disentanglement method to improve the UAV-OD generalization. Specifically, we first verified that the spectrum of different bands in the image has different effects to the UAV-OD generalization. Based on this conclusion, we design two learnable filters to extract domain-invariant spectrum and domain-specific spectrum, respectively. The former can be used to train the UAV-OD network and improve its capacity for generalization. In addition, we design a new instance-level contrastive loss to guide the network training. This loss enables the network to concentrate on extracting domaininvariant spectrum and domain-specific spectrum, so as to achieve better disentangling results. Experimental results on three unseen target domains demonstrate that our method has better generalization ability than both the baseline method and state-of-the-art methods. Xueyang Fu, Chengzhi Cao, Gege Shi, Zhengjun Zha |
CVPR | 2 |
| 2023 | Learning Weather-General and Weather-Specific Features for Image Restoration Under Multiple Adverse Weather ConditionsabstractImage restoration under multiple adverse weather conditions aims to remove weather-related artifacts by using a single set of network parameters. In this paper, we find that image degradations under different weather conditions contain general characteristics as well as their specific characteristics. Inspired by this observation, we design an efficient unified framework with a two-stage training strategy to explore the weather-general and weather-specific features. The first training stage aims to learn the weather-general features by taking the images under various weather conditions as inputs and outputting the coarsely restored results. The second training stage aims to learn to adaptively expand the specific parameters for each weather type in the deep model, where the requisite positions for expanding weather-specific parameters are automatically learned. Hence, we can obtain an efficient and unified model for image restoration under multiple adverse weather conditions. Moreover, we build the first real-world benchmark dataset with multiple weather conditions to better deal with realworld weather scenarios. Experimental results show that our method achieves superior performance on all the synthetic and real-world benchmarks. Codes and datasets are available at this repository. Yurui Zhu, Tianyu Wang 0003, Xueyang Fu, Xuanyu Yang, Xin Guo 0018, Jifeng Dai, Yu Qiao 0001, Xiaowei Hu 0001 |
CVPR | 3 |
| 2023 | Random Shuffle Transformer for Image RestorationabstractNon-local interactions play a vital role in boosting performance for image restoration. However, local window Transformer has been preferred due to its efficiency for processing high-resolution images. The superiority in efficiency comes at the cost of sacrificing the ability to model non-local interactions. In this paper, we present that local window Transformer can also function as modeling non-local interactions. The counterintuitive function is based on the permutation-equivariance of self-attention. The basic principle is quite simple: by *randomly shuffling* the input, local self-attention also has the potential to model non-local interactions without introducing extra parameters. Our random shuffle strategy enjoys elegant theoretical guarantees in extending the local scope. The resulting Transformer dubbed *ShuffleFormer* is capable of processing high-resolution images efficiently while modeling non-local interactions. Extensive experiments demonstrate the effectiveness of ShuffleFormer across a variety of image restoration tasks, including image denoising, deraining, and deblurring. Code is available at https://github.com/jiexiaou/ShuffleFormer. Jie Xiao 0002, Xueyang Fu, Man Zhou 0003, Hongjian Liu, Zhengjun Zha |
ICML | 2 |
| 2023 | Accurate MRI Reconstruction via Multi-Domain Recurrent NetworksabstractIn recent years, deep convolutional neural networks (CNNs) have become dominant in MRI reconstruction from undersampled k-space. However, most existing CNNs methods reconstruct the undersampled images either in the spatial domain or in the frequency domain, and neglecting the correlation between these two domains. This hinders the further reconstruction performance improvement. To tackle this issue, in this work, we propose a new multi-domain recurrent network (MDR-Net) with multi-domain learning (MDL) blocks as its basic units to reconstruct the undersampled MR image progressively. Specifically, the MDL block interactively processes the local spatial features and the global frequency information to facilitate complementary learning, leading to fine-grained features generation. Furthermore, we introduce an effective frequency-based loss to narrow the frequency spectrum gap, compensating for over-smoothness caused by the widely used spatial reconstruction loss. Extensive experiments on public fastMRI datasets demonstrate that our MDR-Net consistently outperforms other competitive methods and is able to provide more details. Jinbao Wei, Kongqiao Wang, Xueyang Fu, Xun Chen 0001 |
IJCAI | 5 |
| 2023 | Alleviating Spatial Misalignment and Motion Interference for UAV-based Video RecognitionabstractRecognizing activities with Unmanned Aerial Vehicles (UAVs) is essential for many applications, while existing video recognition methods are mainly designed for ground cameras and do not account for UAV changing attitudes and fast motion. This creates spatial misalignment of small objects between frames, leading to inaccurate visual movement in drone videos. Additionally, camera motion relative to objects in the video causes relative movements that visually affect object motion and can result in misunderstandings of video content. To address these issues, we present a novel framework named Attentional Spatial and Adaptive Temporal Relations Modeling. First, to mitigate the spatial misalignment of small objects between frames, we design an Attentional Patch-level Spatial Enrichment (APSE) module that models dependencies among patches and enhances patch-level features. Then, we propose a Multi-scale Temporal and Spatial Mixer (MTSM) module that is capable of adapting to disturbances caused by the UAV flight and modeling various temporal clues. By integrating APSE and MTSM into a single model, our network can effectively and accurately capture spatiotemporal relations for UAV videos. Extensive experiments on several benchmarks demonstrate the superiority of our method over state-of-the-art approaches. For instance, our network achieves a classification accuracy of 68.1% with an absolute gain of 1.3% compared to FuTH-Net on the ERA dataset. Gege Shi, Xueyang Fu, Chengzhi Cao, Zhengjun Zha |
ACM Multimedia | 2 |
| 2023 | Continual Image Deraining With Hypergraph Convolutional NetworksabstractImage deraining is a challenging task since rain streaks have the characteristics of a spatially long structure and have a complex diversity. Existing deep learning-based methods mainly construct the deraining networks by stacking vanilla convolutional layers with local relations, and can only handle a single dataset due to catastrophic forgetting, resulting in a limited performance and insufficient adaptability. To address these issues, we propose a new image deraining framework to effectively explore nonlocal similarity, and to continuously learn on multiple datasets. Specifically, we first design a patchwise hypergraph convolutional module, which aims to better extract the nonlocal properties with higher-order constraints on the data, to construct a new backbone and to improve the deraining performance. Then, to achieve better generalizability and adaptability in real-world scenarios, we propose a biological brain-inspired continual learning algorithm. By imitating the plasticity mechanism of brain synapses during the learning and memory process, our continual learning process allows the network to achieve a subtle stability-plasticity tradeoff. This it can effectively alleviate catastrophic forgetting and enables a single network to handle multiple datasets. Compared with the competitors, our new deraining network with unified parameters attains a state-of-the-art performance on seen synthetic datasets and has a significantly improved generalizability on unseen real rainy images. Xueyang Fu, Jie Xiao 0002, Yurui Zhu, Aiping Liu, Feng Wu 0001, Zhengjun Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Image De-Raining TransformerabstractExisting deep learning based de-raining approaches have resorted to the convolutional architectures. However, the intrinsic limitations of convolution, including local receptive fields and independence of input content, hinder the model's ability to capture long-range and complicated rainy artifacts. To overcome these limitations, we propose an effective and efficient transformer-based architecture for the image de-raining. First, we introduce general priors of vision tasks, i.e., locality and hierarchy, into the network architecture so that our model can achieve excellent de-raining performance without costly pre-training. Second, since the geometric appearance of rainy artifacts is complicated and of significant variance in space, it is essential for de-raining models to extract both local and non-local features. Therefore, we design the complementary window-based transformer and spatial transformer to enhance locality while capturing long-range dependencies. Besides, to compensate for the positional blindness of self-attention, we establish a separate representative space for modeling positional relationship, and design a new relative position enhanced multi-head self-attention. In this way, our model enjoys powerful abilities to capture dependencies from both content and position, so as to achieve better image content recovery while removing rainy artifacts. Experiments substantiate that our approach attains more appealing results than state-of-the-art methods quantitatively and qualitatively. Jie Xiao 0002, Xueyang Fu, Aiping Liu, Feng Wu 0001, Zhengjun Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Low-Light Stereo Image EnhancementabstractStereo cameras are now commonly used in more and more devices. Nevertheless, visually unpleasant images captured under low-light conditions hinder their practical application. As an initial attempt at low-light stereo image enhancement, we propose a novel Dual-View Enhancement Network (DVENet) based on the Retinex theory, which consists of two stages. The first stage estimates an illumination map to obtain a coarse enhancement result, which boosts the correlation of two views, while the second stage recovers details by integrating the information from two views to achieve fine image quality improvement with the guidance of the illumination map. To fully utilize the dual-view correlation, we further design a wavelet-based view transfer module to efficiently carry out multi-scale detail recovery. Then, we design an illumination-aware attention fusion module to exploit the complementarity between the fused features from two views and the single-view features. Experiments on both synthetic and real-world stereo datasets demonstrate the superiority of our proposed method over existing solutions. The code and model are publicly available at:https://github.com/KevinJ-Huang/Stereo-Low-Light. Jie Huang 0017, Xueyang Fu, Zeyu Xiao 0002, Feng Zhao 0004, Zhiwei Xiong |
IEEE Trans. Multim. | 2 |
| 2023 | Unsupervised Underexposed Image Enhancement via Self-Illuminated and Perceptual GuidanceabstractUnderexposed images inevitably suffer severe degradation due to light distortion and noise corruption. Motivated by the limited samples of paired datasets, several unsupervised enhancement methods have been developed. However, these techniques heavily rely on pre-defined fixed lightness and noise removal constraints. Correspondingly, they cannot match the image-specific lightness when performing enhancement and can only refine details in a non-perceptual way. In this paper, we propose an Unsupervised Underexposed Image Enhancement Network (U2IENet) with self-illuminated and perceptual guidance. Specifically, to adjust the illumination for matching the image-specific lightness adaptively, we utilize the bright area of the underexposed image as the self-illuminated guidance to constrain the training process and modulate the features. Meanwhile, we introduce the perceptual guidance as a constraint to remove the noise based on illumination distribution, thus refining the details perceptually. Experiments on both underexposed datasets and public low-light datasets demonstrate the superiority of the proposed approach with higher flexibility over state-of- the-art solutions. In addition, our U2IENet also provides a side function that enables users to adjust the lightness via interactive tuning of a single parameter. Naishan Zheng, Jie Huang 0017, Feng Zhao 0004, Xueyang Fu, Feng Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Enhanced Deep Blind Hyperspectral Image FusionabstractThe goal of hyperspectral image fusion (HIF) is to reconstruct high spatial resolution hyperspectral images (HR-HSI) via fusing low spatial resolution hyperspectral images (LR-HSI) and high spatial resolution multispectral images (HR-MSI) without loss of spatial and spectral information. Most existing HIF methods are designed based on the assumption that the observation models are known, which is unrealistic in many scenarios. To address this blind HIF problem, we propose a deep learning-based method that optimizes the observation model and fusion processes iteratively and alternatively during the reconstruction to enforce bidirectional data consistency, which leads to better spatial and spectral accuracy. However, general deep neural network inherently suffers from information loss, preventing us to achieve this bidirectional data consistency. To settle this problem, we enhance the blind HIF algorithm by making part of the deep neural network invertible via applying a slightly modified spectral normalization to the weights of the network. Furthermore, in order to reduce spatial distortion and feature redundancy, we introduce a Content-Aware ReAssembly of FEatures module and an SE-ResBlock model to our network. The former module helps to boost the fusion performance, while the latter make our model more compact. Experiments demonstrate that our model performs favorably against compared methods in terms of both nonblind HIF fusion and semiblind HIF fusion. Xueyang Fu, Weihong Zeng, Liyan Sun, Ronghui Zhan, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Learning to Model Pixel-Embedded Affinity for Homogeneous Instance SegmentationabstractHomogeneous instance segmentation aims to identify each instance in an image where all interested instances belong to the same category, such as plant leaves and microscopic cells. Recently, proposal-free methods, which straightforwardly generate instance-aware information to group pixels into different instances, have received increasing attention due to their efficient pipeline. However, they often fail to distinguish adjacent instances due to similar appearances, dense distribution and ambiguous boundaries of instances in homogeneous images. In this paper, we propose a pixel-embedded affinity modeling method for homogeneous instance segmentation, which is able to preserve the semantic information of instances and improve the distinguishability of adjacent instances. Instead of predicting affinity directly, we propose a self-correlation module to explicitly model the pairwise relationships between pixels, by estimating the similarity between embeddings generated from the input image through CNNs. Based on the self-correlation module, we further design a cross-correlation module to maintain the semantic consistency between instances. Specifically, we map the transformed input images with different views and appearances into the same embedding space, and then mutually estimate the pairwise relationships of embeddings generated from the original input and its transformed variants. In addition, to integrate the global instance information, we introduce an embedding pyramid module to model affinity on different scales. Extensive experiments demonstrate the versatile and superior performance of our method on three representative datasets. Code and models are available at https://github.com/weih527/Pixel-Embedded-Affinity. Wei Huang 0036, Shiyu Deng, Chang Chen 0004, Xueyang Fu, Zhiwei Xiong |
AAAI | 4 |
| 2022 | Efficient Model-Driven Network for Shadow RemovalabstractDeep Convolutional Neural Networks (CNNs) based methods have achieved significant breakthroughs in the task of single image shadow removal. However, the performance of these methods remains limited for several reasons. First, the existing shadow illumination model ignores the spatially variant property of the shadow images, hindering their further performance. Second, most deep CNNs based methods directly estimate the shadow free results from the input shadow images like a black box, thus losing the desired interpretability. To address these issues, we first propose a new shadow illumination model for the shadow removal task. This new shadow illumination model ensures the identity mapping among unshaded regions, and adaptively performs fine grained spatial mapping between shadow regions and their references. Then, based on the shadow illumination model, we reformulate the shadow removal task as a variational optimization problem. To effectively solve the variational problem, we design an iterative algorithm and unfold it into a deep network, naturally increasing the interpretability of the deep model. Experiments show that our method could achieve SOTA performance with less than half parameters, one-fifth of floating-point of operations (FLOPs), and over seventeen times faster than SOTA method (DHAN). Yurui Zhu, Zeyu Xiao 0002, Yanchi Fang, Xueyang Fu, Zhiwei Xiong, Zhengjun Zha |
AAAI | 4 |
| 2022 | Exposure Normalization and Compensation for Multiple-Exposure CorrectionabstractImages captured with improper exposures usually bring unsatisfactory visual effects. Previous works mainly focus on either underexposure or overexposure correction, resulting in poor generalization to various exposures. An alternative solution is to mix the multiple exposure data for training a single network. However, the procedures of correcting underexposure and overexposure to normal exposures are much different from each other, leading to large discrepancies for the network in correcting multiple-exposures, thus resulting in poor performance. The key point to address this issue lies in bridging different exposure representations. To achieve this goal, we design a multiple exposure correction framework based on an Exposure Normalization and Compensation (ENC) module. Specifically, the ENC module consists of an exposure normalization part for mapping different exposure features to the exposure-invariant feature space, and a compensation part for integrating the initial features unprocessed by the exposure normalization part to ensure the completeness of information. Besides, to further alleviate the imbalanced performance caused by variations in the optimization process, we introduce a parameter regularization fine-tuning strategy to improve the performance of the worst-performed exposure without degrading other exposures. Our model empowered by ENC outperforms the existing methods by more than 2dB and is robust to multiple image enhancement tasks, demonstrating its effectiveness and generalization capability for real-world applications. Code: https://github.com/KevinJ-Huang/ExposureNorm-Compensation. Jie Huang 0017, Xueyang Fu, Man Zhou 0003, Yang Wang 0015, Feng Zhao 0004, Zhiwei Xiong |
CVPR | 3 |
| 2022 | Memory-augmented Deep Conditional Unfolding Network for PansharpeningabstractPansharpening aims to obtain high-resolution multispectral (MS) images for remote sensing systems and deep learning-based methods have achieved remarkable success. However, most existing methods are designed in a black-box principle, lacking sufficient interpretability. Additionally, they ignore the different characteristics of each band of MS images and directly concatenate them with panchromatic (PAN) images, leading to severe copy artifacts [9]. To address the above issues, we propose an interpretable deep neural network, namely Memory-augmented Deep Conditional Unfolding Network with two specified core designs. Firstly, considering the degradation process, it formulates the Pansharpening problem as the minimization of a variational model with denoising-based prior and non-local auto-regression prior which is capable of searching the similarities between long-range patches, benefiting the texture enhancement. A novel iteration algorithm with built-in CNNs is exploited for transparent model design. Secondly, to fully explore the potentials of different bands of MS images, the PAN image is combined with each band of MS images, selectively providing the high-frequency details and alleviating the copy artifacts. Extensive experimental results validate the superiority of the proposed algorithm against other state-of-the-art methods. Man Zhou 0003, Aiping Liu, Xueyang Fu, Fan Wang 0005 |
CVPR | 5 |
| 2022 | Mutual Information-driven Pan-sharpeningabstractPan-sharpening aims to integrate the complementary information of texture-rich PAN images and multi-spectral (MS) images to produce the texture-rich MS images. Despite the remarkable progress, existing state-of-the-art Pansharpening methods don't explicitly enforce the complementary information learning between two modalities of PAN and MS images. This leads to information redundancy not being handled well, which further limits the performance of these methods. To address the above issue, we propose a novel mutual information-driven Pan-sharpening framework in this paper. To be specific, we first project the PAN and MS image into modality-aware feature space independently, and then impose the mutual information minimization over them to explicitly encourage the complementary information learning. Such operation is capable of reducing the information redundancy and improving the model performance. Extensive experimental results over multiple satellite datasets demonstrate that the proposed algorithm outperforms other state-of-the-art methods qualitatively and quantitatively with great generalization ability to real-world scenes. Man Zhou 0003, Jie Huang 0017, Zihe Yang, Xueyang Fu, Feng Zhao 0004 |
CVPR | 5 |
| 2022 | Bijective Mapping Network for Shadow RemovalabstractShadow removal, which aims to restore the background in the shadow regions, is challenging due to its highly ill-posed nature. Most existing deep learning-based methods individually remove the shadow by only considering the content of the matched paired images, barely taking into account the auxiliary supervision of shadow generation in the shadow removal procedure. In this work, we argue that shadow removal and generation are interrelated and could provide useful informative supervision for each other. Specifically, we propose a new Bijective Mapping Network (BMNet), which couples the learning procedures of shadow removal and shadow generation in a unified parameter-shared framework. With consistent two way constraints and synchronous optimization of the two procedures, BMNet could effectively recover the underlying background contents during the forward shadow removal procedure. In addition, through statistical analysis of real world datasets, we observe and verify that shadow appearances under different color spectrums are inconsistent. This motivates us to design a Shadow-Invariant Color Guidance Module (SICGM), which can explicitly utilize the learned shadow-invariant color information to guide network color restoration, thereby further reducing color-bias effects. Experiments on the representative ISTD, ISTD+ and SRD benchmarks show that our proposed network outperforms the state-of-the-art method [11] in de-shadowing performance, while only using its 0.25% network parameters and 6.25% floating point operations (FLOPs). Yurui Zhu, Jie Huang 0017, Xueyang Fu, Feng Zhao 0004, Qibin Sun, Zhengjun Zha |
CVPR | 3 |
| 2022 | Dreaming to Prune Image Deraining NetworksabstractConvolutional image deraining networks have achieved great success while suffering from tremendous computational and memory costs. Most model compression methods require original data for iterative fine-tuning, which is limited in real-world applications due to storage, privacy, and transmission constraints. We note that it is overstretched to fine-tune the compressed model using self-collected data, as it exhibits poor generalization over images with different degradation characteristics. To address this problem, we propose a novel data-free compression framework for de-raining networks. It is based on our observation that deep degradation representations can be clustered by degradation characteristics (types of rain) while independent of image content. Therefore, in our framework, we “dream” diverse in-distribution degraded images using a deep inversion paradigm, thus leveraging them to distill the pruned model. Specifically, we preserve the performance of the pruned model in a dual-branch way. In one branch, we invert the pre-trained model (teacher) to reconstruct the degraded inputs that resemble the original distribution and employ the orthogonal regularization for deep features to yield degradation diversity. In the other branch, the pruned model (student) is distilled to fit the teacher's original statistical modeling on these dreamed inputs. Further, an adaptive pruning scheme is proposed to determine the hierarchical sparsity, which alleviates the regression drift of the initial pruned model. Experiments on various deraining datasets demonstrate that our method can reduce about 40% FLOPs of the state-of-the-art models while maintaining comparable performance without original data. Weiqi Zou, Yang Wang 0015, Xueyang Fu, Yang Cao 0010 |
CVPR | 3 |
| 2022 | JPEG Artifacts Removal via Contrastive Representation Learning
Xi Wang 0018, Xueyang Fu, Yurui Zhu, Zhengjun Zha |
ECCV (17) | 2 |
| 2022 | Spatial-Frequency Domain Information Integration for Pan-Sharpening
Man Zhou 0003, Jie Huang 0017, Hu Yu 0001, Xueyang Fu, Aiping Liu, Xian Wei, Feng Zhao 0004 |
ECCV (18) | 5 |
| 2022 | Event-driven Video Deblurring via Spatio-Temporal Relation-Aware NetworkabstractVideo deblurring with event information has attracted considerable attention. To help deblur each frame, existing methods usually compress a specific event sequence into a feature tensor with the same size as the corresponding video. However, this strategy neither considers the pixel-level spatial brightness changes nor the temporal correlation between events at each time step, resulting in insufficient use of spatio-temporal information. To address this issue, we propose a new Spatio-Temporal Relation-Attention network (STRA), for the specific event-based video deblurring. Concretely, to utilize spatial consistency between the frame and event, we model the brightness changes as an extra prior to aware blurring contexts in each frame; to record temporal relationship among different events, we develop a temporal memory block to restore long-range dependencies of event sequences continuously. In this way, the complementary information contained in the events and frames, as well as the correlation of neighboring events, can be fully utilized to recover spatial texture from events constantly. Experiments show that our STRA significantly outperforms several competing methods, e.g., on the HQF dataset, our network achieves up to 1.3 dB in terms of PSNR over the most advanced method. The code is available at https://github.com/Chengzhi-Cao/STRA. Chengzhi Cao, Xueyang Fu, Yurui Zhu, Gege Shi, Zhengjun Zha |
IJCAI | 2 |
| 2022 | Exploring Fourier Prior for Single Image Rain RemovalabstractDeep convolutional neural networks (CNNs) have become dominant in the task of single image rain removal. Most of current CNN methods, however, suffer from the problem of overfitting on one single synthetic dataset as they neglect the intrinsic prior of the physical properties of rain streaks. To address this issue, we propose a simple but effective prior - Fourier prior to improve the generalization ability of an image rain removal model. The Fourier prior is a kind of property of rainy images. It is based on a key observation of us - replacing the Fourier amplitude of rainy images with that of clean images greatly suppresses the synthetic and real-world rain streaks. This means the amplitude contains most of the rain streak information and the phase keeps the similar structures of the background. So it is natural for single image rain removal to process the amplitude and phase information of the rainy images separately. In this paper, we develop a two-stage model where the first stage restores the amplitude of rainy images to clean rain streaks, and the second stage restores the phase information to refine fine-grained background structures. Extensive experiments on synthetic rainy data demonstrate the power of Fourier prior. Moreover, when trained on synthetic data, a robust generalization ability to real-world images can also be obtained. The code will be publicly available at https://github.com/willinglucky/ExploringFourier-Prior-for-Single-Image-Rain-Removal. Xin Guo 0018, Xueyang Fu, Man Zhou 0003, Zhen Huang 0007, Jialun Peng, Zhengjun Zha |
IJCAI | 2 |
| 2022 | JPEG Compression-aware Image Forgery LocalizationabstractImage forgery localization, which aims to find suspicious regions tampered with splicing, copy-move or removal manipulations, has attracted increasing attention. Existing image forgery localization methods have made great progress on public datasets. However, these methods suffer a severe performance drop when the forged images are JPEG compressed, which is widely applied in social media transmission. To tackle this issue, we propose a wavelet-based compression representation learning scheme for the specific JPEG-resistant image forgery localization. Specifically, to improve the performance against JPEG compression, we first learn the abstract representations to distinguish various compression levels through wavelet integrated contrastive learning strategy. Then, based on the learned representations, we introduce a JPEG compression-aware image forgery localization network to flexibly handle forged images compressed with various JPEG quality factors. Moreover, a boundary correction branch is designed to alleviate the edge artifacts caused by JPEG compression. Extensive experiments demonstrate the superiority of our method to existing state-of-the-art approaches, not only on standard datasets, but also on the JPEG forged images with multiple compression quality factors. Menglu Wang 0003, Xueyang Fu, Jiawei Liu 0001, Zhengjun Zha |
ACM Multimedia | 2 |
| 2022 | Learning Dual Convolutional Dictionaries for Image De-rainingabstractRain removal is a vital and highly ill-posed low-level vision task. While currently existing deep convolutional neural networks (CNNs) based image de-raining methods have achieved remarkable results, they still possess apparent shortcomings: First, most of the CNNs based models are lack of interpretability. Second, these models are not embedded with physical structures of rain streaks and background images. Third, they omit useful information in the background images. These deficiencies result in unsatisfied de-raining results in some sophisticated scenarios. To solve the above problems, we propose a Deep Dual Convolutional Dictionary Learning Network (DDCDNet) for these specific tasks. We firstly propose a new dual dictionary learning objective function, and then unfold it into the form of neural networks to learn prior knowledge from the data automatically. This network tries to learn the rain-streaks layer and the clean background using two dictionary learning networks instead of merely predicting the rain-streaks layer like most of the de-raining methods. To further increase the interpretability and generalization capability, we add sparsity and adaptive dictionary to our network to generate dynamic dictionary for each image based on content. Experimental results reveal that our model possesses outstanding de-raining ability on both synthetic and real-world data sets in terms of PSNR and SSIM as well as visual appearance. Chengjie Ge, Xueyang Fu, Zhengjun Zha |
ACM Multimedia | 2 |
| 2022 | Cross-modal Semantic Alignment Pre-training for Vision-and-Language NavigationabstractVision-and-Language Navigation needs an agent to navigate to a target location by progressively grounding and following the relevant instruction conditioning on its memory and current observation. Existing works utilize the cross-modal transformer to pass the message between visual modality and textual modality. However, they are still limited to mining the fine-grained matching between the underlying components of trajectories and instructions. Inspired by the significant progress achieved by large-scale pre-training methods, in this paper, we propose CSAP, a new method of Cross-modal Semantic Alignment Pre-training for Vision-and-Language Navigation. It is designed to learn the alignment from trajectory-instruction pairs through two novel tasks, including trajectory-conditioned masked fragment modeling and contrastive semantic-alignment modeling. Specifically, the trajectory-conditioned masked fragment modeling encourages the agent to extract useful visual information to reconstruct the masked fragment. The contrastive semantic-alignment modeling is designed to align the visual representation with corresponding phrase embeddings. By showing experimental results on the benchmark dataset, we demonstrate that transformer architecture-based navigation agent pre-trained with our proposed CSAP outperforms existing methods on both SR and SPL scores. Siying Wu, Xueyang Fu, Feng Wu 0001, Zhengjun Zha |
ACM Multimedia | 2 |
| 2022 | Single Image Shadow Detection via Complementary MechanismabstractIn this paper, we present a novel shadow detection framework by investigating the mutual complementary mechanisms contained in this specific task. Our method is based on a key observation: in a single shadow image, shadow regions and non-shadow counterparts are complementary to each other in nature, thus a better estimation on one side leads to an improved estimation on the other, and vice versa. Motivated by this observation, we first leverage two parallel interactive branches to jointly produce shadow and non-shadow masks. The interaction between two parallel branches is to retain the deactivated intermediate features of one branch by introducing the negative activation technique, which could serve as complementary features to the other branch. Besides, we also apply identity reconstruction loss as complementary training guidance at the image level. Finally, we design two discriminative losses to satisfy the complementary requirements of shadow detection, i.e., neither missing any shadow regions nor falsely detecting non-shadow regions. By fully exploring and exploiting the complementary mechanism of shadow detection, our method can confidently predict more accurate shadow detection results. Extensive experiments on the three widely-used benchmarks demonstrate our proposed method achieves superior shadow detection performance against state-of-the-art methods with a relatively low computational cost. Yurui Zhu, Xueyang Fu, Chengzhi Cao, Xi Wang 0018, Qibin Sun, Zhengjun Zha |
ACM Multimedia | 2 |
| 2022 | Stochastic Window Transformer for Image RestorationabstractThanks to the powerful representation capabilities, transformers have made impressive progress in image restoration. However, existing transformers-based methods do not carefully consider the particularities of image restoration. In general, image restoration requires that an ideal approach should be translation-invariant to the degradation, i.e., the undesirable degradation should be removed irrespective of its position within the image. Furthermore, the local relationships also play a vital role, which should be faithfully exploited for recovering clean images. Nevertheless, most transformers either adopt local attention with the fixed local window strategy or global attention, which unfortunately breaks the translation invariance and causes huge loss of local relationships. To address these issues, we propose an elegant stochastic window strategy for transformers. Specifically, we first introduce the window partition with stochastic shift to replace the original fixed window partition for training. Then, we design a new layer expectation propagation algorithm to efficiently approximate the expectation of the induced stochastic transformer for testing. Our stochastic window transformer not only enjoys powerful representation but also maintains the desired property of translation invariance and locality. Experiments validate the stochastic window strategy consistently improves performance on various image restoration tasks (deraining, denoising and deblurring) by significant margins. The code is available at https://github.com/jiexiaou/Stoformer. Jie Xiao 0002, Xueyang Fu, Feng Wu 0001, Zhengjun Zha |
NeurIPS | 2 |
| 2022 | Learning Degradation-Invariant Representation for Robust Real-World Person Re-Identification
Xueyang Fu, Liang Li 0003, Zhengjun Zha |
Int. J. Comput. Vis. | 2 |
| 2022 | Progressive Pan-Sharpening via Cross-Scale Collaboration NetworksabstractPan-sharpening aims to produce a high-quality image by fusing a low-resolution multispectral (LRMS) image and a high-resolution panchromatic (PAN) image. Although deep-learning-based methods have dominated the pan-sharpening, they fail to fully utilize spatial and spectral information from the cross-scale perspective. In this work, we propose a novel cross-scale collaboration network to achieve accurate pan-sharpening. Specifically, we first design a progressive framework in a pyramid fashion to achieve a gradual pan-sharpening process, which consists of several subnetworks to handle specific pyramid levels. Then, in each subnetwork, we deploy two cross-scale attention modules to, respectively, capture global and local spatial interactions from MS and PAN images. To better utilize spectral information from different pyramid levels, we further deploy a fusion module between subnetworks to extract cross-scale representations. Through the above-mentioned cross-scale collaboration, our network can fully consider the cross-scale nature of MS and PAN images to better adapt to this specific task. Extensive experiments confirmed that our method outperforms several state-of-the-art (SOTA) methods both qualitatively and quantitatively. Zihe Yang, Xueyang Fu, Aiping Liu, Zhengjun Zha |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Twice Mixing: A rank learning based quality assessment approach for underwater image enhancement
Zhenqi Fu, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
Signal Process. Image Commun. | 2 |
| 2022 | PanCSC-Net: A Model-Driven Deep Unfolding Method for PansharpeningabstractRecently, deep learning (DL) approaches have been widely applied to the pansharpening problem, which is defined as fusing a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image to obtain a high-resolution multispectral (HRMS) image. However, most DL-based methods handle this task by designing black-box network architectures to model the mapping relationship from LRMS and PAN to HRMS. These network architectures always lack sufficient interpretability, which limits their further performance improvements. To address this issue, we adopt the model-driven method to design an interpretable deep network structure for pansharpening. First, we present a new pansharpening model using the convolutional sparse coding (CSC), which is quite different from the current pansharpening frameworks. Second, an alternative algorithm is developed to optimize this model. This algorithm is further unfolded to a network, where each network module corresponds to a specific operation of the iterative algorithm. Therefore, the proposed network has clear physical interpretations, and all the learnable modules can be automatically learned in an end-to-end way from the given dataset. Experimental results on some benchmark datasets show that our network performs better than other advanced methods both quantitatively and qualitatively. Xiangyong Cao, Xueyang Fu, Danfeng Hong, Zongben Xu, Deyu Meng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Deep Spatial-Spectral Global Reasoning Network for Hyperspectral Image DenoisingabstractAlthough deep neural networks (DNNs) have been widely applied to hyperspectral image (HSI) denoising, most DNN-based HSI denoising methods are designed by stacking convolution layer, which can only model and reason local relations, and thus ignore the global contextual information. To address this issue, we propose a deep spatial-spectral global reasoning network to consider both the local and global information for HSI noise removal. Specifically, two novel modules are proposed to model and reason global relational information. The first one aims to model global spatial relations between pixels in feature maps, and the second one models the global relations across the channels. Compared to traditional convolution operations, the two proposed modules enable the network to extract representations from new dimensions. For the HSI denoising task, the two modules, as well as the densely connected structures, are embedded into the U-Net architecture. Thus, the new-designed global reasoning network can help tackle complex noise by exploiting multiple representations, e.g., hierarchical local feature, global spatial coherence, cross-channel correlation, and multi-scale abstract representation. Experiments on both synthetic and real HSI data demonstrate that our proposed network can obtain comparable or even better denoising results than other state-of-the-art methods. Xiangyong Cao, Xueyang Fu, Chen Xu 0007, Deyu Meng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Effective Pan-Sharpening With Transformer and Invertible Neural NetworkabstractIn remote sensing imaging systems, pan-sharpening is an important technique to obtain high-resolution multispectral images from a high-resolution panchromatic image and its corresponding low-resolution multispectral image. Due to the powerful learning capability of convolution neural networks (CNNs), CNN-based methods have dominated this field. However, due to the limitation of the convolution operator, long-range spatial features are often not accurately obtained, thus limiting the overall performance. To this end, we propose a novel and effective method by exploiting a customized transformer architecture and information-lossless invertible neural module for long-range dependencies modeling and effective feature fusion in this article. Specifically, the customized transformer formulates the panchromatic (PAN) and multispectral (MS) features as queries and keys to encourage joint feature learning across two modalities, while the designed invertible neural module enables effective feature fusion to generate the expected pan-sharpened results. To the best of our knowledge, this is the first attempt to introduce a transformer and a invertible neural network into the pan-sharpening field. Extensive experiments over different kinds of satellite datasets demonstrate that our method outperforms state-of-the-art algorithms both visually and quantitatively with fewer parameters and flops. Furthermore, the ablation experiments also prove the effectiveness of the proposed customized long-range transformer and effective invertible neural feature fusion module for pan-sharpening. Man Zhou 0003, Xueyang Fu, Jie Huang 0017, Feng Zhao 0004, Aiping Liu, Rujing Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Effective Pan-Sharpening by Multiscale Invertible Neural Network and Heterogeneous Task DistillingabstractAs recognized, the ground truth multi-spectral (MS) images possess the complementary information (e.g., high-frequency component) of low-resolution (LR) MS images, which can be considered as privileged information to alleviate the spectral distortion and insufficient spatial texture enhancement. Since existing supervised pan-sharpening methods only utilize the ground truth MS image to supervise the network training, its potential value has not been fully explored. To accomplish this, we propose a heterogeneous knowledge-distilling pan-sharpening framework that distills pan-sharpening by imitating the ground truth reconstruction task in both the feature space and network output. In our work, the teacher network performs as a variational auto-encoder to extract effective features of the ground truth MS. The student network, acting as pan-sharpening, is trained by the assistance of the teacher network with the process-oriented feature imitation learning. Moreover, we design a customized information-lossless multi-scale invertible neural module to effectively fuse LR-MS and panchromatic (PAN) images, producing expected pan-sharpened results. To reduce the artifacts generated by the knowledge distillation process, a knowledge-driven refinement sub-network is further devised according to the pan-sharpening imaging model. Extensive experimental results on different satellite datasets validate that the proposed network outperforms the state-of-the-art methods both visually and quantitatively. The source code will be released at https://github.com/manman1995/pansharpening. Man Zhou 0003, Jie Huang 0017, Xueyang Fu, Feng Zhao 0004, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Joint Constrained CCA Model for Network-Dependent Brain Subregion ParcellationabstractConnectivity-based brain region parcellation from functional magnetic resonance imaging (fMRI) data is complicated by heterogeneity among aged and diseased subjects, particularly when the data are spatially transformed to a common space. Here, we propose a group-guided functional brain region parcellation model capable of obtaining subregions from a target region with consistent connectivity profiles across multiple subjects, even when the fMRI signals are kept in their native spaces. The model is based on a joint constrained canonical correlation analysis (JC-CCA) method that achieves group-guided parcellation while allowing the data dimension of the parcellated regions for each subject to vary. We performed extensive experiments on synthetic and real data to demonstrate the superiority of the proposed model compared to other classical methods. When applied to fMRI data of subjects with and without Parkinson's disease (PD) to estimate the subregions in the Putamen, significant between-group differences were found in the derived subregions and the connectivity patterns. Superior classification and regression results were obtained, demonstrating its potential in clinical practice. Qinrui Ling, Aiping Liu, Yu Li 0027, Xueyang Fu, Xun Chen 0001, Martin J. McKeown, Feng Wu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | A Unified Deep Learning Framework for ssTEM Image RestorationabstractSerial section transmission electron micro-scopy (ssTEM) reveals biological information at a scale of nanometer and plays an important role in the ultrastructural analysis. However, due to the imperfect preparation of biological samples, ssTEM images are usually degraded with various artifacts that greatly challenge the subsequent analysis and visualization. In this paper, we introduce a unified deep learning framework for ssTEM image restoration which addresses three main types of artifacts, i.e., Support Film Folds (SFF), Staining Precipitates (SP), and Missing Sections (MS). To achieve this goal, we first model the appearance of SFF and SP artifacts by conducting comprehensive analyses on the statistics of real degraded images, relying on which we can then simulate a large number of paired images (degraded/artifacts-free) for training a deep restoration network. Then, we design a coarse-to-fine restoration network consisting of three modules, i.e., interpolation, correction, and fusion. The interpolation module exploits the adjacent artifacts-free images for an initial restoration, while the correction module resorts to the degraded image itself to rectify the artifacts. Finally, the fusion module jointly utilizes the above two results to further improve the restoration fidelity. Experimental results on both synthetic and real test data validate the significantly improved performance of our proposed framework over existing solutions, in terms of both image restoration fidelity and neuron segmentation accuracy. To the best of our knowledge, this is the first unified deep learning framework for ssTEM image restoration from different types of artifacts. Code is available at https://github.com/sydeng99/ssTEM-restoration. Shiyu Deng, Wei Huang 0036, Chang Chen 0004, Xueyang Fu, Zhiwei Xiong |
IEEE Trans. Medical Imaging | 4 |
| 2022 | A Model-Driven Deep Unfolding Method for JPEG Artifacts RemovalabstractDeep learning-based methods have achieved notable progress in removing blocking artifacts caused by lossy JPEG compression on images. However, most deep learning-based methods handle this task by designing black-box network architectures to directly learn the relationships between the compressed images and their clean versions. These network architectures are always lack of sufficient interpretability, which limits their further improvements in deblocking performance. To address this issue, in this article, we propose a model-driven deep unfolding method for JPEG artifacts removal, with interpretable network structures. First, we build a maximum posterior (MAP) model for deblocking using convolutional dictionary learning and design an iterative optimization algorithm using proximal operators. Second, we unfold this iterative algorithm into a learnable deep network structure, where each module corresponds to a specific operation of the iterative algorithm. In this way, our network inherits the benefits of both the powerful model ability of data-driven deep learning method and the interpretability of traditional model-driven method. By training the proposed network in an end-to-end manner, all learnable modules can be automatically explored to well characterize the representations of both JPEG artifacts and image content. Experiments on synthetic and real-world datasets show that our method is able to generate competitive or even better deblocking results, compared with state-of-the-art methods both quantitatively and qualitatively. Xueyang Fu, Menglu Wang 0003, Xiangyong Cao, Xinghao Ding, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Rain Streak Removal via Dual Graph Convolutional NetworkabstractDeep convolutional neural networks (CNNs) have become dominant in the single image de-raining area. However, most deep CNNs-based de-raining methods are designed by stacking vanilla convolutional layers, which can only be used to model local relations. Therefore, long-range contextual information is rarely considered for this specific task. To address the above problem, we propose a simple yet effective dual graph convolutional network (GCN) for single image rain removal. Specifically, we design two graphs to perform global relational modeling and reasoning. The first GCN is used to explore global spatial relations among pixels in feature maps, while the second GCN models the global relations across the channels. Compared to standard convolutional operations, the proposed two graphs enable the network to extract representations from new dimensions. To achieve the image rain removal, we further embed these two graphs and multi-scale dilated convolution into a symmetrically skip-connected network architecture. Therefore, our dual graph convolutional network is able to well handle complex and spatially long rain streaks by exploring multiple representations, e.g., multi-scale local feature, global spatial coherence and cross-channel correlation. Meanwhile, our model is easy to implement, end-to-end trainable and computationally efficient. Extensive experiments on synthetic and real data demonstrate that our method achieves significant improvements over the recent state-of-the-art methods. Xueyang Fu, Qi Qi 0005, Zhengjun Zha, Yurui Zhu, Xinghao Ding |
AAAI | 1 |
| 2021 | Space-Time Distillation for Video Super-ResolutionabstractCompact video super-resolution (VSR) networks can be easily deployed on resource-limited devices, e.g., smartphones and wearable devices, but have considerable performance gaps compared with complicated VSR networks that require a large amount of computing resources. In this paper, we aim to improve the performance of compact VSR networks without changing their original architectures, through a knowledge distillation approach that transfers knowledge from a complicated VSR network to a compact one. Specifically, we propose a space-time distillation (STD) scheme to exploit both spatial and temporal knowledge in the VSR task. For space distillation, we extract spatial attention maps that hint the high-frequency video content from both networks, which are further used for transferring spatial modeling capabilities. For time distillation, we narrow the performance gap between compact models and complicated models by distilling the feature similarity of the temporal memory cells, which are encoded from the sequence of feature maps generated in the training clips using ConvLSTM. During the training process, STD can be easily incorporated into any network without changing the original network architecture. Experimental results on standard benchmarks demonstrate that, in resource-constrained situations, the proposed method notably improves the performance of existing VSR networks without increasing the inference time. Zeyu Xiao 0002, Xueyang Fu, Jie Huang 0017, Zhen Cheng 0002, Zhiwei Xiong |
CVPR | 2 |
| 2021 | Image De-Raining via Continual LearningabstractWhile deep convolutional neural networks (CNNs) have achieved great success on image de-raining task, most existing methods can only learn fixed mapping rules between paired rainy/clean images on a single dataset. This limits their applications in practical situations with multiple and incremental datasets where the mapping rules may change for different types of rain streaks. However, the catastrophic forgetting of traditional deep CNN model challenges the design of generalized framework for multiple and incremental datasets. A strategy of sharing the network structure but in-dependently updating and storing the network parameters on each dataset has been developed as a potential solution. Nevertheless, this strategy is not applicable to compact systems as it dramatically increases the overall training time and parameter space. To alleviate such limitation, in this study, we propose a parameter importance guided weights modification approach, named PIGWM. Specifically, with new dataset (e.g. new rain dataset), the well-trained network weights are updated according to their importance evaluated on previous training dataset. With extensive experimental validation, we demonstrate that a single network with a single parameter set of our proposed method can process multiple rain datasets almost without performance degradation. The proposed model is capable of achieving superior performance on both inhomogeneous and incremental datasets, and is promising for highly compact systems to gradually learn myriad regularities of the different types of rain streaks. The results indicate that our proposed method has great potential for other computer vision tasks with dynamic learning environments. Man Zhou 0003, Jie Xiao 0002, Yifan Chang, Xueyang Fu, Aiping Liu, Jinshan Pan, Zhengjun Zha |
CVPR | 4 |
| 2021 | Learning Dual Priors for JPEG Compression Artifacts RemovalabstractDeep learning (DL)-based methods have achieved great success in solving the ill-posed JPEG compression artifacts removal problem. However, as most DL architectures are designed to directly learn pixel-level mapping relationship-s, they largely ignore semantic-level information and lack sufficient interpretability. To address the above issues, in this work, we propose an interpretable deep network to learn both pixel-level regressive prior and semantic-level discriminative prior. Specifically, we design a variational model to formulate the image de-blocking problem and propose two prior terms for the image content and gradient, respectively. The content-relevant prior is formulated as a DL-based image-to-image regressor to perform as a de-blocker from the pixel-level. The gradient-relevant prior serves as a DL-based classifier to distinguish whether the image is compressed from the semantic-level. To effectively solve the variational model, we design an alternating minimization algorithm and unfold it into a deep network architecture. In this way, not only the interpretability of the deep network is increased, but also the dual priors can be well estimated from training samples. By integrating the two priors into a single framework, the image de-blocking problem can be well-constrained, leading to a better performance. Experiments on benchmarks and real-world use cases demonstrate the superiority of our method to the existing state-of-the-art approaches. Xueyang Fu, Xi Wang 0018, Aiping Liu, Junwei Han 0001, Zhengjun Zha |
ICCV | 1 |
| 2021 | Attack-Guided Perceptual Data Generation for Real-world Re-IdentificationabstractIn unconstrained real-world surveillance scenarios, person re-identification (Re-ID) models usually suffer from different low-level perceptual variations, e.g., cross-resolution and insufficient lighting. Due to the limited variation range of training data, existing models are difficult to generalize to scenes with unknown perceptual interference types. To address the above problem, in this paper, we propose two disjoint data-generation ways to complement existing training samples to improve the robustness of Re-ID mod-els. Firstly, considering the sparsity and imbalance of samples in the perceptual space, a dense resampling method from the estimated perceptual distribution is performed. Secondly, to dig more representative generated samples for identity representation learning, we introduce a graph-based white-box attacker to guide the data generation process with intra-batch ranking and discriminate attention. In addition, two synthetic-to-real feature constraints are introduced into the Re-ID training to prevent the generated data from bringing domain bias. Our method is effective, easy-to-implement, and independent of the specific network architecture. Applying our approach to a ResNet-50 base-line can already achieve competitive results, surpassing state-of-the-art methods by +1.2% at Rank-1 on the MLR-CUHK03 dataset. Xueyang Fu, Zhengjun Zha |
ICCV | 2 |
| 2021 | Cross-Patch Graph Convolutional Network for Image DenoisingabstractRecently, deep learning-based image denoising methods have achieved significant improvements over traditional methods. Due to the hardware limitation, most deep learning-based image denoising methods utilize cropped small patches to train a convolutional neural network to infer the clean images. However, the real noisy images in practical are mostly of high resolution rather than the cropped small patches and the vanilla training strategies ignore the cross-patch contextual dependency in the whole image. In this paper, we propose Cross-Patch Net (CPNet), which is the first deep-learning-based real image denoising method for HR (high resolution) input. Furthermore, we design a novel loss guided by the noise level map to obtain better performance. Compared with the vanilla patch-based training strategies, our approach effectively exploits the cross-patch contextual dependency. Besides, owing to the difficulty in capturing real noisy and noise-free image paired training data, we propose an effective method to generate realistic sRGB noisy images from their corresponding clean sRGB images for denoiser training. Denoising experiments on real-world sRGB images show the effectiveness of the proposed method. More importantly, our method achieves state-of-the-art performance on practical sRGB noisy image denoising. Xueyang Fu, Zhengjun Zha |
ICCV | 2 |
| 2021 | Improving De-raining Generalization via Neural ReorganizationabstractMost existing image de-raining networks could only learn fixed mapping rules between paired rainy/clean images on single synthetic dataset and then stay static for lifetime. However, since single synthetic dataset merely provides a partial view for the distribution of rain streaks, deep models well trained on an individual synthetic dataset tend to overfit on this biased distribution. This leads to the inability of these methods to well generalize to complex and changeable real-world rainy scenes, thus limiting their practical applications. In this paper, we try for the first time to accumulate the de-raining knowledge from multiple synthetic datasets on a single network parameter set to improve the de-raining generalization of deep networks. To achieve this goal, we explore Neural Reorganization (NR) to allow the de-raining network to keep a subtle stability-plasticity trade-off rather than naive stabilization after training phase. Specifically, we design our NR algorithm by borrowing the synaptic consolidation mechanism in the biological brain and knowledge distillation. Equipped with our NR algorithm, the deep model can be trained on a list of synthetic rainy datasets by overcoming catastrophic forgetting, making it a general-version de-raining network. Extensive experimental validation shows that due to the successful accumulation of de-raining knowledge, our proposed method can not only process multiple synthetic datasets consistently, but also achieve state-of-the-art results when dealing with real-world rainy images. Jie Xiao 0002, Man Zhou 0003, Xueyang Fu, Aiping Liu, Zhengjun Zha |
ICCV | 3 |
| 2021 | Multifocal Attention-Based Cross-Scale Network for Image De-rainingabstractAlbeit existing deep learning-based image de-raining methods have achieved promising results, most of them only extract single scale features, and neglect the fact that similar rain streaks appear repeatedly across different scales. Therefore, this paper aims to explore the cross-scale cues in a multi-scale fashion. Specifically, we first introduce an adaptive-kernel pyramid to provide effective multi-scale information. Then, we design two cross-scale similarity attention blocks (CSSABs) to search spatial and channel relationships between two scales, respectively. The spatial CSSAB explores the spatial similarity between pixels of cross-scale features, while the channel CSSAB emphasizes the interdependencies among cross-scale features. To further improve the diversity of features, we adopt the wavelet transformation and multi-head mechanism in CSSABs to generate multifocal features which focus on different areas. Finally, based on our CSSABs, we construct an effective multifocal attention-based cross-scale network, which exhaustively utilizes the cross-scale correlations of both rain streaks and background, to achieve image de-raining. Experiments show the superiority of our network over state-of-the-art image de-raining approaches both qualitatively and quantitatively. The source code and pre-trained models are available at https://github.com/zhangzheyu0/Multifocal_derain. Zheyu Zhang 0002, Yurui Zhu, Xueyang Fu, Zhiwei Xiong, Zhengjun Zha, Feng Wu 0001 |
ACM Multimedia | 3 |
| 2021 | Unfolding Taylor's Approximations for Image RestorationabstractDeep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing methods empirically construct encapsulated end-to-end mapping networks without deepening into the rationality, and neglect the intrinsic prior knowledge of restoration task. To solve the above problems, inspired by Taylor’s Approximations, we unfold Taylor’s Formula to construct a novel framework for image restoration. We find the main part and the derivative part of Taylor’s Approximations take the same effect as the two competing goals of high-level contextualized information and spatial details of image restoration respectively. Specifically, our framework consists of two steps, which are correspondingly responsible for the mapping and derivative functions. The former first learns the high-level contextualized information and the later combines it with the degraded input to progressively recover local high-order spatial details. Our proposed framework is orthogonal to existing methods and thus can be easily integrated with them for further improvement, and extensive experiments demonstrate the effectiveness and scalability of our proposed framework. Man Zhou 0003, Xueyang Fu, Zeyu Xiao 0002, Aiping Liu, Zhiwei Xiong |
NeurIPS | 2 |
| 2021 | Successive Graph Convolutional Network for Image De-raining
Xueyang Fu, Qi Qi 0005, Zhengjun Zha, Xinghao Ding, Feng Wu 0001, John W. Paisley |
Int. J. Comput. Vis. | 1 |
| 2021 | An Enhanced 3-D Discrete Wavelet Transform for Hyperspectral Image ClassificationabstractIn the classification of hyperspectral image (HSI), there exists a common issue that the collected HSI data set is always contaminated by various noise (e.g., Gaussian, stripe, and deadline), degrading the classification results. To tackle this issue, we modify the 3-dimensional discrete wavelet transform (3DDWT) method by considering the noise effect on feature quality and propose an enhanced 3DDWT (E-3DDWT) approach to extract the feature and meanwhile alleviate the noise. Specifically, the proposed E-3DDWT method first applies classical 3DDWT method to the HSI data cube and thus can generate eight subcubes in each level. Then, the stripe noise is concentrated into several subcubes due to its spatial vertical property. Finally, we abandon these subcubes and obtain the feature cube by stacking the remaining ones. After acquiring the feature, we then adopt the convolutional neural network (CNN) model with an active learning strategy for classification since CNN has been verified to be a state-of-the-art feature extraction method for HSI classification, and active learning strategy can alleviate the insufficient labeled sample issue to some extent. In addition, we apply the Markov random field to enhance the final categorized results. Experiments on two synthetically striped data sets show that our proposed approach achieves better categorized results than other advanced methods. Xiangyong Cao, Jing Yao 0002, Xueyang Fu, Haixia Bi, Danfeng Hong |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Laplacian Pyramid Neural Network for Dense Continuous-Value Regression for Complex ScenesabstractMany computer vision tasks, such as monocular depth estimation and height estimation from a satellite orthophoto, have a common underlying goal, which is regression of dense continuous values for the pixels given a single image. We define them as dense continuous-value regression (DCR) tasks. Recent approaches based on deep convolutional neural networks significantly improve the performance of DCR tasks, particularly on pixelwise regression accuracy. However, it still remains challenging to simultaneously preserve the global structure and fine object details in complex scenes. In this article, we take advantage of the efficiency of Laplacian pyramid on representing multiscale contents to reconstruct high-quality signals for complex scenes. We design a Laplacian pyramid neural network (LAPNet), which consists of a Laplacian pyramid decoder (LPD) for signal reconstruction and an adaptive dense feature fusion (ADFF) module to fuse features from the input image. More specifically, we build an LPD to effectively express both global and local scene structures. In our LPD, the upper and lower levels, respectively, represent scene layouts and shape details. We introduce a residual refinement module to progressively complement high-frequency details for signal prediction at each level. To recover the signals at each individual level in the pyramid, an ADFF module is proposed to adaptively fuse multiscale image features for accurate prediction. We conduct comprehensive experiments to evaluate a number of variants of our model on three important DCR tasks, i.e., monocular depth estimation, single-image height estimation, and density map estimation for crowd counting. Experiments demonstrate that our method achieves new state-of-the-art performance in both qualitative and quantitative evaluation on the NYU-D V2 and KITTI for monocular depth estimation, the challenging Urban Semantic 3D (US3D) for satellite height estimation, and four challenging benchmarks for crowd counting. These results demonstrate that the proposed LAPNet is a universal and effective architecture for DCR problems. Xuejin Chen, Xiaotian Chen, Xueyang Fu, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Deep Multiscale Detail Networks for Multiband Spectral Image SharpeningabstractWe introduce a new deep detail network architecture with grouped multiscale dilated convolutions to sharpen images contain multiband spectral information. Specifically, our end-to-end network directly fuses low-resolution multispectral and panchromatic inputs to produce high-resolution multispectral results, which is the same goal of the pansharpening in remote sensing. The proposed network architecture is designed by utilizing our domain knowledge and considering the two aims of the pansharpening: spectral and spatial preservations. For spectral preservation, the up-sampled multispectral images are directly added to the output for lossless spectral information propagation. For spatial preservation, we train the proposed network in the high-frequency domain instead of the commonly used image domain. Different from conventional network structures, we remove pooling and batch normalization layers to preserve spatial information and improve generalization to new satellites, respectively. To effectively and efficiently obtain multiscale contextual features at a fine-grained level, we propose a grouped multiscale dilated network structure to enlarge the receptive fields for each network layer. This structure allows the network to capture multiscale representations without increasing the parameter burden and network complexity. These representations are finally utilized to reconstruct the residual images which contain spatial details of PAN. Our trained network is able to generalize different satellite images without the need for parameter tuning. Moreover, our model is a general framework, which can be directly used for other kinds of multiband spectral image sharpening, e.g., hyperspectral image sharpening. Experiments show that our model performs favorably against compared methods in terms of both qualitative and quantitative qualities. Xueyang Fu, Yue Huang 0001, Xinghao Ding, John W. Paisley |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Real-World Person Re-Identification via Degradation Invariance LearningabstractPerson re-identification (Re-ID) in real-world scenarios usually suffers from various degradation factors, e.g., low-resolution, weak illumination, blurring and adverse weather. On the one hand, these degradations lead to severe discriminative information loss, which significantly obstructs identity representation learning; on the other hand, the feature mismatch problem caused by low-level visual variations greatly reduces retrieval performance. An intuitive solution to this problem is to utilize low-level image restoration methods to improve the image quality. However, existing restoration methods cannot directly serve to real-world Re-ID due to various limitations, e.g., the requirements of reference samples, domain gap between synthesis and reality, and incompatibility between low-level and high-level methods. In this paper, to solve the above problem, we propose a degradation invariance learning framework for real-world person Re-ID. By introducing a self-supervised disentangled representation learning strategy, our method is able to simultaneously extract identity-related robust features and remove real-world degradations without extra supervision. We use low-resolution images as the main demonstration, and experiments show that our approach is able to achieve state-of-the-art performance on several Re-ID benchmarks. In addition, our framework can be easily extended to other real-world degradation factors, such as weak illumination, with only a few modifications. Zhengjun Zha, Xueyang Fu, Richang Hong, Liang Li 0003 |
CVPR | 3 |
| 2020 | JPEG Artifacts Removal via Compression Quality Ranker-Guided NetworksabstractExisting deep learning-based image de-blocking methods use only pixel-level loss functions to guide network training. The JPEG compression factor, which reflects the degradation degree, has not been fully utilized. However, due to the non-differentiability, the compression factor cannot be directly utilized to train deep networks. To solve this problem, we propose compression quality ranker-guided networks for this specific JPEG artifacts removal. We first design a quality ranker to measure the compression degree, which is highly correlated with the JPEG quality. Based on this differentiable ranker, we then propose one quality-related loss and one feature matching loss to guide de-blocking and perceptual quality optimization. In addition, we utilize dilated convolutions to extract multi-scale features, which enables our single model to handle multiple compression quality factors. Our method can implicitly use the information contained in the compression factors to produce better results. Experiments demonstrate that our model can achieve comparable or even better performance in both quantitative and qualitative measurements. Menglu Wang 0003, Xueyang Fu, Zepei Sun, Zhengjun Zha |
IJCAI | 2 |
| 2020 | Isotropic Reconstruction of 3D EM Images with Unsupervised Degradation Learning
Shiyu Deng, Xueyang Fu, Zhiwei Xiong, Chang Chen 0004, Dong Liu 0002, Xuejin Chen, Qing Ling 0001, Feng Wu 0001 |
MICCAI (5) | 2 |
| 2020 | Space-Time Video Super-Resolution Using Temporal ProfilesabstractIn this paper, we propose a novel space-time video super-resolution method, which aims to recover a high-frame-rate and high-resolution video from its low-frame-rate and low-resolution observation. Existing solutions seldom consider the spatial-temporal correlation and the long-term temporal context simultaneously and thus are limited in the restoration performance. Inspired by the epipolar-plane image used in multi-view computer vision tasks, we first propose the concept of temporal-profile super-resolution to directly exploit the spatial-temporal correlation in the long-term temporal context. Then, we specifically design a feature shuffling module for spatial retargeting and spatial-temporal information fusion, which is followed by a refining module for artifacts alleviation and detail enhancement. Different from existing solutions, our method does not require any explicit or implicit motion estimation, making it lightweight and flexible to handle any number of input frames. Comprehensive experimental results demonstrate that our method not only generates superior space-time video super-resolution results but also retains competitive implementation efficiency. Zeyu Xiao 0002, Zhiwei Xiong, Xueyang Fu, Dong Liu 0002, Zhengjun Zha |
ACM Multimedia | 3 |
| 2020 | Underwater image enhancement with global-local networks and compressed-histogram equalization
Xueyang Fu, Xiangyong Cao |
Signal Process. Image Commun. | 1 |
| 2020 | Learning Dual Transformation Networks for Image Contrast EnhancementabstractIn this work, we introduce a dual transformation network for single image contrast enhancement, which usually aims to improve global contrast and enrich local details. To this end, we propose two parallel branches to respectively handle the two goals by learning different kinds of transformations. Specifically, one branch aims to construct a global transformation curve to improve global contrast, while the other one directly predicts pixel offsets to enrich local details. In addition, we further design a differentiable histogram loss to provide supervised information related to the global contrast. In this way, the network training can be guided by different constraints, e.g., pixel-level mean squared error and statistics-level histogram error. Experiments demonstrate that our method can be effectively applied to various contrast conditions with favorable performance against the state-of-the-art methods. Yurui Zhu, Xueyang Fu, Aiping Liu |
IEEE Signal Process. Lett. | 2 |
| 2020 | Rain O'er Me: Synthesizing Real Rain to Derain With Data DistillationabstractWe present a weakly-supervised technique for learning to remove rain from images without using synthetic rain software. The method is based on a two-stage data distillation approach, which requires only some unpaired rainy and clean images to generate supervision. First, a rainy image is paired with a coarsely derained version using on a simple filtering technique (“rain-to-clean”). Then a clean image is randomly matched with the rainy soft-labeled pair. Through a shared deep neural network, the rain that is removed from the first image is then added to the clean image to generate a second pair (“clean-to-rain”). The neural network simultaneously learns to map both images such that high resolution structure in the clean images can inform the deraining of the rainy images. Demonstrations show that this approach can address those visual characteristics of rain not easily synthesized by software in the usual way. Huangxing Lin, Xueyang Fu, Xinghao Ding, Yue Huang 0001, John W. Paisley |
IEEE Trans. Image Process. | 3 |
| 2020 | Lightweight Pyramid Networks for Image DerainingabstractExisting deep convolutional neural networks (CNNs) have found major success in image deraining, but at the expense of an enormous number of parameters. This limits their potential applications, e.g., in mobile devices. In this paper, we propose a lightweight pyramid networt (LPNet) for single-image deraining. Instead of designing a complex network structure, we use domain-specific knowledge to simplify the learning process. In particular, we find that by introducing the mature Gaussian-Laplacian image pyramid decomposition technology to the neural network, the learning problem at each pyramid level is greatly simplified and can be handled by a relatively shallow network with few parameters. We adopt recursive and residual network structures to build the proposed LPNet, which has less than 8K parameters while still achieving the state-of-the-art performance on rain removal. We also discuss the potential value of LPNet for other low- and high-level vision tasks. Xueyang Fu, Borong Liang, Yue Huang 0001, Xinghao Ding, John W. Paisley |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | A Variational Pan-Sharpening With Local Gradient ConstraintsabstractPan-sharpening aims at fusing spectral and spatial information, which are respectively contained in the multispectral (MS) image and panchromatic (PAN) image, to produce a high resolution multi-spectral (HRMS) image. In this paper, a new variational model based on a local gradient constraint for pan-sharpening is proposed. Different with previous methods that only use global constraints to preserve spatial information, we first consider gradient difference of PAN and HRMS images in different local patches and bands. Then a more accurate spatial preservation based on local gradient constraints is incorporated into the objective to fully utilize spatial information contained in the PAN image. The objective is formulated as a convex optimization problem which minimizes two leastsquares terms and thus very simple and easy to implement. A fast algorithm is also designed to improve efficiency. Experiments show that our method outperforms previous variational algorithms and achieves better generalization than recent deep learning methods. Xueyang Fu, Zihuang Lin, Yue Huang 0001, Xinghao Ding |
CVPR | 1 |
| 2019 | JPEG Artifacts Reduction via Deep Convolutional Sparse CodingabstractTo effectively reduce JPEG compression artifacts, we propose a deep convolutional sparse coding (DCSC) network architecture. We design our DCSC in the framework of classic learned iterative shrinkage-threshold algorithm. To focus on recognizing and separating artifacts only, we sparsely code the feature maps instead of the raw image. The final de-blocked image is directly reconstructed from the coded features. We use dilated convolution to extract multi-scale image features, which allows our single model to simultaneously handle multiple JPEG compression levels. Since our method integrates model-based convolutional sparse coding with a learning-based deep neural network, the entire network structure is compact and more explainable. The resulting lightweight model generates comparable or better de-blocking results when compared with state-of-the-art methods. Xueyang Fu, Zhengjun Zha, Feng Wu 0001, Xinghao Ding, John W. Paisley |
ICCV | 1 |
| 2019 | Hybrid Image Enhancement With Progressive Laplacian Enhancing UnitabstractIn this paper, we propose a novel hybrid network with Laplacian enhancing unit for image enhancement. We combine the merits of two representative enhancement methods, i.e., the scaling scheme and the generative scheme, by forming a hybrid enhancing module. Meanwhile, we model image enhancement in a progressive manner with a deep cascading CNN architecture, in which the previous feature maps are used to enhance subsequent features to get an improved performance. Specifically, we propose a Laplacian enhancing unit, which can adjustably enhance the detail information by adding the residual of previous feature maps. This unit is embedded across layers for progressively enhancing the features. We build our network on the U-Net architecture and name it Hybrid Progressive Enhancing U-Net. Experiments show that our method achieves superior image enhancement results compared with the state-of-the-arts, while retaining competitive implementation efficiency. Jie Huang 0017, Zhiwei Xiong, Xueyang Fu, Dong Liu 0002, Zhengjun Zha |
ACM Multimedia | 3 |
| 2019 | Illumination-Invariant Person Re-IdentificationabstractDue to the effect of weak illumination, person images captured by surveillance cameras usually contain various degradations such as color shift, low contrast and noise. These degradations result in severe discriminant information loss, which makes the person re-identification (re-id) more challenging. However, existing person re-identification approaches are designed based on the assumption that the pedestrians images are under well lighting conditions, which is impractical in real-world scenarios. Inspired by the Retinex theory, we propose a illumination-invariant person re-identification framework which is able to simultaneously achieve Retinex illumination decomposition and person re-identification. We first verify that directly using weak illuminated images can greatly reduce the performance of person re-id. We then design a bottom-up attention network to remove the effect of weak illumination and obtain the enhanced image without introducing over-enhancement. To effectively connect low-level and high-level vision tasks, a joint training strategy is further introduced to boost the performance of person re-id under weak illumination conditions. Experiments have demonstrated the advantages of our method on benchmarks with severe lighting changes and low light conditions. Zhengjun Zha, Xueyang Fu, Wei Zhang 0021 |
ACM Multimedia | 3 |
| 2019 | A Deep Information Sharing Network for Multi-Contrast Compressed Sensing MRI ReconstructionabstractCompressed sensing (CS) theory can accelerate multi-contrast magnetic resonance imaging (MRI) by sampling fewer measurements within each contrast. However, conventional optimization-based reconstruction models suffer several limitations, including a strict assumption of shared sparse support, time-consuming optimization, and "shallow" models with difficulties in encoding the patterns contained in massive MRI data. In this paper, we propose the first deep learning model for multi-contrast CS-MRI reconstruction. We achieve information sharing through feature sharing units, which significantly reduces the number of model parameters. The feature sharing unit combines with a data fidelity unit to comprise an inference block, which are then cascaded with dense connections, allowing for efficient information transmission across different depths of the network. Experiments on various multi-contrast MRI datasets show that the proposed model outperforms both state-of-the-art single-contrast and multi-contrast MRI methods in accuracy and efficiency. We demonstrate that improved reconstruction quality can bring benefits to subsequent medical image analysis. Furthermore, the robustness of the proposed model to misregistration shows its potential in real MRI applications. Liyan Sun, Zhiwen Fan, Xueyang Fu, Yue Huang 0001, Xinghao Ding, John W. Paisley |
IEEE Trans. Image Process. | 3 |
| 2018 | Man-Made Object Recognition from Underwater Optical Images Using Deep Learning and Transfer LearningabstractWith the development of underwater optical sensors, manmade object recognition from underwater optical images has attracted wide attention. Deep learning methods have demonstrated impressive performance in object recognition tasks from natural images. However, it is difficult to collect large-scale labeled underwater optical images for training such a model. Based on the assumption that it is possible to acquire sufficient labeled in-air images, the proposed work leverages a combination of deep learning and transfer learning to develop a novel recognition system for man-made object from underwater optical images. The extracted features from the proposed network have high representative power, and demonstrate robustness in both in-air and underwater imaging modalities. Therefore, our proposed framework has the ability to recognize underwater man-made objects using only labeled in -air images. The results of experiments on simulated data demonstrate that the proposed method outperforms traditional deep learning methods in the task of underwater man-made object recognition. Xiangrui Xing, Han Zheng 0004, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
ICASSP | 4 |
| 2018 | Residual-Guide Network for Single Image DerainingabstractSingle image rain streaks removal is extremely important since rainy condition adversely affects many computer vision systems. Deep learning based methods have great success in image deraining tasks. In this paper, we propose a novel residual-guide feature fusion network, called ResGuideNet, for single image deraining that progressively predicts high-quality reconstruction while using fewer parameters than previous methods. Specifically, we propose a cascaded network and adopt residuals from shallower blocks to guide deeper blocks. We can obtain a coarse-to-fine estimation of negative residual as the blocks go deeper with this strategy. The outputs of different blocks are merged into the final reconstruction. We adopt recursive convolution to build each block and apply supervision to intermediate de-rained results. ResGuideNet is detachable to meet different rainy conditions. For images with light rain streaks and limited computational resource at test time, we can obtain a decent performance even with several building blocks. Experiments validate that ResGuideNet can benefit other low- and high-level vision tasks. Zhiwen Fan, Huafeng Wu, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
ACM Multimedia | 3 |
| 2017 | Removing Rain from Single Images via a Deep Detail NetworkabstractWe propose a new deep network architecture for removing rain streaks from individual images based on the deep convolutional neural network (CNN). Inspired by the deep residual network (ResNet) that simplifies the learning process by changing the mapping form, we propose a deep detail network to directly reduce the mapping range from input to output, which makes the learning process easier. To further improve the de-rained result, we use a priori image domain knowledge by focusing on high frequency detail during training, which removes background interference and focuses the model on the structure of rain in images. This demonstrates that a deep architecture not only has benefits for high-level vision tasks but also can be used to solve low-level imaging problems. Though we train the network on synthetic data, we find that the learned network generalizes well to real-world test images. Experiments show that the proposed method significantly outperforms state-of-the-art methods on both synthetic and real-world images in terms of both qualitative and quantitative measures. We discuss applications of this structure to denoising and JPEG artifact reduction at the end of the paper. Xueyang Fu, Jiabin Huang 0003, Delu Zeng, Yue Huang 0001, Xinghao Ding, John W. Paisley |
CVPR | 1 |
| 2017 | PanNet: A Deep Network Architecture for Pan-SharpeningabstractWe propose a deep network architecture for the pan-sharpening problem called PanNet. We incorporate domain-specific knowledge to design our PanNet architecture by focusing on the two aims of the pan-sharpening problem: spectral and spatial preservation. For spectral preservation, we add up-sampled multispectral images to the network output, which directly propagates the spectral information to the reconstructed image. To preserve spatial structure, we train our network parameters in the high-pass filtering domain rather than the image domain. We show that the trained network generalizes well to images from different satellites without needing retraining. Experiments show significant improvement over state-of-the-art methods visually and in terms of standard quality metrics. Xueyang Fu, Yuwen Hu, Yue Huang 0001, Xinghao Ding, John W. Paisley |
ICCV | 2 |
| 2017 | Image enhancement using divide-and-conquer strategy
Peixian Zhuang, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Clearing the Skies: A Deep Network Architecture for Single-Image Rain RemovalabstractWe introduce a deep network architecture called DerainNet for removing rain streaks from an image. Based on the deep convolutional neural network (CNN), we directly learn the mapping relationship between rainy and clean image detail layers from data. Because we do not possess the ground truth corresponding to real-world rainy images, we synthesize images with rain for training. In contrast to other common strategies that increase depth or breadth of the network, we use image processing domain knowledge to modify the objective function and improve deraining with a modestly sized CNN. Specifically, we train our DerainNet on the detail (high-pass) layer rather than in the image domain. Though DerainNet is trained on synthetic data, we find that the learned network translates very effectively to real-world images for testing. Moreover, we augment the CNN framework with image enhancement to improve the visual results. Compared with the state-of-the-art single image de-raining methods, our method has improved rain removal and much faster computation time after network training. Xueyang Fu, Jiabin Huang 0003, Xinghao Ding, Yinghao Liao, John W. Paisley |
IEEE Trans. Image Process. | 1 |
| 2016 | A Weighted Variational Model for Simultaneous Reflectance and Illumination EstimationabstractWe propose a weighted variational model to estimate both the reflectance and the illumination from an observed image. We show that, though it is widely adopted for ease of modeling, the log-transformed image for this task is not ideal. Based on the previous investigation of the logarithmic transformation, a new weighted variational model is proposed for better prior representation, which is imposed in the regularization terms. Different from conventional variational models, the proposed model can preserve the estimated reflectance with more details. Moreover, the proposed model can suppress noise to some extent. An alternating minimization scheme is adopted to solve the proposed model. Experimental results demonstrate the effectiveness of the proposed model with its algorithm. Compared with other variational methods, the proposed method yields comparable or better results on both subjective and objective assessments. Xueyang Fu, Delu Zeng, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
CVPR | 1 |
| 2016 | A fusion-based method for single backlit image enhancementabstractIn this work, a new simple but effective fusion-based strategy for enhancing single backlit image is proposed. The fundamental idea of proposed strategy is to blend different features into a single one to improve the specific quality of image. Most of existing methods are based on the modification of histogram to enhance the contrast of low light images. However, the backlit images are different from low light images, which have wide dynamic ranges of light regions, thus the existing methods cannot achieve good enhanced results of backlit images. To improve performance of enhanced results, the proposed method considers numerous features of images and processes the dark and bright regions, respectively. Furthermore, proposed method introduces weight maps to increase the visibility. Experimental results show that proposed method is superior to existing methods, which achieves better results both in visual effects and processing time. Xueyang Fu, Xiao-Ping Zhang 0002, Xinghao Ding |
ICIP | 2 |
| 2016 | A fusion-based enhancing method for weakly illuminated images
Xueyang Fu, Delu Zeng, Yue Huang 0001, Yinghao Liao, Xinghao Ding, John W. Paisley |
Signal Process. | 1 |
| 2016 | A novel framework method for non-blind deconvolution using subspace images priors
Peixian Zhuang, Xueyang Fu, Yue Huang 0001, Delu Zeng, Xinghao Ding |
Signal Process. Image Commun. | 2 |
| 2015 | Remote Sensing Image Enhancement Using Regularized-Histogram Equalization and DCTabstractIn this letter, an effective enhancement method for remote sensing images is introduced to improve the global contrast and the local details. The proposed method constitutes an empirical approach by using the regularized-histogram equalization (HE) and the discrete cosine transform (DCT) to improve the image quality. First, a new global contrast enhancement method by regularizing the input histogram is introduced. More specifically, this technique uses the sigmoid function and the histogram to generate a distribution function for the input image. The distribution function is then used to produce a new image with improved global contrast by adopting the standard lookup table-based HE technique. Second, the DCT coefficients of the previous contrast improved image are automatically adjusted to further enhance the local details of the image. Compared with conventional methods, the proposed method can generate enhanced remote sensing images with higher contrast and richer details without introducing saturation artifacts. Xueyang Fu, Jiye Wang, Delu Zeng, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | A Probabilistic Method for Image Enhancement With Simultaneous Illumination and Reflectance EstimationabstractIn this paper, a new probabilistic method for image enhancement is presented based on a simultaneous estimation of illumination and reflectance in the linear domain. We show that the linear domain model can better represent prior information for better estimation of reflectance and illumination than the logarithmic domain. A maximum a posteriori (MAP) formulation is employed with priors of both illumination and reflectance. To estimate illumination and reflectance effectively, an alternating direction method of multipliers is adopted to solve the MAP problem. The experimental results show the satisfactory performance of the proposed method to obtain reflectance and illumination with visually pleasing enhanced results and a promising convergence rate. Compared with other testing methods, the proposed method yields comparable or better results on both subjective and objective assessments. Xueyang Fu, Yinghao Liao, Delu Zeng, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
IEEE Trans. Image Process. | 1 |
| 2014 | A novel retinex based approach for image enhancement with illumination adjustmentabstractRetinex based algorithms have been widely used among in image enhancement. Since many retinex based algorithms remove illumination and regard the reflectance as enhancement, over-enhancement and unnaturalness are inevitable. In this paper, a novel retinex based image enhancement using illumination adjustment is proposed. Different from existing variational retinex models, a new model without the logarithmic transformation is established and can well preserve the edge. A fast alternating direction optimization method is used to solve this problem. After the decomposition of illumination and reflectance, a simple and effective post-processing method for illumination adjustment is adopted for the enhancement to make the result more natural. The proposed method can deal with many kinds of image, such as high dynamic range (HDR) images and non-uniform illumination images. Experimental results illustrate that the naturalness can be preserved while details are enhanced by the presented new approach. Xueyang Fu, Minghui LiWang, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
ICASSP | 1 |
| 2014 | A retinex-based enhancing approach for single underwater imageabstractSince the light is absorbed and scattered while traveling in water, color distortion, under-exposure and fuzz are three major problems of underwater imaging. In this paper, a novel retinex-based enhancing approach is proposed to enhance single underwater image. The proposed approach has mainly three steps to solve the problems mentioned above. First, a simple but effective color correction strategy is adopted to address the color distortion. Second, a variational framework for retinex is proposed to decompose the reflectance and the illumination, which represent the detail and brightness respectively, from single underwater image. An effective alternating direction optimization strategy is adopted to solve the proposed model. Third, the reflectance and the illumination are enhanced by different strategies to address the under-exposure and fuzz problem. The final enhanced image is obtained by combining use the enhanced reflectance and illumination. The enhanced result is improved by color correction, lightens dark regions, naturalness preservation, and well enhanced edges and details. Moreover, the proposed approach is a general method that can enhance other kinds of degraded image, such as sandstorm image. Xueyang Fu, Peixian Zhuang, Yue Huang 0001, Yinghao Liao, Xiao-Ping Zhang 0002, Xinghao Ding |
ICIP | 1 |
| 2014 | A fusion-based enhancing approach for single sandstorm imageabstractIn this paper, a novel image enhancing approach focuses on single sandstorm image is proposed. The degraded image has some problems, such as color distortion, low-visibility, fuzz and non-uniform luminance, due to the light is absorbed and scattered by particles in sandstorm. The proposed approach based on fusion principles aims to overcome the aforementioned limitations. First, the degraded image is color corrected by adopting a statistical strategy. Then two inputs, which represent different brightness, are derived only from the color corrected image by applying Gamma correction. Three weighted maps (sharpness, chromaticity and prominence), which contain important features to increase the quality of the degraded image, are computed from the derived inputs. Finally, the enhanced image is obtained by fusing the inputs with the weight maps. The proposed method is the first to adopt a fusion-based method for enhancing single sandstorm image. Experimental results show that enhanced results can be improved by color correction, well enhanced details and local contrast while promoted global brightness, increasing the visibility, naturalness preservation. Moreover, the proposed algorithm is mostly calculated by per-pixel operation, which is appropriate for real-time applications. Xueyang Fu, Yue Huang 0001, Delu Zeng, Xiao-Ping Zhang 0002, Xinghao Ding |
MMSP | 1 |
| 2014 | Bayesian Nonparametric Dictionary Learning for Compressed Sensing MRIabstractWe develop a Bayesian nonparametric model for reconstructing magnetic resonance images (MRIs) from highly undersampled k -space data. We perform dictionary learning as part of the image reconstruction process. To this end, we use the beta process as a nonparametric dictionary learning prior for representing an image patch as a sparse combination of dictionary elements. The size of the dictionary and patch-specific sparsity pattern are inferred from the data, in addition to other dictionary learning variables. Dictionary learning is performed directly on the compressed image, and so is tailored to the MRI being considered. In addition, we investigate a total variation penalty term in combination with the dictionary learning model, and show how the denoising property of dictionary learning removes dependence on regularization parameters in the noisy setting. We derive a stochastic optimization algorithm based on Markov chain Monte Carlo for the Bayesian model, and use the alternating direction method of multipliers for efficiently performing total variation minimization. We present empirical results on several MRI, which show that the proposed regularization framework can improve reconstruction accuracy over other methods. Yue Huang 0001, John W. Paisley, Xinghao Ding, Xueyang Fu, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 5 |
| 2013 | Single-Image-Based Rain and Snow Removal Using Multi-guided Filter
Xianhui Zheng, Yinghao Liao, Xueyang Fu, Xinghao Ding |
ICONIP (3) | 4 |