Xin Feng 0005

dblp:30/5155-5 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2024
0000-0002-9187-5552ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Contrastive feature decomposition for single image layer separation
Xin Feng 0005, Haobo Ji, Wenjie Pei, Guangming Lu 0002, David Zhang 0001
Neural Comput. Appl.1
2024 U²-Former: Nested U-Shaped Transformer for Image Restoration via Multi-View Contrastive Learning
abstract
While Transformer has achieved remarkable performance in various high-level vision tasks, it is still challenging to exploit the full potential of Transformer in image restoration. The crux lies in the limited depth of applying Transformer in the typical encoder-decoder framework for image restoration, resulting from heavy self-attention computation load and inefficient communications across different depth (scales) of layers. In this paper, we present a deep and effective Transformer-based network for image restoration, termed as U2-Former, which is able to employ self-attention of Transformer as the core operation for feature learning to perform image restoration in a deep encoding and decoding space. Specifically, it leverages the nested U-shaped structure to facilitate the interactions across different layers with different scales of feature maps. Furthermore, we optimize the computational efficiency for the basic Transformer block by introducing a simple yet effective feature-filtering mechanism to compress the token representation. Apart from the typical supervision ways for image restoration, our U2-Former also performs multi-view contrastive learning, which constructs positive pairs in various aspects, to learn noise-sensitive but content-irrelevant features and further decouple the noise component from the background image. Extensive experiments on various image restoration tasks, including reflection removal, rain streak removal and dehazing respectively, demonstrate the effectiveness of the proposed U2-Former.
Xin Feng 0005, Haobo Ji, Wenjie Pei, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Hierarchical Contrastive Learning for Pattern-Generalizable Image Corruption Detection
abstract
Effective image restoration with large-size corruptions, such as blind image inpainting, entails precise detection of corruption region masks which remains extremely challenging due to diverse shapes and patterns of corruptions. In this work, we present a novel method for automatic corruption detection, which allows for blind corruption restoration without known corruption masks. Specifically, we develop a hierarchical contrastive learning framework to detect corrupted regions by capturing the intrinsic semantic distinctions between corrupted and uncorrupted regions. In particular, our model detects the corrupted mask in a coarse-to-fine manner by first predicting a coarse mask by contrastive learning in low-resolution feature space and then refines the uncertain area of the mask by high-resolution contrastive learning. A specialized hierarchical interaction mechanism is designed to facilitate the knowledge propagation of contrastive learning in different scales, boosting the modeling performance substantially. The detected multi-scale corruption masks are then leveraged to guide the corruption restoration. Detecting corrupted regions by learning the contrastive distinctions rather than the semantic patterns of corruptions, our model has well generalization ability across different corruption patterns. Extensive experiments demonstrate following merits of our model: 1) the superior performance over other methods on both corruption detection and various image restoration tasks including blind inpainting and watermark removal, and 2) strong generalization across different corruption patterns such as graffiti, random noise or other image content. Codes and trained weights are available at https://github.com/xyfJASON/HCL.
Xin Feng 0005, Guangming Lu 0002, Wenjie Pei
ICCV1
2022 Learning Generalizable Latent Representations for Novel Degradations in Super-Resolution
abstract
Typical methods for blind image super-resolution (SR) focus on dealing with unknown degradations by directly estimating them or learning the degradation representations in a latent space. A potential limitation of these methods is that they assume the unknown degradations can be simulated by the integration of various handcrafted degradations (e.g., bicubic downsampling), which is not necessarily true. The real-world degradations can be beyond the simulation scope by the handcrafted degradations, which are referred to as novel degradations. In this work, we propose to learn a latent representation space for degradations, which can be generalized from handcrafted (base) degradations to novel degradations. Furthermore, we perform variational inference to match the posterior of degradations in latent representation space with a prior distribution (e.g., Gaussian distribution). Consequently, we are able to sample more high-quality representations for a novel degradation to augment the training data for SR model. We conduct extensive experiments on both synthetic and real-world datasets to validate the effectiveness and advantages of our method for blind super-resolution with novel degradations.
Fengjun Li, Xin Feng 0005, Fanglin Chen 0001, Guangming Lu 0002, Wenjie Pei
ACM Multimedia2
2022 Learning Sequence Representations by Non-local Recurrent Neural Memory
Wenjie Pei, Xin Feng 0005, Canmiao Fu, Qiong Cao, Guangming Lu 0002, Yu-Wing Tai
Int. J. Comput. Vis.2
2022 Generative Memory-Guided Semantic Reasoning Model for Image Inpainting
abstract
The critical challenge of single image inpainting stems from accurate semantic inference via limited information while maintaining image quality. Typical methods for semantic image inpainting train an encoder-decoder network by learning a one-to-one mapping from the corrupted image to the inpainted version. While such methods perform well on images with small corrupted regions, it is challenging for these methods to deal with images with large corrupted area due to two potential limitations. 1) Such one-to-one mapping paradigm tends to overfit each single training pair of images; 2) The inter-image prior knowledge about the general distribution patterns of visual semantics, which can be transferred across images sharing similar semantics, is not explicitly exploited. In this paper, we propose the Generative Memory-guided Semantic Reasoning Model (GM-SRM), which infers the content of corrupted regions based on not only the known regions of the corrupted image, but also the learned inter-image reasoning priors characterizing the generalizable semantic distribution patterns between similar images. In particular, the proposed GM-SRM first pre-learns a generative memory from the whole training data to explicitly learn the distribution of different semantic patterns. Then the learned memory are leveraged to retrieve the matching semantics for the current corrupted image to perform semantic reasoning during image inpainting. While the encoder-decoder network is used for guaranteeing the pixel-level content consistency, our generative priors are favorable for performing high-level semantic reasoning, which is particularly effective for inferring semantic content for large corrupted area. Extensive experiments on Paris Street View, CelebA-HQ, and Places2 benchmarks demonstrate that our GM-SRM outperforms the state-of-the-art methods for image inpainting in terms of both visual quality and quantitative metrics.
Xin Feng 0005, Wenjie Pei, Fengjun Li, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2021 Contrastive Feature Decomposition for Image Reflection Removal
abstract
The crux of image reflection removal stems from the difficulty of recognizing the diverse reflection patterns. Typical methods optimize the modeling of background restoration by performing low-level supervision on the restored image to minimize its per-pixel difference from the groundtruth, which re-lies on substantial training samples to learn diverse reflection patterns robustly and avoid overfitting spurious reflection patterns. In this work, we perform supervision on the contrastive distribution between the predicted background and the reflection image. Specifically, our proposed method restores the background and the reflection images in parallel, and seeks to maximize the distribution consistency between the predicted background-reflection contrast and the groundtruth contrast in the latent space. Such supervision pushes the model to focus on contrastive modeling between the background and reflection image. Extensive experiments on four real-world bench-marks demonstrate that our method consistently outperforms state-of-the-art methods.
Xin Feng 0005, Haobo Ji, Bo Jiang 0017, Wenjie Pei, Fanglin Chen 0001, Guangming Lu 0002
ICME1
2021 Deep-Masking Generative Network: A Unified Framework for Background Restoration From Superimposed Images
abstract
Restoring the clean background from the superimposed images containing a noisy layer is the common crux of a classical category of tasks on image restoration such as image reflection removal, image deraining and image dehazing. These tasks are typically formulated and tackled individually due to diverse and complicated appearance patterns of noise layers within the image. In this work we present the Deep-Masking Generative Network (DMGN), which is a unified framework for background restoration from the superimposed images and is able to cope with different types of noise. Our proposed DMGN follows a coarse-to-fine generative process: a coarse background image and a noise image are first generated in parallel, then the noise image is further leveraged to refine the background image to achieve a higher-quality background image. In particular, we design the novel Residual Deep-Masking Cell as the core operating unit for our DMGN to enhance the effective information and suppress the negative information during image generation via learning a gating mask to control the information flow. By iteratively employing this Residual Deep-Masking Cell, our proposed DMGN is able to generate both high-quality background image and noisy image progressively. Furthermore, we propose a two-pronged strategy to effectively leverage the generated noise image as contrasting cues to facilitate the refinement of the background image. Extensive experiments across three typical tasks for image background restoration, including image reflection removal, image rain steak removal and image dehazing, show that our DMGN consistently outperforms state-of-the-art methods specifically designed for each single task.
Xin Feng 0005, Wenjie Pei, Zihui Jia, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002
IEEE Trans. Image Process.1