Chao Ren 0002

dblp:02/4647-2 · DBLP profile ↗
← Back
61ranked-venue papers
11as first author
49since 2021 · last 2026
0000-0002-5347-2728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 33 · 5 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Data-interactive mamba driven SAR-optical fusion cloud removal
Fajing Liu, En Li 0002, Yuanyuan Wu 0001, Chao Ren 0002
Knowl. Based Syst.6
2026 D2S-RSG-SSD: Dual Double-Sampling With Random Sub-Samples Generation for Self-Supervised Real Image Denoising
abstract
Recent advances in self-supervised image denoising have highlighted the potential of Blind-Spot Networks (BSNs). However, existing methods suffer from three major limitations: (1) Their effectiveness in real-world scenarios is limited by strong assumptions, such as noise independence, which rarely hold in practice. (2) While sampling-based strategies can partially improve performance, BSNs inherently suffer from information loss caused by centroid masking, and removing the blind spot leads to noise overfitting, both of which hinder denoising performance. (3) Sampling-based methods often introduce checkerboard artifacts, yet existing studies typically overlook the fundamental differences between these artifacts and real noise. To address these issues, we propose a novel self-supervised denoising framework, Dual Double-Sampling with Random Sub-samples Generation (D2S-RSG-SSD). To address Limitation 1, we introduce a sampling-based framework that breaks noise dependence by combining Random Sub-samples Generation (RSG) with a cross-paired loss $\mathcal {L}_{RSG}$LRSG. RSG generates diverse sub-samples with inherent variance, referred to as sampling differences, which serve as natural perturbations to augment training data and disrupt spatial noise correlations. The proposed loss function ensures full utilization of these sub-samples while stabilizing optimization. To address Limitation 2, we propose a Dual Double-Sampling (D2S) strategy with fixed sampling patterns and a dual-branch architecture. This design reduces reliance on pixel-level information and leverages complementary features to mitigate both noise overfitting and information loss. A key advantage is its compatibility with various advanced denoising networks, lifting the constraint of using BSNs in self-supervised settings. Additionally, we introduce a fixed sub-image sampling strategy to prevent pattern collapse during inference and ensure stability. To address Limitation 3, we explicitly differentiate checkerboard artifacts from real noise and develop a dedicated artifact remover to correct pixel discontinuities caused by sampling-based operations. This design preserves fine image details while reducing over-smoothing. Experiments on benchmark real-noise datasets and self-captured noisy images demonstrate the robustness and generalizability of our framework, achieving better performance over existing methods.
Xiao Liu 0022, Xiuya Shi, Yizhong Pan, Shuhang Gu, Wei Liu 0044, Chao Ren 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Exploring Semantic Feature Discrimination for Perceptual Image Super-Resolution and Opinion-Unaware No-Reference Image Quality Assessment
abstract
Generative Adversarial Networks (GANs) have been widely applied to image super-resolution (SR) to enhance the perceptual quality. However, most existing GAN-based SR methods typically perform coarse-grained discrimination directly on images and ignore the semantic information of images, making it challenging for the super resolution networks (SRN) to learn fine-grained and semantic-related texture details. To alleviate this issue, we propose a semantic feature discrimination method, SFD, for perceptual SR. Specifically, we first design a feature discriminator (Feat-D), to discriminate the pixel-wise middle semantic features from CLIP, aligning the feature distributions of SR images with that of high-quality images. Additionally, we propose a text-guided discrimination method (TG-D) by introducing learnable prompt pairs (LPP) in an adversarial manner to perform discrimination on the more abstract output feature of CLIP, further enhancing the discriminative ability of our method. With both Feat-D and TG-D, our SFD can effectively distinguish between the semantic feature distributions of low-quality and high-quality images, encouraging SRN to generate more realistic and semantic-relevant textures. Furthermore, based on the trained Feat-D and LPP, we propose a novel opinion-unaware no-reference image quality assessment (OU NR-IQA) method, SFD-IQA, greatly improving OU NR-IQA performance without any additional targeted training. Extensive experiments on classical SISR, real-world SISR, and OU NR-IQA tasks demonstrate the effectiveness of our proposed methods. Code is available at https://github.com/GuangluDong0728/SFD.
Guanglu Dong, Xiangyu Liao, Guihuan Guo, Chao Ren 0002
CVPR5
2025 Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image Deraining
abstract
Recently, deep image deraining models based on paired datasets have made a series of remarkable progress. However, they cannot be well applied in real-world applications due to the difficulty of obtaining real paired datasets and the poor generalization performance. In this paper, we propose a novel Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image Deraining framework, CSUD, to tackle the aforementioned challenges. During training with unpaired data, CSUD is capable of generating high-quality pseudo clean and rainy image pairs which are used to enhance the performance of deraining network. Specifically, to preserve more image background details while transferring rain streaks from rainy images to the unpaired clean images, we propose a novel Channel Consistency Loss (CCLoss) by introducing the Channel Consistency Prior (CCP) of rain streaks into training process, thereby ensuring that the generated pseudo rainy images closely resemble the real ones. Furthermore, we propose a novel Self-Reconstruction (SR) strategy to alleviate the redundant information transfer problem of the generator, further improving the deraining performance and the generalization capability of our method. Extensive experiments on multiple synthetic and real-world datasets demonstrate that the deraining performance of CSUD surpasses other state-of-the-art unsupervised methods and CSUD exhibits superior generalization capability. Code is available at https://github.com/GuangluDong0728/CSUD.
Guanglu Dong, Tianheng Zheng, Yuanzhouhan Cao, Linbo Qing, Chao Ren 0002
CVPR5
2025 Degradation-Aware Feature Perturbation for All-in-One Image Restoration
abstract
All-in-one image restoration aims to recover clear images from various degradation types and levels with a unified model. Nonetheless, the significant variations among degradation types present challenges for training a universal model, often resulting in task interference, where the gradient update directions of different tasks may diverge due to shared parameters. To address this issue, motivated by the routing strategy, we propose DFPIR, a novel all-in-one image restorer that introduces Degradation-aware Feature Perturbations(DFP) to adjust the feature space to align with the unified parameter space. In this paper, the feature perturbations primarily include channel-wise perturbations and attention-wise perturbations. Specifically, channel-wise perturbations are implemented by shuffling the channels in high-dimensional space guided by degradation types, while attention-wise perturbations are achieved through selective masking in the attention space. To achieve these goals, we propose a Degradation-Guided Perturbation Block (DGPB) to implement these two functions, positioned between the encoding and decoding stages of the encoder-decoder architecture. Extensive experimental results demonstrate that DFPIR achieves state-of-the-art performance on several all-in-one image restoration tasks including image denoising, image dehazing, image deraining, motion deblurring, and low-light image enhancement. Our codes are available at https://github.com/TxpHome/DFPIR.
Xiangpeng Tian, Xiangyu Liao, Xiao Liu 0022, Meng Li 0095, Chao Ren 0002
CVPR5
2025 HQGS: High-Quality Novel View Synthesis with Gaussian Splatting in Degraded Scenes
abstract
3D Gaussian Splatting (3DGS) has shown promising results for Novel View Synthesis. However, while it is quite effective when based on high-quality images, its performance declines as image quality degrades, due to lack of resolution, motion blur, noise, compression artifacts, or other factors common in real-world data collection. While some solutions have been proposed for specific types of degradation, general techniques are still missing. To address the problem, we propose a robust HQGS that significantly enhances the 3DGS under various degradation scenarios. We first analyze that 3DGS lacks sufficient attention in some detailed regions in low-quality scenes, leading to the absence of Gaussian primitives in those areas and resulting in loss of detail in the rendered images. To address this issue, we focus on leveraging edge structural information to provide additional guidance for 3DGS, enhancing its robustness. First, we introduce an edge-semantic fusion guidance module that combines rich texture information from high-frequency edge-aware maps with semantic information from images. The fused features serve as prior guidance to capture detailed distribution across different regions, bringing more attention to areas with detailed edge information and allowing for a higher concentration of Gaussian primitives to be assigned to such areas. Additionally, we present a structural cosine similarity loss to complement pixel-level constraints, further improving the quality of the rendered images. Extensive experiments demonstrate that our method offers better robustness and achieves the best results across various degraded scenes. Source code and trained models are publicly available at: \url{https://github.com/linxin0/HQGS}.
Shi Luo, Xiaojun Shan, Chao Ren 0002, Lu Qi 0001, Ming-Hsuan Yang 0001, Nuno Vasconcelos
ICLR5
2025 Dual-Representation Interaction Driven Image Quality Assessment with Restoration Assistance
abstract
No-Reference Image Quality Assessment for distorted images has always been a challenging problem due to image content variance and distortion diversity. Previous IQA models mostly encode explicit single-quality features of synthetic images to obtain quality-aware representations for quality score prediction. However, performance decreases when facing real-world distortion and restored images from restoration models. The reason is that they do not consider the degradation factors of the low-quality images adequately. To address this issue, we first introduce the DRI method to obtain degradation vectors and quality vectors of images, which separately model the degradation and quality information of low-quality images. After that, we add the restoration network to provide the MOS score predictor with degradation information. Then, we design the Representation-based Semantic Loss (RS Loss) to assist in enhancing effective interaction between representations. Extensive experimental results demonstrate that the proposed method performs favorably against existing state-of-the-art models on both synthetic and real-world datasets. The source code will be released at https://github.com/Jingtong0527/DRI-IQA.
Jingtong Yue, Zijiu Yang, Chao Ren 0002
WACV4
2025 Efficient image super resolution via Mixed Window and Dimension Interaction
Xiao Liu 0022, Xiangyu Liao, Chao Ren 0002
Neurocomputing5
2025 PFCPNet: A progressive feature correction and prompt network for robust real-world image denoising
Yizhong Pan, Xiaohai He, Zhengyong Wang, Chao Ren 0002
Neurocomputing5
2025 Real-world blind image super-resolution with mixed and probabilistic scheme based synthetic degradation pipeline
Xiao Liu 0022, Zhengyong Wang, Xiaohai He, Chao Ren 0002
Knowl. Based Syst.5
2025 Re-Boosting Self-Collaboration Parallel Prompt GAN for Unsupervised Image Restoration
abstract
Deep learning methods have demonstrated state-of-the-art performance in image restoration, especially when trained on large-scale paired datasets. However, acquiring paired data in real-world scenarios poses a significant challenge. Unsupervised restoration approaches based on generative adversarial networks (GANs) offer a promising solution without requiring paired datasets. Yet, these GAN-based approaches struggle to surpass the performance of conventional unsupervised GAN-based frameworks without significantly modifying model structures or increasing the computational complexity. To address these issues, we propose a self-collaboration (SC) strategy for existing restoration models. This strategy utilizes information from the previous stage as feedback to guide subsequent stages, achieving significant performance improvement without increasing the framework's inference complexity. The SC strategy comprises a prompt learning (PL) module and a restorer ($Res$Res). It iteratively replaces the previous less powerful fixed restorer $\overline{Res}$Res¯ in the PL module with a more powerful $Res$Res. The enhanced PL module generates better pseudo-degraded/clean image pairs, leading to a more powerful $Res$Res for the next iteration. Our SC can significantly improve the $Res$Res 's performance by over 1.5 dB without adding extra parameters or computational complexity during inference. Meanwhile, existing self-ensemble (SE) and our SC strategies enhance the performance of pre-trained restorers from different perspectives. As SE increases computational complexity during inference, we propose a re-boosting module to the SC (Reb-SC) to improve the SC strategy further by incorporating SE into SC without increasing inference time. This approach further enhances the restorer's performance by approximately 0.3 dB. Additionally, we present a baseline framework that includes parallel generative adversarial branches with complementary "self-synthesis" and "unpaired-synthesis" constraints, ensuring the effectiveness of the training framework. Extensive experimental results on restoration tasks demonstrate that the proposed model performs favorably against existing state-of-the-art unsupervised restoration methods.
Yuyan Zhou, Jingtong Yue, Chao Ren 0002, Kelvin C. K. Chan, Lu Qi 0001, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 QP-adaptive compressed video super-resolution with coding priors
Tingrong Zhang, Zhengxin Chen, Xiaohai He, Chao Ren 0002, Qizhi Teng
Signal Process.4
2025 Enhanced Attention Context Model for Learned Image Compression
abstract
Recently, deep learning has witnessed encouraging advances in image compression. An accurate entropy model, which estimates the probability distribution of the latent representation and reduces the bits required for compressing an image, is one of the keys to the success of learned image compression methods. The latent representation presents potential correlations in local, non-local, and cross-channel contexts. However, most entropy models only consider partial correlations, leading to suboptimal entropy estimation. In this letter, we propose a novel enhanced attention context model (EACM) to make full use of various correlations between latent elements for accurate entropy estimation. The proposed EACM contains a local spatial attention block (LSAB), a local channel attention block (LCAB), a global spatial attention block (GSAB), and a global channel attention block (GCAB). LSAB, LCAB, GSAB, and GCAB are carefully designed to adaptively exploit local spatial, local channel, global spatial, and global channel correlations, respectively. The experimental results on benchmark datasets show that our image compression model with the proposed EACM outperforms several state-of-the-art methods quantitatively and qualitatively.
Zhengxin Chen, Xiaohai He, Chao Ren 0002, Tingrong Zhang, Shuhua Xiong
IEEE Signal Process. Lett.3
2025 Dual Degradation Representation for Joint Deraining and Low-Light Enhancement in the Dark
abstract
Rain in the dark poses significant challenges to deploying real-world applications such as autonomous driving, surveillance systems, and night photography. Existing low-light enhancement or deraining methods struggle to brighten low-light conditions and remove rain simultaneously. Cascade approaches, like “deraining followed by low-light enhancement” and vice versa, often result in problematic rain patterns or overly blurred and overexposed images. To address these challenges, we introduce a novel two-stage model called L2RIRNet, which innovatively integrates low-light enhancement and deraining into a single framework in real-world settings. Our model comprises two key components: a Dual Degradation Representation Network (DDR-Net) and a Restoration Network. The DDR-Net independently learns degradation representations for luminance effects in dark areas and rain patterns in light areas, which are constrained by dual degradation loss and have not been discussed in the previous methods. The Restoration Network restores the degraded image using a Fourier Detail Guidance (FDG) module, which focuses on texture details in frequency and spatial domains to inform the restoration process and leverages near-rainless detailed images. Furthermore, we contribute a dataset containing both synthetic and real-world low-light-rainy images. Extensive experiments demonstrate that our L2RIRNet performs favorably against existing methods in synthetic and complex real-world scenarios. All the code and dataset can be found inhttps://github.com/linxin0/Low_light_rainy.
Jingtong Yue, Sixian Ding, Chao Ren 0002, Lu Qi 0001, Ming-Hsuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Plug-and-Play General Image Registration for Misaligned Multi-Modal Image Fusion
abstract
Contemporary works in multi-modal image fusion often excessively rely on aligned source images, resulting in limited practicality when encountering misaligned data. However, there is still a significant gap in developing effective multimodal image registration methods to address this problem. Moreover, existing multi-modal image registration models are largely restricted to specific types of multi-modal data, lacking a general model applicable to diverse multi-modal data types. To address th above issues, this study introduces a novel method named PGMR, which stands as the first plug-and-play general multi-modal image registration model. PGMR comprises three components: Modality Prompt Module (MPM), Universal Registration Framework (URF), and Detail Enhancement Module (DEM). URF serves as the fundamental registration framework, handling both rigid and non-rigid deformations to achieve basic multi-modal image registration. MPM, one core component of this paper, is embedded within URF. Leveraging prompt learning, MPM dynamically integrates modality prompts into the intermediate output of URF, not only alleviating modality discrepancies but also promoting the ability of the registration model across various multi-modal data types. DEM is a detail enhancement module for multi-modal image registration. It can enrich the details of registration results, thereby enhancing the effectiveness of subsequent tasks. We evaluate the performance of PGMR on four multi-modal types and extensive experiments validate the feasibility of PGMR, demonstrating the superiority of our method compared to state-of-the-art alternatives. The code will be available at https://github.com/stwts/PGMR.
Tianheng Zheng, Guanglu Dong, Xiaohai He, Chao Ren 0002
IEEE Trans. Circuits Syst. Video Technol.5
2025 Transformer-Style Convolutional Network for Efficient Natural and Industrial Image Superresolution
abstract
Single image superresolution (SISR) is a critical task in computer vision with significant applications in both natural and industrial contexts. Although transformer-based approaches for SISR have achieved notable progress due to their exceptional representational capabilities, their quadratic computational complexity poses challenges for deployment on devices with limited resources. Conversely, convolutional networks (ConvNets) are inherently efficient but have difficulty capturing long-range pixel relationships because of their focus on spatial locality. This gives rise to a complementary relationship between the representational power of transformers and the efficiency of ConvNets, both of which are essential for practical applications. Motivated by this, in this article, we introduce TSCN, a novel transformer-style ConvNet. Our analysis highlights the strengths of transformers, including large-range dependencies modeling, two-order features interaction, input self-adaptation, and incorporating advanced components. Based on these insights, we guide the design of ConvNets to fully exploit these characteristics. Specifically, we rethink spatial convolution to enhance the modeling of spatial features and modify the macrostructure of the transformer by replacing self-attention and feed-forward network with the large-range multiorder convolution modulation (LMCM) layer and spatial awareness dynamic feature flow (SADFF) layer. The LMCM integrates reweighting into the large-range convolutional modulation technology, allowing self-adaptive recalibration of input representations using convolutional features as weight matrices and multiorder features interaction. In addition, the SADFF introduces spatial awareness, locality, and dynamic information flow modulation between layers. Experimental results demonstrate that our TSCN outperforms the state-of-the-art method SRFormer on multiple benchmarks by 0.03$\sim$0.17 dB, while using fewer parameters and computations.
Xiao Liu 0022, Zhengyong Wang, Xiaohai He, Haosong Gou, Chao Ren 0002
IEEE Trans. Ind. Informatics6
2025 Ada3DLane: Adaptive 3D Lane Detection From Monocular Images
abstract
3D lane detection is a fundamental task in autonomous driving, and detecting 3D lanes from monocular images has been widely adopted due to the low computational cost and the property of lanes. Recent progresses have been made based on surrogate representations such as bird’s eye view (BEV) features. However, monocular BEV construction strictly relies on flat groud assumption, and the misalignment between perspective view and BEV is inevitable. In this paper, we propose Ada3DLane, a BEV-free, query based 3D lane detector, which adaptively generates queries with rich semantic and geometric information as well as lane interactions; adaptively samples perspective view image features in spatial and temporal domain; adaptively decode the sampled features for fast and accurate 3D lane detection. We conduct extensive experiments on benchmark lane detection datasets and outperforms previous state-of-the-art methods.
Zhiming Hou, Yuanzhouhan Cao, Naiyue Chen, Chao Ren 0002, Chunyu Lin, Congyan Lang, Yidong Li
IEEE Trans. Intell. Transp. Syst.4
2025 GBPG-Net: Global Background Prior-Guided Rain and Snow Image Restoration
abstract
The aim of image restoration in the presence of rain and snow effects is to eliminate these disturbances while retaining the underlying background structure. Most existing methods tend to directly learn the mapping from corrupted images to clean ones, often resulting in residual rain or snow artifacts and compromised background structures. In this work, both theoretical analysis and experimental findings confirm the robustness of the hue channel in HSV color space to rain and snow disturbances, even when extracted from corrupted images. Motivated by this insight, we propose to leverage the global clean background cues inherent in the hue channel to guide the network in preserving the image background structure and removing interference. To this end, we introduce the global background prior-guided network (GBPG-Net) for restoring rain and snow-affected images, which employs a triangular formation to facilitate continuous interaction and updating of the global background prior (GBP) with the image feature within the GBPG-unit, resulting in improved interference removal and background structure preservation. Specifically, the GBPG-Net incorporates the global clean background prior injector (GCBPI) to inject the GBP into the network. Subsequently, the prior-guided local detail excavation (PGLDE) module, built on GCBPI, further refines interference removal and structure preservation to process local details intricately. Finally, the prior-guided local-global aggregation (PGLGA) module aggregates global background features with local detailed features, enabling the network to better understand the overall content and subtle interference for more accurate reconstruction. Quantitative and qualitative evaluations on synthetic and real datasets demonstrate the effectiveness of the proposed GBPG-Net in deraining and desnowing tasks, highlighting its advantages over existing methods. The code and supplementary documentation are available at https://github.com/liux520/GBPG-Net.
Xiao Liu 0022, Haosong Gou, Zhengyong Wang, Chao Ren 0002
IEEE Trans. Neural Networks Learn. Syst.6
2024 Unsupervised Blind Image Deblurring Based on Self-Enhancement
abstract
Significant progress in image deblurring has been achieved by deep learning methods, especially the remarkable performance of supervised models on paired synthetic data. However, real-world quality degradation is more complex than synthetic datasets, and acquiring paired data in real-world scenarios poses significant challenges. To address these challenges, we propose a novel unsupervised image deblurring framework based on self-enhancement. The framework progressively generates improved pseudo-sharp and blurry image pairs without the need for real paired datasets, and the generated image pairs with higher qualities can be used to enhance the performance of the reconstructor. To ensure the generated blurry images are closer to the real blurry images, we propose a novel re-degradation principal component consistency loss, which enforces the principal components of the generated low-quality images to be similar to those of re-degraded images from the original sharp ones. Furthermore, we introduce the self-enhancement strategy that significantly improves deblurring performance without increasing the computational complexity of network during inference. Through extensive experiments on multiple real-world blurry datasets, we demonstrate the superiority of our approach over other state-of-the-art unsupervised methods.
Lufei Chen, Xiangpeng Tian, Shuhua Xiong, Yinjie Lei, Chao Ren 0002
CVPR5
2024 Asymmetric Mask Scheme for Self-supervised Real Image Denoising
Xiangyu Liao, Tianheng Zheng, Jiayu Zhong, Chao Ren 0002
ECCV (25)5
2024 Dual-stage feedback network for lightweight color image compression artifact reduction
Zhengxin Chen, Xiaohai He, Tingrong Zhang, Shuhua Xiong, Chao Ren 0002
Neural Networks5
2024 Depth map super-resolution via learned nonlocal model and enhanced local regularization
Xiaohai He, Honggang Chen, Chao Ren 0002
Signal Process.4
2023 Geometry and Uncertainty-Aware 3D Point Cloud Class-Incremental Semantic Segmentation
abstract
Despite the significant recent progress made on 3D point cloud semantic segmentation, the current methods require training data for all classes at once, and are not suitable for real-life scenarios where new categories are being continuously discovered. Substantial memory storage and expensive re-training is required to update the model to sequentially arriving data for new concepts. In this paper, to continually learn new categories using previous knowledge, we introduce class-incremental semantic segmentation of 3D point cloud. Unlike 2D images, 3D point clouds are disordered and unstructured, making it difficult to store and transfer knowledge especially when the previous data is not available. We further face the challenge of semantic shift, where previous/future classes are indiscriminately collapsed and treated as the background in the current step, causing a dramatic performance drop on past classes. We exploit the structure of point cloud and propose two strategies to address these challenges. First, we design a geometry-aware distillation module that transfers point-wise feature associations in terms of their geometric characteristics. To counter forgetting caused by the semantic shift, we further develop an uncertainty-aware pseudo-labelling scheme that eliminates noise in uncertain pseudo-labels by label propagation within a local neighborhood. Our extensive experiments on S3DIS and ScanNet in a class-incremental setting show impressive results comparable to the joint training strategy (upper bound). Code is available at: https://github.com/leolyj/3DPC-CISS
Munawar Hayat, Chao Ren 0002, Yinjie Lei
CVPR4
2023 Efficient Information Modulation Network for Image Super-Resolution
abstract
Recent researches have shown that the success of Transformers comes from their macro-level framework and advanced components, not just their self-attention (SA) mechanism. Comparable results can be obtained by replacing SA with spatial pooling, shifting, MLP, fourier transform and constant matrix, all of which have spatial information encoding capability like SA. In light of these findings, this work focuses on combining efficient spatial information encoding technology with superior macro architectures in Transformers. We rethink spatial convolution to achieve more efficient encoding of spatial features and dynamic modulation value representations by convolutional modulation techniques. The large-kernel convolution and Hadamard product are utilizated in the proposed Multi-orders Long-range convolutional modulation (MOLRCM) layer to imitate the implementation of SA. Moreover, MOLRCM layer also achieve long-range correlations and self-adaptation behavior, similar to SA, with linear complexity. On the other hand, we also address the sub-optimality of vanilla feed-forward networks (FFN) by introducing spatial awareness and locality, improving feature diversity, and regulating information flow between layers in the proposed Spatial Awareness Dynamic Feature Flow Modulation (SADFFM) layer. Experiment results show that our proposed efficient information modulation network (EIMN) performs better both quantitatively and qualitatively. Codes and supplementary materials link: https://github.com/liux520/EIMN.
Xiao Liu 0022, Xiangyu Liao, Xiuya Shi, Linbo Qing, Chao Ren 0002
ECAI5
2023 Unsupervised Image Denoising in Real-World Scenarios via Self-Collaboration Parallel Generative Adversarial Branches
abstract
Deep learning methods have shown remarkable performance in image denoising, particularly when trained on large-scale paired datasets. However, acquiring such paired datasets for real-world scenarios poses a significant challenge. Although unsupervised approaches based on generative adversarial networks (GANs) offer a promising solution for denoising without paired datasets, they are difficult in surpassing the performance limitations of conventional GAN-based unsupervised frameworks without significantly modifying existing structures or increasing the computational complexity of denoisers. To address this problem, we propose a self-collaboration (SC) strategy for multiple denoisers. This strategy can achieve significant performance improvement without increasing the inference complexity of the GAN-based denoising framework. Its basic idea is to iteratively replace the previous less powerful denoiser in the filter-guided noise extraction module with the current powerful denoiser. This process generates better synthetic clean-noisy image pairs, leading to a more powerful denoiser for the next iteration. In addition, we propose a baseline method that includes parallel generative adversarial branches with complementary "self-synthesis" and "unpaired-synthesis" constraints. This baseline ensures the stability and effectiveness of the training network. The experimental results demonstrate the superiority of our method over state-of-the-art unsupervised methods. https://github.com/linxin0/SCPGabNet
Chao Ren 0002, Xiao Liu 0022, Jie Huang 0036, Yinjie Lei
ICCV2
2023 Random Sub-Samples Generation for Self-Supervised Real Image Denoising
abstract
With sufficient paired training samples, the supervised deep learning methods have attracted much attention in image denoising because of their superior performance. However, it is still very challenging to widely utilize the supervised methods in real cases due to the lack of paired noisy-clean images. Meanwhile, most self-supervised denoising methods are ineffective as well when applied to the real-world denoising tasks because of their strict assumptions in applications. For example, as a typical method for self-supervised denoising, the original blind spot network (BSN) assumes that the noise is pixel-wise independent, which is much different from the real cases. To solve this problem, we propose a novel self-supervised real image denoising framework named Sampling Difference As Perturbation (SDAP) based on Random Sub-samples Generation (RSG) with a cyclic sample difference loss. Specifically, we dig deeper into the properties of BSN to make it more suitable for real noise. Surprisingly, we find that adding an appropriate perturbation to the training images can effectively improve the performance of BSN. Further, we propose that the sampling difference can be considered as perturbation to achieve better results. Finally we propose a new BSN framework in combination with our RSG strategy. The results show that it significantly outperforms other state-of-the-art self-supervised denoising methods on real-world datasets. The code is available at https://github.com/p1y2z3/SDAP.
Yizhong Pan, Xiao Liu 0022, Xiangyu Liao, Yuanzhouhan Cao, Chao Ren 0002
ICCV5
2023 Efficient Parallel Multi-Scale Detail and Semantic Encoding Network for Lightweight Semantic Segmentation
abstract
In this work, we propose PMSDSEN, a parallel multi-scale encoder-decoder network architecture for semantic segmentation, inspired by the human visual perception system's ability to aggregate contextual information in various contexts and scales. Our approach introduces the efficient Parallel Multi-Scale Detail and Semantic Encoding (PMSDSE) unit to extract detailed local information and coarse large-range relationships in parallel, enabling the recognition of object boundaries and object-level areas. By stacking multiple PMSDSEs, our network learns fine-grained details and textures along with abstract category and semantic information, effectively utilizing a larger range of surrounding context information for robust segmentation. To further enhance the network's receptive field without increasing computational complexity, the Multi-Scale Semantic Extractor (MSSE) at the end of the encoder is utilized for multi-scale semantic context extraction and detailed information encoding. Additionally, the Dynamic Weighted Feature Fusion (DWFF) strategy is employed to integrate shallow layer detail information and deep layer semantic information during the decoder stage. Our method can obtain multi-scale context from local to global, achieving efficiently low-level feature extraction to high-level semantic interpretation at different scales and in different contexts. Without bells and whistles, PMSDSEN obtains a better trade-off between accuracy and complexity on popular benchmarks, including Cityscapes and Camvid. Specifically, PMSDSEN attains 73.2% mIoU with only 0.9M parameters on the Cityscapes test set. Codes and supplementary materials link: https://github.com/liux520/PMSDSEN.
Xiao Liu 0022, Xiuya Shi, Lufei Chen, Linbo Qing, Chao Ren 0002
ACM Multimedia5
2023 Nonlocal-guided enhanced interaction spatial-temporal network for compressed video super-resolution
Junxiong Cheng, Shuhua Xiong, Xiaohai He, Chao Ren 0002, Tingrong Zhang, Honggang Chen
Appl. Intell.4
2023 Block-correlation-based intra prediction for VVC
Shuhua Xiong, Xiaohai He, Honggang Chen, Chao Ren 0002
Multim. Tools Appl.6
2023 Mixed Entropy Model Enhanced Residual Attention Network for Remote Sensing Image Compression
Junjun Gao, Qizhi Teng, Xiaohai He, Zhengxin Chen, Chao Ren 0002
Neural Process. Lett.5
2023 Real Image Denoising via Guided Residual Estimation and Noise Correction
abstract
Deep learning-based methods have dominated the field of image denoising with their superior performance. Most of them belong to the non-blind denoising approaches assuming that the noise is known at a specific level. However, real-world noise is complex and usually unknown. Since the distribution and level of noise are often unavailable, it will lead to severe performance degradation for non-blind denoising methods. Therefore, introducing noise levels is crucial for the challenging real-world denoising problem. Meanwhile, we observe that noise level mismatch will bring some artifacts to the denoised images. An intuitive solution is using the intermediate denoised images to correct the inaccurate noise level maps. Thus, we introduce an iterative correction scheme, yielding better results than direct noise prediction. We further propose an effective guided feature domain denoising residual network that can promote denoising for various real-world noises using iteratively denoised features, initial image features, and noise level maps. Experimental results on real-world image datasets show that the proposed method can provide excellent visual and objective performance for the real-world denoising task.
Yizhong Pan, Chao Ren 0002, Jie Huang 0036, Xiaohai He
IEEE Trans. Circuits Syst. Video Technol.2
2023 CasaPuNet: Channel Affine Self-Attention- Based Progressively Updated Network for Real Image Denoising
abstract
Recently, the popularity of deep learning has brought broad applications of computer vision technology in industrial information systems. However, the process of image acquisition will inevitably introduce noise, which may heavily degrade image visual quality. Most of the proposed denoising methods are nonblind and they have limited performance in removing real noise with different noise levels. To overcome this problem, we propose a deep convolutional neural network (CNN)-based blind model, i.e., channel affine self-attention (CASA) based progressively updated network (CasaPuNet) for real image denoising. First, we introduce degradation mapping module (DMM) to extract degradation information, which makes the remaining subnetwork of CasaPuNet perform nonblind denoising. Then, CasaPuNet adopts a multistage architecture, which resolves the large gap between the noisy input and clean output into several small gaps and eliminates these small gaps step by step through progressive inference. Finally, a novel CASA is designed to adaptively fuse the features from multiple stages according to input statistics. Specifically, CASA extracts channel information from different features and converts them into channel weights through an affine structure for adaptive adjustment. CASA brings a significant performance gain with a small number of parameters. Extensive experiments demonstrate that CasaPuNet outperforms state-of-the-art denoising methods both quantitatively and visually.
Jie Huang 0036, Xiao Liu 0022, Yizhong Pan, Xiaohai He, Chao Ren 0002
IEEE Trans. Ind. Informatics5
2022 Enhanced Latent Space Blind Model for Real Image Denoising via Alternative Optimization
abstract
Motivated by the achievements in model-based methods and the advances in deep networks, we propose a novel enhanced latent space blind model based deep unfolding network, namely ScaoedNet, for complex real image denoising. It is derived by introducing latent space, noise information, and guidance constraint into the denoising cost function. A self-correction alternative optimization algorithm is proposed to split the novel cost function into three alternative subproblems, i.e., guidance representation (GR), degradation estimation (DE) and reconstruction (RE) subproblems. Finally, we implement the optimization process by a deep unfolding network consisting of GR, DE and RE networks. For higher performance of the DE network, a novel parameter-free noise feature adaptive enhancement (NFAE) layer is proposed. To synchronously and dynamically realize internal-external feature information mining in the RE network, a novel feature multi-modulation attention (FM2A) module is proposed. Our approach thereby leverages the advantages of deep learning, while also benefiting from the principled denoising provided by the classical model-based formulation. To the best of our knowledge, our enhanced latent space blind model, optimization scheme, NFAE and FM2A have not been reported in the previous literature. Experimental results show the promising performance of ScaoedNet on real image denoising. Code is available at https://github.com/chaoren88/ScaoedNet.
Chao Ren 0002, Yizhong Pan, Jie Huang 0036
NeurIPS1
2022 Medical visual question answering based on question-type reasoning and semantic space constraint
Xiaohai He, Luping Liu, Linbo Qing, Honggang Chen, Yan Liu 0078, Chao Ren 0002
Artif. Intell. Medicine7
2022 A prior-guided deep network for real image denoising and its applications
Jie Huang 0036, Zhibo Zhao, Chao Ren 0002, Qizhi Teng, Xiaohai He
Knowl. Based Syst.3
2022 A quality enhancement network with coding priors for constant bit rate video coding
Weiheng Sun, Xiaohai He, Chao Ren 0002, Shuhua Xiong, Honggang Chen
Knowl. Based Syst.3
2022 Deep Feature Fusion Network for Compressed Video Super-Resolution
Xiaohai He, Chao Ren 0002, Tingrong Zhang
Neural Process. Lett.4
2022 An effective deep network using target vector update modules for image restoration
Sen Zhai, Chao Ren 0002, Zhengyong Wang, Xiaohai He, Linbo Qing
Pattern Recognit.2
2022 CMAN: Leaning Global Structure Correlation for Monocular 3D Object Detection
abstract
The key to 3D object detection is proper utilization of depth data. Compared with LiDAR based approaches, 3D object detection from a single image remains a challenging task due to the lack of structure information. Recent methods leverage monocular depth estimation as a way to produce 2D depth maps, and adopt the depth maps as additional source of input to explore structure information. However, these methods either encode local structure correlations, or encode long range structure correlations by iteratively passing local messages. In this work, we propose a cross modal attention network (CMAN) for monocular 3D object detection. It is built upon the self-attention module which learns attention map from single modal data. Our CMAN is able to encode structure correlations from depth data, and embed the structure correlations with appearance information which is learned from RGB data. Thanks to the attention learning mechanism, our CMAN learns global structure correlations without iteration. In order to reduce the computational burden, our CMAN adopts a novel node sampler to eliminate redundant nodes during the attention map calculation. Experiment results on benchmark KITTI3D dataset show that our proposed CMAN outperforms the state-of-the-art methods.
Yuanzhouhan Cao, Hui Zhang 0091, Yidong Li, Chao Ren 0002, Congyan Lang
IEEE Trans. Intell. Transp. Syst.4
2021 Adaptive Consistency Prior Based Deep Network for Image Denoising
abstract
Recent studies have shown that deep networks can achieve promising results for image denoising. However, how to simultaneously incorporate the valuable achievements of traditional methods into the network design and improve network interpretability is still an open problem. To solve this problem, we propose a novel model-based denoising method to inform the design of our denoising network. First, by introducing a non-linear filtering operator, a reliability matrix, and a high-dimensional feature transformation function into the traditional consistency prior, we propose a novel adaptive consistency prior (ACP). Second, by incorporating the ACP term into the maximum a posteriori framework, a model-based denoising method is proposed. This method is further used to inform the network design, leading to a novel end-to-end trainable and interpretable deep denoising network, called DeamNet. Note that the unfolding process leads to a promising module called dual element-wise attention mechanism (DEAM) module. To the best of our knowledge, both our ACP constraint and DEAM module have not been reported in the previous literature. Extensive experiments verify the superiority of DeamNet on both synthetic and real noisy image datasets.
Chao Ren 0002, Xiaohai He, Chuncheng Wang, Zhibo Zhao
CVPR1
2021 Deep Deblocker Driven Adaptive Iteration Scheme for Compressed Image Recovery
abstract
It is challenging to propose a flexible and effective framework for various JPEG compressed image recovery (CIR) tasks. In this paper, we propose a novel deep deblocker-driven adaptive iteration scheme, which can quickly and flexibly address various CIR tasks. First, a novel fidelity (NF) is introduced into CIR, and then the CIR problem is divided into inversion and deblocking subproblems by our improved split Bregman iteration (ISBI) algorithm. Next, we design a set of compact yet effective deep deblockers. These deblockers are used as implicit priors and also used for NF in the CIR problem. The convergence of our method is proved as well. To the best of our knowledge, our method is the first work to use deblockers as implicit priors. Extensive experiments demonstrate the effectiveness of our CIR method.
Chao Ren 0002, Xiaohai He, Linbo Qing, Yuanzhouhan Cao
ICME1
2021 Learning Structure Affinity for Video Depth Estimation
abstract
Depth estimation is a structure learning problem. The affinity among neighbouring pixels plays an important role in inferring depth values. In this paper, we propose to learn structure affinity in both spatial and temporal domain for accurate depth estimation from monocular videos. Specifically, we first propose a convolutional spatial temporal propagation network (CSTPN) that learns affinity among neighbouring video frames. Secondly, we employ a structure knowledge distillation scheme that transfers the spatial temporal affinity learned by cumbersome network to compact network. By calculating pixel-wise similarities between neighboring frames and neighbouring sequences, our knowledge distillation scheme efficiently captures both short-term and long-term spatial temporal affinity. Finally, we apply a warping loss based on optical flow between video frames to further enforce the temporal affinity. Experiment results show that our proposed depth estimation approach outperform the state-of-the-art methods on both indoor and outdoor benchmark datasets.
Yuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren 0002, Yifan Liu 0001
ACM Multimedia4
2021 Deep recursive network for image denoising with global non-linear smoothness constraint prior
Chuncheng Wang, Chao Ren 0002, Xiaohai He, Linbo Qing
Neurocomputing2
2021 Remote sensing image recovery via enhanced residual learning and dual-luminance scheme
Chao Ren 0002, Xiaohai He, Linbo Qing, Yuanyuan Wu 0001, Yi-Fei Pu
Knowl. Based Syst.1
2021 Compressed image restoration via deep deblocker driven unified framework
Chao Ren 0002, Qizhi Teng, Xiaohai He, Linbo Qing, Truong Q. Nguyen
Knowl. Based Syst.1
2021 Enhanced wide-activated residual network for efficient and accurate image deblocking
Zhengxin Chen, Xiaohai He, Chao Ren 0002, Pradeep Karn, Shuhua Xiong
Signal Process. Image Commun.3
2021 Single depth map super-resolution via joint non-local self-similarity modeling and local multi-directional gradient-guided regularization
Chao Ren 0002, Honggang Chen, Ce Zhu, Kai Liu 0012
Signal Process. Image Commun.2
2021 Enhanced Separable Convolution Network for Lightweight JPEG Compression Artifacts Reduction
abstract
JPEG images are usually corrupted by various undesirable compression artifacts resulted from block-wise coarse quantization on discrete cosine transform coefficients. In recent years, deep convolutional neural networks (CNNs) have made spectacular achievements in compression artifacts reduction. However, most deep CNNs are difficult to be implemented on mobile devices due to their large number of parameters and operations. In this letter, we propose a novel deep CNN called ESCNet for lightweight JPEG compression artifacts reduction, in which enhanced separable convolution (ESConv) is carefully designed to make full use of image multi-scale information for better dense pixel value predictions. Specifically, ESConv consists of a grouped multi-scale dual depth-wise convolution (GMDDConv) and a wide-activated dual point-wise convolution (WDPConv). GMDDConv is dedicated to efficiently extracting abundant image multi-scale spatial features, which will be sent to WDPConv for effective non-linear feature fusion. The experimental results on benchmark datasets show that compared with state-of-the-art methods, our ESCNet not only achieves better performance in both objective indices and subjective quality but also greatly reduces network parameters and operations.
Zhengxin Chen, Xiaohai He, Chao Ren 0002, Honggang Chen, Tingrong Zhang
IEEE Signal Process. Lett.3
2021 Learning Image Profile Enhancement and Denoising Statistics Priors for Single-Image Super-Resolution
abstract
Single-image super-resolution (SR) has been widely used in computer vision applications. The reconstruction-based SR methods are mainly based on certain prior terms to regularize the SR problem. However, it is very challenging to further improve the SR performance by the conventional design of explicit prior terms. Because of the powerful learning ability, deep convolutional neural networks (CNNs) have been widely used in single-image SR task. However, it is difficult to achieve further improvement by only designing the network architecture. In addition, most existing deep CNN-based SR methods learn a nonlinear mapping function to directly map low-resolution (LR) images to desirable high-resolution (HR) images, ignoring the observation models of input images. Inspired by the split Bregman iteration (SBI) algorithm, which is a powerful technique for solving the constrained optimization problems, the original SR problem is divided into two subproblems: 1) inversion subproblem and 2) denoising subproblem. Since the inversion subproblem can be regarded as an inversion step to reconstruct an intermediate HR image with sharper edges and finer structures, we propose to use deep CNN to capture low-level explicit image profile enhancement prior (PEP). Since the denoising subproblem aims to remove the noise in the intermediate image, we adopt a simple and effective denoising network to learn implicit image denoising statistics prior (DSP). Furthermore, the penalty parameter in SBI is adaptively tuned during the iterations for better performance. Finally, we also prove the convergence of our method. Thus, the deep CNNs are exploited to capture both implicit and explicit image statistics priors. Due to SBI, the SR observation model is also leveraged. Consequently, it bridges between two popular SR approaches: 1) learning-based method and 2) reconstruction-based method. Experimental results show that the proposed method achieves the state-of-the-art SR results.
Chao Ren 0002, Xiaohai He, Yi-Fei Pu, Truong Q. Nguyen
IEEE Trans. Cybern.1
2020 Single depth map super-resolution via joint non-local and local modeling
abstract
Depth maps are widely used in 3D imaging techniques because of the appearance of the consumer depth cameras. However, the practical application of the depth map is limited by the poor image quality. In this paper, we propose a novel framework for the single depth map super-resolution via joint the local and non-local constraints simultaneously in the depth map. For the non-local constraint, we use the group-based sparse representation to explore the non-local self-similarity of the depth map. For the local constraint, we first estimate gradient images in different directions of the desired high-resolution (HR) depth map, and then build a multi-directional gradient guided regularizer using these estimated gradient images to describe depth gradients with different orientations. Finally, the two complementary regularizers are cast into a unified optimization framework to obtain the desired HR image. The experimental results show that the proposed method can achieve better depth super-resolution performance than state-of-the-art methods.
Chao Ren 0002, Honggang Chen, Ce Zhu
MMSP2
2020 Reduction of JPEG compression artifacts based on DCT coefficients prediction
Mengdi Sun, Xiaohai He, Shuhua Xiong, Chao Ren 0002, Xinglong Li
Neurocomputing4
2020 Image deblocking via shape-adaptive low-rank prior and sparsity-based detail enhancement
Chao Ren 0002, Xinglong Li, Xiaohai He
Signal Process. Image Commun.3
2019 Enhanced Non-Local Total Variation Model and Multi-Directional Feature Prediction Prior for Single Image Super Resolution
abstract
It is widely acknowledged that single image super-resolution (SISR) methods play a critical role in recovering the missing high-frequencies in an input low-resolution image. As SISR is severely ill-conditioned, image priors are necessary to regularize the solution spaces and generate the corresponding high-resolution image. In this paper, we propose an effective SISR framework based on the enhanced non-local similarity modeling and learning-based multi-directional feature prediction (ENLTV-MDFP). Since both the modeled and learned priors are exploited, the proposed ENLTV-MDFP method benefits from the complementary properties of the reconstruction-based and learning-based SISR approaches. Specifically, for the non-local similarity-based modeled prior [enhanced non-local total variation, (ENLTV)], it is characterized via the decaying kernel and stable group similarity reliability schemes. For the learned prior [multi-directional feature prediction prior, (MDFP)], it is learned via the deep convolutional neural network. The modeled prior performs well in enhancing edges and suppressing visual artifacts, while the learned prior is effective in hallucinating details from external images. Combining these two complementary priors in the MAP framework, a combined SR cost function is proposed. Finally, the combined SR problem is solved via the split Bregman iteration algorithm. Based on the extensive experiments, the proposed ENLTV-MDFP method outperforms many state-of-the-art algorithms visually and quantitatively.
Chao Ren 0002, Xiaohai He, Yi-Fei Pu, Truong Q. Nguyen
IEEE Trans. Image Process.1
2019 Adjusted Non-Local Regression and Directional Smoothness for Image Restoration
abstract
Image restoration (IR) problems are very important in many low-level vision tasks. Due to their ill-posed natures, image priors are widely used to regularize the solution spaces. Recently, patch-based non-local self-similarity has shown great potential in IR problems, leading to many effective non-local priors. Their performance largely depends on whether the non-local self-similarity of the underlying image can be fully exploited. However, most of these priors, including non-local regression (NLR), only utilize the center pixel of each patch to model the non-local feature, which is suboptimal. We propose an effective overlap-based non-local regression (ONLR) to fully exploit the non-local similar patches: first, the concept of overlap-based similar pixels group (OSPG) is introduced; second, for each pixel within an OSPG, the non-local weight is obtained via a novel similarity measurement method; third, based on the consistency assumption, the non-local fitting deviations (NLFDs) by using OSPGs are uniformly constrained. Because of the uniform constraints, the restoration may be poor in regions where OSPGs are not reliable. Consequently, a weighting scheme is proposed to measure the OSPG reliability, leading to a novel adjusted non-local regression (ANLR). In addition, the integral image technique (IIT) is adopted to speed up the similar patches search process. To further boost the ANLR, a local directional smoothness (DS) prior is proposed as a good complement of the non-local feature. Finally, a fast split Bregman iteration algorithm is designed to solve the ANLR-DS minimization problem. Extensive experiments on two typical IR problems, that is, image deblurring and super resolution, demonstrate the superiority of the proposed method compared to many state-of-the-art IR methods.
Chao Ren 0002, Xiaohai He, Truong Q. Nguyen
IEEE Trans. Multim.1
2018 CISRDCNN: Super-resolution of compressed images using deep convolutional neural networks
Honggang Chen, Xiaohai He, Chao Ren 0002, Linbo Qing, Qizhi Teng
Neurocomputing3
2018 SGCRSR: Sequential gradient constrained regression for single image super-resolution
Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng, Chao Ren 0002
Signal Process. Image Commun.5
2018 Nonlocal Similarity Modeling and Deep CNN Gradient Prior for Super Resolution
abstract
This letter presents a novel super-resolution (SR) method via nonlocal similarity modeling and deep convolutional neural network (CNN) gradient prior (GP). Specifically, on the one hand, the group similarity reliability (GSR) strategy is proposed for improving the adaptive high-dimensional nonlocal total variation (AHNLTV) model [statistical prior, GSR-based AHNLTV (GA)], which captures the structures of the underlying high-resolution (HR) image via the image itself. On the other hand, the GP is learned by using the deep CNN (learned prior), which predicts the gradients from external images. Finally, the GA-GP approach is proposed by incorporating the two complementary priors. The results show that GA-GP achieves better performance than other state-of-the-art SR methods.
Chao Ren 0002, Xiaohai He, Yi-Fei Pu
IEEE Signal Process. Lett.1
2017 Single Image Super-Resolution via Adaptive High-Dimensional Non-Local Total Variation and Adaptive Geometric Feature
abstract
Single image super-resolution (SR) is very important in many computer vision systems. However, as a highly ill-posed problem, its performance mainly relies on the prior knowledge. Among these priors, the non-local total variation (NLTV) prior is very popular and has been thoroughly studied in recent years. Nevertheless, technical challenges remain. Because NLTV only exploits a fixed non-shifted target patch in the patch search process, a lack of similar patches is inevitable in some cases. Thus, the non-local similarity cannot be fully characterized, and the effectiveness of NLTV cannot be ensured. Based on the motivation that more accurate non-local similar patches can be found by using shifted target patches, a novel multishifted similar-patch search (MSPS) strategy is proposed. With this strategy, NLTV is extended as a newly proposed super-high-dimensional NLTV (SHNLTV) prior to fully exploit the underlying non-local similarity. However, as SHNLTV is very high-dimensional, applying it directly to SR is very difficult. To solve this problem, a novel statistics-based dimension reduction strategy is proposed and then applied to SHNLTV. Thus, SHNLTV becomes a more computationally effective prior that we call adaptive high-dimensional non-local total variation (AHNLTV). In AHNLTV, a novel joint weight strategy that fully exploits the potential of the MSPS-based non-local similarity is proposed. To further boost the performance of AHNLTV, the adaptive geometric duality (AGD) prior is also incorporated. Finally, an efficient split Bregman iteration-based algorithm is developed to solve the AHNLTV-AGD-driven minimization problem. Extensive experiments validate the proposed method achieves better results than many state-of-the-art SR methods in terms of both objective and subjective qualities.
Chao Ren 0002, Xiaohai He, Truong Q. Nguyen
IEEE Trans. Image Process.1
2016 Single image super resolution using local smoothness and nonlocal self-similarity priors
Honggang Chen, Xiaohai He, Qizhi Teng, Chao Ren 0002
Signal Process. Image Commun.4
2016 Single Image Super-Resolution Using Local Geometric Duality and Non-Local Similarity
abstract
Super-resolution (SR) from a single image plays an important role in many computer vision applications. It aims to estimate a high-resolution (HR) image from an input low- resolution (LR) image. To ensure a reliable and robust estimation of the HR image, we propose a novel single image SR method that exploits both the local geometric duality (GD) and the non-local similarity of images. The main principle is to formulate these two typically existing features of images as effective priors to constrain the super-resolved results. In consideration of this principle, the robust soft-decision interpolation method is generalized as an outstanding adaptive GD (AGD)-based local prior. To adaptively design weights for the AGD prior, a local non-smoothness detection method and a directional standard-deviation-based weights selection method are proposed. After that, the AGD prior is combined with a variational-framework-based non-local prior. Furthermore, the proposed algorithm is speeded up by a fast GD matrices construction method, which primarily relies on the selective pixel processing. The extensive experimental results verify the effectiveness of the proposed method compared with several state-of-the-art SR algorithms.
Chao Ren 0002, Xiaohai He, Qizhi Teng, Yuanyuan Wu 0001, Truong Q. Nguyen
IEEE Trans. Image Process.1
2015 Space-time super-resolution with patch group cuts prior
Tao Li 0014, Xiaohai He, Qizhi Teng, Zhengyong Wang, Chao Ren 0002
Signal Process. Image Commun.5