VLDB 2026 Research / reviewers in the wild / expert
Wenqi Ren
dblp:126/3420
· DBLP profile ↗
189ranked-venue papers
14as first author
153since 2021 · last 2027
0000-0001-5481-653XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 124 · 9 first-author · 96 since 2021Artificial intelligence and machine learning · 105 · 9 first-author · 85 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Security and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Training-free multi-scale super-resolution with diffusion models
Aiping Zhang, Yuning Cui 0001, Jianhou Gan, Wenqi Ren |
Expert Syst. Appl. | 4 |
| 2026 | Decision-Driven Orthogonal Learning with Complementary Feature Mining for Robust Synthetic Image DetectionabstractThe widespread and inconsistent compression applied by Online Social Networks severely degrades the performance of synthetic image detectors. We attribute this degradation to two main issues: 1) the model confuses forgery artifacts with compression artifacts, and 2) compression erodes crucial discriminative high-frequency details. Existing methods suppress compression features during training but overlook the overlap between compression features and forgery-related features, leading to the unintended removal of forgery traces. To address artifact confusion, we introduce a Decision-Driven Orthogonal Constraint, which defines a classification decision axis pointing from the real class centroid to the forged class centroid. This constraint enforces compression artifacts to be orthogonal to the decision axis, mitigating their interference with forgery detection without entirely removing them, thus preventing the suppression of forgery-related features. To mitigate the erosion of high-frequency details, we propose to mine complementary forgery cues from both low-frequency information and compressed high-frequency components. A bidirectional update strategy and an adaptive global-local modulator are proposed to facilitate the utilization of forgery cues. Extensive experiments demonstrate that our method achieves state-of-the-art generalization performance in challenging open-world detection scenarios. Wei Wang 0335, Linchao Zhang, Wenqi Ren |
AAAI | 5 |
| 2026 | G-Cap: A Game Character Caption GeneratorabstractWhile Large Vision-Language Models (LVLMs) have demonstrated remarkable proficiency in image captioning, existing research primarily focuses on real-world scenarios, leaving surreal, highly stylized, and semantically hybrid virtual-world scenarios significantly underexplored. In this work, we introduce Game Character Captioning, a novel task designed to evaluate LVLMs’ capability to perceive and describe game character from the virtual-world. To facilitate evaluation, we establish GC-Bench, a manually annotated benchmark, and propose Graph-F1 to effectively assess performance on this task. Our evaluation reveals that: (1) current state-of-the-art LVLMs, including closed-source giants such as Gemini 3 Pro and GPT-5.1, struggle to maintain the high performance seen in real-world scenarios; and (2) a notable gap exists between open-source and closed-source models. To bridge this gap, we construct GC-148K, a large-scale dataset generated via a specialized data pipeline, and develop the G-Cap series. Experiments demonstrate that G-Cap series rivals the performance of advanced closed-source models at a lower cost, offering an efficient solution for industrial-grade production environment. X. U. Cheng, Gui Zheng, Wenqi Ren |
ACL (1) | 7 |
| 2026 | Shape-aware and feature fused power line detection network
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, LinLin Shen, Jun Zhang 0011 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Misalignment-tolerant perceptual similarity metric for full reference image dehazing quality assessment
Jiyou Chen, Gaobo Yang, Wenqi Ren |
Expert Syst. Appl. | 5 |
| 2026 | Structure-semantic fusion contrastive transformer for robust knowledge graph embedding
Jianhou Gan, Wenqi Ren, Jun Wang 0101 |
Expert Syst. Appl. | 3 |
| 2026 | Focal Modulation for Image RestorationabstractAbstract Image restoration aims to recover a sharp image from its degraded counterpart by removing degradations ( e.g., noise, haze, and blur) and restoring missing details. It plays an important role in many fields, such as remote sensing and medical imaging. How to effectively capture critical information parsimoniously for high-quality reconstruction has long been a pivotal problem in this domain. This study aims to develop an efficient and effective focal modulation scheme for image restoration. Inspired by the fact that different regions in a corrupted image always undergo degradations in various degrees, we introduce a dual-domain selection mechanism to emphasize crucial information for restoration, such as edge signals and hard regions. Moreover, a channel modulation module is developed to facilitate channel interactions by exploring the utility of the Fourier transform in channel dimensions. In addition, we split high-resolution features to insert multi-scale receptive fields into the network, improving efficiency and performance. Incorporating these designs into a U-shaped convolutional backbone, the network achieves state-of-the-art performance on 13 different datasets for five general image restoration tasks, including dehazing, desnowing, deraining, motion/defocus deblurring, and low-light enhancement. To further demonstrate the effectiveness of our focal modulation strategy, we apply it to the all-in-one image restoration setting, and the obtained model performs favorably against state-of-the-art all-in-one algorithms. Moreover, our module extends effectively to tasks such as composite degradation, medical imaging, and ultra-high-definition image restoration. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
Int. J. Comput. Vis. | 2 |
| 2026 | Automated data synthesis and retrieval-augmented generation for legal large language models
Wenqi Ren, Lixing Shen, Yinxia Hong, Jiawei Wang 0025, Da Cao |
Knowl. Based Syst. | 1 |
| 2026 | Visual-in-Visual: A Unified and Efficient Baseline for Image RestorationabstractRecent years have witnessed remarkable progress in image restoration, yet achieving both high performance and efficiency remains a persistent challenge. To address this issue, we present VIVNet, a strong and efficient unified baseline designed to balance accuracy and practicality. Drawing inspiration from the high efficiency of the human visual system, VIVNet embeds a biologically inspired micro visual module into each block of a macro U-shaped vision architecture. This module mimics key perceptual processes such as retinal encoding, lateral inhibition, and high-order processing by combining lightweight depth-wise convolutions for multi-receptive-field feature extraction, a similarity-aware weighting mechanism to emphasize informative signals, and high-order interactions implemented via iterative element-wise multiplication to capture complex dependencies. This design enhances the model's representational capacity while maintaining computational efficiency. Unlike most existing methods that are limited to narrow task settings, we evaluate VIVNet across a wide range of scenarios, including general, all-in-one, and composite degradation tasks, as well as ultra-high-definition (UHD), underwater, medical, and remote sensing datasets. Extensive experiments show that VIVNet delivers competitive performance with high efficiency. Yuning Cui 0001, Wenqi Ren, Boxin Shi, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Follow your prompts: Controllable image dehazing via latent space manipulation
Jiyou Chen, Gaobo Yang, Wenqi Ren |
Pattern Recognit. | 5 |
| 2026 | Bidomain multi-order modeling for image dehazing
Chenxu Wu, Junling Li, Wei Wang 0335, Wenqi Ren |
Pattern Recognit. | 6 |
| 2026 | Dual-strategy retrieval-augmented generation for anomaly recovery in large language models
Jinxin Lv, Jianhou Gan, Wenqi Ren, Jun Wang 0101 |
Pattern Recognit. | 4 |
| 2026 | Unpaired overwater image defogging using inverted dark channel prior-guided cycle-consistent generative adversarial network
Yaozong Mo, Tuxin Guan, Qiuping Jiang, Wenqi Ren, Wenwu Wang 0001 |
Pattern Recognit. | 5 |
| 2026 | Dynamic and Consistent Doubly Stochastic similarity learning for multi-view and multi-order clustering
Nian Wang 0001, Zhigao Cui, Yanzhao Su, Aihua Li, Yuanliang Xue, Wenqi Ren |
Pattern Recognit. | 6 |
| 2026 | Beyond similarity: Mutual information-guided retrieval for in-context learning in VQA
Zezhong Lv, Jian Zhao 0006, Yan Wang 0122, Yuchen Yuan, Yuchu Jiang, Wenqi Ren, Xuelong Li 0001 |
Pattern Recognit. | 9 |
| 2026 | Wavelet-based physically guided normalization network for real-time traffic dehazing
Shengdong Zhang, Xiaoqin Zhang 0002, LinLin Shen, Shaohua Wan 0001, Wenqi Ren |
Pattern Recognit. | 5 |
| 2026 | APDiff: An Adaptive Physics-Guided Diffusion Framework for efficient unpaired image dehazing
Li Zhao 0005, Hanqi Wang, Chenxiang Fan, Haigen Hu, Wenqi Ren, Zhonglong Zheng |
Pattern Recognit. | 5 |
| 2026 | Hierarchical Multi-Modal Enhancement for Robust Transmission Line Detection
Shengdong Zhang, Xiaoqin Zhang 0002, Shaohua Wan 0001, Yujing M. Jiang, Wujie Zhou, LinLin Shen, Wenqi Ren |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2026 | Real-World Nighttime Dehazing via Score-Guided Multi-Scale Fusion and Dual-Channel Enhancement
Yun Liu 0002, Shirui Luo, Wenqi Ren, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Deep Unfolding Dehazing Network via Iterative Refinement and Self-Prompted Correction LearningabstractExisting one-time dehazing approaches struggle to simultaneously restore high-fidelity scene content and suppress haze-induced artifacts. Although recursive iterations can enhance the modeling capacity for complex degradations, the associated estimation errors are often amplified during propagation, thereby constraining the overall dehazing performance. To address these challenges, this work proposes a deep unfolding dehazing network, termed I3-Net, which integrates an iterative refinement strategy (IRS) and self-prompted correction learning (SCL). Specifically, we first develop a baseline dehazing network, termed I-Net, which is constructed around a physically-aware feature enhancement module (PFEM). By embedding physical priors into the feature space, PFEM enforces consistency between learned representations and the haze degradation process, thereby providing a reliable foundation for subsequent refinement. Building upon I-Net, IRS is designed to recursively unfold the baseline, progressively improving its dehazing outputs and enabling the construction of a deep unfolding dehazing network, I3-Net, with cross-stage feature association capability. To mitigate the accumulation of minor estimation errors inherent in iterative frameworks, we further propose SCL mechanism inspired by the corrective behavior of the human visual system. By integrating IRS and SCL, I3-Net adaptively identifies and rectifies residual haze regions at each unfolding stage, effectively suppressing error propagation and achieving high-quality image restoration. Extensive experiments demonstrate that the proposed I3-Net consistently outperforms existing SOTA methods in both quantitative metrics and visual perception across multiple benchmark datasets. Yuting Pang, Shilong Wang 0005, Wenqi Ren, Jiaming Niu, Jiguo Yu, Jianlei Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | ZRID-Net: Zero-Reference Real-World Image Dehazing Framework via Deep Self-Decoupling and Reverse Knowledge TransferabstractThis paper investigates one of the most challenging problems in single image dehazing: how to restore haze-free scenes solely from the input observed image without relying on paired or unpaired images and how to extract useful prior information from the observed image to guide the dehazing process. To address these challenges, this paper introduces a novel zero-reference real-world image dehazing method via deep self-decoupling and reverse knowledge transfer (ZRID-Net). Specifically, we first employ a model-driven approach to preliminarily decouple the observed image into coarse-grained components: the haze-free image, transmission map, and atmospheric light. Subsequently, we refine the haze-free image and transmission map separately via a data-driven approach. In addition, we propose a novel reverse knowledge transfer method to exploit latent prior information within hazy images thoroughly for dehazing guidance. This method combines knowledge transfer and contrastive learning to reverse guide the refinement network away from haze characteristics. Finally, a perceptual fusion strategy is employed to obtain haze-free images with high visibility and realism. Extensive experiments demonstrate that the proposed ZRID-Net effectively restores image clarity, enhances structural details, and improves color fidelity across various challenging haze conditions without relying on paired or unpaired supervision. On multiple benchmark datasets, ZRID-Net outperforms existing SOTA approaches in terms of both quantitative metrics and visual quality. The results also confirm its strong generalizability and practical applicability to real-world scenarios. The relevant implementation code can be found at https://github.com/cswangshilong/ZRID-Net. Shilong Wang 0005, Wenqi Ren, Peng Gao 0005, Jiguo Yu, Jianlei Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Corrections to "Exploring Fuzzy Priors From Multimapping GAN for Robust Image Dehazing"
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2026 | Efficient All-in-One Image Restoration With Adaptive Frequency EnhancementabstractAll-in-one image restoration has recently attracted considerable attention for its ability to address multiple degradation types within a single, unified framework. However, existing methods often incur substantial computational overhead, especially when incorporating explicit degradation priors via complex auxiliary branches, hindering their practical deployment. In this paper, we propose AdaptIR, an efficient all-in-one image restoration network equipped with adaptive frequency enhancement. Recognizing that different degradations impact distinct frequency subbands and exhibit spatially varying restoration demands, we design an Adaptive Frequency Enhancement Module (AFEM) that couples frequency learning with adaptive convolutions to better capture frequency-aware information. Specifically, AFEM learns pixel-wise adaptive attention weights to modulate the spectra of dynamic convolutions, enabling spatially adaptive and content-aware restoration. Furthermore, we introduce a lightweight backbone featuring a Receptive Field Expansion Module (RFEM), which enlarges the receptive field of a convolutional U-shaped architecture by convolving wavelet-transform coefficients. By integrating the plug-and-play AFEM into the bottleneck of the baseline model, AdaptIR achieves state-of-the-art performance on all-in-one image restoration tasks involving multiple degradations, while maintaining high computational efficiency. Moreover, the proposed model can be readily extended to single-degradation tasks (e.g., dehazing, desnowing, and deraining) and domain-specific applications, including ultra-high-definition (UHD), medical, and remote sensing image restoration. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Image Process. | 2 |
| 2026 | CDIR: LoRA-Inspired Attention for Efficient Composite Degradation Image RestorationabstractSpecialized image restoration methods have been extensively explored, each targeting a specific type of degradation. However, real-world images often suffer from composite degradations, prompting growing interest in unified restoration approaches. While recent unified models have shown promising results, many are hindered by high computational complexity, limiting their deployment in resource-constrained settings. Motivated by the parameter-efficient design of Low-Rank Adaptation (LoRA), we propose an efficient attention module specifically designed for composite degradation image restoration. The proposed method adopts a dual-branch architecture, where one branch processes features at full resolution, and the other operates with reduced spatial and channel dimensions to improve efficiency. To better adapt to diverse degradation patterns, the latter branch is further divided into two sub-branches, each incorporating dynamic operations guided by local and contextual priors. These context priors are iteratively updated within each module, drawing inspiration from feedback mechanisms in reinforcement learning, thereby enabling the model to effectively perceive and handle multiple degradation types within a unified structure. Additionally, we introduce a multi-scale feed-forward network to further enhance both performance and computational efficiency. Extensive experiments on two composite degradation benchmarks demonstrate that our proposed network, CDIR, achieves state-of-the-art performance with significantly reduced complexity and fast inference speed. In addition, CDIR shows strong adaptability to various task-specific image restoration scenarios, such as dehazing, desnowing, and deraining. It also performs robustly on domain-specific applications such as ultra-high-definition (UHD), remote sensing, and medical image restoration, highlighting its versatility and practical applicability. Yuning Cui 0001, Wenqi Ren, Boxin Shi, Jianhou Gan, Alois C. Knoll |
IEEE Trans. Image Process. | 2 |
| 2026 | Noise-Induced Cross-Modal Information Interaction and Dual-Prompt Learning for Medical Image SegmentationabstractAccurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling is both labor-intensive and reliant on domain-specific expertise. To address this limitation without requiring additional annotations, we propose a novel multimodal segmentation framework that leverages medical text annotations as an auxiliary modality to complement visual information. In particular, our approach introduces a learnable encoding strategy for joint distribution modeling of image and text, which enables discriminative fusion and effectively suppresses cross-modal redundancy. Moreover, we innovatively design a frequency-domain prompt encoder based on the discrete wavelet transform (DWT) to capture multi-frequency features, thereby significantly enhancing the model's ability to delineate fine-grained boundaries. Overall, our framework integrates cross-attention for effective cross-modal interaction, employs joint distribution modeling to enable discriminative and redundancy-reduced multimodal fusion, and incorporates auxiliary supervision to strengthen the learning of task-relevant features. Extensive experiments on nine public datasets across three clinical tasks-including cell, lung infection, and polyp segmentation-demonstrate that our method achieves competitive segmentation performance while maintaining favorable computational efficiency. Comprehensive ablation studies and feature distribution visualizations further validate the effectiveness and robustness of our proposed components. The code will be made publicly available at https://github.com/chenpeng052/MDFP. Chao Huang 0008, Jie Wen 0001, Wei Wang 0335, Li Shen 0008, Wenqi Ren, Xiaochun Cao, Chengliang Liu 0003 |
IEEE Trans. Image Process. | 6 |
| 2026 | IHDCP: Single Image Dehazing Using Inverted Haze Density Correction PriorabstractImage dehazing, a crucial task in low-level vision, supports numerous practical applications, such as autonomous driving, remote sensing, and surveillance. This paper proposes IHDCP, a novel Inverted Haze Density Correction Prior for efficient single image dehazing. It is observed that the medium transmission can be effectively modeled from the inverted haze density map using correction functions with various gamma coefficients. Based on this observation, a pixel-wise gamma correction coefficient is introduced to formulate the transmission as a function of the inverted haze density map. To estimate the transmission, IHDCP is first incorporated into the classic atmospheric scattering model (ASM), leading to a transcendental equation that is subsequently simplified to a quadratic form with a single unknown parameter using the Taylor expansion. Then, boundary constraints are designed to estimate this model parameter, and the gamma correction coefficient map is derived via the Vieta theorem. Finally, the haze-free result is recovered through ASM inversion. Experimental results on diverse synthetic and real-world datasets verify that our algorithm not only provides visually appealing dehazing performance with high computational efficiency, but also outperforms several state-of-the-art dehazing approaches in both subjective and objective evaluations. Moreover, our IHDCP generalizes well to various types of degraded scenes. Our code is available at https://github.com/TaoLi-TL/IHDCP. Yun Liu 0002, Chunping Tan, Wenqi Ren, Cosmin Ancuti, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2026 | Real-World Nighttime Image Dehazing via Bayesian-Based Fractional-Order Variational ModelabstractImages captured under real-world nighttime haze conditions often suffer from severe degradations, including low visibility, color distortion, and reduced contrast, which not only impair visual perception but also degrade the performance of vision-based tasks. However, existing dehazing methods are mainly designed for daytime scenarios and struggle to cope with the complex illumination and scattering characteristics of nighttime hazy images. In this paper, we propose a novel Bayesian-based variational framework with fractional-order constraints for real-world nighttime image dehazing. First, a simplified physical model is constructed to characterize nighttime hazy images, accounting for haze, low-light conditions, Poisson noise, and glow degradations. An anisotropic pre-processing strategy is iteratively applied in the Lab color space to remove glow effects. Subsequently, illumination and reflectance estimation within our constructed physical model is formulated as a maximum a-posteriori (MAP) problem, which is then approximated as a unified variational optimization function. To impose prior constraints, two fractional-order terms are introduced as priors to regulate the illumination and reflectance, promoting piecewise smoothness in illumination and preserving sharp edges and fine textures in reflectance. The resulting variational model is efficiently solved using the alternating direction minimization method. Finally, the estimated illumination and reflectance are enhanced via spatial-domain gamma correction for brightness adjustment and frequency-domain processing for texture detail enhancement. Extensive experiments on real-world datasets demonstrate that the proposed framework outperforms state-of-the-art dehazing methods in both qualitative and quantitative evaluations. Besides, our algorithm generalizes effectively to both other degraded scenes and high-level vision tasks. Yun Liu 0002, Zichen Zhou, Wenqi Ren, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2026 | Disentangle to Fuse: Toward Content Preservation and Cross-Modality Consistency for Multi-Modality Image FusionabstractMulti-modal image fusion (MMIF) aims to integrate complementary information from heterogeneous sensor modalities. However, substantial cross-modality discrepancies hinder joint scene representation and lead to semantic degradation in the fused output. To address this limitation, we propose C2MFuse, a novel framework designed to preserve content while ensuring cross-modality consistency. To the best of our knowledge, this is the first MMIF approach to explicitly disentangle style and content representations across modalities for image fusion. C2MFuse introduces a content-preserving style normalization mechanism that suppresses modality-specific variations while maintaining the underlying scene structure. The normalized features are then progressively aggregated to enhance fine-grained details and improve content completeness. In light of the lack of ground truth and the inherent ambiguity of the fused distribution, we further align the fused representation with a well-defined source modality, thereby enhancing semantic consistency and reducing distributional uncertainty. Additionally, we introduce an adaptive consistency loss with learnable transformation, which provides dynamic, modality-aware supervision by enforcing global consistency across heterogeneous inputs. Extensive experiments on five datasets across three representative MMIF tasks demonstrate that C2MFuse achieves efficient and high-quality fusion, surpasses existing methods, and generalizes effectively to downstream visual applications. Xinran Qin, Yuning Cui 0001, Shangquan Sun, Ruoyu Chen 0001, Wenqi Ren, Alois C. Knoll, Xiaochun Cao |
IEEE Trans. Image Process. | 5 |
| 2026 | Toward a Completely Blind Attacker for No-Reference Image Quality Assessment ModelsabstractNo-reference image quality assessment (NR-IQA) models are critically vulnerable to adversarial attacks, posing significant risks to downstream vision systems. However, existing attack methods suffer from high computational costs, reliance on Mean Opinion Score (MOS) annotations, and poor cross-model transferability. To overcome these limitations, we propose Degrade-to-OverReconstruct (DOR), a novel prior knowledge-driven black-box attack framework operating in a "completely blind" manner, requiring neither MOS labels nor surrogate models, inducing significant prediction bias solely based on distortion statistics. Specifically, DOR generates universal adversarial examples by first applying mild degradation to preserve global structure and then employing aggressive over-reconstruction using a Residual Denoising Diffusion Model (RDDM) to adaptively disrupt intrinsic Natural Scene Statistics (NSS)-a shared foundation across NR-IQA models. Extensive experiments on synthetic (LIVE, TID2013) and authentic (CLIVE) datasets demonstrate DOR's strong attack performance and superior transferability against leading NR-IQA models that cover diverse deep neural network architectures. Our work pioneers a diffusion model-based "completely blind" attack paradigm, offering a practical, MOS-free solution for adversarial robustness assessment of NR-IQA models in real-world deployments. Xinyu Ruan, Hangwei Chen, Chao Huang 0008, Wenqi Ren, Qiuping Jiang |
IEEE Trans. Image Process. | 4 |
| 2026 | FreeDehaze: Towards Training-Free Real-World Image Dehazing via Diffusion Degradation PriorabstractRestoring high-quality images from degraded hazy images is a challenging task, particularly in real-world scenarios. Recent investigations seek to address this limitation by exploring advanced methods for synthesizing haze and incorporating real-world hazy images. Due to the inherent diversity and complexity of real-world haze, these methods struggle to accurately model haze representations. Based on our observation that the hazy images generated by advanced text-to-image diffusion models exhibit a remarkable resemblance to real-world haze, it suggests that these diffusion models effectively internalize haze representations. Hence, we propose FreeDehaze, a novel training-free diffusion method for real-world image dehazing. FreeDehaze is a posterior-based framework capable of addressing non-linear dehazing challenges without relying on additional degradation estimation networks. It follows the human cognition for image restoration, beginning with perception and subsequently enhancing the image. The core method initially generates pseudo-clean images based on abstract textual descriptions. Subsequently, optimal transport aligns the denoising network output with the pseudo-clean image within a PCA-based haze subspace, facilitating high-fidelity dehazing. Extensive experiments demonstrate that FreeDehaze outperforms comparative methods in subjective metrics on challenging datasets (e.g., RTTS, URHI, and O-HAZE) and achieves competitive objective metrics, demonstrating strong generalization although without additional training. Jiawei Wu 0001, Yikun Ma, Wenqi Ren, Zhi Jin 0002, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2026 | Virtual Consistency Model for All-in-One Image RestorationabstractAll-in-one Image Restoration (AIR) seeks to address diverse degradations using a unified model trained only once. Existing methods often rely on degradation-specific guidance, leading to conflicting gradients during training. In contrast, diffusion models offer a promising alternative by operating in a high-noise space where diverse degradations exhibit a homogeneous Gaussian distribution. This characteristic alleviates gradient conflicts associated with task-specific degradations. However, existing diffusion-based AIR methods often suffer from a lack of direct supervision in the image space, leading to error accumulation during the iterative denoising process and image fidelity compromisation. This highlights a fundamental dilemma for AIR: the optimal space for modeling degradations is inherently suboptimal for preserving image fidelity. To address this issue, we propose a Virtual Consistency Model for AIR (VCMAIR), which restores images in the high-noise space while employing a novel consistency function to enforce accurate supervision in the image space. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art methods across a comprehensive benchmark of diverse degradation scenarios, including both standard AIR tasks and challenging real-world image restoration tasks. Jiawei Wu 0001, Luwei Tu, Zhi Jin 0002, Kaihao Zhang, Wenqi Ren, Xiaochun Cao |
IEEE Trans. Image Process. | 6 |
| 2026 | SGNet: Style-Guided Network With Temporal Compensation for Unpaired Low-Light Colonoscopy Video EnhancementabstractA low-light colonoscopy video enhancement method is needed as poor illumination in colonoscopy can hinder accurate disease diagnosis and adversely affect surgical procedures. Existing low-light video enhancement methods usually apply a frame-by-frame enhancement strategy without considering the temporal correlation between them, which often causes a flickering problem. In addition, most methods are designed for endoscopic devices with fixed imaging styles and cannot be easily adapted to different devices. In this paper, we propose a Style-Guided Network (SGNet) for unpaired Low-Light Colonoscopy Video Enhancement (LLCVE). Given that collecting content-consistent paired videos is difficult, SGNet adopts a CycleGAN-based framework to convert low-light videos to normal-light videos, in which a Temporal Compensation (TC) module and a Style Guidance (SG) module are proposed to alleviate the flickering problem and achieve flexible style transfer, respectively. The TC module compensates for a low-light frame by learning the correlated feature of its adjacent frames, thereby improving the temporal smoothness of the enhanced video. The SG module encodes the text of the imaging style and adaptively explores its intrinsic relationships with video features to obtain style representations, which are then used to guide the subsequent enhancement process. Extensive experiments on a curated database show that SGNet achieves promising performance on the LLCVE task, outperforming state-of-the-art methods in both quantitative metrics and visual quality. Guanghui Yue 0001, Wanqing Liu, Jingfeng Du, Tianwei Zhou, Hanhe Lin, Qiuping Jiang, Wenqi Ren |
IEEE Trans. Image Process. | 8 |
| 2026 | Low-Light Image Enhancement Using a Retinex-Based Variational Model With Weighted $L_{p}$ Norm ConstraintabstractImages taken in low-light conditions are frequently affected by limited visibility, diminished contrast and severe noise, adversely impacting the performance of various computer vision tasks. Most variational-based Retinex decomposition methods mainly depend on integer norms to constrain the illumination and reflectance components. However, this strategy may fail to achieve the ideal Retinex decomposition. In this paper, we propose a Retinex-based variational model that incorporates flexible constraints for both illumination and reflectance. Specifically, we impose the Lpnorm constraints with varying values of p to ensure the piece-wise smoothness of the illumination and promote the presence of abundant textures in the reflectance. Moreover, we develop two effective pixel-wise weight matrices that consider variance and gradients of the input image respectively, with the objective of preserving the structural edges of the illumination and retaining more details in the reflectance. In addition, we use an L2norm to estimate the overall noise level and avoid noise amplification. Through incorporating these above constraints, our proposed variational model can obtain a structure-aware illumination and a detail-revealed reflectance. Qualitative and quantitative comparisons on real-world and synthetic datasets indicate that our approach yields results with superior visual quality and outperforms several state-of-the-art algorithms on objective metrics. Besides, our algorithm can also address similar low-level computer vision challenges, such as image dehazing and underwater image enhancement. The source code is available at https://github.com/Enping-Hu/dual weighted lp. Enping Hu, Yun Liu 0002, Anzhi Wang, Babak Shiri, Wenqi Ren, Weisi Lin |
IEEE Trans. Multim. | 5 |
| 2026 | VDMamba: Vector Decomposition in Vision Mamba for Image Deraining and BeyondabstractImage deraining aims to remove rain perturbations from rainy images and restore clear backgrounds. Recent research has employed the Mamba technique for image restoration, achieving exceptional results due to its effectiveness and efficiency in modeling long-range sequence relationships. However, a significant challenge remains: developing a comprehensive framework that considers the intrinsic coupling characteristics between image deraining and the Mamba architecture is largely unexplored. We propose that introducing a 1D sequential representation of Mamba could enhance image deraining by characterizing the direction-aware distribution of rain perturbations. This motivates us to introduce a new vector decomposition-based vision Mamba approach (VDMamba). This method investigates vector decomposition within the context of vision Mamba, addressing the challenging task of image deraining and beyond in the frequency embedding space. The key innovation of VDMamba is the Mamba-based vector decomposition and synthesis module (VDSM). This module derives 1D basic vectors (vertical and horizontal) from the frequency components via vector decomposition and employs the single-direction scanning of Mamba to eliminate the direction-specific degradation perturbation. This transformation allows the incipient Mamba to explore directionspecific global relationships for accurate perturbation learning, without requiring an elaborate design of the Mamba scanning. Additionally, the vertical and horizontal components in VDSM are encoded jointly in a bidirectional coupling manner, enabling the exploration of complementary and redundant components for refinement. Experiments on various image enhancement tasks, including image deraining, raindrop removal, rain haze removal, image dehazing, low-light image enhancement, and underwater image enhancement, demonstrate that VDMamba delivers competitive performance compared to the NeRD method. Specifically, it achieves a 0.58 dB improvement in PSNR for the image deraining task while reducing model parameters by 94.3%, computational cost by 88.3%, and inference time by 77.5%. Kui Jiang, Junjun Jiang, Shiqi Wang 0001, Wenqi Ren, Chia-Wen Lin, Zhengguo Li |
IEEE Trans. Multim. | 4 |
| 2025 | Unsupervised Diffusion-Based Degradation Modeling for Real-World Super-ResolutionabstractSingle image super-solution (SR) aims to restore a high-resolution (HR) image from a degraded low-resolution (LR) image. However, existing SR models still face a significant domain gap between synthetic and real-world datasets due to the mismatched degradation distributions, hindering SR models from achieving optimal results. In this paper, we propose an unsupervised diffusion-based degradation modeling framework (UDDM) to effectively capture real-world degradation distributions. Specifically, given unpaired LR and HR images, a diffusion-based degradation module (DDM) first models the degradation distribution by diffusing real-world LR images to downsampled LR images, which does not require HR images. It then applies reverse diffusion to generate real-world LR images from extremely downsampled HR images. This approach allows DDM to model and generate real-world degradation distributions without requiring paired data, by using extreme downsampling to link unpaired LR and HR images. Additionally, we introduce a physics-based dynamic degradation module (P-DDM) that adaptively models content-aware degradation, ensuring both content and structural accuracy. Finally, the LR images generated by DDM and P-DDM are adaptively weighted to produce the final LR images, which are paired with the given HR images for training the SR network. Extensive experiments across multiple real-world datasets demonstrate that our framework achieves state-of-the-art performance in both qualitative and quantitative comparison. Yuying Chen, Mingde Yao, Renjing Pei, Jinjing Zhao, Wenqi Ren |
AAAI | 6 |
| 2025 | Ultra-High-Definition Dynamic Multi-Exposure Image Fusion via Infinite Pixel LearningabstractWith the continuous improvement of device imaging resolution, the popularity of Ultra-High-Definition (UHD) images is increasing. Unfortunately, existing methods for fusing multi-exposure images in dynamic scenes are designed for low-resolution images, which makes them inefficient for generating high-quality UHD images on a resource-constrained device. To alleviate the limitations of extremely long-sequence inputs, inspired by the Large Language Model (LLM) for processing infinitely long texts, we propose a novel learning paradigm to achieve UHD multi-exposure dynamic scene image fusion on a single consumer-grade GPU, named Infinite Pixel Learning (IPL). The design of our approach comes from three key components: The first step is to slice the input sequences to relieve the pressure generated by the model processing the data stream; Second, we develop an attention cache technique, which is similar to the KV cache for infinite data stream processing; Finally, we design a method for attention cache compression to alleviate the storage burden of the cache on the device. In addition, we provide a new UHD benchmark to evaluate the effectiveness of our method. Extensive experimental results show that our method maintains high-quality visual performance while fusing UHD dynamic multi-exposure images in real-time (>40fps) on a single consumer-grade GPU. Xingchi Chen, Zhuoran Zheng, Xuerui Li, Yuying Chen, Wenqi Ren |
AAAI | 6 |
| 2025 | Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake DetectionabstractRecent face forgery detection methods based on disentangled representation learning utilize paired images for cross-reconstruction, aiming to extract forgery-relevant attributes and forgery-irrelevant content. However, there still exist the following issues that may comprise the detector performance: 1) using information-dense images as the decoupling targets increases the decoupling difficulty; 2) the extracted attribute features are reconstruction-irrelevant rather than forgery-relevant, and single-scale forgery representation decoupling cannot capture sufficient discriminative information; 3) the generalization performance of decoupled attribute features is poor as the detector focuses on learning specific artifact types in the training set. To address these issues, we propose a novel disentangled representation learning framework for deepfake detection. First, we extract features by partitioning the dense information within the image, focusing independently on texture, color, or edges. These features are then used as the decoupling targets rather than the images themselves, which could mitigate the decoupling difficulty. Second, we extend reconstruction loss from image-level to feature-level, thus extending the forgery representation decoupling from single-scale to multi-scale. Third, we propose a critical forgetting mechanism that forces the detector to forget the most salient features during training, which correspond to specific forgery artifact types in the training set. Extensive experimental results validate the efficacy of the proposed method. Wenqi Ren, Jianshu Li, Wei Wang 0335, Xiaochun Cao |
AAAI | 2 |
| 2025 | RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-ResolutionabstractBenefiting from their powerful generative capabilities, pretrained diffusion models have garnered significant attention for real-world image super-resolution (Real-SR). Existing diffusion-based SR approaches typically utilize semantic information from degraded images and restoration prompts to activate prior for producing realistic high-resolution images. However, general-purpose pretrained diffusion models, not designed for restoration tasks, often have suboptimal prior, and manually defined prompts may fail to fully exploit the generated potential. To address these limitations, we introduce RAP-SR, a novel restoration prior enhancement approach in pretrained diffusion models for Real-SR. First, we develop the High-Fidelity Aesthetic Image Dataset (HFAID), curated through a Quality-Driven Aesthetic Image Selection Pipeline (QDAISP). Our dataset not only surpasses existing ones in fidelity but also excels in aesthetic quality. Second, we propose the Restoration Priors Enhancement Framework, which includes Restoration Priors Refinement (RPR) and Restoration-Oriented Prompt Optimization (ROPO) modules. RPR refines the restoration prior using the HFAID, while ROPO optimizes the unique restoration identifier, improving the quality of the resulting images. RAP-SR effectively bridges the gap between general-purpose models and the demands of Real-SR by enhancing restoration prior. Leveraging the plug-and-play nature of RAP-SR, our approach can be seamlessly integrated into existing diffusion-based SR methods, boosting their performance. Extensive experiments demonstrate its broad applicability and state-of-the-art results. Qingnan Fan, Feng Huang 0007, Wenqi Ren |
AAAI | 6 |
| 2025 | LVPTrack: High Performance Domain Adaptive UAV Tracking with Label Aligned Visual Prompt TuningabstractVisual object tracking is essentially crucial for unmanned aerial vehicles (UAVs). Despite the substantial progress, most of the existing UAV trackers are designed for well-conditioned daytime data, while for the scenarios in challenging weather condition, e.g. foggy or nighttime environment, the tremendous domain gap leads to significant performance degradation. To address this issue, in this paper, we propose a novel robust UAV tracker termed LVPTrack, which conducts high quality label-aligned visual prompt tuning to adapt to various challenging weather conditions. Specifically, we first synthesize the sequential foggy and nighttime video frames to assist the model training. A domain adaptive teacher-student network is utilized to distill the hierarchical visual semantic of the target objects in cross-domain scenarios. Then we propose a target-aware pseudo-label voting (PLV) strategy to alleviate the target-level misalignment in the dual domains. Furthermore, we propose a dynamic aggregated prompt (DAP) module to facilitate the appearance variation adaptation of the target object in challenging scenarios. Extensive experiments demonstrate that our tracker achieves superior performance over existing state-of-the-art UAV trackers. Hongjing Wu, Siyuan Yao, Feng Huang 0007, Linchao Zhang, Zhuoran Zheng, Wenqi Ren |
AAAI | 7 |
| 2025 | Gaze Label Alignment: Alleviating Domain Shift for Gaze EstimationabstractGaze estimation methods encounter significant performance deterioration when being evaluated across different domains, because of the domain gap between the testing and training data. Existing methods try to solve this issue by reducing the deviation of data distribution, however, they ignore the existence of label deviation in the data due to the acquisition mechanism of the gaze label and the individual physiological differences. In this paper, we first point out that the influence brought by the label deviation cannot be ignored, and propose a gaze label alignment algorithm (GLA) to eliminate the label distribution deviation. Specifically, we first train the feature extractor on all domains to get domain invariant features, and then select an anchor domain to train the gaze regressor. We predict the gaze label on remaining domains and use a mapping function to align the labels. Finally, these aligned labels can be used to train gaze estimation models. Therefore, our method can be combined with any existing method. Experimental results show that our GLA method can effectively alleviate the label distribution shift, and SOTA gaze estimation methods can be further improved obviously. Guanzhong Zeng, Zefu Xu, Pengwei Yin, Wenqi Ren, Di Xie |
AAAI | 5 |
| 2025 | Dual Prompting Image Restoration with Diffusion TransformersabstractRecent state-of-the-art image restoration methods mostly adopt latent diffusion models with U-Net backbones, yet still facing challenges in achieving high-quality restoration due to their limited capabilities. Diffusion transformers (DiTs), like SD3, are emerging as a promising alternative because of their better quality with scalability. In this paper, we introduce DPIR (Dual Prompting Image Restoration), a novel image restoration method that effectivly extracts conditional information of low-quality images from multiple perspectives. Specifically, DPIR consits of two branches: a low-quality image conditioning branch and a dual prompting control branch. The first branch utilizes a lightweight module to incorporate image priors into the DiT with high efficiency. More importantly, we believe that in image restoration, textual description alone cannot fully capture its rich visual characteristics. Therefore, a dual prompting module is designed to provide DiT with additional visual cues, capturing both global context and local appearance. The extracted global-local visual prompts as extra conditional control, alongside textual prompts to form dual prompts, greatly enhance the quality of the restoration. Extensive experimental results demonstrate that DPIR delivers superior image restoration performance. Dehong Kong, Zhixin Wang, Renjing Pei, Wenqi Ren |
CVPR | 7 |
| 2025 | Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video DerainingabstractSignificant progress has been made in video restoration under rainy conditions over the past decade, largely propelled by advancements in deep learning. Nevertheless, existing methods that depend on paired data struggle to generalize effectively to real-world scenarios, primarily due to the disparity between synthetic and authentic rain effects. To address these limitations, we propose a dual-branch spatiotemporal state-space model to enhance rain streak removal in video sequences. Specifically, we design spatial and temporal state-space model layers to extract spatial features and incorporate temporal dependencies across frames, respectively. To improve multi-frame feature fusion, we derive a dynamic stacking filter, which adaptively approximates statistical filters for superior pixel-wise feature refinement. Moreover, we develop a median stacking loss to enable semi-supervised learning by generating pseudo-clean patches based on the sparsity prior of rain. To further explore the capacity of deraining models in supporting other vision-based tasks in rainy environments, we introduce a novel real-world benchmark focused on object detection and tracking in rainy conditions. Our method is extensively evaluated across multiple benchmarks containing numerous synthetic and real-world rainy videos, consistently demonstrating its superiority in quantitative metrics, visual quality, efficiency, and its utility for downstream tasks. Shangquan Sun, Wenqi Ren, Juxiang Zhou, Jianhou Gan, Xiaochun Cao |
CVPR | 2 |
| 2025 | ECSNN: Spiking Neural Networks for Efficient Exposure Correction in Endoscopy ImagingabstractThe quality of endoscopic images is critical to the success of polyp segmentation, highlighting the need for accurate exposure correction in endoscopy. While traditional deep learning methods are effective, they demand substantial computational resources during inference. To address this, we propose the Endoscopic Exposure Correction Spiking Neural Network (ECSNN), an efficient framework designed for resource-limited devices. Our approach features a Positive Incentive Learning Module that reduces noise in input images. These enhanced features are then processed by U-Shape Networks (USNet), which leverages spiking neural networks to learn deep representations for exposure correction. Additionally, we introduce a Brightness Prompt Module consisting of two components: the Brightness Spike Encoding Module (BSEM), which encodes brightness information into spike signals, and the Brightness-Aware Prompt Block (BAPB), which adjusts exposure by guiding the network through brightness-aware attention. We evaluate ECSNN on the Endo4IE and ECSEG datasets, where it outperforms six state-of-the-art methods and demonstrates its practical utility in clinical diagnosis. Jun Zhang 0011, Zhuoran Zheng, Jingang Zhang, Wenqi Ren |
ICASSP | 4 |
| 2025 | Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
Xiaojun Jia, Ranjie Duan, Xinfeng Li, Yihao Huang 0001, Xiaoshuang Jia, Zhixuan Chu, Wenqi Ren |
ICCV | 8 |
| 2025 | UMDATrack: Unified Multi-Domain Adaptive Tracking under Adverse Weather ConditionsabstractVisual object tracking has gained promising progress in past decades. Most of the existing approaches focus on learning target representation in well-conditioned daytime data, while for the unconstrained real-world scenarios with adverse weather conditions, e.g. nighttime or foggy environment, the tremendous domain shift leads to significant performance degradation. In this paper, we propose UMDATrack, which is capable of maintaining high-quality target state prediction under various adverse weather conditions within a unified domain adaptation framework. Specifically, we first use a controllable scenario generator to synthesize a small amount of unlabeled videos (less than 2% frames in source daytime datasets) in multiple weather conditions under the guidance of different text prompts. Afterwards, we design a simple yet effective domain-customized adapter (DCA), allowing the target objects' representation to rapidly adapt to various weather conditions without redundant model updating. Furthermore, to enhance the localization consistency between source and target domains, we propose a target-aware confidence alignment module (TCA) following optimal transport theorem. Extensive experiments demonstrate that UMDATrack can surpass existing advanced visual trackers and lead new state-of-the-art performance by a significant margin. Our code is available at https://github.com/Z-Z188/UMDATrack. Siyuan Yao, Wenqi Ren, Yanyang Yan, Xiaochun Cao |
ICCV | 4 |
| 2025 | LLDNet: Joint Low-light Enhancement and Local Motion Deblurring in the DarkabstractLocal motion blur in the dark often occurs in the real world due to long exposure, leading to serious challenges in real-world activities, such as night photography and autonomous driving. Although existing local motion deblurring methods and low-light enhancement methods can solve each problem separately. The simple cascade of these methods cannot handle the joint degradation of the low-light and local motion blur. Therefore, in this paper, we jointly address both low-light conditions and local motion blur, aiming to achieve efficient restoration of low-light local motion blur images. Specifically, to solve the attention noise issue, we introduce the channel differential Transformer that guides the model to focus on essential regions and channels during image enhancement. Besides, we present the window differential Transformer that interpolates the window differential self-attention and the multi-scale feed-forward module to focus on local blurry regions. In addition, the phase content-aware fusion module is employed to enhance the transmission of phase information from encoder to decoder. Extensive experimental results demonstrate the effectiveness of our method for low-light local motion deblurring on both synthetic and real-world datasets. Haigen Liu, Yanyang Yan, Wenqi Ren |
ICME | 3 |
| 2025 | FaceInsight: A Multimodal Large Language Model for Face PerceptionabstractRecent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing inaccurate or misleading responses to face-specific queries. To address this gap, we propose FaceInsight, a versatile face perception MLLM that provides fine-grained information. Our approach introduces visual textual alignment of facial knowledge to model both uncertain dependencies and deterministic relationships among facial information, mitigating the limitations of language-driven reasoning. Additionally, we incorporate face segmentation maps as an auxiliary perceptual modality, enriching visual input with localized structural cues to enhance semantic understanding. Comprehensive experiments show that FaceInsight consistently outperforms nine compared MLLMs under both training-free and fine-tuned settings. Jingzhi Li 0002, Changjiang Luo, Ruoyu Chen 0001, Hua Zhang 0008, Wenqi Ren, Jianhou Gan, Xiaochun Cao |
ACM Multimedia | 5 |
| 2025 | Detecting Synthetic Image by Cross-Modal Commonality InteractionabstractExisting synthetic image detection approaches can be categorized into three paradigms: spatial, frequency, and fingerprint-based methods. Our analysis reveals a fundamental commonality across these paradigms: a significant reliance on high-frequency image components. This observation highlights the discriminative power of high-frequency information for this task and provides a strong rationale for learning generalized artifact representations based on multi-modal fusion strategies. Building on this insight, we introduce a multi-modal high-frequency interactive detection framework for general synthetic image detection. This framework explicitly integrates high-frequency information from both the spatial and frequency domains. Specifically, its spatial processing branch incorporates a novel high-frequency self-enhancement module to bolster local high-frequency representations. Concurrently, the frequency processing branch utilizes a multi-scale frequency information enhancement module to capture diverse contextual cues. At the feature fusion stage, we propose a pooling-guided cross-modal high-frequency interaction module, which dynamically weights cross-modal information to further reinforce salient high-frequency representations. Extensive experiments on public datasets demonstrate that our proposed framework achieves state-of-the-art performance in real-world detection scenarios. Wenqi Ren, Wei Wang 0335, Linchao Zhang, Xiaochun Cao |
ACM Multimedia | 2 |
| 2025 | Bio-Inspired Image RestorationabstractImage restoration aims to recover sharp, high-quality images from degraded, low-quality inputs. Existing methods have progressively advanced from task-specific designs to general architectures, all-in-one frameworks, and composite degradation handling. Despite these advances, computational efficiency remains a critical factor for practical deployment. In this work, we present BioIR, an efficient and universal image restoration framework inspired by the human visual system. Specifically, we design two bio-inspired modules, Peripheral-to-Foveal (P2F) and Foveal-to-Peripheral (F2P), to emulate the perceptual processes of human vision, with a particular focus on the functional interplay between foveal and peripheral pathways. P2F delivers large-field contextual signals to foveal regions based on pixel-to-region affinity, while F2P propagates fine-grained spatial details through a static-to-dynamic two-stage integration strategy. Leveraging the biologically motivated design, BioIR achieves state-of-the-art performance across three representative image restoration settings: single-degradation, all-in-one, and composite degradation. Moreover, BioIR maintains high computational efficiency and fast inference speed, making it highly suitable for real-world applications. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
NeurIPS | 2 |
| 2025 | NDFormer: A Mixed-Scale Transformer with Enhanced Nonlinearity for Nighttime Image Deraining
Zhirui Liu, Shangquan Sun, Yuning Cui 0001, Dehong Kong, Wenqi Ren, Kin-Man Lam 0001 |
PRCV (9) | 6 |
| 2025 | DI-Retinex: Digital-Imaging Retinex Model for Low-Light Image Enhancement
Shangquan Sun, Wenqi Ren, Jingyang Peng, Fenglong Song, Xiaochun Cao |
Int. J. Comput. Vis. | 2 |
| 2025 | Underwater Camera: Improving Visual Perception Via Adaptive Dark Pixel Prior and Color Correction
Jingchun Zhou, Qiuping Jiang, Wenqi Ren, Kin-Man Lam 0001, Weishi Zhang |
Int. J. Comput. Vis. | 4 |
| 2025 | An unsupervised medical image registration network for intelligent medical education
Jie Mu, Jing Zhang 0037, Tiantian Yan, Wei Wang 0335, Hua Zhang 0008, Wenqi Ren |
Neural Comput. Appl. | 8 |
| 2025 | EENet: An effective and efficient network for single image dehazing
Yuning Cui 0001, Chaopeng Li, Wenqi Ren, Alois C. Knoll |
Pattern Recognit. | 4 |
| 2025 | All-Inclusive Image Enhancement for Degraded Images Exhibiting Low-Frequency CorruptionabstractIn this paper, a novel image enhancement method, called the all-inclusive image enhancement (AIIE), is proposed that can effectively enhance the degraded images for improving the visibility of image content. These imageries were acquired under various types of weather conditions such as haze, low-light, underwater, and sandstorm, etc. One commonality shared by this class of noise is that the resulted degradations on visual quality or visibility are caused by low-frequency interference. Existing image enhancement methods lack the ability to deal with all types of degradations from this class, while our proposed AIIE offers a unified treatment for them. To achieve this goal, a statistical property is obtained from the study of the discrete cosine transform (DCT) of 1,000 high- and 1000 low-quality images on their DCT domains. It shows that the normalized DCT coefficients (between 0 and 1) of high-quality images has about 95% fall in the interval [0, 0.2]; for low-quality images, almost all the coefficients are in the same interval. This fundamental property, called the DCT prior (DCT-P), is instrumental to the development of our AIIE algorithm proposed in this paper. Since the proposed DCT-P delineates the attributes of high- and low-quality images clearly, it becomes a highly effective ‘tool’ to convert low-quality images to its enhanced version. Extensive experimental results have clearly validated the superior performance of the AIIE conducted on different types of deteriorated images in terms of visual quality and efficiency as well as significant advantages on computational complexity, which is essential for real-time applications. Mingye Ju, Chunming He, Can Ding 0002, Wenqi Ren, Lin Zhang 0014, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Single Image Dehazing Using Fuzzy Region Segmentation and Haze Density DecompositionabstractImages captured under haze weather conditions usually suffer from visual quality degradations, such as blurred details, faded colors, and decreased saturation. Existing physicsbased dehazing methods mainly have two drawbacks: 1) the atmospheric light is treated as a constant for the entire image, and 2) pixel-or patch-based strategies are employed to estimate the model parameters, resulting in inaccurate haze density estimations. Therefore, these methods may lead to over-dehazing or under-dehazing due to insufficient utilization of features from regions with similar haze densities. To address these issues, a novel single image dehazing framework based on fuzzy region segmentation and haze density decomposition is proposed. Specifically, a region-based physical model that considers the non-uniform atmospheric light is first constructed based on the classic atmospheric scattering model. Then, a fuzzy segmentation algorithm is improved to divide the input hazy image into several separate regions. Subsequently, we formulate a simple linear relationship between the atmospheric light and brightness to estimate region-based atmospheric light. On the other hand, we develop a novel haze density decomposition algorithm based on boundary constraints to separate the atmospheric veil into two components: thin part and dense part. Three haze-related features, contrast, gradient and clarity, are extracted from the input hazy image to construct weight maps and a multi-scale fusion is further exploited to combine weight maps and boundary veils to acquire the refined atmospheric veil. Finally, the model inversion is performed to acquire the haze-free result. Experiments on six diverse hazy datasets demonstrate that the proposed algorithm outperforms several state-of-the-art dehazing methods in both visual quality and objective evaluation. Yun Liu 0002, Wenqi Ren, Babak Shiri, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | LIEDNet: A Lightweight Network for Low-Light Enhancement and DeblurringabstractImages captured at nighttime often face challenges such as low light and blur, primarily caused by dim environments and the frequent use of long exposure. Existing methods either handle the two types of degradations independently or rely on carefully designed priors generated by complex mechanisms, resulting in poor generalization ability and high model complexity. To address these challenges, we propose an end-to-end framework named LIEDNet to efficiently and effectively restore high-quality images on both real-world and synthetic data. Specifically, the introduced LIEDNet consists of three essential components: the Visual State Space Module (VSSM), the Local Feature Module (LFM), and the Dual Gated-Dconv Feedforward Network (DGDFFN). The integration of VSSM and LFM enables the model to capture both global and local features while maintaining low computational overhead. Additionally, the DGDFFN improves image fidelity by extracting multi-scale structural information. Extensive experiments on real-world and synthetic datasets demonstrate the superior performance of LIEDNet in restoring low-light, blurry images. The code is available athttps://github.com/MingyuLiu1/LIEDNethttps://github.com/MingyuLiu1/LIEDNet. Yuning Cui 0001, Wenqi Ren, Juxiang Zhou, Alois C. Knoll |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | From Dynamic to Static: Stepwisely Generate HDR Image for Ghost RemovalabstractGenerating high-quality high dynamic range (HDR) images in dynamic scenes is particularly challenging due to the influence of large motion. Despite the effectiveness of existing deep learning methods, they still suffer from ghosting artifacts when saturation and motion coexist. Inspired by fusion on static scenes, we propose an inpainting and fusion strategy to enhance the quality of the generated HDR images. The proposed method consists of pseudo-static LDR generation and detail-guided HDR generation, which creates pseudo-static images and then generates ghost-free HDR images. Specifically, the pseudo-static LDR generation network utilizes semantic information to identify the motion regions, and employs a diffusion model-based inpainting approach to produce pseudo-static LDR images that closely resemble real scenes. In the detail-guided HDR generation network, we employ a detail enhancement module to refine diverse high-frequency features with detailed information extracted from pseudo-static LDR images, which effectively enhances the visual quality. Extensive experiments on four public datasets demonstrate the superiority of the proposed method, both quantitatively and qualitatively. Qingsen Yan, Kangzhen Yang, Tao Hu 0013, Genggeng Chen, Kexin Dai, Peng Wu 0015, Wenqi Ren, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Exploring Fuzzy Priors From Multimapping GAN for Robust Image DehazingabstractSingle image dehazing has been extensively studied. While convolutional neural networks (CNNs) have driven notable progress in single image dehazing, their performance remains fundamentally constrained by the limited local receptive fields of convolutional operations, which impede the capture of global structural dependencies. In contrast, generative adversarial networks (GANs) have demonstrated exceptional capabilities in image synthesis, offering global insights into structure, texture, and color. The fuzzy prior, a probabilistic knowledge acquired through adversarial training in GANs, plays a pivotal role in robust dehazing. Motivated by this, we propose the fuzzy prior guided dehazing network (FPGDN). Our framework begins with a novel module that distills the fuzzy prior by translating an edge map into a color image, simultaneously capturing global structural, local textural, and color information. Subsequently, a dehazing network is constructed, leveraging this fuzzy prior. While the fuzzy prior captures rich color and texture features, the generated images may exhibit color shifts relative to the original scene. To remedy this, a CNN network is employed to capture local nuances and refine the dehazing outcome. Extensive experiments substantiate that the proposed FPGDN achieves superior dehazing performance on a variety of real and synthetic hazy images. Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, Li Zhao 0005, En Fan, Feng Huang 0007 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | Exploring the Potential of Pooling Techniques for Universal Image RestorationabstractImage restoration involves recovering a clean image from its degraded counterpart. In recent years, we have witnessed a paradigm shift from convolutional neural networks to Transformers, which have quadratic complexity with respect to the input size. Instead of designing more complex modules based on recent techniques, this paper presents an efficient and effective mechanism for image restoration by exploring the potential of ubiquitous pooling techniques. We leverage different pooling operators as tools for implicit dual-domain representation learning. Specifically, the average and max pooling can be used as extractors for implicit low- and high-frequency signals, respectively. Then, we utilize lightweight learnable parameters to modulate the resulting frequency components. Furthermore, the intermediate high-frequency features can serve as attention maps to highlight the spatial edge information. Our pooling module is built by incorporating the aforementioned dual-domain modulation across multiple scales and various shapes. We demonstrate the effectiveness of our module in single-degradation, composite-degradation, and all-in-one image restoration tasks. Extensive experimental results show that the resulting network achieves state-of-the-art performance on 15 datasets for five single-degradation and two composite-degradation image restoration tasks by deploying our module. Moreover, our method can be extended to all-in-one scenarios and performs favorably against state-of-the-art all-in-one algorithms under two settings. The code is available at https://github.com/c-yn/PoolNet. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Image Process. | 2 |
| 2025 | Toward Better Than Pseudo-Reference in Underwater Image EnhancementabstractSince degraded underwater images are not always accompanied with distortion-free counterparts in real-world situations, existing underwater image enhancement (UIE) methods are mostly learned on a paired set consisting of raw underwater images and their corresponding pseudo-reference labels. Although the existing UIE datasets manually select the best model-generated results as pseudo-References, such pseudo-reference labels do not always exhibit perfect visual quality. Therefore, it would be interesting to investigate whether it is possible to break through the performance bottleneck of UIE networks trained with imperfect pseudo-references. Motivated by these facts, this paper focuses on innovating more advanced loss functions rather than designing more complex network architectures. Specifically, a plug-and-play hybrid Performance SurPassing Loss (PSPL), consisting of a Quality Score Comparison Loss (QSCL) and a scene Depth-aware Unpaired Contrastive Loss (DUCL), is formulated to guide the training of UIE network. Functionally, QSCL aims to guide the UIE network to generate enhanced results with better visual quality than pseudo-references by constructing image quality score comparison losses from both image-level and region-level. Nevertheless, only using QSCL cannot guarantee obtaining desired results for those severely degraded distant regions. Therefore, we also design a tailored DUCL to handle this challenging issue from the scene depth perspective, i.e., DUCL encourages the distant regions of the enhanced results to be closer to the high-quality nearby regions (pull) and far away from the low-quality distant regions (push) of the pseudo-references. Extensive experimental results demonstrate the advantage of using PSPL over the state-of-the-arts even with an extremely simple and lightweight UIE network. The source code will be released at https://github.com/lewis081/PSPL. Yi Liu 0085, Qiuping Jiang, Xingbo Li, Ting Luo 0001, Wenqi Ren |
IEEE Trans. Image Process. | 5 |
| 2025 | NDMamba: Dual-Prior State-Space Model for Nighttime DerainingabstractRecent advancements in deep learning, particularly through Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have led to significant progress in nighttime image deraining. However, current architectures still struggle to strike an optimal balance between computational efficiency and restoration performance. Moreover, existing methods often fail to fully exploit the intrinsic characteristics of low-light conditions and inadequately model the interaction between rain and illumination. To overcome these challenges, we propose NDMamba, a dual-prior-guided state-space model that addresses nighttime deraining by incorporating degradation cues related to both lighting and rain distribution. Inspired by the Retinex theory, which suggests that rain streak distribution is influenced by the reflectance component of a scene, we propose a Prior Extraction Module (PEM) to jointly model lighting conditions and rain degradation. Furthermore, we design a Prior-Guided Mamba Block (PGMB), which comprises a Lighting-Adaptive Vision State-Space Module (LVSSM) that incorporates illumination priors, and a Rain Distribution Guidance Module (RDGM) to enhance local features in a more refined manner. Extensive experiments demonstrate that NDMamba outperforms state-of-the-art methods on both synthetic and real-world benchmark datasets. Our code is publicly available at https://github.com/tandaily/NDMamba. Zhirui Liu, Shangquan Sun, Chaopeng Li, Wenqi Ren |
IEEE Trans. Image Process. | 6 |
| 2025 | CAN: Cascade Augmentations Against Noise for Image RestorationabstractImage restoration aims to recover the latent clean image from a degraded counterpart. In general, the prevailing state-of-the-art image restoration methods concentrate on solving only a specific degradation type according to the task, e.g., deblurring or deraining. However, if the corresponding well-trained frameworks confront other real-world image corruptions, i.e., the corruptions are not covered in the training phase, and state-of-the-art restoration models will suffer from a lack of generalization ability. We have observed that an image restoration model can be easily confused by noise corruption. Towards improving the robustness of image restoration networks, in this paper, we focus on alleviating the corruption of noise in various image restoration tasks, which is almost inevitable in real-world scenes. To this end, we devise a novel Cascade Augmentation strategy against Noise (CAN) to enhance the robustness of specific image restoration. Specifically, the given degraded images are sequentially augmented from different perspectives, i.e., noise-aware augmentation and model-aware augmentation. The noise-aware augmentation is proposed to enrich the samples by introducing various noise operations. Moreover, to adapt to more unknown corruptions, we propose a novel model-aware augmentation mechanism, which enhances the scalability by exploring useful both spatial and frequency clues with the help of model randomness. It is worth noting that the proposed augmentation scheme is model-agnostic, and it can plug and play into arbitrary state-of-the-art image restoration architectures. In addition, we construct noise corruption benchmark datasets, derived from the validation set of standard image restoration datasets, to assist us in evaluating the robustness of restoration networks. Extensive quantitative and qualitative evaluations demonstrate that the proposed method has strong generalization capability, which can enhance the robustness of various image restoration frameworks when facing diverse noises. Yanyang Yan, Siyuan Yao, Wenqi Ren, Rui Zhang 0040, Qi Guo 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2025 | UncTrack: Reliable Visual Object Tracking With Uncertainty-Aware Prototype Memory NetworkabstractTransformer-based trackers have achieved promising success and become the dominant tracking paradigm because of their accuracy and efficiency. Despite the substantial progress, most of the existing approaches handle object tracking as a deterministic coordinate regression problem, while the target localization uncertainty has been largely overlooked, which hampers trackers' ability to maintain reliable target state prediction in challenging scenarios. To address this issue, we propose UncTrack, a novel uncertainty-aware transformer-based tracker that predicts the target localization uncertainty and incorporates this uncertainty information for accurate target state inference. Specifically, UncTrack uses a transformer encoder to perform feature interactions between the template and search images. The output features are passed into an uncertainty-aware localization decoder (ULD) to coarsely predict the corner-based localization and the corresponding localization uncertainty. Then, the localization uncertainty is sent into a prototype memory network (PMN) to excavate valuable historical information to identify whether the target state prediction is reliable. To enhance the template representation, the samples with high confidence are fed back into the prototype memory bank for memory updating, which makes the tracker more robust to challenging appearance variations. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. Our code is available at https://github.com/ManOfStory/UncTrack. Siyuan Yao, Yanyang Yan, Wenqi Ren, Xiaochun Cao |
IEEE Trans. Image Process. | 4 |
| 2025 | You Only Need Clear Images: Self-Supervised Single Image DehazingabstractImage hazing refers to adding haze to a clear image, which is important for improving the data amount and diversity of synthetic hazy images that are required to train deep image dehazing models. However, existing image hazing works generate hazy images from a given clear image with a single transmission map. This violates the fact that hazy images are diverse for a natural scene at different times. The domain shift issue between synthetic and real-world hazy images constrains the robustness of deep dehazing models when dealing with real-world hazy images. In this work, we propose an unsupervised haze generation work to synthesize multiple hazy images with diverse haze distributions from a clear image, which requires only an atmospheric scattering model without extra labeling information. Instead of estimating a transmission map from a clear image, we propose to customize the transmission maps by redefining the transmission function. In such a controllable way, hazy images with diverse haze distributions are generated, which avoids the labor-intensive collection of paired data and alleviates the common domain-shift issue of deep image dehazing. Incorporating the unsupervised hazy images generator, we also construct a generalizable self-supervised image dehazing (SSID) framework, where deep image dehazing models can be trained without any human annotations. Extensive experiments on real-world hazy images show that the proposed approach is superior to state-of-the-art unsupervised dehazing works, and achieves competitive performance with the supervised works. Moreover, the proposed SSID framework can be easily generalized to the existing deep dehazing models, greatly improving dehazing robustness on real-world hazy images. Jiyou Chen, Wenqi Ren, Qunbing Xia, Gaobo Yang |
IEEE Trans. Multim. | 2 |
| 2025 | Modumer: Modulating Transformer for Image RestorationabstractImage restoration aims to recover clean images from degraded counterparts. While Transformer-based approaches have achieved significant advancements in this field, they are limited by high complexity and their inability to capture omni-range dependencies, hindering their overall performance. In this work, we develop Modumer for effective and efficient image restoration by revisiting the Transformer block and modulation design, which processes input through a convolutional block and projection layers and fuses features via elementwise multiplication. Specifically, within each unit of Modumer, we integrate the cascaded modulation design with the downsampled Transformer block to build the attention layers, enabling omni-kernel modulation and mapping inputs into high-dimensional feature spaces. Moreover, we introduce a bioinspired parameter-sharing mechanism to attention layers, which not only enhances efficiency but also improves performance. In addition, a dual-domain feed-forward network (DFFN) strengthens the representational power of the model. Extensive experimental evaluations demonstrate that the proposed Modumer achieves state-of-the-art performance across ten datasets in five single-degradation image restoration tasks, including image motion deblurring, deraining, dehazing, desnowing, and low-light enhancement. Moreover, the model exhibits strong generalization capabilities in all-in-one image restoration tasks. Additionally, it demonstrates competitive performance in composite-degradation image restoration. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Boundary-Based Active Domain Adaptation for Semantic Segmentation Under Adverse ConditionsabstractExisting domain adaptation semantic segmentation (DASS) methods under adverse conditions often depend on pseudo-labels for network training. However, these pseudo-labels are frequently plagued by noise and bias toward high-confidence predictions, thereby impeding the enhancement of segmentation performance. This article tackles the above challenge by proposing a novel boundary-based active domain adaptation (ADA) framework, which efficiently selects both informative low-confidence samples and high-confident but misclassified samples to be labeled while maximizing the segmentation performance under a limited annotation budget. For the evaluation of sample confidence and informativeness, we first propose ranking weighted feature space impurity (RWFSI) metric to quantify category distribution among a sample's nearest neighbors within the feature space and consider the samples with higher RWFSI values as low-confidence samples around the decision boundary, which can also alleviate the category imbalance of active labels. Subsequently, we apply Gaussian mixture models (GMMs) to model the distribution across source and target domains. Using the spatial arrangement of each GMM component, we define the intraclass domain shift score (ICDSS), which identifies samples with high ICDSS values as those more likely to be high-confidence but misclassified, aiding in refining sample selection. Extensive experiments demonstrate that our method is superior to the existing state-of-the-art domain adaptation and active learning (AL) methods and comparable with those of full supervision. The code will be released at https://github.com/1061018609/BADA. Gary G. Yen, Chaoqiang Zhao, Qiyu Sun, Wenqi Ren, Lu Sheng, Yang Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Omni-Kernel Network for Image RestorationabstractImage restoration aims to reconstruct a high-quality image from a degraded low-quality observation. Recently, Transformer models have achieved promising performance on image restoration tasks due to their powerful ability to model long-range dependencies. However, the quadratically growing complexity with respect to the input size makes them inapplicable to practical applications. In this paper, we develop an efficient convolutional network for image restoration by enhancing multi-scale representation learning. To this end, we propose an omni-kernel module that consists of three branches, i.e., global, large, and local branches, to learn global-to-local feature representations efficiently. Specifically, the global branch achieves a global perceptive field via the dual-domain channel attention and frequency-gated mechanism. Furthermore, to provide multi-grained receptive fields, the large branch is formulated via different shapes of depth-wise convolutions with unusually large kernel sizes. Moreover, we complement local information using a point-wise depth-wise convolution. Finally, the proposed network, dubbed OKNet, is established by inserting the omni-kernel module into the bottleneck position for efficiency. Extensive experiments demonstrate that our network achieves state-of-the-art performance on 11 benchmark datasets for three representative image restoration tasks, including image dehazing, image desnowing, and image defocus deblurring. The code is available at https://github.com/c-yn/OKNet. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
AAAI | 2 |
| 2024 | Omnidirectional Image Super-resolution via Bi-projection FusionabstractWith the rapid development of virtual reality, omnidirectional images (ODIs) have attracted much attention from both the industrial community and academia. However, due to storage and transmission limitations, the resolution of current ODIs is often insufficient to provide an immersive virtual reality experience. Previous approaches address this issue using conventional 2D super-resolution techniques on equirectangular projection without exploiting the unique geometric properties of ODIs. In particular, the equirectangular projection (ERP) provides a complete field-of-view but introduces significant distortion, while the cubemap projection (CMP) can reduce distortion yet has a limited field-of-view. In this paper, we present a novel Bi-Projection Omnidirectional Image Super-Resolution (BPOSR) network to take advantage of the geometric properties of the above two projections. Then, we design two tailored attention methods for these projections: Horizontal Striped Transformer Block (HSTB) for ERP and Perspective Shift Transformer Block (PSTB) for CMP. Furthermore, we propose a fusion module to make these projections complement each other. Extensive experiments demonstrate that BPOSR achieves state-of-the-art performance on omnidirectional image super-resolution. The code is available at https://github.com/W-JG/BPOSR. Yuning Cui 0001, Yawen Li 0001, Wenqi Ren, Xiaochun Cao |
AAAI | 4 |
| 2024 | Insert or Attach: Taxonomy Completion via Box EmbeddingabstractTaxonomy completion, enriching existing taxonomies by inserting new concepts as parents or attaching them as children, has gained significant interest.Previous approaches embed concepts as vectors in Euclidean space, which makes it difficult to model asymmetric relations in taxonomy.In addition, they introduce pseudo-leaves to convert attachment cases into insertion cases, leading to an incorrect bias in network learning dominated by numerous pseudo-leaves.Addressing these, our framework, TAXBOX, leverages box containment and center closeness to design two specialized geometric scorers within the box embedding space.These scorers are tailored for insertion and attachment operations and can effectively capture intrinsic relationships between concepts by optimizing on a granular box constraint loss.We employ a dynamic ranking loss mechanism to balance the scores from these scorers, allowing adaptive adjustments of insertion and attachment scores.Experiments on four real-world datasets show that TAXBOX significantly outperforms previous methods, yielding substantial improvements over prior methods in real-world datasets, with average performance boosts of 6.7%, 34.9%, and 51.4% in MRR, Hit@1, and Prec@1, respectively. Yongliang Shen 0001, Wenqi Ren, Jietian Guo, Shiliang Pu, Weiming Lu 0001 |
ACL (1) | 3 |
| 2024 | PAD: Patch-Agnostic Defense against Adversarial Patch AttacksabstractAdversarial patch attacks present a significant threat to real-world object detectors due to their practical feasibil-ity. Existing defense methods, which rely on attack data or prior knowledge, struggle to effectively address a wide range of adversarial patches. In this paper, we show two inherent characteristics of adversarial patches, semantic in-dependence and spatial heterogeneity, independent of their appearance, shape, size, quantity, and location. Seman-tic independence indicates that adversarial patches oper-ate autonomously within their semantic context, while spatial heterogeneity manifests as distinct image quality of the patch area that differs from original clean image due to the independent generation process. Based on these observations, we propose PAD, a novel adversarial patch localization and removal method that does not require prior knowledge or additional training. PAD offers patch-agnostic de-fense against various adversarial patches, compatible with any pretrained object detectors. Our comprehensive digital and physical experiments involving diverse patch types, such as localized noise, printable, and naturalistic patches, ex-hibit notable improvements over state-of-the-art works. Our code is available at https://github.com/Lihua-Jing/PAD. Lihua Jing, Rui Wang 0032, Wenqi Ren, Xin Dong 0015, Cong Zou |
CVPR | 3 |
| 2024 | Logit Standardization in Knowledge DistillationabstractKnowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between their logits in terms of logit range and variance. This side-effect limits the performance of student, considering the capacity discrepancy between them and the finding that the innate logit relations of teacher are sufficient for student to learn. To address this issue, we propose setting the temperature as the weighted standard deviation of logit and performing a plug-and-play Z-score pre-process of logit standardization before applying softmax and Kullback-Leibler divergence. Our pre-process enables student to focus on essential logit relationsfrom teacher rather than requiring a magnitude match, and can improve the performance of existing logit-based distillation methods. We also show a typical case where the conventional setting of sharing temperature between teacher and student cannot reliably yield the authentic dis-tillation evaluation; nonetheless, this challenge is success-fully alleviated by our Z-score. We extensively evaluate our method for various student and teacher models on CIFAR-100 and ImageNet, showing its significant superiority. The vanilla knowledge distillation powered by our pre-process can achieve favorable performance against state-of-the-art methods, and other distillation variants can obtain considerable gain with the assistance of our pre-process. The codes, pre-trained models and logs are released on Github. Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Rui Wang 0032, Xiaochun Cao |
CVPR | 2 |
| 2024 | CountFormer: Multi-view Crowd Counting Transformer
Hong Mo, Jianchao Tan, Qiong Gu, Bo Hang, Wenqi Ren |
ECCV (52) | 7 |
| 2024 | Restoring Images in Adverse Weather Conditions via Histogram Transformer
Shangquan Sun, Wenqi Ren, Xinwei Gao, Rui Wang 0032, Xiaochun Cao |
ECCV (22) | 2 |
| 2024 | Frequency Aware and Graph Fusion Network for Polyp SegmentationabstractPolyp segmentation plays a crucial role in the prevention of colon cancer. However, the diverse shapes of polyps and their similarity to normal areas in terms of color and texture make polyp segmentation a challenging task. Currently, most polyp segmentation methods solely focus on spatial domain features, ignoring the valuable features in the frequency domain. Consequently, many polyp segmentation algorithms struggle with the camouflage of polyps. To tackle this issue, we propose the Frequency Aware and Graph Fusion Network (FAGF-Net). Specifically, it begins with a Frequency-based Global Extraction Module (FGEM), which provides an initial estimation of the polyp regions to guide subsequent modules. Next, we design a Frequency-based Feature Attention Module (FFAM) that leverages amplitude and phase information to amplify appearance differences and enhance semantic representations. Moreover, we present a Graph-based Fusion Module (GFM), which infers the geometric characteristic of polyps through aggregating and interacting with enhanced features. Extensive experiments show that our method outperforms state-of-the-art methods with better quantitative and qualitative evaluations. Yan Li 0196, Zhuoran Zheng, Wenqi Ren, Yunfeng Nie, Jingang Zhang, Xiuyi Jia |
ICASSP | 3 |
| 2024 | Hybrid Frequency Modulation Network for Image Restoration
Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IJCAI | 3 |
| 2024 | SensitiveHUE: Multivariate Time Series Anomaly Detection by Enhancing the Sensitivity to Normal PatternsabstractUnsupervised anomaly detection in multivariate time series (MTS) has always been a challenging problem, and the modeling based on reconstruction has garnered significant attention. The insensitivity of these methods towards normal patterns poses challenges in distinguishing between normal and abnormal points. Firstly, the general reconstruction strategies may exhibit limited sensitivity to spatio-temporal dependencies, and their performance remains largely unaffected by such dependencies. Secondly, most methods fail to model the heteroscedastic uncertainty in MTS, hindering their abilities to derive a distinguishable criterion. For instance, normal data with high noise levels may lead to detection failure due to excessively high reconstruction errors. In this work, we emphasize the necessity of sensitivity to normal patterns, which could improve the discrimination between normal and abnormal points remarkably. To this end, we propose SensitiveHUE, a probabilistic network by implementing both reconstruction and heteroscedastic uncertainty estimation. Its core includes a statistical feature removal strategy to ensure the dependency sensitive property, and a novel MTS-NLL loss for modeling the normal patterns in important regions. Experimental results demonstrate that SensitiveHUE exhibits nontrivial sensitivity to normal patterns and outperforms the existing state-of-the-art alternatives by a large margin. Code is publicly available at this URL\footnotehttp://github.com/yuesuoqingqiu/SensitiveHUE. Yuye Feng, Wei Zhang 0387, Yao Fu 0006, Wenqi Ren |
KDD | 6 |
| 2024 | 3D Priors-Guided Diffusion for Blind Face RestorationabstractBlind face restoration endeavors to restore a clear face image from a degraded counterpart. Recent approaches employing Generative Adversarial Networks (GANs) as priors have demonstrated remarkable success in this field. However, these methods encounter challenges in achieving a balance between realism and fidelity, particularly in complex degradation scenarios. To inherit the exceptional realism generative ability of the diffusion model and also constrained by the identity-aware fidelity, we propose a novel diffusion-based framework by embedding the 3D facial priors as structure and identity constraints into a denoising diffusion process. Specifically, in order to obtain more accurate 3D prior representations, the 3D facial image is reconstructed by a 3D Morphable Model (3DMM) using an initial restored face image that has been processed by a pretrained restoration network. A customized multi-level feature extraction method is employed to exploit both structural and identity information of 3D facial images, which are then mapped into the noise estimation process. In order to enhance the fusion of identity information into the noise estimation, we propose a Time-Aware Fusion Block (TAFB). This module offers a more efficient and adaptive fusion of weights for denoising, considering the dynamic nature of the denoising process in the diffusion model, which involves initial structure refinement followed by texture detail enhancement. Extensive experiments demonstrate that our network performs favorably against state-of-the-art algorithms on synthetic and real-world datasets for blind face restoration. Xiaobin Lu, Xiaobin Hu, Jun Luo 0012, Ben Zhu, Yaping Ruan, Wenqi Ren |
ACM Multimedia | 6 |
| 2024 | Cross-Class Domain Adaptive Semantic Segmentation with Visual Language ModelsabstractThis paper addresses the issue of cross-class domain adaptation (CCDA) in semantic segmentation, where the target domain contains both shared and novel classes that are either unlabeled or unseen in the source domain. This problem is challenging, as the absence of labels for novel classes hampers the effective solutions of both cross-domain and cross-class problems. Since Visual Language Models (VLMs) have exhibited impressive generalization across diverse data distributions and are capable of generating zero-shot predictions without requiring task-specific training examples, we propose a label alignment method by leveraging VLMs to relabel pseudo labels for novel classes. Considering that VLMs typically provide only image-level predictions, we embed a two-stage method to enable fine-grained semantic segmentation and design a threshold based on the uncertainty of pseudo labels to exclude noisy VLM predictions. To further augment the supervision of novel classes, we devise memory banks with an adaptive update scheme to effectively manage accurate VLM predictions, which are then resampled to increase the sampling probability of novel classes. Through comprehensive experiments, we demonstrate the effectiveness and versatility of our proposed method across various CCDA scenarios. Wenqi Ren, Ruihao Xia, Meng Zheng 0002, Ziyan Wu 0001, Yang Tang 0001, Nicu Sebe |
ACM Multimedia | 1 |
| 2024 | RFFNet: Towards Robust and Flexible Fusion for Low-Light Image DenoisingabstractLow-light environments will introduce high-intensity noise into images. Containing fine details with reduced noise, near-infrared/flash images can serve as guidance to facilitate noise removal. However, existing fusion-based methods fail to effectively suppress artifacts caused by inconsistency between guidance/noisy image pairs and do not fully excavate the useful information contained in guidance images. In this paper, we propose a robust and flexible fusion network (RFFNet) for low-light image denoising. Specifically, we present a multi-scale inconsistency calibration module to address inconsistency before fusion by first mapping the guidance features to multi-scale spaces and calibrating them with the aid of pre-denoising features in a coarse-to-fine manner. Furthermore, we develop a dual-domain adaptive fusion module to adaptively extract useful high-/low-frequency signals from the guidance features and then highlight the informative frequencies. Extensive experimental results demonstrate that our method achieves state-of-the-art performance on NIR-guided RGB image denoising and flash-guided no-flash image denoising. Yuning Cui 0001, Yawen Li 0001, Yaping Ruan, Ben Zhu, Wenqi Ren |
ACM Multimedia | 6 |
| 2024 | EnsIR: An Ensemble Algorithm for Image Restoration via Gaussian Mixture ModelsabstractImage restoration has experienced significant advancements due to the development of deep learning. Nevertheless, it encounters challenges related to ill-posed problems, resulting in deviations between single model predictions and ground-truths. Ensemble learning, as a powerful machine learning technique, aims to address these deviations by combining the predictions of multiple base models. Most existing works adopt ensemble learning during the design of restoration models, while only limited research focuses on the inference-stage ensemble of pre-trained restoration models. Regression-based methods fail to enable efficient inference, leading researchers in academia and industry to prefer averaging as their choice for post-training ensemble. To address this, we reformulate the ensemble problem of image restoration into Gaussian mixture models (GMMs) and employ an expectation maximization (EM)-based algorithm to estimate ensemble weights for aggregating prediction candidates. We estimate the range-wise ensemble weights on a reference set and store them in a lookup table (LUT) for efficient ensemble inference on the test set. Our algorithm is model-agnostic and training-free, allowing seamless integration and enhancement of various pre-trained image restoration models. It consistently outperforms regression-based methods and averaging ensemble approaches on 14 benchmarks across 3 image restoration tasks, including super-resolution, deblurring and deraining. The codes and all estimated weights have been released in Github. Shangquan Sun, Wenqi Ren, Zikun Liu 0001, Hyunhee Park, Rui Wang 0032, Xiaochun Cao |
NeurIPS | 2 |
| 2024 | Photo realistic synthetic dataset and multi-scale attention dehazing network
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, LinLin Shen, Li Zhao 0005, Jun Zhang 0011 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Fast Ultra High-Definition Video Deblurring via Multi-scale Separable Network
Wenqi Ren, Senyou Deng, Kaihao Zhang, Fenglong Song, Xiaochun Cao, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Image Restoration via Frequency SelectionabstractImage restoration aims to reconstruct the latent sharp image from its corrupted counterpart. Besides dealing with this long-standing task in the spatial domain, a few approaches seek solutions in the frequency domain by considering the large discrepancy between spectra of sharp/degraded image pairs. However, these algorithms commonly utilize transformation tools, e.g., wavelet transform, to split features into several frequency parts, which is not flexible enough to select the most informative frequency component to recover. In this paper, we exploit a multi-branch and content-aware module to decompose features into separate frequency subbands dynamically and locally, and then accentuate the useful ones via channel-wise attention weights. In addition, to handle large-scale degradation blurs, we propose an extremely simple decoupling and modulation module to enlarge the receptive field via global and window-based average pooling. Furthermore, we merge the paradigm of multi-stage networks into a single U-shaped network to pursue multi-scale receptive fields and improve efficiency. Finally, integrating the above designs into a convolutional backbone, the proposed Frequency Selection Network (FSNet) performs favorably against state-of-the-art algorithms on 20 different benchmark datasets for 6 representative image restoration tasks, including single-image defocus deblurring, image dehazing, image motion deblurring, image desnowing, image deraining, and image denoising. Yuning Cui 0001, Wenqi Ren, Xiaochun Cao, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Revitalizing Convolutional Network for Image RestorationabstractImage restoration aims to reconstruct a high-quality image from its corrupted version, playing essential roles in many scenarios. Recent years have witnessed a paradigm shift in image restoration from convolutional neural networks (CNNs) to Transformer-based models due to their powerful ability to model long-range pixel interactions. In this paper, we explore the potential of CNNs for image restoration and show that the proposed simple convolutional network architecture, termed ConvIR, can perform on par with or better than the Transformer counterparts. By re-examing the characteristics of advanced image restoration algorithms, we discover several key factors leading to the performance improvement of restoration models. This motivates us to develop a novel network for image restoration based on cheap convolution operators. Comprehensive experiments demonstrate that our ConvIR delivers state-of-the-art performance with low computation complexity among 20 benchmark datasets on five representative image restoration tasks, including image dehazing, image motion/defocus deblurring, image deraining, and image desnowing. Yuning Cui 0001, Wenqi Ren, Xiaochun Cao, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Correcting Optical Aberration via Depth-Aware Point Spread FunctionsabstractOptical aberration is a ubiquitous degeneration in realistic lens-based imaging systems. Optical aberrations are caused by the differences in the optical path length when light travels through different regions of the camera lens with different incident angles. The blur and chromatic aberrations manifest significant discrepancies when the optical system changes. This work designs a transferable and effective image simulation system of simple lenses via multi-wavelength, depth-aware, spatially-variant four-dimensional point spread functions (4D-PSFs) estimation by changing a small amount of lens-dependent parameters. The image simulation system can alleviate the overhead of dataset collecting and exploiting the principle of computational imaging for effective optical aberration correction. With the guidance of domain knowledge about the image formation model provided by the 4D-PSFs, we establish a multi-scale optical aberration correction network for degraded image reconstruction, which consists of a scene depth estimation branch and an image restoration branch. Specifically, we propose to predict adaptive filters with the depth-aware PSFs and carry out dynamic convolutions, which facilitate the model's generalization in various scenes. We also employ convolution and self-attention mechanisms for global and local feature extraction and realize a spatially-variant restoration. The multi-scale feature extraction complements the features across different scales and provides fine details and contextual features. Extensive experiments demonstrate that our proposed algorithm performs favorably against state-of-the-art restoration methods. Jun Luo 0012, Yunfeng Nie, Wenqi Ren, Xiaochun Cao, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Omni-Kernel Modulation for Universal Image RestorationabstractImage restoration is the process of recovering a clean image from a degraded observation. In order to achieve this, it is essential to refine features at multiple scales. This paper develops an effective omni-kernel modulation module to enhance multi-scale representation learning for image restoration. The module consists of three branches, namely global, large, and local branches, which are designed to learn global-to-local feature representations efficiently. Specifically, the global branch achieves a global perceptive field via the dual-domain channel attention and frequency-gated mechanism. Furthermore, to provide multi-grained receptive fields, the large branch is formulated using different shapes of depth-wise convolutions with unusually large kernel sizes. Moreover, we complement local information with a point-wise depth-wise convolution. Finally, we demonstrate the effectiveness of our omni-kernel modulation module in two cases: general image restoration and all-in-one image restoration tasks. Incorporating our method into a convolutional backbone results in a model that achieves state-of-the-art performance on the 15 datasets for three representative image restoration tasks, including image dehazing, desnowing, and defocus deblurring. Moreover, by integrating our module into a pure Transformer-based backbone, the model demonstrates competitive performance against state-of-the-art algorithms in two all-in-one image restoration settings: the three-task and five-task settings. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Granularity-Aware Single-Point Scene Text Spotting With Sequential Recurrence Self-AttentionabstractScene text spotting, a unified framework between text detection and text recognition, has made great progress in recent years. Existing methods usually adopt the fully-supervised learning strategy, which relies on time-consuming location annotations, particularly for scene texts with arbitrary shapes. In this paper, we propose a weakly-supervised scene text spotting method via the location labels of single points with the corresponding text transcriptions. Due to the weak location annotations for challenging scene texts, previous weakly-supervised methods adopting the convolution neural network structure make it hard to model the different-scale text feature representations under blurring or nosing scenarios. In addition, as the single-point location can only cover part of the text instance, it will burden the confusion of sequential-like scene text recognition. To address these issues, we present a novel sequential recurrence self-attention for granularity-aware single-point scene text spotting. Specifically, we first enhance the scene text feature representations with different scales by integrating the global intra-interaction of high-level features with the low-level local features. Then, based on the granularity-aware text features, we decode them into text transcriptions in the sequential recurrence self-attention manner to capture the sequence-dependent relation in character-level semantics and locations. Extensive experiments show that our proposed method outperforms existing state-of-the-art weakly-supervised scene text spotters by a large margin. Xunquan Tong, Pengwen Dai, Xugong Qin, Rui Wang 0032, Wenqi Ren |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | MC-Blur: A Comprehensive Benchmark for Image DeblurringabstractBlur artifacts can seriously degrade the visual quality of images, and numerous deblurring methods have been proposed for specific scenarios. However, in most real-world images, blur is caused by different factors, e.g., motion, and defocus. In this paper, we address how other deblurring methods perform in the case of multiple types of blur. For in-depth performance evaluation, we construct a new large-scale multi-cause image deblurring dataset (MC-Blur), including real-world and synthesized blurry images with different blur factors. The images in the proposed MC-Blur dataset are collected using other techniques: averaging sharp images captured by a 1000-fps high-speed camera, convolving Ultra-High-Definition (UHD) sharp images with large-size kernels, adding defocus to images, and real-world blurry images captured by various camera models. Based on the MC-Blur dataset, we conduct extensive benchmarking studies to compare SOTA methods in different scenarios, analyze their efficiency, and investigate the buildataset’s capacity. These benchmarking results provide a comprehensive overview of the advantages and limitations of current deblurring methods, revealing our dataset’s advances. The dataset is available to the public athttps://github.com/HDCVLab/MC-Blur-Dataset. Kaihao Zhang, Tao Wang 0052, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Generative Adversarial and Self-Supervised Dehazing NetworkabstractOwing to the fast developments of economics, a lot of devices and objects have been connected and have formed the Internet of Things (IoT). Visual sensors have been applied in vehicle navigation, traffic situational awareness, and traffic safety management. However, the particles in the air degrade the imaging quality, which affects the performance of vehicle navigation, traffic situational awareness, and traffic safety management. Deep-learning-based dehazing methods were proposed to address this issue. However, these methods are trained with simulated hazy images and cannot generalize to natural haze images well. To address the domain shift problem, some methods resort to zero-shot learning or domain adaption to boost the generalization of the model on natural haze images. However, the relevance between dehazed results and clean images is ignored by zero-shot dehazing methods. Domain-adaption-based dehazing methods ignore the relationship between the dehazed results and the hazy images. To overcome these issues, a generative adversarial and self-supervised dehazing network is introduced to boost the dehazing performance on real haze images. First, generative adversarial is employed to construct the relevance between dehazed results and haze-free images, which can boost the natural appearance of dehazed results. Second, self-supervised learning is employed to construct the relevance between the dehazed results and hazy images, which can restrict the solution space of dehazing. To show the effectiveness of the proposed model, we conduct extensive experiments on real and simulated haze images. Compared with state-of-the-art methods, the proposed model achieves state-of-the-art dehazing performance. Shengdong Zhang, Xiaoqin Zhang 0002, Shaohua Wan 0001, Wenqi Ren, Liping Zhao 0005, LinLin Shen |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Cylin-Painting: Seamless 360° Panoramic Image Outpainting and BeyondabstractImage outpainting gains increasing attention since it can generate the complete scene from a partial view, providing a valuable solution to construct 360° panoramic images. As image outpainting suffers from the intrinsic issue of unidirectional completion flow, previous methods convert the original problem into inpainting, which allows a bidirectional flow. However, we find that inpainting has its own limitations and is inferior to outpainting in certain situations. The question of how they may be combined for the best of both has as yet remained under-explored. In this paper, we provide a deep analysis of the differences between inpainting and outpainting, which essentially depends on how the source pixels contribute to the unknown regions under different spatial arrangements. Motivated by this analysis, we present a Cylin-Painting framework that involves meaningful collaborations between inpainting and outpainting and efficiently fuses the different arrangements, with a view to leveraging their complementary benefits on a seamless cylinder. Nevertheless, straightforwardly applying the cylinder-style convolution often generates visually unpleasing results as it discards important positional information. To address this issue, we further present a learnable positional embedding strategy to incorporate the missing component of positional encoding into the cylinder convolution, which significantly improves the panoramic results. It is noted that while developed for image outpainting, the proposed algorithm can be effectively extended to other panoramic vision tasks, such as object detection, depth estimation, and image super-resolution. Code will be made available at https://github.com/KangLiao929/Cylin-Painting. Kang Liao, Xiangyu Xu 0002, Chunyu Lin, Wenqi Ren, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | INformer: Inertial-Based Fusion Transformer for Camera Shake DeblurringabstractInertial measurement units (IMU) in the capturing device can record the motion information of the device, with gyroscopes measuring angular velocity and accelerometers measuring acceleration. However, conventional deblurring methods seldom incorporate IMU data, and existing approaches that utilize IMU information often face challenges in fully leveraging this valuable data, resulting in noise issues from the sensors. To address these issues, in this paper, we propose a multi-stage deblurring network named INformer, which combines inertial information with the Transformer architecture. Specifically, we design an IMU-image Attention Fusion (IAF) block to merge motion information derived from inertial measurements with blurry image features at the attention level. Furthermore, we introduce an Inertial-Guided Deformable Attention (IGDA) block for utilizing the motion information features as guidance to adaptively adjust the receptive field, which can further refine the corresponding blur kernel for pixels. Extensive experiments on comprehensive benchmarks demonstrate that our proposed method performs favorably against state-of-the-art deblurring approaches. Wenqi Ren, Linrui Wu, Yanyang Yan, Shengyao Xu, Feng Huang 0007, Xiaochun Cao |
IEEE Trans. Image Process. | 1 |
| 2024 | Multi-Scale Fusion and Decomposition Network for Single Image DerainingabstractConvolutional neural networks (CNNs) and self-attention (SA) have demonstrated remarkable success in low-level vision tasks, such as image super-resolution, deraining, and dehazing. The former excels in acquiring local connections with translation equivariance, while the latter is better at capturing long-range dependencies. However, both CNNs and Transformers suffer from individual limitations, such as limited receptive field and weak diversity representation of CNNs during low efficiency and weak local relation learning of SA. To this end, we propose a multi-scale fusion and decomposition network (MFDNet) for rain perturbation removal, which unifies the merits of these two architectures while maintaining both effectiveness and efficiency. To achieve the decomposition and association of rain and rain-free features, we introduce an asymmetrical scheme designed as a dual-path mutual representation network that enables iterative refinement. Additionally, we incorporate high-efficiency convolutions throughout the network and use resolution rescaling to balance computational complexity with performance. Comprehensive evaluations show that the proposed approach outperforms most of the latest SOTA deraining methods and is versatile and robust in various image restoration tasks, including underwater image enhancement, image dehazing, and low-light image enhancement. The source codes and pretrained models are available at https://github.com/qwangg/MFDNet. Kui Jiang, Zheng Wang 0007, Wenqi Ren, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2024 | Adaptive Blind Super-Resolution Network for Spatial-Specific and Spatial-Agnostic DegradationsabstractPrior methodologies have disregarded the diversities among distinct degradation types during image reconstruction, employing a uniform network model to handle multiple deteriorations. Nevertheless, we discover that prevalent degradation modalities, including sampling, blurring, and noise, can be roughly categorized into two classes. We classify the first class as spatial-agnostic dominant degradations, less affected by regional changes in image space, such as downsampling and noise degradation. The second class degradation type is intimately associated with the spatial position of the image, such as blurring, and we identify them as spatial-specific dominant degradations. We introduce a dynamic filter network integrating global and local branches to address these two degradation types. This network can greatly alleviate the practical degradation problem. Specifically, the global dynamic filtering layer can perceive the spatial-agnostic dominant degradation in different images by applying weights generated by the attention mechanism to multiple parallel standard convolution kernels, enhancing the network's representation ability. Meanwhile, the local dynamic filtering layer converts feature maps of the image into a spatially specific dynamic filtering operator, which performs spatially specific convolution operations on the image features to handle spatial-specific dominant degradations. By effectively integrating both global and local dynamic filtering operators, our proposed method outperforms state-of-the-art blind super-resolution algorithms in both synthetic and real image datasets. Weilei Wen, Chunle Guo, Wenqi Ren, Hongpeng Wang 0001, Xiuli Shao |
IEEE Trans. Image Process. | 3 |
| 2024 | Real-Time Multi-Scene Visibility Enhancement for Promoting Navigational Safety of Vessels Under Complex Weather ConditionsabstractThe visible-light camera, which is capable of environment perception and navigation assistance, has emerged as an essential imaging sensor for marine surface vessels in intelligent waterborne transportation systems (IWTS). However, the visual imaging quality inevitably suffers from several kinds of degradations (e.g., limited visibility, low contrast, color distortion, etc.) under complex weather conditions (e.g., haze, rain, and low-lightness). The degraded visual information will accordingly result in inaccurate environment perception and delayed operations for navigational risk. To promote the navigational safety of vessels, many computational methods have been presented to perform visual quality enhancement under poor weather conditions. However, most of these methods are essentially specific-purpose implementation strategies, only available for one specific weather type. To overcome this limitation, we propose to develop a general-purpose multi-scene visibility enhancement method, i.e., edge reparameterization- and attention-guided neural network (ERANet), to adaptively restore the degraded images captured under different weather conditions. In particular, our ERANet simultaneously exploits the channel attention, spatial attention, and reparameterization technology to enhance the visual quality while maintaining low computational cost. Extensive experiments conducted on standard and IWTS-related datasets have demonstrated that our ERANet could outperform several representative visibility enhancement methods in terms of both imaging quality and computational efficiency. The superior performance of IWTS-related object detection and scene segmentation could also be steadily obtained after ERANet-based visibility enhancement under complex weather conditions. Ryan Wen Liu, Yuxu Lu, Yuan Gao 0015, Yu Guo 0008, Wenqi Ren, Fenghua Zhu, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Perception-Driven Deep Underwater Image Enhancement Without Paired SupervisionabstractUnderwater image enhancement (UIE) aims to improve the visual quality of raw underwater images. Current UIE algorithms primarily train a deep neural network (DNN) on synthetic datasets or datasets with pseudo labels by minimizing the reconstruction loss between enhanced images and ground truth images. However, there is a domain gap between synthetic and real-world underwater images, and the widely used$\ell _{1}$or$\ell _{2}$loss tends to overlook the importance of human perception, resulting in unsatisfactory perceptual quality of the final enhanced results. In this paper, we propose an unsupervised perception-driven DNN called PDD-Net for generalizable UIE. Instead of relying on paired images for training, we resort to an unsupervised generative adversarial network (GAN) with a large-scale set of easily available natural images as the target domain. This enables training on larger image sets collected from various domains while avoiding over-fitted to any specific data generation protocol. Additionally, to make the visual quality of enhanced underwater images more in line with human perception, we pre-train a DNN-based pairwise quality ranking (PQR) model based on which a PQR loss is formulated to progressively guides the enhancement of raw underwater image toward the higher quality direction. In addition, we introduce a global attention module (GAM) that integrates modulation and attention mechanisms to enable capturing rich global and local information, leading to improvements in both brightness and contrast. Extensive experiments demonstrate that our proposed PDD-Net exhibits excellent generalization capabilities and outperforms existing methods in terms of both visual perception quality and quantitative indicators across different datasets. Qiuping Jiang, Yaozu Kang, Zhihua Wang 0002, Wenqi Ren, Chongyi Li |
IEEE Trans. Multim. | 4 |
| 2023 | Dual-Domain Attention for Image DeblurringabstractAs a long-standing and challenging task, image deblurring aims to reconstruct the latent sharp image from its degraded counterpart. In this study, to bridge the gaps between degraded/sharp image pairs in the spatial and frequency domains simultaneously, we develop the dual-domain attention mechanism for image deblurring. Self-attention is widely used in vision tasks, however, due to the quadratic complexity, it is not applicable to image deblurring with high-resolution images. To alleviate this issue, we propose a novel spatial attention module by implementing self-attention in the style of dynamic group convolution for integrating information from the local region, enhancing the representation learning capability and reducing computational burden. Regarding frequency domain learning, many frequency-based deblurring approaches either treat the spectrum as a whole or decompose frequency components in a complicated manner. In this work, we devise a frequency attention module to compactly decouple the spectrum into distinct frequency parts and accentuate the informative part with extremely lightweight learnable parameters. Finally, we incorporate attention modules into a U-shaped network. Extensive comparisons with prior arts on the common benchmarks show that our model, named Dual-domain Attention Network (DDANet), obtains comparable results with a significantly improved inference speed. Yuning Cui 0001, Wenqi Ren, Alois C. Knoll |
AAAI | 3 |
| 2023 | High-Resolution Iterative Feedback Network for Camouflaged Object DetectionabstractSpotting camouflaged objects that are visually assimilated into the background is tricky for both object detection algorithms and humans who are usually confused or cheated by the perfectly intrinsic similarities between the foreground objects and the background surroundings. To tackle this challenge, we aim to extract the high-resolution texture details to avoid the detail degradation that causes blurred vision in edges and boundaries. We introduce a novel HitNet to refine the low-resolution representations by high-resolution features in an iterative feedback manner, essentially a global loop-based connection among the multi-scale resolutions. To design better feedback feature flow and avoid the feature corruption caused by recurrent path, an iterative feedback strategy is proposed to impose more constraints on each feedback connection. Extensive experiments on four challenging datasets demonstrate that our HitNet breaks the performance bottleneck and achieves significant improvements compared with 29 state-of-the-art methods. In addition, to address the data scarcity in camouflaged scenarios, we provide an application example to convert the salient objects to camouflaged objects, thereby generating more camouflaged training samples from the diverse salient object datasets. Code will be made publicly available. Xiaobin Hu, Xuebin Qin, Hang Dai, Wenqi Ren, Donghao Luo 0001, Ying Tai, Ling Shao 0001 |
AAAI | 5 |
| 2023 | Robust Single Image Reflection Removal Against Adversarial AttacksabstractThis paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustness study, we first conduct diverse adversarial attacks specifically for the SIRR problem, i.e. towards different attacking targets and regions. Then we propose a robust SIRR model, which integrates the cross-scale attention module, the multi-scale fusion module, and the adversarial image discriminator. By exploiting the multi-scale mechanism, the model narrows the gap between features from clean and adversarial images. The image discriminator adaptively distinguishes clean or noisy inputs, and thus further gains reliable robustness. Extensive experiments on Nature, SIR2, and Real datasets demonstrate that our model remarkably improves the robustness of SIRR across disparate scenes. Zhenbo Song, Zhenyuan Zhang 0001, Kaihao Zhang, Wenhan Luo, Zhaoxin Fan, Wenqi Ren, Jianfeng Lu 0003 |
CVPR | 6 |
| 2023 | MProto: Multi-Prototype Network with Denoised Optimal Transport for Distantly Supervised Named Entity RecognitionabstractDistantly supervised named entity recognition (DS-NER) aims to locate entity mentions and classify their types with only knowledge bases or gazetteers and unlabeled corpus.However, distant annotations are noisy and degrade the performance of NER models.In this paper, we propose a noise-robust prototype network named MProto for the DS-NER task.Different from previous prototype-based NER methods, MProto represents each entity type with multiple prototypes to characterize the intra-class variance among entity representations.To optimize the classifier, each token should be assigned an appropriate ground-truth prototype and we consider such token-prototype assignment as an optimal transport (OT) problem.Furthermore, to mitigate the noise from incomplete labeling, we propose a novel denoised optimal transport (DOT) algorithm.Specifically, we utilize the assignment result between Other class tokens and all prototypes to distinguish unlabeled entity tokens from true negatives.Experiments on several DS-NER benchmarks demonstrate that our MProto achieves state-of-the-art performance.The source code is now available on Github 1 . Shuhui Wu, Yongliang Shen 0001, Zeqi Tan, Wenqi Ren, Jietian Guo, Shiliang Pu, Weiming Lu 0001 |
EMNLP | 4 |
| 2023 | Focal Network for Image RestorationabstractImage restoration aims to reconstruct a sharp image from its degraded counterpart, which plays an important role in many fields. Recently, Transformer models have achieved promising performance on various image restoration tasks. However, their quadratic complexity remains an intractable issue for practical applications. The aim of this study is to develop an efficient and effective framework for image restoration. Inspired by the fact that different regions in a corrupted image always undergo degradations in various degrees, we propose to focus more on the important areas for reconstruction. To this end, we introduce a dual-domain selection mechanism to emphasize crucial information for restoration, such as edge signals and hard regions. In addition, we split high-resolution features to insert multi-scale receptive fields into the network, which improves both efficiency and performance. Finally, the proposed network, dubbed FocalNet, is built by incorporating these designs into a U-shaped backbone. Extensive experiments demonstrate that our model achieves state-of-the-art performance on ten datasets for three tasks, including single-image defocus deblurring, image dehazing, and image desnowing. Our code is available at https://github.com/c-yn/FocalNet. Yuning Cui 0001, Wenqi Ren, Xiaochun Cao, Alois C. Knoll |
ICCV | 2 |
| 2023 | Lightweight Image Super-Resolution with Superpixel Token InteractionabstractTransformer-based methods have demonstrated impressive results on single-image super-resolution (SISR) task. However, self-attention mechanism is computationally expensive when applied to the entire image. As a result, current approaches divide low-resolution input images into small patches, which are processed separately and then fused to generate high-resolution images. Nevertheless, this conventional regular patch division is too coarse and lacks interpretability, resulting in artifacts and non-similar structure interference during attention operations. To address these challenges, we propose a novel super token interaction network (SPIN). Our method employs superpixels to cluster local similar pixels to form the explicable local regions and utilizes intra-superpixel attention to enable local information interaction. It is interpretable because only similar regions complement each other and dissimilar regions are excluded. Moreover, we design a super-pixel cross-attention module to facilitate information propagation via the surrogation of superpixels. Extensive experiments demonstrate that the proposed SPIN model performs favorably against the state-of-the-art SR methods in terms of accuracy and lightweight. Code is available at https://github.com/ArcticHare105/SPIN. Aiping Zhang, Wenqi Ren, Yi Liu 0085, Xiaochun Cao |
ICCV | 2 |
| 2023 | Selective Frequency Network for Image Restoration
Yuning Cui 0001, Zhenshan Bing, Wenqi Ren, Xinwei Gao, Xiaochun Cao, Kai Huang 0001, Alois C. Knoll |
ICLR | 4 |
| 2023 | IRNeXt: Rethinking Convolutional Network Design for Image RestorationabstractWe present IRNeXt, a simple yet effective convolutional network architecture for image restoration. Recently, Transformer models have dominated the field of image restoration due to the powerful ability of modeling long-range pixels interactions. In this paper, we excavate the potential of the convolutional neural network (CNN) and show that our CNN-based model can receive comparable or better performance than Transformer models with low computation overhead on several image restoration tasks. By re-examining the characteristics possessed by advanced image restoration algorithms, we discover several key factors leading to the performance improvement of restoration models. This motivates us to develop a novel network for image restoration based on cheap convolution operators. Comprehensive experiments demonstrate that IRNeXt delivers state-of-the-art performance among numerous datasets on a range of image restoration tasks with low computational complexity, including image dehazing, single-image defocus/motion deblurring, image deraining, and image desnowing. https://github.com/c-yn/IRNeXt. Yuning Cui 0001, Wenqi Ren, Sining Yang, Xiaochun Cao, Alois C. Knoll |
ICML | 2 |
| 2023 | Rethinking Unsupervised Domain Adaptation for Nighttime Tracking
Qiyu Sun, Chaoqiang Zhao, Wenqi Ren, Yang Tang 0001 |
ICONIP (14) | 4 |
| 2023 | Dual-Domain Learning Network for Polyp Segmentation
Yan Li 0196, Zhuoran Zheng, Wenqi Ren, Yunfeng Nie, Jingang Zhang, Xiuyi Jia |
IWDW | 3 |
| 2023 | SIGMA-DF: Single-Side Guided Meta-Learning for Deepfake DetectionabstractThe current challenge of Deepfake detection is the cross-domain performance on unseen Deepfake data. Instead of extracting forgery artifacts that are robust to the cross-domain scenarios as most previous works, we propose a novel method named Single-sIde Guided Meta-leArning framework for DeepFake detection (SIGMA-DF) which simulates the cross-domain scenarios during training by synthesizing virtual testing domain through meta-learning. In addition, SIGMA-DF integrates the meta-learning algorithm with a new ensemble meta-learning framework, which separately trains multiple meta-learners in the meta-train phase to aggregate multiple domain shifts in each iteration. Hence multiple cross-domain scenarios are simulated, better leveraging the domain knowledge. In addition, considering the contribution of hard samples in single-side distribution optimization, a novel weighted single-side loss function is proposed to only narrow the intra-class distance between real faces and enlarge the inter-class distance for both real and fake faces in embedding space with the awareness of sample weights. Extensive experiments are conducted on several standard Deepfake detection datasets to demonstrate that the proposed SIGMA-DF achieves state-of-the-art performance. In particular, in the cross-domain evaluation from FF++ to Celeb-DF and DFDC, our SIGMA-DF outperforms the baselines by 4.4% and 4.5% in terms of AUC, respectively. Jianshu Li, Wenqi Ren, Jian Liu 0012, Xiaochun Cao |
ICMR | 3 |
| 2023 | NightHazeFormer: Single Nighttime Haze Removal Using Prior Query TransformerabstractNighttime image dehazing is a challenging task due to the presence of multiple types of adverse degrading effects including glow, haze, blur, noise, color distortion, and so on. However, most previous studies mainly focus on daytime image dehazing or partial degradations presented in nighttime hazy scenes, which may lead to unsatisfactory restoration results. In this paper, we propose an end-to-end transformer-based framework for nighttime haze removal, called NightHazeFormer. Our proposed approach consists of two stages: supervised pre-training and semi-supervised fine-tuning. During the pre-training stage, we introduce two powerful priors into the transformer decoder to generate the non-learnable prior queries, which guide the model to extract specific degradations. For the fine-tuning, we combine the generated pseudo ground truths with input real-world nighttime hazy images as paired images and feed into the synthetic domain to fine-tune the pre-trained model. This semi-supervised fine-tuning paradigm helps improve the generalization to real domain. In addition, we also propose a large-scale synthetic dataset called UNREAL-NH, to simulate the real-world nighttime haze scenarios comprehensively. Extensive experiments on several synthetic and real-world datasets demonstrate the superiority of our NightHazeFormer over state-of-the-art nighttime haze removal methods in terms of both visually and quantitatively. Yun Liu 0002, Zhongsheng Yan, Sixiang Chen, Tian Ye 0001, Wenqi Ren, Erkang Chen |
ACM Multimedia | 5 |
| 2023 | Memory-Augmented Deep Unfolding Network for Guided Image Super-resolution
Man Zhou 0003, Jinshan Pan, Wenqi Ren, Qi Xie 0002, Xiangyong Cao |
Int. J. Comput. Vis. | 4 |
| 2023 | Physical-priors-guided DehazeFormer
Hao Zhou 0038, Yun Liu 0002, Yongpan Sheng, Wenqi Ren, Hailing Xiong |
Knowl. Based Syst. | 5 |
| 2023 | Enhanced Spatio-Temporal Interaction Learning for Video Deraining: Faster and BetterabstractVideo deraining is an important task in computer vision as the unwanted rain hampers the visibility of videos and deteriorates the robustness of most outdoor vision systems. Despite the significant success which has been achieved for video deraining recently, two major challenges remain: 1) how to exploit the vast information among successive frames to extract powerful spatio-temporal features across both the spatial and temporal domains, and 2) how to restore high-quality derained videos with a high-speed approach. In this paper, we present a new end-to-end video deraining framework, dubbed Enhanced Spatio-Temporal Interaction Network (ESTINet), which considerably boosts current state-of-the-art video deraining quality and speed. The ESTINet takes the advantage of deep residual networks and convolutional long short-term memory, which can capture the spatial features and temporal correlations among successive frames at the cost of very little computational resource. Extensive experiments on three public datasets show that the proposed ESTINet can achieve faster speed than the competitors, while maintaining superior performance over the state-of-the-art methods. https://github.com/HDCVLab/Enhanced-Spatio-Temporal-Interaction-Learning-for-Video-Deraining. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren, Wei Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A Perception-Aware Decomposition and Fusion Framework for Underwater Image EnhancementabstractThis paper presents a perception-aware decomposition and fusion framework for underwater image enhancement (UIE). Specifically, a general structural patch decomposition and fusion (SPDF) approach is introduced. SPDF is built upon the fusion of two complementary pre-processed inputs in a perception-aware and conceptually independent image space. First, a raw underwater image is pre-processed to produce two complementary versions including a contrast-corrected image and a detail-sharpened image. Then, each of them is decomposed into three conceptually independent components, i.e., mean intensity, contrast, and structure, via structural patch decomposition (SPD). Afterwards, the corresponding components are fused using tailored strategies. The three components after fusion are finally integrated via inverting the decomposition to reconstruct a final enhanced underwater image. The main advantage of SPDF is that two complementary pre-processed images are fused in a perception-aware and conceptually independent image space and the fusions of different components can be performed separately without any interactions and information loss. Comprehensive comparisons on two benchmark datasets demonstrate that SPDF outperforms several state-of-the-art UIE algorithms qualitatively and quantitatively. Moreover, the effectiveness of SPDF is also verified on another two relevant tasks, i.e., low-light image enhancement and single image dehazing. The code will be made available soon. Yaozu Kang, Qiuping Jiang, Chongyi Li, Wenqi Ren, Hantao Liu, Pengjun Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Semantic-Aware Dehazing Network With Adaptive Feature FusionabstractDespite that convolutional neural networks (CNNs) have shown high-quality reconstruction for single image dehazing, recovering natural and realistic dehazed results remains a challenging problem due to semantic confusion in the hazy scene. In this article, we show that it is possible to recover textures faithfully by incorporating semantic prior into dehazing network since objects in haze-free images tend to show certain shapes, textures, and colors. We propose a semantic-aware dehazing network (SDNet) in which the semantic prior is taken as a color constraint for dehazing, benefiting the acquisition of a reasonable scene configuration. In addition, we design a densely connected block to capture global and local information for dehazing and semantic prior estimation. To eliminate the unnatural appearance of some objects, we propose to fuse the features from shallow and deep layers adaptively. Experimental results demonstrate that our proposed model performs favorably against the state-of-the-art single image dehazing approaches. Shengdong Zhang, Wenqi Ren, Xin Tan 0002, Zhi-Jie Wang 0009, Yong Liu 0018, Jingang Zhang, Xiaoqin Zhang 0002, Xiaochun Cao |
IEEE Trans. Cybern. | 2 |
| 2023 | PFONet: A Progressive Feedback Optimization Network for Lightweight Single Image DehazingabstractImage dehazing is an effective means to enhance the quality of images captured in foggy or hazy weather conditions. However, existing image dehazing methods are either ineffective in dealing with complex haze scenes, or incurring too much computation. To overcome these deficiencies, we propose a progressive feedback optimization network (PFONet) which is lightweight yet effective for image dehazing. The PFONet consists of a multi-stream dehazing module and a progressive feedback module. The progressive feedback module feeds the output dehazed image back to the intermedia features extracted by the network, thus enabling the network to gradually reconstruct a complex degraded image. Considering both the effectiveness and efficiency of the network, we also design a lightweight hybrid residual dense block serving as the basic feature extraction module of the proposed PFONet. Extensive experimental results are presented to demonstrate that the proposed model outperforms its state-of-the-art single-image dehazing competitors for both synthetic and real-world images. Shuoshi Li, Yuan Zhou 0006, Wenqi Ren, Wei Xiang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Multi-Exposure Image Fusion via Deformable Self-AttentionabstractMost multi-exposure image fusion (MEF) methods perform unidirectional alignment within limited and local regions, which ignore the effects of augmented locations and preserve deficient global features. In this work, we propose a multi-scale bidirectional alignment network via deformable self-attention to perform adaptive image fusion. The proposed network exploits differently exposed images and aligns them to the normal exposure in varying degrees. Specifically, we design a novel deformable self-attention module that considers variant long-distance attention and interaction and implements the bidirectional alignment for image fusion. To realize adaptive feature alignment, we employ a learnable weighted summation of different inputs and predict the offsets in the deformable self-attention module, which facilitates that the model generalizes well in various scenes. In addition, the multi-scale feature extraction strategy makes the features across different scales complementary and provides fine details and contextual features. Extensive experiments demonstrate that our proposed algorithm performs favorably against state-of-the-art MEF methods. Jun Luo 0012, Wenqi Ren, Xinwei Gao, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2023 | Event-Aware Video Deraining via Multi-Patch Progressive LearningabstractIn this paper, we address the problem of video-based rain streak removal by developing an event-aware multi-patch progressive neural network. Rain streaks in video exhibit correlations in both temporal and spatial dimensions. Existing methods have difficulties in modeling the characteristics. Based on the observation, we propose to develop a module encoding events from neuromorphic cameras to facilitate deraining. Events are captured asynchronously at pixel-level only when intensity changes by a margin exceeding a certain threshold. Due to this property, events contain considerable information about moving objects including rain streaks passing though the camera across adjacent frames. Thus we suggest that utilizing it properly facilitates deraining performance non-trivially. In addition, we develop a multi-patch progressive neural network. The multi-patch manner enables various receptive fields by partitioning patches and the progressive learning in different patch levels makes the model emphasize each patch level to a different extent. Extensive experiments show that our method guided by events outperforms the state-of-the-art methods by a large margin in synthetic and real-world datasets. Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Kaihao Zhang, Meiyu Liang, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2022 | Image Dehazing Transformer with Transmission-Aware 3D Position EmbeddingabstractDespite single image dehazing has been made promising progress with Convolutional Neural Networks (CNNs), the inherent equivariance and locality of convolution still bottleneck deharing performance. Though Transformer has occupied various computer vision tasks, directly leveraging Transformer for image dehazing is challenging: 1) it tends to result in ambiguous and coarse details that are undesired for image reconstruction; 2) previous position embedding of Transformer is provided in logic or spatial position order that neglects the variational haze densities, which results in the sub-optimal dehazlng performance. The key insight of this study is to investigate how to combine CNN and Transformer for image dehazing. To solve the feature inconsistency issue between Transformer and CNN, we propose to modulate CNN features via learning modulation matrices (i.e., coefficient matrix and bias matrix) conditioned on Transformer features instead of simple feature addition or concatenation. The feature modulation naturally inherits the global context modeling capability of Transformer and the local representation capability of CNN. We bring a haze density-related prior into Trans-former via a novel transmission-aware 3D position embedding module, which not only provides the relative position but also suggests the haze density of different spatial regions. Extensive experiments demonstrate that our method, DeHamer, attains state-of-the-art performance on several image dehazing benchmarks. Chunle Guo, Qixin Yan, Saeed Anwar, Runmin Cong, Wenqi Ren, Chongyi Li |
CVPR | 5 |
| 2022 | Self-supervised Learning and Adaptation for Single Image DehazingabstractExisting deep image dehazing methods usually depend on supervised learning with a large number of hazy-clean image pairs which are expensive or difficult to collect. Moreover, dehazing performance of the learned model may deteriorate significantly when the training hazy-clean image pairs are insufficient and are different from real hazy images in applications. In this paper, we show that exploiting large scale training set and adapting to real hazy images are two critical issues in learning effective deep dehazing models. Under the depth guidance estimated by a well-trained depth estimation network, we leverage the conventional atmospheric scattering model to generate massive hazy-clean image pairs for the self-supervised pre-training of dehazing network. Furthermore, self-supervised adaptation is presented to adapt pre-trained network to real hazy images. Learning without forgetting strategy is also deployed in self-supervised adaptation by combining self-supervision and model adaptation via contrastive learning. Experiments show that our proposed method performs favorably against the state-of-the-art methods, and is quite efficient, i.e., handling a 4K image in 23 ms. The codes are available at https://github.com/DongLiangSXU/SLAdehazing. Yudong Liang, Bin Wang 0071, Wangmeng Zuo, Jiaying Liu 0001, Wenqi Ren |
IJCAI | 5 |
| 2022 | Learning Hierarchical Dynamics with Spatial Adjacency for Image EnhancementabstractIn various real-world image enhancement applications, the degradations are always non-uniform or non-homogeneous and diverse, which challenges most deep networks with fixed parameters during the inference phase. Inspired by the dynamic deep networks that adapt the model structures or parameters conditioned on the inputs, we propose a DCP-guided hierarchical dynamic mechanism for image enhancement to adapt the model parameters and features from local to global as well as to keep spatial adjacency within the region. Specifically, channel-spatial-level, structure-level, and region-level dynamic components are sequentially applied. Channel-spatial-level dynamics obtain channel- and spatial-wise representation variations, and structure-level dynamics enable modeling geometric transformations and augment sampling locations for the varying local features to better describe the structures. In addition, a novel region-level dynamic is proposed to generate spatially continuous masks for dynamic features which capitalizes on the Dark Channel Priors (DCP). The proposed region-level dynamics benefit from exploiting the statistical differences between distorted and undistorted images. Moreover, the DCP-guided region generations are inherently spatial coherent which facilitates capturing local coherence of the images. The proposed method achieves state-of-the-art performance and generates visually pleasing images for multiple enhancement tasks,i.e. , image dehazing, image deraining and low-light image enhancement. The codes are available at https://github.com/DongLiangSXU/HDM. Yudong Liang, Bin Wang 0071, Wenqi Ren, Jiaying Liu 0001, Wangmeng Zuo |
ACM Multimedia | 3 |
| 2022 | Rethinking Image Restoration for Object DetectionabstractAlthough image restoration has achieved significant progress, its potential to assist object detectors in adverse imaging conditions lacks enough attention. It is reported that the existing image restoration methods cannot improve the object detector performance and sometimes even reduce the detection performance. To address the issue, we propose a targeted adversarial attack in the restoration procedure to boost object detection performance after restoration. Specifically, we present an ADAM-like adversarial attack to generate pseudo ground truth for restoration training. Resultant restored images are close to original sharp images, and at the same time, lead to better results of object detection. We conduct extensive experiments in image dehazing and low light enhancement and show the superiority of our method over conventional training and other domain adaptation and multi-task methods. The proposed pipeline can be applied to all restoration methods and detectors in both one- and two-stage. Shangquan Sun, Wenqi Ren, Tao Wang 0053, Xiaochun Cao |
NeurIPS | 2 |
| 2022 | Integrating deep learning and traditional image enhancement techniques for underwater image enhancementabstractAbstract Underwater images usually suffer from colour distortion, blur, and low contrast, which hinder the subsequent processing of underwater information. To address these problems, this paper proposes a novel approach for single underwater images enhancement by integrating data‐driven deep learning and hand‐crafted image enhancement techniques. First, a statistical analysis is made on the average deviation of each channel of input underwater images to that of its corresponding ground truths, and it is found that both the red channel and the green channel of an underwater image contribute to its colour distortion. Concretely, the red channel of an underwater image is usually seriously attenuated, and the green channel is usually over strengthened. Motivated by such an observation, an attention mechanism guided residual module for underwater image colour correction is proposed, where the colour of the red channel of the underwater image and that of the green channel is compensated in a different way, respectively. Coupled with an attention mechanism, the residual module can adaptively extract and integrate the most discriminative features for colour correction. For scene contrast enhancement and scene deblurring, the traditional image enhancement techniques such as CLAHE (contrast limited adaptive histogram equalization) and Gamma correction are coupled with a multi‐scale convolutional neural network (MSCNN), where CLAHE and Gamma correction are used as complement to deal with the complex and changeable underwater imaging environment. Experiments on synthetic and real underwater images demonstrate that the proposed method performs favourably against the state‐of‐the‐art underwater image enhancement methods. Zhenghao Shi, Yongli Wang 0003, Zhaorun Zhou, Wenqi Ren |
IET Image Process. | 4 |
| 2022 | Beyond Monocular Deraining: Parallel Stereo Deraining Network Via Semantic Prior
Kaihao Zhang, Wenhan Luo, Yanjiang Yu, Wenqi Ren, Fang Zhao 0006, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 4 |
| 2022 | Deep Image Deblurring: A Survey
Kaihao Zhang, Wenqi Ren, Wenhan Luo, Wei-Sheng Lai, Björn Stenger, Ming-Hsuan Yang 0001, Hongdong Li |
Int. J. Comput. Vis. | 2 |
| 2022 | Auto Color Correction of Underwater Images Utilizing Depth InformationabstractThe red spectrum is saliently attenuated due to the absorption and scattering properties of water. The acquired underwater images show severe color cast in underwater scenes. In this letter, we propose a novel color correction method for underwater images, which removes color cast on single pixels based on scene depth. The experimental results demonstrate that our approach can significantly improve the color effect and provide a correct input for the subsequent underwater image defogging methods. Jingchun Zhou, Dehuan Zhang, Wenqi Ren, Weishi Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Face Restoration via Plug-and-Play 3D Facial PriorsabstractState-of-the-art face restoration methods employ deep convolutional neural networks (CNNs) to learn a mapping between degraded and sharp facial patterns by exploring local appearance knowledge. However, most of these methods do not well exploit facial structures and identity information, and only deal with task-specific face restoration (e.g., face super-resolution or deblurring). In this paper, we propose cross-tasks and cross-models plug-and-play 3D facial priors to explicitly embed the network with the sharp facial structures for general face restoration tasks. Our 3D priors are the first to explore 3D morphable knowledge based on the fusion of parametric descriptions of face attributes (e.g., identity, facial expression, texture, illumination, and face pose). Furthermore, the priors can easily be incorporated into any network and are very efficient in improving the performance and accelerating the convergence speed. Firstly, a 3D face rendering branch is set up to obtain 3D priors of salient facial structures and identity knowledge. Secondly, for better exploiting this hierarchical information (i.e., intensity similarity, 3D facial structure, and identity content), a spatial attention module is designed for the image restoration problems. Extensive face restoration experiments including face super-resolution and deblurring demonstrate that the proposed 3D priors achieve superior face restoration results over the state-of-the-art algorithms. Xiaobin Hu, Wenqi Ren, Jiaolong Yang, Xiaochun Cao, David P. Wipf, Bjoern Menze, Xin Tong 0001, Hongbin Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Deblurring Dynamic Scenes via Spatially Varying Recurrent Neural NetworksabstractDeblurring images captured in dynamic scenes is challenging as the motion blurs are spatially varying caused by camera shakes and object movements. In this paper, we propose a spatially varying neural network to deblur dynamic scenes. The proposed model is composed of three deep convolutional neural networks (CNNs) and a recurrent neural network (RNN). The RNN is used as a deconvolution operator on feature maps extracted from the input image by one of the CNNs. Another CNN is used to learn the spatially varying weights for the RNN. As a result, the RNN is spatial-aware and can implicitly model the deblurring process with spatially varying kernels. To better exploit properties of the spatially varying RNN, we develop both one-dimensional and two-dimensional RNNs for deblurring. The third component, based on a CNN, reconstructs the final deblurred feature maps into a restored image. In addition, the whole network is end-to-end trainable. Quantitative and qualitative evaluations on benchmark datasets demonstrate that the proposed method performs favorably against the state-of-the-art deblurring algorithms. Wenqi Ren, Jiawei Zhang 0002, Jinshan Pan, Sifei Liu, Jimmy S. J. Ren, Junping Du 0001, Xiaochun Cao, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | MTRBNet: Multi-Branch Topology Residual Block-Based Network for Low-Light EnhancementabstractThe learning-based low-light image enhancement methods have remarkable performance due to the robust feature learning and mapping capabilities. This paper proposes a multi-branch topology residual block (MTRB)-based network (MTRBNet), which can alleviate training difficulties and more efficiently use the parameters between neurons. Compared with the previous residual block, the proposed MTRB increases the width of the network and simultaneously transmits information along with the depth and width directions, which can effectively select network nodes to promote the network learning capacity. Meanwhile, the feature information of neighbor nodes is transferred to each other, thereby maximizing the information flow of the convolution unit. The proposed information connection and feedback mechanism can improve the network’s ability to capture the global and local features. We analyze the pros and cons of two multi-feature fusion strategies (i.e., addition and concatenation) and three normalization methods on the quantitative results. In addition, we embed our MTRB into traditional Encoder-Decoder structure to improve the image enhancement results under different low-light imaging conditions. Experiments on the LOL image dataset have demonstrated that our MTRBNet achieves superior performance compared with several state-of-the-art methods. Yuxu Lu, Yu Guo 0008, Ryan Wen Liu, Wenqi Ren |
IEEE Signal Process. Lett. | 4 |
| 2022 | IDBP: Image Dehazing Using Blended Priors Including Non-Local, Local, and Global PriorsabstractIn this letter, a robust and promising atmospheric scattering model (ASM)-based image dehazing technique called IDBP is developed, which overcomes the intrinsic limitation of available techniques based on single priors. It consists of two modules, i.e., an atmospheric light estimation (ALE) module and a multiple prior constraint (MPC) module. The ALE module is based on a new global brightening strategy of enhancing the brightness of image with minimum information loss. The MPC smartly blends the constrains of non-local prior, local prior, and global prior to shrink the solution space of haze removal, which avoids the limitation of using any single priors. Unlike previous works, IDBP does not require any training process, but is based on multiple priors and minimal information loss principle to impose the ASM, thereby making it easy to implement and ensuring its robustness. Numerous experiments reveal that the proposed IDBP outperforms the state-of-the-art alternates. Mingye Ju, Can Ding 0002, Wenqi Ren, Yi Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Hierarchical Density-Aware Dehazing NetworkabstractThe commonly used atmospheric model in image dehazing cannot hold in real cases. Although deep end-to-end networks were presented to solve this problem by disregarding the physical model, the transmission map in the atmospheric model contains significant haze density information, which cannot simply be ignored. In this article, we propose a novel hierarchical density-aware dehazing network, which consists of a the densely connected pyramid encoder, a density generator, and a Laplacian pyramid decoder. The proposed network incorporates density estimation but alleviates the constraint of the atmospheric model. The predicted haze density then guides the Laplacian pyramid decoder to generate a haze-free image in a coarse-to-fine fashion. In addition, we introduce a multiscale discriminator to preserve global and local consistency for dehazing. We conduct extensive experiments on natural and synthetic hazy images, which prove that the proposed model performs favorably against the state-of-the-art dehazing approaches. Jingang Zhang, Wenqi Ren, Shengdong Zhang, He Zhang 0004, Yunfeng Nie, Zhe Xue, Xiaochun Cao |
IEEE Trans. Cybern. | 2 |
| 2022 | Under-Display Camera Image Enhancement via Cascaded Curve EstimationabstractThe new trend of full-screen devices encourages manufacturers to position a camera behind a screen, i.e., the newly-defined Under-Display Camera (UDC). Therefore, UDC image restoration has been a new realistic single image enhancement problem. In this work, we propose a curve estimation network operating on the hue (H) and saturation (S) channels to perform adaptive enhancement for degraded images captured by UDCs. The proposed network aims to match the complicated relationship between the images captured by under-display and display-free cameras. To extract effective features, we cascade the proposed curve estimation network with sharing weights, and we introduce a spatial and channel attention module in each curve estimation network to exploit attention-aware features. In addition, we learn the curve estimation network in a semi-supervised manner to alleviate the restriction of the requirement for amounts of labeled images and improve the generalization ability for unseen degraded images in various realistic scenes. The semi-supervised network consists of a supervised branch trained on labeled data and an unsupervised branch trained on unlabeled data. To train the proposed model, we build a new dataset comprised of real-world labeled and unlabeled images. Extensive experiments demonstrate that our proposed algorithm performs favorably against state-of-the-art image enhancement methods for UDC images in terms of accuracy and speed, especially on ultra-high-definition (UHD) images. Jun Luo 0012, Wenqi Ren, Tao Wang 0074, Chongyi Li, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2022 | Attention-Guided Collaborative CountingabstractExisting crowd counting designs usually exploit multi-branch structures to address the scale diversity problem. However, branches in these structures work in a competitive rather than collaborative way. In this paper, we focus on promoting collaboration between branches. Specifically, we propose an attention-guided collaborative counting module (AGCCM) comprising an attention-guided module (AGM) and a collaborative counting module (CCM). The CCM promotes collaboration among branches by recombining each branch's output into an independent count and joint counts with other branches. The AGM capturing the global attention map through a transformer structure with a pair of foreground-background related loss functions can distinguish the advantages of different branches. The loss functions do not require additional labels and crowd division. In addition, we design two kinds of bidirectional transformers (Bi-Transformers) to decouple the global attention to row attention and column attention. The proposed Bi-Transformers are able to reduce the computational complexity and handle images in any resolution without cropping the image into small patches. Extensive experiments on several public datasets demonstrate that the proposed algorithm performs favorably against the state-of-the-art crowd counting methods. Hong Mo, Wenqi Ren, Feihu Yan, Zhong Zhou, Xiaochun Cao, Wei Wu 0008 |
IEEE Trans. Image Process. | 2 |
| 2022 | Video Super-Resolution via a Spatio-Temporal Alignment NetworkabstractDeep convolutional neural network based video super-resolution (SR) models have achieved significant progress in recent years. Existing deep video SR methods usually impose optical flow to wrap the neighboring frames for temporal alignment. However, accurate estimation of optical flow is quite difficult, which tends to produce artifacts in the super-resolved results. To address this problem, we propose a novel end-to-end deep convolutional network that dynamically generates the spatially adaptive filters for the alignment, which are constituted by the local spatio-temporal channels of each pixel. Our method avoids generating explicit motion compensation and utilizes spatio-temporal adaptive filters to achieve the operation of alignment, which effectively fuses the multi-frame information and improves the temporal consistency of the video. Capitalizing on the proposed adaptive filter, we develop a reconstruction network and take the aligned frames as input to restore the high-resolution frames. In addition, we employ residual modules embedded with channel attention as the basic unit to extract more informative features for video SR. Both quantitative and qualitative evaluation results on three public video datasets demonstrate that the proposed method performs favorably against state-of-the-art super-resolution methods in terms of clearness and texture details. Weilei Wen, Wenqi Ren, Yinghuan Shi, Yunfeng Nie, Jingang Zhang, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2022 | Unsupervised Decomposition and Correction Network for Low-Light Image EnhancementabstractVision-based intelligent driving assistance systems and transportation systems can be improved by enhancing the visibility of the scenes captured in extremely challenging conditions. In particular, many low-image image enhancement (LIE) algorithms have been proposed to facilitate such applications in low-light conditions. While deep learning-based methods have achieved substantial success in this field, most of them require paired training data, which is difficult to be collected. This paper advocates a novel Unsupervised Decomposition and Correction Network (UDCN) for LIE without depending on paired data for training. Inspired by the Retinex model, our method first decomposes images into illumination and reflectance components with an image decomposition network (IDN). Then, the decomposed illumination is processed by an illumination correction network (ICN) and fused with the reflectance to generate a primary enhanced result. In contrast with fully supervised learning approaches, UDCN is an unsupervised one which is trained only with low-light images and corresponding histogram equalized (HE) counterparts (can be derived from the low-light image itself) as input. Both the decomposition and correction networks are optimized under the guidance of hybrid no-reference quality-aware losses and inter-consistency constraints between the low-light image and its HE counterpart. In addition, we also utilize an unsupervised noise removal network (NRN) to remove the noise previously hidden in the darkness for further improving the primary result. Qualitative and quantitative comparison results are reported to demonstrate the efficacy of UDCN and its superiority over several representative alternatives in the literature. The results and code will be made public available athttps://github.com/myd945/UDCN. Qiuping Jiang, Yudong Mao, Runmin Cong, Wenqi Ren, Chao Huang 0008, Feng Shao 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Disentangled Feature Networks for Facial Portrait and Caricature GenerationabstractFacial portrait is an artistic form which draws faces by emphasizing discriminative or prominent parts of faces via various kinds of drawing tools. However, the complex interplay between the different facial factors, such as facial parts, background, and drawing styles, and the significant domain gap between natural facial images and their portrait counterparts makes the task challenging. In this paper, a flexible four-stream Disentangled Feature Networks (DFN) is proposed to learn disentangled feature representation of different facial factors and generate plausible portraits with reasonable exaggerations and richness in style. Four factors are encoded as embedding features, and combined to reconstruct facial portraits. Meanwhile, to make the process fully automatic (without manually specifying either portrait style or exaggerating form), we propose a new Adversarial Portrait Mapping Module (APMM) to map noise to the embedding feature space, as proxies for portrait style and exaggerating. Thanks to the proposedDFNandAPMM, we are able to manipulate the portrait style and facial geometric structures to generate a large number of portraits. Extensive experiments on two public datasets show that our proposed methods can generate a diverse set of artistic portraits. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wenqi Ren, Hongdong Li |
IEEE Trans. Multim. | 4 |
| 2021 | ARVo: Learning All-Range Volumetric Correspondence for Video DeblurringabstractVideo deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly on homography or optical flows to spatially align neighboring blurry frames. However, such explicit approaches are less effective in the presence of fast motions with large pixel displacements. In this work, we propose a novel implicit method to learn spatial correspondence among blurry frames in the feature space. To construct distant pixel correspondences, our model builds a correlation volume pyramid among all the pixel-pairs between neigh-boring frames. To enhance the features of the reference frame, we design a correlative aggregation module that maximizes the pixel-pair correlations with its neighbors based on the volume pyramid. Finally, we feed the aggregated features into a reconstruction module to obtain the restored frame. We design a generative adversarial paradigm to optimize the model progressively. Our proposed method is evaluated on the widely-adopted DVD dataset, along with a newly collected High-Frame-Rate (1000 fps) Dataset for Video Deblurring (HFR-DVD). Quantitative and qualitative experiments show that our model performs favorably on both datasets against previous state-of-the-art methods, confirming the benefit of modeling all-range spatial correspondence for video deblurring. Dongxu Li 0003, Kaihao Zhang, Xin Yu 0002, Yiran Zhong, Wenqi Ren, Hanna Suominen, Hongdong Li |
CVPR | 6 |
| 2021 | DCNAS: Densely Connected Neural Architecture Search for Semantic Image SegmentationabstractExisting NAS methods for dense image prediction tasks usually compromise on restricted search space or search on proxy task to meet the achievable computational demands. To allow as wide as possible network architectures and avoid the gap between realistic and proxy setting, we propose a novel Densely Connected NAS (DCNAS) framework, which directly searches the optimal network structures for the multi-scale representations of visual information, over a large-scale target dataset without proxy. Specifically, by connecting cells with each other using learnable weights, we introduce a densely connected search space to cover an abundance of mainstream network designs. Moreover, by combining both path-level and channel-level sampling strategies, we design a fusion module and mixture layer to reduce the memory consumption of ample search space, hence favoring the proxyless searching. Compared with contemporary works, experiments reveal that the proxyless searching scheme is capable of bridging the gap between searching and training environments. Further, DCNAS achieves new state-of-the-art performances on public semantic image segmentation benchmarks, including 84.3% on Cityscapes, and 86.9% on PASCAL VOC 2012. We also retain leading performances when evaluating the architecture on the more challenging ADE20K and PASCAL-Context dataset. Hongmin Xu, Hong Mo, Jianchao Tan, Wenqi Ren |
CVPR | 7 |
| 2021 | Ultra-High-Definition Image Dehazing via Multi-Guided Bilateral LearningabstractConvolutional neural networks (CNNs) have achieved significant success in the single image dehazing task. Unfortunately, most existing deep dehazing models have high computational complexity, which hinders their application to high-resolution images, especially for UHD (ultra-high-definition) or 4K resolution images. To address the problem, we propose a novel network capable of real-time dehazing of 4K images on a single GPU, which consists of three deep CNNs. The first CNN extracts haze-relevant features at a reduced resolution of the hazy input and then fits locally-affine models in the bilateral space. Another CNN is used to learn multiple full-resolution guidance maps corresponding to the learned bilateral model. As a result, the feature maps with high-frequency can be reconstructed by multi-guided bilateral upsampling. Finally, the third CNN fuses the high-quality feature maps into a dehazed image. In addition, we create a large-scale 4K image dehazing dataset to support the training and testing of compared models. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art dehazing approaches on various benchmarks. Zhuoran Zheng, Wenqi Ren, Xiaochun Cao, Xiaobin Hu, Tao Wang 0053, Fenglong Song, Xiuyi Jia |
CVPR | 2 |
| 2021 | Multi-Scale Separable Network for Ultra-High-Definition Video DeblurringabstractAlthough recent research has witnessed a significant progress on the video deblurring task, these methods struggle to reconcile inference efficiency and visual quality simultaneously, especially on ultra-high-definition (UHD) videos (e.g., 4K resolution). To address the problem, we propose a novel deep model for fast and accurate UHD Video Deblurring (UHDVD). The proposed UHDVD is achieved by a separable-patch architecture, which collaborates with a multi-scale integration scheme to achieve a large receptive field without adding the number of generic convolutional layers and kernels. Additionally, we design a residual channel-spatial attention (RCSA) module to improve accuracy and reduce the depth of the network appropriately. The proposed UHDVD is the first real-time deblurring model for 4K videos at 35 fps. To train the proposed model, we build a new dataset comprised of 4K blurry videos and corresponding sharp frames using three different smartphones. Comprehensive experimental results show that our network performs favorably against the state-of-the-art methods on both the 4K dataset and public benchmarks in terms of accuracy, speed, and model size. Senyou Deng, Wenqi Ren, Yanyang Yan, Fenglong Song, Xiaochun Cao |
ICCV | 2 |
| 2021 | Pyramid Architecture Search for Real-Time Image DeblurringabstractMulti-scale and multi-patch deep models have been shown effective in removing blurs of dynamic scenes. However, these methods still suffer from one major obstacle: manually designing a lightweight and high-efficiency network is challenging and time-consuming. To tackle this obstacle, we propose a novel deblurring method, dubbed PyNAS (pyramid neural architecture search network), towards automatically designing hyper-parameters including the scales, patches, and standard cell operators. The proposed PyNAS adopts gradient-based search strategies and innovatively searches the hierarchy patch and scale scheme not limited to cell searching. Specifically, we introduce a hierarchical search strategy tailored to the multi-scale and multi-patch deblurring task. The strategy follows the principle that the first distinguishes between the top-level (pyramid-scales and pyramid-patches) and bottom-level variables (cell operators) and then searches multi-scale variables using the top-to-bottom principle. During the search stage, PyNAS employs an early stopping strategy to avoid the collapse and computational issues. Furthermore, we use a path-level binarization mechanism for multi-scale cell searching to save the memory consumption. Our primary contribution is a real-time deblurring algorithm (around 58 fps) for 720p images while achieves state-of-the-art deblurring performance on the GoPro and Video Deblurring datasets. Xiaobin Hu, Wenqi Ren, Kaicheng Yu, Kaihao Zhang, Xiaochun Cao, Wei Liu 0005, Bjoern Menze |
ICCV | 2 |
| 2021 | Benchmarking Ultra-High-Definition Image Super-resolutionabstractIncreasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-resolution UHD images. To explore their performance on UHD images, in this paper, we first introduce two large-scale image datasets, UHDSR4K and UHDSR8K, to benchmark existing SISR methods. With 70,000 V100 GPU hours of training, we benchmark these methods on 4K and 8K resolution images under seven different settings to provide a set of baseline models. Moreover, we propose a baseline model, called Mesh Attention Network (MANet) for SISR. The MANet applies the attention mechanism in both different depths (horizontal) and different levels of receptive field (vertical). In this way, correlations among feature maps are learned, enabling the network to focus on more important features. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001 |
ICCV | 4 |
| 2021 | Ultra-High-Definition Image HDR Reconstruction via Collaborative Bilateral LearningabstractExisting single image high dynamic range (HDR) reconstruction methods attempt to expand the range of illuminance. They are not effective in generating plausible textures and colors in the reconstructed results, especially for high-density pixels in ultra-high-definition (UHD) images. To address these problems, we propose a new HDR reconstruction network for UHD images by collaboratively learning color and texture details. First, we propose a dual-path network to extract the content and chromatic features at a reduced resolution of the low dynamic range (LDR) input. These two types of features are used to fit bilateral-space affine models for real-time HDR reconstruction. To extract the main data structure of the LDR input, we propose to use 3D Tucker decomposition and reconstruction to prevent pseudo edges and noise amplification in the learned bilateral grid. As a result, the high-quality content and chromatic features can be reconstructed capitalized on guided bilateral upsampling. Finally, we fuse these two full-resolution feature maps into the HDR reconstructed results. Our proposed method can achieve real-time processing for UHD images (about 160 fps). Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art HDR reconstruction approaches on public benchmarks and real-world UHD images. Zhuoran Zheng, Wenqi Ren, Xiaochun Cao, Tao Wang 0053, Xiuyi Jia |
ICCV | 2 |
| 2021 | Reciprocal Feature Learning via Explicit and Implicit Tasks in Scene Text Recognition
Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Wenqi Ren, Fei Wu 0001, Wenming Tan |
ICDAR (1) | 6 |
| 2021 | LGPMA: Complicated Table Structure Recognition with Local and Global Pyramid Mask Alignment
Liang Qiao 0001, Zaisheng Li, Zhanzhan Cheng, Peng Zhang 0075, Shiliang Pu, Wenqi Ren, Wenming Tan, Fei Wu 0001 |
ICDAR (1) | 7 |
| 2021 | A Comprehensive Survey on Image Dehazing Based on Deep LearningabstractThe presence of haze significantly reduces the quality of images. Researchers have designed a variety of algorithms for image dehazing (ID) to restore the quality of hazy images. However, there are few studies that summarize the deep learning (DL) based dehazing technologies. In this paper, we conduct a comprehensive survey on the recent proposed dehazing methods. Firstly, we conclude the commonly used datasets, loss functions and evaluation metrics. Secondly, we group the existing researches of ID into two major categories: supervised ID and unsupervised ID. The core ideas of various influential dehazing models are introduced. Finally, the open issues for future research on ID are pointed out. Jie Gui, Xiaofeng Cong, Yuan Cao 0005, Wenqi Ren, Jun Zhang 0011, Jing Zhang 0037, Dacheng Tao |
IJCAI | 4 |
| 2021 | Clustering-Induced Adaptive Structure Enhancing Network for Incomplete Multi-View DataabstractIncomplete multi-view clustering aims to cluster samples with missing views, which has drawn more and more research interest. Although several methods have been developed for incomplete multi-view clustering, they fail to extract and exploit the comprehensive global and local structure of multi-view data, so their clustering performance is limited. This paper proposes a Clustering-induced Adaptive Structure Enhancing Network (CASEN) for incomplete multi-view clustering, which is an end-to-end trainable framework that jointly conducts multi-view structure enhancing and data clustering. Our method adopts multi-view autoencoder to infer the missing features of the incomplete samples. Then, we perform adaptive graph learning and graph convolution on the reconstructed complete multi-view data to effectively extract data structure. Moreover, we use multiple kernel clustering to integrate the global and local structure for clustering, and the clustering results in turn are used to enhance the data structure. Extensive experiments on several benchmark datasets demonstrate that our method can comprehensively obtain the structure of incomplete multi-view data and achieve superior performance compared to the other methods. Zhe Xue, Junping Du 0001, Changwei Zheng, Wenqi Ren, Meiyu Liang |
IJCAI | 5 |
| 2021 | A Comprehensive Benchmark Analysis of Single Image Deraining: Current Challenges and Future Perspectives
Siyuan Li 0001, Wenqi Ren, Feng Wang 0063, Iago Breno Araujo, Eric K. Tokuda, Roberto Hirata Jr., Roberto Marcondes Cesar Junior, Zhangyang Wang, Xiaochun Cao |
Int. J. Comput. Vis. | 2 |
| 2021 | Kernel Stability for Model Selection in Kernel-Based AlgorithmsabstractModel selection is one of the fundamental problems in kernel-based algorithms, which is commonly done by minimizing an estimation of generalization error. The notion of stability and cross-validation (CV) error of learning machines consists of two widely used tools for analyzing the generalization performance. However, there are some disadvantages to both tools when applied for model selection: 1) the stability of learning machines is not practical due to the difficulty of the estimation of its specific value and 2) the CV-based estimate of generalization error usually has a relatively high variance, so it is prone to overfitting. To overcome these two limitations, we present a novel notion of kernel stability (KS) for deriving the generalization error bounds and variance bounds of CV and provide an effective approach to the application of KS for practical model selection. Unlike the existing notions of stability of the learning machine, KS is defined on the kernel matrix; hence, it can avoid the difficulty of the estimation of its value. We manifest the relationship between the KS and the popular uniform stability of the learning algorithm, and further propose several KS-based generalization error bounds and variance bounds of CV. By minimizing the proposed bounds, we present two novel KS-based criteria that can ensure good performance. Finally, we empirically analyze the performance of the proposed criteria on many benchmark data, which demonstrates that our KS-based criteria are sound and effective. Yong Liu 0018, Shizhong Liao, Hua Zhang 0008, Wenqi Ren, Weiping Wang 0005 |
IEEE Trans. Cybern. | 4 |
| 2021 | IDRLP: Image Dehazing Using Region Line PriorabstractIn this work, a novel and ultra-robust single image dehazing method called IDRLP is proposed. It is observed that when an image is divided into n regions, with each region having a similar scene depth, the brightness of both the hazy image and its haze-free correspondence are positively related with the scene depth. Based on this observation, this work determines that the hazy input and its haze-free correspondence exhibit a quasi-linear relationship after performing this region segmentation, which is named as region line prior (RLP). By combining RLP and the atmospheric scattering model (ASM), a recovery formula (RF) can be easily obtained with only two unknown parameters, i.e., the slope of the linear function and the atmospheric light. A 2D joint optimization function considering two constraints is then designed to seek the solution of RF. Unlike other comparable works, this "joint optimization" strategy makes efficient use of the information across the entire image, leading to more accurate results with ultra-high robustness. Finally, a guided filter is introduced in RF to eliminate the adverse interference caused by the region segmentation. The proposed RLP and IDRLP are evaluated from various perspectives and compared with related state-of-the-art techniques. Extensive analysis verifies the superiority of IDRLP over state-of-the-art image dehazing techniques in terms of both the recovery quality and efficiency. A software release is available at https://sites.google.com/site/renwenqi888/. Mingye Ju, Can Ding 0002, Charles A. Guo, Wenqi Ren, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2021 | IDE: Image Dehazing and Exposure Using an Enhanced Atmospheric Scattering ModelabstractAtmospheric scattering model (ASM) is one of the most widely used model to describe the imaging processing of hazy images. However, we found that ASM has an intrinsic limitation which leads to a dim effect in the recovered results. In this paper, by introducing a new parameter, i.e., light absorption coefficient, into ASM, an enhanced ASM (EASM) is attained, which can address the dim effect and better model outdoor hazy scenes. Relying on this EASM, a simple yet effective gray-world-assumption-based technique called IDE is then developed to enhance the visibility of hazy images. Experimental results show that IDE eliminates the dim effect and exhibits excellent dehazing performance. It is worth mentioning that IDE does not require any training process or extra information related to scene depth, which makes it very fast and robust. Moreover, the global stretch strategy used in IDE can effectively avoid some undesirable effects in recovery results, e.g., over-enhancement, over-saturation, and mist residue, etc. Comparison between the proposed IDE and other state-of-the-art techniques reveals the superiority of IDE in terms of both dehazing quality and efficiency over all the comparable techniques. Mingye Ju, Can Ding 0002, Wenqi Ren, Yi Yang 0001, Dengyin Zhang, Y. Jay Guo |
IEEE Trans. Image Process. | 3 |
| 2021 | Underwater Image Enhancement via Medium Transmission-Guided Multi-Color Space EmbeddingabstractUnderwater images suffer from color casts and low contrast due to wavelength- and distance-dependent attenuation and scattering. To solve these two degradation issues, we present an underwater image enhancement network via medium transmission-guided multi-color space embedding, called Ucolor. Concretely, we first propose a multi-color space encoder network, which enriches the diversity of feature representations by incorporating the characteristics of different color spaces into a unified structure. Coupled with an attention mechanism, the most discriminative features extracted from multiple color spaces are adaptively integrated and highlighted. Inspired by underwater imaging physical models, we design a medium transmission (indicating the percentage of the scene radiance reaching the camera)-guided decoder network to enhance the response of network towards quality-degraded regions. As a result, our network can effectively improve the visual quality of underwater images by exploiting multiple color spaces embedding and the advantages of both physical model-based and learning-based methods. Extensive experiments demonstrate that our Ucolor achieves superior performance against state-of-the-art methods in terms of both visual quality and quantitative metrics. The code is publicly available at: https://li-chongyi.github.io/Proj_Ucolor.html. Chongyi Li, Saeed Anwar, Junhui Hou, Runmin Cong, Chunle Guo, Wenqi Ren |
IEEE Trans. Image Process. | 6 |
| 2021 | SRGAT: Single Image Super-Resolution With Graph Attention NetworkabstractDeep neural networks have demonstrated remarkable reconstruction for single-image super-resolution (SISR). However, most existing CNN-based SISR methods directly learn the relation between low-resolution (LR) and high-resolution (HR) images, neglecting to explore the recurrence of internal patches, hence hindering the representational power of CNNs. In this paper, we propose a novel single image Super-Resolution network based on Graph ATtention network (SRGAT) to make full use of the internal patch-recurrence in a natural image. The proposed model employs a feature mapping block with a recurrent structure to refine low-level representations with high-level information. Especifically, the feature mapping block contains a parallel graph similarity branch and a content branch, where the graph similarity branch aims at exploiting the similarity and symmetry across different image patches in low-resolution feature space and provides additional priors for the content branch to enhance texture details. Specifically, we consider the internal patch-recurrence of an image by constructing a graph network on image feature patches. In this way, the information from neighboring patches can be interacted using graph attention network (GAT) to help it recover additional textures, which complements the textures learned from the content branch. Extensive quantitative and qualitative evaluations on five benchmark datasets demonstrate that the proposed algorithm performs favorably against the state-of-the-art super-resolution methods. Yanyang Yan, Wenqi Ren, Xiaobin Hu, Kun Li 0029, Haifeng Shen, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2021 | Robust Online Tracking via Contrastive Spatio-Temporal Aware NetworkabstractExisting tracking-by-detection approaches using deep features have achieved promising results in recent years. However, these methods mainly exploit feature representations learned from individual static frames, thus paying little attention to the temporal smoothness between frames. This easily leads trackers to drift in the presence of large appearance variations and occlusions. To address this issue, we propose a two-stream network to learn discriminative spatio-temporal feature representations to represent the target objects. The proposed network consists of a Spatial ConvNet module and a Temporal ConvNet module. Specifically, the Spatial ConvNet adopts 2D convolutions to encode the target-specific appearance in static frames, while the Temporal ConvNet models the temporal appearance variations using 3D convolutions and learns consistent temporal patterns in a short video clip. Then we propose a proposal refinement module to adjust the predicted bounding box, which can make the target localizing outputs to be more consistent in video sequences. In addition, to improve the model adaptation during online update, we propose a contrastive online hard example mining (OHEM) strategy, which selects hard negative samples and enforces them to be embedded in a more discriminative feature space. Extensive experiments conducted on the OTB, Temple Color and VOT benchmarks demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods. Siyuan Yao, Hua Zhang 0008, Wenqi Ren, Chao Ma 0004, Xiaoguang Han 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2021 | Dual Attention-in-Attention Model for Joint Rain Streak and Raindrop RemovalabstractRain streaks and raindrops are two natural phenomena, which degrade image capture in different ways. Currently, most existing deep deraining networks take them as two distinct problems and individually address one, and thus cannot deal adequately with both simultaneously. To address this, we propose a Dual Attention-in-Attention Model (DAiAM) which includes two DAMs for removing both rain streaks and raindrops. Inside the DAM, there are two attentive maps - each of which attends to the heavy and light rainy regions, respectively, to guide the deraining process differently for applicable regions. In addition, to further refine the result, a Differential-driven Dual Attention-in-Attention Model (D-DAiAM) is proposed with a "heavy-to-light" scheme to remove rain via addressing the unsatisfying deraining regions. Extensive experiments on one public raindrop dataset, one public rain streak and our synthesized joint rain streak and raindrop (JRSRD) dataset have demonstrated that the proposed method not only is capable of removing rain streaks and raindrops simultaneously, but also achieves the state-of-the-art performance on both tasks. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren |
IEEE Trans. Image Process. | 4 |
| 2020 | Domain Adaptation for Image DehazingabstractImage dehazing using learning-based methods has achieved state-of-the-art performance in recent years. However, most existing methods train a dehazing model on synthetic hazy images, which are less able to generalize well to real hazy images due to domain shift. To address this issue, we propose a domain adaptation paradigm, which consists of an image translation module and two image dehazing modules. Specifically, we first apply a bidirectional translation network to bridge the gap between the synthetic and real domains by translating images from one domain to another. And then, we use images before and after translation to train the proposed two image dehazing networks with a consistency constraint. In this phase, we incorporate the real hazy image into the dehazing training via exploiting the properties of the clear image (e.g., dark channel prior and image gradient smoothing) to further improve the domain adaptivity. By training image translation and dehazing network in an end-to-end manner, we can obtain better effects of both image translation and dehazing. Experimental results on both synthetic and real-world images demonstrate that our model performs favorably against the state-of-the-art dehazing algorithms. Yuanjie Shao, Lerenhan Li, Wenqi Ren, Changxin Gao, Nong Sang |
CVPR | 3 |
| 2020 | Face Super-Resolution Guided by 3D Facial Priors
Xiaobin Hu, Wenqi Ren, John LaMaster, Xiaochun Cao, Xiaoming Li 0002, Zechao Li, Bjoern Menze, Wei Liu 0005 |
ECCV (4) | 2 |
| 2020 | Single Image Super-Resolution via a Holistic Attention Network
Weilei Wen, Wenqi Ren, Xiangde Zhang, Lianping Yang, Shuzhen Wang, Kaihao Zhang, Xiaochun Cao, Haifeng Shen |
ECCV (12) | 3 |
| 2020 | Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding
Kaihao Zhang, Wenhan Luo, Wenqi Ren, Jingwen Wang 0003, Fang Zhao 0006, Lin Ma 0002, Hongdong Li |
ECCV (27) | 3 |
| 2020 | Every Moment Matters: Detail-Aware Networks to Bring a Blurry Image AliveabstractMotion-blurred images are the result of light accumulation over the period of camera exposure time, during which the camera and objects in the scene are in relative motion to each other. The inverse process of extracting an image sequence from a single motion-blurred image is an ill-posed vision problem. One key challenge is that the motions across frames are subtle, which makes the generating networks difficult to capture them and thus the recovery sequences lack motion details. In order to alleviate this problem, we propose a detail-aware network with three consecutive stages to improve the reconstruction quality by addressing specific aspects in the recovery process. The detail-aware network firstly models the dynamics using a cycle flow loss, resolving the temporal ambiguity of the reconstruction in the first stage. Then, a GramNet is proposed in the second stage to refine subtle motion between continuous frames using Gram matrices as motion representation. Finally, we introduce a HeptaGAN in the third stage to bridge the continuous and discrete nature of exposure time and recovered frames, respectively, in order to maintain rich detail. Experiments show that the proposed detail-aware networks produce sharp image sequences with rich details and subtle motion, outperforming the state-of-the-art methods. Kaihao Zhang, Wenhan Luo, Björn Stenger, Wenqi Ren, Lin Ma 0002, Hongdong Li |
ACM Multimedia | 4 |
| 2020 | Single Image Dehazing via Multi-scale Convolutional Neural Networks with Holistic Edges
Wenqi Ren, Jinshan Pan, Hua Zhang 0008, Xiaochun Cao, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2020 | NLDN: Non-local dehazing network for dense haze removal
Shengdong Zhang, Fazhi He, Wenqi Ren |
Neurocomputing | 3 |
| 2020 | Photo-realistic dehazing via contextual generative adversarial networks
Shengdong Zhang, Fazhi He, Wenqi Ren |
Mach. Vis. Appl. | 3 |
| 2020 | Semi-Supervised Image DehazingabstractWe present an effective semi-supervised learning algorithm for single image dehazing. The proposed algorithm applies a deep Convolutional Neural Network (CNN) containing a supervised learning branch and an unsupervised learning branch. In the supervised branch, the deep neural network is constrained by the supervised loss functions, which are mean squared, perceptual, and adversarial losses. In the unsupervised branch, we exploit the properties of clean images via sparsity of dark channel and gradient priors to constrain the network. We train the proposed network on both the synthetic data and real-world images in an end-to-end manner. Our analysis shows that the proposed semi-supervised learning algorithm is not limited to synthetic training datasets and can be generalized well to real-world images. Extensive experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art single image dehazing algorithms on both benchmark datasets and real-world images. Lerenhan Li, Yunlong Dong, Wenqi Ren, Jinshan Pan, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | An Underwater Image Enhancement Benchmark Dataset and BeyondabstractUnderwater image enhancement has been attracting much attention due to its significance in marine engineering and aquatic robotics. Numerous underwater image enhancement algorithms have been proposed in the last few years. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real-world images. It is thus unclear how these algorithms would perform on images acquired in the wild and how we could gauge the progress in the field. To bridge this gap, we present the first comprehensive perceptual study and analysis of underwater image enhancement using large-scale real-world images. In this paper, we construct an Underwater Image Enhancement Benchmark (UIEB) including 950 real-world underwater images, 890 of which have the corresponding reference images. We treat the rest 60 underwater images which cannot obtain satisfactory reference images as challenging data. Using this dataset, we conduct a comprehensive study of the state-of-the-art underwater image enhancement algorithms qualitatively and quantitatively. In addition, we propose an underwater image enhancement network (called Water-Net) trained on this benchmark as a baseline, which indicates the generalization of the proposed UIEB for training Convolutional Neural Networks (CNNs). The benchmark evaluations and the proposed Water-Net demonstrate the performance and limitations of state-of-the-art algorithms, which shed light on future research in underwater image enhancement. The dataset and code are available at. Chongyi Li, Chunle Guo, Wenqi Ren, Runmin Cong, Junhui Hou, Sam Kwong, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2020 | Background Noise Filtering and Distribution Dividing for Crowd CountingabstractCrowd counting is a challenging problem due to the diverse crowd distribution and background interference. In this paper, we propose a new approach for head size estimation to reduce the impact of different crowd scale and background noise. Different from just using local information of distance between human heads, the global information of the people distribution in the whole image is also under consideration. We obey the order of far- to near-region (small to large) to spread head size, and ensure that the propagation is uninterrupted by inserting dummy head points. The estimated head size is further exploited, such as dividing the crowd into parts of different densities and generating a high-fidelity head mask. On the other hand, we design three different head mask usage mechanisms and the corresponding head masks to analyze where and which mask could lead to better background filtering1. Based on the learned masks, two competitive models are proposed which can perform robust crowd estimation against background noise and diverse crowd scale. We evaluate the proposed method on three public crowd counting datasets of ShanghaiTech [2], UCFQNRF [3] and UCFCC_50 [4]. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art crowd counting approaches. Hong Mo, Wenqi Ren, Yuan Xiong, Xiaoqi Pan, Zhong Zhou, Xiaochun Cao, Wei Wu 0008 |
IEEE Trans. Image Process. | 2 |
| 2020 | Learning Interleaved Cascade of Shrinkage Fields for Joint Image Dehazing and DenoisingabstractMost existing image dehazing methods deteriorate to different extents when processing hazy inputs with noise. The main reason is that the commonly adopted two-step strategy tends to amplify noise in the inverse operation of division by the transmission. To address this problem, we learn an interleaved Cascade of Shrinkage Fields (CSF) to reduce noise in jointly recovering the transmission map and the scene radiance from a single hazy image. Specifically, an auxiliary shrinkage field (SF) model is integrated into each cascade of the proposed scheme to reduce undesirable artifacts during the transmission estimation. Different from conventional CSF, our learned SF models have special visual patterns, which facilitate the specific task of noise reduction in haze removal. Furthermore, a numerical algorithm is proposed to efficiently update the scene radiance and the transmission map in each cascade. Extensive experiments on synthetic and real-world data demonstrate that the proposed algorithm performs favorably against state-of-the-art dehazing methods on hazy and noisy images. Qingbo Wu 0002, Wenqi Ren, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2020 | Accurate Transmission Estimation for Removing Haze and Noise From a Single ImageabstractImage noise usually causes depth-dependent visual artifacts in single image dehazing. Most existing dehazing methods exploit a two-step strategy in the restoration, which inevitably leads to inaccurate transmission maps and low-quality scene radiance for noisy and hazy inputs. To address these problems, we present a novel variational model for joint recovery of the transmission map and the scene radiance from a single image. In the model, we propose a transmission-aware non-local regularization to avoid noise amplification by adaptively suppressing noise and preserving fine details in the recovered image. Meanwhile, to improve the accuracy of transmission estimation, we introduce a semantic-guided regularization to smooth out the transmission map while keeping depth inconsistency at the boundaries of different objects. Furthermore, we design an alternating scheme to jointly optimize the transmission map and the scene radiance as well as the segmentation map. Extensive experiments on synthetic and real-world data demonstrate that the proposed algorithm performs favorably against state-of-the-art dehazing methods on noisy and hazy images. Qingbo Wu 0002, Jingang Zhang, Wenqi Ren, Wangmeng Zuo, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 3 |
| 2020 | Joint learning of image detail and transmission map for single image dehazing
Shengdong Zhang, Fazhi He, Wenqi Ren, Jian Yao 0002 |
Vis. Comput. | 3 |
| 2019 | Semantic Correlations Loss: Improving Model Interpretability for Multi-class ClassificationabstractDespite that convolutional neural networks (CNNs) have recently demonstrated high-quality object classification, the trained models suffer from their extreme unexplainability. In this paper, we propose a general method, named as semantic correlation loss, for introducing common-sense knowledge to CNN architectures. In contrast to traditional cross-entropy loss which only considers the ground-truth class, we exploit to be aware of the accuracy of all classes. By adding this simple add-on, current multi-class classification models are able to improve on the ability of “making mistakes reasonably”. In addition, a slight performance gain is also achieved. Experimental results on CUB-200-2011, CIFAR -10 and 100 are provided to demonstrate the efficacy of our proposed method. Moreover, this novel loss is able to be applied in any setting as long as the labels of training data are included in the common sense knowledge base. Xuezhi Tong, Rui Wang 0032, Xiaochun Cao, Wenqi Ren |
IEEE BigData | 4 |
| 2019 | Single Image Deraining: A Comprehensive Benchmark AnalysisabstractWe present a comprehensive study and evaluation of existing single image deraining algorithms, using a new large-scale benchmark consisting of both synthetic and real-world rainy images.This dataset highlights diverse data sources and image contents, and is divided into three subsets (rain streak, rain drop, rain and mist), each serving different training or evaluation purposes. We further provide a rich variety of criteria for dehazing algorithm evaluation, ranging from full-reference metrics, to no-reference metrics, to subjective evaluation and the novel task-driven evaluation. Experiments on the dataset shed light on the comparisons and limitations of state-of-the-art deraining algorithms, and suggest promising future directions. Siyuan Li 0001, Iago Breno Araujo, Wenqi Ren, Zhangyang Wang, Eric K. Tokuda, Roberto Hirata Jr., Roberto Marcondes Cesar Junior, Jiawan Zhang, Xiaojie Guo 0001, Xiaochun Cao |
CVPR | 3 |
| 2019 | Face Video Deblurring Using 3D Facial PriorsabstractExisting face deblurring methods only consider single frames and do not account for facial structure and identity information. These methods struggle to deblur face videos that exhibit significant pose variations and misalignment. In this paper we propose a novel face video deblurring network capitalizing on 3D facial priors. The model consists of two main branches: i) a face video deblurring sub-network based on an encoder-decoder architecture, and ii) a 3D face reconstruction and rendering branch for predicting 3D priors of salient facial structures and identity knowledge. These structures encourage the deblurring branch to generate sharp faces with detailed structures. Our method not only uses low-level information (i.e., image intensity), but also middle-level information (i.e., 3D facial structure) and high-level knowledge (i.e., identity content) to further explore spatial constraints of facial components from blurry face frames. Extensive experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods. Wenqi Ren, Jiaolong Yang, Senyou Deng, David P. Wipf, Xiaochun Cao, Xin Tong 0001 |
ICCV | 1 |
| 2019 | Deep Correlated Predictive Subspace Learning for Incomplete Multi-View Semi-Supervised ClassificationabstractIncomplete view information often results in failure cases of the conventional multi-view methods. To address this problem, we propose a Deep Correlated Predictive Subspace Learning (DCPSL) method for incomplete multi-view semi-supervised classification. Specifically, we integrate semi-supervised deep matrix factorization, correlated subspace learning, and multi-view label prediction into a unified framework to jointly learn the deep correlated predictive subspace and multi-view shared and private label predictors. DCPSL is able to learn proper subspace representation that is suitable for class label prediction, which can further improve the performance of classification. Extensive experimental results on various practical datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods. Zhe Xue, Junping Du 0001, Dawei Du, Wenqi Ren, Siwei Lyu |
IJCAI | 4 |
| 2019 | Single image rain removal via a deep decomposition-composition network
Siyuan Li 0001, Wenqi Ren, Jiawan Zhang, Jinke Yu, Xiaojie Guo 0001 |
Comput. Vis. Image Underst. | 2 |
| 2019 | An efficient privacy protection scheme for data security in video surveillance
Wei Zhang 0031, Huazhu Fu, Wenqi Ren, Xinpeng Zhang 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Learning to Deblur Images with ExemplarsabstractHuman faces are one interesting object class with numerous applications. While significant progress has been made in the generic deblurring problem, existing methods are less effective for blurry face images. The success of the state-of-the-art image deblurring algorithms stems mainly from implicit or explicit restoration of salient edges for kernel estimation. However, existing methods are less effective as only few edges can be restored from blurry face images for kernel estimation. In this paper, we address the problem of deblurring face images by exploiting facial structures. We propose a deblurring algorithm based on an exemplar dataset without using coarse-to-fine strategies or heuristic edge selections. In addition, we develop a convolutional neural network to restore sharp edges from blurry images for deblurring. Extensive experiments against the state-of-the-art methods demonstrate the effectiveness of the proposed algorithm for deblurring face images. In addition, we show that the proposed algorithms can be applied to image deblurring for other object classes. Jinshan Pan, Wenqi Ren, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Recolored Image Detection via a Deep Discriminative ModelabstractImage recoloring is a technique that can transfer image color or theme and result in an imperceptible change in human eyes. Although image recoloring is one of the most important image manipulation techniques, there is no special method designed for detecting this kind of forgery. In this paper, we propose a trainable end-to-end system for distinguishing recolored images from natural images. The proposed network takes the original image and two derived inputs based on illumination consistency and inter-channel correlation of the original input into consideration and outputs the probability that it is recolored. Our algorithm adopts a convolutional neural network (CNN)- based deep architecture, which consists of three feature extraction blocks and a feature fusion module. To train the deep neural network, we synthesize a data set comprised of recolored images and corresponding ground truth using different recoloring methods. Extensive experimental results on the recolored images generated by various methods show that our proposed network is well generalized and very robust. Yanyang Yan, Wenqi Ren, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Benchmarking Single-Image Dehazing and BeyondabstractIn this paper, we present a comprehensive study and evaluation of existing single image dehazing algorithms, using a new large-scale benchmark consisting of both synthetic and real-world hazy images, called REalistic Single Image DEhazing (RESIDE). RESIDE highlights diverse data sources and image contents, and is divided into five subsets, each serving different training or evaluation purposes. We further provide a rich variety of criteria for dehazing algorithm evaluation, ranging from full-reference metrics, to no-reference metrics, to subjective evaluation and the novel task-driven evaluation. Experiments on RESIDE shed light on the comparisons and limitations of stateof- the-art dehazing algorithms, and suggest promising future directions. Boyi Li 0001, Wenqi Ren, Dengpan Fu, Dacheng Tao, Dan Feng 0001, Wenjun Zeng 0001, Zhangyang Wang |
IEEE Trans. Image Process. | 2 |
| 2019 | Low-Light Image Enhancement via a Deep Hybrid NetworkabstractCamera sensors often fail to capture clear images or videos in a poorly lit environment. In this paper, we propose a trainable hybrid network to enhance the visibility of such degraded images. The proposed network consists of two distinct streams to simultaneously learn the global content and the salient structures of the clear image in a unified network. More specifically, the content stream estimates the global content of the low-light input through an encoder-decoder network. However, the encoder in the content stream tends to lose some structure details. To remedy this, we propose a novel spatially variant recurrent neural network (RNN) as an edge stream to model edge details, with the guidance of another auto-encoder. The experimental results show that the proposed network favorably performs against the state-of-the-art low-light image enhancement algorithms. Wenqi Ren, Sifei Liu, Lin Ma 0002, Qianqian Xu 0001, Xiangyu Xu 0002, Xiaochun Cao, Junping Du 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Deep Video Dehazing With Semantic SegmentationabstractRecent research have shown the potential of using convolutional neural networks (CNNs) to accomplish single image dehazing. In this work, we take one step further to explore the possibility of exploiting a network to perform haze removal for videos. Unlike single image dehazing, video based approaches can take advantage of the abundant information that exists across neighboring frames. In this work, assuming that a scene point yields highly correlated transmission values between adjacent video frames, we develop a deep learning solution for video dehazing, where a CNN is trained end-to-end to learn how to accumulate information across frames for transmission estimation. The estimated transmission map is subsequently used to recover a haze-free frame via atmospheric scattering model. In addition, as the semantic information of a scene provides a strong prior for image restoration, we propose to incorporate global semantic priors as input to regularize the transmission maps so that the estimated maps can be smooth in the regions of the same object and only discontinuous across the boundaries of different objects. To train this network, we generate a dataset consisted of synthetic hazy and haze-free videos for supervision based on the NYU depth dataset. We show that the features learned from this dataset are capable of removing haze that arises in outdoor scenes in a wide range of videos. Extensive experiments demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods on both synthetic and real-world videos. Wenqi Ren, Jingang Zhang, Xiangyu Xu 0002, Lin Ma 0002, Xiaochun Cao, Gaofeng Meng, Wei Liu 0005 |
IEEE Trans. Image Process. | 1 |
| 2018 | Gated Fusion Network for Single Image DehazingabstractIn this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while the decoder is employed to estimate the contribution of each input to the final dehazed result using the learned representations attributed to the encoder. The constructed network adopts a novel fusion-based strategy which derives three inputs from an original hazy image by applying White Balance (WB), Contrast Enhancing (CE), and Gamma Correction (GC). We compute pixel-wise confidence maps based on the appearance differences between these different inputs to blend the information of the derived inputs and preserve the regions with pleasant visibility. The final dehazed image is yielded by gating the important features of the derived inputs. To train the network, we introduce a multi-scale approach such that the halo artifacts can be avoided. Extensive experimental results on both synthetic and real-world images demonstrate that the proposed algorithm performs favorably against the state-of-the-art algorithms. Wenqi Ren, Lin Ma 0002, Jiawei Zhang 0002, Jinshan Pan, Xiaochun Cao, Wei Liu 0005, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2018 | Rendering Portraitures from Monocular Camera and Beyond
Xiangyu Xu 0002, Deqing Sun, Sifei Liu, Wenqi Ren, Yu-Jin Zhang, Ming-Hsuan Yang 0001 |
ECCV (9) | 4 |
| 2018 | Feed-Net: Fully End-to-End DehazingabstractThis paper proposes an image dehazing model built with a fully convolutional neural network (CNN), called Fully End-to-End Dehazing Network (FEED-Net). In contrast to estimate the transmission map and the atmospheric light separately as most previous deep learning methods, FEED-Net recovers the hazy-free image directly from a hazy image via a light-weight CNN. In addition, we introduce contextual information into dehazing via dilated convolution and use dense skip connection for feature fusion, which makes end-to-end dehazing possible. Experimental results show our method outperforms the state-of-the-art algorithms on both synthetic dataset and real-world images. Shengdong Zhang, Wenqi Ren, Jian Yao 0002 |
ICME | 2 |
| 2018 | Deep Non-Blind Deconvolution via Generalized Low-Rank ApproximationabstractIn this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We first compute a generalized low-rank approximation for a large number of blur kernels, and then use separable filters to initialize the convolutional parameters in the network. Our analysis shows that the estimated decomposed matrices contain the most essential information of the input kernel, which ensures the proposed network to handle various blurs in a unified framework and generate high-quality deblurring results. Experimental results on benchmark datasets with noise and saturated pixels demonstrate that the proposed algorithm performs favorably against state-of-the-art methods. Wenqi Ren, Jiawei Zhang 0002, Lin Ma 0002, Jinshan Pan, Xiaochun Cao, Wangmeng Zuo, Wei Liu 0005, Ming-Hsuan Yang 0001 |
NeurIPS | 1 |
| 2017 | Image Deblurring via Extreme Channels PriorabstractCamera motion introduces motion blur, affecting many computer vision tasks. Dark Channel Prior (DCP) helps the blind deblurring on scenes including natural, face, text, and low-illumination images. However, it has limitations and is less likely to support the kernel estimation while bright pixels dominate the input image. We observe that the bright pixels in the clear images are not likely to be bright after the blur process. Based on this observation, we first illustrate this phenomenon mathematically and define it as the Bright Channel Prior (BCP). Then, we propose a technique for deblurring such images which elevates the performance of existing motion deblurring algorithms. The proposed method takes advantage of both Bright and Dark Channel Prior. This joint prior is named as extreme channels prior and is crucial for achieving efficient restorations by leveraging both the bright and dark information. Extensive experimental results demonstrate that the proposed method is more robust and performs favorably against the state-of-the-art image deblurring methods on both synthesized and natural images. Yanyang Yan, Wenqi Ren, Yuanfang Guo, Rui Wang 0032, Xiaochun Cao |
CVPR | 2 |
| 2017 | Video Deblurring via Semantic Segmentation and Pixel-Wise Non-linear KernelabstractVideo deblurring is a challenging problem as the blur is complex and usually caused by the combination of camera shakes, object motions, and depth variations. Optical flow can be used for kernel estimation since it predicts motion trajectories. However, the estimates are often inaccurate in complex scenes at object boundaries, which are crucial in kernel estimation. In this paper, we exploit semantic segmentation in each blurry frame to understand the scene contents and use different motion models for image regions to guide optical flow estimation. While existing pixel-wise blur models assume that the blur kernel is the same as optical flow during the exposure time, this assumption does not hold when the motion blur trajectory at a pixel is different from the estimated linear optical flow. We analyze the relationship between motion blur trajectory and optical flow, and present a novel pixel-wise non-linear kernel model to account for motion blur. The proposed blur model is based on the non-linear optical flow, which describes complex motion blur more effectively. Extensive experiments on challenging blurry videos demonstrate the proposed algorithm performs favorably against the state-of-the-art methods. Wenqi Ren, Jinshan Pan, Xiaochun Cao, Ming-Hsuan Yang 0001 |
ICCV | 1 |
| 2016 | SketchNet: Sketch Classification with Web ImagesabstractIn this study, we present a weakly supervised approach that discovers the discriminative structures of sketch images, given pairs of sketch images and web images. In contrast to traditional approaches that use global appearance features or relay on keypoint features, our aim is to automatically learn the shared latent structures that exist between sketch images and real images, even when there are significant appearance differences across its relevant real images. To accomplish this, we propose a deep convolutional neural network, named SketchNet. We firstly develop a triplet composed of sketch, positive and negative real image as the input of our neural network. To discover the coherent visual structures between the sketch and its positive pairs, we introduce the softmax as the loss function. Then a ranking mechanism is introduced to make the positive pairs obtain a higher score comparing over negative ones to achieve robust representation. Finally, we formalize above-mentioned constrains into the unified objective function, and create an ensemble feature representation to describe the sketch images. Experiments on the TUBerlin sketch benchmark demonstrate the effectiveness of our model and show that deep feature representation brings substantial improvements over other state-of-the-art methods on sketch classification. Hua Zhang 0008, Si Liu 0001, Changqing Zhang 0002, Wenqi Ren, Rui Wang 0032, Xiaochun Cao |
CVPR | 4 |
| 2016 | Single Image Dehazing via Multi-scale Convolutional Neural Networks
Wenqi Ren, Si Liu 0001, Hua Zhang 0008, Jinshan Pan, Xiaochun Cao, Ming-Hsuan Yang 0001 |
ECCV (2) | 1 |
| 2016 | Image Deblurring via Enhanced Low-Rank PriorabstractLow-rank matrix approximation has been successfully applied to numerous vision problems in recent years. In this paper, we propose a novel low-rank prior for blind image deblurring. Our key observation is that directly applying a simple low-rank model to a blurry input image significantly reduces the blur even without using any kernel information, while preserving important edge information. The same model can be used to reduce blur in the gradient map of a blurry input. Based on these properties, we introduce an enhanced prior for image deblurring by combining the low rank prior of similar patches from both the blurry image and its gradient map. We employ a weighted nuclear norm minimization method to further enhance the effectiveness of low-rank prior for image deblurring, by retaining the dominant edges and eliminating fine texture and slight edges in intermediate images, allowing for better kernel estimation. In addition, we evaluate the proposed enhanced low-rank prior for both the uniform and the non-uniform deblurring. Quantitative and qualitative experimental evaluations demonstrate that the proposed algorithm performs favorably against the state-of-the-art deblurring methods. Wenqi Ren, Xiaochun Cao, Jinshan Pan, Xiaojie Guo 0001, Wangmeng Zuo, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Scene Text Deblurring Using Text-Specific Multiscale DictionariesabstractTexts in natural scenes carry critical semantic clues for understanding images. When capturing natural scene images, especially by handheld cameras, a common artifact, i.e., blur, frequently happens. To improve the visual quality of such images, deblurring techniques are desired, which also play an important role in character recognition and image understanding. In this paper, we study the problem of recovering the clear scene text by exploiting the text field characteristics. A series of text-specific multiscale dictionaries (TMD) and a natural scene dictionary is learned for separately modeling the priors on the text and nontext fields. The TMD-based text field reconstruction helps to deal with the different scales of strings in a blurry image effectively. Furthermore, an adaptive version of nonuniform deblurring method is proposed to efficiently solve the real-world spatially varying problem. Dictionary learning allows more flexible modeling with respect to the text field property, and the combination with the nonuniform method is more appropriate in real situations where blur kernel sizes are depth dependent. Experimental results show that the proposed method achieves the deblurring results with better visual quality than the state-of-the-art methods. Xiaochun Cao, Wenqi Ren, Wangmeng Zuo, Xiaojie Guo 0001, Hassan Foroosh |
IEEE Trans. Image Process. | 2 |