EDBT 2026 Demo / reviewers in the wild / expert
Yuhui Quan
dblp:62/10771
· DBLP profile ↗
90ranked-venue papers
32as first author
55since 2021 · last 2026
0000-0002-2564-7703ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 27 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 57 · 20 first-author · 35 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nighttime image dehazing via a physics-aware dynamic neural model with progressive contrastive regularization
Yun Liang 0003, Xinjie Xiao, Zihan Zhou 0007, Lianghui Li, Yuhui Quan |
Pattern Recognit. | 5 |
| 2025 | Zero-Shot Low-Light Image Enhancement via Latent Diffusion ModelsabstractLow-light image enhancement (LLIE) aims to improve visibility and signal-to-noise ratio in images captured under poor lighting conditions. While deep learning has shown promise in this domain, current approaches require extensive paired training data, limiting their practical utility. We present a novel framework that reformulates low-light image enhancement as a zero-shot inference problem using pre-trained latent diffusion models (LDMs), eliminating the need for task-specific training data. Our key insight is that the rich natural image priors encoded in LDMs can be leveraged to recover well-lit images through a carefully designed optimization process. To address the ill-posed nature of low-light degradation and the complexity of latent space optimization, our framework introduces an exposure-aware degradation module that adaptively models illumination variations and a principled latent regularization scheme with adaptive guidance that ensures both enhancement quality and natural image statistics. Experimental results demonstrate that our framework outperforms existing zero-shot methods across diverse real-world scenarios. Yan Huang 0031, Xiaoshan Liao, Jinxiu Liang, Yuhui Quan, Boxin Shi, Yong Xu 0007 |
AAAI | 4 |
| 2025 | Multi-Focus Image Fusion via Explicit Defocus Blur ModellingabstractMulti-focus image fusion (MFIF) enhances depth of field in photography by generating an all-in-focus image from multiple images captured at different focal lengths. While deep learning has shown promise in MFIF, most existing methods overlooked the physical properties of defocus blurring in their network design, limiting their interoperability and generalization. This paper introduces a novel framework that integrates explicit defocus blur modelling into the MFIF process, improving both interpretability and performance. Using an atom-based spatially-varying parameterized defocus blurring model, our approach calculates pixel-wise defocus descriptors and initial focused images from multi-focus source images in a scale-recurrent manner to estimate soft decision maps. Fusion is then performed using masks derived from these decision maps, with special treatment for pixels likely defocused in all source images or near boundaries of defocused/focused regions. The model is trained with a fusion loss and a cross-scale defocus estimation loss. Extensive experiments on benchmark datasets demonstrated the effectiveness of our approach. Yuhui Quan, Xi Wan, Zitao Tang, Jinxiu Liang, Hui Ji 0002 |
AAAI | 1 |
| 2025 | A Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse ArtifactsabstractStructured artifacts are semi-regular, repetitive patterns that closely intertwine with genuine image content, making their removal highly challenging. In this paper, we introduce the Scale-Adaptive Deformable Transformer, an network architecture specifically designed to eliminate such artifacts from images. The proposed network features two key components: a scale-enhanced deformable convolution module for modeling scale-varying patterns with abundant orientations and potential distortions, and a scale-adaptive deformable attention mechanism for capturing long-range relationships among repetitive patterns with different sizes and non-uniform spatial distributions. Extensive experiments show that our network consistently outperforms state-of-the-art methods in diverse artifact removal tasks, including image deraining, image demoiréing, and image debanding. Xuyi He, Yuhui Quan, Ruotao Xu, Hui Ji 0002 |
CVPR | 2 |
| 2025 | Zero-Shot Blind-spot Image Denoising via Implicit Neural SamplingabstractThe blind-spot principle has been a widely used tool in zero-shot image denoising but faces challenges with real-world noise that exhibits strong local correlations. Existing methods focus on reducing noise correlation, which also weaken the pixel correlations needed for accurately estimating missing pixels. In this paper, we first present a rigorous analysis of how noise correlation and pixel correlation impact the statistical risk of a linear blind-spot denoiser. We then propose using an implicit neural representation to resample noisy pixels, effectively reducing noise correlation while preserving the essential pixel correlations for successful blind-spot denoising. Extensive experiments show our method surpasses existing zero-shot de-noising techniques on real-world noisy images. Yuhui Quan, Hui Ji 0002 |
CVPR | 1 |
| 2025 | Fingerprinting Denoising Diffusion Probabilistic ModelsabstractDiffusion models, especially denoising diffusion probabilistic models (DDPMs), are prevalent tools in generative AI, making their intellectual property (IP) protection increasingly important. Most existing IP protection methods for DDPMs are invasive, e.g., model watermarking, which alter model parameters and raise concerns about performance degradation, also with requirement for extra computational resources for retraining or fine-tuning. In this paper, we propose the first non-invasive fingerprinting scheme for DDPMs, requiring no parameter changes or fine-tuning, and keeping generation quality intact. We introduce a discriminative and robust fingerprint latent space based on the well-designed "crossing route" of noisy samples that span the performance border-zone of DDPMs, with only black-box access required for the diffusion denoiser in ownership verification. Extensive experiments demonstrate that our fingerprinting approach enjoys both robustness against the often-seen attacks and distinctiveness on various DDPMs, providing an alternative for protecting DDPMs’ IP rights without compromising their performance or integrity1. Huan Teng, Yuhui Quan, Chengyu Wang 0001, Jun Huang 0007, Hui Ji 0002 |
CVPR | 2 |
| 2025 | Image debanding using cross-scale invertible networks with banded deformable convolutions
Yuhui Quan, Xuyi He, Ruotao Xu, Yong Xu 0007, Hui Ji 0002 |
Neural Networks | 1 |
| 2025 | Image shadow removal via multi-scale deep Retinex decomposition
Yan Huang 0031, Xinchang Lu, Yuhui Quan, Yong Xu 0007, Hui Ji 0002 |
Pattern Recognit. | 3 |
| 2025 | Model Extraction for Image Denoising Networks
Huan Teng, Yuhui Quan, Yong Xu 0007, Jun Huang 0007, Hui Ji 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Dual-Path Deep Unsupervised Learning for Multi-Focus Image FusionabstractMulti-focus image fusion (MFIF) aims at merging multiple images captured at different focal lengths to create an all-in-focus image. This paper introduces a fully unsupervised learning approach for MFIF that uses only pairs of defocused images for end-to-end training, bypassing the need for ground-truths in supervised learning. Unlike existing methods training via a similarity loss between fused and source images, we propose a dual-path learning framework comprising two networks: an image fuser and a mask predictor. The mask predictor is modeled as a self-supervised denoising network on imperfect fusion masks, trained with a masking-based unsupervised learning scheme. The image fuser, crafted with deep unrolling, leverages the output from the mask predictor to supervise its mask generation at each unrolled step. Moreover, we introduce a fusion consistency loss to ensure the alignment between the image fuser and the mask predictor. In extensive experiments, our proposed approach shows superiority over existing end-to-end unsupervised methods and competitive performance against the supervised ones. Yuhui Quan, Xi Wan, Yan Huang 0031, Hui Ji 0002 |
IEEE Trans. Multim. | 1 |
| 2025 | Enhancing CNN-Based Blind Image Quality Assessment via Deep Cross-Layer Pattern EncodingabstractEvaluating image quality without reference images, known as blind image quality assessment (BIQA), is crucial for image communication. Recently, convolutional neural networks (CNNs) have emerged as a prominent BIQA approach due to their feature learning power. Usually, both high-level semantic information and low-level details significantly impact perceived visual quality. However, most existing CNN-based methods focus on high-level semantic information via aggregating features on top of the last convolutional layer into a global descriptor, neglecting the importance of shallow, low-level cues. To address this limitation, this paper proposes a novel approach that exploits local encoding and histogram-based pyramid pooling on crosslayer features produced by a CNN, achieving a joint local and global analysis. Specifically, we introduce a cross-layer pattern encoding model that characterizes features generated along convolutional layers via a soft histogram of local 3D binary patterns. This leads to a highly informative yet compact descriptor for score regression. By building this module into a ResNet backbone, we present an effective BIQA model demonstrating state-ofthe-art performance in extensive experiments on synthetic and authentic datasets. Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Yun Liang 0003, Jing Li 0026, Patrick Le Callet |
IEEE Trans. Multim. | 3 |
| 2025 | Robust Unsupervised Deep Learning for Nonblind Image Deconvolution With Inaccurate KernelsabstractNonblind image deconvolution/deblurring aims at restoring sharp images from their noisy blurred versions using an associated blur kernel with potential inaccuracy. Current deep learning (DL) models of nonblind image deconvolution (NBID) predominantly reply on ground truth (GT) images for supervision, which restricts their applicability to certain real-world scenarios such as scientific imaging. This article proposes a fully unsupervised DL approach for NBID, utilizing a GT-free end-to-end training process that adeptly handles both measurement noise and kernel error. Specifically, in the absence of GT images, a self-reconstruction loss is proposed to handle measurement noise, by effectively emulating its supervised counterpart. Recognizing the likely occurrence of kernel error during both training and testing data, we introduce a self-ensemble loss function and an ensemble inference scheme, anchored by a phase-keeping kernel perturbation strategy. Furthermore, a shifting mechanism is integrated so as to the loss functions to resolve the shift ambiguity caused by kernel error. Extensive experiments show the superiority of our proposed approach over existing unsupervised NBID methods, as well as its competitive performance against some of the recent supervised methods. Xinran Qin, Yuhui Quan, Zhuojie Chen, Hui Ji 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Unsupervised Deep Unrolling Networks for Phase UnwrappingabstractPhase unwrapping (PU) is a technique to reconstruct original phase images from their noisy wrapped counterparts, finding many applications in scientific imaging. Although supervised learning has shown promise in PU, its utility is limited in ground-truth (GT) scarce scenarios. This paper presents an unsupervised learning approach that eliminates the need for GTs during end-to-end training. Our approach leverages the insight that both the gradients and wrapped gradients of wrapped phases serve as noisy labels for GT phase gradients, along with sparse outliers induced by the wrapping operation. A recorruption-based self-reconstruction loss in the gradient domain is proposed to mitigate the adverse effects of label noise, complemented with a self-distillation loss for improved generalization. Additionally, by unfolding a variational model of PU that utilizes wrapped gradients of wrapped phases for its data-fitting term, we develop a deep unrolling network that encodes physics of phase wrapping and incorporates special treatments on outliers. In the experiments on three types of phase data, our approach outperforms existing GT-free methods and competes well against the supervised ones. Zhile Chen, Yuhui Quan, Hui Ji 0002 |
CVPR | 2 |
| 2024 | Enhancing Underwater Images via Asymmetric Multi-Scale Invertible NetworksabstractUnderwater images, often plagued by complex degradation, pose significant challenges for image enhancement. To address these challenges, the paper redefines underwater image enhancement as an image decomposition problem and proposes a deep invertible neural network (INN) that accurately predicts both the latent image and the degradation effects. Instead of using an explicit formation model to describe the degradation process, the INN adheres to the constraints of the image decomposition model, providing necessary regularization for model training, particularly in the absence of supervision on degradation effects. Taking into account the diverse scales of degradation factors, the INN is structured on a multi-scale basis to effectively manage the varied scales of degradation factors. Moreover, the INN incorporates several asymmetric design elements that are specifically optimized for the decomposition model and the unique physics of underwater imaging. Comprehensive experiments show that our approach provides significant performance improvement over existing methods. Yuhui Quan, Xiaoheng Tan, Yan Huang 0031, Yong Xu 0007, Hui Ji 0002 |
ACM Multimedia | 1 |
| 2024 | Highly Efficient No-reference 4K Video Quality Assessment with Full-Pixel Covering Sampling and Training StrategyabstractDeep Video Quality Assessment (VQA) methods have shown impressive high-performance capabilities. Notably, no-reference (NR) VQA methods play a vital role in situations where obtaining reference videos is restricted or not feasible. Nevertheless, as more streaming videos are being created in ultra-high definition (e.g., 4K) to enrich viewers' experiences, the current deep VQA methods face unacceptable computational costs. Furthermore, the resizing, cropping, and local sampling techniques employed in these methods can compromise the details and content of original 4K videos, thereby negatively impacting quality assessment. In this paper, we propose a highly efficient and novel NR 4K VQA technology. Specifically, first, a novel data sampling and training strategy is proposed to tackle the problem of excessive resolution. This strategy allows the VQA Swin Transformer-based model to effectively train and make inferences using the full data of 4K videos on standard consumer-grade GPUs without compromising content or details. Second, a weighting and scoring scheme is developed to mimic the human subjective perception mode, which is achieved by considering the distinct impact of each sub-region within a 4K frame on the overall perception. Third, we incorporate the frequency domain information of video frames to better capture the details that affect video quality, consequently further improving the model's generalizability. To our knowledge, this is the first technology for the NR 4K VQA task. Thorough empirical studies demonstrate it not only significantly outperforms existing methods on a specialized 4K VQA dataset but also achieves state-of-the-art performance across multiple open-source NR video quality datasets. Xiaoheng Tan, Jiabin Zhang, Yuhui Quan, Jing Li 0026, Yajing Wu, Zilin Bian |
ACM Multimedia | 3 |
| 2024 | Pseudo-Siamese Blind-spot Transformers for Self-Supervised Real-World DenoisingabstractReal-world image denoising remains a challenge task. This paper studies self-supervised image denoising, requiring only noisy images captured in a single shot. We revamping the blind-spot technique by leveraging the transformer’s capability for long-range pixel interactions, which is crucial for effectively removing noise dependence in relating pixel–a requirement for achieving great performance for the blind-spot technique. The proposed method integrates these elements with two key innovations: a directional self-attention (DSA) module using a half-plane grid for self-attention, creating a sophisticated blind-spot structure, and a Siamese architecture with mutual learning to mitigate the performance impacts
from the restricted attention grid in DSA. Experiments on benchmark datasets demonstrate that our method outperforms existing self-supervised and clean-image-free methods. This combination of blind-spot and transformer techniques provides a natural synergy for tackling real-world image denoising challenges. Yuhui Quan, Hui Ji 0002 |
NeurIPS | 1 |
| 2024 | Cross-Scale Self-Supervised Blind Image Deblurring via Implicit Neural RepresentationabstractBlind image deblurring (BID) is an important yet challenging image recovery problem. Most existing deep learning methods require supervised training with ground truth (GT) images. This paper introduces a self-supervised method for BID that does not require GT images. The key challenge is to regularize the training to prevent over-fitting due to the absence of GT images. By leveraging an exact relationship among the blurred image, latent image, and blur kernel across consecutive scales, we propose an effective cross-scale consistency loss. This is implemented by representing the image and kernel with implicit neural representations (INRs), whose resolution-free property enables consistent yet efficient computation for network training across multiple scales. Combined with a progressively coarse-to-fine training scheme, the proposed method significantly outperforms existing self-supervised methods in extensive experiments. Tianjing Zhang, Yuhui Quan, Hui Ji 0002 |
NeurIPS | 2 |
| 2024 | Enhanced deep unrolling networks for snapshot compressive hyperspectral imaging
Xinran Qin, Yuhui Quan, Hui Ji 0002 |
Neural Networks | 2 |
| 2024 | 3D Snapshot: Invertible Embedding of 3D Neural Representations in a Single Imageabstract3D neural rendering enables photo-realistic reconstruction of a specific scene by encoding discontinuous inputs into a neural representation. Despite the remarkable rendering results, the storage of network parameters is not transmission-friendly and not extendable to metaverse applications. In this paper, we propose an invertible neural rendering approach that enables generating an interactive 3D model from a single image (i.e., 3D Snapshot). Our idea is to distill a pre-trained neural rendering model (e.g., NeRF) into a visualizable image form that can then be easily inverted back to a neural network. To this end, we first present a neural image distillation method to optimize three neural planes for representing the original neural rendering model. However, this representation is noisy and visually meaningless. We thus propose a dynamic invertible neural network to embed this noisy representation into a plausible image representation of the scene. We demonstrate promising reconstruction quality quantitatively and qualitatively, by comparing to the original neural rendering model, as well as video-based invertible methods. On the other hand, our method can store dozens of NeRFs with a compact restoration network (5 MB), and embedding each 3D scene takes up only 160 KB of storage. More importantly, our approach is the first solution that allows embedding a neural rendering model into image representations, which enables applications like creating an interactive 3D model from a printed image in the metaverse. Yuqin Lu, Bailin Deng, Zhixuan Zhong, Yuhui Quan, Hongmin Cai, Shengfeng He |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Siamese Cooperative Learning for Unsupervised Image Reconstruction From Incomplete MeasurementsabstractImage reconstruction from incomplete measurements is one basic task in imaging. While supervised deep learning has emerged as a powerful tool for image reconstruction in recent years, its applicability is limited by its prerequisite on a large number of latent images for model training. To extend the application of deep learning to the imaging tasks where acquisition of latent images is challenging, this article proposes an unsupervised deep learning method that trains a deep model for image reconstruction with the access limited to measurement data. We develop a Siamese network whose twin sub-networks perform reconstruction cooperatively on a pair of complementary spaces: the null space of the measurement matrix and the range space of its pseudo inverse. The Siamese network is trained by a self-supervised loss with three terms: a data consistency loss over available measurements in the range space, a data consistency loss between intermediate results in the null space, and a mutual consistency loss on the predictions of the twin sub-networks in the full space. The proposed method is applied to four imaging tasks from different applications, and extensive experiments have shown its advantages over existing unsupervised solutions. Yuhui Quan, Xinran Qin, Tongyao Pang, Hui Ji 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Deep Single Image Defocus Deblurring via Gaussian Kernel Mixture LearningabstractThis paper proposes an end-to-end deep learning approach for removing defocus blur from a single defocused image. Defocus blur is a common issue in digital photography that poses a challenge due to its spatially-varying and large blurring effect. The proposed approach addresses this challenge by employing a pixel-wise Gaussian kernel mixture (GKM) model to accurately yet compactly parameterize spatially-varying defocus point spread functions (PSFs), which is motivated by the isotropy in defocus PSFs. We further propose a grouped GKM (GGKM) model that decouples the coefficients in GKM, so as to improve the modeling accuracy with an economic manner. Afterward, a deep neural network called GGKMNet is then developed by unrolling a fixed-point iteration process of GGKM-based image deblurring, which avoids the efficiency issues in existing unrolling DNNs. Using a lightweight scale-recurrent architecture with a coarse-to-fine estimation scheme to predict the coefficients in GGKM, the GGKMNet can efficiently recover an all-in-focus image from a defocused one. Such advantages are demonstrated with extensive experiments on five benchmark datasets, where the GGKMNet outperforms existing defocus deblurring methods in restoration quality, as well as showing advantages in terms of model complexity and computational efficiency. Yuhui Quan, Zicong Wu, Ruotao Xu, Hui Ji 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Enhancing texture representation with deep tracing pattern encodingabstractTexture representation is a challenging problem due to the complex underlying physics of texture as well as the variations caused by changes in viewpoint. Recent progress in texture analysis has been made by the power of convolutional neural networks (CNNs) in feature learning . However, most current methods aggregate the features from the last convolutional layer of the CNN to obtain a global feature vector, which fails to leverage shallow low-level visual cues and cross-layer feature patterns, limiting their performance. In this paper, we propose to trace the features generated along the convolutional layers via a histogram of local 3D invariant binary patterns, called deep tracing patterns. This leads to a highly discriminative yet robust global feature representation module. Building such a module into a CNN backbone, we develop an effective approach for texture recognition. Extensive experiments on six benchmark datasets show that the proposed approach provides a discriminative and robust texture descriptor , with state-of-the-art performance achieved. Zhile Chen, Yuhui Quan, Ruotao Xu, Yong Xu 0007 |
Pattern Recognit. | 2 |
| 2024 | Image Smoothing via Multiscale Global PerceptionabstractImage smoothing provides a fundamental operation for image processing, with a broad spectrum of applications. It is a challenging task which requires global analysis on image patterns with scale awareness. Existing deep models for image smoothing are insufficiently efficient in global perception and multi-scale processing. This paper proposes a deep model with an efficient multi-scale fusion architecture and a series of global processing blocks. The architecture enhances multi-scale feature flow by incorporating features of different scales into both the encoder and decoder blocks of a U-shape network, with multi-scale feature fusion modules inserted between the encoder and the decoder. The global processing blocks leverage the multi-axis processing mechanism to achieve joint local and global perception. Benefiting from these two key designs, our proposed model enjoys superiority in both smoothing performance and computational complexity, as demonstrated in the experiments on two benchmark datasets. Xuyi He, Yuhui Quan, Yong Xu 0007, Ruotao Xu |
IEEE Signal Process. Lett. | 2 |
| 2023 | Unsupervised Deep Learning for Phase Retrieval via Teacher-Student DistillationabstractPhase retrieval (PR) is a challenging nonlinear inverse problem in scientific imaging that involves reconstructing the phase of a signal from its intensity measurements. Recently, there has been an increasing interest in deep learning-based PR. Motivated by the challenge of collecting ground-truth (GT) images in many domains, this paper proposes a fully-unsupervised learning approach for PR, which trains an end-to-end deep model via a GT-free teacher-student online distillation framework. Specifically, a teacher model is trained using a self-expressive loss with noise resistance, while a student model is trained with a consistency loss on augmented data to exploit the teacher's dark knowledge. Additionally, we develop an enhanced unfolding network for both the teacher and student models. Extensive experiments show that our proposed approach outperforms existing unsupervised PR methods with higher computational efficiency and performs competitively against supervised methods. Yuhui Quan, Zhile Chen, Tongyao Pang, Hui Ji 0002 |
AAAI | 1 |
| 2023 | Ground-Truth Free Meta-Learning for Deep Compressive SamplingabstractCompressive sampling (CS) is an efficient technique for imaging. This paper proposes a ground-truth (GT) free meta-learning method for CS, which leverages both ex-ternal and internal deep learning for unsupervised high-quality image reconstruction. The proposed method first trains a deep neural network (NN) via external meta-learning using only CS measurements, and then efficiently adapts the trained model to a test sample for exploiting sample-specific internal characteristic for performance gain. The meta-learning and model adaptation are built on an improved Stein's unbiased risk estimator (iSURE) that provides efficient computation and effective guidance for accurate prediction in the range space of the adjoint of the measurement matrix. To improve the learning and adaption on the null space of the measurement matrix, a modi-fied model-agnostic meta-learning scheme and a null-space consistency loss are proposed. In addition, a bias tuning scheme for unrolling NNs is introduced for further acceler-ation of model adaption. Experimental results have demonstrated that the proposed GT-free method performs well and can even compete with supervised methods. Xinran Qin, Yuhui Quan, Tongyao Pang, Hui Ji 0002 |
CVPR | 2 |
| 2023 | Neumann Network with Recursive Kernels for Single Image Defocus DeblurringabstractSingle image defocus deblurring (SIDD) refers to recovering an all-in-focus image from a defocused blurry one. It is a challenging recovery task due to the spatially-varying defocus blurring effects with significant size variation. Motivated by the strong correlation among defocus kernels of different sizes and the blob-type structure of defocus kernels, we propose a learnable recursive kernel representation (RKR) for defocus kernels that expresses a defocus kernel by a linear combination of recursive, separable and positive atom kernels, leading to a compact yet effective and physics-encoded parametrization of the spatially-varying defocus blurring process. Afterwards, a physics-driven and efficient deep model with a cross-scale fusion structure is presented for SIDD, with inspirations from the truncated Neumann series for approximating the matrix inversion of the RKR-based blurring operator. In addition, a reblurring loss is proposed to regularize the RKR learning. Extensive experiments show that, our proposed approach significantly outperforms existing ones, with a model size comparable to that of the top methods. Yuhui Quan, Zicong Wu, Hui Ji 0002 |
CVPR | 1 |
| 2023 | Diffuse3D: Wide-Angle 3D Photography via Bilateral DiffusionabstractThis paper aims to resolve the challenging problem of wide-angle novel view synthesis from a single image, a.k.a. wide-angle 3D photography. Existing approaches rely on local context and treat them equally to inpaint occluded RGB and depth regions, which fail to deal with large-region occlusion (i.e., observing from an extreme angle) and foreground layers might blend into background inpainting. To address the above issues, we propose Diffuse3D which employs a pre-trained diffusion model for global synthesis, while amending the model to activate depth-aware inference. Our key insight is to alter the convolution mechanism in the denoising process. We inject depth information into the denoising convolution operation with bilateral kernels, i.e., a depth kernel and a spatial kernel, to consider layered correlations among pixels. In this way, foreground regions are overlooked in background inpainting and only pixels close in depth are leveraged. On the other hand, we propose a global-local balancing approach to maximize both contextual understandings. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in novel view synthesis, especially in wide-angle scenarios. More importantly, our method does not require any training and is a plug-and-play module that can be integrated with any diffusion model. Our code can be found at https://github.com/yutaojiang1/Diffuse3D. Yutao Jiang, Yang Zhou 0038, Wenxi Liu, Jianbo Jiao, Yuhui Quan, Shengfeng He |
ICCV | 6 |
| 2023 | Deep Video Demoiréing via Compact Invertible Dyadic DecompositionabstractRemoving moiré patterns from videos recorded on screens or complex textures is known as video demoiréing. It is a challenging task as both structures and textures of an image usually exhibit strong periodic patterns, which thus are easily confused with moiré patterns and can be significantly erased in the removal process. By interpreting video demoiréing as a multi-frame decomposition problem, we propose a compact invertible dyadic network called CIDNet that progressively decouples latent frames and the moiré patterns from an input video sequence. Using a dyadic cross-scale coupling structure with coupling layers tailored for multi-scale processing, CIDNet aims at disentangling the features of image patterns from that of moiré patterns at different scales, while retaining all latent image features to facilitate reconstruction. In addition, a compressed form for the network’s output is introduced to reduce computational complexity and alleviate overfitting. The experiments show that CIDNet outperforms existing methods and enjoys the advantages in model size and computational efficiency. Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu |
ICCV | 1 |
| 2023 | Fingerprinting Deep Image Restoration ModelsabstractFingerprinting is a promising non-invasive method for protecting the intellectual property rights (IPR) of deep neural network (DNN) models. It extracts a feature called a fingerprint from a DNN model to identify its ownership. Existing fingerprinting methods focus only on classification-related models that map images to labels, while inapplicable to models for image restoration that map images to images. This paper proposes a fingerprinting framework for DNN models of image restoration. The proposed framework defines the fingerprint using a critical image, which exhibits strongly discriminative patterns and is robust to modest model modifications. Model ownership is then verified by comparing the distance of color histograms and local gradient pattern histograms of critical images between the suspect and source models. We apply the proposed framework to two representative tasks, denoising and super-resolution. It outperforms the baselines of fingerprinting and competes against existing invasive model watermarking methods. Yuhui Quan, Huan Teng, Ruotao Xu, Jun Huang 0007, Hui Ji 0002 |
ICCV | 1 |
| 2023 | Single Image Defocus Deblurring via Implicit Neural Inverse KernelsabstractSingle image defocus deblurring (SIDD) is a challenging task due to the spatially-varying nature of defocus blur, characterized by per-pixel point spread functions (PSFs). Existing deep-learning-based methods for SIDD are limited by either over-fitting due to the lack of model constraints or under-parametrization that restricts their applicability to real-world images. To address the limitations, this paper proposes an interpretable approach that explicitly predicts inverse kernels with structural regularization. Motivated by the observation that defocus PSFs within an image often have similar shapes but different sizes, we represent the inverse kernels linearly over a multi-scale dictionary parameterized by implicit neural representations. We predict the corresponding representation coefficients via a duplex scale-recurrent neural network that jointly performs fine-to-coarse and coarse-to-fine estimations. Extensive experiments demonstrate that our approach achieves excellent performance using a lightweight model. Yuhui Quan, Hui Ji 0002 |
ICCV | 1 |
| 2023 | Video Noise Removal Using Progressive Decomposition With Conditional InvertibilityabstractVideo denoising aims at removing noise from noisy video frames and meanwhile preserving their structures and details. It is a challenging task, as both noise and video structures/details correspond to high-frequency components of a noisy video which are hard to distinguish. This paper proposes a deep video denoiser using a progressive decomposition process with conditional invertibility. Noisy video frames are first decomposed into two latent codes via a forward process of conditional invertible coupling layers, where one latent code carries the maximal information regarding the noise-free reference frame while the other encodes the information regarding noise, misalignment and content difference. The clean video is then reconstructed from the latent codes of noise-free frames using the reverse pass of the coupling layers. To improve the robustness to variant noise levels, the coupling layers are conditioned on noise level. In addition, memory units are introduced to the conditioned coupling layers to better exploit temporal correlation among frames for feature disentanglement. Experiments on two benchmark datasets have demonstrated the effectiveness of our method. Haoran Huang, Yuhui Quan, Zhenghua Lei, Jinlong Hu 0002, Yan Huang 0031 |
ICME | 2 |
| 2023 | Self-Supervised Blind Image Deconvolution via Deep Generative Ensemble LearningabstractBlind image deconvolution (BID) is about recovering a latent image with sharp details from its blurred observation generated by the convolution with an unknown smoothing kernel. Recently, deep generative priors from untrained neural networks (NNs) have emerged as a promising deep learning approach for BID, with the benefit of being free of external training samples. However, existing untrained-NN-based BID methods may suffer from under-deblurring or overfitting. In this paper, we propose an ensemble approach to better exploit the priors from untrained NNs for BID, which aggregates the deblurring results of multiple untrained NNs for improvement. To enjoy both the effectiveness and computational efficiency in ensemble learning, the untrained NNs are designed with a specific shared-base and multi-head architecture. In addition, a kernel-centering layer is proposed for handling the shift ambiguity among different predictions during ensemble, which also improves the robustness of kernel prediction to the setting of the kernel size parameter. Extensive experiments show that the proposed approach noticeably outperforms both exiting dataset-free methods and dataset-based methods. Mingqin Chen, Yuhui Quan, Yong Xu 0007, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Image Desnowing via Deep Invertible SeparationabstractImages taken on snowy days often suffer from severe negative visual effects caused by snowflakes. The task of removing snowflakes from a snowy image is known as image desnowing, which is challenging as image details are easily mistakenly treated and thus may be significantly lost during snowflake removal. Leveraging invertible neural networks (INNs), this paper presents a deep learning-based method for single image desnowing, which can remove snowflakes accurately while preserving image details well. Interpreting desnowing as an image decomposition problem, we propose an INN composed of two asymmetric interactive paths for predicting a latent image and a snowflake layer respectively. Such an INN is able to progressively refine the features of both latent images and snowflake layers for disentanglement, while retaining all information possibly relevant to latent image reconstruction. In addition, an attentive coupling layer supervised by snowflake masks is introduced to enhance feature dismantlement and a coupling-in-coupling structure is developed for further improvement. Extensive experiments show that, the proposed method outperforms existing ones on three benchmark datasets of synthetic and real-world images, and meanwhile it also shows advantages in terms of model size and computational efficiency. Yuhui Quan, Xiaoheng Tan, Yan Huang 0031, Yong Xu 0007, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Learning Deep Non-blind Image Deconvolution Without Ground Truths
Yuhui Quan, Zhuojie Chen, Hui Ji 0002 |
ECCV (6) | 1 |
| 2022 | Dual-Domain Self-supervised Learning and Model Adaption for Deep Compressive Imaging
Yuhui Quan, Xinran Qin, Tongyao Pang, Hui Ji 0002 |
ECCV (30) | 1 |
| 2022 | High-Quality Self-Supervised Snapshot Hyperspectral ImagingabstractHyperspectral image (HSI) reconstruction is about recovering a 3D HSI from its 2D snapshot measurements, to which deep models have become a promising approach. However, most existing studies train deep models on large amounts of organized data, the collection of which can be difficult in many applications. This paper leverages the image priors encoded in untrained neural networks (NNs) to have a self-supervised learning method which is free from training datasets while adaptive to the statistics of a test sample. To induce better image priors and prevent the NN overfitting undesired solutions, we construct an unrolling-based NN equipped with fractional max pooling (FMP). Furthermore, the FMP is used with randomness to enable self-ensemble for reconstruction accuracy improvement. In the experiments, our self-supervised learning approach enjoys high-quality reconstruction and outperforms recent methods including the supervised ones. Yuhui Quan, Xinran Qin, Mingqin Chen, Yan Huang 0031 |
ICASSP | 1 |
| 2022 | Deep Blind Image Quality Assessment Using Dual-Order StatisticsabstractDeep convolutional neural networks (CNNs) have become a promising approach to blind image quality assessment (BIQA). Existing CNN-based BIQA methods often employ global average pooling (GAP) to aggregate feature maps into a fixed-size representation for regression, so as to handle input images with varying sizes. However, GAP is only capable of extracting the first-order statistics of feature distributions, which is ineffective for distinguishing complex distortions that cause local degradation or preserve global features. To tackle this problem, we introduce the second-order global covariance pooling (GCP) for aggregating feature maps, leading to a more distortion-sensitive and more discriminative global representation. By incorporating GCP and GAP into a ResNet backbone, we propose an effective deep model for BIQA. The experimental results on five BIQA benchmark datasets, including both the synthetic and authentic ones, have demon-strated the excellent performance of the proposed method. Zihan Zhou 0007, Yong Xu 0007, Yuhui Quan, Ruotao Xu |
ICME | 3 |
| 2022 | No-Reference Image Quality Assessment Using Dynamic Complex-Valued Neural ModelabstractDeep convolutional neural networks (CNNs) have become a promising approach to no-reference image quality assessment (NR-IQA). This paper aims at improving the power of CNNs for NR-IQA in two aspects. Firstly, motivated by the deep connection between complex-valued transforms and human visual perception, we introduce complex-valued convolutions and phase-aware activations beyond traditional real-valued CNNs, which improves the accuracy of NR-IQA without bringing noticeable additional computational costs. Secondly, considering the content-awareness of visual quality perception, we include a dynamic filtering module for better extracting content-aware features, which predicts features based on both local content and global semantics. These two improvements lead to a complex-valued content-aware neural NR-IQA model with good generalization. Extensive experiments on both synthetically and authentically distorted data have demonstrated the state-of-the-art performance of the proposed approach. Zihan Zhou 0007, Yong Xu 0007, Ruotao Xu, Yuhui Quan |
ACM Multimedia | 4 |
| 2022 | Nonblind Image Deconvolution via Leveraging Model Uncertainty in An Untrained Deep Neural Network
Mingqin Chen, Yuhui Quan, Tongyao Pang, Hui Ji 0002 |
Int. J. Comput. Vis. | 2 |
| 2022 | Unsupervised knowledge transfer for nonblind image deconvolution
Zhuojie Chen, Yong Xu 0007, Junle Wang, Yuhui Quan |
Pattern Recognit. Lett. | 5 |
| 2022 | Self-Supervised Low-Light Image Enhancement Using Discrepant Untrained Network PriorsabstractThis paper proposes a deep learning method for low-light image enhancement, which exploits the generation capability of Neural Networks (NNs) while requiring no training samples except the input image itself. Based on the Retinex decomposition model, the reflectance and illumination of a low-light image are parameterized by two untrained NNs. The ambiguity between the two layers is resolved by the discrepancy between the two NNs in terms of architecture and capacity, while the complex noise with spatially-varying characteristics is handled by an illumination-adaptive self-supervised denoising module. The enhancement is done by jointly optimizing the Retinex decomposition and the illumination adjustment. Extensive experiments show that the proposed method not only outperforms existing non-learning-based and unsupervised-learning-based methods, but also competes favorably with some supervised-learning-based methods in extreme low-light conditions. Jinxiu Liang, Yong Xu 0007, Yuhui Quan, Boxin Shi, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Unsupervised Deep Background Matting Using Deep Matte PriorabstractBackground matting is a recently developed image matting approach, with applications to image and video editing. It refers to estimating both the alpha matte and foreground from a pair of images with and without foreground objects. Recent work has applied deep learning to background matting, with very promising performance achieved. However, existing deep models are supervised which require a large dataset with ground truth alpha mattes for training. To avoid the cost of data collection and possible bias in training data, this paper proposes a dataset-free unsupervised deep learning-based approach for background matting. Observing that the local smoothness of alpha matte can be well characterized by the untrained network prior called deep matte prior, we model the foreground and alpha matte using the priors encoded by two generative convolutional neural networks. To avoid possible overfitting during unsupervised learning, a two-stage learning scheme is developed which contains projection-based training and Bayesian post refinement. An alpha-matte-driven initialization scheme is also developed for performance boost. Even without calling external training data, the proposed approach provides competitive performance to recent supervised learning-based methods in the experiments. Yong Xu 0007, Baoling Liu, Yuhui Quan, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Recurrent Exposure Generation for Low-Light Face DetectionabstractFace detection from low-light images is challenging due to limited photons and inevitable noise, which, to make the task even harder, are often spatially unevenly distributed. A natural solution is to borrow the idea frommulti-exposure, which captures multiple shots to obtain well-exposed images under challenging conditions. High-quality implementation/approximation of multi-exposure from a single image is however nontrivial. Fortunately, as shown in this paper, neither is such high-quality necessary since our task isface detectionrather thanimage enhancement. Specifically, we propose a novelRecurrent Exposure Generation (REG)module and couple it seamlessly with aMulti-Exposure Detection (MED)module, and thus significantly improve face detection performance by effectively inhibiting non-uniform illumination and noise issues. REG produces progressively and efficiently intermediate images corresponding to various exposure settings, and such pseudo-exposures are then fused by MED to detect faces across different lighting conditions. The proposed method, namedREGDet, is the first ‘detection-with-enhancement’ framework for low-light face detection. It not only encourages rich interaction and feature fusion across different illumination levels, but also enables effective end-to-end learning of the REG component to be better tailored for face detection. Moreover, as clearly shown in our experiments, REG can be flexibly coupled with different face detectors without extra low/normal-light image pairs for training. We tested REGDet on the DARK FACE low-light face benchmark with thorough ablation study, where REGDet outperforms previous state-of-the-arts by a significant margin, with only negligible extra parameters. Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Jiaying Liu 0001, Haibin Ling, Yong Xu 0007 |
IEEE Trans. Multim. | 3 |
| 2022 | Nonblind Image Deblurring via Deep Learning in Complex FieldabstractNonblind image deblurring is about recovering the latent clear image from a blurry one generated by a known blur kernel, which is an often-seen yet challenging inverse problem in imaging. Its key is how to robustly suppress noise magnification during the inversion process. Recent approaches made a breakthrough by exploiting convolutional neural network (CNN)-based denoising priors in the image domain or the gradient domain, which allows using a CNN for noise suppression. The performance of these approaches is highly dependent on the effectiveness of the denoising CNN in removing magnified noise whose distribution is unknown and varies at different iterations of the deblurring process for different images. In this article, we introduce a CNN-based image prior defined in the Gabor domain. The prior not only utilizes the optimal space-frequency resolution and strong orientation selectivity of the Gabor transform but also enables using complex-valued (CV) representations in intermediate processing for better denoising. A CV CNN is developed to exploit the benefits of the CV representations, with better generalization to handle unknown noises over the real-valued ones. Combining our Gabor-domain CV CNN-based prior with an unrolling scheme, we propose a deep-learning-based approach to nonblind image deblurring. Extensive experiments have demonstrated the superior performance of the proposed approach over the state-of-the-art ones. Yuhui Quan, Peikang Lin, Yong Xu 0007, Yuesong Nan, Hui Ji 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Deep Texture Recognition via Exploiting Cross-Layer Statistical Self-SimilarityabstractIn recent years, convolutional neural networks (CNNs) have become a prominent tool for texture recognition. The key of existing CNN-based approaches is aggregating the convolutional features into a robust yet discriminative description. This paper presents a novel feature aggregation module called CLASS (Cross-Layer Aggregation of Statistical Self-similarity) for texture recognition. We model the CNN feature maps across different layers, as a dynamic process which carries the statistical self-similarity (SSS), one well-known property of texture, from input image along the network depth dimension. The CLASS module characterizes the cross-layer SSS using a soft histogram of local differential box-counting dimensions of cross-layer features. The resulting descriptor encodes both cross-layer dynamics and local SSS of input image, providing additional discrimination over the often-used global average pooling. Integrating CLASS into a ResNet backbone, we develop CLASSNet, an effective deep model for texture recognition, which shows state-of-the-art performance in the experiments. Zhile Chen, Yuhui Quan, Yong Xu 0007, Hui Ji 0002 |
CVPR | 3 |
| 2021 | Recorrupted-to-Recorrupted: Unsupervised Deep Learning for Image DenoisingabstractDeep denoiser, the deep network for denoising, has been the focus of the recent development on image denoising. In the last few years, there is an increasing interest in developing unsupervised deep denoisers which only call unorganized noisy images without ground truth for training. Nevertheless, the performance of these unsupervised deep denoisers is not competitive to their supervised counterparts. Aiming at developing a more powerful unsupervised deep denoiser, this paper proposed a data augmentation technique, called recorrupted-to-recorrupted (R2R), to address the overfitting caused by the absence of truth images. For each noisy image, we showed that the cost function defined on the noisy/noisy image pairs constructed by the R2R method is statistically equivalent to its supervised counterpart defined on the noisy/truth image pairs. Extensive experiments showed that the proposed R2R method noticeably outperformed existing unsupervised deep denoisers, and is competitive to representative supervised deep denoisers. Tongyao Pang, Yuhui Quan, Hui Ji 0002 |
CVPR | 3 |
| 2021 | Gaussian Kernel Mixture Network for Single Image Defocus DeblurringabstractDefocus blur is one kind of blur effects often seen in images, which is challenging to remove due to its spatially variant amount. This paper presents an end-to-end deep learning approach for removing defocus blur from a single image, so as to have an all-in-focus image for consequent vision tasks. First, a pixel-wise Gaussian kernel mixture (GKM) model is proposed for representing spatially variant defocus blur kernels in an efficient linear parametric form, with higher accuracy than existing models. Then, a deep neural network called GKMNet is developed by unrolling a fixed-point iteration of the GKM-based deblurring. The GKMNet is built on a lightweight scale-recurrent architecture, with a scale-recurrent attention module for estimating the mixing coefficients in GKM for defocus deblurring. Extensive experiments show that the GKMNet not only noticeably outperforms existing defocus deblurring methods, but also has its advantages in terms of model complexity and computational efficiency. Yuhui Quan, Zicong Wu, Hui Ji 0002 |
NeurIPS | 1 |
| 2021 | Encoding Spatial Distribution of Convolutional Features for Texture RepresentationabstractExisting convolutional neural networks (CNNs) often use global average pooling (GAP) to aggregate feature maps into a single representation. However, GAP cannot well characterize complex distributive patterns of spatial features while such patterns play an important role in texture-oriented applications, e.g., material recognition and ground terrain classification. In the context of texture representation, this paper addressed the issue by proposing Fractal Encoding (FE), a feature encoding module grounded by multi-fractal geometry. Considering a CNN feature map as a union of level sets of points lying in the 2D space, FE characterizes their spatial layout via a local-global hierarchical fractal analysis which examines the multi-scale power behavior on each level set. This enables a CNN to encode the regularity on the spatial arrangement of image features, leading to a robust yet discriminative spectrum descriptor. In addition, FE has trainable parameters for data adaptivity and can be easily incorporated into existing CNNs for end-to-end training. We applied FE to ResNet-based texture classification and retrieval, and demonstrated its effectiveness on several benchmark datasets. Yong Xu 0007, Zhile Chen, Jinxiu Liang, Yuhui Quan |
NeurIPS | 5 |
| 2021 | Attentive deep network for blind motion deblurring on dynamic scenes
Yong Xu 0007, Ye Zhu 0003, Yuhui Quan, Hui Ji 0002 |
Comput. Vis. Image Underst. | 3 |
| 2021 | Image denoising using complex-valued deep CNN
Yuhui Quan, Yizhen Shao, Huan Teng, Yong Xu 0007, Hui Ji 0002 |
Pattern Recognit. | 1 |
| 2021 | Structure-Texture Image Decomposition Using Discriminative Patch RecurrenceabstractMorphology component analysis provides an effective framework for structure-texture image decomposition, which characterizes the structure and texture components by sparsifying them with certain transforms respectively. Due to the complexity and randomness of texture, it is challenging to design effective sparsifying transforms for texture components. This paper aims at exploiting the recurrence of texture patterns, one important property of texture, to develop a nonlocal transform for texture component sparsification. Since the plain patch recurrence holds for both cartoon contours and texture regions, the nonlocal sparsifying transform constructed based on such patch recurrence sparsifies both the structure and texture components well. As a result, cartoon contours could be wrongly assigned to the texture component, yielding ambiguity in decomposition. To address this issue, we introduce a discriminative prior on patch recurrence, that the spatial arrangement of recurrent patches in texture regions exhibits isotropic structure which differs from that of cartoon contours. Based on the prior, a nonlocal transform is constructed which only sparsifies texture regions well. Incorporating the constructed transform into morphology component analysis, we propose an effective approach for structure-texture decomposition. Extensive experiments have demonstrated the superior performance of our approach over existing ones. Ruotao Xu, Yong Xu 0007, Yuhui Quan |
IEEE Trans. Image Process. | 3 |
| 2021 | Multi-View 3D Shape Recognition via Correspondence-Aware Deep LearningabstractIn recent years, multi-view learning has emerged as a promising approach for 3D shape recognition, which identifies a 3D shape based on its 2D views taken from different viewpoints. Usually, the correspondences inside a view or across different views encode the spatial arrangement of object parts and the symmetry of the object, which provide useful geometric cues for recognition. However, such view correspondences have not been explicitly and fully exploited in existing work. In this paper, we propose a correspondence-aware representation (CAR) module, which explicitly finds potential intra-view correspondences and cross-view correspondences via k NN search in semantic space and then aggregates the shape features from the correspondences via learned transforms. Particularly, the spatial relations of correspondences in terms of their viewpoint positions and intra-view locations are taken into account for learning correspondence-aware features. Incorporating the CAR module into a ResNet-18 backbone, we propose an effective deep model called CAR-Net for 3D shape classification and retrieval. Extensive experiments have demonstrated the effectiveness of the CAR module as well as the excellent performance of the CAR-Net. Yong Xu 0007, Chaoda Zheng, Ruotao Xu, Yuhui Quan, Haibin Ling |
IEEE Trans. Image Process. | 4 |
| 2021 | Factorized Tensor Dictionary Learning for Visual Tensor Data CompletionabstractThis paper aims at developing a dictionary-learning-based method for completing the visual tensor data with missing elements. Traditional dictionary learning approaches suffer from very high computational costs when processing high-dimensional tensor data. Some existing approaches for acceleration impose orthogonality constraints or rank-one decompositions on dictionary atoms; however, the expressibility of the resulting dictionary is rather limited. To address such issues, we propose a convolutional analysis model for tensor dictionary learning, where the update of sparse coefficients during dictionary learning is simple and fast. Furthermore, we propose an orthogonality-constrained convolutional factorization scheme for dictionary construction, in which each tensor dictionary atom is factorized by the convolution of two atoms selected from two orthogonal factor dictionaries respectively. This factorization scheme enables us to efficiently learn an expressive dictionary with over-completeness and non-rank-one atoms. Based on our convolutional analysis model and factorization scheme, an effective yet efficient dictionary learning method is proposed for visual tensor completion. Extensive experiments show that, our method not only outperforms existing dictionary-based approaches with relatively-low time cost, but also outperforms recent low-rank approaches. Ruotao Xu, Yong Xu 0007, Yuhui Quan |
IEEE Trans. Multim. | 3 |
| 2021 | Image Quality Assessment Using Kernel Sparse CodingabstractOne key in image quality assessment (IQA) is the design of image representations that can capture the changes of image structures caused by distortions. Recent studies show that sparse coding has emerged as a promising approach to analyzing image structures for IQA. However, existing sparse-coding-based IQA approaches use linear coding models, which ignore the nonlinearities of manifolds of image patches and thus cannot analyze complex image structures well. To overcome such a weakness, in this paper, we introduce nonlinear sparse coding to IQA. A kernel dictionary construction scheme is proposed, which combines analytic dictionaries and learnable dictionaries to guarantee both the stability and effectiveness of kernel sparse coding in the context of IQA. Built upon the kernel dictionary construction, an effective full-reference IQA metric is developed. Benefiting from the considerations on nonlinearities during sparse coding, the proposed IQA metric not only characterizes image distortions better, but also achieves improvement on the consistency with subjective perception, when compared to the metrics built upon linear sparse coding. Such benefits are demonstrated with the experimental results on eight benchmark datasets in terms of common criteria. Zihan Zhou 0007, Jing Li 0026, Yuhui Quan, Ruotao Xu |
IEEE Trans. Multim. | 3 |
| 2021 | Watermarking Deep Neural Networks in Image ProcessingabstractPublishing/sharing pretrained deep neural network (DNN) models is a common practice in the community of computer vision. The increasing popularity of pretrained models has made it a serious concern: how to protect the intellectual properties of model owners and avert illegal usages by malicious attackers. This article aims at developing a framework for watermarking DNNs, with a particular focus on low-level image processing tasks that map images to images. Using image denoising and superresolution as case studies, we develop a black-box watermarking method for pretrained models, which exploits the overparameterization of the DNNs in image processing. In addition, an auxiliary module for visualizing the watermark information is proposed for further verification. Extensive experiments show that the proposed watermarking framework has no noticeable impact on model performance and enjoys the robustness against the often-seen attacks. Yuhui Quan, Huan Teng, Hui Ji 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Variational-EM-Based Deep Learning for Noise-Blind Image DeblurringabstractNon-blind deblurring is an important problem encountered in many image restoration tasks. The focus of non-blind deblurring is on how to suppress noise magnification during deblurring. In practice, it often happens that the noise level of input image is unknown and varies among different images. This paper aims at developing a deep learning framework for deblurring images with unknown noise level. Based on the framework of variational expectation maximization (EM), an iterative noise-blind deblurring scheme is proposed which integrates the estimation of noise level and the quantification of image prior uncertainty. Then, the proposed scheme is unrolled to a neural network (NN) where image prior is modeled by NN with uncertainty quantification. Extensive experiments showed that the proposed method not only outperformed existing noise-blind deblurring methods by a large margin, but also outperformed those state-of-the-art image deblurring methods designed/trained with known noise level. Yuesong Nan, Yuhui Quan, Hui Ji 0002 |
CVPR | 2 |
| 2020 | Self2Self With Dropout: Learning Self-Supervised Denoising From Single ImageabstractIn last few years, supervised deep learning has emerged as one powerful tool for image denoising, which trains a denoising network over an external dataset of noisy/clean image pairs. However, the requirement on a high-quality training dataset limits the broad applicability of the denoising networks. Recently, there have been a few works that allow training a denoising network on the set of external noisy images only. Taking one step further, this paper proposes a self-supervised learning method which only uses the input noisy image itself for training. In the proposed method, the network is trained with dropout on the pairs of Bernoulli-sampled instances of the input image, and the result is estimated by averaging the predictions generated from multiple instances of the trained model with dropout. The experiments show that the proposed method not only significantly outperforms existing single-image learning or non-learning methods, but also is competitive to the denoising networks trained on external datasets. Yuhui Quan, Mingqin Chen, Tongyao Pang, Hui Ji 0002 |
CVPR | 1 |
| 2020 | Self-supervised Bayesian Deep Learning for Image Recovery with Applications to Compressive Sensing
Tongyao Pang, Yuhui Quan, Hui Ji 0002 |
ECCV (11) | 2 |
| 2020 | Full-reference image quality metric for blurry images and compressed images using hybrid dictionary learning
Zihan Zhou 0007, Jing Li 0026, Yong Xu 0007, Yuhui Quan |
Neural Comput. Appl. | 4 |
| 2020 | Cartoon-Texture Image Decomposition using Orientation Characteristics in Patch RecurrenceabstractCartoon-texture image decomposition is about decomposing an image into the linear sum of two layers: cartoon and texture, where the key challenge is how to resolve the ambiguity between two layers. It is observed that the recurrence of texture patches occurs along multiple orientations, and the recurrence of cartoon patches only occurs along certain orientations. This paper proposes to separate these two layers by exploiting their orientation characteristics of image patch recurrence, i.e., isotropy property of texture patch recurrence versus anisotropy property of cartoon patch recurrence. Together with the sparsity-based regularizations in the image domain, a variational method is then developed in this paper for cartoon-texture decomposition. The experiments show that the proposed method noticeably outperforms many well-established ones on test images. Ruotao Xu, Yong Xu 0007, Yuhui Quan, Hui Ji 0002 |
SIAM J. Imaging Sci. | 3 |
| 2020 | Weakly-Supervised Sparse Coding With Geometric Prior for Interactive Texture SegmentationabstractTexture segmentation is about dividing a texture-dominant image into multiple homogeneous texture regions. The existing unsupervised approaches for texture segmentation are annotation-free but often yield unsatisfactory results. In contrast, supervised approaches such as deep learning may have better performance but require a large amount of annotated data. In this letter, we propose a user-interactive approach to win the trade-off between unsupervised approaches and supervised deep approaches. Our approach requires the user to mark one pixel in each texture region, whose label is directly propagated to its neighbor region. Such labeled data are of very small amount and even partially erroneous. To effectively exploit such weakly-labeled data, we construct a weakly-supervised sparse coding model that jointly conducts feature learning and segmentation. In addition, the geometric constraints are developed for the model to exploit the geometric prior on the local connectivity of region boundaries. The experiments on two benchmark datasets have validated the effectiveness of the proposed approach. Yuhui Quan, Huan Teng, Yan Huang 0031 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Image Denoising via Sequential Ensemble LearningabstractImage denoising is about removing measurement noise from input image for better signal-to-noise ratio. In recent years, there has been great progress on the development of data-driven approaches for image denoising, which introduce various techniques and paradigms from machine learning in the design of image denoisers. This paper aims at investigating the application of ensemble learning in image denoising, which combines a set of simple base denoisers to form a more effective image denoiser. Based on different types of image priors, two types of base denoisers in the form of transform-shrinkage are proposed for constructing the ensemble. Then, with an effective re-sampling scheme, several ensemble-learning-based image denoisers are constructed using different sequential combinations of multiple proposed base denoisers. The experiments showed that sequential ensemble learning can effectively boost the performance of image denoising. Xuhui Yang, Yong Xu 0007, Yuhui Quan, Hui Ji 0002 |
IEEE Trans. Image Process. | 3 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 32 |
| 2019 | Deep Learning for Seeing Through Window With RaindropsabstractWhen taking pictures through glass window in rainy day, the images are comprised and corrupted by the raindrops adhered to glass surfaces. It is a challenging problem to remove the effect of raindrops from an image. The key task is how to accurately and robustly identify the raindrop regions in an image. This paper develops a convolutional neural network (CNN) for removing the effect of raindrops from an image. In the proposed CNN, we introduce a double attention mechanism that concurrently guides the CNN using shape-driven attention and channel re-calibration. The shape-driven attention exploits physical shape priors of raindrops, i.e. convexness and contour closedness, to accurately locate raindrops, and the channel re-calibration improves the robustness when processing raindrops with varying appearances. The experimental results show that the proposed CNN outperforms the state-of-the-art approaches in terms of both quantitative metrics and visual quality. Yuhui Quan, Hui Ji 0002 |
ICCV | 1 |
| 2019 | Multi-view Rank Pooling for 3D Object Recognition*abstract3D shape recognition via deep learning is drawing more and more attention due to huge industry interests. As 3D deep learning methods emerged, the view-based approaches have gained considerable success in object classification. Most of these methods focus on designing a pooling scheme to aggregate CNN features of multi-view images into a single compact one. However, these view-wise pooling techniques suffer from loss of visual information. To deal with this issue, an adaptive rank pooling layer is introduced in this paper. Unlike max-pooling which only considers the maximum or mean-pooling that treats each element indiscriminately, the proposed pooling layer takes all the elements into account and dynamically adjusts their importances during the training. Experiments conducted on ModelNet40 and ModelNet10 shows both efficiency and accuracy gain when inserting such a layer into a baseline CNN architecture. Chaoda Zheng, Yong Xu 0007, Ruotao Xu, Hongyu Chi, Yuhui Quan |
VCIP | 5 |
| 2019 | Attention with structure regularization for action recognition
Yuhui Quan, Ruotao Xu, Hui Ji 0002 |
Comput. Vis. Image Underst. | 1 |
| 2019 | Exploiting label consistency in structured sparse representation for classification
Yan Huang 0031, Yuhui Quan, Yong Xu 0007 |
Neural Comput. Appl. | 2 |
| 2019 | Barzilai-Borwein-based adaptive learning rate for deep learning
Jinxiu Liang, Yong Xu 0007, Chenglong Bao, Yuhui Quan, Hui Ji 0002 |
Pattern Recognit. Lett. | 4 |
| 2019 | Supervised Sparse Coding With Decision ForestabstractBy jointly conducting sparse coding and classifier training, supervised sparse coding has shown its effectiveness in a variety of recognition tasks. However, the existing supervised sparse coding methods often consider linear classification, which limits their discrimination in handling highly nonlinear data. In this letter, we propose a new supervised sparse coding model by incorporating decision tree classifiers. Since decision trees can well deal with the non-linear properties of data, the introduction of decision trees to sparse coding can noticeably improve the discrimination of coding. Meanwhile, sparse coding is able to produce sparse de-correlated features that decision tree is in favor of. For further improvement, we close the loop of sparse coding and decision tree learning with an ensemble framework, which alternatively learns a dictionary for sparse coding and a decision tree for classification. The resulting series of decision trees as well as series of dictionaries are used to construct a decision forest for classification. The proposed method was applied to face recognition and scene classification, and the experimental results have demonstrated its power in comparison with recent supervised sparse coding methods. Yan Huang 0031, Yuhui Quan |
IEEE Signal Process. Lett. | 2 |
| 2019 | Exploiting Global Low-Rank Structure and Local Sparsity Nature for Tensor CompletionabstractIn the era of data science, a huge amount of data has emerged in the form of tensors. In many applications, the collected tensor data are incomplete with missing entries, which affects the analysis process. In this paper, we investigate a new method for tensor completion, in which a low-rank tensor approximation is used to exploit the global structure of data, and sparse coding is used for elucidating the local patterns of data. Regarding the characterization of low-rank structures, a weighted nuclear norm for the tensor is introduced. Meanwhile, an orthogonal dictionary learning process is incorporated into sparse coding for more effective discovery of the local details of data. By simultaneously using the global patterns and local cues, the proposed method can effectively and efficiently recover the lost information of incomplete tensor data. The capability of the proposed method is demonstrated with several experiments on recovering MRI data and visual data, and the experimental results have shown the excellent performance of the proposed method in comparison with recent related methods. Yong Du 0003, Guoqiang Han 0002, Yuhui Quan, Zhiwen Yu 0002, Hau-San Wong, C. L. Philip Chen, Jun Zhang 0003 |
IEEE Trans. Cybern. | 3 |
| 2018 | Sparse coding and dictionary learning with class-specific group sparsity
Yuping Sun, Yuhui Quan |
Neural Comput. Appl. | 2 |
| 2017 | Estimating Defocus Blur via Rank of Local PatchesabstractThis paper addresses the problem of defocus map estimation from a single image. We present a fast yet effective approach to estimate the spatially varying amounts of defocus blur at edge locations, which is based on the maximum ranks of the corresponding local patches with different orientations in gradient domain. Such an approach is motivated by the theoretical analysis which reveals the connection between the rank of a local patch blurred by a defocus-blur kernel and the blur amount by the kernel. After the amounts of defocus blur at edge locations are obtained, a complete defocus map is generated by a standard propagation procedure. The proposed method is extensively evaluated on real image datasets, and the experimental results show its superior performance to existing approaches. Yuhui Quan, Hui Ji 0002 |
ICCV | 2 |
| 2017 | Spatiotemporal lacunarity spectrum for dynamic texture classification
Yuhui Quan, Yuping Sun, Yong Xu 0007 |
Comput. Vis. Image Underst. | 1 |
| 2017 | Image-based action recognition using hint-enhanced deep neural networks
Tangquan Qi, Yong Xu 0007, Yuhui Quan, Haibin Ling |
Neurocomputing | 3 |
| 2016 | Equiangular Kernel Dictionary Learning with Applications to Dynamic Texture AnalysisabstractMost existing dictionary learning algorithms consider a linear sparse model, which often cannot effectively characterize the nonlinear properties present in many types of visual data, e.g. dynamic texture (DT). Such nonlinear properties can be exploited by the so-called kernel sparse coding. This paper proposed an equiangular kernel dictionary learning method with optimal mutual coherence to exploit the nonlinear sparsity of high-dimensional visual data. Two main issues are addressed in the proposed method: (1) coding stability for redundant dictionary of infinite-dimensional space, and (2) computational efficiency for computing kernel matrix of training samples of high-dimensional data. The proposed kernel sparse coding method is applied to dynamic texture analysis with both local DT pattern extraction and global DT pattern characterization. The experimental results showed its performance gain over existing methods. Yuhui Quan, Chenglong Bao, Hui Ji 0002 |
CVPR | 1 |
| 2016 | Sparse Coding for Classification via Discrimination EnsembleabstractDiscriminative sparse coding has emerged as a promising technique in image analysis and recognition, which couples the process of classifier training and the process of dictionary learning for improving the discriminability of sparse codes. Many existing approaches consider only a simple single linear classifier whose discriminative power is rather weak. In this paper, we proposed a discriminative sparse coding method which jointly learns a dictionary for sparse coding and an ensemble classifier for discrimination. The ensemble classifier is composed of a set of linear predictors and constructed via both subsampling on data and subspace projection on sparse codes. The advantages of the proposed method over the existing ones are multi-fold: better discriminability of sparse codes, weaker dependence on peculiarities of training data, and more expressibility of classifier for classification. These advantages are also justified in the experiments, as our method outperformed several recent methods in several recognition tasks. Yuhui Quan, Yong Xu 0007, Yuping Sun, Yan Huang 0031, Hui Ji 0002 |
CVPR | 1 |
| 2016 | Dictionary Learning for Sparse Coding: Algorithms and Convergence AnalysisabstractIn recent years, sparse coding has been widely used in many applications ranging from image processing to pattern recognition. Most existing sparse coding based applications require solving a class of challenging non-smooth and non-convex optimization problems. Despite the fact that many numerical methods have been developed for solving these problems, it remains an open problem to find a numerical method which is not only empirically fast, but also has mathematically guaranteed strong convergence. In this paper, we propose an alternating iteration scheme for solving such problems. A rigorous convergence analysis shows that the proposed method satisfies the global convergence property: the whole sequence of iterates is convergent and converges to a critical point. Besides the theoretical soundness, the practical benefit of the proposed method is validated in applications including image restoration and recognition. Experiments show that the proposed method achieves similar results with less computation when compared to widely used methods such as K-SVD. Chenglong Bao, Hui Ji 0002, Yuhui Quan, Zuowei Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Supervised dictionary learning with multiple classifier integration
Yuhui Quan, Yong Xu 0007, Yuping Sun, Yan Huang 0031 |
Pattern Recognit. | 1 |
| 2015 | Dynamic Texture Recognition via Orthogonal Tensor Dictionary LearningabstractDynamic textures (DTs) are video sequences with stationary properties, which exhibit repetitive patterns over space and time. This paper aims at investigating the sparse coding based approach to characterizing local DT patterns for recognition. Owing to the high dimensionality of DT sequences, existing dictionary learning algorithms are not suitable for our purpose due to their high computational costs as well as poor scalability. To overcome these obstacles, we proposed a structured tensor dictionary learning method for sparse coding, which learns a dictionary structured with orthogonality and separability. The proposed method is very fast and more scalable to high-dimensional data than the existing ones. In addition, based on the proposed dictionary learning method, a DT descriptor is developed, which has better adaptivity, discriminability and scalability than the existing approaches. These advantages are demonstrated by the experiments on multiple datasets. Yuhui Quan, Yan Huang 0031, Hui Ji 0002 |
ICCV | 1 |
| 2015 | Characterizing dynamic textures with space-time lacunarity analysisabstractThis paper addresses the challenge of reliably capturing the temporal characteristics of local space-time patterns in dynamic texture (DT). A powerful DT descriptor is proposed, which enjoys strong robustness to viewpoint changes, illumination changes, and video deformation. Observing that local DT patterns are spatial-temporally distributed with stationary irregularities, we proposed to characterize the distributions of local binarized DT patterns along both the temporal and the spatial axes via lacunarity analysis. We also observed such irregularities are similar on the DT slices along the same axis but distinct between axes. Thus, the resulting lacunarity based features are averaged along each axis and concatenated as the final DT descriptor. We applied the proposed DT descriptor to DT classification and evaluated its performance on several benchmark datasets. The experimental results have demonstrated the power of the proposed descriptor in comparison with existing ones. Yuping Sun, Yong Xu 0007, Yuhui Quan |
ICME | 3 |
| 2015 | Discriminative structured dictionary learning with hierarchical group sparsity
Yong Xu 0007, Yuping Sun, Yuhui Quan |
Comput. Vis. Image Underst. | 3 |
| 2015 | Classifying dynamic textures via spatiotemporal fractal analysis
Yong Xu 0007, Yuhui Quan, Zhuming Zhang, Haibin Ling, Hui Ji 0002 |
Pattern Recognit. | 2 |
| 2015 | Directional regularity for visual quality estimation
Delei Liu, Yong Xu 0007, Yuhui Quan, Zhiwen Yu 0002, Patrick Le Callet |
Signal Process. | 3 |
| 2014 | L0 Norm Based Dictionary Learning by Proximal Methods with Global ConvergenceabstractSparse coding and dictionary learning have seen their applications in many vision tasks, which usually is formulated as a non-convex optimization problem. Many iterative methods have been proposed to tackle such an optimization problem. However, it remains an open problem to have a method that is not only practically fast but also is globally convergent. In this paper, we proposed a fast proximal method for solving ℓ0norm based dictionary learning problems, and we proved that the whole sequence generated by the proposed method converges to a stationary point with sub-linear convergence rate. The benefit of having a fast and convergent dictionary learning method is demonstrated in the applications of image recovery and face recognition. Chenglong Bao, Hui Ji 0002, Yuhui Quan, Zuowei Shen |
CVPR | 3 |
| 2014 | Lacunarity Analysis on Image Patterns for Texture ClassificationabstractBased on the concept of lacunarity in fractal geometry, we developed a statistical approach to texture description, which yields highly discriminative feature with strong robustness to a wide range of transformations, including pho- tometric changes and geometric changes. The texture feature is constructed by concatenating the lacunarity-related parameters estimated from the multi-scale local binary patterns of image. Benefiting from the ability of lacunarity analysis to distinguish spatial patterns, our method is able to characterize the spatial distribution of local image structures from multiple scales. The proposed feature was applied to texture classification and has demonstrated excellent performance in comparison with several state-of-the- art approaches on four benchmark datasets. Yuhui Quan, Yong Xu 0007, Yuping Sun, Yu Luo 0004 |
CVPR | 1 |
| 2014 | A Convergent Incoherent Dictionary Learning Algorithm for Sparse Coding
Chenglong Bao, Yuhui Quan, Hui Ji 0002 |
ECCV (6) | 2 |
| 2014 | A distinct and compact texture descriptor
Yuhui Quan, Yong Xu 0007, Yuping Sun |
Image Vis. Comput. | 1 |
| 2014 | Reduced reference image quality assessment using regularity of phase congruency
Delei Liu, Yong Xu 0007, Yuhui Quan, Patrick Le Callet |
Signal Process. Image Commun. | 3 |
| 2012 | Contour-based recognitionabstractContour is an important cue for object recognition. In this paper, built upon the concept of torque in image space, we propose a new contour-related feature to detect and describe local contour information in images. There are two components for our proposed feature: One is a contour patch detector for detecting image patches with interesting information of object contour, which we call the Maximal/Minimal Torque Patch (MTP) detector. The other is a contour patch descriptor for characterizing a contour patch by sampling the torque values, which we call the Multi-scale Torque (MST) descriptor. Experiments for object recognition on the Caltech-101 dataset showed that the proposed contour feature outperforms other contour-related features and is on a par with many other types of features. When combing our descriptor with the complementary SIFT descriptor, impressive recognition results are observed. Yong Xu 0007, Yuhui Quan, Zhuming Zhang, Hui Ji 0002, Cornelia Fermüller, Morimichi Nishigaki, Daniel DeMenthon |
CVPR | 2 |
| 2011 | Dynamic texture classification using dynamic fractal analysisabstractIn this paper, we developed a novel tool called dynamic fractal analysis for dynamic texture (DT) classification, which not only provides a rich description of DT but also has strong robustness to environmental changes. The resulting dynamic fractal spectrum (DFS) for DT sequences consists of two components: One is the volumetric dynamic fractal spectrum component (V-DFS) that captures the stochastic self-similarities of DT sequences as 3D volume datasets; the other is the multi-slice dynamic fractal spectrum component (S-DFS) that encodes fractal structures of DT sequences on 2D slices along different views of the 3D volume. Various types of measures of DT sequences are collected in our approach to analyze DT sequences from different perspectives. The experimental evaluation is conducted on three widely used benchmark datasets. In all the experiments, our method demonstrated excellent performance in comparison with state-of-the-art approaches. Yong Xu 0007, Yuhui Quan, Haibin Ling, Hui Ji 0002 |
ICCV | 2 |