VLDB 2026 Research / reviewers in the wild / expert
Hui Ji 0002
dblp:52/6573-2
· DBLP profile ↗
94ranked-venue papers
11as first author
47since 2021 · last 2026
0000-0002-1674-6056ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 8 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 67 · 9 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep intrinsic image decomposition via physics-aware neural networks
Yan Huang 0031, Kangjie Liu, Tengyue Chen, Yong Xu 0007, Hui Ji 0002 |
Pattern Recognit. | 5 |
| 2025 | Multi-Focus Image Fusion via Explicit Defocus Blur ModellingabstractMulti-focus image fusion (MFIF) enhances depth of field in photography by generating an all-in-focus image from multiple images captured at different focal lengths. While deep learning has shown promise in MFIF, most existing methods overlooked the physical properties of defocus blurring in their network design, limiting their interoperability and generalization. This paper introduces a novel framework that integrates explicit defocus blur modelling into the MFIF process, improving both interpretability and performance. Using an atom-based spatially-varying parameterized defocus blurring model, our approach calculates pixel-wise defocus descriptors and initial focused images from multi-focus source images in a scale-recurrent manner to estimate soft decision maps. Fusion is then performed using masks derived from these decision maps, with special treatment for pixels likely defocused in all source images or near boundaries of defocused/focused regions. The model is trained with a fusion loss and a cross-scale defocus estimation loss. Extensive experiments on benchmark datasets demonstrated the effectiveness of our approach. Yuhui Quan, Xi Wan, Zitao Tang, Jinxiu Liang, Hui Ji 0002 |
AAAI | 5 |
| 2025 | A Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse ArtifactsabstractStructured artifacts are semi-regular, repetitive patterns that closely intertwine with genuine image content, making their removal highly challenging. In this paper, we introduce the Scale-Adaptive Deformable Transformer, an network architecture specifically designed to eliminate such artifacts from images. The proposed network features two key components: a scale-enhanced deformable convolution module for modeling scale-varying patterns with abundant orientations and potential distortions, and a scale-adaptive deformable attention mechanism for capturing long-range relationships among repetitive patterns with different sizes and non-uniform spatial distributions. Extensive experiments show that our network consistently outperforms state-of-the-art methods in diverse artifact removal tasks, including image deraining, image demoiréing, and image debanding. Xuyi He, Yuhui Quan, Ruotao Xu, Hui Ji 0002 |
CVPR | 4 |
| 2025 | Zero-Shot Blind-spot Image Denoising via Implicit Neural SamplingabstractThe blind-spot principle has been a widely used tool in zero-shot image denoising but faces challenges with real-world noise that exhibits strong local correlations. Existing methods focus on reducing noise correlation, which also weaken the pixel correlations needed for accurately estimating missing pixels. In this paper, we first present a rigorous analysis of how noise correlation and pixel correlation impact the statistical risk of a linear blind-spot denoiser. We then propose using an implicit neural representation to resample noisy pixels, effectively reducing noise correlation while preserving the essential pixel correlations for successful blind-spot denoising. Extensive experiments show our method surpasses existing zero-shot de-noising techniques on real-world noisy images. Yuhui Quan, Hui Ji 0002 |
CVPR | 4 |
| 2025 | Fingerprinting Denoising Diffusion Probabilistic ModelsabstractDiffusion models, especially denoising diffusion probabilistic models (DDPMs), are prevalent tools in generative AI, making their intellectual property (IP) protection increasingly important. Most existing IP protection methods for DDPMs are invasive, e.g., model watermarking, which alter model parameters and raise concerns about performance degradation, also with requirement for extra computational resources for retraining or fine-tuning. In this paper, we propose the first non-invasive fingerprinting scheme for DDPMs, requiring no parameter changes or fine-tuning, and keeping generation quality intact. We introduce a discriminative and robust fingerprint latent space based on the well-designed "crossing route" of noisy samples that span the performance border-zone of DDPMs, with only black-box access required for the diffusion denoiser in ownership verification. Extensive experiments demonstrate that our fingerprinting approach enjoys both robustness against the often-seen attacks and distinctiveness on various DDPMs, providing an alternative for protecting DDPMs’ IP rights without compromising their performance or integrity1. Huan Teng, Yuhui Quan, Chengyu Wang 0001, Jun Huang 0007, Hui Ji 0002 |
CVPR | 5 |
| 2025 | Robust Unfolding Network for HDR Imaging with Modulo Cameras
Zhile Chen, Hui Ji 0002 |
ICCV | 2 |
| 2025 | Zero-Shot Blind-Spot Image Denoising via Cross-Scale Non-Local Pixel RefillingabstractBlind-spot denoising (BSD) method is a powerful paradigm for zero-shot image denoising by training models to predict masked target pixels from their neighbors. However, they struggle with real-world noise exhibiting strong local correlations, where efforts to suppress noise correlation often weaken pixel-value dependencies, adversely affecting denoising performance. This paper presents a theoretical analysis quantifying the impact of replacing masked pixels with observations exhibiting weaker noise correlation but potentially reduced similarity, revealing a trade-off that impacts the statistical risk of the estimation. Guided by this insight, we propose a computational scheme that replaces masked pixels with distant ones of similar appearance and lower noise correlation. This strategy improves the prediction by balancing noise suppression and structural consistency. Experiments confirm the effectiveness of our method, outperforming existing zero-shot BSD methods. Qilong Guo, Tianjing Zhang, Hui Ji 0002 |
NeurIPS | 4 |
| 2025 | Image debanding using cross-scale invertible networks with banded deformable convolutions
Yuhui Quan, Xuyi He, Ruotao Xu, Yong Xu 0007, Hui Ji 0002 |
Neural Networks | 5 |
| 2025 | Image shadow removal via multi-scale deep Retinex decomposition
Yan Huang 0031, Xinchang Lu, Yuhui Quan, Yong Xu 0007, Hui Ji 0002 |
Pattern Recognit. | 5 |
| 2025 | Model Extraction for Image Denoising Networks
Huan Teng, Yuhui Quan, Yong Xu 0007, Jun Huang 0007, Hui Ji 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Dual-Path Deep Unsupervised Learning for Multi-Focus Image FusionabstractMulti-focus image fusion (MFIF) aims at merging multiple images captured at different focal lengths to create an all-in-focus image. This paper introduces a fully unsupervised learning approach for MFIF that uses only pairs of defocused images for end-to-end training, bypassing the need for ground-truths in supervised learning. Unlike existing methods training via a similarity loss between fused and source images, we propose a dual-path learning framework comprising two networks: an image fuser and a mask predictor. The mask predictor is modeled as a self-supervised denoising network on imperfect fusion masks, trained with a masking-based unsupervised learning scheme. The image fuser, crafted with deep unrolling, leverages the output from the mask predictor to supervise its mask generation at each unrolled step. Moreover, we introduce a fusion consistency loss to ensure the alignment between the image fuser and the mask predictor. In extensive experiments, our proposed approach shows superiority over existing end-to-end unsupervised methods and competitive performance against the supervised ones. Yuhui Quan, Xi Wan, Yan Huang 0031, Hui Ji 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | Robust Unsupervised Deep Learning for Nonblind Image Deconvolution With Inaccurate KernelsabstractNonblind image deconvolution/deblurring aims at restoring sharp images from their noisy blurred versions using an associated blur kernel with potential inaccuracy. Current deep learning (DL) models of nonblind image deconvolution (NBID) predominantly reply on ground truth (GT) images for supervision, which restricts their applicability to certain real-world scenarios such as scientific imaging. This article proposes a fully unsupervised DL approach for NBID, utilizing a GT-free end-to-end training process that adeptly handles both measurement noise and kernel error. Specifically, in the absence of GT images, a self-reconstruction loss is proposed to handle measurement noise, by effectively emulating its supervised counterpart. Recognizing the likely occurrence of kernel error during both training and testing data, we introduce a self-ensemble loss function and an ensemble inference scheme, anchored by a phase-keeping kernel perturbation strategy. Furthermore, a shifting mechanism is integrated so as to the loss functions to resolve the shift ambiguity caused by kernel error. Extensive experiments show the superiority of our proposed approach over existing unsupervised NBID methods, as well as its competitive performance against some of the recent supervised methods. Xinran Qin, Yuhui Quan, Zhuojie Chen, Hui Ji 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Unsupervised Deep Unrolling Networks for Phase UnwrappingabstractPhase unwrapping (PU) is a technique to reconstruct original phase images from their noisy wrapped counterparts, finding many applications in scientific imaging. Although supervised learning has shown promise in PU, its utility is limited in ground-truth (GT) scarce scenarios. This paper presents an unsupervised learning approach that eliminates the need for GTs during end-to-end training. Our approach leverages the insight that both the gradients and wrapped gradients of wrapped phases serve as noisy labels for GT phase gradients, along with sparse outliers induced by the wrapping operation. A recorruption-based self-reconstruction loss in the gradient domain is proposed to mitigate the adverse effects of label noise, complemented with a self-distillation loss for improved generalization. Additionally, by unfolding a variational model of PU that utilizes wrapped gradients of wrapped phases for its data-fitting term, we develop a deep unrolling network that encodes physics of phase wrapping and incorporates special treatments on outliers. In the experiments on three types of phase data, our approach outperforms existing GT-free methods and competes well against the supervised ones. Zhile Chen, Yuhui Quan, Hui Ji 0002 |
CVPR | 3 |
| 2024 | Test-Time Model Adaptation for Image Reconstruction Using Self-supervised Adaptive Layers
Yutian Zhao, Tianjing Zhang, Hui Ji 0002 |
ECCV (37) | 3 |
| 2024 | Enhancing Underwater Images via Asymmetric Multi-Scale Invertible NetworksabstractUnderwater images, often plagued by complex degradation, pose significant challenges for image enhancement. To address these challenges, the paper redefines underwater image enhancement as an image decomposition problem and proposes a deep invertible neural network (INN) that accurately predicts both the latent image and the degradation effects. Instead of using an explicit formation model to describe the degradation process, the INN adheres to the constraints of the image decomposition model, providing necessary regularization for model training, particularly in the absence of supervision on degradation effects. Taking into account the diverse scales of degradation factors, the INN is structured on a multi-scale basis to effectively manage the varied scales of degradation factors. Moreover, the INN incorporates several asymmetric design elements that are specifically optimized for the decomposition model and the unique physics of underwater imaging. Comprehensive experiments show that our approach provides significant performance improvement over existing methods. Yuhui Quan, Xiaoheng Tan, Yan Huang 0031, Yong Xu 0007, Hui Ji 0002 |
ACM Multimedia | 5 |
| 2024 | Pseudo-Siamese Blind-spot Transformers for Self-Supervised Real-World DenoisingabstractReal-world image denoising remains a challenge task. This paper studies self-supervised image denoising, requiring only noisy images captured in a single shot. We revamping the blind-spot technique by leveraging the transformer’s capability for long-range pixel interactions, which is crucial for effectively removing noise dependence in relating pixel–a requirement for achieving great performance for the blind-spot technique. The proposed method integrates these elements with two key innovations: a directional self-attention (DSA) module using a half-plane grid for self-attention, creating a sophisticated blind-spot structure, and a Siamese architecture with mutual learning to mitigate the performance impacts
from the restricted attention grid in DSA. Experiments on benchmark datasets demonstrate that our method outperforms existing self-supervised and clean-image-free methods. This combination of blind-spot and transformer techniques provides a natural synergy for tackling real-world image denoising challenges. Yuhui Quan, Hui Ji 0002 |
NeurIPS | 3 |
| 2024 | Cross-Scale Self-Supervised Blind Image Deblurring via Implicit Neural RepresentationabstractBlind image deblurring (BID) is an important yet challenging image recovery problem. Most existing deep learning methods require supervised training with ground truth (GT) images. This paper introduces a self-supervised method for BID that does not require GT images. The key challenge is to regularize the training to prevent over-fitting due to the absence of GT images. By leveraging an exact relationship among the blurred image, latent image, and blur kernel across consecutive scales, we propose an effective cross-scale consistency loss. This is implemented by representing the image and kernel with implicit neural representations (INRs), whose resolution-free property enables consistent yet efficient computation for network training across multiple scales. Combined with a progressively coarse-to-fine training scheme, the proposed method significantly outperforms existing self-supervised methods in extensive experiments. Tianjing Zhang, Yuhui Quan, Hui Ji 0002 |
NeurIPS | 3 |
| 2024 | Enhanced deep unrolling networks for snapshot compressive hyperspectral imaging
Xinran Qin, Yuhui Quan, Hui Ji 0002 |
Neural Networks | 3 |
| 2024 | Siamese Cooperative Learning for Unsupervised Image Reconstruction From Incomplete MeasurementsabstractImage reconstruction from incomplete measurements is one basic task in imaging. While supervised deep learning has emerged as a powerful tool for image reconstruction in recent years, its applicability is limited by its prerequisite on a large number of latent images for model training. To extend the application of deep learning to the imaging tasks where acquisition of latent images is challenging, this article proposes an unsupervised deep learning method that trains a deep model for image reconstruction with the access limited to measurement data. We develop a Siamese network whose twin sub-networks perform reconstruction cooperatively on a pair of complementary spaces: the null space of the measurement matrix and the range space of its pseudo inverse. The Siamese network is trained by a self-supervised loss with three terms: a data consistency loss over available measurements in the range space, a data consistency loss between intermediate results in the null space, and a mutual consistency loss on the predictions of the twin sub-networks in the full space. The proposed method is applied to four imaging tasks from different applications, and extensive experiments have shown its advantages over existing unsupervised solutions. Yuhui Quan, Xinran Qin, Tongyao Pang, Hui Ji 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Deep Single Image Defocus Deblurring via Gaussian Kernel Mixture LearningabstractThis paper proposes an end-to-end deep learning approach for removing defocus blur from a single defocused image. Defocus blur is a common issue in digital photography that poses a challenge due to its spatially-varying and large blurring effect. The proposed approach addresses this challenge by employing a pixel-wise Gaussian kernel mixture (GKM) model to accurately yet compactly parameterize spatially-varying defocus point spread functions (PSFs), which is motivated by the isotropy in defocus PSFs. We further propose a grouped GKM (GGKM) model that decouples the coefficients in GKM, so as to improve the modeling accuracy with an economic manner. Afterward, a deep neural network called GGKMNet is then developed by unrolling a fixed-point iteration process of GGKM-based image deblurring, which avoids the efficiency issues in existing unrolling DNNs. Using a lightweight scale-recurrent architecture with a coarse-to-fine estimation scheme to predict the coefficients in GGKM, the GGKMNet can efficiently recover an all-in-focus image from a defocused one. Such advantages are demonstrated with extensive experiments on five benchmark datasets, where the GGKMNet outperforms existing defocus deblurring methods in restoration quality, as well as showing advantages in terms of model complexity and computational efficiency. Yuhui Quan, Zicong Wu, Ruotao Xu, Hui Ji 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Unsupervised Deep Learning for Phase Retrieval via Teacher-Student DistillationabstractPhase retrieval (PR) is a challenging nonlinear inverse problem in scientific imaging that involves reconstructing the phase of a signal from its intensity measurements. Recently, there has been an increasing interest in deep learning-based PR. Motivated by the challenge of collecting ground-truth (GT) images in many domains, this paper proposes a fully-unsupervised learning approach for PR, which trains an end-to-end deep model via a GT-free teacher-student online distillation framework. Specifically, a teacher model is trained using a self-expressive loss with noise resistance, while a student model is trained with a consistency loss on augmented data to exploit the teacher's dark knowledge. Additionally, we develop an enhanced unfolding network for both the teacher and student models. Extensive experiments show that our proposed approach outperforms existing unsupervised PR methods with higher computational efficiency and performs competitively against supervised methods. Yuhui Quan, Zhile Chen, Tongyao Pang, Hui Ji 0002 |
AAAI | 4 |
| 2023 | Self-Supervised Blind Motion Deblurring with Deep Expectation MaximizationabstractWhen taking a picture, any camera shake during the shutter time can result in a blurred image. Recovering a sharp image from the one blurred by camera shake is a challenging yet important problem. Most existing deep learning methods use supervised learning to train a deep neural network (DNN) on a dataset of many pairs of blurred/latent images. In contrast, this paper presents a dataset-free deep learning method for removing uniform and non-uniform blur effects from images of static scenes. Our method involves a DNN-based re-parametrization of the latent image, and we propose a Monte Carlo Expectation Maximization (MCEM) approach to train the DNN without requiring any latent images. The Monte Carlo simulation is implemented via Langevin dynamics. Experiments showed that the proposed method outperforms existing methods significantly in removing motion blur from images of static scenes. Weixi Wang, Yuesong Nan, Hui Ji 0002 |
CVPR | 4 |
| 2023 | Ground-Truth Free Meta-Learning for Deep Compressive SamplingabstractCompressive sampling (CS) is an efficient technique for imaging. This paper proposes a ground-truth (GT) free meta-learning method for CS, which leverages both ex-ternal and internal deep learning for unsupervised high-quality image reconstruction. The proposed method first trains a deep neural network (NN) via external meta-learning using only CS measurements, and then efficiently adapts the trained model to a test sample for exploiting sample-specific internal characteristic for performance gain. The meta-learning and model adaptation are built on an improved Stein's unbiased risk estimator (iSURE) that provides efficient computation and effective guidance for accurate prediction in the range space of the adjoint of the measurement matrix. To improve the learning and adaption on the null space of the measurement matrix, a modi-fied model-agnostic meta-learning scheme and a null-space consistency loss are proposed. In addition, a bias tuning scheme for unrolling NNs is introduced for further acceler-ation of model adaption. Experimental results have demonstrated that the proposed GT-free method performs well and can even compete with supervised methods. Xinran Qin, Yuhui Quan, Tongyao Pang, Hui Ji 0002 |
CVPR | 4 |
| 2023 | Neumann Network with Recursive Kernels for Single Image Defocus DeblurringabstractSingle image defocus deblurring (SIDD) refers to recovering an all-in-focus image from a defocused blurry one. It is a challenging recovery task due to the spatially-varying defocus blurring effects with significant size variation. Motivated by the strong correlation among defocus kernels of different sizes and the blob-type structure of defocus kernels, we propose a learnable recursive kernel representation (RKR) for defocus kernels that expresses a defocus kernel by a linear combination of recursive, separable and positive atom kernels, leading to a compact yet effective and physics-encoded parametrization of the spatially-varying defocus blurring process. Afterwards, a physics-driven and efficient deep model with a cross-scale fusion structure is presented for SIDD, with inspirations from the truncated Neumann series for approximating the matrix inversion of the RKR-based blurring operator. In addition, a reblurring loss is proposed to regularize the RKR learning. Extensive experiments show that, our proposed approach significantly outperforms existing ones, with a model size comparable to that of the top methods. Yuhui Quan, Zicong Wu, Hui Ji 0002 |
CVPR | 3 |
| 2023 | Fingerprinting Deep Image Restoration ModelsabstractFingerprinting is a promising non-invasive method for protecting the intellectual property rights (IPR) of deep neural network (DNN) models. It extracts a feature called a fingerprint from a DNN model to identify its ownership. Existing fingerprinting methods focus only on classification-related models that map images to labels, while inapplicable to models for image restoration that map images to images. This paper proposes a fingerprinting framework for DNN models of image restoration. The proposed framework defines the fingerprint using a critical image, which exhibits strongly discriminative patterns and is robust to modest model modifications. Model ownership is then verified by comparing the distance of color histograms and local gradient pattern histograms of critical images between the suspect and source models. We apply the proposed framework to two representative tasks, denoising and super-resolution. It outperforms the baselines of fingerprinting and competes against existing invasive model watermarking methods. Yuhui Quan, Huan Teng, Ruotao Xu, Jun Huang 0007, Hui Ji 0002 |
ICCV | 5 |
| 2023 | Single Image Defocus Deblurring via Implicit Neural Inverse KernelsabstractSingle image defocus deblurring (SIDD) is a challenging task due to the spatially-varying nature of defocus blur, characterized by per-pixel point spread functions (PSFs). Existing deep-learning-based methods for SIDD are limited by either over-fitting due to the lack of model constraints or under-parametrization that restricts their applicability to real-world images. To address the limitations, this paper proposes an interpretable approach that explicitly predicts inverse kernels with structural regularization. Motivated by the observation that defocus PSFs within an image often have similar shapes but different sizes, we represent the inverse kernels linearly over a multi-scale dictionary parameterized by implicit neural representations. We predict the corresponding representation coefficients via a duplex scale-recurrent neural network that jointly performs fine-to-coarse and coarse-to-fine estimations. Extensive experiments demonstrate that our approach achieves excellent performance using a lightweight model. Yuhui Quan, Hui Ji 0002 |
ICCV | 3 |
| 2023 | Self-Supervised Deep Learning for Image Reconstruction: A Langevin Monte Carlo ApproachabstractAbstract. Deep learning has proved to be a powerful tool for solving inverse problems in imaging, and most of the related work is based on supervised learning. In many applications, collecting truth images is a challenging and costly task, and the prerequisite of having a training dataset of truth images limits its applicability. This paper proposes a self-supervised deep learning method for solving inverse imaging problems that does not require any training samples. The proposed approach is built on a reparametrization of latent images using a convolutional neural network, and the reconstruction is motivated by approximating the minimum mean square error estimate of the latent image using a Langevin dynamics–based Monte Carlo (MC) method. To efficiently sample the network weights in the context of image reconstruction, we propose a Langevin MC scheme called Adam-LD, inspired by the well-known optimizer in deep learning, Adam. The proposed method is applied to solve linear and nonlinear inverse problems, specifically, sparse-view computed tomography image reconstruction and phase retrieval. Our experiments demonstrate that the proposed method outperforms existing unsupervised or self-supervised solutions in terms of reconstruction quality. Weixi Wang, Hui Ji 0002 |
SIAM J. Imaging Sci. | 3 |
| 2023 | Self-Supervised Blind Image Deconvolution via Deep Generative Ensemble LearningabstractBlind image deconvolution (BID) is about recovering a latent image with sharp details from its blurred observation generated by the convolution with an unknown smoothing kernel. Recently, deep generative priors from untrained neural networks (NNs) have emerged as a promising deep learning approach for BID, with the benefit of being free of external training samples. However, existing untrained-NN-based BID methods may suffer from under-deblurring or overfitting. In this paper, we propose an ensemble approach to better exploit the priors from untrained NNs for BID, which aggregates the deblurring results of multiple untrained NNs for improvement. To enjoy both the effectiveness and computational efficiency in ensemble learning, the untrained NNs are designed with a specific shared-base and multi-head architecture. In addition, a kernel-centering layer is proposed for handling the shift ambiguity among different predictions during ensemble, which also improves the robustness of kernel prediction to the setting of the kernel size parameter. Extensive experiments show that the proposed approach noticeably outperforms both exiting dataset-free methods and dataset-based methods. Mingqin Chen, Yuhui Quan, Yong Xu 0007, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Image Desnowing via Deep Invertible SeparationabstractImages taken on snowy days often suffer from severe negative visual effects caused by snowflakes. The task of removing snowflakes from a snowy image is known as image desnowing, which is challenging as image details are easily mistakenly treated and thus may be significantly lost during snowflake removal. Leveraging invertible neural networks (INNs), this paper presents a deep learning-based method for single image desnowing, which can remove snowflakes accurately while preserving image details well. Interpreting desnowing as an image decomposition problem, we propose an INN composed of two asymmetric interactive paths for predicting a latent image and a snowflake layer respectively. Such an INN is able to progressively refine the features of both latent images and snowflake layers for disentanglement, while retaining all information possibly relevant to latent image reconstruction. In addition, an attentive coupling layer supervised by snowflake masks is introduced to enhance feature dismantlement and a coupling-in-coupling structure is developed for further improvement. Extensive experiments show that, the proposed method outperforms existing ones on three benchmark datasets of synthetic and real-world images, and meanwhile it also shows advantages in terms of model size and computational efficiency. Yuhui Quan, Xiaoheng Tan, Yan Huang 0031, Yong Xu 0007, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Self-supervised Deep Image Restoration via Adaptive Stochastic Gradient Langevin DynamicsabstractWhile supervised deep learning has been a prominent tool for solving many image restoration problems, there is an increasing interest on studying self-supervised or un-supervised methods to address the challenges and costs of collecting truth images. Based on the neuralization of a Bayesian estimator of the problem, this paper presents a self-supervised deep learning approach to general image restoration problems. The key ingredient of the neuralized estimator is an adaptive stochastic gradient Langevin dy-namics algorithm for efficiently sampling the posterior distri-bution of network weights. The proposed method is applied on two image restoration problems: compressed sensing and phase retrieval. The experiments on these applications showed that the proposed method not only outperformed existing non-learning and unsupervised solutions in terms of image restoration quality, but also is more computationally efficient. Weixi Wang, Hui Ji 0002 |
CVPR | 3 |
| 2022 | Learning Deep Non-blind Image Deconvolution Without Ground Truths
Yuhui Quan, Zhuojie Chen, Hui Ji 0002 |
ECCV (6) | 4 |
| 2022 | Dual-Domain Self-supervised Learning and Model Adaption for Deep Compressive Imaging
Yuhui Quan, Xinran Qin, Tongyao Pang, Hui Ji 0002 |
ECCV (30) | 4 |
| 2022 | Deep Scale-Aware Image SmoothingabstractImage smoothing, a technique for smoothing out insignificant textures while preserving meaningful structures, is an important component in many vision and graphics applications. Scale-awareness plays a fundamental role in image smoothing, as insignificant textures and noise usually are at fine scales while meaningful boundary objects are at coarse scales. This paper proposes a deep-learning-based scale-aware image smoothing method, which is built on a downscaling-upscaling mechanism with attention. The downscaling mechanism is for predicting large-scale salient structures, and the upscaling mechanism is for identifying and inferring insignificant small-scale details from the salient structures. In the experiments, the proposed one provides a noticeable performance improvement over recent methods. Kunkun Qin, Ruotao Xu, Hui Ji 0002 |
ICASSP | 4 |
| 2022 | Nonblind Image Deconvolution via Leveraging Model Uncertainty in An Untrained Deep Neural Network
Mingqin Chen, Yuhui Quan, Tongyao Pang, Hui Ji 0002 |
Int. J. Comput. Vis. | 4 |
| 2022 | $L_1$-Norm Regularization for Short-and-Sparse Blind Deconvolution: Point Source Separability and Region SelectionabstractBlind deconvolution is about estimating both the convolution kernel and the latent signal from their convolution. Many blind deconvolution problems have a short-and-sparse (SaS) structure; i.e., the signal (or its gradient) is sparse and the kernel size is much smaller than the signal size. While $\ell_1$-norm relating regularizations have been widely used for solving SaS blind deconvolution problems, the so-called region/edge selection technique brings great empirical improvement to such $\ell_1$-norm relating regularizations in image deblurring. The essence of region/edge selection is during an alternative iterative scheme of SaS blind deconvolution: one estimates the kernel on an estimate of the latent image with well-separated image edges instead of the one with the least fitting error. In this paper, we first examine the validity and soundness of $\ell_1$-norm relating regularization in the setting of 1D SaS blind deconvolution. The analysis reveals the importance of the separation of nonzero signal entries toward the soundness of such a regularization. The studies laid out the foundation of region selection technique; i.e., during the iteration, an estimate of the latent image with well-separated edges is a better candidate for estimating the kernel than the one with the least fitting error. Based on the studies conducted in this paper, an alternating iterative scheme with region selection model is developed for SaS blind deconvolution, which is then applied to blind motion deblurring. The experiments show its effectiveness over many existing $\ell_1$-norm relating approaches. Weixi Wang, Hui Ji 0002 |
SIAM J. Imaging Sci. | 3 |
| 2022 | Self-Supervised Low-Light Image Enhancement Using Discrepant Untrained Network PriorsabstractThis paper proposes a deep learning method for low-light image enhancement, which exploits the generation capability of Neural Networks (NNs) while requiring no training samples except the input image itself. Based on the Retinex decomposition model, the reflectance and illumination of a low-light image are parameterized by two untrained NNs. The ambiguity between the two layers is resolved by the discrepancy between the two NNs in terms of architecture and capacity, while the complex noise with spatially-varying characteristics is handled by an illumination-adaptive self-supervised denoising module. The enhancement is done by jointly optimizing the Retinex decomposition and the illumination adjustment. Extensive experiments show that the proposed method not only outperforms existing non-learning-based and unsupervised-learning-based methods, but also competes favorably with some supervised-learning-based methods in extreme low-light conditions. Jinxiu Liang, Yong Xu 0007, Yuhui Quan, Boxin Shi, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Unsupervised Deep Background Matting Using Deep Matte PriorabstractBackground matting is a recently developed image matting approach, with applications to image and video editing. It refers to estimating both the alpha matte and foreground from a pair of images with and without foreground objects. Recent work has applied deep learning to background matting, with very promising performance achieved. However, existing deep models are supervised which require a large dataset with ground truth alpha mattes for training. To avoid the cost of data collection and possible bias in training data, this paper proposes a dataset-free unsupervised deep learning-based approach for background matting. Observing that the local smoothness of alpha matte can be well characterized by the untrained network prior called deep matte prior, we model the foreground and alpha matte using the priors encoded by two generative convolutional neural networks. To avoid possible overfitting during unsupervised learning, a two-stage learning scheme is developed which contains projection-based training and Bayesian post refinement. An alpha-matte-driven initialization scheme is also developed for performance boost. Even without calling external training data, the proposed approach provides competitive performance to recent supervised learning-based methods in the experiments. Yong Xu 0007, Baoling Liu, Yuhui Quan, Hui Ji 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Nonblind Image Deblurring via Deep Learning in Complex FieldabstractNonblind image deblurring is about recovering the latent clear image from a blurry one generated by a known blur kernel, which is an often-seen yet challenging inverse problem in imaging. Its key is how to robustly suppress noise magnification during the inversion process. Recent approaches made a breakthrough by exploiting convolutional neural network (CNN)-based denoising priors in the image domain or the gradient domain, which allows using a CNN for noise suppression. The performance of these approaches is highly dependent on the effectiveness of the denoising CNN in removing magnified noise whose distribution is unknown and varies at different iterations of the deblurring process for different images. In this article, we introduce a CNN-based image prior defined in the Gabor domain. The prior not only utilizes the optimal space-frequency resolution and strong orientation selectivity of the Gabor transform but also enables using complex-valued (CV) representations in intermediate processing for better denoising. A CV CNN is developed to exploit the benefits of the CV representations, with better generalization to handle unknown noises over the real-valued ones. Combining our Gabor-domain CV CNN-based prior with an unrolling scheme, we propose a deep-learning-based approach to nonblind image deblurring. Extensive experiments have demonstrated the superior performance of the proposed approach over the state-of-the-art ones. Yuhui Quan, Peikang Lin, Yong Xu 0007, Yuesong Nan, Hui Ji 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Deep Texture Recognition via Exploiting Cross-Layer Statistical Self-SimilarityabstractIn recent years, convolutional neural networks (CNNs) have become a prominent tool for texture recognition. The key of existing CNN-based approaches is aggregating the convolutional features into a robust yet discriminative description. This paper presents a novel feature aggregation module called CLASS (Cross-Layer Aggregation of Statistical Self-similarity) for texture recognition. We model the CNN feature maps across different layers, as a dynamic process which carries the statistical self-similarity (SSS), one well-known property of texture, from input image along the network depth dimension. The CLASS module characterizes the cross-layer SSS using a soft histogram of local differential box-counting dimensions of cross-layer features. The resulting descriptor encodes both cross-layer dynamics and local SSS of input image, providing additional discrimination over the often-used global average pooling. Integrating CLASS into a ResNet backbone, we develop CLASSNet, an effective deep model for texture recognition, which shows state-of-the-art performance in the experiments. Zhile Chen, Yuhui Quan, Yong Xu 0007, Hui Ji 0002 |
CVPR | 5 |
| 2021 | Recorrupted-to-Recorrupted: Unsupervised Deep Learning for Image DenoisingabstractDeep denoiser, the deep network for denoising, has been the focus of the recent development on image denoising. In the last few years, there is an increasing interest in developing unsupervised deep denoisers which only call unorganized noisy images without ground truth for training. Nevertheless, the performance of these unsupervised deep denoisers is not competitive to their supervised counterparts. Aiming at developing a more powerful unsupervised deep denoiser, this paper proposed a data augmentation technique, called recorrupted-to-recorrupted (R2R), to address the overfitting caused by the absence of truth images. For each noisy image, we showed that the cost function defined on the noisy/noisy image pairs constructed by the R2R method is statistically equivalent to its supervised counterpart defined on the noisy/truth image pairs. Extensive experiments showed that the proposed R2R method noticeably outperformed existing unsupervised deep denoisers, and is competitive to representative supervised deep denoisers. Tongyao Pang, Yuhui Quan, Hui Ji 0002 |
CVPR | 4 |
| 2021 | Learnable Multi-scale Fourier Interpolation for Sparse View CT Image Reconstruction
Qiaoqiao Ding, Hui Ji 0002, Hao Gao 0003, Xiaoqun Zhang |
MICCAI (6) | 2 |
| 2021 | Gaussian Kernel Mixture Network for Single Image Defocus DeblurringabstractDefocus blur is one kind of blur effects often seen in images, which is challenging to remove due to its spatially variant amount. This paper presents an end-to-end deep learning approach for removing defocus blur from a single image, so as to have an all-in-focus image for consequent vision tasks. First, a pixel-wise Gaussian kernel mixture (GKM) model is proposed for representing spatially variant defocus blur kernels in an efficient linear parametric form, with higher accuracy than existing models. Then, a deep neural network called GKMNet is developed by unrolling a fixed-point iteration of the GKM-based deblurring. The GKMNet is built on a lightweight scale-recurrent architecture, with a scale-recurrent attention module for estimating the mixing coefficients in GKM for defocus deblurring. Extensive experiments show that the GKMNet not only noticeably outperforms existing defocus deblurring methods, but also has its advantages in terms of model complexity and computational efficiency. Yuhui Quan, Zicong Wu, Hui Ji 0002 |
NeurIPS | 3 |
| 2021 | Attentive deep network for blind motion deblurring on dynamic scenes
Yong Xu 0007, Ye Zhu 0003, Yuhui Quan, Hui Ji 0002 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Fast vertex-based graph convolutional neural network and its application to brain images
Chaoqiang Liu, Hui Ji 0002, Anqi Qiu |
Neurocomputing | 2 |
| 2021 | Rethinking medical image reconstruction via shape prior, going deeper and faster: Deep joint indirect registration and reconstruction
Jiulong Liu, Angelica I. Avilés-Rivero, Hui Ji 0002, Carola-Bibiane Schönlieb |
Medical Image Anal. | 3 |
| 2021 | Image denoising using complex-valued deep CNN
Yuhui Quan, Yizhen Shao, Huan Teng, Yong Xu 0007, Hui Ji 0002 |
Pattern Recognit. | 6 |
| 2021 | Watermarking Deep Neural Networks in Image ProcessingabstractPublishing/sharing pretrained deep neural network (DNN) models is a common practice in the community of computer vision. The increasing popularity of pretrained models has made it a serious concern: how to protect the intellectual properties of model owners and avert illegal usages by malicious attackers. This article aims at developing a framework for watermarking DNNs, with a particular focus on low-level image processing tasks that map images to images. Using image denoising and superresolution as case studies, we develop a black-box watermarking method for pretrained models, which exploits the overparameterization of the DNNs in image processing. In addition, an auxiliary module for visualizing the watermark information is proposed for further verification. Extensive experiments show that the proposed watermarking framework has no noticeable impact on model performance and enjoys the robustness against the often-seen attacks. Yuhui Quan, Huan Teng, Hui Ji 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Deep Learning for Handling Kernel/model Uncertainty in Image DeconvolutionabstractMost existing non-blind image deconvolution methods assume that the given blurring kernel is error-free. In practice, blurring kernel often is estimated via some blind deblurring algorithm which is not exactly the truth. Also, the convolution model is only an approximation to practical blurring effect. It is known that non-blind deconvolution is susceptible to such a kernel/model error. Based on an error-in-variable (EIV) model of image blurring that takes kernel error into consideration, this paper presents a deep learning method for deconvolution, which unrolls a total-least-squares (TLS) estimator whose relating priors are learned by neural networks (NNs). The experiments showed that the proposed method is robust to kernel/model error. It noticeably outperformed existing solutions when deblurring images using noisy kernels, e.g. the ones estimated from existing blind motion deblurring methods. Yuesong Nan, Hui Ji 0002 |
CVPR | 2 |
| 2020 | Variational-EM-Based Deep Learning for Noise-Blind Image DeblurringabstractNon-blind deblurring is an important problem encountered in many image restoration tasks. The focus of non-blind deblurring is on how to suppress noise magnification during deblurring. In practice, it often happens that the noise level of input image is unknown and varies among different images. This paper aims at developing a deep learning framework for deblurring images with unknown noise level. Based on the framework of variational expectation maximization (EM), an iterative noise-blind deblurring scheme is proposed which integrates the estimation of noise level and the quantification of image prior uncertainty. Then, the proposed scheme is unrolled to a neural network (NN) where image prior is modeled by NN with uncertainty quantification. Extensive experiments showed that the proposed method not only outperformed existing noise-blind deblurring methods by a large margin, but also outperformed those state-of-the-art image deblurring methods designed/trained with known noise level. Yuesong Nan, Yuhui Quan, Hui Ji 0002 |
CVPR | 3 |
| 2020 | Self2Self With Dropout: Learning Self-Supervised Denoising From Single ImageabstractIn last few years, supervised deep learning has emerged as one powerful tool for image denoising, which trains a denoising network over an external dataset of noisy/clean image pairs. However, the requirement on a high-quality training dataset limits the broad applicability of the denoising networks. Recently, there have been a few works that allow training a denoising network on the set of external noisy images only. Taking one step further, this paper proposes a self-supervised learning method which only uses the input noisy image itself for training. In the proposed method, the network is trained with dropout on the pairs of Bernoulli-sampled instances of the input image, and the result is estimated by averaging the predictions generated from multiple instances of the trained model with dropout. The experiments show that the proposed method not only significantly outperforms existing single-image learning or non-learning methods, but also is competitive to the denoising networks trained on external datasets. Yuhui Quan, Mingqin Chen, Tongyao Pang, Hui Ji 0002 |
CVPR | 4 |
| 2020 | Self-supervised Bayesian Deep Learning for Image Recovery with Applications to Compressive Sensing
Tongyao Pang, Yuhui Quan, Hui Ji 0002 |
ECCV (11) | 3 |
| 2020 | Cartoon-Texture Image Decomposition using Orientation Characteristics in Patch RecurrenceabstractCartoon-texture image decomposition is about decomposing an image into the linear sum of two layers: cartoon and texture, where the key challenge is how to resolve the ambiguity between two layers. It is observed that the recurrence of texture patches occurs along multiple orientations, and the recurrence of cartoon patches only occurs along certain orientations. This paper proposes to separate these two layers by exploiting their orientation characteristics of image patch recurrence, i.e., isotropy property of texture patch recurrence versus anisotropy property of cartoon patch recurrence. Together with the sparsity-based regularizations in the image domain, a variational method is then developed in this paper for cartoon-texture decomposition. The experiments show that the proposed method noticeably outperforms many well-established ones on test images. Ruotao Xu, Yong Xu 0007, Yuhui Quan, Hui Ji 0002 |
SIAM J. Imaging Sci. | 4 |
| 2020 | Image Denoising via Sequential Ensemble LearningabstractImage denoising is about removing measurement noise from input image for better signal-to-noise ratio. In recent years, there has been great progress on the development of data-driven approaches for image denoising, which introduce various techniques and paradigms from machine learning in the design of image denoisers. This paper aims at investigating the application of ensemble learning in image denoising, which combines a set of simple base denoisers to form a more effective image denoiser. Based on different types of image priors, two types of base denoisers in the form of transform-shrinkage are proposed for constructing the ensemble. Then, with an effective re-sampling scheme, several ensemble-learning-based image denoisers are constructed using different sequential combinations of multiple proposed base denoisers. The experiments showed that sequential ensemble learning can effectively boost the performance of image denoising. Xuhui Yang, Yong Xu 0007, Yuhui Quan, Hui Ji 0002 |
IEEE Trans. Image Process. | 4 |
| 2019 | A Variational EM Framework With Adaptive Edge Selection for Blind Motion DeblurringabstractBlind motion deblurring is an important problem that receives enduring attention in last decade. Based on the observation that a good intermediate estimate of latent image for estimating motion-blur kernel is not necessarily the one closest to latent image, edge selection has proven itself a very powerful technique for achieving state-of-the-art performance in blind deblurring. This paper presented an interpretation of edge selection/reweighting in terms of variational Bayes inference, and therefore developed a novel variational expectation maximization (VEM) algorithm with built-in adaptive edge selection for blind deblurring. Together with a restart strategy for avoiding undesired local convergence, the proposed VEM method not only has a solid mathematical foundation but also noticeably outperformed the state-of-the-art methods on benchmark datasets. Liuge Yang, Hui Ji 0002 |
CVPR | 2 |
| 2019 | Deep Learning for Seeing Through Window With RaindropsabstractWhen taking pictures through glass window in rainy day, the images are comprised and corrupted by the raindrops adhered to glass surfaces. It is a challenging problem to remove the effect of raindrops from an image. The key task is how to accurately and robustly identify the raindrop regions in an image. This paper develops a convolutional neural network (CNN) for removing the effect of raindrops from an image. In the proposed CNN, we introduce a double attention mechanism that concurrently guides the CNN using shape-driven attention and channel re-calibration. The shape-driven attention exploits physical shape priors of raindrops, i.e. convexness and contour closedness, to accurately locate raindrops, and the channel re-calibration improves the robustness when processing raindrops with varying appearances. The experimental results show that the proposed CNN outperforms the state-of-the-art approaches in terms of both quantitative metrics and visual quality. Yuhui Quan, Hui Ji 0002 |
ICCV | 4 |
| 2019 | Attention with structure regularization for action recognition
Yuhui Quan, Ruotao Xu, Hui Ji 0002 |
Comput. Vis. Image Underst. | 4 |
| 2019 | Barzilai-Borwein-based adaptive learning rate for deep learning
Jinxiu Liang, Yong Xu 0007, Chenglong Bao, Yuhui Quan, Hui Ji 0002 |
Pattern Recognit. Lett. | 5 |
| 2018 | Coherence Retrieval Using Trace RegularizationabstractThe mutual intensity and its equivalent phase-space representations quantify an optical field's state of coherence and are important tools in the study of light propagation and dynamics, but they can only be estimated indirectly from measurements through a process called coherence retrieval, otherwise known as phase-space tomography. As practical considerations often rule out the availability of a complete set of measurements, coherence retrieval is usually a challenging high-dimensional ill-posed inverse problem. In this paper, we propose a trace-regularized optimization model for coherence retrieval and a provably convergent adaptive accelerated proximal gradient algorithm for solving the resulting problem. Applying our model and algorithm to both simulated and experimental data, we demonstrate an improvement in reconstruction quality over previous models as well as an increase in convergence speed compared to existing first-order methods. Chenglong Bao, George Barbastathis, Hui Ji 0002, Zuowei Shen, Zhengyun Zhang |
SIAM J. Imaging Sci. | 3 |
| 2017 | Estimating Defocus Blur via Rank of Local PatchesabstractThis paper addresses the problem of defocus map estimation from a single image. We present a fast yet effective approach to estimate the spatially varying amounts of defocus blur at edge locations, which is based on the maximum ranks of the corresponding local patches with different orientations in gradient domain. Such an approach is motivated by the theoretical analysis which reveals the connection between the rank of a local patch blurred by a defocus-blur kernel and the blur amount by the kernel. After the amounts of defocus blur at edge locations are obtained, a complete defocus map is generated by a standard propagation procedure. The proposed method is extensively evaluated on real image datasets, and the experimental results show its superior performance to existing approaches. Yuhui Quan, Hui Ji 0002 |
ICCV | 3 |
| 2016 | Equiangular Kernel Dictionary Learning with Applications to Dynamic Texture AnalysisabstractMost existing dictionary learning algorithms consider a linear sparse model, which often cannot effectively characterize the nonlinear properties present in many types of visual data, e.g. dynamic texture (DT). Such nonlinear properties can be exploited by the so-called kernel sparse coding. This paper proposed an equiangular kernel dictionary learning method with optimal mutual coherence to exploit the nonlinear sparsity of high-dimensional visual data. Two main issues are addressed in the proposed method: (1) coding stability for redundant dictionary of infinite-dimensional space, and (2) computational efficiency for computing kernel matrix of training samples of high-dimensional data. The proposed kernel sparse coding method is applied to dynamic texture analysis with both local DT pattern extraction and global DT pattern characterization. The experimental results showed its performance gain over existing methods. Yuhui Quan, Chenglong Bao, Hui Ji 0002 |
CVPR | 3 |
| 2016 | Sparse Coding for Classification via Discrimination EnsembleabstractDiscriminative sparse coding has emerged as a promising technique in image analysis and recognition, which couples the process of classifier training and the process of dictionary learning for improving the discriminability of sparse codes. Many existing approaches consider only a simple single linear classifier whose discriminative power is rather weak. In this paper, we proposed a discriminative sparse coding method which jointly learns a dictionary for sparse coding and an ensemble classifier for discrimination. The ensemble classifier is composed of a set of linear predictors and constructed via both subsampling on data and subspace projection on sparse codes. The advantages of the proposed method over the existing ones are multi-fold: better discriminability of sparse codes, weaker dependence on peculiarities of training data, and more expressibility of classifier for classification. These advantages are also justified in the experiments, as our method outperformed several recent methods in several recognition tasks. Yuhui Quan, Yong Xu 0007, Yuping Sun, Yan Huang 0031, Hui Ji 0002 |
CVPR | 5 |
| 2016 | Dictionary Learning for Sparse Coding: Algorithms and Convergence AnalysisabstractIn recent years, sparse coding has been widely used in many applications ranging from image processing to pattern recognition. Most existing sparse coding based applications require solving a class of challenging non-smooth and non-convex optimization problems. Despite the fact that many numerical methods have been developed for solving these problems, it remains an open problem to find a numerical method which is not only empirically fast, but also has mathematically guaranteed strong convergence. In this paper, we propose an alternating iteration scheme for solving such problems. A rigorous convergence analysis shows that the proposed method satisfies the global convergence property: the whole sequence of iterates is convergent and converges to a critical point. Besides the theoretical soundness, the practical benefit of the proposed method is validated in applications including image restoration and recognition. Experiments show that the proposed method achieves similar results with less computation when compared to widely used methods such as K-SVD. Chenglong Bao, Hui Ji 0002, Yuhui Quan, Zuowei Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Removing Rain from a Single Image via Discriminative Sparse CodingabstractVisual distortions on images caused by bad weather conditions can have a negative impact on the performance of many outdoor vision systems. One often seen bad weather is rain which causes significant yet complex local intensity fluctuations in images. The paper aims at developing an effective algorithm to remove visual effects of rain from a single rainy image, i.e. separate the rain layer and the de-rained image layer from an rainy image. Built upon a non-linear generative model of rainy image, namely screen blend mode, we proposed a dictionary learning based algorithm for single image de-raining. The basic idea is to sparsely approximate the patches of two layers by very high discriminative codes over a learned dictionary with strong mutual exclusivity property. Such discriminative sparse codes lead to accurate separation of two layers from their non-linear composite. The experiments showed that the proposed method outperformed the existing single image de-raining methods on tested rain images. Yu Luo 0004, Yong Xu 0007, Hui Ji 0002 |
ICCV | 3 |
| 2015 | Dynamic Texture Recognition via Orthogonal Tensor Dictionary LearningabstractDynamic textures (DTs) are video sequences with stationary properties, which exhibit repetitive patterns over space and time. This paper aims at investigating the sparse coding based approach to characterizing local DT patterns for recognition. Owing to the high dimensionality of DT sequences, existing dictionary learning algorithms are not suitable for our purpose due to their high computational costs as well as poor scalability. To overcome these obstacles, we proposed a structured tensor dictionary learning method for sparse coding, which learns a dictionary structured with orthogonality and separability. The proposed method is very fast and more scalable to high-dimensional data than the existing ones. In addition, based on the proposed dictionary learning method, a DT descriptor is developed, which has better adaptivity, discriminability and scalability than the existing approaches. These advantages are demonstrated by the experiments on multiple datasets. Yuhui Quan, Yan Huang 0031, Hui Ji 0002 |
ICCV | 3 |
| 2015 | Classifying dynamic textures via spatiotemporal fractal analysis
Yong Xu 0007, Yuhui Quan, Zhuming Zhang, Haibin Ling, Hui Ji 0002 |
Pattern Recognit. | 5 |
| 2014 | L0 Norm Based Dictionary Learning by Proximal Methods with Global ConvergenceabstractSparse coding and dictionary learning have seen their applications in many vision tasks, which usually is formulated as a non-convex optimization problem. Many iterative methods have been proposed to tackle such an optimization problem. However, it remains an open problem to have a method that is not only practically fast but also is globally convergent. In this paper, we proposed a fast proximal method for solving ℓ0norm based dictionary learning problems, and we proved that the whole sequence generated by the proposed method converges to a stationary point with sub-linear convergence rate. The benefit of having a fast and convergent dictionary learning method is demonstrated in the applications of image recovery and face recognition. Chenglong Bao, Hui Ji 0002, Yuhui Quan, Zuowei Shen |
CVPR | 2 |
| 2014 | A Convergent Incoherent Dictionary Learning Algorithm for Sparse Coding
Chenglong Bao, Yuhui Quan, Hui Ji 0002 |
ECCV (6) | 3 |
| 2013 | Fast Sparsity-Based Orthogonal Dictionary Learning for Image RestorationabstractIn recent years, how to learn a dictionary from input images for sparse modelling has been one very active topic in image processing and recognition. Most existing dictionary learning methods consider an over-complete dictionary, e.g. the K-SVD method. Often they require solving some minimization problem that is very challenging in terms of computational feasibility and efficiency. However, if the correlations among dictionary atoms are not well constrained, the redundancy of the dictionary does not necessarily improve the performance of sparse coding. This paper proposed a fast orthogonal dictionary learning method for sparse image representation. With comparable performance on several image restoration tasks, the proposed method is much more computationally efficient than the over-complete dictionary based learning methods. Chenglong Bao, Jian-Feng Cai 0001, Hui Ji 0002 |
ICCV | 3 |
| 2013 | Recovering Over-/Underexposed Regions in PhotographsabstractWhen taking pictures using a commodity camera in a scene with strong or harsh lighting, such as a sunny day outdoors, we often see a loss of highlight details (overexposure) in some bright regions and a loss of shadow details (underexposure) in some dark regions. In this paper, we develop a wavelet tight frame--based approach to reconstruct a well-exposed image with better visibility of details than that with over-/underexposed regions. There are two modules in the proposed approach: one in lightness channels that inpaints the clipped lightness and adjusts image contrast, and the other in chromatic channels that inpaints the saturated color regions. The experiments show that our method can effectively repair over-/underexposed regions, and it performs better than other existing methods on tested real photographs. Likun Hou, Hui Ji 0002, Zuowei Shen |
SIAM J. Imaging Sci. | 2 |
| 2013 | Wavelet Domain Multifractal Analysis for Static and Dynamic Texture ClassificationabstractIn this paper, we propose a new texture descriptor for both static and dynamic textures. The new descriptor is built on the wavelet-based spatial-frequency analysis of two complementary wavelet pyramids: standard multiscale and wavelet leader. These wavelet pyramids essentially capture the local texture responses in multiple high-pass channels in a multiscale and multiorientation fashion, in which there exists a strong power-law relationship for natural images. Such a power-law relationship is characterized by the so-called multifractal analysis. In addition, two more techniques, scale normalization and multiorientation image averaging, are introduced to further improve the robustness of the proposed descriptor. Combining these techniques, the proposed descriptor enjoys both high discriminative power and robustness against many environmental changes. We apply the descriptor for classifying both static and dynamic textures. Our method has demonstrated excellent performance in comparison with the state-of-the-art approaches in several public benchmark datasets. Hui Ji 0002, Haibin Ling, Yong Xu 0007 |
IEEE Trans. Image Process. | 1 |
| 2012 | Real time robust L1 tracker using accelerated proximal gradient approachabstractRecently sparse representation has been applied to visual tracker by modeling the target appearance using a sparse approximation over a template set, which leads to the so-called L1 trackers as it needs to solve an ℓ1norm related minimization problem for many times. While these L1 trackers showed impressive tracking accuracies, they are very computationally demanding and the speed bottleneck is the solver to ℓ1norm minimizations. This paper aims at developing an L1 tracker that not only runs in real time but also enjoys better robustness than other L1 trackers. In our proposed L1 tracker, a new ℓ1norm related minimization model is proposed to improve the tracking accuracy by adding an ℓ1norm regularization on the coefficients associated with the trivial templates. Moreover, based on the accelerated proximal gradient approach, a very fast numerical solver is developed to solve the resulting ℓ1norm related minimization problem with guaranteed quadratic convergence. The great running time efficiency and tracking accuracy of the proposed tracker is validated with a comprehensive evaluation involving eight challenging sequences and five alternative state-of-the-art trackers. Chenglong Bao, Yi Wu 0001, Haibin Ling, Hui Ji 0002 |
CVPR | 4 |
| 2012 | A two-stage approach to blind spatially-varying motion deblurringabstractMany blind motion deblur methods model the motion blur as a spatially invariant convolution process. However, motion blur caused by the camera movement in 3D space during shutter time often leads to spatially varying blurring effect over the image. In this paper, we proposed an efficient two-stage approach to remove spatially-varying motion blurring from a single photo. There are three main components in our approach: (i) a minimization method of estimating region-wise blur kernels by using both image information and correlations among neighboring kernels, (ii) an interpolation scheme of constructing pixel-wise blur matrix from region-wise blur kernels, and (iii) a non-blind deblurring method robust to kernel errors. The experiments showed that the proposed method outperformed the existing software based approaches on tested real images. Hui Ji 0002 |
CVPR | 1 |
| 2012 | Contour-based recognitionabstractContour is an important cue for object recognition. In this paper, built upon the concept of torque in image space, we propose a new contour-related feature to detect and describe local contour information in images. There are two components for our proposed feature: One is a contour patch detector for detecting image patches with interesting information of object contour, which we call the Maximal/Minimal Torque Patch (MTP) detector. The other is a contour patch descriptor for characterizing a contour patch by sampling the torque values, which we call the Multi-scale Torque (MST) descriptor. Experiments for object recognition on the Caltech-101 dataset showed that the proposed contour feature outperforms other contour-related features and is on a par with many other types of features. When combing our descriptor with the complementary SIFT descriptor, impressive recognition results are observed. Yong Xu 0007, Yuhui Quan, Zhuming Zhang, Hui Ji 0002, Cornelia Fermüller, Morimichi Nishigaki, Daniel DeMenthon |
CVPR | 4 |
| 2012 | Scale-space texture description on SIFT-like textons
Yong Xu 0007, Si-Bin Huang, Hui Ji 0002, Cornelia Fermüller |
Comput. Vis. Image Underst. | 3 |
| 2012 | Framelet-Based Blind Motion Deblurring From a Single ImageabstractHow to recover a clear image from a single motion-blurred image has long been a challenging open problem in digital imaging. In this paper, we focus on how to recover a motion-blurred image due to camera shake. A regularization-based approach is proposed to remove motion blurring from the image by regularizing the sparsity of both the original image and the motion-blur kernel under tight wavelet frame systems. Furthermore, an adapted version of the split Bregman method is proposed to efficiently solve the resulting minimization problem. The experiments on both synthesized images and real images show that our algorithm can effectively remove complex motion blurring from natural images without requiring any prior information of the motion-blur kernel. Jian-Feng Cai 0001, Hui Ji 0002, Chaoqiang Liu, Zuowei Shen |
IEEE Trans. Image Process. | 2 |
| 2012 | Robust Image Deblurring With an Inaccurate Blur KernelabstractMost existing nonblind image deblurring methods assume that the blur kernel is free of error. However, it is often unavoidable in practice that the input blur kernel is erroneous to some extent. Sometimes, the error could be severe, e.g., for images degraded by nonuniform motion blurring. When an inaccurate blur kernel is used as the input, significant distortions will appear in the image recovered by existing methods. In this paper, we present a novel convex minimization model that explicitly takes account of error in the blur kernel. The resulting minimization problem can be efficiently solved by the so-called accelerated proximal gradient method. In addition, a new boundary extension scheme is incorporated in the proposed model to further improve the results. The experiments on both synthesized and real images showed the efficiency and robustness of our algorithm to both the image noise and the model error in the blur kernel. Hui Ji 0002 |
IEEE Trans. Image Process. | 1 |
| 2011 | Dynamic texture classification using dynamic fractal analysisabstractIn this paper, we developed a novel tool called dynamic fractal analysis for dynamic texture (DT) classification, which not only provides a rich description of DT but also has strong robustness to environmental changes. The resulting dynamic fractal spectrum (DFS) for DT sequences consists of two components: One is the volumetric dynamic fractal spectrum component (V-DFS) that captures the stochastic self-similarities of DT sequences as 3D volume datasets; the other is the multi-slice dynamic fractal spectrum component (S-DFS) that encodes fractal structures of DT sequences on 2D slices along different views of the 3D volume. Various types of measures of DT sequences are collected in our approach to analyze DT sequences from different perspectives. The experimental evaluation is conducted on three widely used benchmark datasets. In all the experiments, our method demonstrated excellent performance in comparison with state-of-the-art approaches. Yong Xu 0007, Yuhui Quan, Haibin Ling, Hui Ji 0002 |
ICCV | 4 |
| 2011 | Robust Video Restoration by Joint Sparse and Low Rank Matrix ApproximationabstractThis paper presents a new patch-based video restoration scheme. By grouping similar patches in the spatiotemporal domain, we formulate the video restoration problem as a joint sparse and low-rank matrix approximation problem. The resulting nuclear norm and $\ell_1$ norm related minimization problem can also be efficiently solved by many recently developed numerical methods. The effectiveness of the proposed video restoration scheme is illustrated on two applications: video denoising in the presence of random-valued noise, and video in-painting for archived films. The numerical experiments indicate that the proposed video restoration method compares favorably against many existing algorithms. Hui Ji 0002, Si-Bin Huang, Zuowei Shen, Yuhong Xu |
SIAM J. Imaging Sci. | 1 |
| 2010 | Robust video denoising using low rank matrix completionabstractMost existing video denoising algorithms assume a single statistical model of image noise, e.g. additive Gaussian white noise, which often is violated in practice. In this paper, we present a new patch-based video denoising algorithm capable of removing serious mixed noise from the video data. By grouping similar patches in both spatial and temporal domain, we formulate the problem of removing mixed noise as a low-rank matrix completion problem, which leads to a denoising scheme without strong assumptions on the statistical properties of noise. The resulting nuclear norm related minimization problem can be efficiently solved by many recently developed methods. The robustness and effectiveness of our proposed denoising algorithm on removing mixed noise, e.g. heavy Gaussian noise mixed with impulsive noise, is validated in the experiments and our proposed approach compares favorably against some existing video denoising algorithms. Hui Ji 0002, Chaoqiang Liu, Zuowei Shen, Yuhong Xu |
CVPR | 1 |
| 2010 | Learning shift-invariant sparse representation of actionsabstractA central problem in the analysis of motion capture (MoCap) data is how to decompose motion sequences into primitives. Ideally, a description in terms of primitives should facilitate the recognition, synthesis, and characterization of actions. We propose an unsupervised learning algorithm for automatically decomposing joint movements in human motion capture (MoCap) sequences into shift-invariant basis functions. Our formulation models the time series data of joint movements in actions as a sparse linear combination of short basis functions (snippets), which are executed (or “activated”) at different positions in time. Given a set of MoCap sequences of different actions, our algorithm finds the decomposition of MoCap sequences in terms of basis functions and their activations in time. Using the tools of L1minimization, the procedure alternately solves two large convex minimizations: Given the basis functions, a variant of Orthogonal Matching Pursuit solves for the activations, and given the activations, the Split Bregman Algorithm solves for the basis functions. Experiments demonstrate the power of the decomposition in a number of applications, including action recognition, retrieval, MoCap data compression, and as a tool for classification in the diagnosis of Parkinson (a motion disorder disease). Cornelia Fermüller, Yiannis Aloimonos, Hui Ji 0002 |
CVPR | 4 |
| 2010 | A new texture descriptor using multifractal analysis in multi-orientation wavelet pyramidabstractBased on multifractal analysis in wavelet pyramids of texture images, a new texture descriptor is proposed in this paper that implicitly combines information from both spatial and frequency domains. Beyond the traditional wavelet transform, a multi-oriented wavelet leader pyramid is used in our approach that robustly encodes the multi-scale information of texture edgels. Moreover, the resulting texture model shows empirically a strong power law relationship for nature textures, which can be characterized well by multifractal analysis. Combined with a statistics on affine invariant local patches, our proposed texture descriptor is robust to scale and rotation changes, more general geometrical transforms and illumination variations. In addition, the proposed texture descriptor is computationally efficient since it does not require many expensive processing steps, e.g., texton generation and cross-bin comparisons, which are often used by existing methods. As an application, the proposed descriptor is applied to texture classification and the experimental results on several public texture datasets verified the accuracy and efficiency of our descriptor. Yong Xu 0007, Haibin Ling, Hui Ji 0002 |
CVPR | 4 |
| 2009 | Blind motion deblurring from a single image using sparse approximationabstractRestoring a clear image from a single motion-blurred image due to camera shake has long been a challenging problem in digital imaging. Existing blind deblurring techniques either only remove simple motion blurring, or need user interactions to work on more complex cases. In this paper, we present an approach to remove motion blurring from a single image by formulating the blind blurring as a new joint optimization problem, which simultaneously maximizes the sparsity of the blur kernel and the sparsity of the clear image under certain suitable redundant tight frame systems (curvelet system for kernels and framelet system for images). Without requiring any prior information of the blur kernel as the input, our proposed approach is able to recover high-quality images from given blurred images. Furthermore, the new sparsity constraints under tight frame systems enable the application of a fast algorithm called linearized Bregman iteration to efficiently solve the proposed minimization problem. The experiments on both simulated images and real images showed that our algorithm can effectively removing complex motion blurring from nature images. Jian-Feng Cai 0001, Hui Ji 0002, Chaoqiang Liu, Zuowei Shen |
CVPR | 2 |
| 2009 | High-quality curvelet-based motion deblurring from an image pairabstractOne promising approach to remove motion deblurring is to recover one clear image using an image pair. Existing dual-image methods require an accurate image alignment between the image pair, which could be very challenging even with the help of user interactions. Based on the observation that typical motion-blur kernels will have an extremely sparse representation in the redundant curvelet system, we propose a new minimization model to recover a clear image from the blurred image pair by enhancing the sparsity of blur kernels in the curvelet system. The sparsity prior on the motion-blur kernels improves the robustness of our algorithm to image alignment errors and image formation noise. Also, a numerical method is presented to efficiently solve the resulted minimization problem. The experiments showed that our proposed algorithm is capable of accurately estimating the blur kernels of complex camera motions with low requirement on the accuracy of image alignment, which in turn led to a high-quality recovered image from the blurred image pair. Jian-Feng Cai 0001, Hui Ji 0002, Chaoqiang Liu, Zuowei Shen |
CVPR | 2 |
| 2009 | Combining powerful local and global statistics for texture descriptionabstractA texture descriptor is proposed, which combines local highly discriminative features with the global statistics of fractal geometry to achieve high descriptive power, but also invariance to geometric and illumination transformations. As local measurements SIFT features are estimated densely at multiple window sizes and discretized. On each of the discretized measurements the fractal dimension is computed to obtain the so-called multifractal spectrum, which is invariant to geometric transformations and illumination changes. Finally to achieve robustness to scale changes, a multi-scale representation of the multifractal spectrum is developed using a framelet system, that is, a redundant tight wavelet frame system. Experiments on classification demonstrate that the descriptor outperforms existing methods on the UIUC as well as the UMD high-resolution dataset. Yong Xu 0007, Si-Bin Huang, Hui Ji 0002, Cornelia Fermüller |
CVPR | 3 |
| 2009 | Wavelet Leader multifractal analysis for texture classificationabstractImage classification often relies on texture characterization. Yet texture characterization has so far rarely been based on a true 2D multifractal analysis. Recently, a 2D wavelet Leader based multifractal formalism has been proposed. It allows to perform an accurate, complete and low computational and memory costs multifractal characterization of textures in images. This contribution describes the first application of such a formalism to a real large size (publicly available) image database, consisting of 25 classes of non traditional textures, with 40 high resolution images in each class. Multifractal attributes are estimated from each image and used as classification features within a standard k nearest neighbor classification procedure. The results reported here show that this Leader based multifractal analysis enables the effective discrimination of different textures, as performances in both classification scores and computational costs compare favorably against those of procedures previously proposed in the literature on the same database. Herwig Wendt, Patrice Abry, Stéphane Jaffard, Hui Ji 0002, Zuowei Shen |
ICIP | 4 |
| 2009 | Integrating local feature and global statistics for texture analysisabstractA main challenge for texture analysis is to construct a compact texture descriptor which is not only highly discriminative to intra-class textures, but also robust to inter-class variations, geometric and photometric changes. In this paper, a new texture descriptor is developed by integrating the local affine-invariant texture features and the global viewpoint-invariant statistics. Based on the pixel clustering using two state-ofart robust local texture descriptors (i.e. SIFT and SPIN), the proposed texture descriptor enables impressive invariance to a wide range of environmental changes (e.g. view changes, illumination changes, surface distortions) by characterizing the spatial distribution of pixel sets using multi-fractal analysis. Experiments on some real datasets (publicly available) showed that the proposed texture descriptor achieved better performance than some state-of-art techniques in texture retrieval and texture classification while the computation cost is significantly reduced. Yong Xu 0007, Si-Bin Huang, Hui Ji 0002 |
ICIP | 3 |
| 2009 | Viewpoint Invariant Texture Description Using Fractal Analysis
Yong Xu 0007, Hui Ji 0002, Cornelia Fermüller |
Int. J. Comput. Vis. | 2 |
| 2009 | Robust Wavelet-Based Super-Resolution Reconstruction: Theory and AlgorithmabstractWe present an analysis and algorithm for the problem of super-resolution imaging, that is the reconstruction of HR (high-resolution) images from a sequence of LR (low-resolution) images. Super-resolution reconstruction entails solutions to two problems. One is the alignment of image frames. The other is the reconstruction of a HR image from multiple aligned LR images. Both are important for the performance of super-resolution imaging. Image alignment is addressed with a new batch algorithm, which simultaneously estimates the homographies between multiple image frames by enforcing the surface normal vectors to be the same. This approach can handle longer video sequences quite well. Reconstruction is addressed with a wavelet-based iterative reconstruction algorithm with an efficient denoising scheme. The technique is based on a new analysis of video formation. At a high level our method could be described as a better-conditioned iterative back projection scheme with an efficient regularization criteria in each iteration step. Experiments with both simulated and real data demonstrate that our approach has better performance than existing super-resolution methods. It can remove even large amounts of mixed noise without creating artifacts. Hui Ji 0002, Cornelia Fermüller |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Motion blur identification from image gradientsabstractRestoration of a degraded image from motion blurring is highly dependent on the estimation of the blurring kernel. Most of the existing motion deblurring techniques model the blurring kernel with a shift-invariant box filter, which holds true only if the motion among images is of uniform velocity. In this paper, we present a spectral analysis of image gradients, which leads to a better configuration for identifying the blurring kernel of more general motion types (uniform velocity motion, accelerated motion and vibration). Furthermore, we introduce a hybrid Fourier-Radon transform to estimate the parameters of the blurring kernel with improved robustness to noise over available techniques. The experiments on both simulated images and real images show that our algorithm is capable of accurately identifying the blurring kernel for a wider range of motion types. Hui Ji 0002, Chaoqiang Liu |
CVPR | 1 |
| 2006 | A Projective Invariant for TexturesabstractImage texture analysis has received a lot of attention in the past years. Researchers have developed many texture signatures based on texture measurements, for the purpose of uniquely characterizing the texture. Existing texture signatures, in general, are not invariant to 3D transforms such as view-point changes and non-rigid deformations of the texture surface, which is a serious limitation for many applications. In this paper, we introduce a new texture signature, called the multifractal spectrum (MFS). It provides an efficient framework combining global spatial invariance and local robust measurements. The MFS is invariant under the bi-Lipschitz map, which includes view-point changes and non-rigid deformations of the texture surface, as well as local affine illumination changes. Experiments demonstrate that the MFS captures the essential structure of textures with quite low dimension. Yong Xu 0007, Hui Ji 0002, Cornelia Fermüller |
CVPR (2) | 2 |
| 2006 | Wavelet-Based Super-Resolution Reconstruction: Theory and Algorithm
Hui Ji 0002, Cornelia Fermüller |
ECCV (4) | 1 |
| 2006 | A 3D Shape Constraint on VideoabstractWe propose to combine the information from multiple motion fields by enforcing a constraint on the surface normals (3D shape) of the scene in view. The fact that the shape vectors in the different views are related only by rotation can be formulated as a rank = 3 constraint. This constraint is implemented in an algorithm which solves 3D motion and structure estimation as a practical constrained minimization. Experiments demonstrate its usefulness as a tool in structure from motion providing very accurate estimates of 3D motion. Hui Ji 0002, Cornelia Fermüller |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Integration of Motion Fields through ShapeabstractStructure from motion from single flow fields has been studied intensively, but the integration of information from multiple flow fields has not received much attention. Here we address this problem by enforcing constraints on the shape (surface normals) of the scene in view, as opposed to constraints on the structure (depth). The advantage of integrating shape is two-fold. First, we do not need to estimate feature correspondences over multiple frames, but we only need to match patches. Second, the shape vectors in the different views are related only by rotation. This constraint on shape can be combined easily with motion estimation, thus formulating motion and structure estimation from multiple views as a practical constrained minimization problem using a rank-3 constraint. Based on this constraint, we develop a 3D motion technique, which locates through color and motion segmentation, planar patches in the scene, matches patches over multiple frames, and estimates the motion between multiple frames and the shape of the selected scene patches using the image gradients. Experiments evaluate the accuracy of the 3D motion estimation and demonstrate the motion and shape estimation of the technique by super-resolving an image sequence. Hui Ji 0002, Cornelia Fermüller |
CVPR (2) | 1 |
| 2004 | Bias in Shape Estimation
Hui Ji 0002, Cornelia Fermüller |
ECCV (3) | 1 |