Yunjin Chen

dblp:134/3031 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-4428-2797ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach
abstract
Predicting video saliency is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking database and learning-based approach, particularly tailored for sports videos. In this paper, we establish a large-scale eye-tracking database dubbed audio-visual sports (AVS). AVS consists of 1,000 high-quality sports videos with eye fixations from 60 participants. Through the data analysis on AVS, we observe that human attention patterns exhibit significant variations based on the specific scene context of the sports. Motivated by this, we propose a sport-aware audiovisual saliency model, which can adaptively learn the scene context in a hyper manner. Specifically, a new audio-visual fusion (AVF) block is developed to effectively fuse features from the visual and audio backbone. After that, a hyper network is introduced to learn sport-aware priors, which are then adopted to guide the self-adaptive saliency predictor for predicting saliency map. Experimental results demonstrate that our approach outperforms other state-of-the-art saliency prediction models over the only two sports video eye-tracking databases.
Minglang Qiao, Mai Xu, Shijie Wen, Lai Jiang 0004, Shengxi Li, Yunjin Chen, Leonid Sigal
ICASSP7
2024 HyperSOR: Context-Aware Graph Hypernetwork for Salient Object Ranking
abstract
Salient object ranking (SOR) aims to segment salient objects in an image and simultaneously predict their saliency rankings, according to the shifted human attention over different objects. The existing SOR approaches mainly focus on object-based attention, e.g., the semantic and appearance of object. However, we find that the scene context plays a vital role in SOR, in which the saliency ranking of the same object varies a lot at different scenes. In this paper, we thus make the first attempt towards explicitly learning scene context for SOR. Specifically, we establish a large-scale SOR dataset of 24,373 images with rich context annotations, i.e., scene graphs, segmentation, and saliency rankings. Inspired by the data analysis on our dataset, we propose a novel graph hypernetwork, named HyperSOR, for context-aware SOR. In HyperSOR, an initial graph module is developed to segment objects and construct an initial graph by considering both geometry and semantic information. Then, a scene graph generation module with multi-path graph attention mechanism is designed to learn semantic relationships among objects based on the initial graph. Finally, a saliency ranking prediction module dynamically adopts the learned scene context through a novel graph hypernetwork, for inferring the saliency rankings. Experimental results show that our HyperSOR can significantly improve the performance of SOR.
Minglang Qiao, Mai Xu, Lai Jiang 0004, Shijie Wen, Yunjin Chen, Leonid Sigal
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 DINN360: Deformable Invertible Neural Network for Latitude-aware 360° Image Rescaling
abstract
With the rapid development of virtual reality, 360° images have gained increasing popularity. Their wide field of view necessitates high resolution to ensure image quality. This, however, makes it harder to acquire, store and even process such 360° images. To alleviate this issue, we propose the first attempt at 360° image rescaling, which refers to downscaling a 360° image to a visually valid lowresolution (LR) counterpart and then upscaling to a highresolution (HR) 360° image given the LR variant. Specifically, we first analyze two 360° image datasets and observe several findings that characterize how 360° images typically change along their latitudes. Inspired by these findings, we propose a novel deformable invertible neural network (INN), named DINN360, for latitude-aware 360° image rescaling. In DINN360, a deformable INN is designed to downscale the LR image, and project the high-frequency (HF) component to the latent space by adaptively handling various deformations occurring at different latitude regions. Given the downscaled LR image, the high-quality HR image is then reconstructed in a conditional latitude-aware manner by recovering the structure-related HF component from the latent space. Extensive experiments over four public datasets show that our DINN360 method performs considerably better than other state-of-the-art methods for 2 x, 4 x and 8 x 360° image rescaling.
Mai Xu, Lai Jiang 0004, Leonid Sigal, Yunjin Chen
CVPR5
2023 Optimizing DNN based quality assessment metric for image compression: A novel rate control method
abstract
In the existing coding standards, rate control (RC) plays a critical role in optimally allocating bit-rates to each coding unit, for improving rate-distortion performance under the limited bandwidth. However, the existing RC methods are mainly based on traditional distortion metrics, which fail to take the advantage of the emerging DNN based image quality assessment (IQA) metrics. In this paper, we set up the first attempt to achieve IQA score based RC for image compression. Specifically, a novel visualization based score-distortion (VSD) model and ρ-slope model are proposed to explicitly establish the relationship between IQA score and bit-rates. Then, by solving optimal rate-distortion optimization based on the IQA score, we propose a novel RC method for the HEVC standard. The experimental results show that, given the target bit-rates, the proposed RC method can accurately control the bit-rates and generate the compressed images with higher IQA score and better perceptual quality. More importantly, the proposed RC method is evaluated to be effective over two DNN based IQA metrics and four image datasets, exhibiting the potential in practical use. The code is available at https://github.com/Ffangqy/IQA-RC.
Qiuyue Fang, Lai Jiang 0004, Shengxi Li, Mai Xu, Yunjin Chen, Leonid Sigal
ICME6
2022 Self-supervised Learning for Real-World Super-Resolution from Dual Zoomed Observations
Zhilu Zhang 0001, Ruohao Wang, Yunjin Chen, Wangmeng Zuo
ECCV (18)4
2022 Semi-Supervised Image Deraining Using Knowledge Distillation
abstract
Image deraining has achieved considerable progress based on supervised learning with synthetic training pairs, but is usually limited in handling real-world rainy images. Although semi-supervised methods are suggested to exploit real-world rainy images when training deep deraining models, their performances are still notably inferior. To address this crucial issue, this work proposes a semi-supervised image deraining network with knowledge distillation (SSID-KD) for better exploiting real-world rainy images. In particular, the consistency of feature distribution of rain streaks extracted from synthetic and real-world rainy images is enforced by adopting knowledge distillation. Moreover, as for the backbone in SSID-KD, we propose the multi-scale feature fusion module and the pyramid fusion module to better extract deep features of rainy images. SSID-KD can relieve the problem of over-deraining or under-deraining for real-world rainy images, while it can keep comparable performance with supervised deraining methods on several benchmark datasets. Extensive experiments on both synthetic and real-world rainy images have validated that our SSID-KD not only can achieve better deraining results than existing semi-supervised deraining methods but also are quantitatively comparable with state-of-the-art supervised deraining methods. Benefiting from the well exploration of real-world rainy images, our SSID-KD can obtain more visually plausible deraining results. The source code and trained models are publicly available athttps://github.com/cuiyixin555/SSID-KD.
Cong Wang 0018, Dongwei Ren, Yunjin Chen, Pengfei Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Learned Dynamic Guidance for Depth Image Reconstruction
abstract
The depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate a high resolution RGB camera and exploit the statistical correlation of its data and depth. In recent years, both optimization-based and learning-based approaches have been proposed to deal with the guided depth reconstruction problems. In this paper, we introduce a weighted analysis sparse representation (WASR) model for guided depth image enhancement, which can be considered a generalized formulation of a wide range of previous optimization-based models. We unfold the optimization by the WASR model and conduct guided depth reconstruction with dynamically changed stage-wise operations. Such a guidance strategy enables us to dynamically adjust the stage-wise operations that update the depth image, thus improving the reconstruction quality and speed. To learn the stage-wise operations in a task-driven manner, we propose two parameterizations and their corresponding methods: dynamic guidance with Gaussian RBF nonlinearity parameterization (DG-RBF) and dynamic guidance with CNN nonlinearity parameterization (DG-CNN). The network structures of the proposed DG-RBF and DG-CNN methods are designed with the the objective function of our WASR model in mind and the optimal network parameters are learned from paired training data. Such optimization-inspired network architectures enable our models to leverage the previous expertise as well as take benefit from training data. The effectiveness is validated for guided depth image super-resolution and for realistic depth image reconstruction tasks using standard benchmarks. Our DG-RBF and DG-CNN methods achieve the best quantitative results (RMSE) and better visual quality than the state-of-the-art approaches at the time of writing. The code is available at https://github.com/ShuhangGu/GuidedDepthSR.
Shuhang Gu, Shi Guo, Wangmeng Zuo, Yunjin Chen, Radu Timofte, Luc Van Gool, Lei Zhang 0006
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Intelligent Caching Algorithms in Heterogeneous Wireless Networks with Uncertainty
abstract
A burgeoning number of wireless devices connecting to the Internet tend to impose a heavy traffic load on the network backbone. Caching the most popular content at the heterogeneous wireless network edge is a promising way to alleviate the network overload. However, to cache the diverse content effectively, a file popularity profile that may not be known in advance to network operators has to be utilized. To tackle the challenge caused by this uncertainty, online learning techniques can be considered. Additionally, in practice, dense small-cell networks are often deployed to maximize spectral efficiency, which will naturally bring overlapping coverage areas among individual small cells. In this paper, we propose to address the content caching problem in a scenario of overlapping coverage areas among small cells while further allowing users distributed in the overlapping area to stochastically choose to connect to the small-cell base station they can reach. We propose two effective and efficient online learning algorithms to address the aforementioned problem and also provide theoretical guarantees. Finally, experiments are conducted to verify the performance of the proposed algorithms practically.
Bingshan Hu, Yunjin Chen, Zhiming Huang 0002, Nishant A. Mehta, Jianping Pan 0001
ICDCS2
2018 Learning Generic Diffusion Processes for Image Restoration
Peng Qiao, Yong Dou, Yunjin Chen, WenSen Feng
BMVC3
2018 Fast and Accurate Poisson Denoising With Trainable Nonlinear Diffusion
abstract
The degradation of the acquired signal by Poisson noise is a common problem for various imaging applications, such as medical imaging, night vision, and microscopy. Up to now, many state-of-the-art Poisson denoising techniques mainly concentrate on achieving utmost performance, with little consideration for the computation efficiency. Therefore, in this paper we aim to propose an efficient Poisson denoising model with both high computational efficiency and recovery quality. To this end, we exploit the newly developed trainable nonlinear reaction diffusion (TNRD) model which has proven an extremely fast image restoration approach with performance surpassing recent state-of-the-arts. However, the straightforward direct gradient descent employed in the original TNRD-based denoising task is not applicable in this paper. To solve this problem, we resort to the proximal gradient descent method. We retrain the model parameters, including the linear filters and influence functions by taking into account the Poisson noise statistics, and end up with a well-trained nonlinear diffusion model specialized for Poisson denoising. The trained model provides strongly competitive results against state-of-the-art approaches, meanwhile bearing the properties of simple structure and high efficiency. Furthermore, our proposed model comes along with an additional advantage, that the diffusion process is well-suited for parallel computation on graphics processing units (GPUs). For images of size , our GPU implementation takes less than 0.1 s to produce state-of-the-art Poisson denoising performance.
WenSen Feng, Peng Qiao, Yunjin Chen
IEEE Trans. Cybern.3
2018 LEARN: Learned Experts' Assessment-Based Reconstruction Network for Sparse-Data CT
abstract
Compressive sensing (CS) has proved effective for tomographic reconstruction from sparsely collected data or under-sampled measurements, which are practically important for few-view computed tomography (CT), tomosynthesis, interior tomography, and so on. To perform sparse-data CT, the iterative reconstruction commonly uses regularizers in the CS framework. Currently, how to choose the parameters adaptively for regularization is a major open problem. In this paper, inspired by the idea of machine learning especially deep learning, we unfold the state-of-the-art "fields of experts"-based iterative reconstruction scheme up to a number of iterations for data-driven training, construct a learned experts' assessment-based reconstruction network (LEARN) for sparse-data CT, and demonstrate the feasibility and merits of our LEARN network. The experimental results with our proposed LEARN network produces a superior performance with the well-known Mayo Clinic low-dose challenge data set relative to the several state-of-the-art methods, in terms of artifact reduction, feature preservation, and computational speed. This is consistent to our insight that because all the regularization terms and parameters used in the iterative reconstruction are now learned from the training data, our LEARN network utilizes application-oriented knowledge more effectively and recovers underlying images more favorably than competing algorithms. Also, the number of layers in the LEARN network is only 50, reducing the computational complexity of typical iterative algorithms by orders of magnitude.
Hu Chen 0002, Yi Zhang 0018, Yunjin Chen, Huaiqiang Sun, Yang Lu 0011, Peixi Liao, Jiliu Zhou, Ge Wang 0001
IEEE Trans. Medical Imaging3
2017 Correlation Filter Tracking: Beyond an Open-loop System
Qingyong Hu, Yulan Guo, Yunjin Chen, Wei An 0003
BMVC3
2017 Learning Dynamic Guidance for Depth Image Enhancement
abstract
The depth images acquired by consumer depth sensors (e.g., Kinect and ToF) usually are of low resolution and insufficient quality. One natural solution is to incorporate with high resolution RGB camera for exploiting their statistical correlation. However, most existing methods are intuitive and limited in characterizing the complex and dynamic dependency between intensity and depth images. To address these limitations, we propose a weighted analysis representation model for guided depth image enhancement, which advances the conventional methods in two aspects: (i) task driven learning and (ii) dynamic guidance. First, we generalize the analysis representation model by including a guided weight function for dependency modeling. And the task-driven learning formulation is introduced to obtain the optimized guidance tailored to specific enhancement task. Second, the depth image is gradually enhanced along with the iterations, and thus the guidance should also be dynamically adjusted to account for the updating of depth image. To this end, stage-wise parameters are learned for dynamic guidance. Experiments on guided depth image upsampling and noisy depth image restoration validate the effectiveness of our method.
Shuhang Gu, Wangmeng Zuo, Shi Guo, Yunjin Chen, Chongyu Chen, Lei Zhang 0006
CVPR4
2017 Learning Non-local Image Diffusion for Image Denoising
abstract
Image diffusion plays a fundamental role for the task of image denoising. The recently proposed trainable nonlinear reaction diffusion (TNRD) model defines a simple but very effective framework for image denoising. However, as the TNRD model is a local model, whose diffusion behavior is purely controlled by information of local patches, it is prone to create artifacts in the homogenous regions and over-smooth highly textured regions, especially in the case of strong noise levels. Meanwhile, it is widely known that the non-local self-similarity (NSS) prior stands as an effective image prior for image denoising, which has been widely exploited in many non-local methods. In this work, we are highly motivated to embed the NSS prior into the TNRD model to tackle its weaknesses. In order to preserve the expected property that end-to-end training remains available, we exploit the NSS prior by defining a set of non-local filters, and derive our proposed trainable non-local reaction diffusion (TNLRD) model for image denoising. Together with the local filters and influence functions, the non-local filters are learned by employing loss-specific training. The experimental results show that the trained TNLRD model produces visually plausible recovered images with more textures and less artifacts, compared to its local versions. Moreover, the trained TNLRD model can achieve strongly competitive performance to recent state-of-the-art image denoising methods in terms of peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM).
Peng Qiao, Yong Dou, WenSen Feng, Rongchun Li, Yunjin Chen
ACM Multimedia5
2017 Trainable Nonlinear Reaction Diffusion: A Flexible Framework for Fast and Effective Image Restoration
abstract
Image restoration is a long-standing problem in low-level computer vision with many interesting applications. We describe a flexible learning framework based on the concept of nonlinear reaction diffusion models for various image restoration problems. By embodying recent improvements in nonlinear diffusion models, we propose a dynamic nonlinear reaction diffusion model with time-dependent parameters (i.e., linear filters and influence functions). In contrast to previous nonlinear diffusion models, all the parameters, including the filters and the influence functions, are simultaneously learned from training data through a loss based approach. We call this approach TNRD-Trainable Nonlinear Reaction Diffusion. The TNRD approach is applicable for a variety of image restoration tasks by incorporating appropriate reaction force. We demonstrate its capabilities with three representative applications, Gaussian image denoising, single image super resolution and JPEG deblocking. Experiments show that our trained nonlinear diffusion models largely benefit from the training of the parameters and finally lead to the best reported performance on common test datasets for the tested applications. Our trained models preserve the structural simplicity of diffusion models and take only a small number of diffusion steps, thus are highly efficient. Moreover, they are also well-suited for parallel computation on GPUs, which makes the inference procedure extremely fast.
Yunjin Chen, Thomas Pock
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Image Denoising via Multiscale Nonlinear Diffusion Models
abstract
Image denoising is a fundamental operation in image processing and holds considerable practical importance for various real-world applications. Arguably several thousands of papers are dedicated to image denoising. In the past decade, state-of-the-art denoising algorithms have been clearly dominated by nonlocal patch-based methods, which explicitly exploit patch self-similarity within the targeted image. However, in the past two years, discriminatively trained local approaches have started to outperform previous nonlocal models and have been attracting increasing attention due to the additional advantage of computational efficiency. Successful approaches include cascade of shrinkage fields (CSF) and trainable nonlinear reaction diffusion (TNRD). These two methods are built on the filter response of linear filters of small size using feed forward architectures. Due to the locality inherent in local approaches, the CSF and TNRD models become less effective when the noise level is high and consequently introduce some noise artifacts. In order to overcome this problem, in this paper we introduce a multiscale strategy. To be specific, we build on our newly developed TNRD model, adopting the multiscale pyramid image representation to devise a multiscale nonlinear diffusion process. As expected, all the parameters in the proposed multiscale diffusion model, including the filters and the influence functions across scales, are learned from training data through a loss-based approach. Numerical results on Gaussian and Poisson denoising substantiate that the exploited multiscale strategy can successfully boost the performance of the original TNRD model with a single scale. As a consequence, the resulting multiscale diffusion models can significantly suppress the typical incorrect features for those noisy images with heavy noise. It turns out that multiscale TNRD variants achieve better performance than state-of-the-art denoising methods.
WenSen Feng, Peng Qiao, Xuanyang Xi, Yunjin Chen
SIAM J. Imaging Sci.4
2017 Variational JPEG artifacts suppression based on high-order MRFs
Yunjin Chen
Signal Process. Image Commun.1
2017 Variational single image interpolation with time-varying regularization
Peng Qiao, Yunjin Chen, Yong Dou
Signal Process. Image Commun.2
2017 Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising
abstract
The discriminative model learning for image denoising has been recently attracting considerable attentions due to its favorable denoising performance. In this paper, we take one step forward by investigating the construction of feed-forward denoising convolutional neural networks (DnCNNs) to embrace the progress in very deep architecture, learning algorithm, and regularization method into image denoising. Specifically, residual learning and batch normalization are utilized to speed up the training process as well as boost the denoising performance. Different from the existing discriminative denoising models which usually train a specific model for additive white Gaussian noise at a certain noise level, our DnCNN model is able to handle Gaussian denoising with unknown noise level (i.e., blind Gaussian denoising). With the residual learning strategy, DnCNN implicitly removes the latent clean image in the hidden layers. This property motivates us to train a single DnCNN model to tackle with several general image denoising tasks, such as Gaussian denoising, single image super-resolution, and JPEG image deblocking. Our extensive experiments demonstrate that our DnCNN model can not only exhibit high effectiveness in several general image denoising tasks, but also be efficiently implemented by benefiting from GPU computing.
Kai Zhang 0008, Wangmeng Zuo, Yunjin Chen, Deyu Meng, Lei Zhang 0006
IEEE Trans. Image Process.3
2016 Higher-order MRFs based image super resolution: why not MAP?
abstract
A trainable filter‐based higher‐order Markov random fields model – the so called fields of experts (FoE), has proved a highly effective image prior model for many classic image restoration problems. Generally, two options are available to incorporate the learned FoE prior in the inference procedure: (i) sampling‐based minimum mean square error (MMSE) estimate, and (ii) energy minimisation‐based maximum a posteriori (MAP) estimate. This study is devoted to the FoE prior based single image super resolution (SR) problem, and the author suggest to make use of the MAP estimate for inference based on two facts: (i) It is well‐known that the MAP inference has a remarkable advantage of high computational efficiency, while the sampling‐based MMSE estimate is very time consuming. (ii) Practical SR experiment results demonstrate that the MAP estimate works equally well compared with the MMSE estimate with exactly the same FoE prior model. Moreover, it can lead to even further improvements by incorporating the discriminatively trained FoE prior model. In summary, the author hold that for higher‐order natural image prior based SR problem, it is better to employ the MAP estimate for inference.
Yunjin Chen
IET Image Process.1
2016 Poisson Noise Reduction with Higher-Order Natural Image Prior Model
abstract
Poisson denoising is an essential issue for various imaging applications, such as night vision, medical imaging, and microscopy. State-of-the-art approaches are clearly dominated by patch-based non-local methods in recent years. In this paper, we aim to propose a local Poisson denoising model with both structural simplicity and good performance. To this end, we consider a variational modeling to integrate the so-called fields of experts (FoE) image prior, that has proven an effective higher-order Markov random fields model for many classic image restoration problems. We exploit several feasible variational variants for this task. We start with a direct modeling in the original image domain by taking into account the Poisson noise statistics, which performs generally well for the cases of high signal-to-noise ratio (SNR). However, this strategy encounters problem in cases of low SNR. Then we turn to an alternative modeling strategy by using the Anscombe transform and Gaussian statistics derived data term. We retrain the FoE prior model directly in the transform domain. With the newly trained FoE model, we end up with a local variational model providing strongly competitive results against state-of-the-art nonlocal approaches, meanwhile bearing the property of simple structure. Furthermore, our proposed model comes along with an additional advantage, that the inference is very efficient as it is well suited for parallel computation on GPUs. For images of size $512 \times 512$, our GPU implementation takes less than 1 second to produce state-of-the-art Poisson denoising performance.
WenSen Feng, Hong Qiao, Yunjin Chen
SIAM J. Imaging Sci.3
2016 A New Algorithm for Optimizing TV-Based PolSAR Despeckling Model
abstract
The Wishart fidelity and total variation (TV) based variational model (WisTV) with the positive definite (PD) constraint has shown to be effective for the whole PolSAR covariance data speckle reduction. However, the existing algorithms for solving the WisTV model only give approximation solutions by projecting the results onto the set of PD matrices, and their parameters depend strongly on the data. The purpose of this letter is to propose a new optimization algorithm to address the issues. To keep the uniformity of the parameters for different PolSAR data, a sigmoid function-based normalization method is designed, which ensures the applicability of the WisTV model for the normalized data. By using the orthogonal decomposition of the PD variables, the WisTV model is converted into an unconstrained optimization problem which is further transformed into a multivariable problem based on the equivalent representations of the trace and logdet functions. The alternative minimization technique is then utilized to solve the final optimization problem. The subproblems for each individual variable are convex and their solutions have explicit expressions. Moreover, the computational complexity of the algorithm is discussed. Experimental results on both synthetic and real PolSAR data demonstrate the validity of the proposed algorithm.
Xiangli Nie, Bo Zhang 0006, Yunjin Chen, Hong Qiao
IEEE Signal Process. Lett.3
2015 On learning optimized reaction diffusion processes for effective image restoration
abstract
For several decades, image restoration remains an active research topic in low-level computer vision and hence new approaches are constantly emerging. However, many recently proposed algorithms achieve state-of-the-art performance only at the expense of very high computation time, which clearly limits their practical relevance. In this work, we propose a simple but effective approach with both high computational efficiency and high restoration quality. We extend conventional nonlinear reaction diffusion models by several parametrized linear filters as well as several parametrized influence functions. We propose to train the parameters of the filters and the influence functions through a loss based approach. Experiments show that our trained nonlinear reaction diffusion models largely benefit from the training of the parameters and finally lead to the best reported performance on common test datasets for image restoration. Due to their structural simplicity, our trained models are highly efficient and are also well-suited for parallel computation on GPUs.
Yunjin Chen, Wei Yu 0014, Thomas Pock
CVPR1
2014 iPiano: Inertial Proximal Algorithm for Nonconvex Optimization
abstract
In this paper we study an algorithm for solving a minimization problem composed of a differentiable (possibly nonconvex) and a convex (possibly nondifferentiable) function. The algorithm iPiano combines forward-backward splitting with an inertial force. It can be seen as a nonsmooth split version of the Heavy-ball method from Polyak. A rigorous analysis of the algorithm for the proposed class of problems yields global convergence of the function values and the arguments. This makes the algorithm robust for usage on nonconvex problems. The convergence result is obtained based on the Kurdyka--Łojasiewicz inequality. This is a very weak restriction, which was used to prove convergence for several other gradient methods. First, an abstract convergence theorem for a generic algorithm is proved, and then iPiano is shown to satisfy the requirements of this theorem. Furthermore, a convergence rate is established for the general problem class. We demonstrate iPiano on computer vision problems---image denoising with learned priors and diffusion based image compression.
Peter Ochs, Yunjin Chen, Thomas Brox, Thomas Pock
SIAM J. Imaging Sci.2
2014 A Higher-Order MRF Based Variational Model for Multiplicative Noise Reduction
abstract
The Fields of Experts (FoE) image prior model, a filter-based higher-order Markov Random Fields (MRF) model, has been shown to be effective for many image restoration problems. Motivated by the successes of FoE-based approaches, in this letter we propose a novel variational model for multiplicative noise reduction based on the FoE image prior model. The resulting model corresponds to a non-convex minimization problem, which can be efficiently solved by a recently published non-convex optimization algorithm. Experimental results based on synthetic speckle noise and real synthetic aperture radar (SAR) images suggest that the performance of our proposed method is on par with the best published despeckling algorithm. Besides, our proposed model comes along with an additional advantage, that the inference is extremely efficient. Our GPU based implementation takes less than 1s to produce state-of-the-art despeckling performance.
Yunjin Chen, WenSen Feng, René Ranftl, Hong Qiao, Thomas Pock
IEEE Signal Process. Lett.1
2014 Insights Into Analysis Operator Learning: From Patch-Based Sparse Models to Higher Order MRFs
abstract
This paper addresses a new learning algorithm for the recently introduced co-sparse analysis model. First, we give new insights into the co-sparse analysis model by establishing connections to filter-based MRF models, such as the field of experts model of Roth and Black. For training, we introduce a technique called bi-level optimization to learn the analysis operators. Compared with existing analysis operator learning approaches, our training procedure has the advantage that it is unconstrained with respect to the analysis operator. We investigate the effect of different aspects of the co-sparse analysis model and show that the sparsity promoting function (also called penalty function) is the most important factor in the model. In order to demonstrate the effectiveness of our training approach, we apply our trained models to various classical image restoration problems. Numerical experiments show that our trained models clearly outperform existing analysis operator learning approaches and are on par with state-of-the-art image denoising algorithms. Our approach develops a framework that is intuitive to understand and easy to implement.
Yunjin Chen, René Ranftl, Thomas Pock
IEEE Trans. Image Process.1