VLDB 2026 Research / reviewers in the wild / expert
Peyman Milanfar
dblp:48/6882
· DBLP profile ↗
127ranked-venue papers
12as first author
32since 2021 · last 2025
0000-0003-1455-7662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 111 · 11 first-author · 27 since 2021Artificial intelligence and machine learning · 34 · 1 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Power of Context: How Multimodality Improves Image Super-ResolutionabstractSingle-image super-resolution (SISR) remains challenging due to the inherent difficulty of recovering fine-grained details and preserving perceptual quality from low-resolution inputs. Existing methods often rely on limited image priors, leading to suboptimal results. We propose a novel approach that leverages the rich contextual information available in multiple modalities - including depth, segmentation, edges, and text prompts-to learn a powerful generative prior for SISR within a diffusion model framework. We introduce a flexible network architecture that effectively fuses multimodal information, accommodating an arbitrary number of input modalities without requiring significant modifications to the diffusion process. Crucially, we mitigate hallucinations, often introduced by text prompts, by using spatial information from other modalities to guide regional text-based conditioning. Each modality’s guidance strength can also be controlled independently, allowing steering outputs toward different directions, such as increasing bokeh through depth or adjusting object prominence via segmentation. Extensive experiments demonstrate that our model surpasses state-of-the-art generative SISR methods, achieving superior visual quality and fidelity. Kangfu Mei, Hossein Talebi, Mojtaba Ardakani, Vishal M. Patel, Peyman Milanfar, Mauricio Delbracio |
CVPR | 5 |
| 2025 | UniRes: Universal Image Restoration for Complex DegradationsabstractReal-world image restoration is hampered by diverse degradations stemming from varying capture conditions, capture devices and post-processing pipelines. Existing works make improvements through simulating those degradations and leveraging image generative priors, however generalization to in-the-wild data remains an unresolved problem. In this paper, we focus on complex degradations, i.e., arbitrary mixtures of multiple types of known degradations, which is frequently seen in the wild. A simple yet flexible diffusionbased framework, named UniRes, is proposed to address such degradations in an end-to-end manner. It combines several specialized models during the diffusion sampling steps, hence transferring the knowledge from several well-isolated restoration tasks to the restoration of complex in-the-wild degradations. This only requires well-isolated training data for several degradation types. The framework is flexible as extensions can be added through a unified formulation, and the fidelity-quality trade-off can be adjusted through a new paradigm. Our proposed method is evaluated on both complex-degradation and single-degradation image restoration datasets. Extensive qualitative and quantitative experimental results show consistent performance gain especially for images with complex degradations. Keren Ye, Mauricio Delbracio, Peyman Milanfar, Vishal M. Patel, Hossein Talebi |
ICCV | 4 |
| 2025 | Stochastic Deep Restoration Priors for Imaging Inverse ProblemsabstractDeep neural networks trained as image denoisers are widely used as priors for solving imaging inverse problems. We introduce Stochastic deep Restoration Priors (ShaRP), a novel framework that stochastically leverages an ensemble of deep restoration models beyond denoisers to regularize inverse problems. By using generalized restoration models trained on a broad range of degradations beyond simple Gaussian noise, ShaRP effectively addresses structured artifacts and enables self-supervised training without fully sampled data. We prove that ShaRP minimizes an objective function involving a regularizer derived from the score functions of minimum mean square error (MMSE) restoration operators. We also provide theoretical guarantees for learning restoration operators from incomplete measurements. ShaRP achieves state-of-the-art performance on tasks such as magnetic resonance imaging reconstruction and single-image super-resolution, surpassing both denoiser- and diffusion-model-based methods without requiring retraining. Albert Peng, Weijie Gan, Peyman Milanfar, Mauricio Delbracio, Ulugbek Kamilov |
ICML | 4 |
| 2025 | Kernel Density Steering: Inference-Time Scaling via Mode Seeking for Image RestorationabstractDiffusion models show promise for image restoration, but existing methods often struggle with inconsistent fidelity and undesirable artifacts. To address this, we introduce Kernel Density Steering (KDS), a novel inference-time framework promoting robust, high-fidelity outputs through explicit local mode-seeking. KDS employs an $N$-particle ensemble of diffusion samples, computing patch-wise kernel density estimation gradients from their collective outputs. These gradients steer patches in each particle towards shared, higher-density regions identified within the ensemble. This collective local mode-seeking mechanism, acting as "collective wisdom", steers samples away from spurious modes prone to artifacts, arising from independent sampling or model imperfections, and towards more robust, high-fidelity structures. This allows us to obtain better quality samples at the expense of higher compute by simultaneously sampling multiple particles. As a plug-and-play framework, KDS requires no retraining or external verifiers, seamlessly integrating with various diffusion samplers. Extensive numerical validations demonstrate KDS substantially improves both quantitative and qualitative performance on challenging real-world super-resolution and image inpainting tasks. Kangfu Mei, Mojtaba Sahraee-Ardakan, Ulugbek Kamilov, Peyman Milanfar, Mauricio Delbracio |
NeurIPS | 5 |
| 2025 | DVMark: A Deep Multiscale Framework for Video WatermarkingabstractVideo watermarking embeds a message into a cover video in an imperceptible manner, which can be retrieved even if the video undergoes certain modifications or distortions. Traditional watermarking methods are often manually designed for particular types of distortions and thus cannot simultaneously handle a broad spectrum of distortions. To this end, we propose a robust deep learning-based solution for video watermarking that is end-to-end trainable. Our model consists of a novel multiscale design where the watermarks are distributed across multiple spatial-temporal scales. Extensive evaluations on a wide variety of distortions show that our method outperforms traditional video watermarking methods as well as deep image watermarking models by a large margin. We further demonstrate the practicality of our method on a realistic video-editing application. Xiyang Luo, Yinxiao Li, Huiwen Chang, Ce Liu 0001, Peyman Milanfar, Feng Yang 0008 |
IEEE Trans. Image Process. | 5 |
| 2024 | CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image GenerationabstractLarge generative diffusion models have revolution-ized text-to-image generation and offer immense po-tential for conditional generation tasks such as im-age enhancement, restoration, editing, and compositing. However, their widespread adoption is hindered by the high computational cost, which limits their real-time application. To address this challenge, we in-troduce a novel method dubbed CoDi, that adapts a pre-trained latent diffusion model to accept additional image conditioning inputs while significantly reducing the sampling steps required to achieve high-quality results. Our method can leverage architectures such as ControlNet to incorporate conditioning inputs with-out compromising the model's prior knowledge gained during large scale pre-training. Additionally, a con-ditional consistency loss enforces consistent predictions across diffusion steps, effectively compelling the model to generate high-quality images with conditions in a few steps. Our conditional-task learning and distil-lation approach outperforms previous distillation meth-ods, achieving a new state-of-the-art in producing high-quality images with very few steps (e.g., 1–4) across multiple tasks, including super-resolution, text-guided image editing, and depth-to-image generation. Kangfu Mei, Mauricio Delbracio, Hossein Talebi, Zhengzhong Tu, Vishal M. Patel, Peyman Milanfar |
CVPR | 6 |
| 2024 | SPIRE: Semantic Prompt-Driven Image Restoration
Zhengzhong Tu, Keren Ye, Mauricio Delbracio, Peyman Milanfar, Qifeng Chen 0001, Hossein Talebi |
ECCV (40) | 5 |
| 2024 | ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling
William Yicheng Zhu, Keren Ye, Junjie Ke, Leonidas J. Guibas, Peyman Milanfar, Feng Yang 0008 |
ECCV (27) | 6 |
| 2024 | A Restoration Network as an Implicit PriorabstractImage denoisers have been shown to be powerful priors for solving inverse problems in imaging. In this work, we introduce a generalization of these methods that allows any image restoration network to be used as an implicit prior. The proposed method uses priors specified by deep neural networks pre-trained as general restoration operators. The method provides a principled approach for adapting state-of-the-art restoration models for other inverse problems. Our theoretical result analyzes its convergence to a stationary point of a global functional associated with the restoration operator. Numerical results show that the method using a super-resolution prior achieves state-of-the-art performance both quantitatively and qualitatively. Overall, this work offers a step forward for solving inverse problems by enabling the use of powerful pre-trained restoration models as priors. Mauricio Delbracio, Peyman Milanfar, Ulugbek Kamilov |
ICLR | 3 |
| 2024 | Prompt-tuning Latent Diffusion Models for Inverse ProblemsabstractWe propose a new method for solving imaging inverse problems using text-to-image latent diffusion models as general priors. Existing methods using latent diffusion models for inverse problems typically rely on simple null text prompts, which can lead to suboptimal performance. To improve upon this, we introduce a method for prompt tuning, which jointly optimizes the text embedding on-the-fly while running the reverse diffusion. This allows us to generate images that are more faithful to the diffusion prior. Specifically, our approach involves a unified optimization framework that simultaneously considers the prompt, latent, and pixel values through alternating minimization. This significantly diminishes image artifacts - a major problem when using latent diffusion models instead of pixel-based diffusion ones. Our method, called P2L, outperforms both pixel- and latent-diffusion model-based inverse problem solvers on a variety of tasks, such as super-resolution, deblurring, and inpainting. Furthermore, P2L demonstrates remarkable scalability to higher resolutions without artifacts. Hyungjin Chung, Jong Chul Ye, Peyman Milanfar, Mauricio Delbracio |
ICML | 3 |
| 2023 | VILA: Learning Image Aesthetics from User Comments with Vision-Language PretrainingabstractAssessing the aesthetics of an image is challenging, as it is influenced by multiple factors including composition, color, style, and high-level semantics. Existing image aesthetic assessment (IAA) methods primarily rely on human-labeled rating scores, which oversimplify the visual aesthetic information that humans perceive. Conversely, user comments offer more comprehensive information and are a more natural way to express human opinions and preferences regarding image aesthetics. In light of this, we propose learning image aesthetics from user comments, and exploring vision-language pretraining methods to learn multimodal aesthetic representations. Specifically, we pretrain an image-text encoder-decoder model with image-comment pairs, using contrastive and generative objectives to learn rich and generic aesthetic semantics without human labels. To efficiently adapt the pretrained model for downstream IAA tasks, we further propose a lightweight rank-based adapter that employs text as an anchor to learn the aesthetic ranking concept. Our results show that our pretrained aesthetic vision-language model outperforms prior works on image aesthetic captioning over the AVA-Captions dataset, and it has powerful zero-shot capability for aesthetic tasks such as zero-shot style classification and zero-shot IAA, surpassing many supervised baselines. With only minimal finetuning parameters using the proposed adapter module, our model achieves state-of-the-art IAA performance over the AVA dataset.11Our model is available at https://github.com/google-research/google-research/tree/master/VILA Junjie Ke, Keren Ye, Peyman Milanfar, Feng Yang 0008 |
CVPR | 5 |
| 2023 | SVDiff: Compact Parameter Space for Diffusion Fine-TuningabstractDiffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts or other modalities. However, existing methods for customizing these models are limited by handling multiple personalized subjects and the risk of overfitting. Moreover, their large number of parameters is inefficient for model storage. In this paper, we propose a novel approach to address these limitations in existing text-to-image diffusion models for personalization. Our method involves fine-tuning the singular values of the weight matrices, leading to a compact and efficient parameter space that reduces the risk of overfitting and language-drifting. We also propose a Cut-Mix-Unmix data-augmentation technique to enhance the quality of multi-subject image generation and a simple text-based image editing framework. Our proposed SVDiff method has a significantly smaller model size compared to existing methods (≈2,200 times fewer parameters compared with vanilla DreamBooth), making it more practical for real-world applications. Ligong Han, Yinxiao Li, Han Zhang 0010, Peyman Milanfar, Dimitris N. Metaxas, Feng Yang 0008 |
ICCV | 4 |
| 2023 | Multiscale Structure Guided Diffusion for Image DeblurringabstractDiffusion Probabilistic Models (DPMs) have recently been employed for image deblurring, formulated as an image-conditioned generation process that maps Gaussian noise to the high-quality image, conditioned on the blurry input. Image-conditioned DPMs (icDPMs) have shown more realistic results than regression-based methods when trained on pairwise in-domain data. However, their robustness in restoring images is unclear when presented with out-of-domain images as they do not impose specific degradation models or intermediate constraints. To this end, we introduce a simple yet effective multiscale structure guidance as an implicit bias that informs the icDPM about the coarse structure of the sharp image at the intermediate layers. This guided formulation leads to a significant improvement of the deblurring results, particularly on unseen domain. The guidance is extracted from the latent space of a regression network trained to predict the clean-sharp target at multiple lower resolutions, thus maintaining the most salient sharp structures. With both the blurry input and multiscale guidance, the icDPM model can better understand the blur and recover the clean image. We evaluate a single-dataset trained model on diverse datasets and demonstrate more robust deblurring results with fewer artifacts on unseen data. Our method outperforms existing baselines, achieving state-of-the-art perceptual quality while keeping competitive distortion metrics. Mengwei Ren, Mauricio Delbracio, Hossein Talebi, Guido Gerig, Peyman Milanfar |
ICCV | 5 |
| 2023 | MULLER: Multilayer Laplacian Resizer for VisionabstractImage resizing operation is a fundamental preprocessing module in modern computer vision. Throughout the deep learning revolution, researchers have overlooked the potential of alternative resizing methods beyond the commonly used resizers that are readily available, such as nearest-neighbors, bilinear, and bicubic. The key question of our interest is whether the front-end resizer affects the performance of deep vision models? In this paper, we present an extremely lightweight multilayer Laplacian resizer with only a handful of trainable parameters, dubbed MULLER resizer. MULLER has a bandpass nature in that it learns to boost details in certain frequency subbands that benefit the downstream recognition models. We show that MULLER can be easily plugged into various training pipelines, and it effectively boosts the performance of the underlying vision task with little to no extra cost. Specifically, we select a state-of-the-art vision Transformer, MaxViT [50], as the baseline, and show that, if trained with MULLER, MaxViT gains up to 0.6% top-1 accuracy, and meanwhile enjoys 36% inference cost saving to achieve similar top-1 accuracy on ImageNet-1k, as compared to the standard training scheme. Notably, MULLER’s performance also scales with model size and training data size such as ImageNet-21k and JFT, and it is widely applicable to multiple vision tasks, including image classification, object detection and segmentation, as well as image quality assessment. The code is available at https://github.com/google-research/google-research/tree/master/muller. Zhengzhong Tu, Peyman Milanfar, Hossein Talebi |
ICCV | 2 |
| 2022 | MAXIM: Multi-Axis MLP for Image ProcessingabstractRecent progress on Transformers and multilayer perceptron (MLP) models provide new network architectural designs for computer vision tasks. Although these models proved to be effective in many vision tasks such as image recognition, there remain challenges in adapting them for lowlevel vision. The inflexibility to support high-resolution images and limitations of local attention are perhaps the main bottlenecks. In this work, we present a multi-axis MLP based architecture called MAXIM, that can serve as an efficient and flexible general-purpose vision backbone for image processing tasks. MAXIM uses a UNet-shaped hierarchical structure and supports long-range interactions enabled by spatially-gated MLPs. Specifically, MAXIM contains two MLP-based building blocks: a multi-axis gated MLP that allows for efficient and scalable spatial mixing of local and global visual cues, and a cross-gating block, an alternative to cross-attention, which accounts for cross-feature conditioning. Both these modules are exclusively based on MLPs, but also benefit from being both global and ‘fully-convolutional’, two properties that are desirable for image processing. Our extensive experimental results show that the proposed MAXIM model achieves state-of-the-art performance on more than ten benchmarks across a range of image processing tasks, including denoising, deblurring, de raining, dehazing, and enhancement while requiring fewer or comparable numbers of parameters and FLOPs than competitive models. The source code and trained models will be available at https://github.com/google-research/maxim. Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li |
CVPR | 5 |
| 2022 | Deblurring via Stochastic Refinement
Jay Whang, Mauricio Delbracio, Hossein Talebi, Chitwan Saharia, Alexandros G. Dimakis, Peyman Milanfar |
CVPR | 6 |
| 2022 | Deep 3D-to-2D Watermarking: Embedding Messages in 3D Meshes and Extracting Them from 2D RenderingsabstractDigital watermarking is widely used for copyright protection. Traditional 3D watermarking approaches or commercial software are typically designed to embed messages into 3D meshes, and later retrieve the messages directly from distorted/undistorted watermarked 3D meshes. However, in many cases, users only have access to rendered 2D images instead of 3D meshes. Unfortunately, retrieving messages from 2D renderings of 3D meshes is still challenging and underexplored. We introduce a novel end-to-end learning framework to solve this problem through: 1) an encoder to covertly embed messages in both mesh geometry and textures; 2) a differentiable renderer to render watermarked 3D objects from different camera angles and under varied lighting conditions; 3) a decoder to recover the messages from 2D rendered images. From our experiments, we show that our model can learn to embed information visually imperceptible to humans, and to retrieve the embedded information from 2D renderings that undergo 3D distortions. In addition, we demonstrate that our method can also work with other renderers, such as ray tracers and real-time renderers with and without fine-tuning. Innfarn Yoo, Huiwen Chang, Xiyang Luo, Ondrej Stava, Ce Liu 0001, Peyman Milanfar, Feng Yang 0008 |
CVPR | 6 |
| 2022 | MaxViT: Multi-axis Vision Transformer
Zhengzhong Tu, Hossein Talebi, Han Zhang 0010, Feng Yang 0008, Peyman Milanfar, Alan C. Bovik, Yinxiao Li |
ECCV (24) | 5 |
| 2022 | Interpretable Unsupervised Diversity Denoising and Artefact Removal
Mangal Prakash, Mauricio Delbracio, Peyman Milanfar, Florian Jug |
ICLR | 3 |
| 2021 | Rich Features for Perceptual Quality Assessment of UGC VideosabstractVideo quality assessment for User Generated Content (UGC) is an important topic in both industry and academia. Most existing methods only focus on one aspect of the perceptual quality assessment, such as technical quality or compression artifacts. In this paper, we create a large scale dataset to comprehensively investigate characteristics of generic UGC video quality. Besides the subjective ratings and content labels of the dataset, we also propose a DNN-based framework to thoroughly analyze importance of content, technical quality, and compression level in perceptual quality. Our model is able to provide quality scores as well as human-friendly quality indicators, to bridge the gap between low level video signals to human perceptual quality. Experimental results show that our model achieves state-of-the-art correlation with Mean Opinion Scores (MOS). Yilin Wang 0001, Junjie Ke, Hossein Talebi, Joong Gon Yim, Neil Birkbeck, Balu Adsumilli, Peyman Milanfar, Feng Yang 0008 |
CVPR | 7 |
| 2021 | The Rate-Distortion-Accuracy Tradeoff: JPEG Case StudyabstractHandling digital images is almost always accompanied by a lossy compression in order to facilitate efficient transmission and storage. This introduces an unavoidable tension between the allocated bit-budget (rate) and the faithfulness of the resulting image to the original one (distortion). An additional complicating consideration is the effect of the compression on recognition performance by given classifiers (accuracy). This work aims to explore this rate-distortion-accuracy tradeoff. As a case study, we focus on the design of the quantization tables in the JPEG compression standard, offering a novel optimal tuning of these tables, leveraging a differential implementation of both the JPEG encoder-decoder and an entropy estimator. This enables us to offer a unified framework that considers the interplay between rate, distortion and classification accuracy. In all these fronts, we report a substantial boost in performance by a simple and easily implemented modification of these tables. Xiyang Luo, Hossein Talebi, Feng Yang 0008, Michael Elad, Peyman Milanfar |
DCC | 5 |
| 2021 | Projected Distribution Loss for Image EnhancementabstractFeatures obtained from object recognition CNNs have been widely used for measuring perceptual similarities between images. Such differentiable metrics can be used as perceptual learning losses to train image enhancement models. However, the choice of the distance function between input and target features may have a consequential impact on the performance of the trained model. While using the norm of the difference between extracted features leads to limited hallucination of details, measuring the distance between distributions of features may generate more textures; yet also more unrealistic details and artifacts. In this paper, we demonstrate that aggregating 1D-Wasserstein distances between CNN activations is more reliable than the existing approaches, and it can significantly improve the perceptual performance of enhancement models. More explicitly, we show that in imaging applications such as denoising, super-resolution, demosaicing, deblurring and JPEG artifact removal, the proposed learning loss outperforms the current state-of-the-art on reference-based perceptual losses. This means that the proposed learning loss can be plugged into different imaging frameworks and produce perceptually realistic results. Mauricio Delbracio, Hossein Talebi, Peyman Milanfar |
ICCP | 3 |
| 2021 | Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel DataabstractRecent work has shown impressive results on data-driven defocus deblurring using the two-image views available on modern dual-pixel (DP) sensors. One significant challenge in this line of research is access to DP data. Despite many cameras having DP sensors, only a limited number provide access to the low-level DP sensor images. In addition, capturing training data for defocus deblurring involves a time-consuming and tedious setup requiring the camera’s aperture to be adjusted. Some cameras with DP sensors (e.g., smartphones) do not have adjustable apertures, further limiting the ability to produce the necessary training data. We address the data capture bottleneck by proposing a procedure to generate realistic DP data synthetically. Our synthesis approach mimics the optical image formation found on DP sensors and can be applied to virtual scenes rendered with standard computer software. Leveraging these realistic synthetic DP images, we introduce a recurrent convolutional network (RCN) architecture that improves deblurring results and is suitable for use with single-frame and multi-frame data (e.g., video) captured by DP sensors. Finally, we show that our synthetic DP data is useful for training DNN models targeting video deblurring applications where access to DP data remains challenging. Abdullah Abuolaim, Mauricio Delbracio, Damien Kelly, Michael S. Brown, Peyman Milanfar |
ICCV | 5 |
| 2021 | MUSIQ: Multi-scale Image Quality TransformerabstractImage quality assessment (IQA) is an important research topic for understanding and improving visual experience. The current state-of-the-art IQA methods are based on convolutional neural networks (CNNs). The performance of CNN-based models is often compromised by the fixed shape constraint in batch training. To accommodate this, the input images are usually resized and cropped to a fixed shape, causing image quality degradation. To address this, we design a multi-scale image quality Transformer (MUSIQ) to process native resolution images with varying sizes and aspect ratios. With a multi-scale image representation, our proposed method can capture image quality at different granularities. Furthermore, a novel hash-based 2D spatial embedding and a scale embedding is proposed to support the positional embedding in the multi-scale representation. Experimental results verify that our method can achieve state-of-the-art performance on multiple large scale IQA datasets such as PaQ-2-PiQ [41], SPAQ [11], and KonIQ-10k [16].1 Junjie Ke, Qifei Wang, Yilin Wang 0001, Peyman Milanfar, Feng Yang 0008 |
ICCV | 4 |
| 2021 | COMISR: Compression-Informed Video Super-ResolutionabstractMost video super-resolution methods focus on restoring high-resolution video frames from low-resolution videos without taking into account compression. However, most videos on the web or mobile devices are compressed, and the compression can be severe when the bandwidth is limited. In this paper, we propose a new compression-informed video super-resolution model to restore high-resolution content without introducing artifacts caused by compression. The proposed model consists of three modules for video super-resolution: bi-directional recurrent warping, detail-preserving flow estimation, and Laplacian enhancement. All these three modules are used to deal with compression properties such as the location of the intra-frames in the input and smoothness in the output frames. For thorough performance evaluation, we conducted extensive experiments on standard datasets with a wide range of compression rates, covering many real video use cases. We showed that our method not only recovers high-resolution content on uncompressed frames from the widely-used benchmark datasets, but also achieves state-of-the-art performance in super-resolving compressed videos based on numerous quantitative metrics. We also evaluated the proposed method by simulating streaming from YouTube to demonstrate its effectiveness and robustness. The source codes and trained models are available at https://github.com/google-research/googleresearch/tree/master/comisr. Yinxiao Li, Pengchong Jin, Feng Yang 0008, Ce Liu 0001, Ming-Hsuan Yang 0001, Peyman Milanfar |
ICCV | 6 |
| 2021 | Learning to Resize Images for Computer Vision TasksabstractFor all the ways convolutional neural nets have revolutionized computer vision in recent years, one important aspect has received surprisingly little attention: the effect of image size on the accuracy of tasks being trained for. Typically, to be efficient, the input images are resized to a relatively small spatial resolution (e.g. 224 × 224), and both training and inference are carried out at this resolution. The actual mechanism for this re-scaling has been an afterthought: Namely, off-the-shelf image re sizers such as bilinear and bicubic are commonly used in most machine learning software frameworks. But do these re sizers limit the on-task performance of the trained networks? The answer is yes. Indeed, we show that the typical linear re sizer can be replaced with learned resizers that can substantially improve performance. Importantly, while the classical re-sizers typically result in better perceptual quality of the downscaled images, our proposed learned resizers do not necessarily give better visual quality, but instead improve task performance.Our learned image resizer is jointly trained with a base-line vision model. This learned CNN-based resizer creates machine friendly visual manipulations that lead to a consistent improvement of the end task metric over the baseline model. Specifically, here we focus on the classification task with the ImageNet dataset [26], and experiment with four different models to learn resizers adapted to each model. Moreover, we show that the proposed resizer can also be useful for fine-tuning the classification baselines for other vision tasks. To this end, we experiment with three different baselines to develop image quality assessment (IQA) models on the AVA dataset [24]. Hossein Talebi, Peyman Milanfar |
ICCV | 2 |
| 2021 | Patch Craft: Video Denoising by Deep Modeling and Patch MatchingabstractThe non-local self-similarity property of natural images has been exploited extensively for solving various image processing problems. When it comes to video sequences, harnessing this force is even more beneficial due to the temporal redundancy. In the context of image and video denoising, many classically-oriented algorithms employ self-similarity, splitting the data into overlapping patches, gathering groups of similar ones and processing these together somehow. With the emergence of convolutional neural networks (CNN), the patch-based framework has been abandoned. Most CNN denoisers operate on the whole image, leveraging non-local relations only implicitly by using a large receptive field. This work proposes a novel approach for leveraging self-similarity in the context of video denoising, while still relying on a regular convolutional architecture. We introduce a concept of patch-craft frames – artificial frames that are similar to the real ones, built by tiling matched patches. Our algorithm augments video sequences with patch-craft frames and feeds them to a CNN. We demonstrate the substantial boost in denoising performance obtained with the proposed approach. Gregory Vaksman, Michael Elad, Peyman Milanfar |
ICCV | 3 |
| 2021 | Deep Perceptual Image Quality Assessment for CompressionabstractLossy Image compression is necessary for efficient storage and transfer of data. Typically the trade-off between bit-rate and quality determines the optimal compression level. This makes the image quality metric an integral part of any imaging system. While the existing full-reference metrics such as PSNR and SSIM may be less sensitive to perceptual quality, the recently introduced learning methods may fail to generalize to unseen data. In this paper we propose the largest image compression quality dataset to date with human perceptual preferences, enabling the use of deep learning, and we develop a full reference perceptual quality assessment metric for lossy image compression that outperforms the existing state-of-the-art methods. We show that the proposed model can effectively learn from thousands of examples available in the new dataset, and consequently it generalizes better to other unseen datasets of human perceptual preference. The CIQA dataset can be found at https://github.com/googleresearch/google-research/tree/master/CIQA Juan Carlos Mier, Eddie Huang, Hossein Talebi, Feng Yang 0008, Peyman Milanfar |
ICIP | 5 |
| 2021 | Multi-path Neural Networks for On-device Multi-domain Visual ClassificationabstractLearning multiple domains/tasks with a single model is important for improving data efficiency and lowering inference cost for numerous vision tasks, especially on resource-constrained mobile devices. However, hand-crafting a multi-domain/task model can be both tedious and challenging. This paper proposes a novel approach to automatically learn a multi-path network for multi-domain visual classification on mobile devices. The proposed multi-path network is learned from neural architecture search by applying one reinforcement learning controller for each domain to select the best path in the super-network created from a MobileNetV3-like search space. An adaptive balanced domain prioritization algorithm is proposed to balance optimizing the joint model on multiple domains simultaneously. The determined multi-path model selectively shares parameters across domains in shared nodes while keeping domain-specific parameters within non-shared nodes in individual domain paths. This approach effectively reduces the total number of parameters and FLOPS, encouraging positive knowledge transfer while mitigating negative interference across domains. Extensive evaluations on the Visual Decathlon dataset demonstrate that the proposed multi-path model achieves state-of-the-art performance in terms of accuracy, model size, and FLOPS against other approaches using MobileNetV3-like architectures. Furthermore, the proposed method improves average accuracy over learning single-domain models individually, and reduces the total number of parameters and FLOPS by 78% and 32% respectively, compared to the approach that simply bundles single-domain models for multi-domain learning. Qifei Wang, Junjie Ke, Joshua Greaves, Grace Chu, Gabriel Bender, Luciano Sbaiz, Alec Go, Andrew G. Howard, Ming-Hsuan Yang 0001, Jeff Gilbert, Peyman Milanfar, Feng Yang 0008 |
WACV | 11 |
| 2021 | Regularization by Denoising via Fixed-Point Projection (RED-PRO)abstractInverse problems in image processing are typically cast as optimization tasks, consisting of data fidelity and stabilizing regularization terms. A recent regularization strategy of great interest utilizes the power of denoising engines. Two such methods are the plug-and-play prior (PnP) and regularization by denoising (RED). While both have shown state-of-the-art results in various recovery tasks, their theoretical justification is incomplete. In this paper, we aim to bridge RED and PnP, enriching the understanding of both frameworks. Toward that end, we reformulate RED as a convex optimization problem utilizing a projection (RED-PRO) onto the fixed-point set of demicontractive denoisers. We offer a simple iterative solution to this problem, by which we show that under certain conditions the PnP proximal gradient method is a special case of RED-PRO, while providing guarantees for the convergence of both frameworks to globally optimal solutions. In addition, we present relaxations of RED-PRO that allow for handling denoisers with limited fixed-point sets. Finally, we demonstrate RED-PRO for the tasks of image deblurring and superresolution, showing improved results with respect to the original RED framework. Regev Cohen, Michael Elad, Peyman Milanfar |
SIAM J. Imaging Sci. | 3 |
| 2021 | Better Compression With Deep Pre-EditingabstractCould we compress images via standard codecs while avoiding visible artifacts? The answer is obvious - this is doable as long as the bit budget is generous enough. What if the allocated bit-rate for compression is insufficient? Then unfortunately, artifacts are a fact of life. Many attempts were made over the years to fight this phenomenon, with various degrees of success. In this work we aim to break the unholy connection between bit-rate and image quality, and propose a way to circumvent compression artifacts by pre-editing the incoming image and modifying its content to fit the given bits. We design this editing operation as a learned convolutional neural network, and formulate an optimization problem for its training. Our loss takes into account a proximity between the original image and the edited one, a bit-budget penalty over the proposed image, and a no-reference image quality measure for forcing the outcome to be visually pleasing. The proposed approach is demonstrated on the popular JPEG compression, showing savings in bits and/or improvements in visual quality, obtained with intricate editing effects. Hossein Talebi Esfandarani, Damien Kelly, Xiyang Luo, Ignacio Garcia-Dorado, Feng Yang 0008, Peyman Milanfar, Michael Elad |
IEEE Trans. Image Process. | 6 |
| 2021 | Deep K-SVD DenoisingabstractThis work considers noise removal from images, focusing on the well-known K-SVD denoising algorithm. This sparsity-based method was proposed in 2006, and for a short while it was considered as state-of-the-art. However, over the years it has been surpassed by other methods, including the recent deep-learning-based newcomers. The question we address in this paper is whether K-SVD was brought to its peak in its original conception, or whether it can be made competitive again. The approach we take in answering this question is to redesign the algorithm to operate in a supervised manner. More specifically, we propose an end-to-end deep architecture with the exact K-SVD computational path, and train it for optimized denoising. Our work shows how to overcome difficulties arising in turning the K-SVD scheme into a differentiable, and thus learnable, machine. With a small number of parameters to learn and while preserving the original K-SVD essence, the proposed architecture is shown to outperform the classical K-SVD algorithm substantially, and getting closer to recent state-of-the-art learning-based denoising methods. Adopting a broader context, this work touches on themes around the design of deep-learning solutions for image processing tasks, while paving a bridge between classic methods and novel deep-learning-based ones. Meyer Scetbon, Michael Elad, Peyman Milanfar |
IEEE Trans. Image Process. | 3 |
| 2020 | Distortion Agnostic Deep WatermarkingabstractWatermarking is the process of embedding information into an image that can survive under distortions, while requiring the encoded image to have little or no perceptual difference with the original image. Recently, deep learning-based methods achieved impressive results in both visual quality and message payload under a wide variety of image distortions. However, these methods all require differentiable models for the image distortions at training time, and may generalize poorly to unknown distortions. This is undesirable since the types of distortions applied to watermarked images are usually unknown and non-differentiable. In this paper, we propose a new framework for distortion-agnostic watermarking, where the image distortion is not explicitly modeled during training. Instead, the robustness of our system comes from two sources: adversarial training and channel coding. Compared to training on a fixed set of distortions and noise levels, our method achieves comparable or better results on distortions available during training, and better performance overall on unknown distortions. Xiyang Luo, Ruohan Zhan, Huiwen Chang, Feng Yang 0008, Peyman Milanfar |
CVPR | 5 |
| 2020 | GIFnets: Differentiable GIF Encoding FrameworkabstractGraphics Interchange Format (GIF) is a widely used image file format. Due to the limited number of palette colors, GIF encoding often introduces color banding artifacts. Traditionally, dithering is applied to reduce color banding, but introducing dotted-pattern artifacts. To reduce artifacts and provide a better and more efficient GIF encoding, we introduce a differentiable GIF encoding pipeline, which includes three novel neural networks: PaletteNet, DitherNet, and BandingNet. Each of these three networks provides an important functionality within the GIF encoding pipeline. PaletteNet predicts a near-optimal color palette given an input image. DitherNet manipulates the input image to reduce color banding artifacts and provides an alternative to traditional dithering. Finally, BandingNet is designed to detect color banding, and provides a new perceptual loss specifically for GIF images. As far as we know, this is the first fully differentiable GIF encoding pipeline based on deep neural networks and compatible with existing GIF decoders. User study shows that our algorithm is better than Floyd-Steinberg based GIF encoding. Innfarn Yoo, Xiyang Luo, Yilin Wang 0001, Feng Yang 0008, Peyman Milanfar |
CVPR | 5 |
| 2020 | Rank-Smoothed Pairwise Learning In Perceptual Quality AssessmentabstractConducting pairwise comparisons is a widely used approach in curating human perceptual preference data. Typically raters are instructed to make their choices according to a specific set of rules that address certain dimensions of image quality and aesthetics. The outcome of this process is a dataset of sampled image pairs with their associated empirical preference probabilities. Training a model on these pairwise preferences is a common deep learning approach. However, optimizing by gradient descent through mini-batch learning means that the “global” ranking of the images is not explicitly taken into account. In other words, each step of the gradient descent relies only on a limited number of pairwise comparisons. In this work, we demonstrate that regularizing the pairwise empirical probabilities with aggregated rankwise probabilities leads to a more reliable training loss. We show that training a deep image quality assessment model with our rank-smoothed loss consistently improves the accuracy of predicting human preferences. Hossein Talebi, Ehsan Amid, Peyman Milanfar, Manfred K. Warmuth |
ICIP | 3 |
| 2020 | Super-Resolving Commercial Satellite Imagery Using Realistic Training DataabstractIn machine learning based single image super-resolution, the degradation model is embedded in training data generation. However, most existing satellite image super-resolution methods use a simple down-sampling model with a fixed kernel to create training images. These methods work fine on synthetic data, but do not perform well on real satellite images. We propose a realistic training data generation model for commercial satellite imagery products, which includes not only the imaging process on satellites but also the post-process on the ground. We also propose a convolutional neural network optimized for satellite images. Experiments show that the proposed training data generation model is able to improve super-resolution performance on real satellite images. Hossein Talebi, Xinwei Shi, Feng Yang 0008, Peyman Milanfar |
ICIP | 5 |
| 2020 | Image stylisation: from predefined to personalisedabstractThe authors present a framework for interactive design of new image stylisations using a wide range of predefined filter blocks. Both novel and off‐the‐shelf image filtering and rendering techniques are extended and combined to allow the user to unleash their creativity to intuitively invent, modify, and tune new styles from a given set of filters. In parallel to this manual design, they propose a novel procedural approach that automatically assembles sequences of filters, leading to unique and novel styles. An important aim of the authors’ framework is to allow for interactive exploration and design, as well as to enable videos and camera streams to be stylised on the fly. In order to achieve this real‐time performance, they use the Best Linear Adaptive Enhancement (BLADE) framework – an interpretable shallow machine learning method that simulates complex filter blocks in real time. Their representative results include over a dozen styles designed using their interactive tool, a set of styles created procedurally, and new filters trained with their BLADE approach. Ignacio Garcia-Dorado, Pascal Getreuer, Bartlomiej Wronski, Peyman Milanfar |
IET Comput. Vis. | 4 |
| 2019 | Local Kernels That Approximate Bayesian Regularization and Proximal OperatorsabstractIn this work, we broadly connect kernel-based filtering (e.g. approaches such as the bilateral filter and nonlocal means, but also many more) with general variational formulations of Bayesian regularized least squares, and the related concept of proximal operators. Variational/Bayesian/proximal formulations often result in optimization problems that do not have closed-form solutions, and therefore typically require global iterative solutions. Our main contribution here is to establish how one can approximate the solution of the resulting global optimization problems using locally adaptive filters with specific kernels. Our results are valid for small regularization strength (i.e. weak noise) but the approach is powerful enough to be useful for a wide range of applications because we expose how to derive a "kernelized" solution to these problems that approximates the global solution in one shot, using only local operations. As another side benefit in the reverse direction, given a local data-adaptive filter constructed with a particular choice of kernel, we enable the interpretation of such filters in the variational/Bayesian/proximal framework. Frank Ong, Peyman Milanfar, Pascal Getreuer |
IEEE Trans. Image Process. | 2 |
| 2019 | Handheld multi-frame super-resolutionabstractCompared to DSLR cameras, smartphone cameras have smaller sensors, which limits their spatial resolution; smaller apertures, which limits their light gathering ability; and smaller pixels, which reduces their signal-to-noise ratio. The use of color filter arrays (CFAs) requires demosaicing, which further degrades resolution. In this paper, we supplant the use of traditional demosaicing in single-frame and burst photography pipelines with a multiframe super-resolution algorithm that creates a complete RGB image directly from a burst of CFA raw images. We harness natural hand tremor, typical in handheld photography, to acquire a burst of raw frames with small offsets. These frames are then aligned and merged to form a single image with red, green, and blue values at every pixel site. This approach, which includes no explicit demosaicing step, serves to both increase image resolution and boost signal to noise ratio. Our algorithm is robust to challenging scene conditions: local motion, occlusion, or scene changes. It runs at 100 milliseconds per 12-megapixel RAW input burst frame on mass-produced mobile phones. Specifically, the algorithm is the basis of the Super-Res Zoom feature, as well as the default merge method in Night Sight mode (whether zooming or not) on Google's flagship phone. Bartlomiej Wronski, Ignacio Garcia-Dorado, Manfred Ernst, Damien Kelly, Michael Krainin, Chia-Kai Liang, Marc Levoy, Peyman Milanfar |
ACM Trans. Graph. | 8 |
| 2018 | RED-UCATION: A Novel CNN Architecture Based on Denoising NonlinearitiesabstractImage denoising is the most fundamental image enhancement task, and many algorithms have been proposed over the years for its solution. Interestingly, such an image denoising “engine” can be used to solve general inverse problems. Indeed, in our recent work we have presented the Regularization by Denoising (RED) framework: using a denoising engine in defining the regularization of any inverse problem. We have shown how this scheme leads to well-founded iterative algorithms in which the denoiser is applied in each iteration. In this work we describe how a learned version of RED defines a novel convolutional neural network architecture, where the commonly used point-wise nonlinearities are replaced by a denoising engine. We show how this network can be optimized end- to-end using a back - propagation that relies on guided denoising algorithms. As a case-study, we concentrate on the image deblurring problem and show the superiority of the trainable variant of RED over its analytic form. Yaniv Romano, Michael Elad, Peyman Milanfar |
ICASSP | 3 |
| 2018 | BLADE: Filter learning for general purpose computational photographyabstractThe Rapid and Accurate Image Super Resolution (RAISR) method of Romano, Isidoro, and Milanfar is a computationally efficient image upscaling method using a trained set of filters. We describe a generalization of RAISR, which we name Best Linear Adaptive Enhancement (BLADE). This approach is a trainable edge-adaptive filtering framework that is general, simple, computationally efficient, and useful for a wide range of problems in computational photography. We show applications to operations which may appear in a camera pipeline including denoising, demosaicking, and stylization. Pascal Getreuer, Ignacio Garcia-Dorado, John Isidoro, Sungjoon Choi 0001, Frank Ong, Peyman Milanfar |
ICCP | 6 |
| 2018 | Learned perceptual image enhancementabstractLearning a typical image enhancement pipeline involves minimization of a loss function between enhanced and reference images. While L1and L2losses are perhaps the most widely used functions for this purpose, they do not necessarily lead to perceptually compelling results. In this paper, we show that adding a learned no-reference image quality metric to the loss can significantly improve enhancement operators. This metric is implemented using a CNN (convolutional neural network) trained on a large-scale dataset labelled with aesthetic preferences ofhuman raters. This loss allows us to conveniently perform back-propagation in our learning framework to simultaneously optimize for similarity to a given ground truth reference and perceptual quality. This perceptual loss is only used to train parameters of image processing operators, and does not impose any extra complexity at inference time. Our experiments demonstrate that this loss can be effective for tuning a variety of operators such as local tone mapping and dehazing. Hossein Talebi Esfandarani, Peyman Milanfar |
ICCP | 2 |
| 2018 | Fast, Trainable, Multiscale DenoisingabstractDenoising is a fundamental imaging problem. Versatile but fast filtering has been demanded for mobile camera systems. We present an approach to multiscale filtering which allows real-time applications on low-powered devices. The key idea is to learn a set of kernels that upscales, filters, and blends patches of different scales guided by local structure analysis. This approach is trainable so that learned filters are capable of treating diverse noise patterns and artifacts. Experimental results show that the presented approach produces comparable results to state-of-the-art algorithms while processing time is orders of magnitude faster. Sungjoon Choi 0001, John Isidoro, Pascal Getreuer, Peyman Milanfar |
ICIP | 4 |
| 2018 | Rendition: Reclaiming What a Black Box Takes AwayabstractThe premise of our work is deceptively familiar: A black box $f(\cdot)$ has altered an image $\mathbf{x} \rightarrow f(\mathbf{x})$, but only mildly. Recover the image $\mathbf{x}$. This black box might be any number of simple or complicated things: a linear or nonlinear filter, some app on your phone, etc. The latter is a good canonical example for the problem we address: Given only “the app” and an image produced by the app, find the image that was fed to the app. We can run the given image (or any other image) through the app as many times as we like, but we cannot look inside the (code for the) app to see how it works At first blush, the problem sounds a lot like a standard inverse problem [16, 19], but it is not in the following sense: While we have access to the black box $f(\cdot)$ and can run any image through it and observe the output, we do not know how the block box alters the image. Therefore we have no explicit form or model of $f(\cdot)$. Nor are we necessarily interested in the internal workings of the black box. We assume only that the effect of the black box is mild in the sense that $f(\mathbf{x}) -\mathbf{x}$ has a small Lipschitz seminorm---that is, $\|f(\mathbf{x}) -\mathbf{x}\|\leq \delta \|\mathbf{x} \|$ for a sufficiently small $\delta$. And as such, we are seeking to reverse its effect on a particular image, to whatever extent possible. This is what we call the “rendition” (rather than restoration) problem, as it does not fit the mold of an inverse problem (blind or otherwise). We describe general conditions under which this rendition is possible and provide a remarkably simple algorithm that works for both contractive and expansive black box operators. The principal and novel takeaway message from our work is this surprising fact: One simple algorithm can reliably undo a wide class of (not too violent) image distortions. Peyman Milanfar |
SIAM J. Imaging Sci. | 1 |
| 2018 | NIMA: Neural Image AssessmentabstractAutomatically learned quality assessment for images has recently become a hot topic due to its usefulness in a wide variety of applications such as evaluating image capture pipelines, storage techniques and sharing media. Despite the subjective nature of this problem, most existing methods only predict the mean opinion score provided by datasets such as AVA [1] and TID2013 [2]. Our approach differs from others in that we predict the distribution of human opinion scores using a convolutional neural network. Our architecture also has the advantage of being significantly simpler than other methods with comparable performance. Our proposed approach relies on the success (and retraining) of proven, state-of-the-art deep object recognition networks. Our resulting network can be used to not only score images reliably and with high correlation to human perception, but also to assist with adaptation and optimization of photo editing/enhancement algorithms in a photographic pipeline. All this is done without need for a "golden" reference image, consequently allowing for single-image, semantic- and perceptually-aware, no-reference quality assessment. Hossein Talebi Esfandarani, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2018 | Statistical Models of Signal and Noise and Fundamental Limits of Segmentation Accuracy in Retinal Optical Coherence TomographyabstractOptical coherence tomography (OCT) has revolutionized diagnosis and prognosis of ophthalmic diseases by visualization and measurement of retinal layers. To speed up the quantitative analysis of disease biomarkers, an increasing number of automatic segmentation algorithms have been proposed to estimate the boundary locations of retinal layers. While the performance of these algorithms has significantly improved in recent years, a critical question to ask is how far we are from a theoretical limit to OCT segmentation performance. In this paper, we present the Cramèr-Rao lower bounds (CRLBs) for the problem of OCT layer segmentation. In deriving the CRLBs, we address the important problem of defining statistical models that best represent the intensity distribution in each layer of the retina. Additionally, we calculate the bounds under an optimal affine bias, reflecting the use of prior knowledge in many segmentation algorithms. Experiments using in vivo images of human retina from a commercial spectral domain OCT system are presented, showing potential for improvement of automated segmentation accuracy. Our general mathematical model can be easily adapted for virtually any OCT system. Furthermore, the statistical models of signal and noise developed in this paper can be utilized for the future improvements of OCT image denoising, reconstruction, and many other applications. Theodore B. Dubose, David Cunefare, Elijah Cole, Peyman Milanfar, Joseph A. Izatt, Sina Farsiu |
IEEE Trans. Medical Imaging | 4 |
| 2017 | The Little Engine That Could: Regularization by Denoising (RED)abstractRemoval of noise from an image is an extensively studied problem in image processing. Indeed, the recent advent of sophisticated and highly effective denoising algorithms has led some to believe that existing methods are touching the ceiling in terms of noise removal performance. Can we leverage this impressive achievement to treat other tasks in image processing? Recent work has answered this question positively, in the form of the Plug-and-Play Prior ($P^3$) method, showing that any inverse problem can be handled by sequentially applying image denoising steps. This relies heavily on the ADMM optimization technique in order to obtain this chained denoising interpretation. Is this the only way in which tasks in image processing can exploit the image denoising engine? In this paper we provide an alternative, more powerful, and more flexible framework for achieving the same goal. As opposed to the $P^3$ method, we offer Regularization by Denoising (RED): using the denoising engine in defining the regularization of the inverse problem. We propose an explicit image-adaptive Laplacian-based regularization functional, making the overall objective functional clearer and better defined. With a complete flexibility to choose the iterative optimization procedure for minimizing the above functional, RED is capable of incorporating any image denoising algorithm, can treat general inverse problems very effectively, and is guaranteed to converge to the globally optimal result. We test this approach and demonstrate state-of-the-art results in the image deblurring and super-resolution problems. Yaniv Romano, Michael Elad, Peyman Milanfar |
SIAM J. Imaging Sci. | 3 |
| 2017 | Linear Support Tensor Machine With LSK Channels: Pedestrian Detection in Thermal Infrared ImagesabstractPedestrian detection in thermal infrared images poses unique challenges because of the low resolution and noisy nature of the image. Here, we propose a mid-level attribute in the form of the multidimensional template, or tensor, using local steering kernel (LSK) as low-level descriptors for detecting pedestrians in far infrared images. LSK is specifically designed to deal with intrinsic image noise and pixel level uncertainty by capturing local image geometry succinctly instead of collecting local orientation statistics (e.g., histograms in histogram of oriented gradients). In order to learn the LSK tensor, we introduce a new image similarity kernel following the popular maximum margin framework of support vector machines facilitating a relatively short and simple training phase for building a rigid pedestrian detector. Tensor representation has several advantages, and indeed, LSK templates allow exact acceleration of the sluggish but de facto sliding window-based detection methodology with multichannel discrete Fourier transform, facilitating very fast and efficient pedestrian localization. The experimental studies on publicly available thermal infrared images justify our proposals and model assumptions. In addition, the proposed work also involves the release of our in-house annotations of pedestrians in more than 17 000 frames of OSU color thermal database for the purpose of sharing with the research community. Sujoy Kumar Biswas, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2017 | Style Transfer Via Texture SynthesisabstractStyle transfer is a process of migrating a style from a given image to the content of another, synthesizing a new image, which is an artistic mixture of the two. Recent work on this problem adopting convolutional neural-networks (CNN) ignited a renewed interest in this field, due to the very impressive results obtained. There exists an alternative path toward handling the style transfer task, via the generalization of texture synthesis algorithms. This approach has been proposed over the years, but its results are typically less impressive compared with the CNN ones. In this paper, we propose a novel style transfer algorithm that extends the texture synthesis work of Kwatra et al. (2005), while aiming to get stylized images that are closer in quality to the CNN ones. We modify Kwatra's algorithm in several key ways in order to achieve the desired transfer, with emphasis on a consistent way for keeping the content intact in selected regions, while producing hallucinated and rich style in others. The results obtained are visually pleasing and diverse, shown to be competitive with the recent CNN style transfer algorithms. The proposed algorithm is fast and flexible, being able to process any pair of content + style images. Michael Elad, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2016 | A pull-push method for fast non-local means filteringabstractNon-local means filtering (NLM), has garnered a large amount of interest in the image processing community due to its capability to exploit image patch self-similarity in order to effectively filter noisy images. However, the computational complexity of non-local means filtering is the product of three different factors; namely, O(NDK), where K is the number of filter kernel taps (e.g. search window size), D is the number of patch taps, and N is number of pixels. We propose a fast approximation of non-local means filtering using the multiscale methodology of the pull-push scattered data interpolation method. By using NLM with a small filter kernel to selectively propagate filtering results and noise variance estimates from fine to coarse scales and back, the process can be used to provide comparable filtering capability to brute force NLM but with algorithmic complexity that is linear in the number of image pixels and the patch comparison taps, O(ND). In practical application, we demonstrate its denoising capability is comparable to NLM with much larger filter kernels, but at a fraction of the computational cost. John Isidoro, Peyman Milanfar |
ICIP | 2 |
| 2016 | A new class of image filters without normalizationabstractWhen applying a filter to an image, it often makes practical sense to maintain the local brightness level from input to output image. This is achieved by normalizing the filter coefficients so that they sum to one. This concept is generally taken for granted, but is particularly important where nonlinear filters such as the bilateral or and non-local means are concerned, where the effect on local brightness and contrast can be complex. Here we present a method for achieving the same level of control over the local filter behavior without the need for this normalization. Namely, we show how to closely approximate any normalized filter without in fact needing this normalization step. This yields a new class of filters. We derive a closed-form expression for the approximating filter and analyze its behavior, showing it to be easily controlled for quality and nearness to the exact filter, with a single parameter. Our experiments demonstrate that the un-normalized affinity weights can be effectively used in applications such as image smoothing, sharpening and detail enhancement. Peyman Milanfar, Hossein Talebi Esfandarani |
ICIP | 1 |
| 2016 | Turbo denoising for mobile photographic applicationsabstractWe propose a new denoising algorithm for camera pipelines and other photographic applications. We aim for a scheme that is (1) fast enough to be practical even for mobile devices, and (2) handles the realistic content dependent noise in real camera captures. Our scheme consists of a simple two-stage non-linear processing. We introduce a new form of boosting/blending which proves to be very effective in restoring the details lost in the first denoising stage. We also employ IIR filtering to significantly reduce the computation time. Further, we incorporate a novel noise model to address the content dependent noise. For realistic camera noise, our results are competitive with BM3D, but with nearly 400 times speedup. Tak-Shing Wong, Peyman Milanfar |
ICIP | 2 |
| 2016 | One Shot Detection with Laplacian Object and Fast Matrix Cosine SimilarityabstractOne shot, generic object detection involves searching for a single query object in a larger target image. Relevant approaches have benefited from features that typically model the local similarity patterns. In this paper, we combine local similarity (encoded by local descriptors) with a global context (i.e., a graph structure) of pairwise affinities among the local descriptors, embedding the query descriptors into a low dimensional but discriminatory subspace. Unlike principal components that preserve global structure of feature space, we actually seek a linear approximation to the Laplacian eigenmap that permits us a locality preserving embedding of high dimensional region descriptors. Our second contribution is an accelerated but exact computation of matrix cosine similarity as the decision rule for detection, obviating the computationally expensive sliding window search. We leverage the power of Fourier transform combined with integral image to achieve superior runtime efficiency that allows us to test multiple hypotheses (for pose estimation) within a reasonably short time. Our approach to one shot detection is training-free, and experiments on the standard data sets confirm the efficacy of our model. Besides, low computation cost of the proposed (codebook-free) object detector facilitates rather straightforward query detection in large data sets including movie videos. Sujoy Kumar Biswas, Peyman Milanfar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Asymptotic Performance of Global DenoisingabstractWe provide an upper bound on the rate of convergence of the mean-squared error for global image denoising and illustrate that this upper bound decays with increasing image size. Hence, global denoising is asymptotically optimal. At least in an oracle scenario this property does not hold for patch-based methods such as BM3D, thereby limiting their performance for large images. As observed in practice and shown in this work, this gap in performance is small for moderate size images, but it can grow quickly with image size. Hossein Talebi Esfandarani, Peyman Milanfar |
SIAM J. Imaging Sci. | 2 |
| 2014 | Laplacian object: One-shot object detection by locality preserving projectionabstractOne shot, generic object detection involves detecting a single query image in a target image. Relevant approaches have benefitted from features that typically model the local similarity patterns. Also important is the global matching of local features along the object detection process. In this paper, we consider such global information early in the feature extraction stage by combining local geodesic structure (encoded by LARK descriptors) with a global context (i.e., graph structure) of pairwise affinities among the local descriptors. The result is an embedding of the LARK descriptors (extracted from query image) into a discriminatory subspace (obtained using locality preserving projection [1]) that preserves the local intrinsic geometry of the query image patterns. Experiments on standard data sets demonstrate efficacy of our proposed approach. Sujoy Kumar Biswas, Peyman Milanfar |
ICIP | 2 |
| 2014 | A smartphone application for removing handshake blur and compensating rolling shutterabstractSmartphones are now widely used as photographic devices. Equipped with cheap cameras they are prone to many degradations, most notably handshake in combination with rolling shutter causes severe space-variant blur. Removing blur without any information about the camera motion is a computationally demanding and unstable process. We use built-in gyroscopes to record the motion trajectory of the camera during exposure and then remove blur from the acquired photograph based on the reconstructed trajectory. The proposed deblurring application is implemented on Android smartphones with close-to-real-time performance. Ondrej Sindelar, Filip Sroubek, Peyman Milanfar |
ICIP | 3 |
| 2014 | Global denoising is asymptotically optimalabstractIn this paper an upper bound on the decay rate of the mean-squared error for global image denoising is derived. As image size increases, this upper bound decays to zero; that is, the global denoising is asymptotically optimal. Unlike patch-based methods such as BM3D, this property only holds for global denoising schemes. In practice, and as demonstrated in this work, this performance gap between patch-based and global denoisers can grow rapidly with image size. Hossein Talebi Esfandarani, Peyman Milanfar |
ICIP | 2 |
| 2014 | Global Image DenoisingabstractMost existing state-of-the-art image denoising algorithms are based on exploiting similarity between a relatively modest number of patches. These patch-based methods are strictly dependent on patch matching, and their performance is hamstrung by the ability to reliably find sufficiently similar patches. As the number of patches grows, a point of diminishing returns is reached where the performance improvement due to more patches is offset by the lower likelihood of finding sufficiently close matches. The net effect is that while patch-based methods, such as BM3D, are excellent overall, they are ultimately limited in how well they can do on (larger) images with increasing complexity. In this paper, we address these shortcomings by developing a paradigm for truly global filtering where each pixel is estimated from all pixels in the image. Our objectives in this paper are two-fold. First, we give a statistical analysis of our proposed global filter, based on a spectral decomposition of its corresponding operator, and we study the effect of truncation of this spectral decomposition. Second, we derive an approximation to the spectral (principal) components using the Nyström extension. Using these, we demonstrate that this global filter can be implemented efficiently by sampling a fairly small percentage of the pixels in the image. Experiments illustrate that our strategy can effectively globalize any existing denoising filters to estimate each pixel using all pixels in the image, hence improving upon the best patch-based methods. Hossein Talebi Esfandarani, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2014 | Nonlocal Image EditingabstractIn this paper, we introduce a new image editing tool based on the spectrum of a global filter computed from image affinities. Recently, it has been shown that the global filter derived from a fully connected graph representing the image can be approximated using the Nyström extension. This filter is computed by approximating the leading eigenvectors of the filter. These orthonormal eigenfunctions are highly expressive of the coarse and fine details in the underlying image, where each eigenvector can be interpreted as one scale of a data-dependent multiscale image decomposition. In this filtering scheme, each eigenvalue can boost or suppress the corresponding signal component in each scale. Our analysis shows that the mapping of the eigenvalues by an appropriate polynomial function endows the filter with a number of important capabilities, such as edge-aware sharpening, denoising, tone manipulation, and abstraction, to name a few. Furthermore, the edits can be easily propagated across the image. Hossein Talebi Esfandarani, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2014 | A General Framework for Regularized, Similarity-Based Image RestorationabstractAny image can be represented as a function defined on a weighted graph, in which the underlying structure of the image is encoded in kernel similarity and associated Laplacian matrices. In this paper, we develop an iterative graph-based framework for image restoration based on a new definition of the normalized graph Laplacian. We propose a cost function, which consists of a new data fidelity term and regularization term derived from the specific definition of the normalized graph Laplacian. The normalizing coefficients used in the definition of the Laplacian and associated regularization term are obtained using fast symmetry preserving matrix balancing. This results in some desired spectral properties for the normalized Laplacian such as being symmetric, positive semidefinite, and returning zero vector when applied to a constant image. Our algorithm comprises of outer and inner iterations, where in each outer iteration, the similarity weights are recomputed using the previous estimate and the updated objective function is minimized using inner conjugate gradient iterations. This procedure improves the performance of the algorithm for image deblurring, where we do not have access to a good initial estimate of the underlying image. In addition, the specific form of the cost function allows us to render the spectral analysis for the solutions of the corresponding linear equations. In addition, the proposed approach is general in the sense that we have shown its effectiveness for different restoration problems, including deblurring, denoising, and sharpening. Experimental results verify the effectiveness of the proposed algorithm on both synthetic and real examples. Amin Kheradmand, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2013 | Blind Deconvolution Using Alternating Maximum a Posteriori Estimation with Heavy-Tailed Priors
Jan Kotera, Filip Sroubek, Peyman Milanfar |
CAIP (2) | 3 |
| 2013 | Removing Atmospheric Turbulence via Space-Invariant DeconvolutionabstractTo correct geometric distortion and reduce space and time-varying blur, a new approach is proposed in this paper capable of restoring a single high-quality image from a given image sequence distorted by atmospheric turbulence. This approach reduces the space and time-varying deblurring problem to a shift invariant one. It first registers each frame to suppress geometric deformation through B-spline-based nonrigid registration. Next, a temporal regression process is carried out to produce an image from the registered frames, which can be viewed as being convolved with a space invariant near-diffraction-limited blur. Finally, a blind deconvolution algorithm is implemented to deblur the fused image, generating a final output. Experiments using real data illustrate that this approach can effectively alleviate blur and distortions, recover details of the scene, and significantly improve visual quality. Peyman Milanfar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Symmetrizing Smoothing FiltersabstractWe study a general class of nonlinear and shift-varying smoothing filters that operate based on averaging. This important class of filters includes many well-known examples such as the bilateral filter, nonlocal means, general adaptive moving average filters, and more. (Many linear filters such as linear minimum mean-squared error smoothing filters, Savitzky--Golay filters, smoothing splines, and wavelet smoothers can be considered special cases.) They are frequently used in both signal and image processing as they are elegant, computationally simple, and high performing. The operators that implement such filters, however, are not symmetric in general. The main contribution of this paper is to provide a provably stable method for symmetrizing the smoothing operators. Specifically, we propose a novel approximation of smoothing operators by symmetric doubly stochastic matrices and show that this approximation is stable and accurate, even more so in higher dimensions. We demonstrate that there are several important advantages to this symmetrization, particularly in image processing/filtering applications such as denoising. In particular, (1) doubly stochastic filters generally lead to improved performance over the baseline smoothing procedure; (2) when the filters are applied iteratively, the symmetric ones can be guaranteed to lead to stable algorithms; and (3) symmetric smoothers allow an orthonormal eigendecomposition which enables us to peer into the complex behavior of such nonlinear and shift-varying filters in a locally adapted basis using principal components. Finally, a doubly stochastic filter has a simple and intuitive interpretation. Namely, it implies the very natural property that every pixel in the given input image has the same sum total contribution to the output image. Peyman Milanfar |
SIAM J. Imaging Sci. | 1 |
| 2013 | How to SAIF-ly Boost Denoising PerformanceabstractSpatial domain image filters (e.g., bilateral filter, non-local means, locally adaptive regression kernel) have achieved great success in denoising. Their overall performance, however, has not generally surpassed the leading transform domain-based filters (such as BM3-D). One important reason is that spatial domain filters lack efficiency to adaptively fine tune their denoising strength; something that is relatively easy to do in transform domain method with shrinkage operators. In the pixel domain, the smoothing strength is usually controlled globally by, for example, tuning a regularization parameter. In this paper, we propose spatially adaptive iterative filtering (SAIF) is the Middle Eastern/Arabic name for sword. This acronym somehow seems appropriate for what the algorithm does by precisely tuning the value of the iteration number. a new strategy to control the denoising strength locally for any spatial domain method. This approach is capable of filtering local image content iteratively using the given base filter, and the type of iteration and the iteration number are automatically optimized with respect to estimated risk (i.e., mean-squared error). In exploiting the estimated local signal-to-noise-ratio, we also present a new risk estimator that is different from the often-employed SURE method, and exceeds its performance in many cases. Experiments illustrate that our strategy can significantly relax the base algorithm's sensitivity to its tuning (smoothing) parameters, and effectively boost the performance of several existing denoising filters to generate state-of-the-art results under both simulated and practical conditions. Hossein Talebi Esfandarani, Peyman Milanfar |
IEEE Trans. Image Process. | 3 |
| 2013 | Estimating Spatially Varying Defocus Blur From A Single ImageabstractEstimating the amount of blur in a given image is important for computer vision applications. More specifically, the spatially varying defocus point-spread-functions (PSFs) over an image reveal geometric information of the scene, and their estimate can also be used to recover an all-in-focus image. A PSF for a defocus blur can be specified by a single parameter indicating its scale. Most existing algorithms can only select an optimal blur from a finite set of candidate PSFs for each pixel. Some of those methods require a coded aperture filter inserted in the camera. In this paper, we present an algorithm estimating a defocus scale map from a single image, which is applicable to conventional cameras. This method is capable of measuring the probability of local defocus scale in the continuous domain. It also takes smoothness and color edge information into consideration to generate a coherent blur map indicating the amount of blur at each pixel. Simulated and real data experiments illustrate excellent performance and its successful applications in foreground/background segmentation. Scott Cohen, Stephen Schiller, Peyman Milanfar |
IEEE Trans. Image Process. | 4 |
| 2012 | Deconvolving PSFs for a Better Motion Deblurring Using Multiple Images
Filip Sroubek, Peyman Milanfar |
ECCV (5) | 3 |
| 2012 | Improving denoising filters by optimal diffusionabstractKernel based methods have recently been used widely in image denoising. Tuning the parameters of these algorithms directly affects their performance. In this paper, an iterative method is proposed which optimizes the performance of any kernel based denoising algorithm in the mean-squared error (MSE) sense, even with arbitrary parameters. In this work we estimate the MSE in each image patch, and use this estimate to guide the iterative application to a stop, hence leading to improve performance. We propose a new estimator for the risk (i.e. MSE) which is different than the often-employed SURE method. We illustrate that the proposed risk estimate can outperform SURE in many instances. Hossein Talebi Esfandarani, Peyman Milanfar |
ICIP | 2 |
| 2012 | Patch-Based Near-Optimal Image DenoisingabstractIn this paper, we propose a denoising method motivated by our previous analysis of the performance bounds for image denoising. Insights from that study are used here to derive a high-performance practical denoising algorithm. We propose a patch-based Wiener filter that exploits patch redundancy for image denoising. Our framework uses both geometrically and photometrically similar patches to estimate the different filter parameters. We describe how these parameters can be accurately estimated directly from the input noisy image. Our denoising approach, designed for near-optimal performance (in the mean-squared error sense), has a sound statistical foundation that is analyzed in detail. The performance of our approach is experimentally verified on a variety of images and noise levels. The results presented here demonstrate that our proposed method is on par or exceeding the current state of the art, both visually and quantitatively. Priyam Chatterjee, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2012 | Robust Multichannel Blind Deconvolution via Fast Alternating MinimizationabstractBlind deconvolution, which comprises simultaneous blur and image estimations, is a strongly ill-posed problem. It is by now well known that if multiple images of the same scene are acquired, this multichannel (MC) blind deconvolution problem is better posed and allows blur estimation directly from the degraded images. We improve the MC idea by adding robustness to noise and stability in the case of large blurs or if the blur size is vastly overestimated. We formulate blind deconvolution as an l(1) -regularized optimization problem and seek a solution by alternately optimizing with respect to the image and with respect to blurs. Each optimization step is converted to a constrained problem by variable splitting and then is addressed with an augmented Lagrangian method, which permits simple and fast implementation in the Fourier domain. The rapid convergence of the proposed method is illustrated on synthetically blurred data. Applicability is also demonstrated on the deconvolution of real photos taken by a digital camera. Filip Sroubek, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2011 | Stabilizing and deblurring atmospheric turbulenceabstractA new approach is proposed to correct geometric distortion and reduce space and time-variant blur in videos that suffer from atmospheric turbulence. We first register the frames to suppress geometric deformation using a B-spline based non-rigid registration method. Next, a fusion process is carried out to produce an image from the registered frames, which can be viewed as being convolved with a space invariant near-diffraction-limited blur. Finally, a blind deconvolution algorithm is implemented to deblur the fused image. Experiments using real data illustrate that this approach is capable of alleviating blur and geometric deformation caused by turbulence, recovering details of the scene and significantly improving visual quality. Peyman Milanfar |
ICCP | 2 |
| 2011 | Patch-based locally optimal denoisingabstractIn our previous work [1], we formulated the fundamental limits of image denoising. In this paper, we propose a practical algorithm where the motivation is to realize a locally optimal denoising filter that achieves the lower bound. The proposed method is a patch-based Wiener filter that takes advantage of both geometrically and photometrically similar patches. The resultant approach has a nice statistical foundation while producing denoising results that are comparable to or exceeding the current state-of-the-art, both visually and quantitatively. Priyam Chatterjee, Peyman Milanfar |
ICIP | 2 |
| 2011 | Superfast superresolutionabstractWe propose a fast algorithm for solving the inverse problem of resolution enhancement (superresolution). Robustness is achieved by a non-linear regularizer and a method based on variable splitting is used to obtain an equivalent linear formulation. Special attention is paid to fast implementation using the Fourier transform. In particular, we show that a degradation operator (downsampling) can be implemented in the frequency domain and that all computations can be performed very efficiently without losing robustness. To our knowledge, this is the first attempt towards a very fast SR algorithm, which retains favorable edge-preserving properties of non-linear regularizers. Filip Sroubek, Jan Kamenický, Peyman Milanfar |
ICIP | 3 |
| 2011 | Restoration for weakly blurred and strongly noisy imagesabstractIn this paper we present an adaptive sharpening algorithm for restoration of an image which has been corrupted by mild blur, and strong noise. Most existing adaptive sharpening algorithms can not handle strong noise well due to the intrinsic contradiction between sharpening and de-noising. To solve this problem we propose an algorithm that is capable of capturing local image structure and sharpness, and adjusting sharpening accordingly so that it effectively combines denoising and sharpening together without either noise magnification or over-sharpening artifacts. It also uses structure information from the luminance channel to remove artifacts in the chrominance channels. Experiments illustrate that compared with other sharpening approaches, our method can produce state of the art results under practical imaging conditions. Peyman Milanfar |
WACV | 2 |
| 2011 | Action Recognition from One ExampleabstractWe present a novel action recognition method based on space-time locally adaptive regression kernels and the matrix cosine similarity measure. The proposed method uses a single example of an action as a query to find similar matches. It does not require prior knowledge about actions, foreground/background segmentation, or any motion estimation or tracking. Our method is based on the computation of novel space-time descriptors from the query video which measure the likeness of a voxel to its surroundings. Salient features are extracted from said descriptors and compared against analogous features from the target video. This comparison is done using a matrix generalization of the cosine similarity measure. The algorithm yields a scalar resemblance volume, with each voxel indicating the likelihood of similarity between the query video and all cubes in the target video. Using nonparametric significance tests by controlling the false discovery rate, we detect the presence and location of actions similar to the query video. High performance is demonstrated on challenging sets of action data containing fast motions, varied contexts, and complicated background. Further experiments on the Weizmann and KTH data sets demonstrate state-of-the-art performance in action categorization. Hae Jong Seo, Peyman Milanfar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Face Verification Using the LARK RepresentationabstractWe present a novel face representation based on locally adaptive regression kernel (LARK) descriptors. Our LARK descriptor measures a self-similarity based on “signal-induced distance” between a center pixel and surrounding pixels in a local neighborhood. By applying principal component analysis (PCA) and a logistic function to LARK consecutively, we develop a new binary-like face representation which achieves state-of-the-art face verification performance on the challenging benchmark “Labeled Faces in the Wild” (LFW) dataset. In the case where training data are available, we employ one-shot similarity (OSS) based on linear discriminant analysis (LDA). The proposed approach achieves state-of-the-art performance on both the unsupervised setting and the image restrictive training setting (72.23% and 78.90% verification rates), respectively, as a single descriptor representation, with no preprocessing step. As opposed to combined 30 distances which achieve 85.13%, we achieve comparable performance (85.1%) with only 14 distances while significantly reducing computational complexity. Hae Jong Seo, Peyman Milanfar |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | Practical Bounds on Image Denoising: From Estimation to InformationabstractRecently, in a previous work, we proposed a way to bound how well any given image can be denoised. The bound was computed directly from the noise-free image that was assumed to be available. In this work, we extend the formulation to the more practical case where no ground truth is available. We show that the parameters of the bounds, namely the cluster covariances and level of redundancy for patches in the image, can be estimated directly from the noise corrupted image. Further, we analyze the bounds formulation to show that these two parameters are interdependent and they, along with the bounds formulation as a whole, have a nice information-theoretic interpretation as well. The results are verified through a variety of well-motivated experiments. Priyam Chatterjee, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2011 | Removing Motion Blur With Space-Time ProcessingabstractAlthough spatial deblurring is relatively well understood by assuming that the blur kernel is shift invariant, motion blur is not so when we attempt to deconvolve on a frame-by-frame basis: this is because, in general, videos include complex, multilayer transitions. Indeed, we face an exceedingly difficult problem in motion deblurring of a single frame when the scene contains motion occlusions. Instead of deblurring video frames individually, a fully 3-D deblurring method is proposed in this paper to reduce motion blur from a single motion-blurred video to produce a high-resolution video in both space and time. Unlike other existing approaches, the proposed deblurring kernel is free from knowledge of the local motions. Most importantly, due to its inherent locally adaptive nature, the 3-D deblurring is capable of automatically deblurring the portions of the sequence, which are motion blurred, without segmentation and without adversely affecting the rest of the spatiotemporal domain, where such blur is not present. Our method is a two-step approach; first we upscale the input video in space and time without explicit estimates of local motions, and then perform 3-D deblurring to obtain the restored sequence. Hiroyuki Takeda, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2010 | Fundamental limits of image denoising: Are we there yet?abstractIn this paper, we study the fundamental performance limits of image denoising where the aim is to recover the original image from its noisy observation. Our study is based on a general class of estimators whose bias can be modeled to be affine. A bound on the performance in terms of mean squared error (MSE) of the recovered image is derived in a Bayesian framework. In this work, we assume that the original image is available, from which we learn the image statistics. Performances of some current state-of-the-art methods are compared to our MSE bounds for some commonly used experimental images. These show that some gain in denoising performance is yet to be achieved. Priyam Chatterjee, Peyman Milanfar |
ICASSP | 2 |
| 2010 | Visual saliency for automatic target detection, boundary detection, and image quality assessmentabstractWe present a visual saliency detection method and its applications. The proposed method does not require prior knowledge (learning) or any pre-processing step. Local visual descriptors which measure the likeness of a pixel to its surroundings are computed from an input image. Self-resemblance measured between local features results in a scalar map where each pixel indicates the statistical likelihood of saliency. Promising experimental results are illustrated for three applications: automatic target detection, boundary detection, and image quality assessment. Hae Jong Seo, Peyman Milanfar |
ICASSP | 2 |
| 2010 | Nonlinear kernel backprojection for computed tomographyabstractIn this paper, we propose a kernel backprojection method for computed tomography. The classical backprojection method estimates an unknown pixel value by the summation of the projection values with linear weights, while our kernel backprojection is a generalized version of the classic approach, in which we compute the weights from a kernel (weight) function. The generalization reveals that the performance of the backprojection operation strongly depends on the choice of the kernel, and a good choice of the kernels effectively suppresses both noise and streak artifacts while preserving major structures of the unknown phantom. The proposed method is a two-step procedure where we first compute a preliminary estimate of the phantom (a “pilot”), from which we compute the kernel weights. From these kernel weights we then reestimate the phantom, arriving at a much improved result. The experimental results show that our approach significantly enhances the backprojection operation not only numerically but also visually. Hiroyuki Takeda, Peyman Milanfar |
ICASSP | 2 |
| 2010 | Learning denoising bounds for noisy imagesabstractIn [1], we derived an expression for the fundamental limit to image denoising assuming that the noise-free image is available. In this paper, we propose an estimator for the bound on the mean squared error given only the noisy image and noise characteristics. To do this, we make use of an assortment of independently collected noise-free images from which prior information about the noisy image is learned. We show that even for reasonably low input signal-to-noise levels, our method can predict the denoising bound with accuracy. Priyam Chatterjee, Peyman Milanfar |
ICIP | 2 |
| 2010 | A no-reference image content metric and its application to denoisingabstractA no-reference image metric based on the singular value decomposition of local image gradients is proposed in this paper. This metric provides a quantitative measure of true image content, and reacts reasonably to both blur and random noise, so that it can be used in the automatic selection of parameters for image restoration algorithms, especially for denoising filters. Compared with GCV or SURE based approaches, this metric costs a small amount of computation, and does not require the noise to be Gaussian. Simulated and real data experiments demonstrated that our metric can capture the trend of quality change during the denoising process, and can yield parameters that show excellent visual performance in balancing between denoising and detail preservation. Peyman Milanfar |
ICIP | 2 |
| 2010 | Training-Free, Generic Object Detection Using Locally Adaptive Regression KernelsabstractWe present a generic detection/localization algorithm capable of searching for a visual object of interest without training. The proposed method operates using a single example of an object of interest to find similar matches, does not require prior knowledge (learning) about objects being sought, and does not require any preprocessing step or segmentation of a target image. Our method is based on the computation of local regression kernels as descriptors from a query, which measure the likeness of a pixel to its surroundings. Salient features are extracted from said descriptors and compared against analogous features from the target image. This comparison is done using a matrix generalization of the cosine similarity measure. We illustrate optimality properties of the algorithm using a naive-Bayes framework. The algorithm yields a scalar resemblance map, indicating the likelihood of similarity between the query and all patches in the target image. By employing nonparametric significance tests and nonmaxima suppression, we detect the presence and location of objects similar to the given query. The approach is extended to account for large variations in scale and rotation. High performance is demonstrated on several challenging data sets, indicating successful detection of objects in diverse contexts and under different imaging conditions. Hae Jong Seo, Peyman Milanfar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Is Denoising Dead?abstractImage denoising has been a well studied problem in the field of image processing. Yet researchers continue to focus attention on it to better the current state-of-the-art. Recently proposed methods take different approaches to the problem and yet their denoising performances are comparable. A pertinent question then to ask is whether there is a theoretical limit to denoising performance and, more importantly, are we there yet? As camera manufacturers continue to pack increasing numbers of pixels per unit area, an increase in noise sensitivity manifests itself in the form of a noisier image. We study the performance bounds for the image denoising problem. Our work in this paper estimates a lower bound on the mean squared error of the denoised result and compares the performance of current state-of-the-art denoising methods with this bound. We show that despite the phenomenal recent progress in the quality of denoising algorithms, some room for improvement still remains for a wide class of general images, and at certain signal-to-noise levels. Therefore, image denoising is not dead--yet. Priyam Chatterjee, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2010 | Automatic Parameter Selection for Denoising Algorithms Using a No-Reference Measure of Image ContentabstractAcross the field of inverse problems in image and video processing, nearly all algorithms have various parameters which need to be set in order to yield good results. In practice, usually the choice of such parameters is made empirically with trial and error if no "ground-truth" reference is available. Some analytical methods such as cross-validation and Stein's unbiased risk estimate (SURE) have been successfully used to set such parameters. However, these methods tend to be strongly reliant on restrictive assumptions on the noise, and also computationally heavy. In this paper, we propose a no-reference metric Q which is based upon singular value decomposition of local image gradient matrix, and provides a quantitative measure of true image content (i.e., sharpness and contrast as manifested in visually salient geometric features such as edges,) in the presence of noise and other disturbances. This measure 1) is easy to compute, 2) reacts reasonably to both blur and random noise, and 3) works well even when the noise is not Gaussian. The proposed measure is used to automatically and effectively set the parameters of two leading image denoising algorithms. Ample simulated and real data experiments support our claims. Furthermore, tests using the TID2008 database show that this measure correlates well with subjective quality evaluations for both blur and noise distortions. Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2009 | Detection of human actions from a single exampleabstractWe present an algorithm for detecting human actions based upon a single given video example of such actions. The proposed method is unsupervised, does not require learning, segmentation, or motion estimation. The novel features employed in our method are based on space-time locally adaptive regression kernels. Our method is based on the dense computation of so-called space-time local regression kernels (i.e. local descriptors) from a query video, which measure the likeness of a voxel to its spatio-temporal surroundings. Salient features are then extracted from these descriptors using principal components analysis (PCA). These are efficiently compared against analogous features from the target video using a matrix generalization of the cosine similarity measure. The algorithm yields a scalar resemblance volume; each voxel indicating the like-lihood of similarity between the query video and all cubes in the target video. By employing non-parametric significance tests and non-maxima suppression, we accurately detect the presence and location of actions similar to the given query video. High performance is demonstrated on a challenging set of action data indicating successful detection of multiple complex actions even in the presence of fast motions. Hae Jong Seo, Peyman Milanfar |
ICCV | 2 |
| 2009 | Optimal Registration Of Aliased Images Using Variable Projection With Applications To Super-ResolutionabstractAccurate registration of images is the most important and challenging aspect of multiframe image restoration problems such as super-resolution. The accuracy of super-resolution algorithms is quite often limited by the ability to register a set of low-resolution images. The main challenge in registering such images is the presence of aliasing. In this paper, we analyse the problem of jointly registering a set of aliased images and its relationship to super-resolution. We describe a statistically optimal approach to multiframe registration which exploits the concept of variable projections to achieve very efficient algorithms. Finally, we demonstrate how the proposed algorithm offers accurate estimation under various conditions when standard approaches fail to provide sufficient accuracy for super-resolution. M. Dirk Robinson, Sina Farsiu, Peyman Milanfar |
Comput. J. | 3 |
| 2009 | Clustering-Based Denoising With Locally Learned DictionariesabstractIn this paper, we propose K-LLD: a patch-based, locally adaptive denoising method based on clustering the given noisy image into regions of similar geometric structure. In order to effectively perform such clustering, we employ as features the local weight functions derived from our earlier work on steering kernel regression . These weights are exceedingly informative and robust in conveying reliable local structural information about the image even in the presence of significant amounts of noise. Next, we model each region (or cluster)-which may not be spatially contiguous-by "learning" a best basis describing the patches within that cluster using principal components analysis. This learned basis (or "dictionary") is then employed to optimally estimate the underlying pixel values using a kernel regression framework. An iterated version of the proposed algorithm is also presented which leads to further performance enhancements. We also introduce a novel mechanism for optimally choosing the local patch size for each cluster using Stein's unbiased risk estimator (SURE). We illustrate the overall algorithm's capabilities with several examples. These indicate that the proposed method appears to be competitive with some of the most recently published state of the art denoising methods. Priyam Chatterjee, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2009 | Generalizing the Nonlocal-Means to Super-Resolution ReconstructionabstractSuper-resolution reconstruction proposes a fusion of several low-quality images into one higher quality result with better optical resolution. Classic super-resolution techniques strongly rely on the availability of accurate motion estimation for this fusion task. When the motion is estimated inaccurately, as often happens for nonglobal motion fields, annoying artifacts appear in the super-resolved outcome. Encouraged by recent developments on the video denoising problem, where state-of-the-art algorithms are formed with no explicit motion estimation, we seek a super-resolution algorithm of similar nature that will allow processing sequences with general motion patterns. In this paper, we base our solution on the Nonlocal-Means (NLM) algorithm. We show how this denoising method is generalized to become a relatively simple super-resolution algorithm with no explicit motion estimation. Results on several test movies show that the proposed method is very successful in providing super-resolution on general sequences. Matan Protter, Michael Elad, Hiroyuki Takeda, Peyman Milanfar |
IEEE Trans. Image Process. | 4 |
| 2009 | Super-Resolution Without Explicit Subpixel Motion EstimationabstractThe need for precise (subpixel accuracy) motion estimates in conventional super-resolution has limited its applicability to only video sequences with relatively simple motions such as global translational or affine displacements. In this paper, we introduce a novel framework for adaptive enhancement and spatiotemporal upscaling of videos containing complex activities without explicit need for accurate motion estimation. Our approach is based on multidimensional kernel regression, where each pixel in the video sequence is approximated with a 3-D local (Taylor) series, capturing the essential local behavior of its spatiotemporal neighborhood. The coefficients of this series are estimated by solving a local weighted least-squares problem, where the weights are a function of the 3-D space-time orientation in the neighborhood. As this framework is fundamentally based upon the comparison of neighboring pixels in both space and time, it implicitly contains information about the local motion of the pixels across time, therefore rendering unnecessary an explicit computation of motions of modest size. The proposed approach not only significantly widens the applicability of super-resolution methods to a broad variety of video sequences containing complex motions, but also yields improved overall performance. Using several examples, we illustrate that the developed algorithm has super-resolution capabilities that provide improved optical resolution in the output, while being able to work on general input video with essentially arbitrary motion. Hiroyuki Takeda, Peyman Milanfar, Matan Protter, Michael Elad |
IEEE Trans. Image Process. | 2 |
| 2008 | Video denoising using higher order optimal space-time adaptationabstractThe optimal spatial adaptation (OSA) method [1] proposed by Boulanger and Kervrann has proven to be quite effective for spatially adaptive image denoising. This method, in addition to extending the Non-Local Means(NLM) method of [2], employs an iteratively growing window scheme, and a local estimate of the mean square error to very effectively remove noise from images. By adopting an iteratively growing space-time window, the method was recently extended to 3-D for video denoising in [3]. In the present paper, we demonstrate a simple, but effective improvement on the OSA method in both 2- and 3-D. We demonstrate that the OSA implicitly relies on a locally constant model of the underlying signal. Thereby, removing this constraint and introducing the possibility of higher order local regression models, we arrive at a relatively simple modification that results in an improvement in performance. While this improvement is observed in both 2-D and 3-D, we concentrate on demonstrating it in 3-D for the application of video denoising. Hae Jong Seo, Peyman Milanfar |
ICASSP | 2 |
| 2008 | Using local regression kernels for statistical object detectionabstractWe present a novel approach to the problem of detection of visual similarity between a template image, and patches in a given image. The method is based on the computation of a local kernel from the template, which measures the likeness of a pixel to its surroundings. This kernel is then used as a descriptor from which features are extracted and compared against analogous features from the target image. Comparison of the features extracted is carried out using canonical correlations analysis. The overall algorithm yields a scalar resemblance map (RM) which indicates the statistical likelihood of similarity between a given template and all target patches in an image being examined. Performing a statistical test on the resulting RM identifies similar objects with high accuracy and is robust to various challenging conditions such as partial occlusion, and illumination change. Hae Jong Seo, Peyman Milanfar |
ICIP | 2 |
| 2008 | Spatio-temporal video interpolation and denoising using motion-assisted steering kernel (MASK) regressionabstractIn this paper, we extend a (2-D) data-adaptive steering kernel regression framework for image processing to a (3-D) spatio-temporal framework for processing video. In particular, we propose a motion- assisted steering kernel (MASK) suitable for interpolating video data spatially, temporally, or spatio-temporally, and for video noise reduction. We present an algorithm for multi-frame interpolation and reconstruction of video data, and present several simulation results on synthetic and real video data. Comparisons between single-frame and multi-frame kernel regression and with other methods demonstrate the effectiveness of our approach. Hiroyuki Takeda, Peter van Beek, Peyman Milanfar |
ICIP | 3 |
| 2008 | On Iterative Regularization and Its ApplicationabstractMany existing techniques for image restoration can be expressed in terms of minimizing a particular cost function. Iterative regularization methods are a novel variation on this theme where the cost function is not fixed, but rather refined iteratively at each step. This provides an unprecedented degree of control over the tradeoff between the bias and variance of the image estimate, which can result in improved overall estimation error. This useful property, along with the provable convergence properties of the sequence of estimates produced by these iterative regularization methods lend themselves to a variety of useful applications. In this paper, we introduce a general set of iterative regularization methods, discuss some of their properties and applications, and include examples to illustrate them. Michael R. Charest, Peyman Milanfar |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Deblurring Using Regularized Locally Adaptive Kernel RegressionabstractKernel regression is an effective tool for a variety of image processing tasks such as denoising and interpolation [1]. In this paper, we extend the use of kernel regression for deblurring applications. In some earlier examples in the literature, such nonparametric deblurring was suboptimally performed in two sequential steps, namely denoising followed by deblurring. In contrast, our optimal solution jointly denoises and deblurs images. The proposed algorithm takes advantage of an effective and novel image prior that generalizes some of the most popular regularization techniques in the literature. Experimental results demonstrate the effectiveness of our method. Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar |
IEEE Trans. Image Process. | 3 |
| 2007 | Multi-Scale Statistical Detection and Ballistic Imaging Through Turbid MediaabstractWe exploit recent advances in the physical design of fast optical systems which enable active imaging with "ballistic" light. In this modality, fast bursts of optical energy are propagated into a medium, and the ballistic component of light (which travels with minimal diffusive distortion) is detected after transmission through the target and the medium. To improve the detection rate of the common single pixel optimal detectors, we exploit sampling at a diversity of locations in space, and develop a multi-scale algorithm based upon the generalized likelihood ratio test (GLRT) framework, which takes advantage of the spatial correlation of nearby samples. Experimental results show that objects of different size and shape that are completely unrecognizable using the common single pixel detection techniques, are detectable with very high accuracy using the said multi-scale GLRT technique. Sina Farsiu, Peyman Milanfar |
ICIP (3) | 2 |
| 2007 | Mask Design for Optical Microlithography-An Inverse Imaging ProblemabstractIn all imaging systems, the forward process introduces undesirable effects that cause the output signal to be a distorted version of the input. A typical example is of course the blur introduced by the aperture. When the input to such systems can be controlled, prewarping techniques can be employed which consist of systematically modifying the input such that it (at least approximately) cancels out (or compensates for) the process losses. In this paper, we focus on the optical proximity correction mask design problem for "optical microlithography," a process similar to photographic printing used for transferring binary circuit patterns onto silicon wafers. We consider the idealized case of an incoherent imaging system and solve an inverse problem which is an approximation of the real-world optical lithography problem. Our algorithm is based on pixel-based mask representation and uses a continuous function formulation. We also employ the regularization framework to control the tone and complexity of the synthesized masks. Finally, we discuss the extension of our framework to coherent and (the more practical) partially coherent imaging systems. Amyn Poonawala, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2007 | Kernel Regression for Image Processing and ReconstructionabstractIn this paper, we make contact with the field of nonparametric statistics and present a development and generalization of tools and results for use in image processing and reconstruction. In particular, we adapt and expand kernel regression ideas for use in image denoising, upscaling, interpolation, fusion, and more. Furthermore, we establish key relationships with some popular existing methods and show how several of these algorithms, including the recently popularized bilateral filter, are special cases of the proposed framework. The resulting algorithms and analyses are amply illustrated with practical examples. Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar |
IEEE Trans. Image Process. | 3 |
| 2006 | Robust Kernel Regression for Restoration and Reconstruction of Images from Sparse Noisy DataabstractWe introduce a class of robust non-parametric estimation methods which are ideally suited for the reconstruction of signals and images from noise-corrupted or sparsely collected samples. The filters derived from this class are locally adapted kernels which take into account both the local density of the available samples, and the actual values of these samples. As such, they are automatically steered and adapted to both the given sampling "geometry", and the samples' "radiometry". As the framework we proposed does not rely upon specific assumptions about noise or sampling distributions, it is applicable to a wide class of problems including efficient image upscaling, high quality reconstruction of an image from as little as 15% of its (irregularly sampled) pixels, super-resolution from noisy and under-determined data sets, state of the art denoising of images corrupted by Gaussian and other noise, effective removal of compression artifacts; and more. Hiroyuki Takeda, Sina Farsiu, Peyman Milanfar |
ICIP | 3 |
| 2006 | Multiframe demosaicing and super-resolution of color imagesabstractIn the last two decades, two related categories of problems have been studied independently in image restoration literature: super-resolution and demosaicing. A closer look at these problems reveals the relation between them, and, as conventional color digital cameras suffer from both low-spatial resolution and color-filtering, it is reasonable to address them in a unified context. In this paper, we propose a fast and robust hybrid method of super-resolution and demosaicing, based on a maximum a posteron estimation technique by minimizing a multiterm cost function. The L1 norm is used for measuring the difference between the projected estimate of the high-resolution image and each low-resolution image, removing outliers in the data and errors due to possibly inaccurate motion estimation. Bilateral regularization is used for spatially regularizing the luminance component, resulting in sharp edges and forcing interpolation along the edges and not across them. Simultaneously, Tikhonov regularization is used to smooth the chrominance components. Finally, an additional regularization term is used to force similar edge location and orientation in different color channels. We show that the minimization of the total cost function is relatively easy and fast. Experimental results on synthetic and real data sets confirm the effectiveness of our method. Sina Farsiu, Michael Elad, Peyman Milanfar |
IEEE Trans. Image Process. | 3 |
| 2006 | Statistical performance analysis of super-resolutionabstractRecently, there has been a great deal of work developing super-resolution algorithms for combining a set of low-quality images to produce a set of higher quality images. Either explicitly or implicitly, such algorithms must perform the joint task of registering and fusing the low-quality image data. While many such algorithms have been proposed, very little work has addressed the performance bounds for such problems. In this paper, we analyze the performance limits from statistical first principles using Cramér-Rao inequalities. Such analysis offers insight into the fundamental super-resolution performance bottlenecks as they relate to the subproblems of image registration, reconstruction, and image restoration. M. Dirk Robinson, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2006 | Statistical and Information-Theoretic Analysis of Resolution in ImagingabstractIn this paper, some detection-theoretic, estimation-theoretic, and information-theoretic methods are investigated to analyze the problem of determining resolution limits in imaging systems. The canonical problem of interest is formulated based on a model of the blurred image of two closely spaced point sources of unknown brightness. To quantify a measure of resolution in statistical terms, the following question is addressed: "What is the minimum detectable separation between two point sources at a given signal-to-noise ratio (SNR), and for prespecified probabilities of detection and false alarm (Pdand Pf)?". Furthermore, asymptotic performance analysis for the estimation of the unknown parameters is carried out using the Crameacuter-Rao bound. Although similar approaches to this problem (for one-dimensional (1-D) and oversampled signals) have been presented in the past, the analyzes presented in this paper are carried out for the general two-dimensional (2-D) model and general sampling scheme. In particular the case of under-Nyquist (aliased) images is studied. Furthermore, the Kullback-Liebler distance is derived to further confirm the earlier results and to establish a link between the detection-theoretic approach and Fisher information. To study the effects of variation in point spread function (PSF) and model mismatch, a perturbation analysis of the detection problem is presented as well Morteza Shahram, Peyman Milanfar |
IEEE Trans. Inf. Theory | 2 |
| 2005 | Improved spectral analysis of nearby tones using local detectorsabstractThis paper concerns the problem of resolvability power in the frequency domain. The canonical case of interest is to distinguish whether the received noise-corrupted signal is a single-frequency sinusoid or a two-frequency sinusoid, where the amplitudes, phases and frequencies are unknown to the receiver. Using a model-based hypothesis testing approach, we quantify a measure of attainable resolution between sinusoids with nearby frequencies, in the presence of noise. An explicit relationship is derived for the minimum detectable difference between the frequencies of two tones, for any particular false alarm and detection rate, and at a given SNR. An associated algorithm is proposed that produces significantly better performance compared to the standard subspace-based methods like MUSIC and can be effectively used in practice as a postprocessing step for the existing spectral estimation methods. Morteza Shahram, Peyman Milanfar |
ICASSP (4) | 2 |
| 2005 | Bias minimizing filter design for gradient-based image registration
M. Dirk Robinson, Peyman Milanfar |
Signal Process. Image Commun. | 2 |
| 2005 | Variable projection for near-optimal filtering in low bit-rate block coders
Yaakov Tsaig, Michael Elad, Peyman Milanfar, Gene H. Golub |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Improved high-definition video by encoding at an intermediate resolutionabstractIn this paper, we consider the compression of high-definition video sequences for bandwidth sensitive applications. We show that down-sampling the image sequence prior to encoding and then up-sampling the decoded frames increases compression efficiency. This is particularly true at lower bit-rates, as direct encoding of the high-definition sequence requires a large number of blocks to be signaled. We survey previous work that combines a resolution change and compression mechanism. We then illustrate the success of our proposed approach through simulations. Both MPEG-2 and H.264 scenarios are considered. Given the benefits of the approach, we also interpret the results within the context of traditional spatial scalability. C. Andrew Segall, Michael Elad, Peyman Milanfar, Richard Webb, Chad Fogg |
VCIP | 3 |
| 2004 | Trained detection of buried mines in SAR images via the deflection-optimal criterionabstractIn this paper, we apply a deflection-optimal linear-quadratic detector to the detection of buried mines in images formed by a forward-looking ground-penetrating synthetic aperture radar. The detector is a linear-quadratic form that maximizes the output SNR (deflection), and its parameters are estimated from a set of training data. We show that this detector is useful when the signal to be detected is expected to be stochastic, with an unknown distribution, and when only a small set of training data is available to estimate its statistics. The detector structure can be understood in terms of the singular value decomposition; the statistical variations of the target signature are modeled using a compact set of orthogonal "eigenmodes" (or principal components) of the training dataset. Because only the largest eigenvalues and associated eigenvectors contribute, statistical variations that are underrepresented in the training data do not significantly corrupt the detector performance. The resulting detection algorithm is tested on data that are not in the training set, which has been collected at government test sites, and the algorithm performance is reported. Russell B. Cosgrove, Peyman Milanfar, Joel Kositsky |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2004 | Fast and robust multiframe super resolutionabstractSuper-resolution reconstruction produces one or a set of high-resolution images from a set of low-resolution images. In the last two decades, a variety of super-resolution methods have been proposed. These methods are usually very sensitive to their assumed model of data and noise, which limits their utility. This paper reviews some of these methods and addresses their short-comings. We propose an alternate approach using L1 norm minimization and robust regularization based on a bilateral prior to deal with different data and noise models. This computationally inexpensive method is robust to errors in motion and blur estimation and results in images with sharp edges. Simulation results confirm the effectiveness of our method and demonstrate its superiority to other super-resolution methods. Sina Farsiu, M. Dirk Robinson, Michael Elad, Peyman Milanfar |
IEEE Trans. Image Process. | 4 |
| 2004 | Fundamental performance limits in image registrationabstractThe task of image registration is fundamental in image processing. It often is a critical preprocessing step to many modern image processing and computer vision tasks, and many algorithms and techniques have been proposed to address the registration problem. Often, the performances of these techniques have been presented using a variety of relative measures comparing different estimators, leaving open the critical question of overall optimality. In this paper, we present the fundamental performance limits for the problem of image registration as derived from the Cramer-Rao inequality. We compare the experimental performance of several popular methods with respect to this performance bound, and explain the fundamental tradeoff between variance and bias inherent to the problem of image registration. In particular, we derive and explore the bias of the popular gradient-based estimator showing how widely used multiscale methods for improving performance can be explained with this bias expression. Finally, we present experimental simulations showing the general rule-of-thumb performance limits for gradient-based image registration techniques. M. Dirk Robinson, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2004 | Imaging below the diffraction limit: a statistical analysisabstractThe present paper is concerned with the statistical analysis of the resolution limit in a so-called "diffraction-limited" imaging system. The canonical case study is that of incoherent imaging of two closely-spaced sources of possibly unequal brightness. The objective is to study how far beyond the classical Rayleigh limit of resolution one can reach at a given signal to noise ratio. The analysis uses tools from statistical detection and estimation theory. Specifically, we will derive explicit relationships between the minimum detectable distance between two closely-spaced point sources imaged incoherently at a given SNR. For completeness, asymptotic performance analysis for the estimation of the unknown parameters is carried out using the Cramér-Rao bound. To gain maximum intuition, the analysis is carried out in one dimension, but can be well extended to the two-dimensional case and to more practical models. Morteza Shahram, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2003 | Fast and robust super-resolutionabstractIn the last two decades, many papers have been published, proposing a variety methods of multiframe resolution enhancement. These methods are usually very sensitive to their assumed model of data and noise, which limits their utility. This paper reviews some of these methods and addresses their shortcomings. We propose a different implementation using L/sub 1/ norm minimization and robust regularization to deal with different data and noise models. This computationally inexpensive method is robust to errors in motion and blur estimation, and results in sharp edges. Simulation results confirm the effectiveness of our method and demonstrate its superiority to other robust super-resolution methods. Sina Farsiu, M. Dirk Robinson, Michael Elad, Peyman Milanfar |
ICIP (2) | 4 |
| 2003 | Fundamental performance limits in image registrationabstractWhile many algorithms have been developed to solve the problem of image registration, their performance has typically been evaluated only by comparing one method with another often in an ad-hoc manner. We propose a statistical performance measure based on the mean square error (MSE) and explore the performance bounds using the Cramer-Rao inequality. We show how these performance bounds depend on image content under observation. By analyzing these bounds we provide insight into the inherent tradeoff between bias and variance found in all image registration algorithms. Specifically, we derive a functional expression for the bias inherent in the popular class of gradient-based image registration algorithms. M. Dirk Robinson, Peyman Milanfar |
ICIP (2) | 2 |
| 2003 | Optimal framework for low bit-rate block codersabstractBlock coders are among the most common compression tools available for still images and video sequences. Their low computational complexity along with their good performance make them a popular choice for compression of natural images. Yet, at low bit-rates, block coders introduce visually annoying artifacts into the image. One approach that alleviates this problem is to downsample the image, apply the coding algorithm, and interpolate back to the original resolution. In this paper, we consider the use of optimal decimation and interpolation filters in this scheme. We first consider only optimization of the interpolation filter, by formulating the problem as least-squares minimization. We then consider the joint optimization over both the decimation and the interpolation filters, using the variable projection method. The experimental results presented clearly exhibit a significant improvement over other approaches. Yaakov Tsaig, Michael Elad, Gene H. Golub, Peyman Milanfar |
ICIP (2) | 4 |
| 2003 | Reconstruction of Convex Bodies from Brightness Functions
Richard J. Gardner, Peyman Milanfar |
Discret. Comput. Geom. | 2 |
| 2002 | A statistical analysis of diffraction-limited imagingabstractThe Rayleigh criterion is generally regarded as a fundamental limit and due to its practical accuracy in predicting the performance of optical imaging systems, it has unfortunately become accepted as a de-facto physical law. We show that this limit is simply a very good rule of thumb, which under proper conditions typically related to the signal-to-noise (SNR) of the sensor, can be overcome. While we do not discuss specific methodology for improving resolution, it is the aim of this work to explore how far beyond the Rayleigh limit one can go, and to identify the theoretical limits to such resolution enhancement. Peyman Milanfar, Ali Shakouri |
ICIP (1) | 1 |
| 2001 | A computationally efficient superresolution image reconstruction algorithmabstractSuperresolution reconstruction produces a high-resolution image from a set of low-resolution images. Previous iterative methods for superresolution had not adequately addressed the computational and numerical issues for this ill-conditioned and typically underdetermined large scale problem. We propose efficient block circulant preconditioners for solving the Tikhonov-regularized superresolution problem by the conjugate gradient method. We also extend to underdetermined systems the derivation of the generalized cross-validation method for automatic calculation of regularization parameters. The effectiveness of our preconditioners and regularization techniques is demonstrated with superresolution results for a simulated sequence and a forward looking infrared (FLIR) camera image sequence. Nhat Nguyen, Peyman Milanfar, Gene H. Golub |
IEEE Trans. Image Process. | 2 |
| 2001 | Efficient generalized cross-validation with applications to parametric image restoration and resolution enhancementabstractIn many image restoration/resolution enhancement applications, the blurring process, i.e., point spread function (PSF) of the imaging system, is not known or is known only to within a set of parameters. We estimate these PSF parameters for this ill-posed class of inverse problem from raw data, along with the regularization parameters required to stabilize the solution, using the generalized cross-validation method (GCV). We propose efficient approximation techniques based on the Lanczos algorithm and Gauss quadrature theory, reducing the computational complexity of the GCV. Data-driven PSF and regularization parameter estimation experiments with synthetic and real image sequences are presented to demonstrate the effectiveness and robustness of our method. Nhat Nguyen, Peyman Milanfar, Gene H. Golub |
IEEE Trans. Image Process. | 2 |
| 2000 | An Efficient Wavelet-Based Algorithm for Image SuperresolutionabstractSuperresolution produces high quality, high resolution images from a set of degraded, low resolution frames. We present a new and efficient wavelet-based algorithm for image superresolution. The algorithm is a combination of interpolation and restoration processes. Unlike previous work, our method exploits the interlaced sampling structure in the low resolution data. Numerical experiments and analysis demonstrate the effectiveness of our approach and illustrate why the computational complexity only doubles for 2-D superresolution versus 1-D case. Nhat Nguyen, Peyman Milanfar |
ICIP | 2 |
| 1999 | Preconditioners for regularized image superresolutionabstractSuperresolution reconstruction produces a high resolution image from a set of low resolution images. Previous work on superresolution had not adequately addressed the computational issues for this problem. In this paper, we propose efficient block circulant preconditioners for solving the regularized superresolution problem by conjugate gradients. The effectiveness of the preconditioners is demonstrated with superresolution results for a simulated image sequence and a FLIR image sequence. Nhat Nguyen, Gene H. Golub, Peyman Milanfar |
ICASSP | 3 |
| 1999 | Two-dimensional matched filtering for motion estimationabstractIn this work, we describe a frequency domain technique for the estimation of multiple superimposed motions in an image sequence. The least-squares optimum approach involves the computation of the three-dimensional (3-D) Fourier transform of the sequence, followed by the detection of one or more planes in this domain with high energy concentration. We present a more efficient algorithm, based on the properties of the Radon transform and the two-dimensional (2-D) fast Fourier transform, which can sacrifice little performance for significant computational savings. We accomplish the motion detection and estimation by designing appropriate matched filters. The performance is demonstrated on two image sequences. Peyman Milanfar |
IEEE Trans. Image Process. | 1 |
| 1999 | A model of the effect of image motion in the Radon transform domainabstractOne of the most fundamental properties of the Radon (projection) transform is that shifting of the image results in shifted projections. This useful property relates translational motion in the image to simple displacement in the projections. It is far from clear, however, how more general types of motion in the image domain will be manifested in the projections. In this paper, we present a model for this phenomenon in the general case; namely, we develop a generalization of the shift property of the Radon transform. We study various properties of the apparent projected motion implied by the model, and study the case of affine motion in particular. We also present illustrative examples, and briefly discuss the inverse problem implied by the forward model developed herein, along with some possible applications. Peyman Milanfar |
IEEE Trans. Image Process. | 1 |
| 1998 | Motion from Projections: A Foward ModelabstractOne of the most fundamental properties of the Radon (projection) transform is that shifting of the image results in shifted projections. This useful property relates translational motion in the image to simple displacement in the projections. It is far from clear, however, how more general types of motion in the image domain will be manifested in the projections. In this paper, we present a model for this phenomenon in the general case; namely, we develop a generalization of the shift property of the Radon transform. We study various properties of the apparent projected motion applied by the model, and study the case of affine motion in particular. We also present illustrative examples, and briefly discuss the inverse problem implied by the forward model developed herein, along with some possible applications. Peyman Milanfar |
ICIP (2) | 1 |
| 1996 | On the hough transform of a polygon
Peyman Milanfar |
Pattern Recognit. Lett. | 1 |
| 1996 | A moment-based variational approach to tomographic reconstructionabstractWe describe a variational framework for the tomographic reconstruction of an image from the maximum likelihood (ML) estimates of its orthogonal moments. We show how these estimated moments and their (correlated) error statistics can be computed directly, and in a linear fashion from given noisy and possibly sparse projection data. Moreover, thanks to the consistency properties of the Radon transform, this two-step approach (moment estimation followed by image reconstruction) can be viewed as a statistically optimal procedure. Furthermore, by focusing on the important role played by the moments of projection data, we immediately see the close connection between tomographic reconstruction of nonnegative valued images and the problem of nonparametric estimation of probability densities given estimates of their moments. Taking advantage of this connection, our proposed variational algorithm is based on the minimization of a cost functional composed of a term measuring the divergence between a given prior estimate of the image and the current estimate of the image and a second quadratic term based on the error incurred in the estimation of the moments of the underlying image from the noisy projection data. We show that an iterative refinement of this algorithm leads to a practical algorithm for the solution of the highly complex equality constrained divergence minimization problem. We show that this iterative refinement results in superior reconstructions of images from very noisy data as compared with the classical filtered back-projection (FBP) algorithm. Peyman Milanfar, W. Clem Karl, Alan S. Willsky |
IEEE Trans. Image Process. | 1 |
| 1994 | Moment-Based Geometric Image ReconstructionabstractIn this paper, we discuss two interesting instantiations of the moment problem in image processing. The first involves the estimation of moments of an image indirectly from projections, and the reconstruction of the image from these moments. The second relates the reconstruction of binary polygons from moments to well-known algorithm in array signal processing. Through these examples, we place the moment problem into a geometric perspective and illustrate how this perspective leads to a number of interesting practical applications in image processing and other fields.> Peyman Milanfar, W. Clem Karl, Alan S. Willsky |
ICIP (2) | 1 |
| 1994 | Modeling and Estimation for a Class of Multiresolution Random FieldsabstractDiscusses a class of multiresolution models of random fields based on a generalization of the midpoint deflection construction of the 1D Brownian motion. The authors then present least squares (LS) algorithms for the estimation of parameters which define these models and hence provide a framework for synthesizing and analyzing images with fractal-like properties such as those found in statistical representation of natural terrain and other geophysical phenomena. The authors also briefly discuss possible applications of this modeling framework to target detection in images.> Peyman Milanfar, Robert R. Tenney, Robert B. Washburn, Alan S. Willsky |
ICIP (3) | 1 |
| 1994 | Reconstructing Binary Polygonal Objects from Projections: A Statistical View
Peyman Milanfar, W. Clem Karl, Alan S. Willsky |
CVGIP Graph. Model. Image Process. | 1 |