EDBT 2026 Demo / reviewers in the wild / expert
Paolo Favaro
dblp:02/4162
· DBLP profile ↗
105ranked-venue papers
17as first author
27since 2021 · last 2025
0000-0003-3546-8247ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 17 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 83 · 14 first-author · 20 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CAGE: Unsupervised Visual Composition and Animation for Controllable Video GenerationabstractThe field of video generation has expanded significantly in recent years, with controllable and compositional video generation garnering considerable interest. Most methods rely on leveraging annotations such as text, objects' bounding boxes, and motion cues, which require substantial human effort and thus limit their scalability. In contrast, we address the challenge of controllable and compositional video generation without any annotations by introducing a novel unsupervised approach. Our model is trained from scratch on a dataset of unannotated videos. At inference time, it can compose plausible novel scenes and animate objects by placing object parts at the desired locations in space and time. The core innovation of our method lies in the unified control format and the training process, where video generation is conditioned on a randomly selected subset of pre-trained self-supervised local features. This conditioning compels the model to learn how to inpaint the missing information in the video both spatially and temporally, thereby learning the inherent compositionality of a scene and the dynamics of moving objects. The abstraction level and the imposed invariance of the conditioning input to minor visual perturbations enable control over object motion by simply using the same features at all the desired future locations. We call our model CAGE, which stands for visual Composition and Animation for video GEneration. We conduct extensive experiments to validate the effectiveness of CAGE across various scenarios, demonstrating its capability to accurately follow the control and to generate high-quality videos that exhibit coherent scene composition and realistic animation. Aram Davtyan, Sepehr Sameni, Björn Ommer, Paolo Favaro |
AAAI | 4 |
| 2025 | GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition ControlabstractWe present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced1. Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Pedro M. B. Rezende, Yasaman Haghighi, David Brüggemann, Isinsu Katircioglu, Xiaoran Chen, Marco Cannici, Elie Aljalbout, Botao Ye, Xi Wang 0021, Aram Davtyan, Mathieu Salzmann, Davide Scaramuzza 0001, Marc Pollefeys, Paolo Favaro, Alexandre Alahi |
CVPR | 19 |
| 2025 | Adapting Dense Matching for Homography Estimation with Grid-based AccelerationabstractCurrent deep homography estimation methods are typically constrained to processing low-resolution image pairs due to network architecture and computational limitations. For high-resolution images, downsampling is often required, which can greatly degrade estimation accuracy. In contrast, image matching methods, which match pixels and compute homography from correspondences, provide greater resolution flexibility. So in this work, we revisit the traditional image matching paradigm for homography estimation and propose GFNet, a Grid Flow regression Network that adapts the high-accuracy dense matching framework for homography estimation while enhancing efficiency through a grid-based strategy—estimating flow only over a coarse grid by leveraging homography’s global smoothness. We demonstrate the effectiveness of GFNet on a wide range of experiments on multiple datasets, including the common scene MSCOCO, multimodal datasets VIS-IR and GoogleMap, and the dynamic scene VIRAT. Notably, on 448×448 GoogleMap, GFNet achieves an improvement of +13.5% in auc@3 while reducing MACs by ~47% compared to the SOTA dense matching method. Additionally, it shows a 1.8× improvement in auc@3 over the SOTA deep homography method. Code is available at https://github.com/KN-Zhang/GFNet. Kaining Zhang, Yuxin Deng 0002, Jiayi Ma 0001, Paolo Favaro |
CVPR | 4 |
| 2025 | Diffusion Image Prior
Hamadi Chihaoui, Paolo Favaro |
ICCV | 2 |
| 2025 | Faster Inference of Flow-Based Generative Models via Improved Data-Noise CouplingabstractConditional Flow Matching (CFM), a simulation-free method for training continuous normalizing flows, provides an efficient alternative to diffusion models for key tasks like image and video generation. The performance of CFM in solving these tasks depends on the way data is coupled with noise. A recent approach uses minibatch optimal transport (OT) to reassign noise-data pairs in each training step to streamline sampling trajectories and thus accelerate inference. However, its optimization is restricted to individual minibatches, limiting its effectiveness on large datasets. To address this shortcoming, we introduce LOOM-CFM (Looking Out Of Minibatch-CFM), a novel method to extend the scope of minibatch OT by preserving and optimizing these assignments across minibatches over training time. Our approach demonstrates consistent improvements in the sampling speed-quality trade-off across multiple datasets. LOOM-CFM also enhances distillation initialization and supports high-resolution synthesis in latent space training. Aram Davtyan, Leello Tadesse Dadi, Volkan Cevher, Paolo Favaro |
ICLR | 4 |
| 2025 | KOALA++: Efficient Kalman-Based Optimization with Gradient-Covariance ProductsabstractWe propose KOALA++, a scalable Kalman-based optimization algorithm that explicitly models structured gradient uncertainty in neural network training. Unlike second-order methods, which rely on expensive second order gradient calculation, our method directly estimates the parameter covariance matrix by recursively updating compact gradient covariance products. This design improves upon the original KOALA framework that assumed diagonal covariance by implicitly capturing richer uncertainty structure without storing the full covariance matrix and avoiding large matrix inversions. Across diverse tasks, including image classification and language modeling, KOALA++ achieves accuracy on par or better than state-of-the-art second-order optimizers while maintaining the efficiency of first-order methods. Zixuan Xia, Aram Davtyan, Paolo Favaro |
NeurIPS | 3 |
| 2024 | Sparse 3D Reconstruction via Object-Centric Ray SamplingabstractWe propose a novel method for 3D object reconstruction from a sparse set of views captured from a 360-degree calibrated camera rig. We represent the object surface through a hybrid model that uses both an MLP-based neural representation and a triangle mesh. A key contribution in our work is a novel object-centric sampling scheme of the neural representation, where rays are shared among all views. This efficiently concentrates and reduces the number of samples used to update the neural model at each iteration. This sampling scheme relies on the mesh representation to ensure also that samples are well-distributed along its normals. The rendering is then performed efficiently by a differentiable renderer. We demonstrate that this sampling scheme results in a more effective training of the neural representation, does not require the additional supervision of segmentation masks, yields state of the art 3D reconstructions, and works with sparse views on the Google’s Scanned Objects, Tank and Temples and MVMC Car datasets. Code available at: https://github.com/llukmancerkezi/ROSTER Llukman Cerkezi, Paolo Favaro |
3DV | 2 |
| 2024 | Learn the Force We Can: Enabling Sparse Motion Control in Multi-Object Video GenerationabstractWe propose a novel unsupervised method to autoregressively generate videos from a single frame and a sparse motion input. Our trained model can generate unseen realistic object-to-object interactions. Although our model has never been given the explicit segmentation and motion of each object in the scene during training, it is able to implicitly separate their dynamics and extents. Key components in our method are the randomized conditioning scheme, the encoding of the input motion control, and the randomized and sparse sampling to enable generalization to out of distribution but realistic correlations. Our model, which we call YODA, has therefore the ability to move objects without physically touching them. Through extensive qualitative and quantitative evaluations on several datasets, we show that YODA is on par with or better than state of the art video generation prior work in terms of both controllability and video quality. Aram Davtyan, Paolo Favaro |
AAAI | 2 |
| 2024 | Masked and Shuffled Blind Spot Denoising for Real-World ImagesabstractWe introduce a novel approach to single image denoising based on the Blind Spot Denoising principle, which we call MAsked and SHuffled Blind Spot Denoising (MASH). We focus on the case of correlated noise, which often plagues real images. MASH is the result of a careful analysis to determine the relationships between the level of blindness (masking) of the input and the (unknown) noise correlation. Moreover, we introduce a shuffling technique to weaken the local correlation of noise, which in turn yields an additional denoising performance improvement. We evaluate MASH via extensive experiments on real-world noisy image datasets. We demonstrate state-of-the-art results compared to existing self-supervised denoising methods. Website: https://hamadichihaoui.github.io/mash. Hamadi Chihaoui, Paolo Favaro |
CVPR | 2 |
| 2024 | When Self-Supervised Pre-Training Meets Single Image DenoisingabstractWe present a self-supervised pre-training scheme for single image denoising based on a novel pretext task. Our work is inspired by the success of self-supervised learning (SSL) methods in transfer learning. These methods have been shown to be extremely effective when used to pre-train a model that is then fine-tuned on small datasets. As pretext task, we propose to train a denoising network on patches of the downsampled input image, which we treat as pseudo-clean image patches, and an adaptive noise estimator to learn the specific noise distribution of the input image. By carrying out the pre-training on the single input image, rather than on a separate dataset, we avoid the well-known noise distribution gap between images in the training dataset and the single input image used at test time. We evaluate our SSL method for single image denoising via extensive experiments on both synthetic and real-world noisy image datasets. We demonstrate SotA results compared to existing unsupervised denoising methods, by transferring our pre-training to IDR [1], thus showing that SSL pre-training is a promising framework also in image denoising. Website: https://hamadichihaoui.github.io/SSL-Denoising. Hamadi Chihaoui, Paolo Favaro |
ICIP | 2 |
| 2024 | Blind Image Restoration via Fast Diffusion InversionabstractImage Restoration (IR) methods based on a pre-trained diffusion model have demonstrated state-of-the-art performance. However, they have two fundamental limitations: 1) they often assume that the degradation operator is completely known and 2) they alter the diffusion sampling process, which may result in restored images that do not lie onto the data manifold. To address these issues, we propose Blind Image Restoration via fast Diffusion inversion (BIRD) a blind IR method that jointly optimizes for the degradation model parameters and the restored image. To ensure that the restored images lie onto the data manifold, we propose a novel sampling technique on a pre-trained diffusion model. A key idea in our method is not to modify the reverse sampling, i.e., not to alter all the intermediate latents, once an initial noise is sampled. This is ultimately equivalent to casting the IR task as an optimization problem in the space of the input noise. Moreover, to mitigate the computational cost associated with inverting a fully unrolled diffusion model, we leverage the inherent capability of these models to skip ahead in the forward diffusion process using large time steps. We experimentally validate BIRD on several image restoration tasks and show that it achieves state of the art performance. Hamadi Chihaoui, Abdelhak Lemkhenter, Paolo Favaro |
NeurIPS | 3 |
| 2024 | Denoising and Selecting Pseudo-Heatmaps for Semi-Supervised Human Pose EstimationabstractWe propose a new semi-supervised learning design for human pose estimation that revisits the popular dual-student framework and enhances it two ways. First, we introduce a denoising scheme to generate reliable pseudo-heatmaps as targets for learning from unlabeled data. This uses multi-view augmentations and a threshold-and-refine procedure to produce a pool of pseudo-heatmaps. Second, we select the learning targets from these pseudo-heatmaps guided by the estimated cross-student uncertainty. We evaluate our proposed method on multiple evaluation setups on the COCO benchmark. Our results show that our model outperforms previous state-of-the-art semi-supervised pose estimators, especially in extreme low-data regime. For example with only 0.5K labeled images our method is capable of surpassing the best competitor by 7.22 mAP (+25% absolute improvement). We also demonstrate that our model can learn effectively from unlabeled data in the wild to further boost its generalization and performance. Zhuoran Yu, Manchen Wang, Yanbei Chen, Paolo Favaro, Davide Modolo |
WACV | 4 |
| 2023 | Representation Learning by Detecting Incorrect Location EmbeddingsabstractIn this paper, we introduce a novel self-supervised learning (SSL) loss for image representation learning. There is a growing belief that generalization in deep neural networks is linked to their ability to discriminate object shapes. Since object shape is related to the location of its parts, we propose to detect those that have been artificially misplaced. We represent object parts with image tokens and train a ViT to detect which token has been combined with an incorrect positional embedding. We then introduce sparsity in the inputs to make the model more robust to occlusions and to speed up the training. We call our method DILEMMA, which stands for Detection of Incorrect Location EMbeddings with MAsked inputs. We apply DILEMMA to MoCoV3, DINO and SimCLR and show an improvement in their performance of respectively 4.41%, 3.97%, and 0.5% under the same training time and with a linear probing transfer on ImageNet-1K. We also show full fine-tuning improvements of MAE combined with our method on ImageNet-100. We evaluate our method via fine-tuning on common SSL benchmarks. Moreover, we show that when downstream tasks are strongly reliant on shape (such as in the YOGA-82 pose dataset), our pre-trained features yield a significant gain over prior work. Sepehr Sameni, Simon Jenni, Paolo Favaro |
AAAI | 3 |
| 2023 | ScaleDet: A Scalable Multi-Dataset Object DetectorabstractMulti-dataset training provides a viable solution for exploiting heterogeneous large-scale datasets without extra annotation cost. In this work, we propose a scalable multi-dataset detector (ScaleDet) that can scale up its generalization across datasets when increasing the number of training datasets. Unlike existing multi-dataset learners that mostly rely on manual relabelling efforts or sophisticated optimizations to unify labels across datasets, we introduce a simple yet scalable formulation to derive a unified semantic label space for multi-dataset training. ScaleDet is trained by visual-textual alignment to learn the label assignment with label semantic similarities across datasets. Once trained, ScaleDet can generalize well on any given upstream and downstream datasets with seen and unseen classes. We conduct extensive experiments using LVIS, COCO, Objects365, OpenImages as upstream datasets, and 13 datasets from Object Detection in the Wild (ODinW) as downstream datasets. Our results show that ScaleDet achieves compelling strong model performance with an mAP of 50.7 on LVIS, 58.8 on COCO, 46.8 on Objects365, 76.2 on OpenImages, and 71.8 on ODinW, surpassing state-of-the-art detectors with the same backbone. Yanbei Chen, Manchen Wang, Abhay Mittal, Zhenlin Xu, Paolo Favaro, Joseph Tighe, Davide Modolo |
CVPR | 5 |
| 2023 | A Meta-Learning Approach to Predicting Performance and Data RequirementsabstractWe propose an approach to estimate the number of samples required for a model to reach a target performance. We find that the power law, the de facto principle to estimate model performance, leads to a large error when using a small dataset (e.g., 5 samples per class) for extrapolation. This is because the log-performance error against the log-dataset size follows a nonlinear progression in the few-shot regime followed by a linear progression in the high-shot regime. We introduce a novel piecewise power law (PPL) that handles the two data regimes differently. To estimate the parameters of the PPL, we introduce a random forest regressor trained via meta learning that generalizes across classification/detection tasks, ResNet/ViT based architectures, and random/pre-trained initializations. The PPL improves the performance estimation on average by 37% across 16 classification and 33% across 10 detection datasets, compared to the power law. We further extend the PPL to provide a confidence bound and use it to limit the prediction horizon that reduces over-estimation of data by 76% on classification and 91% on detection datasets. Achin Jain, Gurumurthy Swaminathan, Paolo Favaro, Hao Yang 0043, Avinash Ravichandran, Hrayr Harutyunyan, Alessandro Achille, Onkar Dabeer, Bernt Schiele, Ashwin Swaminathan, Stefano Soatto |
CVPR | 3 |
| 2023 | Efficient Video Prediction via Sparsely Conditioned Flow MatchingabstractWe introduce a novel generative model for video prediction based on latent flow matching, an efficient alternative to diffusion-based models. In contrast to prior work, we keep the high costs of modeling the past during training and inference at bay by conditioning only on a small random set of past frames at each integration step of the image generation process. Moreover, to enable the generation of high-resolution videos and to speed up the training, we work in the latent space of a pretrained VQGAN. Finally, we propose to approximate the initial condition of the flow ODE with the previous noisy frame. This allows to reduce the number of integration steps and hence, speed up the sampling at inference time. We call our model Random frame conditioned flow Integration for VidEo pRediction, or, in short, RIVER. We show that RIVER achieves superior or on par performance compared to prior work on common video prediction benchmarks, while requiring an order of magnitude fewer computational resources. Project website: https://araachie.github.io/river. Aram Davtyan, Sepehr Sameni, Paolo Favaro |
ICCV | 3 |
| 2023 | Spatio-Temporal Crop Aggregation for Video Representation LearningabstractWe propose Spatio-temporal Crop Aggregation for video representation LEarning (SCALE), a novel method that enjoys high scalability at both training and inference time. Our model builds long-range video features by learning from sets of video clip-level features extracted with a pre-trained backbone. To train the model, we propose a self-supervised objective consisting of masked clip feature predictions. We apply sparsity to both the input, by extracting a random set of video clips, and to the loss function, by only reconstructing the sparse inputs. Moreover, we use dimensionality reduction by working in the latent space of a pre-trained backbone applied to single video clips. These techniques make our method not only extremely efficient to train but also highly effective in transfer learning. We demonstrate that our video representation yields state-of-the-art performance with linear, nonlinear, and k-NN probing on common action classification and video understanding datasets. Sepehr Sameni, Simon Jenni, Paolo Favaro |
ICCV | 3 |
| 2022 | KOALA: A Kalman Optimization Algorithm with Loss AdaptivityabstractOptimization is often cast as a deterministic problem, where the solution is found through some iterative procedure such as gradient descent. However, when training neural networks the loss function changes over (iteration) time due to the randomized selection of a subset of the samples. This randomization turns the optimization problem into a stochastic one. We propose to consider the loss as a noisy observation with respect to some reference optimum. This interpretation of the loss allows us to adopt Kalman filtering as an optimizer, as its recursive formulation is designed to estimate unknown parameters from noisy measurements. Moreover, we show that the Kalman Filter dynamical model for the evolution of the unknown parameters can be used to capture the gradient dynamics of advanced methods such as Momentum and Adam. We call this stochastic optimization method KOALA, which is short for Kalman Optimization Algorithm with Loss Adaptivity. KOALA is an easy to implement, scalable, and efficient method to train neural networks. We provide convergence analysis and show experimentally that it yields parameter estimates that are on par with or better than existing state of the art optimization algorithms across several neural network architectures and machine learning tasks, such as computer vision and language modeling. The project page with the code and the supplementary materials is available at https://araachie.github.io/koala/. Aram Davtyan, Sepehr Sameni, Llukman Cerkezi, Givi Meishvili, Adam Bielski, Paolo Favaro |
AAAI | 6 |
| 2022 | Controllable Video Generation Through Global and Local Motion Dynamics
Aram Davtyan, Paolo Favaro |
ECCV (17) | 2 |
| 2022 | MOVE: Unsupervised Movable Object Segmentation and DetectionabstractWe introduce MOVE, a novel method to segment objects without any form of supervision. MOVE exploits the fact that foreground objects can be shifted locally relative to their initial position and result in realistic (undistorted) new images. This property allows us to train a segmentation model on a dataset of images without annotation and to achieve state of the art (SotA) performance on several evaluation datasets for unsupervised salient object detection and segmentation. In unsupervised single object discovery, MOVE gives an average CorLoc improvement of 7.2% over the SotA, and in unsupervised class-agnostic object detection it gives a relative AP improvement of 53% on average. Our approach is built on top of self-supervised features (e.g. from DINO or MAE), an inpainting network (based on the Masked AutoEncoder) and adversarial training. Adam Bielski, Paolo Favaro |
NeurIPS | 2 |
| 2022 | Semi-supervised Vision Transformers at ScaleabstractWe study semi-supervised learning (SSL) for vision transformers (ViT), an under-explored topic despite the wide adoption of the ViT architectures to different tasks. To tackle this problem, we use a SSL pipeline, consisting of first un/self-supervised pre-training, followed by supervised fine-tuning, and finally semi-supervised fine-tuning. At the semi-supervised fine-tuning stage, we adopt an exponential moving average (EMA)-Teacher framework instead of the popular FixMatch, since the former is more stable and delivers higher accuracy for semi-supervised vision transformers. In addition, we propose a probabilistic pseudo mixup mechanism to interpolate unlabeled samples and their pseudo labels for improved regularization, which is important for training ViTs with weak inductive bias. Our proposed method, dubbed Semi-ViT, achieves comparable or better performance than the CNN counterparts in the semi-supervised classification setting. Semi-ViT also enjoys the scalability benefits of ViTs that can be readily scaled up to large-size models with increasing accuracy. For example, Semi-ViT-Huge achieves an impressive 80\% top-1 accuracy on ImageNet using only 1\% labels, which is comparable with Inception-v4 using 100\% ImageNet labels. The code is available at https://github.com/amazon-science/semi-vit. Zhaowei Cai, Avinash Ravichandran, Paolo Favaro, Manchen Wang, Davide Modolo, Rahul Bhotika, Zhuowen Tu, Stefano Soatto |
NeurIPS | 3 |
| 2022 | Learn to Zoom in Single Image Super-ResolutionabstractIn this letter, we propose a novel solution to the problem of single image super-resolution at multiple scaling factors, including the extreme case of 8x, with a single network architecture. In applications where only a detail needs to be super-resolved, traditional solutions must choose to use as input either the low-resolution detail, thus losing the information about the context, or the whole low-resolution image and then crop the desired output detail, which is quite wasteful in terms of computations and storage. To address both of these issues we propose ZoomGAN, a model that takes as input the whole low-resolution image, which we call context, and a binary mask that specifies with a box which image detail in the low-resolution image to magnify. The output of ZoomGAN has the same size as the inputs so that the scaling factor is implicitly defined by the arbitrary size of the mask box. To encourage a realistic and high-quality output, we combine adversarial training with a perceptual loss. We use two discriminators: one promotes the similarity between the distributions of real and generated details and the other promotes the similarity between the distributions of real and generated (detail, context) pairs. We evaluate ZoomGAN with several experiments on several datasets and show that it achieves state of the art performance on zoomed in details in terms of the LPIPS and PI perceptual metrics, while being on par in terms of the PSNR distortion metric. Zili Zhang 0002, Paolo Favaro, Jianxiang Li |
IEEE Signal Process. Lett. | 2 |
| 2021 | Learning to Deblur and Rotate Motion-Blurred Faces
Givi Meishvili, Attila Szabó, Simon Jenni, Paolo Favaro |
BMVC | 4 |
| 2021 | Real-Time Light Field 3D Microscopy via Sparsity-Driven Learned DeconvolutionabstractLight Field Microscopy (LFM) is a scan-less 3D imaging technique capable of capturing fast biological processes, such as neural activity in zebrafish. However, current methods to recover a 3D volume from the raw data require long reconstruction times hampering the usability of the microscope in a closed-loop system. Moreover, because the main focus of zebrafish brain imaging is to isolate and study neural activity, the ideal volumetric reconstruction should be sparse to reveal the dominant signals. Unfortunately, current sparse decomposition methods are computationally intensive and thus introduce substantial delays. This motivates us to introduce a 3D reconstruction method that recovers the spatio-temporally sparse components of an image sequence in real-time. In this work we propose a combination of a neural network (SLNet) that recovers the sparse components of a light field image sequence and a neural network (XLFMNet) for 3D reconstruction. In particular, XLFMNet is able to achieve high data fidelity and to preserve important signals, such as neural potentials, even on previously unobserved samples. We demonstrate successful sparse 3D volumetric reconstructions of the neural activity of live zebrafish, with an imaging span covering 800×800×250Mm3at an imaging rate of 24 - 88Hz, which provides a 1500 fold speed increase against prior work and enables real-time reconstructions without sacrificing imaging resolution. Josué Page Vizcaíno, Zeguan Wang, Panagiotis Symvoulidis, Paolo Favaro, Burcu Guner-Ataman, Edward S. Boyden, Tobias Lasser |
ICCP | 4 |
| 2021 | ISD: Self-Supervised Learning by Iterative Similarity DistillationabstractRecently, contrastive learning has achieved great results in self-supervised learning, where the main idea is to pull two augmentations of an image (positive pairs) closer compared to other random images (negative pairs). We argue that not all negative images are equally negative. Hence, we introduce a self-supervised learning algorithm where we use a soft similarity for the negative images rather than a binary distinction between positive and negative pairs. We iteratively distill a slowly evolving teacher model to the student model by capturing the similarity of a query image to some random images and transferring that knowledge to the student. Specifically, our method should handle unbalanced and unlabeled data better than existing contrastive learning methods, because the randomly chosen negative set might include many samples that are semantically similar to the query image. In this case, our method labels them as highly similar while standard contrastive methods label them as negatives. Our method achieves comparable results to the state-of-the-art models. Our code is available here: https://github.com/UMBCvision/ISD. Ajinkya Tejankar, Soroush Abbasi Koohpayegani, Vipin Pillai, Paolo Favaro, Hamed Pirsiavash |
ICCV | 4 |
| 2021 | A Unified Generative Adversarial Network Training via Self-Labeling and Self-AttentionabstractWe propose a novel GAN training scheme that can handle any level of labeling in a unified manner. Our scheme introduces a form of artificial labeling that can incorporate manually defined labels, when available, and induce an alignment between them. To define the artificial labels, we exploit the assumption that neural network generators can be trained more easily to map nearby latent vectors to data with semantic similarities, than across separate categories. We use generated data samples and their corresponding artificial conditioning labels to train a classifier. The classifier is then used to self-label real data. To boost the accuracy of the self-labeling, we also use the exponential moving average of the classifier. However, because the classifier might still make mistakes, especially at the beginning of the training, we also refine the labels through self-attention, by using the labeling of real data samples only when the classifier outputs a high classification probability score. We evaluate our approach on CIFAR-10, STL-10 and SVHN, and show that both self-labeling and self-attention consistently improve the quality of generated data. More surprisingly, we find that the proposed scheme can even outperform class-conditional GANs. Tomoki Watanabe, Paolo Favaro |
ICML | 2 |
| 2021 | Editorial for CVIU_DL for image restoration
Jinshan Pan, Deqing Sun, Jian Yang 0003, Wangmeng Zuo, Paolo Favaro, Yasuyuki Matsushita, Ming-Hsuan Yang 0001 |
Comput. Vis. Image Underst. | 5 |
| 2020 | Self-supervised Multi-view Synchronization Learning for 3D Pose Estimation
Simon Jenni, Paolo Favaro |
ACCV (5) | 2 |
| 2020 | Steering Self-Supervised Feature Learning Beyond Local Pixel StatisticsabstractWe introduce a novel principle for self-supervised feature learning based on the discrimination of specific transformations of an image. We argue that the generalization capability of learned features depends on what image neighborhood size is sufficient to discriminate different image transformations: The larger the required neighborhood size and the more global the image statistics that the feature can describe. An accurate description of global image statistics allows to better represent the shape and configuration of objects and their context, which ultimately generalizes better to new tasks such as object classification and detection. This suggests a criterion to choose and design image transformations. Based on this criterion, we introduce a novel image transformation that we call limited context inpainting (LCI). This transformation inpaints an image patch conditioned only on a small rectangular pixel boundary (the limited context). Because of the limited boundary information, the inpainter can learn to match local pixel statistics, but is unlikely to match the global statistics of the image. We claim that the same principle can be used to justify the performance of transformations such as image rotations and warping. Indeed, we demonstrate experimentally that learning to discriminate transformations such as LCI, image warping and rotations, yields features with state of the art generalization capabilities on several datasets such as Pascal VOC, STL-10, CelebA, and ImageNet. Remarkably, our trained features achieve a performance on Places on par with features trained through supervised learning with ImageNet labels. Simon Jenni, Hailin Jin, Paolo Favaro |
CVPR | 3 |
| 2020 | Learning to Have an Ear for Face Super-ResolutionabstractWe propose a novel method to use both audio and a low-resolution image to perform extreme face super-resolution (a 16x increase of the input size). When the resolution of the input image is very low (e.g., 8x8 pixels), the loss of information is so dire that important details of the original identity have been lost and audio can aid the recovery of a plausible high-resolution image. In fact, audio carries information about facial attributes, such as gender and age. To combine the aural and visual modalities, we propose a method to first build the latent representations of a face from the lone audio track and then from the lone low-resolution image. We then train a network to fuse these two representations. We show experimentally that audio can assist in recovering attributes such as the gender, the age and the identity, and thus improve the correctness of the high-resolution image reconstruction process. Our procedure does not make use of human annotation and thus can be easily trained with existing video datasets. Moreover, we show that our model builds a factorized representation of images and audio as it allows one to mix low-resolution images and audio from different videos and to generate realistic faces with semantically meaningful combinations. Givi Meishvili, Simon Jenni, Paolo Favaro |
CVPR | 3 |
| 2020 | Video Representation Learning by Recognizing Temporal Transformations
Simon Jenni, Givi Meishvili, Paolo Favaro |
ECCV (28) | 3 |
| 2020 | Learning to Model and Calibrate Optics Via a Differentiable Wave Optics SimulatorabstractWe present a novel learning-based method to build a differentiable computational model of a real fluorescence microscope. Our model can be used to calibrate a real optical setup directly from data samples and to engineer point spread functions by specifying the desired input-output data. This approach is poised to drastically improve the design of microscopes, because the parameters of current models of optical setups cannot be easily fit to real data. Inspired by the recent progress in deep learning, our solution is to build a differentiable wave optics simulator as a composition of trainable modules, each computing light wave-front (WF) propagation due to a specific optical element. We call our differentiable modules WaveBlocks and show reconstruction results in the case of lenses, wave propagation in air, camera sensors and diffractive elements (e.g., phase-masks). Josué Page Vizcaíno, Paolo Favaro |
ICIP | 2 |
| 2020 | Learning to Take Directions One Step at a TimeabstractWe present a method to generate a video sequence given a single image. Because items in an image can be animated in arbitrarily many different ways, we introduce as control signal a sequence of motion strokes. Such control signal can also be automatically transferred from other videos, e.g., via bounding box tracking. Each motion stroke provides the direction to the moving object in the input image and we aim to train a network to generate an animation following a sequence of such directions. To address this task we design a novel recurrent architecture, which can be trained easily and effectively thanks to an explicit separation of past, future and current states. As we demonstrate in the experiments, our proposed architecture is capable of generating an arbitrary number of frames from a single image and a sequence of motion strokes. Key components of our architecture are an autoencoding constraint to ensure consistency with the past and a generative adversarial scheme to ensure that images look realistic and are temporally smooth. We demonstrate the effectiveness of our approach on the MNIST, KTH, Human3.6M, Push and Weizmann datasets. Qiyang Hu, Adrian Waelchli, Tiziano Portenier, Matthias Zwicker, Paolo Favaro |
ICPR | 5 |
| 2019 | On Stabilizing Generative Adversarial Training With NoiseabstractWe present a novel method and analysis to train generative adversarial networks (GAN) in a stable manner. As shown in recent analysis, training is often undermined by the probability distribution of the data being zero on neighborhoods of the data space. We notice that the distributions of real and generated data should match even when they undergo the same filtering. Therefore, to address the limited support problem we propose to train GANs by using different filtered versions of the real and generated data distributions. In this way, filtering does not prevent the exact matching of the data distribution, while helping training by extending the support of both distributions. As filtering we consider adding samples from an arbitrary distribution to the data, which corresponds to a convolution of the data distribution with the arbitrary one. We also propose to learn the generation of these samples so as to challenge the discriminator in the adversarial training. We show that our approach results in a stable and well-behaved training of even the original minimax GAN formulation. Moreover, our technique can be incorporated in most modern GAN formulations and leads to a consistent improvement on several common datasets. Simon Jenni, Paolo Favaro |
CVPR | 2 |
| 2019 | Learning to Extract Flawless Slow Motion From Blurry VideosabstractIn this paper, we introduce the task of generating a sharp slow-motion video given a low frame rate blurry video. We propose a data-driven approach, where the training data is captured with a high frame rate camera and blurry images are simulated through an averaging process. While it is possible to train a neural network to recover the sharp frames from their average, there is no guarantee of the temporal smoothness for the formed video, as the frames are estimated independently. To address the temporal smoothness requirement we propose a system with two networks: One, DeblurNet, to predict sharp keyframes and the second, InterpNet, to predict intermediate frames between the generated keyframes. A smooth transition is ensured by interpolating between consecutive keyframes using InterpNet. Moreover, the proposed scheme enables further increase in frame rate without retraining the network, by applying InterpNet recursively between pairs of sharp frames. We evaluate the proposed method on several datasets, including a novel dataset captured with a Sony RX V camera. We also demonstrate its performance of increasing the frame rate up to 20 times on real blurry videos. Meiguang Jin, Paolo Favaro |
CVPR | 3 |
| 2019 | Emergence of Object Segmentation in Perturbed Generative ModelsabstractWe introduce a novel framework to build a model that can learn how to segment objects from a collection of images without any human annotation. Our method builds on the observation that the location of object segments can be perturbed locally relative to a given background without affecting the realism of a scene. Our approach is to first train a generative model of a layered scene. The layered representation consists of a background image, a foreground image and the mask of the foreground. A composite image is then obtained by overlaying the masked foreground image onto the background. The generative model is trained in an adversarial fashion against a discriminator, which forces the generative model to produce realistic composite images. To force the generator to learn a representation where the foreground layer corresponds to an object, we perturb the output of the generative model by introducing a random shift of both the foreground image and mask relative to the background. Because the generator is unaware of the shift before computing its output, it must produce layered representations that are realistic for any such random perturbation. Finally, we learn to segment an image by defining an autoencoder consisting of an encoder, which we train, and the pre-trained generator as the decoder, which we freeze. The encoder maps an image to a feature vector, which is fed as input to the generator to give a composite image matching the original input image. Because the generator outputs an explicit layered representation of the scene, the encoder learns to detect and segment objects. We demonstrate this framework on real images of several object categories. Adam Bielski, Paolo Favaro |
NeurIPS | 2 |
| 2019 | Motion Deblurring of FacesabstractFace analysis lies at the heart of computer vision with remarkable progress in the past decades. Face recognition and tracking are tackled by building invariance to fundamental modes of variation such as illumination, 3D pose. A much less standing mode of variation is motion deblurring, which however presents substantial challenges in face analysis. Recent approaches either make oversimplifying assumptions, e.g. in cases of joint optimization with other tasks, or fail to preserve the highly structured shape/identity information. We introduce a two-step architecture tailored to the challenges of motion deblurring: the first step restores the low frequencies; the second restores the high frequencies, while ensuring that the outputs span the natural images manifold. Both steps are implemented with a supervised data-driven method; to train those we devise a method for creating realistic motion blur by averaging a variable number of frames. The averaged images originate from the $$2MF^2$$ dataset with $$19$$ million facial frames, which we introduce for the task. Considering deblurring as an intermediate step, we conduct a thorough experimentation on high-level face analysis tasks, i.e. landmark localization and face verification, on blurred images. The experimental evaluation demonstrates the superiority of our method. Grigorios Chrysos 0002, Paolo Favaro, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 2 |
| 2018 | Disentangling Factors of Variation by Mixing ThemabstractWe propose an approach to learn image representations that consist of disentangled factors of variation without exploiting any manual labeling or data domain knowledge. A factor of variation corresponds to an image attribute that can be discerned consistently across a set of images, such as the pose or color of objects. Our disentangled representation consists of a concatenation of feature chunks, each chunk representing a factor of variation. It supports applications such as transferring attributes from one image to another, by simply mixing and unmixing feature chunks, and classification or retrieval based on one or several attributes, by considering a user-specified subset of feature chunks. We learn our representation without any labeling or knowledge of the data domain, using an autoencoder architecture with two novel training objectives: first, we propose an invariance objective to encourage that encoding of each attribute, and decoding of each chunk, are invariant to changes in other attributes and chunks, respectively; second, we include a classification objective, which ensures that each chunk corresponds to a consistently discernible attribute in the represented image, hence avoiding degenerate feature mappings where some chunks are completely ignored. We demonstrate the effectiveness of our approach on the MNIST, Sprites, and CelebA datasets. Qiyang Hu, Attila Szabó, Tiziano Portenier, Paolo Favaro, Matthias Zwicker |
CVPR | 4 |
| 2018 | Self-Supervised Feature Learning by Learning to Spot ArtifactsabstractWe introduce a novel self-supervised learning method based on adversarial training. Our objective is to train a discriminator network to distinguish real images from images with synthetic artifacts, and then to extract features from its intermediate layers that can be transferred to other data domains and tasks. To generate images with artifacts, we pre-train a high-capacity autoencoder and then we use a damage and repair strategy: First, we freeze the autoencoder and damage the output of the encoder by randomly dropping its entries. Second, we augment the decoder with a repair network, and train it in an adversarial manner against the discriminator. The repair network helps generate more realistic images by inpainting the dropped feature entries. To make the discriminator focus on the artifacts, we also make it predict what entries in the feature were dropped. We demonstrate experimentally that features learned by creating and spotting artifacts achieve state of the art performance in several benchmarks. Simon Jenni, Paolo Favaro |
CVPR | 2 |
| 2018 | Learning to Extract a Video Sequence From a Single Motion-Blurred ImageabstractWe present a method to extract a video sequence from a single motion-blurred image. Motion-blurred images are the result of an averaging process, where instant frames are accumulated over time during the exposure of the sensor. Unfortunately, reversing this process is nontrivial. Firstly, averaging destroys the temporal ordering of the frames. Secondly, the recovery of a single frame is a blind deconvolution task, which is highly ill-posed. We present a deep learning scheme that gradually reconstructs a temporal ordering by sequentially extracting pairs of frames. Our main contribution is to introduce loss functions invariant to the temporal order. This lets a neural network choose during training what frame to output among the possible combinations. We also address the ill-posedness of deblurring by designing a network with a large receptive field and implemented via resampling to achieve a higher computational efficiency. Our proposed method can successfully retrieve sharp image sequences from a single motion blurred image and can generalize well on synthetic and real datasets captured with different cameras. Meiguang Jin, Givi Meishvili, Paolo Favaro |
CVPR | 3 |
| 2018 | Boosting Self-Supervised Learning via Knowledge TransferabstractIn self-supervised learning, one trains a model to solve a so-called pretext task on a dataset without the need for human annotation. The main objective, however, is to transfer this model to a target domain and task. Currently, the most effective transfer strategy is fine-tuning, which restricts one to use the same model or parts thereof for both pretext and target tasks. In this paper, we present a novel framework for self-supervised learning that overcomes limitations in designing and comparing different tasks, models, and data domains. In particular, our framework decouples the structure of the self-supervised model from the final task-specific fine-tuned model. This allows us to: 1) quantitatively assess previously incompatible models including handcrafted features; 2) show that deeper neural network models can learn better representations from the same pretext task; 3) transfer knowledge learned with a deep model to a shallower one and thus boost its learning. We use this framework to design a novel self-supervised task, which achieves state-of-the-art performance on the common benchmarks in PASCAL VOC 2007, ILSVRC12 and Places by a significant margin. Our learned features shrink the mAP gap between models trained via self-supervised learning and supervised learning from 5.9% to 2.6% in object detection on PASCAL VOC 2007. Mehdi Noroozi, Ananth Vinjimoor, Paolo Favaro, Hamed Pirsiavash |
CVPR | 3 |
| 2018 | Deep Bilevel Learning
Simon Jenni, Paolo Favaro |
ECCV (10) | 2 |
| 2018 | Normalized Blind Deconvolution
Meiguang Jin, Stefan Roth 0001, Paolo Favaro |
ECCV (7) | 3 |
| 2018 | Understanding Degeneracies and Ambiguities in Attribute Transfer
Attila Szabó, Qiyang Hu, Tiziano Portenier, Matthias Zwicker, Paolo Favaro |
ECCV (5) | 5 |
| 2018 | Learning to see through reflectionsabstractPictures of objects behind a glass are difficult to interpret and understand due to the superposition of two real images: a reflection layer and a background layer. Separation of these two layers is challenging due to the ambiguities in assigning texture patterns and the average color in the input image to one of the two layers. In this paper, we propose a novel method to reconstruct these layers given a single input image by explicitly handling the ambiguities of the reconstruction. Our approach combines the ability of neural networks to build image priors on large image regions with an image model that accounts for the brightness ambiguity and saturation. We find that our solution generalizes to real images even in the presence of strong reflections. Extensive quantitative and qualitative experimental evaluations on both real and synthetic data show the benefits of our approach over prior work. Moreover, our proposed neural network is computationally and memory efficient. Meiguang Jin, Sabine Süsstrunk, Paolo Favaro |
ICCP | 3 |
| 2018 | Plenoptic Image Motion DeblurringabstractWe propose a method to remove motion blur in a single light field captured with a moving plenoptic camera. Since motion is unknown, we resort to a blind deconvolution formulation, where one aims to identify both the blur point spread function and the latent sharp image. Even in the absence of motion, light field images captured by a plenoptic camera are affected by a non-trivial combination of both aliasing and defocus, which depends on the 3D geometry of the scene. Therefore, motion deblurring algorithms designed for standard cameras are not directly applicable. Moreover, many state of the art blind deconvolution algorithms are based on iterative schemes, where blurry images are synthesized through the imaging model. However, current imaging models for plenoptic images are impractical due to their high dimensionality. We observe that plenoptic cameras introduce periodic patterns that can be exploited to obtain highly parallelizable numerical schemes to synthesize images. These schemes allow extremely efficient GPU implementations that enable the use of iterative methods. We can then cast blind deconvolution of a blurry light field image as a regularized energy minimization to recover a sharp high-resolution scene texture and the camera motion. Furthermore, the proposed formulation can handle non-uniform motion blur due to camera shake as demonstrated on both synthetic and real light field data. Paramanand Chandramouli, Meiguang Jin, Daniele Perrone, Paolo Favaro |
IEEE Trans. Image Process. | 4 |
| 2018 | Faceshop: deep sketch-based face image editingabstractWe present a novel system for sketch-based face image editing, enabling users to edit images intuitively by sketching a few strokes on a region of interest. Our interface features tools to express a desired image manipulation by providing both geometry and color constraints as user-drawn strokes. As an alternative to the direct user input, our proposed system naturally supports a copy-paste mode, which allows users to edit a given image region by using parts of another exemplar image without the need of hand-drawn sketching at all. The proposed interface runs in real-time and facilitates an interactive and iterative workflow to quickly express the intended edits. Our system is based on a novel sketch domain and a convolutional neural network trained end-to-end to automatically learn to render image regions corresponding to the input strokes. To achieve high quality and semantically consistent results we train our neural network on two simultaneous tasks, namely image completion and image translation. To the best of our knowledge, we are the first to combine these two tasks in a unified framework for interactive image editing. Our results show that the proposed sketch domain, network architecture, and training procedure generalize well to real user input and enable high quality synthesis results without additional post-processing. Tiziano Portenier, Qiyang Hu, Attila Szabó, Siavash Arjomand Bigdeli, Paolo Favaro, Matthias Zwicker |
ACM Trans. Graph. | 5 |
| 2017 | Noise-Blind Image DeblurringabstractWe present a novel approach to noise-blind deblurring, the problem of deblurring an image with known blur, but unknown noise level. We introduce an efficient and robust solution based on a Bayesian framework using a smooth generalization of the 0-1 loss. A novel bound allows the calculation of very high-dimensional integrals in closed form. It avoids the degeneracy of Maximum a-Posteriori (MAP) estimates and leads to an effective noise-adaptive scheme. Moreover, we drastically accelerate our algorithm by using Majorization Minimization (MM) without introducing any approximation or boundary artifacts. We further speed up convergence by turning our algorithm into a neural network termed GradNet, which is highly parallelizable and can be efficiently trained. We demonstrate that our noise-blind formulation can be integrated with different priors and significantly improves existing deblurring algorithms in the noise-blind and in the known-noise case. Furthermore, GradNet leads to state-of-the-art performance across different noise levels, while retaining high computational efficiency. Meiguang Jin, Stefan Roth 0001, Paolo Favaro |
CVPR | 3 |
| 2017 | Representation Learning by Learning to CountabstractWe introduce a novel method for representation learning that uses an artificial supervision signal based on counting visual primitives. This supervision signal is obtained from an equivariance relation, which does not require any manual annotation. We relate transformations of images to transformations of the representations. More specifically, we look for the representation that satisfies such relation rather than the transformations that match a given representation. In this paper, we use two image transformations in the context of counting: scaling and tiling. The first transformation exploits the fact that the number of visual primitives should be invariant to scale. The second transformation allows us to equate the total number of visual primitives in each tile to that in the whole image. These two transformations are combined in one constraint and used to train a neural network with a contrastive loss. The proposed task produces representations that perform on par or exceed the state of the art in transfer learning benchmarks. Mehdi Noroozi, Hamed Pirsiavash, Paolo Favaro |
ICCV | 3 |
| 2017 | Deep Mean-Shift Priors for Image RestorationabstractIn this paper we introduce a natural image prior that directly represents a Gaussian-smoothed version of the natural image distribution. We include our prior in a formulation of image restoration as a Bayes estimator that also allows us to solve noise-blind image restoration problems. We show that the gradient of our prior corresponds to the mean-shift vector on the natural image distribution. In addition, we learn the mean-shift vector field using denoising autoencoders, and use it in a gradient descent approach to perform Bayes risk minimization. We demonstrate competitive results for noise-blind deblurring, super-resolution, and demosaicing. Siavash Arjomand Bigdeli, Matthias Zwicker, Paolo Favaro, Meiguang Jin |
NIPS | 3 |
| 2016 | ConvNet-Based Depth Estimation, Reflection Separation and Deblurring of Plenoptic Images
Paramanand Chandramouli, Mehdi Noroozi, Paolo Favaro |
ACCV (3) | 3 |
| 2016 | Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
Mehdi Noroozi, Paolo Favaro |
ECCV (6) | 2 |
| 2016 | A Logarithmic Image Prior for Blind Deconvolution
Daniele Perrone, Paolo Favaro |
Int. J. Comput. Vis. | 2 |
| 2016 | A Clearer Picture of Total Variation Blind DeconvolutionabstractBlind deconvolution is the problem of recovering a sharp image and a blur kernel from a noisy blurry image. Recently, there has been a significant effort on understanding the basic mechanisms to solve blind deconvolution. While this effort resulted in the deployment of effective algorithms, the theoretical findings generated contrasting views on why these approaches worked. On the one hand, one could observe experimentally that alternating energy minimization algorithms converge to the desired solution. On the other hand, it has been shown that such alternating minimization algorithms should fail to converge and one should instead use a so-called Variational Bayes approach. To clarify this conundrum, recent work showed that a good image and blur prior is instead what makes a blind deconvolution algorithm work. Unfortunately, this analysis did not apply to algorithms based on total variation regularization. In this manuscript, we provide both analysis and experiments to get a clearer picture of blind deconvolution. Our analysis reveals the very reason why an algorithm based on total variation works. We also introduce an implementation of this algorithm and show that, in spite of its extreme simplicity, it is very robust and achieves a performance comparable to the top performing algorithms. Daniele Perrone, Paolo Favaro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Synchronization of Independently Moving Cameras via Motion RecoveryabstractThis work addresses the video synchronization problem, which consists in finding the temporal alignment between sequences of images acquired with unsynchronized cameras. This problem has been addressed before, under the assumptions that the cameras are static or jointly moving and that there are correspondences between the features visible in the different sequences. There are some methods in the literature that managed to drop one of these assumptions, but none of them was successful in getting rid of both assumptions simultaneously. In this work, we introduce a new strategy that synchronizes cameras that are allowed to move freely even when there are no correspondences between the features that are visible in the different sequences. Our approach consists in tracking features on two rigid objects that move independently on the scene and use the relative motion between them as a clue for the synchronization. New synchronization algorithms for static or jointly moving cameras that see (possibly) different parts of a common rigidly moving object are also presented. Even though the emphasis of this work is essentially on the theoretical contribution of the proposed methods, rather than on an exhaustive experimental validation, several proof of concept experiments conducted with both real and synthetic data are presented. In the case of static or jointly moving cameras, comparisons with a state-of-the-art approach are also provided. Tiago Gaspar, Paulo Oliveira 0001, Paolo Favaro |
SIAM J. Imaging Sci. | 3 |
| 2014 | Uncalibrated Near-Light Photometric Stereo
Thoma Papadhimitri, Paolo Favaro |
BMVC | 2 |
| 2014 | Total Variation Blind Deconvolution: The Devil Is in the DetailsabstractIn this paper we study the problem of blind deconvolution. Our analysis is based on the algorithm of Chan and Wong [2] which popularized the use of sparse gradient priors via total variation. We use this algorithm because many methods in the literature are essentially adaptations of this framework. Such algorithm is an iterative alternating energy minimization where at each step either the sharp image or the blur function are reconstructed. Recent work of Levin et al. [14] showed that any algorithm that tries to minimize that same energy would fail, as the desired solution has a higher energy than the no-blur solution, where the sharp image is the blurry input and the blur is a Dirac delta. However, experimentally one can observe that Chan and Wong's algorithm converges to the desired solution even when initialized with the no-blur one. We provide both analysis and experiments to resolve this paradoxical conundrum. We find that both claims are right. The key to understanding how this is possible lies in the details of Chan and Wong's implementation and in how seemingly harmless choices result in dramatic effects. Our analysis reveals that the delayed scaling (normalization) in the iterative step of the blur kernel is fundamental to the convergence of the algorithm. This then results in a procedure that eludes the no-blur solution, despite it being a global minimum of the original energy. We introduce an adaptation of this algorithm and show that, in spite of its extreme simplicity, it is very robust and achieves a performance comparable to the state of the art. Daniele Perrone, Paolo Favaro |
CVPR | 2 |
| 2014 | Synchronization of Two Independently Moving Cameras without Feature Correspondences
Tiago Gaspar, Paulo Oliveira 0001, Paolo Favaro |
ECCV (1) | 3 |
| 2014 | Anamorphic pixels for multi-channel superresolutionabstractSuperresolution from plenoptic cameras or camera arrays is usually treated similarly to superresolution from video streams. However, the transformation between the low-resolution views can be determined precisely from camera geometry and parallax. Furthermore, as each low-resolution image originates from a unique physical camera, its sampling properties can also be unique. We exploit this option with a custom design of either the optics or the sensor pixels. This design makes sure that the sampling matrix of the complete system is always well-formed, enabling robust and high-resolution image reconstruction. We show that simply changing the pixel aspect ratio from square to anamorphic is sufficient to achieve that goal, as long as each camera has a unique aspect ratio. We support this claim with theoretical analysis and image reconstruction of real images. We derive the optimal aspect ratios for sets of 2 or 4 cameras. Finally, we verify our solution with a camera system using an anamorphic lens. Alexander Oberdörster, Paolo Favaro, Hendrik P. A. Lensch |
ICCP | 2 |
| 2014 | Which side of the focal plane are you on?abstractDefocus blur is an indicator for the depth structure of a scene. However, given a single input image from a conventional camera one cannot distinguish between blurred objects lying in front or behind the focal plane, as they may be subject to exactly the same amount of blur. In this paper we address this limitation by exploiting coded apertures. Previous work in this area focuses on setups where the scene is placed either entirely in front or entirely behind the focal plane. We demonstrate that asymmetric apertures result in unique blurs for all distances from the camera. To exploit asymmetric apertures we propose an algorithm that can unambiguously estimate scene depth and texture from a single input image. One of the main advantages of our method is that, within the same depth range, we can work with less blurred data than in other methods. The technique is tested on both synthetic and real images. Anita Sellent, Paolo Favaro |
ICCP | 2 |
| 2014 | A Closed-Form, Consistent and Robust Solution to Uncalibrated Photometric Stereo Via Local Diffuse Reflectance Maxima
Thoma Papadhimitri, Paolo Favaro |
Int. J. Comput. Vis. | 2 |
| 2014 | Optimized aperture shapes for depth estimation
Anita Sellent, Paolo Favaro |
Pattern Recognit. Lett. | 2 |
| 2014 | Low rank subspace clustering (LRSC)
René Vidal, Paolo Favaro |
Pattern Recognit. Lett. | 2 |
| 2013 | A New Perspective on Uncalibrated Photometric StereoabstractWe investigate the problem of reconstructing normals, albedo and lights of Lambertian surfaces in uncalibrated photometric stereo under the perspective projection model. Our analysis is based on establishing the integrability constraint. In the orthographic projection case, it is well-known that when such constraint is imposed, a solution can be identified only up to 3 parameters, the so-called generalized bas-relief (GBR) ambiguity. We show that in the perspective projection case the solution is unique. We also propose a closed-form solution which is simple, efficient and robust. We test our algorithm on synthetic data and publicly available real data. Our quantitative tests show that our method outperforms all prior work of uncalibrated photometric stereo under orthographic projection. Thoma Papadhimitri, Paolo Favaro |
CVPR | 2 |
| 2013 | Sparse representation based action and gesture recognitionabstractIn this paper we present a solution to the problem of action and gesture recognition using sparse representations. The dictionary is modelled as a simple concatenation of features computed for each action or gesture class from the training data, and test data is classified by finding sparse representation of the test video features over this dictionary. Our method does not impose any explicit training procedure on the dictionary. We experiment our model with two kinds of features, by projecting (i) Gait Energy Images (GEIs) and (ii) Motion-descriptors, to a lower dimension using Random projection. Experiments have shown 100% recognition rate on standard datasets and are compared to the results obtained with widely used SVM classifier. Sushma Bomma, Paolo Favaro, Neil Robertson 0002 |
ICIP | 2 |
| 2012 | Image Priors for Image Deblurring with Uncertain BlurabstractWe consider the problem of non-blind deconvolution of images corrupted by a blur that is not accurately known. We propose a method that exploits dictionary-based image priors and non Gaussian noise models to improve deblurring accuracy in the presence of an inexact blur. The proposed image priors express each image patch as a linear combination of atoms from a dictionary learned from patches extracted from the same image or from an image database. When applied to blurred images, this model imposes that patches that are similar in the blurred image retain the same similarity when deblurred. We perform image deblurring by imposing this prior model in an energy minimization scheme that also deals with outliers. Experimental results on publicly available databases show that our approach is able to remove artifacts such as oscillations, which are often introduced during the deblurring process when the correct blur is not known. Daniele Perrone, Avinash Ravichandran, René Vidal, Paolo Favaro |
BMVC | 4 |
| 2012 | A closed-form solution to uncalibrated photometric stereo via diffuse maximaabstractIn this paper we propose a novel solution to uncalibrated photometric stereo. Our approach is to eliminate the so-called generalized bas relief (GBR) ambiguity by exploiting points where the Lambertian reflection is maximal. We demonstrate several noteworthy properties of these maxima: 1) Closed-form solution: A single diffuse maximum constrains the GBR ambiguity to a semi-circle in 3D space; 2) Efficiency: As few as two diffuse maxima in different images identify a unique solution; 3) GBR-invariance: The estimation error of the GBR parameters is completely independent of the true parameters. Furthermore, our algorithm is remarkably robust: It can obtain an accurate estimate of the GBR parameters even with extremely high levels of outliers in the detected maxima (up to 80% of the observations). The method is validated on real data and achieves state-of-the-art results. Paolo Favaro, Thoma Papadhimitri |
CVPR | 1 |
| 2012 | The Light Field Camera: Extended Depth of Field, Aliasing, and SuperresolutionabstractPortable light field (LF) cameras have demonstrated capabilities beyond conventional cameras. In a single snapshot, they enable digital image refocusing and 3D reconstruction. We show that they obtain a larger depth of field but maintain the ability to reconstruct detail at high resolution. In fact, all depths are approximately focused, except for a thin slab where blur size is bounded, i.e., their depth of field is essentially inverted compared to regular cameras. Crucial to their success is the way they sample the LF, trading off spatial versus angular resolution, and how aliasing affects the LF. We show that applying traditional multiview stereo methods to the extracted low-resolution views can result in reconstruction errors due to aliasing. We address these challenges using an explicit image formation model, and incorporate Lambertian and texture preserving priors to reconstruct both scene depth and its superresolved texture in a variational Bayesian framework, eliminating aliasing by fusing multiview information. We demonstrate the method on synthetic and real images captured with our LF camera, and show that it can outperform other computational camera systems. Tom E. Bishop, Paolo Favaro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | A closed form solution to robust subspace estimation and clusteringabstractWe consider the problem of fitting one or more subspaces to a collection of data points drawn from the subspaces and corrupted by noise/outliers. We pose this problem as a rank minimization problem, where the goal is to decompose the corrupted data matrix as the sum of a clean, self-expressive, low-rank dictionary plus a matrix of noise/outliers. Our key contribution is to show that, for noisy data, this non-convex problem can be solved very efficiently and in closed form from the SVD of the noisy data matrix. Remarkably, this is true for both one or more subspaces. An important difference with respect to existing methods is that our framework results in a polynomial thresholding of the singular values with minimal shrinkage. Indeed, a particular case of our framework in the case of a single subspace leads to classical PCA, which requires no shrinkage. In the case of multiple subspaces, our framework provides an affinity matrix that can be used to cluster the data according to the sub-spaces. In the case of data corrupted by outliers, a closed-form solution appears elusive. We thus use an augmented Lagrangian optimization framework, which requires a combination of our proposed polynomial thresholding operator with the more traditional shrinkage-thresholding operator. Paolo Favaro, René Vidal, Avinash Ravichandran |
CVPR | 1 |
| 2011 | Fragmented aperture imaging for motion and defocus deblurringabstractIn this paper we present a space-varying deblurring algorithm from a single defocused and motion-blurred image obtained with a fragmented aperture. We show that, for the same overall incoming light, a fragmented aperture leads to better motion and defocus deblurring than a (compact) conventional circular aperture. We demonstrate that not only fragmented apertures preserve more spectrum of an image of the scene than traditional circular apertures, but they also allow a better identification of blur scale. Our algorithm estimates both motion blur magnitude and direction as well as defocus blur scale at each pixel. The estimation of the blur parameters is addressed by using local projections on subspaces and L1regularization, while deblurring is posed as a variational minimization problem and solved via linearization of the Euler-Lagrange equations. The technique produces convincing results on real scenario. Manuel Martinello, Paolo Favaro |
ICIP | 2 |
| 2010 | Full-Resolution Depth Map Estimation from an Aliased Plenoptic Light Field
Tom E. Bishop, Paolo Favaro |
ACCV (2) | 2 |
| 2010 | A Unified Approach to Segmentation and Categorization of Dynamic Textures
Avinash Ravichandran, Paolo Favaro, René Vidal |
ACCV (1) | 2 |
| 2010 | Recovering thin structures via nonlocal-means regularization with application to depth from defocusabstractWe propose a novel scheme to recover depth maps containing thin structures based on nonlocal-means filtering regularization. The scheme imposes a distributed smoothness constraint by relying on the assumption that pixels with similar colors are likely to belong to the same surface, and therefore can be used jointly to obtain a robust estimate of their depth. This scheme can be used to solve shape-from-X problems and we demonstrate its use in the case of depth from defocus. We cast the problem in a variational framework and solve it by linearizing the corresponding Euler-Lagrange equations. The linearized system is then inverted by using efficient numerical methods such as successive overrelaxations or more general methods such as conjugate gradient when the system is not diagonally dominant. One of the main benefits of this formulation is that it can handle the regularization of highly fragmented surfaces, which require large neighborhood structures typically difficult to solve efficiently with graph-based methods. We compare the performance of the proposed algorithm with methods recently proposed in the literature that are analogous to neighborhood filters. Finally, experimental results are shown on synthetic and real data. Paolo Favaro |
CVPR | 1 |
| 2010 | A Bayesian approach to shape from coded apertureabstractIn this paper we present analysis and a novel algorithm to estimate depth from a single image captured by a coded aperture camera. This is a challenging problem which requires new tools and investigations, compared with multi-view reconstruction. Unlike previous approaches, which need to recover both sharp image and depth, we consider directly estimating only depth, whilst still accounting for the statistics of the sharp image. The problem is formulated in a Bayesian framework, which enables us to reduce the estimation of the original sharp image to the local space-varying statistics of the texture. This yields an algorithm that can be solved via graph cuts (without user interaction). Performance and results on both synthetic and real data are reported and compared with previous methods. Manuel Martinello, Tom E. Bishop, Paolo Favaro |
ICIP | 3 |
| 2010 | Greedy scheduling algorithm (GSA) - Design and evaluation of an efficient and flexible WiMAX OFDMA scheduling solution
Anatolij Zubow, Daniel Camps-Mur, Xavier Pérez Costa, Paolo Favaro |
Comput. Networks | 4 |
| 2008 | On the Challenges for the Maximization of Radio Resources Usage in WiMAX NetworksabstractWiMAX is one of the most promising technologies to provide broadband wireless access in the near future. In this paper we identify a key element for the performance of a WiMAX network, the DL-MAP packing algorithm, which mainly determines the usage efficiency of the available radio resources and investigate potential differences that could appear between WiMAX equipment vendors in the maximum capacity of the system due to the packing approach used. Our results show that the performance of simple DL-MAP packing algorithms might be significantly outperformed by more complex ones resulting in a clear differentiation factor among manufacturers. Xavier Pérez Costa, Paolo Favaro, Anatolij Zubow, Daniel Camps-Mur, Julio Aráuz |
CCNC | 2 |
| 2008 | Off-axis aperture camera: 3D shape reconstruction and image restorationabstractIn this paper we present a novel 3D surface and image reconstruction method based on the off-axis aperture camera. The key idea is to change the size or the 3-D location of the aperture of the camera lens so as to extract selected portions of the light field of the scene. We show that this results in an imaging device that blends defocus and stereo information, and present an image formation model that simultaneously captures both phenomena. As this model involves a non trivial deformation of the scene space, we also introduce the concept of scene space rectification and how this helps the reconstruction problem. Finally, we formulate our shape and image reconstruction problem as an energy minimization, and use a gradient flow algorithm to find the solution. Results on both real and synthetic data are shown. Qingxu Dou, Paolo Favaro |
CVPR | 2 |
| 2008 | A theory of defocus via Fourier analysisabstractIn this paper we present a novel theory to analyze defocused images of a volume density by exploiting well-known results in Fourier analysis and the singular value decomposition. This analysis is fundamental in two respects: First, it gives a deep insight into the basic mechanisms of image formation of defocused images, and second, it shows how to incorporate additional a-priori knowledge about the geometry and photometry of the scene in restoration algorithms. For instance, we show that the case of a scene made of a single surface results in a simple constraint in the Fourier domain. We derive two basic types of algorithms for volumetric reconstruction: One based on a dense set of defocused images, and one based on a sparse set of defocused images. While the first one excels in simplicity, the second one is of more practical use. Both algorithms are tested on real and synthetic data. Paolo Favaro, Alessandro Duci |
CVPR | 1 |
| 2008 | Shape from Defocus via DiffusionabstractDefocus can be modeled as a diffusion process and represented mathematically using the heat equation, where image blur corresponds to the diffusion of heat. This analogy can be extended to non-planar scenes by allowing a space-varying diffusion coefficient. The inverse problem of reconstructing 3-D structure from blurred images corresponds to an "inverse diffusion" that is notoriously ill-posed. We show how to bypass this problem by using the notion of relative blur. Given two images, within each neighborhood, the amount of diffusion necessary to transform the sharper image into the blurrier one depends on the depth of the scene. This can be used to devise a global algorithm to estimate the depth profile of the scene without recovering the deblurred image, using only forward diffusion. Paolo Favaro, Stefano Soatto, Martin Burger 0001, Stanley J. Osher |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Autocalibration and Uncalibrated Reconstruction of Shape from DefocusabstractMost algorithms for reconstructing shape from defocus assume that the images are obtained with a camera that has been previously calibrated so that the aperture, focal plane, and focal length are known. In this manuscript we characterize the set of scenes that can be reconstructed from defocused images regardless of calibration parameters. In lack of knowledge about the camera or about the scene, reconstruction is possible only up to an equivalence class that is described analytically. When weak knowledge about the scene is available, however, we show how it can be exploited in order to auto-calibrate the imaging device. This includes imaging a slanted plane or generic assumptions on the restoration of the deblurred images. Yifei Lou, Paolo Favaro, Andrea L. Bertozzi, Stefano Soatto |
CVPR | 2 |
| 2007 | Shape from Focus and Defocus: Convexity, Quasiconvexity and Defocus-Invariant TexturesabstractIn this paper we analyze the convexity and the quasiconvexity of shape from focus/defocus and image restoration. We show that these problems are strictly quasiconvex for a family of Bregman's divergences, and in particular for least-squares. In addition to giving novel analytical insight to these problems, this study can be readily exploited to design algorithms: One can do away with global minimizers and obtain the same optimal solution by employing simple and efficient local methods. We experimentally validate this investigation by comparing two minimization algorithms: one based on a local method (gradient-flow) and another based on a global method (graph cuts). We show that both algorithms find the global optimum. Finally, we fully characterize defocus-invariant textures, a class of textures that do not allow depth recovery. We show how to decompose textures into defocus-invariant and defocus-varying components, and how this decomposition can be used to dramatically improve depth estimates. Paolo Favaro |
ICCV | 1 |
| 2007 | Boosting Invariance and Efficiency in Supervised LearningabstractIn this paper we present a novel boosting algorithm for supervised learning that incorporates invariance to data transformations and has high generalization capabilities. While one can incorporate invariance by adding virtual samples to the data (e.g., by jittering), we adopt a much more efficient strategy and work along the lines of vicinal risk minimization and tangent distance methods. As in vicinal risk minimization, we incorporate invariance to data by applying anisotropic smoothing along the directions of invariance. Moreover, as in tangent distance methods, we provide a simple local approximation to such directions, thus obtaining an efficient computational scheme. We also show that it is possible to automatically design optimal weak classifiers by using gradient descent. To increase efficiency at run time, such optimal weak classifiers are projected on a Haar basis. This results in designing strong classifiers that are more computationally efficient than in the case of exhaustive search. For illustration and validation purposes, we demonstrate the novel algorithm both on synthetic and on real data sets that are publicly available. Andrea Vedaldi, Paolo Favaro, Enrico Grisan |
ICCV | 2 |
| 2007 | DynamicBoost: Boosting Time Series Generated by Dynamical SystemsabstractBoosting is a remarkably simple and flexible classification algorithm with widespread applications in computer vision. However, the application of boosting to non-Euclidean, infinite length, and time-varying data, such as videos, is not straightforward. In dynamic textures, for example, the temporal evolution of image intensities is captured by a linear dynamical system, whose parameters live in a Stiefel manifold, which is clearly non-Euclidean. In this paper, we present a novel boosting method for the recognition of visual dynamical processes. Our key contribution is the design of weak classifiers (features) that are formulated as linear dynamical systems. The main advantage of such features is that they can be applied to infinitely long sequences and that they can be efficiently computed by solving a set of Sylvester equations. We also present an application of our method to dynamic texture classification. René Vidal, Paolo Favaro |
ICCV | 2 |
| 2006 | Defocus Inpainting
Paolo Favaro, Enrico Grisan |
ECCV (2) | 1 |
| 2005 | Visual Tracking in the Presence of Motion BlurabstractWe consider the problem of visual tracking of regions of interest in a sequence of motion blurred images. Traditional methods couple tracking with deblurring in order to correctly account for the effects of motion blur. Such coupling is usually appropriate, but computationally wasteful when visual tracking is the lone objective. Instead of deblurring images, we propose to match regions by blurring them. The matching score for two image regions is governed by a cost function that only involves the region deformation parameters and two motion blur vectors. We present an efficient algorithm to minimize the proposed cost function and demonstrate it on sequences of real blurred images. Hailin Jin, Paolo Favaro, Roberto Cipolla |
CVPR (2) | 2 |
| 2005 | KALMANSAC: Robust Filtering by ConsensusabstractWe propose an algorithm to perform causal inference of the state of a dynamical model when the measurements are corrupted by outliers. While the optimal (maximum-likelihood) solution has doubly exponential complexity due to the combinatorial explosion of possible choices of inliers, we exploit the structure of the problem to design a sampling-based algorithm that has constant complexity. We derive our algorithm from the equations of the optimal filter, which makes our approximation explicit. Our work is motivated by real-time tracking and the estimation of structure from motion (SFM). We test our algorithm for on-line outlier rejection both for tracking and for SFM. We show that our approach can tolerate a large proportion of outliers, whereas previous causal robust statistical inference methods failed with less than half as many. Our work can be thought of as the extension of random sample consensus algorithms to dynamic data, or as the implementation of pseudo-Bayesian filtering algorithms in a sampling framework. Andrea Vedaldi, Hailin Jin, Paolo Favaro, Stefano Soatto |
ICCV | 3 |
| 2005 | Using Frontier Points to Recover Shape, Reflectance and IllumunationabstractWe describe a method to recover the surface reflectance and the 3D shape of a non-Lambertian object as well as illumination, from a collection of images. It is based on the so-called frontier points, which are extracted from the outlines of an object. Frontier points provide 3D locations on the object surface where the surface normal is known. This information is exploited to infer the surface reflectance of the object and the light distribution of the scene both under varying illumination and fixed vantage point, and under varying vantage point and fixed illumination. We also show how to apply frontier points for shape recovery in photometric stereo. The effectiveness of frontier points for recovering reflectance, illumination and shape is confirmed by a number of experiments on both real and synthetic data. George Vogiatzis, Paolo Favaro, Roberto Cipolla |
ICCV | 2 |
| 2005 | A Geometric Approach to Shape from DefocusabstractWe introduce a novel approach to shape from defocus, i.e., the problem of inferring the three-dimensional (3D) geometry of a scene from a collection of defocused images. Typically, in shape from defocus, the task of extracting geometry also requires deblurring the given images. A common approach to bypass this task relies on approximating the scene locally by a plane parallel to the image (the so-called equifocal assumption). We show that this approximation is indeed not necessary, as one can estimate 3D geometry while avoiding deblurring without strong assumptions on the scene. Solving the problem of shape from defocus requires modeling how light interacts with the optics before reaching the imaging surface. This interaction is described by the so-called point spread function (PSF). When the form of the PSF is known, we propose an optimal method to infer 3D geometry from defocused images that involves computing orthogonal operators which are regularized via functional singular value decomposition. When the form of the PSF is unknown, we propose a simple and efficient method that first learns a set of projection operators from blurred images and then uses these operators to estimate the 3D geometry of the scene from novel blurred images. Our experiments on both real and synthetic images show that the performance of the algorithm is relatively insensitive to the form of the PSF. Our general approach is to minimize the Euclidean norm of the difference between the estimated images and the observed images. The method is geometric in that we reduce the minimization to performing projections onto linear subspaces, by using inner product structures on both infinite and finite-dimensional Hilbert spaces. Both proposed algorithms involve only simple matrix-vector multiplications which can be implemented in real-time. Paolo Favaro, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | A Variational Approach to Scene Reconstruction and Image Segmentation from Motion-Blur Cues
Paolo Favaro, Stefano Soatto |
CVPR (1) | 1 |
| 2004 | Scene and Motion Reconstruction from Defocused and Motion-Blurred Images via Anisotropic Diffusion
Paolo Favaro, Martin Burger 0001, Stefano Soatto |
ECCV (1) | 1 |
| 2003 | 3D Shape from Anisotropic DiffusionabstractWe cast the problem of inferring the 3D shape of a scene from a collection of defocused images in the framework of anisotropic diffusion. We propose an algorithm that can estimate the shape of a scene by inferring the diffusion coefficient of a heat equation. The method is optimal, as we pose it as the minimization of a certain cost functional based on the input images, and fast. Furthermore, we also extend our algorithm to the case of multiple images, and derive a 3D scene segmentation algorithm that can work in the presence of pictorial camouflage. Paolo Favaro, Stanley J. Osher, Stefano Soatto, Luminita A. Vese |
CVPR (1) | 1 |
| 2003 | Seeing Beyond Occlusions (and other marvels of a finite lens aperture)abstractWe present a novel algorithm to reconstruct the geometry and photometry of a scene with occlusions from a collection of defocused images. The presence of a finite lens aperture allows us to recover portions of the scene that would be occluded in a pin-hole projection, thus "uncovering" the occlusion. We estimate the shape of each object (a surface, including the occluding boundaries), and its radiance (a positive function defined on the surface, including portions that are occluded by other objects). Paolo Favaro, Stefano Soatto |
CVPR (2) | 1 |
| 2003 | Dynamic Texture SegmentationabstractWe address the problem of segmenting a sequence of images of natural scenes into disjoint regions that are characterized by constant spatio-temporal statistics. We model the spatio-temporal dynamics in each region by Gauss-Markov models, and infer the model parameters as well as the boundary of the regions in a variational optimization framework. Numerical results demonstrate that - in contrast to purely texture-based segmentation schemes - our method is effective in segmenting regions that differ in their dynamics even when spatial statistics are identical. Gianfranco Doretto, Daniel Cremers, Paolo Favaro, Stefano Soatto |
ICCV | 3 |
| 2003 | On Exploiting Occlusions in Multiple-view GeometryabstractOcclusions are commonplace in man-made and natural environments; they often result in photometric features where a line terminates at an occluding boundary, resembling a "T". We show that the 2-D motion of such T-junctions in multiple views carries nontrivial information on the 3-D structure of the scene and its motion relative to the camera. We show how the constraint among multiple views of T-junctions can be used to reliably detect them and differentiate them from ordinary point features. Finally, we propose an integrated algorithm to recursively and causally estimate structure and motion in the presence of T-junctions along with other point-features. Paolo Favaro, Alessandro Duci, Yi Ma 0001, Stefano Soatto |
ICCV | 1 |
| 2003 | Observing Shape from Defocused Images
Paolo Favaro, Andrea Mennucci, Stefano Soatto |
Int. J. Comput. Vis. | 1 |
| 2003 | A semi-direct approach to structure from motion
Hailin Jin, Paolo Favaro, Stefano Soatto |
Vis. Comput. | 2 |
| 2002 | Learning Shape from Defocus
Paolo Favaro, Stefano Soatto |
ECCV (2) | 1 |
| 2002 | A Variational Approach to Shape from Defocus
Hailin Jin, Paolo Favaro |
ECCV (2) | 2 |
| 2002 | Structure from Motion Causally Integrated Over TimeabstractWe describe an algorithm for reconstructing three-dimensional structure and motion causally, in real time from monocular sequences of images. We prove that the algorithm is minimal and stable, in the sense that the estimation error remains bounded with probability one throughout a sequence of arbitrary length. We discuss a scheme for handling occlusions (point features appearing and disappearing) and drift in the scale factor. These issues are crucial for the algorithm to operate in real time on real scenes. We describe in detail the implementation of the algorithm, which runs on a personal computer and has been made available to the community. We report the performance of our implementation on a few representative long sequences of real and synthetic images. The algorithm, which has been tested extensively over the course of the past few years, exhibits honest performance when the scene contains at least 20-40 points with high contrast, when the relative motion is "slow" compared to the sampling frequency of the frame grabber (30 Hz), and the lens aperture is "large enough" (typically more than 30/spl deg/ of visual field). Alessandro Chiuso, Paolo Favaro, Hailin Jin, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Real-time Virtual Object InsertionabstractWe present a system to insert virtual objects into real image sequences in real time. The system consists of offthe- shelf hardware (a camera connected to a Pentium PC) and software to (a) automatically select and track region features despite changes in illumination, (b) estimate threedimensional position and orientation of surface patches relative to an inertial reference frame despite individual pointfeatures appearing and disappearing, (c) insert a texturemapped virtual object into the scene so as to make it appear to be part of the scene and moving with it. This is all done in real time. The multi-thread C++ code, which is readily interfaced with a frame grabber as well as Matlab for development, will be made available to the public at the demonstration. Paolo Favaro, Hailin Jin, Stefano Soatto |
ICCV | 1 |
| 2001 | Real-Time Feature Tracking and Outlier Rejection with Changes in Illumination
Hailin Jin, Paolo Favaro, Stefano Soatto |
ICCV | 2 |
| 2000 | Real-Time 3-D Motion and Structure of Point-Features: A Front-End for Vision-Based Control and InteractionabstractWe present a system that consists of one camera connected to a personal computer that can (a) select and track a number of high-contrast point features on a sequence of images, (b) estimate their three-dimensional motion and position relative to an inertial reference frame, assuming rigidity, (c) handle occlusions that cause point-features to disappear as well as new features to appear. The system can also (d) perform partial self-calibration and (e) check for consistency of the rigidity assumption, although these features are not implemented in the current release. All of this is done automatically and in real-time (30 Hz) for 40-50 point features using commercial off-the-shelf hardware. The system is based on an algorithm presented by Chiuso et al. (2000), the properties of which have been analyzed by Chiuso and Soatto (2000). In particular, the algorithm is provably observable, provably minimal and provably stable- under suitable conditions. The core of the system, consisting of C++ code ready to interface with a frame grabber as well as Matlab code for development, is available at http://ee.wustl.edu/-soatto/research.html. We demonstrate the system by showing its use as (1) an ego-motion estimator, (2) an object tracker, and (3) an interactive input device, all without any modification of the system settings. Hailin Jin, Paolo Favaro, Stefano Soatto |
CVPR | 2 |
| 2000 | A Geometric Approach to Blind Deconvolution with Application to Shape from DefocuabstractWe propose a solution to the generic "bilinear calibration-estimation problem" when using a quadratic cost function and restricting to (locally) translation-invariant imaging models. We apply the solution to the problem of reconstructing the three-dimensional shape and radiance of a scene from a number of defocused images. Since the imaging process maps the continuum of three-dimensional space onto the discrete pixel grid, rather than discretizing the continuum we exploit the structure of maps between (finite-and infinite-dimensional) Hilbert spaces and arrive at a principled algorithm that does not involve any choice of basis or discretization. Rather, these are uniquely determined by the data, and exploited in a functional singular value decomposition in order to obtain a regularized solution. Stefano Soatto, Paolo Favaro |
CVPR | 2 |
| 2000 | 3-D Motion and Structure from 2-D Motion Causally Integrated over Time: Implementation
Alessandro Chiuso, Paolo Favaro, Hailin Jin, Stefano Soatto |
ECCV (2) | 2 |
| 2000 | Shape and Radiance Estimation from the Information-Divergence of Blurred Images
Paolo Favaro, Stefano Soatto |
ECCV (1) | 1 |