EDBT 2026 Demo / reviewers in the wild / expert
Stefano Soatto
dblp:08/1262
· DBLP profile ↗
271ranked-venue papers
18as first author
72since 2021 · last 2025
0000-0003-2902-6362ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 247 · 14 first-author · 71 since 2021Graphics, computer vision, multimedia, augmented reality and games · 184 · 15 first-author · 44 since 2021Systems, architecture and hardware · 9 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Computer networks · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling up Image Segmentation across Data and TasksabstractTraditional segmentation models, while effective in isolated tasks, often fail to generalize to more complex and open-ended segmentation problems, such as free-form, open-vocabulary, and in-the-wild scenarios. To bridge this gap, we propose to scale up image segmentation across diverse datasets and tasks such that the knowledge across different tasks and datasets can be integrated while improving the generalization ability. Mixed-Query Transformer (MQ-Former), a novel segmentation framework, is introduced and designed to scale seamlessly across both data size and task diversity. It is built upon a dynamic object query mechanism called mixed query, which fuses different types of queries using cross-attention. This hybrid approach enables the model to balance between instance- and stuff-level segmentation, providing enhanced scalability for handling diverse object types. We further enhance scalability by leveraging synthetic data-generating segmentation masks and captions for pixel-level and open-vocabulary tasks-drastically reducing the need for costly human annotations. By training on multiple datasets and tasks at scale, MQ-Former continuously improves performance as the volume and diversity of data and tasks increase. It exhibits strong generalization capabilities, boosting performance in open-set segmentation tasks SeginW by 7 points. These advancements mark a key step toward universal, scalable segmentation models capable of addressing the demands of real-world applications. Zhaowei Cai, Hao Yang 0043, Ashwin Swaminathan, R. Manmatha, Stefano Soatto |
CVPR | 6 |
| 2025 | PICASO: Permutation-Invariant Context Composition with State Space ModelsabstractProviding Large Language Models with relevant contextual knowledge at inference time has been shown to greatly improve the quality of their generations. This is often achieved by prepending informative passages of text, or 'contexts', retrieved from external knowledge bases to their input. However, processing additional contexts online incurs significant computation costs that scale with their length. State Space Models (SSMs) offer a promising solution by allowing a database of contexts to be mapped onto fixed-dimensional states from which to start the generation. A key challenge arises when attempting to leverage information present across multiple contexts, since there is no straightforward way to condition generation on multiple independent states in existing SSMs. To address this, we leverage a simple mathematical relation derived from SSM dynamics to compose multiple states into one that efficiently approximates the effect of concatenating raw context tokens. Since the temporal ordering of contexts can often be uninformative, we enforce permutation-invariance by efficiently averaging states obtained via our composition algorithm across all possible context orderings. We evaluate our resulting method on WikiText and MSMARCO in both zero-shot and fine-tuned settings, and show that we can match the strongest performing baseline while enjoying on average $5.4\times$ speedup. Tian Yu Liu, Alessandro Achille, Matthew Trager, Aditya Golatkar, Luca Zancato, Stefano Soatto |
ICLR | 6 |
| 2025 | Robust Planning for Autonomous Driving via Mixed Adversarial Diffusion PredictionsabstractWe describe a robust planning method for autonomous driving that mixes normal and adversarial agent predictions output by a diffusion model trained for motion prediction. We first train a diffusion model to learn an unbiased distribution of normal agent behaviors. We then generate a distribution of adversarial predictions by biasing the diffusion model at test time to generate predictions that are likely to collide with a candidate plan. We score plans using expected cost with respect to a mixture distribution of normal and adversarial predictions, leading to a planner that is robust against adversarial behaviors but not overly conservative when agents behave normally. Unlike current approaches, we do not use risk measures that over-weight adversarial behaviors while placing little to no weight on low-cost normal behaviors or use hard safety constraints that may not be appropriate for all driving scenarios. We show the effectiveness of our method on single-agent and multi-agent jaywalking scenarios as well as a red light violation scenario. Albert Zhao, Stefano Soatto |
ICRA | 2 |
| 2025 | STree: Speculative Tree Decoding for Hybrid State Space ModelsabstractSpeculative decoding is a technique to leverage hardware concurrency in order to enable multiple steps of token generation in a single forward pass, thus improving the efficiency of large-scale autoregressive (AR) Transformer models. State-space models (SSMs) are already more efficient than AR Transformers, since their state summarizes all past data with no need to cache or re-process tokens in the sliding window context. However, their state can also comprise thousands of tokens; so, speculative decoding has recently been extended to SSMs. Existing approaches, however, do not leverage the tree-based verification methods, since current SSMs lack the means to compute a token tree efficiently. We propose the first scalable algorithm to perform tree-based speculative decoding in state-space models (SSMs) and hybrid architectures of SSMs and Transformer layers. We exploit the structure of accumulated state transition matrices to facilitate tree-based speculative decoding with minimal overhead relative to current SSM implementations. Along with the algorithm, we describe a hardware-aware implementation that improves naive application of AR Transformer tree-based speculative decoding methods to SSMs. Furthermore, we outperform vanilla speculative decoding with SSMs even with a baseline drafting model and tree structure on three different benchmarks, opening up opportunities for further speed up with SSM and hybrid model inference. Code can be find at: https://github.com/wyc1997/stree. Yangchao Wu, Zongyue Qin, Alex Wong 0001, Stefano Soatto |
NeurIPS | 4 |
| 2024 | Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-EncodingabstractQuantifying the degree of similarity between images is a key copyright issue for image-based machine learning. In legal doctrine however, determining the degree of similarity between works requires subjective analysis, and fact-finders (judges and juries) can demonstrate considerable variabil-ity in these subjective judgement calls. Images that are structurally similar can be deemed dissimilar, whereas images of completely different scenes can be deemed similar enough to support a claim of copying. We seek to define and compute a notion of ‘conceptual similarity’ among images that captures high-level relations even among images that do not share repeated elements or visually similar components. The idea is to use a base multi-modal model to gen-erate 'explanations' (captions) of visual data at increasing levels of complexity. Then, similarity can be measured by the length of the caption needed to discriminate between the two images: Two highly dissimilar images can be dis-criminated early in their description, whereas conceptually dissimilar ones will need more detail to be distinguished. We operationalize this definition and show that it correlates with subjective (averaged human evaluation) assessment, and beats existing baselines on both image-to-image and text-to-text similarity benchmarks. Beyond just providing a number, our method also offers interpretability by pointing to the specific level of granularity of the description where the source data are differentiated. Alessandro Achille, Greg Ver Steeg, Tian Yu Liu, Matthew Trager, Carson Klingenberg, Stefano Soatto |
CVPR | 6 |
| 2024 | Multi-Modal Hallucination Control by Visual Information GroundingabstractGenerative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phenomenon, usually referred to as “hallucination” and show that it stems from an excessive reliance on the language prior. In particular, we show that as more tokens are generated, the reliance on the visual prompt decreases, and this behavior strongly correlates with the emergence of hallucinations. To reduce hallucinations, we introduce Multi-Modal Mutual-Information Decoding (M3ID), a new sampling method for prompt amplification. M3ID amplifies the influence of the reference image over the language prior, hence favoring the generation of tokens with higher mutual information with the visual prompt. M3ID can be applied to any pre-trained autoregressive VLM at inference time without necessitating further training and with minimal computational overhead. If training is an option, we show that M3ID can be paired with Direct Preference Optimization (DPO) to improve the model's reliance on the prompt image without requiring any labels. Our empirical findings show that our algorithms maintain the fluency and linguistic capabilities of pre-trained VLMs while reducing hallucinations by mitigating visually ungrounded answers. Specifically, for the LLaVA 13B model, M3ID and M3ID+DPO reduce the percentage of hallucinated objects in captioning tasks by 25% and 28%, respectively, and improve the accuracy on VQA benchmarks such as POPE by 21% and 24%. Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto |
CVPR | 8 |
| 2024 | Enhancing Vision-Language Pre-Training with Rich SupervisionsabstractWe propose Strongly Supervised pre-training with ScreenShots (S4) - a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. Using web screenshots unlocks a treasure trove of visual and textual cues that are not present in using image-text pairs. In S4, we leverage the inherent tree-structured hierarchy of HTML elements and the spatial localization to carefully design 10 pre-training tasks with large scale annotated data. These tasks resemble down- stream tasks across different domains and the annotations are cheap to obtain. We demonstrate that, compared to current screenshot pre-training objectives, our innovative pre-training method significantly enhances performance of image-to-text model in nine varied and popular downstream tasks - up to 76.1% improvements on Table Detection, and at least 1 % on Widget Captioning. Kunyu Shi, Pengkai Zhu, Edouard Belval, Oren Nuriel, Srikar Appalaraju, Shabnam Ghadar, Zhuowen Tu, Vijay Mahadevan, Stefano Soatto |
CVPR | 10 |
| 2024 | CPR: Retrieval Augmented Generation for Copyright ProtectionabstractRetrieval Augmented Generation (RAG) is emerging as a flexible and robust technique to adapt models to private users data without training, to handle credit attribution, and to allow efficient machine unlearning at scale. However, RAG techniques for image generation may lead to parts of the retrieved samples being copied in the model's output. To reduce risks of leaking private information contained in the retrieved set, we introduce Copy-Protected generation with Retrieval (CPR), a new method for RAG with strong copyright protection guarantees in a mixed-private setting for diffusion models. CPR allows to condition the output of diffusion models on a set of retrieved images, while also guaranteeing that unique identifiable information about those example is not exposed in the generated outputs. In particular, it does so by sampling from a mixture of public (safe) distribution and private (user) distribution by merging their diffusion scores at inference. We prove that CPR satisfies Near Access Freeness (NAF) which bounds the amount of information an attacker may be able to extract from the generated images. We provide two algorithms for copyright protection, CPR-KL and CPR-Choose. Unlike previously proposed rejection-sampling-based NAF methods, our methods enable efficient copyright-protected sampling with a single run of backward diffusion. We show that our method can be applied to any pre-trained conditional diffusion model, such as Stable Diffusion or unCLIP. In particular, we empirically show that applying CPR on top of unCLIP improves quality and text-to-image alignment of the generated results (81.4 to 83.17 on TIFA benchmark), while enabling credit attribution, copy-right protection, and deterministic, constant time, unlearning. Aditya Golatkar, Alessandro Achille, Luca Zancato, Yu-Xiang Wang 0003, Ashwin Swaminathan, Stefano Soatto |
CVPR | 6 |
| 2024 | THRONE: An Object-Based Hallucination Benchmark for the Free-Form Generations of Large Vision-Language ModelsabstractMitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term “Type I hallucinations”. Instead, they focus on hallucinations responding to very specific question formats-typically a multiple-choice response regarding a particular object or attribute-which we term “Type II hallucinations”. Additionally, such benchmarks often require external API calls to models which are subject to change. In practice, we observe that a reduction in Type II hallucinations does not lead to a reduction in Type I hallucinations but rather that the two forms of halluci-nations are often anti-correlated. To address this, we propose THRONE, a novel object-based automatic framework for quantitatively evaluating Type I hallucinations in LVLM free-form outputs. We use public language models (LMs) to identify hallucinations in LVLM responses and compute informative metrics. By evaluating a large selection of recent LVLMs using public datasets, we show that an improvement in existing metrics do not lead to a reduction in Type I hallucinations, and that established benchmarks for measuring Type I hallucinations are incomplete. Finally, we provide a simple and effective data augmentation method to reduce Type I and Type II hallucinations as a strong baseline. Prannay Kaul, Zhizhong Li 0001, Hao Yang 0043, Yonatan Dukler, Ashwin Swaminathan, C. J. Taylor, Stefano Soatto |
CVPR | 7 |
| 2024 | Diffeomorphic Template Registration for Atmospheric Turbulence MitigationabstractWe describe a method for recovering the irradiance underlying a collection of images corrupted by atmospheric turbulence. Since supervised data is often technically im-possible to obtain, assumptions and biases have to be im-posed to solve this inverse problem, and we choose to model them explicitly. Rather than initializing a latent irradiance (“template”) by heuristics to estimate deformation, we se-lect one of the images as a reference, and model the de-formation in this image by the aggregation of the optical flow from it to other images, exploiting a prior imposed by Central Limit Theorem. Then with a novel flow inversion module, the model registers each image TO the template but WITHOUT the template, avoiding artifacts related to poor template initialization. To illustrate the robustness of the method, we simply (i) select the first frame as the ref-erence and (ii) use the simplest optical flow to estimate the warpings, yet the improvement in registration is decisive in the final reconstruction, as we achieve state-of-the-art per-formance despite its simplicity. The method establishes a strong baseline that can be further improved by integrating it seamlessly into more sophisticated pipelines, or with domain-specific methods if so desired. Dong Lao, Congli Wang, Alex Wong 0001, Stefano Soatto |
CVPR | 4 |
| 2024 | On the Scalability of Diffusion-based Text-to-Image GenerationabstractScaling up model and data size has been quite successful for the evolution of LLMs. However, the scaling law for the diffusion based text-to-image (T2I) models is not fully explored. It is also unclear how to efficiently scale the model for better performance at reduced cost. The different training settings and expensive training cost make a fair model comparison extremely difficult. In this work, we empirically study the scaling properties of diffusion based T2I models by performing extensive and rigours ablations on scaling both denoising backbones and training set, including training scaled UNet and Transformer variants ranging from 0.4B to 4B parameters on datasets upto 600M images. For model scaling, we find the location and amount of cross attention distinguishes the performance of existing UNet designs. And increasing the transformer blocks is more parameter-efficient for improving text-image alignment than increasing channel numbers. We then identify an efficient UNet variant, which is 45% smaller and 28% faster than SDXL's UNet. On the data scaling side, we show the quality and diversity of the training set matters more than simply dataset size. Increasing caption density and diversity improves text-image alignment performance and the learning efficiency. Finally, we provide scaling functions to predict the text-image alignment performance as functions of the scale of model size, compute and dataset size. Orchid Majumder, Yusheng Xie, R. Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, Stefano Soatto |
CVPR | 10 |
| 2024 | Non-autoregressive Sequence-to-Sequence Vision-Language ModelsabstractSequence-to-sequence vision-language models are showing promise, but their applicability is limited by their inference latency due to their autoregressive way of generating predictions. We propose a parallel decoding sequence-to-sequence vision-language model, trained with a Query-CTC loss, that marginalizes over multiple inference paths in the decoder. This allows us to model the joint distribution of tokens, rather than restricting to conditional distribution as in an autoregressive model. The resulting model, NARVL, achieves performance on-par with its state-of-the-art autoregressive counterpart, but is faster at inference time, reducing from the linear complexity associated with the sequential generation of tokens to a paradigm of constant time joint inference. Kunyu Shi, Luis Goncalves, Zhuowen Tu, Stefano Soatto |
CVPR | 5 |
| 2024 | WorDepth: Variational Language Prior for Monocular Depth EstimationabstractThree-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities, i. e. scale. Predicting a 3D scene from text descriptionis) is similarly ill-posed, i. e. spatial arrangements of objects described. We investigate the question of whether two inher-ently ambiguous modalities can be used in conjunction to produce metric-scaled reconstructions. To test this, we fo-cus on monocular depth estimation, the problem of predicting a dense depth map from a single image, but with an additional text caption describing the scene. To this end, we begin by encoding the text caption as a mean and standard deviation; using a variational framework, we learn the distribution of the plausible metric reconstructions of 3D scenes corresponding to the text captions as a prior. To “select” a specific reconstruction or depth map, we encode the given image through a conditional sampler that samples from the latent space of the variational text encoder, which is then decoded to the output depth map. Our approach is trained alternatingly between the text and image branches: in one optimization step, we predict the mean and standard deviation from the text description and sample from a standard Gaussian, and in the other, we sample using a (image) conditional sampler. Once trained, we directly predict depth from the encoded text using the conditional sampler. We demonstrate our approach on indoor (NYUv2) and out-door (KITTI) scenarios, where we show that language can consistently improve performance in both. Code: https://github.com/Adonis-galaxy/WorDepth. Ziyao Zeng, Daniel Wang 0005, Fengyu Yang 0003, Hyoungseob Park, Stefano Soatto, Dong Lao, Alex Wong 0001 |
CVPR | 5 |
| 2024 | Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
Benjamin Biggs, Arjun Seshadri, Achin Jain, Aditya Golatkar, Yusheng Xie, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto |
ECCV (63) | 9 |
| 2024 | On the Viability of Monocular Depth Pre-training for Semantic Segmentation
Dong Lao, Fengyu Yang 0003, Daniel Wang 0005, Hyoungseob Park, Samuel Lu, Alex Wong 0001, Stefano Soatto |
ECCV (37) | 7 |
| 2024 | AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation
Yangchao Wu, Tian Yu Liu, Hyoungseob Park, Stefano Soatto, Dong Lao, Alex Wong 0001 |
ECCV (64) | 4 |
| 2024 | DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding ModelsabstractSungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, Stefano Soatto. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang 0005, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, Stefano Soatto |
EMNLP | 9 |
| 2024 | Critical Learning Periods Emerge Even in Deep Linear NetworksabstractCritical learning periods are periods early in development where temporary sensory deficits can have a permanent effect on behavior and learned representations.
Despite the radical differences between biological and artificial networks, critical learning periods have been empirically observed in both systems. This suggests that critical periods may be fundamental to learning and not an accident of biology.
Yet, why exactly critical periods emerge in deep networks is still an open question, and in particular it is unclear whether the critical periods observed in both systems depend on particular architectural or optimization details. To isolate the key underlying factors, we focus on deep linear network models, and show that, surprisingly, such networks also display much of the behavior seen in biology and artificial networks, while being amenable to analytical treatment. We show that critical periods depend on the depth of the model and structure of the data distribution. We also show analytically and in simulations that the learning of features is tied to competition between sources. Finally, we extend our analysis to multi-task learning to show that pre-training on certain tasks can damage the transfer performance on new tasks, and show how this depends on the relationship between tasks and the duration of the pre-training stage. To the best of our knowledge, our work provides the first analytically tractable model that sheds light into why critical learning periods emerge in biological and artificial networks. Michael Kleinman, Alessandro Achille, Stefano Soatto |
ICLR | 3 |
| 2024 | Tangent Transformers for Composition, Privacy and RemovalabstractWe introduce Tangent Attention Fine-Tuning (TAFT), a method for fine-tuning linearized transformers obtained by computing a First-order Taylor Expansion around a pre-trained initialization. We show that the Jacobian-Vector Product resulting from linearization can be computed efficiently in a single forward pass, reducing training and inference cost to the same order of magnitude as its original non-linear counterpart, while using the same number of parameters. Furthermore, we show that, when applied to various downstream visual classification tasks, the resulting Tangent Transformer fine-tuned with TAFT can perform comparably with fine-tuning the original non-linear network. Since Tangent Transformers are linear with respect to the new set of weights, and the resulting fine-tuning loss is convex, we show that TAFT enjoys several advantages compared to non-linear fine-tuning when it comes to model composition, parallel training, machine unlearning, and differential privacy. Our code is available at: https://github.com/tianyu139/tangent-model-composition Tian Yu Liu, Aditya Golatkar, Stefano Soatto |
ICLR | 3 |
| 2024 | Meaning Representations from Trajectories in Autoregressive ModelsabstractWe propose to extract meaning representations from autoregressive language models by considering the distribution of all possible trajectories extending an input text. This strategy is prompt-free, does not require fine-tuning, and is applicable to any pre-trained autoregressive model. Moreover, unlike vector-based representations, distribution-based representations can also model asymmetric relations (e.g., direction of logical entailment, hypernym/hyponym relations) by using algebraic operations between likelihood functions. These ideas are grounded in distributional perspectives on semantics and are connected to standard constructions in automata theory, but to our knowledge they have not been applied to modern language models. We empirically show that the representations obtained from large models align well with human annotations, outperform other zero-shot and prompt-free methods on semantic similarity tasks, and can be used to solve more complex entailment and containment tasks that standard embeddings cannot handle. Finally, we extend our method to represent data from different modalities (e.g., image and text) using multimodal autoregressive models. Our code is available at: https://github.com/tianyu139/meaning-as-trajectories Tian Yu Liu, Matthew Trager, Alessandro Achille, Pramuditha Perera, Luca Zancato, Stefano Soatto |
ICLR | 6 |
| 2024 | Fewer Truncations Improve Language ModelingabstractIn large language model training, input documents are typically concatenated together and then split into sequences of equal length to avoid padding tokens. Despite its efficiency, the concatenation approach compromises data integrity—it inevitably breaks many documents into incomplete pieces, leading to excessive truncations that hinder the model from learning to compose logically coherent and factually consistent content that is grounded on the complete context. To address the issue, we propose Best-fit Packing, a scalable and efficient method that packs documents into training sequences through length-aware combinatorial optimization. Our method completely eliminates unnecessary truncations while retaining the same training efficiency as concatenation. Empirical results from both text and code pre-training show that our method achieves superior performance (e.g., +4.7% on reading comprehension; +16.8% in context following; and +9.2% on program synthesis), and reduces closed-domain hallucination effectively by up to 58.3%. Hantian Ding, Zijian Wang 0002, Giovanni Paolini, Anoop Deoras, Dan Roth 0001, Stefano Soatto |
ICML | 7 |
| 2024 | Sub-token ViT Embedding via Stochastic Resonance TransformersabstractVision Transformer (ViT) architectures represent images as collections of high-dimensional vectorized tokens, each corresponding to a rectangular non-overlapping patch. This representation trades spatial granularity for embedding dimensionality, and results in semantically rich but spatially coarsely quantized feature maps. In order to retrieve spatial details beneficial to fine-grained inference tasks we propose a training-free method inspired by "stochastic resonance." Specifically, we perform sub-token spatial transformations to the input data, and aggregate the resulting ViT features after applying the inverse transformation. The resulting "Stochastic Resonance Transformer" (SRT) retains the rich semantic information of the original representation, but grounds it on a finer-scale spatial domain, partly mitigating the coarse effect of spatial tokenization. SRT is applicable across any layer of any ViT architecture, consistently boosting performance on several tasks including segmentation, classification, depth estimation, and others by up to 14.9% without the need for any fine-tuning. Code: https://github.com/donglao/srt. Dong Lao, Yangchao Wu, Tian Yu Liu, Alex Wong 0001, Stefano Soatto |
ICML | 5 |
| 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading MemoryabstractWe describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resources for inference. Current architectures use such resources to represent data either eidetically over a finite span ('context' in Transformers), or fading over an infinite span (in State Space Models, or SSMs). Recent hybrid architectures have combined eidetic and fading memory, but with limitations that do not allow the designer or the learning process to seamlessly modulate the two, nor to extend the eidetic memory span. We leverage ideas from Stochastic Realization Theory to develop a class of models called B'MOJO to seamlessly combine eidetic and fading memory within an elementary composable module. The overall architecture can be used to implement models that can access short-term eidetic memory 'in-context,' permanent structural memory 'in-weights,' fading memory 'in-state,' and long-term eidetic memory 'in-storage' by natively incorporating retrieval from an asynchronously updated memory. We show that Transformers, existing SSMs such as Mamba, and hybrid architectures such as Jamba are special cases of B'MOJO and describe a basic implementation that can be stacked and scaled efficiently in hardware. We test B'MOJO on transductive inference tasks, such as associative recall, where it outperforms existing SSMs and Hybrid models; as a baseline, we test ordinary language modeling where B'MOJO achieves perplexity comparable to similarly-sized Transformers and SSMs up to 1.4B parameters, while being up to 10% faster to train. Finally, we test whether models trained inductively on a-priori bounded sequences (up to 8K tokens) can still perform transductive inference on sequences many-fold longer. B'MOJO's ability to modulate eidetic and fading memory results in better inference on longer sequences tested up to 32K tokens, four-fold the length of the longest sequences seen during training. Luca Zancato, Arjun Seshadri, Yonatan Dukler, Aditya Golatkar, Yantao Shen 0002, Benjamin Bowman, Matthew Trager, Alessandro Achille, Stefano Soatto |
NeurIPS | 9 |
| 2024 | RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language DescriptionsabstractWe propose a method for metric-scale monocular depth estimation. Inferring depth from a single image is an ill-posed problem due to the loss of scale from perspective projection during the image formation process. Any scale chosen is a bias, typically stemming from training on a dataset; hence, existing works have instead opted to use relative (normalized, inverse) depth. Our goal is to recover metric-scaled depth maps through a linear transformation. The crux of our method lies in the observation that certain objects (e.g., cars, trees, street signs) are typically found or associated with certain types of scenes (e.g., outdoor). We explore whether language descriptions can be used to transform relative depth predictions to those in metric scale. Our method, RSA , takes as input a text caption describing objects present in an image and outputs the parameters of a linear transformation which can be applied globally to a relative depth map to yield metric-scaled depth predictions. We demonstrate our method on recent general-purpose monocular depth models on indoors (NYUv2, VOID) and outdoors (KITTI). When trained on multiple datasets, RSA can serve as a general alignment module in zero-shot settings. Our method improves over common practices in aligning relative to metric depth and results in predictions that are comparable to an upper bound of fitting relative depth to ground truth via a linear transformation. Code is available at: https://github.com/Adonis-galaxy/RSA. Ziyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang 0005, Fengyu Yang 0003, Stefano Soatto, Dong Lao, Byung-Woo Hong, Alex Wong 0001 |
NeurIPS | 6 |
| 2024 | Elodi: Ensemble Logit Difference Inhibition for Positive-Congruent TrainingabstractNegative flips are errors introduced in a classification system when a legacy model is updated. Existing methods to reduce the negative flip rate (NFR) either do so at the expense of overall accuracy by forcing a new model to imitate the old models, or use ensembles, which multiply inference cost prohibitively. We analyze the role of ensembles in reducing NFR and observe that they remove negative flips that are typically not close to the decision boundary, but often exhibit large deviations in the distance among their logits. Based on the observation, we present a method, called Ensemble Logit Difference Inhibition (ELODI), to train a classification system that achieves paragon performance in both error rate and NFR, at the inference cost of a single model. The method distills a homogeneous ensemble to a single student model which is used to update the classification system. ELODI also introduces a generalized distillation objective, Logit Difference Inhibition (LDI), which only penalizes the logit difference of a subset of classes with the highest logit values. On multiple image classification benchmarks, model updates with ELODI demonstrate superior accuracy retention and NFR reduction. Yue Zhao 0006, Yantao Shen 0002, Yuanjun Xiong, Shuo Yang 0003, Wei Xia 0009, Zhuowen Tu, Bernt Schiele, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Graph Spectral Embedding using the Geodesic Betweenness CentralityabstractWe introduce the Graph Sylvester Embedding (GSE), an unsupervised graph representation of local similarity, connectivity, and global structure. GSE uses the solution of the Sylvester equation to capture both network structure and neighborhood proximity in a single representation. Unlike embeddings based on the eigenvectors of the Laplacian, GSE incorporates two or more basis functions, for instance using the Laplacian and the affinity matrix. Such basis functions are constructed not from the original graph, but from one whose weights measure the centrality of an edge (the fraction of the number of shortest paths that pass through that edge) in the original graph. This allows more flexibility and control to represent complex network structure and shows significant improvements over the state of the art when used for data analysis tasks such as predicting failed edges in material science and network alignment in the human-SARS CoV-2 protein-protein interactome. Shay Deutsch, Stefano Soatto |
AISTATS | 2 |
| 2023 | À-la-carte Prompt Tuning (APT): Combining Distinct Data Via Composable PromptingabstractWe introduce À-la-carte Prompt Tuning (APT), a transformer-based scheme to tune prompts on distinct data so that they can be arbitrarily composed at inference time. The individual prompts can be trained in isolation, possibly on different devices, at different times, and on different distributions or domains. Furthermore each prompt only contains information about the subset of data it was exposed to during training. During inference, models can be assembled based on arbitrary selections of data sources, which we call à-la-carte learning. À-la-carte learning enables constructing bespoke models specific to each user's individual access rights and preferences. We can add or remove information from the model by simply adding or removing the corresponding prompts without retraining from scratch. We demonstrate that à-la-carte built models achieve accuracy within 5% of models trained on the union of the respective sources, with comparable cost in terms of training and inference time. For the continual learning benchmarks Split CIFAR- 100 and CORe50, we achieve state-of-the-art performance. Benjamin Bowman, Alessandro Achille, Luca Zancato, Matthew Trager, Pramuditha Perera, Giovanni Paolini, Stefano Soatto |
CVPR | 7 |
| 2023 | A Meta-Learning Approach to Predicting Performance and Data RequirementsabstractWe propose an approach to estimate the number of samples required for a model to reach a target performance. We find that the power law, the de facto principle to estimate model performance, leads to a large error when using a small dataset (e.g., 5 samples per class) for extrapolation. This is because the log-performance error against the log-dataset size follows a nonlinear progression in the few-shot regime followed by a linear progression in the high-shot regime. We introduce a novel piecewise power law (PPL) that handles the two data regimes differently. To estimate the parameters of the PPL, we introduce a random forest regressor trained via meta learning that generalizes across classification/detection tasks, ResNet/ViT based architectures, and random/pre-trained initializations. The PPL improves the performance estimation on average by 37% across 16 classification and 33% across 10 detection datasets, compared to the power law. We further extend the PPL to provide a confidence bound and use it to limit the prediction horizon that reduces over-estimation of data by 76% on classification and 91% on detection datasets. Achin Jain, Gurumurthy Swaminathan, Paolo Favaro, Hao Yang 0043, Avinash Ravichandran, Hrayr Harutyunyan, Alessandro Achille, Onkar Dabeer, Bernt Schiele, Ashwin Swaminathan, Stefano Soatto |
CVPR | 11 |
| 2023 | Critical Learning Periods for Multisensory Integration in Deep NetworksabstractWe show that the ability of a neural network to integrate information from diverse sources hinges critically on being exposed to properly correlated signals during the early phases of training. Interfering with the learning process during this initial stage can permanently impair the development of a skill, both in artificial and biological systems where the phenomenon is known as a critical learning period. We show that critical periods arise from the complex and unstable early transient dynamics, which are decisive of final performance of the trained system and their learned representations. This evidence challenges the view, engendered by analysis of wide and shallow networks, that early learning dynamics of neural networks are simple, akin to those of a linear model. Indeed, we show that even deep linear networks exhibit critical learning periods for multi-source integration, while shallow networks do not. To better understand how the internal representations change according to disturbances or sensory deficits, we introduce a new measure of source sensitivity, which allows us to track the inhibition and integration of sources during training. Our analysis of inhibition suggests cross-source reconstruction as a natural auxiliary training objective, and indeed we show that architectures trained with cross-sensor reconstruction objectives are remarkably more resilient to critical periods. Our findings suggest that the recent success in self-supervised multi-modal training compared to previous supervised efforts may be in part due to more robust learning dynamics and not solely due to better architectures and/or more data. Michael Kleinman, Alessandro Achille, Stefano Soatto |
CVPR | 3 |
| 2023 | Guided Recommendation for Model Fine-TuningabstractModel selection is essential for reducing the search cost of the best pre-trained model over a large-scale model zoo for a downstream task. After analyzing recent hand-designed model selection criteria with 400+ ImageNet pre-trained models and 40 downstream tasks, we find that they can fail due to invalid assumptions and intrinsic limitations. The prior knowledge on model capacity and dataset also can not be easily integrated into the existing criteria. To address these issues, we propose to convert model selection as a recommendation problem and to learn from the past training history. Specifically, we characterize the meta information of datasets and models as features, and use their transfer learning performance as the guided score. With thousands of historical training jobs, a recommendation system can be learned to predict the model selection score given the features of the dataset and the model as input. Our approach enables integrating existing model selection scores as additional features and scales with more historical data. We evaluate the prediction accuracy with 22 pre-trained models over 40 downstream tasks. With extensive evaluations, we show that the learned approach can outperform prior hand-designed model selection methods significantly when relevant training history is available. Charless C. Fowlkes, Hao Yang 0043, Onkar Dabeer, Zhuowen Tu, Stefano Soatto |
CVPR | 6 |
| 2023 | Depth Estimation from Camera Image and mmWave Radar Point CloudabstractWe present a method for inferring dense depth from a camera image and a sparse noisy radar point cloud. We first describe the mechanics behind mmWave radar point cloud formation and the challenges that it poses, i.e. ambiguous elevation and noisy depth and azimuth components that yields incorrect positions when projected onto the image, and how existing works have overlooked these nuances in camera-radar fusion. Our approach is motivated by these mechanics, leading to the design of a network that maps each radar point to the possible surfaces that it may project onto in the image plane. Unlike existing works, we do not process the raw radar point cloud as an erroneous depth map, but query each raw point independently to associate it with likely pixels in the image – yielding a semi-dense radar depth map. To fuse radar depth with an image, we propose a gated fusion scheme that accounts for the confidence scores of the correspondence so that we selectively combine radar and camera embeddings to yield a dense depth map. We test our method on the NuScenes benchmark and show a 10.3% improvement in mean absolute error and a 9.1% improvement in root-mean-square error over the best method. Code: https://github.com/nesl/radar-camera-fusion-depth. Akash Deep Singh, Yunhao Ba, Ankur Sarker, Howard Zhang, Achuta Kadambi, Stefano Soatto, Mani Srivastava 0001, Alex Wong 0001 |
CVPR | 6 |
| 2023 | Train/Test-Time Adaptation with RetrievalabstractWe introduce Train/Test-Time Adaptation with Retrieval (T3AR), a method to adapt models both at train and test time by means of a retrieval module and a searchable pool of external samples. Before inference, T3AR adapts a given model to the downstream task using refined pseudo-labels and a self-supervised contrastive objective function whose noise distribution leverages retrieved real samples to improve feature adaptation on the target data manifold. The retrieval of real images is key to T3AR since it does not rely solely on synthetic data augmentations to compensate for the lack of adaptation data, as typically done by other adaptation algorithms. Furthermore, thanks to the retrieval module, our method gives the user or service provider the possibility to improve model adaptation on the downstream task by incorporating further relevant data or to fully remove samples that may no longer be available due to changes in user preference after deployment. First, we show that T3AR can be used at training time to improve downstream fine-grained classification over standard fine-tuning baselines, and the fewer the adaptation data the higher the relative improvement (up to 13%). Second, we apply T3ARfor test-time adaptation and show that exploiting a pool of external images at test-time leads to more robust representations over existing methods on DomainNet-126 and VISDA-C, especially when few adaptation data are available (up to 8%). Luca Zancato, Alessandro Achille, Tian Yu Liu, Matthew Trager, Pramuditha Perera, Stefano Soatto |
CVPR | 6 |
| 2023 | SAFE: Machine Unlearning With Shard GraphsabstractWe present Synergy Aware Forgetting Ensemble (SAFE), a method to adapt large models on a diverse collection of data while minimizing the expected cost to remove the influence of training samples from the trained model. This process, also known as selective forgetting or unlearning, is often conducted by partitioning a dataset into shards, training fully independent models on each, then ensembling the resulting models. Increasing the number of shards reduces the expected cost to forget but at the same time it increases inference cost and reduces the final accuracy of the model since synergistic information between samples is lost during the independent model training. Rather than treating each shard as independent, SAFE introduces the notion of a shard graph, which allows incorporating limited information from other shards during training, trading off a modest increase in expected forgetting cost with a significant increase in accuracy, all while still attaining complete removal of residual influence after forgetting. SAFE uses a lightweight system of adapters which can be trained while reusing most of the computations. This allows SAFE to be trained on shards an order-of-magnitude smaller than current state-of-the-art methods (thus reducing the forgetting costs) while also maintaining high accuracy, as we demonstrate empirically on fine-grained computer vision datasets. Yonatan Dukler, Benjamin Bowman, Alessandro Achille, Aditya Golatkar, Ashwin Swaminathan, Stefano Soatto |
ICCV | 6 |
| 2023 | Tangent Model Composition for Ensembling and Continual Fine-tuningabstractTangent Model Composition (TMC) is a method to combine component models independently fine-tuned around a pre-trained point. Component models are tangent vectors to the pre-trained model that can be added, scaled, or subtracted to support incremental learning, ensembling, or unlearning. Component models are composed at inference time via scalar combination, reducing the cost of ensembling to that of a single model. TMC improves accuracy by 4.2% compared to ensembling non-linearly fine-tuned models at a 2.5× to 10× reduction of inference cost, growing linearly with the number of component models. Each component model can be forgotten at zero cost, with no residual effect on the resulting inference. When used for continual fine-tuning, TMC is not constrained by sequential bias and can be executed in parallel on federated data. TMC outperforms recently published continual fine-tuning methods almost uniformly on each setting – task-incremental, class-incremental, and data-incremental – on a total of 13 experiments across 3 benchmark datasets, despite not using any replay buffer. TMC is designed for composing models that are local to a pre-trained embedding, but could be extended to more general settings. The code is available at: https://github.com/tianyu139/tangent-model-composition Tian Yu Liu, Stefano Soatto |
ICCV | 2 |
| 2023 | Linear Spaces of Meanings: Compositional Structures in Vision-Language ModelsabstractWe investigate compositional structures in data embeddings from pre-trained vision-language models (VLMs). Traditionally, compositionality has been associated with algebraic operations on embeddings of words from a preexisting vocabulary. In contrast, we seek to approximate representations from an encoder as combinations of a smaller set of vectors in the embedding space. These vectors can be seen as "ideal words" for generating concepts directly within embedding space of the model. We first present a framework for understanding compositional structures from a geometric perspective. We then explain what these compositional structures entail probabilistically in the case of VLM embeddings, providing intuitions for why they arise in practice. Finally, we empirically explore these structures in CLIP’s embeddings and we evaluate their usefulness for solving different vision-language tasks such as classification, debiasing, and retrieval. Our results show that simple linear algebraic operations on embedding vectors can be used as compositional and interpretable methods for regulating the behavior of VLMs. Matthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille, Parminder Bhatia, Stefano Soatto |
ICCV | 6 |
| 2023 | Masked Vision and Language Modeling for Multi-modal Representation Learning
Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, Stefano Soatto |
ICLR | 6 |
| 2023 | Your representations are in the network: composable and parallel adaptation for large scale modelsabstractWe present a framework for transfer learning that efficiently adapts a large base-model by learning lightweight cross-attention modules attached to its intermediate activations.
We name our approach InCA (Introspective-Cross-Attention) and show that it can efficiently survey a network’s representations and identify strong performing adapter models for a downstream task.
During training, InCA enables training numerous adapters efficiently and in parallel, isolated from the frozen base model. On the ViT-L/16 architecture, our experiments show that a single adapter, 1.3% of the full model, is able to reach full fine-tuning accuracy on average across 11 challenging downstream classification tasks.
Compared with other forms of parameter-efficient adaptation, the isolated nature of the InCA adaptation is computationally desirable for large-scale models. For instance, we adapt ViT-G/14 (1.8B+ parameters) quickly with 20+ adapters in parallel on a single V100 GPU (76% GPU memory reduction) and exhaustively identify its most useful representations.
We further demonstrate how the adapters learned by InCA can be incrementally modified or combined for flexible learning scenarios and our approach achieves state of the art performance on the ImageNet-to-Sketch multi-task benchmark. Yonatan Dukler, Alessandro Achille, Hao Yang 0043, Varsha Vivek, Luca Zancato, Benjamin Bowman, Avinash Ravichandran, Charless C. Fowlkes, Ashwin Swaminathan, Stefano Soatto |
NeurIPS | 10 |
| 2023 | Leveraging sparse and shared feature activations for disentangled representation learningabstractRecovering the latent factors of variation of high dimensional data has so far focused on simple synthetic settings. Mostly building on unsupervised and weakly-supervised objectives, prior work missed out on the positive implications for representation learning on real world data. In this work, we propose to leverage knowledge extracted from a diversified set of supervised tasks to learn a common disentangled representation. Assuming each supervised task only depends on an unknown subset of the factors of variation, we disentangle the feature space of a supervised multi-task model, with features activating sparsely across different tasks and information being shared as appropriate. Importantly, we never directly observe the factors of variations but establish that access to multiple tasks is sufficient for identifiability under sufficiency and minimality assumptions.
We validate our approach on six real world distribution shift benchmarks, and different data modalities (images, text), demonstrating how disentangled representations can be transferred to real settings. Marco Fumero, Florian Wenzel, Luca Zancato, Alessandro Achille, Emanuele Rodolà, Stefano Soatto, Bernhard Schölkopf, Francesco Locatello |
NeurIPS | 6 |
| 2023 | Gacs-Korner Common Information Variational AutoencoderabstractWe propose a notion of common information that allows one to quantify and separate the information that is shared between two random variables from the information that is unique to each. Our notion of common information is defined by an optimization problem over a family of functions and recovers the G\'acs-K\"orner common information as a special case. Importantly, our notion can be approximated empirically using samples from the underlying data distribution. We then provide a method to partition and quantify the common and unique information using a simple modification of a traditional variational auto-encoder. Empirically, we demonstrate that our formulation allows us to learn semantically meaningful common and unique factors of variation even on high-dimensional data such as images and videos. Moreover, on datasets where ground-truth latent factors are known, we show that we can accurately quantify the common information between the random variables. Michael Kleinman, Alessandro Achille, Stefano Soatto, Jonathan C. Kao |
NeurIPS | 3 |
| 2023 | Harnessing Unrecognizable Faces for Improving Face RecognitionabstractThe common implementation of face recognition systems as a cascade of a detection stage and a recognition or verification stage can cause problems beyond failures of the detector. When the detector succeeds, it can detect faces that cannot be recognized, no matter how capable the recognition system is. Recognizability, a latent variable, should therefore be factored into the design and implementation of face recognition systems. We propose a measure of recognizability of a face image that leverages a key empirical observation: An embedding of face images, implemented by a deep neural network trained using mostly recognizable identities, induces a partition of the hypersphere whereby unrecognizable identities cluster together. This occurs regardless of the phenomenon that causes a face to be unrecognizable, be it optical or motion blur, partial occlusion, spatial quantization, or poor illumination. Therefore, we use the distance from such an "unrecognizable identity" as a measure of recognizability, and incorporate it into the design of the overall system. We show that accounting for recognizability reduces the error rate of single-image face recognition by 58% at FAR=1e-5 on the IJB-C Covariate Verification benchmark, and reduces the verification error rate by 24% at FAR=1e-5 in set-based recognition on the IJB-C benchmark. Siqi Deng, Yuanjun Xiong, Wei Xia 0009, Stefano Soatto |
WACV | 5 |
| 2022 | Stereoscopic Universal Perturbations across Different Architectures and DatasetsabstractWe study the effect of adversarial perturbations of images on deep stereo matching networks for the disparity estimation task. We present a method to craft a single set of perturbations that, when added to any stereo image pair in a dataset, can fool a stereo network to significantly alter the perceived scene geometry. Our perturbation images are “universal” in that they not only corrupt estimates of the network on the dataset they are optimized for, but also generalize to different architectures trained on different datasets. We evaluate our approach on multiple benchmark datasets where our perturbations can increase the D1-error (akin to fooling rate) of state-of-the-art stereo networks from 1% to as much as 87%. We investigate the effect of perturbations on the estimated scene geometry and identify object classes that are most vulnerable. Our analysis on the activations of registered points between left and right images led us to find architectural components that can increase robustness against adversaries. By simply designing networks with such components, one can reduce the effect of adversaries by up to 60.5%, which rivals the robustness of networks finetuned with costly adversarial data augmentation. Our design principle also improves their robustness against common image corruptions by an average of 70%. Zachary Berger, Parth Agrawal, Tian Yu Liu, Stefano Soatto, Alex Wong 0001 |
CVPR | 4 |
| 2022 | MeMOT: Multi-Object Tracking with MemoryabstractWe propose an online tracking algorithm that performs the object detection and data association under a common framework, capable of linking objects after a long time span. This is realized by preserving a large spatio-temporal memory to store the identity embeddings of the tracked objects, and by adaptively referencing and aggregating useful information from the memory as needed. Our model, called MeMOT, consists of three main modules that are all Transformer-based: 1) Hypothesis Generation that produce object proposals in the current video frame; 2) Memory Encoding that extracts the core information from the memory for each tracked object; and 3) Memory Decoding that solves the object detection and data association tasks simultaneously for multi-object tracking. When evaluated on widely adopted MOT benchmark datasets, MeMOT observes very competitive performance. Jiarui Cai, Yuanjun Xiong, Wei Xia 0009, Zhuowen Tu, Stefano Soatto |
CVPR | 7 |
| 2022 | Mixed Differential Privacy in Computer VisionabstractWe introduce AdaMix, an adaptive differentially private algorithm for training deep neural network classifiers using both private and public image data. While pre-training language models on large public datasets has enabled strong differential privacy (DP) guarantees with minor loss of accuracy, a similar practice yields punishing trade-offs in vision tasks. A few-shot or even zero-shot learning baseline that ignores private data can outperform fine-tuning on a large private dataset. AdaMix incorporates few-shot training, or cross-modal zero-shot learning, on public data prior to private fine-tuning, to improve the trade-off. AdaMix reduces the error increase from the non-private upper bound from the 167–311% of the baseline, on average across 6 datasets, to 68-92% depending on the desired privacy level selected by the user. AdaMix tackles the trade-off arising in visual classification, whereby the most privacy sensitive data, corresponding to isolated points in representation space, are also critical for high classification accuracy. In addition, AdaMix comes with strong theoretical privacy guarantees and convergence analysis. Aditya Golatkar, Alessandro Achille, Yu-Xiang Wang 0003, Aaron Roth 0001, Michael Kearns, Stefano Soatto |
CVPR | 6 |
| 2022 | Task Adaptive Parameter Sharing for Multi-Task LearningabstractAdapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models for each task is performant, but incurs a substantial memory cost. To efficiently learn multiple down-stream tasks we introduce Task Adaptive Parameter Sharing (TAPS), a simple method for tuning a base model to a new task by adaptively modifying a small, task-specific subset of layers. This enables multi-task learning while minimizing the resources used and avoids catastrophic forgetting and competition between tasks. TAPS solves a joint optimization problem which determines both the layers that are shared with the base model and the value of the task-specific weights. Further, a sparsity penalty on the number of active layers promotes weight sharing with the base model. Compared to other methods, TAPS retains a high accuracy on the target tasks while still introducing only a small number of task-specific parameters. Moreover, TAPS is agnostic to the particular architecture used and requires only minor changes to the training scheme. We evaluate our method on a suite of fine-tuning tasks and architectures (ResNet, DenseNet, ViT) and show that it achieves state-of-the-art performance while being simple to implement. Matthew Wallingford, Alessandro Achille, Avinash Ravichandran, Charless C. Fowlkes, Rahul Bhotika, Stefano Soatto |
CVPR | 7 |
| 2022 | Omni-DETR: Omni-Supervised Object Detection with TransformersabstractWe consider the problem of omni-supervised object detection, which can use unlabeled, fully labeled and weakly labeled annotations, such as image tags, counts, points, etc., for object detection. This is enabled by a unified architecture, Omni-DETR, based on the recent progress on student-teacher framework and end-to-end transformer based object detection. Under this unified architecture, different types of weak labels can be leveraged to generate accurate pseudo labels, by a bipartite matching based filtering mechanism, for the model to learn. In the experiments, Omni-DETR has achieved state-of-the-art results on multiple datasets and settings. And we have found that weak annotations can help to improve detection performance and a mixture of them can achieve a better trade-off between annotation cost and accuracy than the standard complete annotation. These findings could encourage larger object detection datasets with mixture annotations. The code is available at https://github.com/amazon-research/omni-detr. Zhaowei Cai, Hao Yang 0043, Gurumurthy Swaminathan, Nuno Vasconcelos, Bernt Schiele, Stefano Soatto |
CVPR | 7 |
| 2022 | Class-Incremental Learning with Strong Pre-trained ModelsabstractClass-incremental learning (CIL) has been widely stud-ied under the setting of starting from a small number of classes (base classes). Instead, we explore an understud-ied real-world setting of CIL that starts with a strong model pre-trained on a large number of base classes. We hypoth-esize that a strong base model can provide a good repre-sentation for novel classes and incremental learning can be done with small adaptations. We propose a 2-stage training scheme, i) feature augmentation - cloning part of the backbone and fine-tuning it on the novel data, and ii) fusion - combining the base and novel classifiers into a unified classifier. Experiments show that the proposed method sig-nificantly outperforms state-of-the-art CIL methods on the large-scale ImageNet dataset (e.g. + 10% overall accuracy than the best). We also propose and analyze understudied practical CIL scenarios, such as base-novel overlap with distribution shift. Our proposed method is robust and gen-eralizes to all analyzed CIL settings. Tz-Ying Wu, Gurumurthy Swaminathan, Zhizhong Li 0001, Avinash Ravichandran, Nuno Vasconcelos, Rahul Bhotika, Stefano Soatto |
CVPR | 7 |
| 2022 | Not Just Streaks: Towards Ground Truth for Single Image Deraining
Yunhao Ba, Howard Zhang, Ethan Yang, Akira Suzuki 0002, Arnold Pfahnl, Chethan Chinder Chandrappa, Celso de Melo, Suya You, Stefano Soatto, Alex Wong 0001, Achuta Kadambi |
ECCV (7) | 9 |
| 2022 | X-DETR: A Versatile Architecture for Instance-wise Vision-Language Tasks
Zhaowei Cai, Gukyeong Kwon, Avinash Ravichandran, Erhan Bas, Zhuowen Tu, Rahul Bhotika, Stefano Soatto |
ECCV (36) | 7 |
| 2022 | DIVA: Dataset Derivative of a Learning Task
Yonatan Dukler, Alessandro Achille, Giovanni Paolini, Avinash Ravichandran, Marzia Polito, Stefano Soatto |
ICLR | 6 |
| 2022 | Semi-supervised Vision Transformers at ScaleabstractWe study semi-supervised learning (SSL) for vision transformers (ViT), an under-explored topic despite the wide adoption of the ViT architectures to different tasks. To tackle this problem, we use a SSL pipeline, consisting of first un/self-supervised pre-training, followed by supervised fine-tuning, and finally semi-supervised fine-tuning. At the semi-supervised fine-tuning stage, we adopt an exponential moving average (EMA)-Teacher framework instead of the popular FixMatch, since the former is more stable and delivers higher accuracy for semi-supervised vision transformers. In addition, we propose a probabilistic pseudo mixup mechanism to interpolate unlabeled samples and their pseudo labels for improved regularization, which is important for training ViTs with weak inductive bias. Our proposed method, dubbed Semi-ViT, achieves comparable or better performance than the CNN counterparts in the semi-supervised classification setting. Semi-ViT also enjoys the scalability benefits of ViTs that can be readily scaled up to large-size models with increasing accuracy. For example, Semi-ViT-Huge achieves an impressive 80\% top-1 accuracy on ImageNet using only 1\% labels, which is comparable with Inception-v4 using 100\% ImageNet labels. The code is available at https://github.com/amazon-science/semi-vit. Zhaowei Cai, Avinash Ravichandran, Paolo Favaro, Manchen Wang, Davide Modolo, Rahul Bhotika, Zhuowen Tu, Stefano Soatto |
NeurIPS | 8 |
| 2022 | On Leave-One-Out Conditional Mutual Information For GeneralizationabstractWe derive information theoretic generalization bounds for supervised learning algorithms based on a new measure of leave-one-out conditional mutual information (loo-CMI). In contrast to other CMI bounds, which may be hard to evaluate in practice, our loo-CMI bounds are easier to compute and can be interpreted in connection to other notions such as classical leave-one-out cross-validation, stability of the optimization algorithm, and the geometry of the loss-landscape. It applies both to the output of training algorithms as well as their predictions. We empirically validate the quality of the bound by evaluating its predicted generalization gap in scenarios for deep learning. In particular, our bounds are non-vacuous on image-classification tasks. Mohamad Rida Rammal, Alessandro Achille, Aditya Golatkar, Suhas N. Diggavi, Stefano Soatto |
NeurIPS | 5 |
| 2022 | Stochastic batch size for adaptive regularization in deep network optimization
Kensuke Nakamura 0001, Stefano Soatto, Byung-Woo Hong |
Pattern Recognit. | 2 |
| 2021 | Dynamically Grown Generative Adversarial NetworksabstractRecent work introduced progressive network growing as a promising way to ease the training for large GANs, but the model design and architecture-growing strategy still remain under-explored and needs manual design for different image data. In this paper, we propose a method to dynamically grow a GAN during training, optimizing the network architecture and its parameters together with automation. The method embeds architecture search techniques as an interleaving step with gradient-based training to periodically seek the optimal architecture-growing strategy for the generator and discriminator. It enjoys the benefits of both eased training because of progressive growing and improved performance because of broader architecture design space. Experimental results demonstrate new state-of-the-art of image generation. Observations in the search procedure also provide constructive insights into the GAN model design such as generator-discriminator balance and convolutional layer choices. Lanlan Liu, Jia Deng 0001, Stefano Soatto |
AAAI | 4 |
| 2021 | Stereopagnosia: Fooling Stereo Networks with Adversarial PerturbationsabstractWe study the effect of adversarial perturbations of images on the estimates of disparity by deep learning models trained for stereo. We show that imperceptible additive perturbations can significantly alter the disparity map, and correspondingly the perceived geometry of the scene. These perturbations not only affect the specific model they are crafted for, but transfer to models with different architecture, trained with different loss functions. We show that, when used for adversarial data augmentation, our perturbations result in trained models that are more robust, without sacrificing overall accuracy of the model. This is unlike what has been observed in image classification, where adding the perturbed images to the training set makes the model less vulnerable to adversarial perturbations, but to the detriment of overall accuracy. We test our method using the most recent stereo networks and evaluate their performance on public benchmark datasets. Alex Wong 0001, Mukund Mundhra, Stefano Soatto |
AAAI | 3 |
| 2021 | Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model UpdatesabstractYuqing Xie, Yi-An Lai, Yuanjun Xiong, Yi Zhang, Stefano Soatto. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yuqing Xie 0001, Yi-An Lai, Yuanjun Xiong, Stefano Soatto |
ACL/IJCNLP (1) | 5 |
| 2021 | DyStaB: Unsupervised Object Segmentation via Dynamic-Static BootstrappingabstractWe describe an unsupervised method to detect and segment portions of images of live scenes that, at some point in time, are seen moving as a coherent whole, which we refer to as objects. Our method first partitions the motion field by minimizing the mutual information between segments. Then, it uses the segments to learn object models that can be used for detection in a static image. Static and dynamic models are represented by deep neural networks trained jointly in a bootstrapping strategy, which enables extrapolation to previously unseen objects. While the training process requires motion, the resulting object segmentation network can be used on either static images or videos at inference time. As the volume of seen videos grows, more and more objects are seen moving, priming their detection, which then serves as a regularizer for new objects, turning our method into unsupervised continual learning to segment objects. Our models are compared to the state of the art in both video object segmentation and salient object detection. In the six benchmark datasets tested, our models compare favorably even to those using pixel-level supervision, despite requiring no manual annotation. Yanchao Yang 0001, Brian Lai, Stefano Soatto |
CVPR | 3 |
| 2021 | LQF: Linear Quadratic Fine-TuningabstractClassifiers that are linear in their parameters, and trained by optimizing a convex loss function, have predictable behavior with respect to changes in the training data, initial conditions, and optimization. Such desirable properties are absent in deep neural networks (DNNs), typically trained by non-linear fine-tuning of a pre-trained model. Previous attempts to linearize DNNs have led to interesting theoretical insights, but have not impacted the practice due to the substantial performance gap compared to standard non-linear optimization. We present the first method for linearizing a pre-trained model that achieves comparable performance to non-linear fine-tuning on most of real-world image classification tasks tested, thus enjoying the interpretability of linear models without incurring punishing losses in performance. LQF consists of simple modifications to the architecture, loss function and optimization typically used for classification: Leaky-ReLU instead of ReLU, mean squared loss instead of cross-entropy, and pre-conditioning using Kronecker factorization. None of these changes in isolation is sufficient to approach the performance of non-linear fine-tuning. When used in combination, they allow us to reach comparable performance, and even superior in the low-data regime, while enjoying the simplicity, robustness and interpretability of linear-quadratic optimization. Alessandro Achille, Aditya Golatkar, Avinash Ravichandran, Marzia Polito, Stefano Soatto |
CVPR | 5 |
| 2021 | Learning Semantic-Aware Dynamics for Video PredictionabstractWe propose an architecture and training scheme to predict video frames by explicitly modeling dis-occlusions and capturing the evolution of semantically consistent regions in the video. The scene layout (semantic map) and motion (optical flow) are decomposed into layers, which are predicted and fused with their context to generate future layouts and motions. The appearance of the scene is warped from past frames using the predicted motion in co-visible regions; dis-occluded regions are synthesized with content-aware inpainting utilizing the predicted scene layout. The result is a predictive model that explicitly represents objects and learns their class-specific motion, which we evaluate on video prediction benchmarks. Xinzhu Bei, Yanchao Yang 0001, Stefano Soatto |
CVPR | 3 |
| 2021 | Exponential Moving Average Normalization for Self-Supervised and Semi-Supervised LearningabstractWe present a plug-in replacement for batch normalization (BN) called exponential moving average normalization (EMAN), which improves the performance of existing student-teacher based self- and semi-supervised learning techniques. Unlike the standard BN, where the statistics are computed within each batch, EMAN, used in the teacher, updates its statistics by exponential moving average from the BN statistics of the student. This design reduces the intrinsic cross-sample dependency of BN and enhances the generalization of the teacher. EMAN improves strong baselines for self-supervised learning by 4-6/1-2 points and semi-supervised learning by about 7/2 points, when 1%/10% supervised labels are available on ImageNet. These improvements are consistent across methods, network architectures, training duration, and datasets, demonstrating the general effectiveness of this technique. The code will be made available online. Zhaowei Cai, Avinash Ravichandran, Subhransu Maji, Charless C. Fowlkes, Zhuowen Tu, Stefano Soatto |
CVPR | 6 |
| 2021 | Compatibility-Aware Heterogeneous Visual SearchabstractWe tackle the problem of visual search under resource constraints. Existing systems use the same embedding model to compute representations (embeddings) for the query and gallery images. Such systems inherently face a hard accuracy-efficiency trade-off: the embedding model needs to be large enough to ensure high accuracy, yet small enough to enable query-embedding computation on resource-constrained platforms. This trade-off could be mitigated if gallery embeddings are generated from a large model and query embeddings are extracted using a compact model. The key to building such a system is to ensure representation compatibility between the query and gallery models. In this paper, we address two forms of compatibility: One enforced by modifying the parameters of each model that computes the embeddings. The other by modifying the architectures that compute the embeddings, leading to compatibility-aware neural architecture search (Cmp-NAS). We test Cmp-NAS on challenging retrieval tasks for fashion images (DeepFashion2), and face images (IJB-C). Compared to ordinary (homogeneous) visual search using the largest embedding model (paragon), Cmp-NAS achieves 80-fold and 23-fold cost reduction while maintaining accuracy within 0.3% and 1.6% of the paragon on DeepFashion2 and IJB-C respectively. Rahul Duggal, Shuo Yang 0003, Yuanjun Xiong, Wei Xia 0009, Zhuowen Tu, Stefano Soatto |
CVPR | 7 |
| 2021 | Mixed-Privacy Forgetting in Deep NetworksabstractWe show that the influence of a subset of the training samples can be removed – or "forgotten" – from the weights of a network trained on large-scale image classification tasks, and we provide strong computable bounds on the amount of remaining information after forgetting. Inspired by real-world applications of forgetting techniques, we introduce a novel notion of forgetting in mixed-privacy setting, where we know that a "core" subset of the training samples does not need to be forgotten. While this variation of the problem is conceptually simple, we show that working in this setting significantly improves the accuracy and guarantees of forgetting methods applied to vision classification tasks. Moreover, our method allows efficient removal of all information contained in non-core data by simply setting to zero a subset of the weights with minimal loss in performance. We achieve these results by replacing a standard deep network with a suitable linear approximation. With opportune changes to the network architecture and training procedure, we show that such linear approximation achieves comparable performance to the original network and that the forgetting problem becomes quadratic and can be solved efficiently even for large models. Unlike previous forgetting methods on deep networks, ours can achieve close to the state-of-the-art accuracy on large scale vision tasks. In particular, we show that our method allows forgetting without having to trade off the model accuracy. Aditya Golatkar, Alessandro Achille, Avinash Ravichandran, Marzia Polito, Stefano Soatto |
CVPR | 5 |
| 2021 | Positive-Congruent Training: Towards Regression-Free Model UpdatesabstractReducing inconsistencies in the behavior of different versions of an AI system can be as important in practice as reducing its overall error. In image classification, sample-wise inconsistencies appear as "negative flips": A new model incorrectly predicts the output for a test sample that was correctly classified by the old (reference) model. Positive-congruent (PC) training aims at reducing error rate while at the same time reducing negative flips, thus maximizing congruency with the reference model only on positive predictions, unlike model distillation. We propose a simple approach for PC training, Focal Distillation, which enforces congruence with the reference model by giving more weights to samples that were correctly classified. We also found that, if the reference model itself can be chosen as an ensemble of multiple deep neural networks, negative flips can be further reduced without affecting the new model’s accuracy. Sijie Yan, Yuanjun Xiong, Kaustav Kundu, Shuo Yang 0003, Siqi Deng, Wei Xia 0009, Stefano Soatto |
CVPR | 8 |
| 2021 | Unsupervised Depth Completion with Calibrated Backprojection LayersabstractWe propose a deep neural network architecture to infer dense depth from an image and a sparse point cloud. It is trained using a video stream and corresponding synchronized sparse point cloud, as obtained from a LIDAR or other range sensor, along with the intrinsic calibration parameters of the camera. At inference time, the calibration of the camera, which can be different than the one used for training, is fed as an input to the network along with the sparse point cloud and a single image. A Calibrated Backprojection Layer backprojects each pixel in the image to three-dimensional space using the calibration matrix and a depth feature descriptor. The resulting 3D positional encoding is concatenated with the image descriptor and the previous layer output to yield the input to the next layer of the encoder. A decoder, exploiting skip-connections, produces a dense depth map. The resulting Calibrated Backprojection Network, or KBNet, is trained without supervision by minimizing the photometric reprojection error. KBNet imputes missing depth value based on the training set, rather than on generic regularization. We test KBNet on public depth completion benchmarks, where it outperforms the state of the art by 30% indoor and 8% outdoor when the same camera is used for training and testing. When the test camera is different, the improvement reaches 62%. Alex Wong 0001, Stefano Soatto |
ICCV | 2 |
| 2021 | Visual Relationship Detection Using Part-and-Sum Transformers with Composite QueriesabstractComputer vision applications such as visual relationship detection and human object interaction can be formulated as a composite (structured) set detection problem in which both the parts (subject, object, and predicate) and the sum (triplet as a whole) are to be detected in a hierarchical fashion. In this paper, we present a new approach, denoted Part-and-Sum detection Transformer (PST), to perform end-to-end visual composite set detection. Different from existing Transformers in which queries are at a single level, we simultaneously model the joint part and sum hypotheses/interactions with composite queries and attention modules. We explicitly incorporate sum queries to enable better modeling of the part-and-sum relations that are absent in the standard Transformers. Our approach also uses novel tensor-based part queries and vector-based sum queries, and models their joint interaction. We report experiments on two vision tasks, visual relationship detection and human object interaction and demonstrate that PST achieves state of the art results among single-stage models, while nearly matching the results of custom designed two-stage models. Zhuowen Tu, Haofu Liao, Vijay Mahadevan, Stefano Soatto |
ICCV | 6 |
| 2021 | ARCH++: Animation-Ready Clothed Human Reconstruction RevisitedabstractWe present ARCH++, an image-based method to reconstruct 3D avatars with arbitrary clothing styles. Our reconstructed avatars are animation-ready and highly realistic, in both the visible regions from input views and the unseen regions. While prior work shows great promise of reconstructing animatable clothed humans with various topologies, we observe that there exist fundamental limitations resulting in sub-optimal reconstruction quality. In this paper, we revisit the major steps of image-based avatar reconstruction and address the limitations with ARCH++. First, we introduce an end-to-end point based geometry encoder to better describe the semantics of the underlying 3D human body, in replacement of previous hand-crafted features. Second, in order to address the occupancy ambiguity caused by topological changes of clothed humans in the canonical pose, we propose a co-supervising framework with cross-space consistency to jointly estimate the occupancy in both the posed and canonical spaces. Last, we use image-to-image translation networks to further refine detailed geometry and texture on the reconstructed surface, which improves the fidelity and consistency across arbitrary viewpoints. In the experiments, we demonstrate improvements over the state of the art on both public benchmarks and user studies in reconstruction quality and realism. Tong He 0002, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, Tony Tung |
ICCV | 4 |
| 2021 | Learning Hierarchical Graph Neural Networks for Image ClusteringabstractWe propose a hierarchical graph neural network (GNN) model that learns how to cluster a set of images into an unknown number of identities using a training set of images annotated with labels belonging to a disjoint set of identities. Our hierarchical GNN uses a novel approach to merge connected components predicted at each level of the hierarchy to form a new graph at the next level. Unlike fully unsupervised hierarchical clustering, the choice of grouping and complexity criteria stems naturally from supervision in the training set. The resulting method, Hi-LANDER, achieves an average of 49% improvement in F-score and 7% increase in Normalized Mutual Information (NMI) relative to current GNN-based clustering algorithms. Additionally, state-of-the-art GNN-based methods rely on separate models to predict linkage probabilities and node densities as intermediate steps of the clustering process. In contrast, our unified framework achieves a three-fold decrease in computational cost. Our training and inference code are released1. Yifan Xing, Tong He 0002, Tianjun Xiao, Yuanjun Xiong, Wei Xia 0009, David P. Wipf, Zheng Zhang 0001, Stefano Soatto |
ICCV | 9 |
| 2021 | Estimating informativeness of samples with Smooth Unique Information
Hrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ICLR | 7 |
| 2021 | Structured Prediction as Translation between Augmented Natural Languages
Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, Stefano Soatto |
ICLR | 9 |
| 2021 | Learned Uncertainty Calibration for Visual Inertial Localization
Stephanie Tsuei, Stefano Soatto, Paulo Tabuada, Mark B. Milam |
ICRA | 2 |
| 2021 | Uniform Sampling over Episode DifficultyabstractEpisodic training is a core ingredient of few-shot learning to train models on tasks with limited labelled data. Despite its success, episodic training remains largely understudied, prompting us to ask the question: what is the best way to sample episodes? In this paper, we first propose a method to approximate episode sampling distributions based on their difficulty. Building on this method, we perform an extensive analysis and find that sampling uniformly over episode difficulty outperforms other sampling schemes, including curriculum and easy-/hard-mining. As the proposed sampling method is algorithm agnostic, we can leverage these insights to improve few-shot learning accuracies across many episodic training algorithms. We demonstrate the efficacy of our method across popular few-shot learning datasets, algorithms, network architectures, and protocols. Sébastien M. R. Arnold, Guneet S. Dhillon 0001, Avinash Ravichandran, Stefano Soatto |
NeurIPS | 4 |
| 2021 | Long Short-Term Transformer for Online Action DetectionabstractWe present Long Short-term TRansformer (LSTR), a temporal modeling algorithm for online action detection, which employs a long- and short-term memory mechanism to model prolonged sequence data. It consists of an LSTR encoder that dynamically leverages coarse-scale historical information from an extended temporal window (e.g., 2048 frames spanning of up to 8 minutes), together with an LSTR decoder that focuses on a short time window (e.g., 32 frames spanning 8 seconds) to model the fine-scale characteristics of the data. Compared to prior work, LSTR provides an effective and efficient method to model long videos with fewer heuristics, which is validated by extensive empirical analysis. LSTR achieves state-of-the-art performance on three standard online action detection benchmarks, THUMOS'14, TVSeries, and HACS Segment. Code has been made available at: https://xumingze0308.github.io/projects/lstr. Yuanjun Xiong, Hao Chen 0024, Xinyu Li 0003, Wei Xia 0009, Zhuowen Tu, Stefano Soatto |
NeurIPS | 7 |
| 2021 | Block-cyclic stochastic coordinate descent for deep neural networks
Kensuke Nakamura 0001, Stefano Soatto, Byung-Woo Hong |
Neural Networks | 2 |
| 2020 | Zero Shot Learning with the Isoperimetric LossabstractWe introduce the isoperimetric loss as a regularization criterion for learning the map from a visual representation to a semantic embedding, to be used to transfer knowledge to unknown classes in a zero-shot learning setting. We use a pre-trained deep neural network model as a visual representation of image data, a Word2Vec embedding of class labels, and linear maps between the visual and semantic embedding spaces. However, the spaces themselves are not linear, and we postulate the sample embedding to be populated by noisy samples near otherwise smooth manifolds. We exploit the graph structure defined by the sample points to regularize the estimates of the manifolds by inferring the graph connectivity using a generalization of the isoperimetric inequalities from Riemannian geometry to graphs. Surprisingly, this regularization alone, paired with the simplest baseline model, outperforms the state-of-the-art among fully automated methods in zero-shot learning benchmarks such as AwA and CUB. This improvement is achieved solely by learning the structure of the underlying spaces by imposing regularity. Shay Deutsch, Andrea L. Bertozzi, Stefano Soatto |
AAAI | 3 |
| 2020 | Spatial Class Distribution Shift in Unsupervised Domain Adaptation: Local Alignment Comes to Rescue
Safa Cicek, Ning Xu 0007, Hailin Jin, Stefano Soatto |
ACCV (3) | 5 |
| 2020 | DeepVoxels++: Enhancing the Fidelity of Novel View Synthesis from 3D Voxel Embeddings
Tong He 0002, John P. Collomosse, Hailin Jin, Stefano Soatto |
ACCV (1) | 4 |
| 2020 | FDA: Fourier Domain Adaptation for Semantic SegmentationabstractWe describe a simple method for unsupervised domain adaptation, whereby the discrepancy between the source and target distributions is reduced by swapping the low-frequency spectrum of one with the other. We illustrate the method in semantic segmentation, where densely annotated images are aplenty in one domain (synthetic data), but difficult to obtain in another (real images). Current state-of-the-art methods are complex, some requiring adversarial optimization to render the backbone of a neural network invariant to the discrete domain selection variable. Our method does not require any training to perform the domain alignment, just a simple Fourier Transform and its inverse. Despite its simplicity, it achieves state-of-the-art performance in the current benchmarks, when integrated into a relatively standard semantic segmentation model. Our results indicate that even simple procedures can discount nuisance variability in the data that more sophisticated methods struggle to learn away. Yanchao Yang 0001, Stefano Soatto |
CVPR | 2 |
| 2020 | Eternal Sunshine of the Spotless Net: Selective Forgetting in Deep NetworksabstractWe explore the problem of selectively forgetting a particular subset of the data used for training a deep neural network. While the effects of the data to be forgotten can be hidden from the output of the network, insights may still be gleaned by probing deep into its weights. We propose a method for "scrubbing" the weights clean of information about a particular set of training data. The method does not require retraining from scratch, nor access to the data originally used for training. Instead, the weights are modified so that any probing function of the weights is indistinguishable from the same function applied to the weights of a network trained without the data to be forgotten. This condition is a generalized and weaker form of Differential Privacy. Exploiting ideas related to the stability of stochastic gradient descent, we introduce an upper-bound on the amount of information remaining in the weights, which can be estimated efficiently even for deep neural networks. Aditya Golatkar, Alessandro Achille, Stefano Soatto |
CVPR | 3 |
| 2020 | Towards Backward-Compatible Representation LearningabstractWe propose a way to learn visual features that are compatible with previously computed ones even when they have different dimensions and are learned via different neural network architectures and loss functions. Compatible means that, if such features are used to compare images, then ``new'' features can be compared directly to ``old'' features, so they can be used interchangeably. This enables visual search systems to bypass computing new features for all previously seen images when updating the embedding models, a process known as backfilling. Backward compatibility is critical to quickly deploy new embedding models that leverage ever-growing large-scale training datasets and improvements in deep learning architectures and training methods. We propose a framework to train embedding models, called backward-compatible training (BCT), as a first step towards backward compatible representation learning. In experiments on learning embeddings for face recognition, models trained with BCT successfully achieve backward compatibility without sacrificing accuracy, thus enabling backfill-free model updates of visual embeddings. Yantao Shen 0002, Yuanjun Xiong, Wei Xia 0009, Stefano Soatto |
CVPR | 4 |
| 2020 | Learning to Manipulate Individual Objects in an ImageabstractWe describe a method to train a generative model with latent factors that are (approximately) independent and localized. This means that perturbing the latent variables affects only local regions of the synthesized image, corresponding to objects. Unlike other unsupervised generative models, ours enables object-centric manipulation, without requiring object-level annotations, or any form of annotation for that matter. The key to our method is the combination of spatial disentanglement, enforced by a Contextual Information Separation loss, and perceptual cycle-consistency, enforced by a loss that penalizes changes in the image partition in response to perturbations of the latent factors. We test our method's ability to allow independent control of spatial and semantic factors of variability on existing datasets and also introduce two new ones that highlight the limitations of current methods. Yanchao Yang 0001, Stefano Soatto |
CVPR | 3 |
| 2020 | Phase Consistent Ecological Domain AdaptationabstractWe introduce two criteria to regularize the optimization involved in learning a classifier in a domain where no annotated data are available, leveraging annotated data in a different domain, a problem known as unsupervised domain adaptation. We focus on the task of semantic segmentation, where annotated synthetic data are aplenty, but annotating real data is laborious. The first criterion, inspired by visual psychophysics, is that the map between the two image domains be phase-preserving. This restricts the set of possible learned maps, while enabling enough flexibility to transfer semantic information. The second criterion aims to leverage ecological statistics, or regularities in the scene which are manifest in any image of it, regardless of the characteristics of the illuminant or the imaging sensor. It is implemented using a deep neural network that scores the likelihood of each possible segmentation given a single un-annotated image. Incorporating these two priors in a standard domain adaptation framework improves performance across the board in the most common unsupervised domain adaptation benchmarks for semantic segmentation. Yanchao Yang 0001, Dong Lao, Ganesh Sundaramoorthi, Stefano Soatto |
CVPR | 4 |
| 2020 | Forgetting Outside the Box: Scrubbing Deep Networks of Information Accessible from Input-Output Observations
Aditya Golatkar, Alessandro Achille, Stefano Soatto |
ECCV (29) | 3 |
| 2020 | Incremental Few-Shot Meta-learning via Indirect Discriminant Alignment
Orchid Majumder, Alessandro Achille, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ECCV (7) | 6 |
| 2020 | A Baseline for Few-Shot Image Classification
Guneet S. Dhillon 0001, Pratik Chaudhari, Avinash Ravichandran, Stefano Soatto |
ICLR | 4 |
| 2020 | Meta-Q-Learning
Rasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. Smola |
ICLR | 3 |
| 2020 | Rethinking the Hyperparameters for Fine-tuning
Pratik Chaudhari, Hao Yang 0043, Michael Lam, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ICLR | 7 |
| 2020 | Risk-Averse MPC via Visual-Inertial Input and Recurrent Networks for Online Collision AvoidanceabstractIn this paper, we propose an online path planning architecture that extends the model predictive control (MPC) formulation to consider future location uncertainties for safer navigation through cluttered environments. Our algorithm combines an object detection pipeline with a recurrent neural network (RNN) which infers the covariance of state estimates through each step of our MPC's finite time horizon. The RNN model is trained on a dataset that comprises of robot and landmark poses generated from camera images and inertial measurement unit (IMU) readings via a state-of-the-art visualinertial odometry framework. To detect and extract object locations for avoidance, we use a custom-trained convolutional neural network model in conjunction with a feature extractor to retrieve 3D centroid and radii boundaries of nearby obstacles. The robustness of our methods is validated on complex quadruped robot dynamics and can be generally applied to most robotic platforms, demonstrating autonomous behaviors that can plan fast and collision-free paths towards a goal point. Alexander Schperberg, Kenny Chen, Stephanie Tsuei, Michael Jewett, Joshua Hooks, Stefano Soatto, Ankur Mehta, Dennis W. Hong |
IROS | 6 |
| 2020 | Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human ReconstructionabstractWe propose Geo-PIFu, a method to recover a 3D mesh from a monocular color image of a clothed person. Our method is based on a deep implicit function-based representation to learn latent voxel features using a structure-aware 3D U-Net, to constrain the model in two ways: first, to resolve feature ambiguities in query point encoding, second, to serve as a coarse human shape proxy to regularize the high-resolution mesh and encourage global shape regularity. We show that, by both encoding query points and constraining global shape using latent voxel features, the reconstruction we obtain for clothed human meshes exhibits less shape distortion and improved surface details compared to competing methods. We evaluate Geo-PIFu on a recent human mesh public dataset that is 10x larger than the private commercial dataset used in PIFu and previous derivative work. On average, we exceed the state of the art by 42.7% reduction in Chamfer and Point-to-Surface Distances, and 19.4% reduction in normal estimation errors. Tong He 0002, John P. Collomosse, Hailin Jin, Stefano Soatto |
NeurIPS | 4 |
| 2020 | Targeted Adversarial Perturbations for Monocular Depth PredictionabstractWe study the effect of adversarial perturbations on the task of monocular depth prediction. Specifically, we explore the ability of small, imperceptible additive perturbations to selectively alter the perceived geometry of the scene. We show that such perturbations can not only globally re-scale the predicted distances from the camera, but also alter the prediction to match a different target scene. We also show that, when given semantic or instance information, perturbations can fool the network to alter the depth of specific categories or instances in the scene, and even remove them while preserving the rest of the scene. To understand the effect of targeted perturbations, we conduct experiments on state-of-the-art monocular depth prediction methods. Our experiments reveal vulnerabilities in monocular depth prediction networks, and shed light on the biases and context learned by them. Alex Wong 0001, Safa Cicek, Stefano Soatto |
NeurIPS | 3 |
| 2020 | Predicting Training Time Without TrainingabstractWe tackle the problem of predicting the number of optimization steps that a pre-trained deep network needs to converge to a given value of the loss function. To do so, we leverage the fact that the training dynamics of a deep network during fine-tuning are well approximated by those of a linearized model. This allows us to approximate the training loss and accuracy at any point during training by solving a low-dimensional Stochastic Differential Equation (SDE) in function space. Using this result, we are able to predict the time it takes for Stochastic Gradient Descent (SGD) to fine-tune a model to a given loss without having to perform any training. In our experiments, we are able to predict training time of a ResNet within a 20\% error margin on a variety of datasets and hyper-parameters, at a 30 to 45-fold reduction in cost compared to actual training. We also discuss how to further reduce the computational and memory cost of our method, and in particular we show that by exploiting the spectral properties of the gradients' matrix it is possible to predict training time on a large dataset while processing only a subset of the samples. Luca Zancato, Alessandro Achille, Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
NeurIPS | 5 |
| 2020 | Adaptive Regularization of Some Inverse Problems in Image AnalysisabstractWe present an adaptive regularization scheme for optimizing composite energy functionals arising in image analysis problems. The scheme automatically trades off data fidelity and regularization depending on the current data fit during the iterative optimization, so that regularization is strongest initially, and wanes as data fidelity improves, with the weight of the regularizer being minimized at convergence. We also introduce a Huber loss function in both data fidelity and regularization terms, and present an efficient convex optimization algorithm based on the alternating direction method of multipliers (ADMM) using the equivalent relation between the Huber function and the proximal operator of the one-norm. We illustrate and validate our adaptive Huber-Huber model on synthetic and real images in segmentation, motion estimation, and denoising problems. Byung-Woo Hong, Jakeoung Koo, Martin Burger 0001, Stefano Soatto |
IEEE Trans. Image Process. | 4 |
| 2019 | Mono3D++: Monocular 3D Vehicle Detection with Two-Scale 3D Hypotheses and Task PriorsabstractWe present a method to infer 3D pose and shape of vehicles from a single image. To tackle this ill-posed problem, we optimize two-scale projection consistency between the generated 3D hypotheses and their 2D pseudo-measurements. Specifically, we use a morphable wireframe model to generate a fine-scaled representation of vehicle shape and pose. To reduce its sensitivity to 2D landmarks, we jointly model the 3D bounding box as a coarse representation which improves robustness. We also integrate three task priors, including unsupervised monocular depth, a ground plane constraint as well as vehicle shape priors, with forward projection errors into an overall energy function. Tong He 0002, Stefano Soatto |
AAAI | 2 |
| 2019 | Unsupervised Moving Object Detection via Contextual Information SeparationabstractWe propose an adversarial contextual model for detecting moving objects in images. A deep neural network is trained to predict the optical flow in a region using information from everywhere else but that region (context), while another network attempts to make such context as uninformative as possible. The result is a model where hypotheses naturally compete with no need for explicit regularization or hyper-parameter tuning. Although our method requires no supervision whatsoever, it outperforms several methods that are pre-trained on large annotated datasets. Our model can be thought of as a generalization of classical variational generative region-based segmentation, but in a way that avoids explicit regularization or solution of partial differential equations at run-time. Yanchao Yang 0001, Antonio Loquercio, Davide Scaramuzza 0001, Stefano Soatto |
CVPR | 4 |
| 2019 | Dense Depth Posterior (DDP) From Single Image and Sparse RangeabstractWe present a deep learning system to infer the posterior distribution of a dense depth map associated with an image, by exploiting sparse range measurements, for instance from a lidar. While the lidar may provide a depth value for a small percentage of the pixels, we exploit regularities reflected in the training set to complete the map so as to have a probability over depth for each pixel in the image. We exploit a Conditional Prior Network, that allows associating a probability to each depth value given an image, and combine it with a likelihood term that uses the sparse measurements. Optionally we can also exploit the availability of stereo during training, but in any case only require a single image and a sparse point cloud at run-time. We test our approach on both unsupervised and supervised depth completion using the KITTI benchmark, and improve the state-of-the-art in both. Yanchao Yang 0001, Alex Wong 0001, Stefano Soatto |
CVPR | 3 |
| 2019 | GeoNet: Deep Geodesic Networks for Point Cloud AnalysisabstractSurface-based geodesic topology provides strong cues for object semantic analysis and geometric modeling. However, such connectivity information is lost in point clouds. Thus we introduce GeoNet, the first deep learning architecture trained to model the intrinsic structure of surfaces represented as point clouds. To demonstrate the applicability of learned geodesic-aware representations, we propose fusion schemes which use GeoNet in conjunction with other baseline or backbone networks, such as PU-Net and PointNet++, for down-stream point cloud analysis. Our method improves the state-of-the-art on multiple representative tasks that can benefit from understandings of the underlying surface topology, including point upsampling, normal estimation, mesh reconstruction and non-rigid shape classification. Tong He 0002, Li Yi 0001, Yuqian Zhou, Chihao Wu 0001, Jue Wang 0001, Stefano Soatto |
CVPR | 7 |
| 2019 | Meta-Learning With Differentiable Convex OptimizationabstractMany meta-learning approaches for few-shot learning rely on simple base learners such as nearest-neighbor classifiers. However, even in the few-shot regime, discriminatively trained linear predictors can offer better generalization. We propose to use these predictors as base learners to learn representations for few-shot learning and show they offer better tradeoffs between feature size and performance across a range of few-shot recognition benchmarks. Our objective is to learn feature embeddings that generalize well under a linear classification rule for novel categories. To efficiently solve the objective, we exploit two properties of linear classifiers: implicit differentiation of the optimality conditions of the convex problem and the dual formulation of the optimization problem. This allows us to use high-dimensional embeddings with improved generalization at a modest increase in computational overhead. Our approach, named MetaOptNet, achieves state-of-the-art performance on miniImageNet, tieredImageNet, CIFAR-FS, and FC100 few-shot learning benchmarks. Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, Stefano Soatto |
CVPR | 4 |
| 2019 | Bilateral Cyclic Constraint and Adaptive Regularization for Unsupervised Monocular Depth PredictionabstractSupervised learning methods to infer (hypothesize) depth of a scene from a single image require costly per-pixel ground-truth. We follow a geometric approach that exploits abundant stereo imagery to learn a model to hypothesize scene structure without direct supervision. Although we train a network with stereo pairs, we only require a single image at test time to hypothesize disparity or depth. We propose a novel objective function that exploits the bilateral cyclic relationship between the left and right disparities and we introduce an adaptive regularization scheme that allows the network to handle both the co-visible and occluded regions in a stereo pair. This process ultimately produces a model to generate hypotheses for the 3-dimensional structure of the scene as viewed in a single image. When used to generate a single (most probable) estimate of depth, our method outperforms state-of-the-art unsupervised monocular depth prediction methods on the KITTI benchmarks. We show that our method generalizes well by applying our models trained on KITTI to the Make3d dataset. Alex Wong 0001, Stefano Soatto |
CVPR | 2 |
| 2019 | Task2Vec: Task Embedding for Meta-LearningabstractWe introduce a method to generate vectorial representations of visual classification tasks which can be used to reason about the nature of those tasks and their relations. Given a dataset with ground-truth labels and a loss function, we process images through a "probe network" and compute an embedding based on estimates of the Fisher information matrix associated with the probe network parameters. This provides a fixed-dimensional embedding of the task that is independent of details such as the number of classes and requires no understanding of the class label semantics. We demonstrate that this embedding is capable of predicting task similarities that match our intuition about semantic and taxonomic relations between different visual tasks. We demonstrate the practical value of this framework for the meta-task of selecting a pre-trained feature extractor for a novel task. We present a simple meta-learning framework for learning a metric on embeddings that is capable of predicting which feature extractors will perform well on which task. Selecting a feature extractor with task embedding yields performance close to the best available feature extractor, with substantially less computational effort than exhaustively training and evaluating all available models. Alessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran, Subhransu Maji, Charless C. Fowlkes, Stefano Soatto, Pietro Perona |
ICCV | 7 |
| 2019 | Unsupervised Domain Adaptation via Regularized Conditional AlignmentabstractWe propose a method for unsupervised domain adaptation that trains a shared embedding to align the joint distributions of inputs (domain) and outputs (classes), making any classifier agnostic to the domain. Joint alignment ensures that not only the marginal distributions of the domains are aligned, but the labels as well. We propose a novel objective function that encourages the class-conditional distributions to have disjoint support in feature space. We further exploit adversarial regularization to improve the performance of the classifier on the domain for which no annotated data is available. Safa Cicek, Stefano Soatto |
ICCV | 2 |
| 2019 | Few-Shot Learning With Embedded Class Models and Shot-Free Meta TrainingabstractWe propose a method for learning embeddings for few-shot learning that is suitable for use with any number of shots (shot-free). Rather than fixing the class prototypes to be the Euclidean average of sample embeddings, we allow them to live in a higher-dimensional space (embedded class models) and learn the prototypes along with the model parameters. The class representation function is defined implicitly, which allows us to deal with a variable number of shots per class with a simple constant-size architecture. The class embedding encompasses metric learning, that facilitates adding new classes without crowding the class representation space. Despite being general and not tuned to the benchmark, our approach achieves state-of-the-art performance on the standard few-shot benchmark datasets. Avinash Ravichandran, Rahul Bhotika, Stefano Soatto |
ICCV | 3 |
| 2019 | Critical Learning Periods in Deep Networks
Alessandro Achille, Matteo Rovere, Stefano Soatto |
ICLR (Poster) | 3 |
| 2019 | Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near ConvergenceabstractRegularization is typically understood as improving generalization by altering the landscape of local extrema to which the model eventually converges. Deep neural networks (DNNs), however, challenge this view: We show that removing regularization after an initial transient period has little effect on generalization, even if the final loss landscape is the same as if there had been no regularization. In some cases, generalization even improves after interrupting regularization. Conversely, if regularization is applied only after the initial transient, it has no effect on the final solution, whose generalization gap is as bad as if regularization never happened. This suggests that what matters for training deep networks is not just whether or how, but when to regularize. The phenomena we observe are manifest in different datasets (CIFAR-10, CIFAR-100, SVHN, ImageNet), different architectures (ResNet-18, All-CNN), different regularization methods (weight decay, data augmentation, mixup), different learning rate schedules (exponential, piece-wise constant). They collectively suggest that there is a "critical period'' for regularizing deep networks that is decisive of the final performance. More analysis should, therefore, focus on the transient rather than asymptotic behavior of learning. Aditya Golatkar, Alessandro Achille, Stefano Soatto |
NeurIPS | 3 |
| 2018 | Empirical Study of the Topology and Geometry of Deep NetworksabstractThe goal of this paper is to analyze the geometric properties of deep neural network image classifiers in the input space. We specifically study the topology of classification regions created by deep networks, as well as their associated decision boundary. Through a systematic empirical study, we show that state-of-the-art deep nets learn connected classification regions, and that the decision boundary in the vicinity of datapoints is flat along most directions. We further draw an essential connection between two seemingly unrelated properties of deep networks: their sensitivity to additive perturbations of the inputs, and the curvature of their decision boundary. The directions where the decision boundary is curved in fact characterize the directions to which the classifier is the most vulnerable. We finally leverage a fundamental asymmetry in the curvature of the decision boundary of deep nets, and propose a method to discriminate between original images, and images perturbed with small adversarial examples. We show the effectiveness of this purely geometric approach for detecting small adversarial perturbations in images, and for recovering the labels of perturbed images. Alhussein Fawzi, Seyed-Mohsen Moosavi-Dezfooli, Pascal Frossard, Stefano Soatto |
CVPR | 4 |
| 2018 | OATM: Occlusion Aware Template Matching by Consensus Set MaximizationabstractWe present a novel approach to template matching that is efficient, can handle partial occlusions, and comes with provable performance guarantees. A key component of the method is a reduction that transforms the problem of searching a nearest neighbor among N high-dimensional vectors, to searching neighbors among two sets of order √N vectors, which can be found efficiently using range search techniques. This allows for a quadratic improvement in search complexity, and makes the method scalable in handling large search spaces. The second contribution is a hashing scheme based on consensus set maximization, which allows us to handle occlusions. The resulting scheme can be seen as a randomized hypothesize-and-test algorithm, which is equipped with guarantees regarding the number of iterations required for obtaining an optimal solution with high probability. The predicted matching rates are validated empirically and the algorithm shows a significant improvement over the state-of-the-art in both speed and robustness to occlusions. Simon Korman, Mark Milam, Stefano Soatto |
CVPR | 3 |
| 2018 | SaaS: Speed as a Supervisor for Semi-supervised Learning
Safa Cicek, Alhussein Fawzi, Stefano Soatto |
ECCV (2) | 3 |
| 2018 | Visual-Inertial Object Detection and Mapping
Xiaohan Fei, Stefano Soatto |
ECCV (11) | 2 |
| 2018 | Reinforced Temporal Attention and Split-Rate Transfer for Depth-Based Person Re-identification
Nikolaos Karianakis, Zicheng Liu 0001, Yinpeng Chen, Stefano Soatto |
ECCV (5) | 4 |
| 2018 | Conditional Prior Networks for Optical Flow
Yanchao Yang 0001, Stefano Soatto |
ECCV (15) | 2 |
| 2018 | Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari, Stefano Soatto |
ICLR (Poster) | 2 |
| 2018 | Robustness of Classifiers to Universal Perturbations: A Geometric Perspective
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, Pascal Frossard, Stefano Soatto |
ICLR (Poster) | 5 |
| 2018 | Emergence of Invariance and Disentanglement in Deep RepresentationsabstractUsing established principles from Statistics and Information Theory, we show that invariance to nuisance factors in a deep neural network is equivalent to information minimality of the learned representation, and that stacking layers and injecting noise during training naturally bias the network towards learning invariant representations. We then decompose the cross-entropy loss used during training and highlight the presence of an inherent overfitting term. We propose regularizing the loss by bounding such a term in two equivalent ways: One with a Kullbach-Leibler term, which relates to a PAC-Bayes perspective; the other using the information in the weights as a measure of complexity of a learned model, yielding a novel Information Bottleneck for the weights. Finally, we show that invariance and independence of the components of the representation learned by the network are bounded above and below by the information in the weights, and therefore are implicitly optimized during training. The theory enables us to quantify and predict sharp phase transitions between underfitting and overfitting of random labels when using our regularized loss, which we verify in experiments, and sheds light on the relation between the geometry of the loss function, invariance properties of the learned representation, and generalization error. Alessandro Achille, Stefano Soatto |
J. Mach. Learn. Res. | 2 |
| 2018 | Information Dropout: Learning Optimal Representations Through Noisy ComputationabstractThe cross-entropy loss commonly used in deep learning is closely related to the defining properties of optimal representations, but does not enforce some of the key properties. We show that this can be solved by adding a regularization term, which is in turn related to injecting multiplicative noise in the activations of a Deep Neural Network, a special case of which is the common practice of dropout. We show that our regularized loss function can be efficiently minimized using Information Dropout, a generalization of dropout rooted in information theoretic principles that automatically adapts to the data and can better exploit architectures of limited capacity. When the task is the reconstruction of the input, we show that our loss function yields a Variational Autoencoder as a special case, thus providing a link between representation learning, information theory and variational inference. Finally, we prove that we can promote the creation of optimal disentangled representations simply by enforcing a factorized prior, a fact that has been observed empirically in recent work. Our experiments validate the theoretical intuitions behind our method, and we find that Information Dropout achieves a comparable or better generalization performance than binary dropout, especially on smaller models, since it can automatically adapt the noise to the structure of the network, as well as to the test sample. Alessandro Achille, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Zero Shot Learning via Multi-scale Manifold RegularizationabstractWe address zero-shot learning using a new manifold alignment framework based on a localized multi-scale transform on graphs. Our inference approach includes a smoothness criterion for a function mapping nodes on a graph (visual representation) onto a linear space (semantic representation), which we optimize using multi-scale graph wavelets. The robustness of the ensuing scheme allows us to operate with automatically generated semantic annotations, resulting in an algorithm that is entirely free of manual supervision, and yet improves the state-of-the-art as measured on benchmark datasets. Shay Deutsch, Soheil Kolouri, Kyungnam Kim, Yuri Owechko, Stefano Soatto |
CVPR | 5 |
| 2017 | Visual-Inertial-Semantic Scene Representation for 3D Object DetectionabstractWe describe a system to detect objects in three-dimensional space using video and inertial sensors (accelerometer and gyrometer), ubiquitous in modern mobile platforms from phones to drones. Inertials afford the ability to impose class-specific scale priors for objects, and provide a global orientation reference. A minimal sufficient representation, the posterior of semantic (identity) and syntactic (pose) attributes of objects in space, can be decomposed into a geometric term, which can be maintained by a localization-and-mapping filter, and a likelihood function, which can be approximated by a discriminatively-trained convolutional neural network The resulting system can process the video stream causally in real time, and provides a representation of objects in the scene that is persistent: Confidence in the presence of objects grows with evidence, and objects previously seen are kept in memory even when temporarily occluded, with their return into view automatically predicted to prime re-detection. Jingming Dong, Xiaohan Fei, Stefano Soatto |
CVPR | 3 |
| 2017 | S2F: Slow-to-Fast Interpolator FlowabstractWe introduce a method to compute optical flow at multiple scales of motion, without resorting to multi-resolution or combinatorial methods. It addresses the key problem of small objects moving fast, and resolves the artificial binding between how large an object is and how fast it can move before being diffused away by classical scale-space. Even with no learning, it achieves top performance on the most challenging optical flow benchmark. Moreover, the results are interpretable, and indeed we list the assumptions underlying our method explicitly. The key to our approach is the matching progression from slow to fast, as well as the choice of in-terpolation method, or equivalently the prior, to fill in regions where the data allows it. We use several off-the-shelf components, with relatively low sensitivity to parameter tuning. Computational cost is comparable to the state-of-the-art. Yanchao Yang 0001, Stefano Soatto |
CVPR | 2 |
| 2017 | Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, Riccardo Zecchina |
ICLR (Poster) | 3 |
| 2016 | An Empirical Evaluation of Current Convolutional Architectures' Ability to Manage Nuisance Location and Scale VariabilityabstractWe conduct an empirical study to test the ability of convolutional neural networks (CNNs) to reduce the effects of nuisance transformations of the input data, such as location, scale and aspect ratio. We isolate factors by adopting a common convolutional architecture either deployed globally on the image to compute class posterior distributions, or restricted locally to compute class conditional distributions given location, scale and aspect ratios of bounding boxes determined by proposal heuristics. In theory, averaging the latter should yield inferior performance compared to proper marginalization. Yet empirical evidence suggests the converse, leading us to conclude that - at the current level of complexity of convolutional architectures and scale of the data sets used to train them - CNNs are not very effective at marginalizing nuisance variability. We also quantify the effects of context on the overall classification task and its impact on the performance of CNNs, and propose improved sampling techniques for heuristic proposal schemes that improve end-to-end performance to state-of-the-art levels. We test our hypothesis on a classification task using the ImageNet Challenge benchmark and on a wide-baseline matching task using the Oxford and Fischer's datasets. Nikolaos Karianakis, Jingming Dong, Stefano Soatto |
CVPR | 3 |
| 2016 | A Simple Hierarchical Pooling Data Structure for Loop Closure
Xiaohan Fei, Konstantine Tsotsos, Stefano Soatto |
ECCV (3) | 3 |
| 2016 | ShapeFit and ShapeKick for Robust, Scalable Structure from Motion
Tom Goldstein, Paul Hand, Choongbum Lee, Vladislav Voroninski, Stefano Soatto |
ECCV (7) | 5 |
| 2016 | Intent-aware long-term prediction of pedestrian motionabstractWe present a method to predict long-term motion of pedestrians, modeling their behavior as jump-Markov processes with their goal a hidden variable. Assuming approximately rational behavior, and incorporating environmental constraints and biases, including time-varying ones imposed by traffic lights, we model intent as a policy in a Markov decision process framework. We infer pedestrian state using a Rao-Blackwellized filter, and intent by planning according to a stochastic policy, reflecting individual preferences in aiming at the same goal. Vasiliy Karasev, Alper Ayvaci, Bernd Heisele, Stefano Soatto |
ICRA | 4 |
| 2016 | Observability, Identifiability and Sensitivity of Vision-Aided Inertial Navigation
Joshua Hernandez, Konstantine Tsotsos, Stefano Soatto |
IJCAI | 3 |
| 2016 | A mid-level representation of visual structures for video compressionabstractA video coding system is presented that partitions the scene into "visual structures" and a residual "background" layer. A low-level representation ("track-template") of visual structures is proposed that exploits their temporal redundancy. A dictionary of track-templates is constructed that is used to encode video frames. We make optimal use of the dictionary in terms of rate-distortion by choosing a subset of the dictionary's elements for encoding using a Markov Random Field (MRF) formulation that places the track-templates in "depth" layers. The selected "track-templates" form the mid-level representation of the "visual structure" regions of the video. Our video coding system offers improvements over H.265/H.264 and other methods in a rate-distortion comparison. Georgios Georgiadis, Stefano Soatto |
WACV | 2 |
| 2016 | Robust Surface ReconstructionabstractWe propose a method to reconstruct surfaces from oriented point clouds corrupted by errors arising from range imaging sensors. The core of this technique is the formulation of the problem as a convex minimization that reconstructs the indicator function of the surface's interior and substitutes the usual least-squares fidelity terms by Huber penalties to be robust to outliers, recover sharp corners, and avoid the shrinking bias of least-squares models. To achieve both flexibility and accuracy, we couple an implicit parametrization that reconstructs surfaces of unknown topology with adaptive discretizations that avoid the high memory and computational cost of volumetric representations. The hierarchical structure of the discretizations speeds minimization through multiresolution, while the proposed splitting algorithm minimizes nondifferentiable functionals and is easy to parallelize. In experiments, our model improves reconstruction from synthetic and real data, while the choice of discretization affects both the accuracy of the reconstruction and its computational cost. Virginia Estellers, M. A. Scott, Stefano Soatto |
SIAM J. Imaging Sci. | 3 |
| 2015 | Multi-view feature engineering and learningabstractWe frame the problem of local representation of imaging data as the computation of minimal sufficient statistics that are invariant to nuisance variability induced by viewpoint and illumination. We show that, under very stringent conditions, these are related to “feature descriptors” commonly used in Computer Vision. Such conditions can be relaxed if multiple views of the same scene are available. We propose a sampling-based and a point-estimate based approximation of such a representation, compared empirically on image-to-(multiple)image matching, for which we introduce a multi-view wide-baseline matching benchmark, consisting of a mixture of real and synthetic objects with ground truth camera motion and dense three-dimensional geometry. Jingming Dong, Nikolaos Karianakis, Damek Davis, Joshua Hernandez, Jonathan Balzer, Stefano Soatto |
CVPR | 6 |
| 2015 | Domain-size pooling in local descriptors: DSP-SIFTabstractWe introduce a simple modification of local image descriptors, such as SIFT, based on pooling gradient orientations across different domain sizes, in addition to spatial locations. The resulting descriptor, which we call DSP-SIFT, outperforms other methods in wide-baseline matching benchmarks, including those based on convolutional neural networks, despite having the same dimension of SIFT and requiring no training. Jingming Dong, Stefano Soatto |
CVPR | 2 |
| 2015 | Texture representations for image and video synthesisabstractIn texture synthesis and classification, algorithms require a small texture to be provided as an input, which is assumed to be representative of a larger region to be re-synthesized or categorized. We focus on how to characterize such textures and automatically retrieve them. Most works generate these small input textures manually by cropping, which does not ensure maximal compression, nor that the selection is the best representative of the original. We construct a new representation that compactly summarizes a texture, while using less storage, that can be used for texture compression and synthesis. We also demonstrate how the representation can be integrated in our proposed video texture synthesis algorithm to generate novel instances of textures and video hole-filling. Finally, we propose a novel criterion that measures structural and statistical dissimilarity between textures. Georgios Georgiadis, Alessandro Chiuso, Stefano Soatto |
CVPR | 3 |
| 2015 | Efficient minimal-surface regularization of perspective depth maps in variational stereoabstractWe propose a method for dense three-dimensional surface reconstruction that leverages the strengths of shape-based approaches, by imposing regularization that respects the geometry of the surface, and the strength of depth-map-based stereo, by avoiding costly computation of surface topology. The result is a near real-time variational reconstruction algorithm free of the staircasing artifacts that affect depth-map and plane-sweeping approaches. This is made possible by exploiting the gauge ambiguity to design a novel representation of the regularizer that is linear in the parameters and hence amenable to be optimized with state-of-the-art primal-dual numerical schemes. Gottfried Munda, Jonathan Balzer, Stefano Soatto, Thomas Pock |
CVPR | 3 |
| 2015 | Causal video object segmentation from persistence of occlusionsabstractOcclusion relations inform the partition of the image domain into “objects” but are difficult to determine from a single image or short-baseline video. We show how long-term occlusion relations can be robustly inferred from video, and used within a convex optimization framework to segment the image domain into regions. We highlight the challenges in determining these occluder/occluded relations and ensuring regions remain temporally consistent, propose strategies to overcome them, and introduce an efficient numerical scheme to perform the partition directly on the pixel grid, without the need for superpixelization or other preprocessing steps. Brian Taylor 0001, Vasiliy Karasev, Stefano Soatto |
CVPR | 3 |
| 2015 | Exploiting Temporal Redundancy of Visual Structures for Video CompressionabstractSummary form only given. We present a video coding system that partitions the scene into "visual structures" and a residual "background" layer. The system exploits the temporal redundancy of visual structures to compress video sequences. We construct a dictionary of track-templates, which correspond to a representation of visual structures. We subsequently choose a subset of the dictionary's elements to encode video frames using a Markov Random Field (MRF) formulation that places the track-templates in "depth" layers. Our video coding system offers an improvement over H.265/H.264 and other methods in a rate-distortion comparison. Georgios Georgiadis, Stefano Soatto |
DCC | 2 |
| 2015 | Self-Occlusions and Disocclusions in Causal Video Object SegmentationabstractWe propose a method to detect disocclusion in video sequences of three-dimensional scenes and to partition the disoccluded regions into objects, defined by coherent deformation corresponding to surfaces in the scene. Our method infers deformation fields that are piecewise smooth by construction without the need for an explicit regularizer and the associated choice of weight. It then partitions the disoccluded region and groups its components with objects by leveraging on the complementarity of motion and appearance cues: Where appearance changes within an object, motion can usually be reliably inferred and used for grouping. Where appearance is close to constant, it can be used for grouping directly. We integrate both cues in an energy minimization framework, incorporate prior assumptions explicitly into the energy, and propose a numerical scheme. Yanchao Yang 0001, Ganesh Sundaramoorthi, Stefano Soatto |
ICCV | 3 |
| 2015 | A Power-Performance Approach to Comparing Sensor Families, with application to comparing neuromorphic to traditional vision sensorsabstractThere is considerable freedom in choosing the sensors to be equipped on a robot. Currently many sensing technologies are available (radar, lidar, vision sensors, time-of-flight cameras, etc.). For each class, there are additional choices regarding the exact sensor parameters (spatial resolution, frame rate, etc.). Which sensor is best? In general, this question needs to be qualified. It depends on the task. In an estimation task, the answer depends on the prior for the signal. In a control task, the answer depends exactly on which are the sufficient statistics for computing the control signal. This paper shows that an ulterior qualification that needs to be made: the answer depends on the power available for sensing, even when the task is fixed. We define the “power-performance” curve as the performance attainable on a task for a given level of sensing power. We show that this approach is well suited to comparing a traditional CMOS sensor with the recently available “neuromorphic” sensors. We discuss estimation tasks with different priors for the signal. We find priors for which one sensor dominates the other and vice-versa, priors for which they are equivalent, and priors for which the answer depends on the power available. This shows that comparing sensors is a quite delicate problem. It also suggests that the optimal architecture might have more that one sensor, and would switch sensors on and off according to the performance level required instantaneously. Andrea Censi, Erich Mueller, Emilio Frazzoli, Stefano Soatto |
ICRA | 4 |
| 2015 | Observability, identifiability and sensitivity of vision-aided inertial navigationabstractWe analyze the observability of 3-D pose from the fusion of visual and inertial sensors. Because the model contains unknown parameters, such as sensor biases, the problem is usually cast as a mixed filtering/identification, with the resulting observability analysis providing necessary conditions for convergence to a unique point estimate. Most models treat sensor bias rates as “noise,” independent of other states, including biases themselves, an assumption that is patently violated in practice. We show that, when this assumption is lifted, the resulting model is not observable, and therefore existing analyses cannot be used to conclude that the set of states that are indistinguishable from the measurements is a singleton. In other words, the resulting model is not observable. We therefore re-cast the analysis as one of sensitivity: Rather than attempting to prove that the set of indistinguishable trajectories is a singleton, we derive bounds on its volume, as a function of characteristics of the sensor and other sufficient excitation conditions. This provides an explicit characterization of the indistinguishable set that can be used for analysis and validation purposes. Joshua Hernandez, Konstantine Tsotsos, Stefano Soatto |
ICRA | 3 |
| 2015 | Robust inference for visual-inertial sensor fusionabstractInference of three-dimensional motion from the fusion of inertial and visual sensory data has to contend with the preponderance of outliers in the latter. Robust filtering deals with the joint inference and classification task of selecting which data fits the model, and estimating its state. We derive the optimal discriminant and propose several approximations, some used in the literature, others new. We compare them analytically, by pointing to the assumptions underlying their approximations, and empirically. We show that the best performing method improves the performance of state-of-the-art visual-inertial sensor fusion systems, while retaining the same computational complexity. Konstantine Tsotsos, Alessandro Chiuso, Stefano Soatto |
ICRA | 3 |
| 2015 | Shape Matching Using Multiscale Integral InvariantsabstractWe present a shape descriptor based on integral kernels. Shape is represented in an implicit form and it is characterized by a series of isotropic kernels that provide desirable invariance properties. The shape features are characterized at multiple scales which form a signature that is a compact description of shape over a range of scales. The shape signature is designed to be invariant with respect to group transformations which include translation, rotation, scaling, and reflection. In addition, the integral kernels that characterize local shape geometry enable the shape signature to be robust with respect to undesirable perturbations while retaining discriminative power. Use of our shape signature is demonstrated for shape matching based on a number of synthetic and real examples. Byung-Woo Hong, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Adaptive Regularization With the Structure TensorabstractNatural images exhibit geometric structures that are informative of the properties of the underlying scene. Modern image processing algorithms respect such characteristics by employing regularizers that capture the statistics of natural images. For instance, total variation (TV) respects the highly kurtotic distribution of the pointwise gradient by allowing for large magnitude outlayers. However, the gradient magnitude alone does not capture the directionality and scale of local structures in natural images. The structure tensor provides a more meaningful description of gradient information as it describes both the size and orientation of the image gradients in a neighborhood of each point. Based on this observation, we propose a variational model for image reconstruction that employs a regularization functional adapted to the local geometry of image by means of its structure tensor. Our method alternates two minimization steps: 1) robust estimation of the structure tensor as a semidefinite program and 2) reconstruction of the image with an adaptive regularizer defined from this tensor. This two-step procedure allows us to extend anisotropic diffusion into the convex setting and develop robust, efficient, and easy-to-code algorithms for image denoising, deblurring, and compressed sensing. Our method extends naturally to nonlocal regularization, where it exploits the local self-similarity of natural images to improve nonlocal TV and diffusion operators. Our experiments show a consistent accuracy improvement over classic regularization. Virginia Estellers, Stefano Soatto, Xavier Bresson |
IEEE Trans. Image Process. | 2 |
| 2014 | Cavlectometry: Towards Holistic Reconstruction of Large Mirror ObjectsabstractWe introduce a method based on the deflectometry principle for the reconstruction of specular objects exhibiting significant size and geometric complexity. A key feature of our approach is the deployment of an Automatic Virtual Environment (CAVE) as pattern generator. To unfold the full power of this experimental setup, an optical encoding scheme is developed which accounts for the distinctive topology of the CAVE. Furthermore, we devise an algorithm for detecting the object of interest in raw deflect metric images. The segmented foreground is used for single-view reconstruction, the background for estimation of the camera pose, necessary for calibrating the sensor system. Experiments suggest a significant gain of coverage in single measurements compared to previous methods. Jonathan Balzer, Daniel Acevedo Feliz, Stefano Soatto, Sebastian Höfer, Markus Hadwiger, Jürgen Beyerer |
3DV | 3 |
| 2014 | Second-Order Shape Optimization for Geometric Inverse Problems in VisionabstractWe develop a method for optimization in shape spaces, i.e., sets of surfaces modulo re-parametrization. Unlike previously proposed gradient flows, we achieve superlinear convergence rates through an approximation of the shape Hessian, which is generally hard to compute and suffers from a series of degeneracies. Our analysis highlights the role of mean curvature motion in comparison with first-order schemes: instead of surface area, our approach penalizes deformation, either by its Dirichlet energy or total variation, and hence does not suffer from shrinkage. The latter regularizer sparks the development of an alternating direction method of multipliers on triangular meshes. Therein, a conjugate-gradient solver enables us to bypass formation of the Gaussian normal equations appearing in the course of the overall optimization. We combine all of these ideas in a versatile geometric variation-regularized Levenberg-Marquardt-type method applicable to a variety of shape functionals, depending on intrinsic properties of the surface such as normal field and curvature as well as its embedding into space. Promising experimental results are reported. Jonathan Balzer, Stefano Soatto |
CVPR | 2 |
| 2014 | Asymmetric Sparse Kernel Approximations for Large-Scale Visual SearchabstractWe introduce an asymmetric sparse approximate embedding optimized for fast kernel comparison operations arising in large-scale visual search. In contrast to other methods that perform an explicit approximate embedding using kernel PCA followed by a distance compression technique in Rd, which loses information at both steps, our method utilizes the implicit kernel representation directly. In addition, we empirically demonstrate that our method needs no explicit training step and can operate with a dictionary of random exemplars from the dataset. We evaluate our method on three benchmark image retrieval datasets: SIFT1M, ImageNet, and 80M-TinyImages. Damek Davis, Jonathan Balzer, Stefano Soatto |
CVPR | 3 |
| 2014 | Active Frame, Location, and Detector Selection for Automated and Manual Video AnnotationabstractWe describe an information-driven active selection approach to determine which detectors to deploy at which location in which frame of a video to minimize semantic class label uncertainty at every pixel, with the smallest computational cost that ensures a given uncertainty bound. We show minimal performance reduction compared to a "paragon" algorithm running all detectors at all locations in all frames, at a small fraction of the computational cost. Our method can handle uncertainty in the labeling mechanism, so it can handle both "oracles" (manual annotation) or noisy detectors (automated annotation). Vasiliy Karasev, Avinash Ravichandran, Stefano Soatto |
CVPR | 3 |
| 2014 | Volumetric reconstruction applied to perceptual studies of size and weightabstractWe explore the application of volumetric reconstruction from structured-light sensors in cognitive neuroscience, specifically in the quantification of the size-weight illusion, whereby humans tend to systematically perceive smaller objects as heavier. We investigate the performance of two commercial structured-light scanning systems in comparison to one we developed specifically for this application. Our method has two main distinct features: First, it only samples a sparse series of viewpoints, unlike other systems such as the Kinect Fusion. Second, instead of building a distance field for the purpose of points-to-surface conversion directly, we pursue a first-order approach: the distance function is recovered from its gradient by a screened Poisson reconstruction, which is very resilient to noise and yet preserves high-frequency signal components. Our experiments show that the quality of metric reconstruction from structured light sensors is subject to systematic biases, and highlights the factors that influence it. Our main performance index rates estimates of volume (a proxy of size), for which we review a well-known formula applicable to incomplete meshes. Our code and data will be made publicly available upon completion of the anonymous review process. Jonathan Balzer, Megan A. K. Peters, Stefano Soatto |
WACV | 3 |
| 2013 | CLAM: Coupled Localization and Mapping with Efficient Outlier HandlingabstractWe describe a method to efficiently generate a model (map) of small-scale objects from video. The map encodes sparse geometry as well as coarse photometry, and could be used to initialize dense reconstruction schemes as well as to support recognition and localization of three-dimensional objects. Self-occlusions and the predominance of outliers present a challenge to existing online Structure From Motion and Simultaneous Localization and Mapping systems. We propose a unified inference criterion that encompasses map building and localization (object detection) relative to the map in a coupled fashion. We establish correspondence in a computationally efficient way without resorting to combinatorial matching or random-sampling techniques. Instead, we use a simpler M-estimator that exploits putative correspondence from tracking after photometric and topological validation. We have collected a new dataset to benchmark model building in the small scale, which we test our algorithm on in comparison to others. Although our system is significantly leaner than previous ones, it compares favorably to the state of the art in terms of accuracy and robustness. Jonathan Balzer, Stefano Soatto |
CVPR | 2 |
| 2013 | Nonlinearly Constrained MRFs: Exploring the Intrinsic Dimensions of Higher-Order CliquesabstractThis paper introduces an efficient approach to integrating non-local statistics into the higher-order Markov Random Fields (MRFs) framework. Motivated by the observation that many non-local statistics (e.g., shape priors, color distributions) can usually be represented by a small number of parameters, we reformulate the higher-order MRF model by introducing additional latent variables to represent the intrinsic dimensions of the higher-order cliques. The resulting new model, called NC-MRF, not only provides the flexibility in representing the configurations of higher-order cliques, but also automatically decomposes the energy function into less coupled terms, allowing us to design an efficient algorithmic framework for maximum a posteriori (MAP) inference. Based on this novel modeling/ inference framework, we achieve state-of-the-art solutions to the challenging problems of class-specific image segmentation and template-based 3D facial expression tracking, which demonstrate the potential of our approach. Chaohui Wang, Stefano Soatto, Shing-Tung Yau |
CVPR | 3 |
| 2013 | Texture CompressionabstractWe characterize ``visual textures'' as realizations of a stationary, ergodic, Markovian process, and propose using its approximate minimal sufficient statistics for compressing texture images. We propose inference algorithms for estimating the ``state'' of such process and its ``variability''. These represent the encoding stage. We also propose a non-parametric sampling scheme for decoding, by synthesizing textures from their encoding. While these are not faithful reproductions of the original textures (so they would fail a comparison test based on PSNR), they capture the statistical properties of the underlying process, as we demonstrate empirically. We also quantify the tradeoff between fidelity (measured by a proxy of a perceptual score) and complexity. Georgios Georgiadis, Alessandro Chiuso, Stefano Soatto |
DCC | 3 |
| 2012 | Actionable saliency detection: Independent motion detection without independent motion estimationabstractWe present a model and an algorithm to detect salient regions in video taken from a moving camera. In particular, we are interested in capturing small objects that move independently in the scene, such as vehicles and people as seen from aerial or ground vehicles. Many of the scenarios of interest challenge existing schemes based on background subtraction (background motion too complex), multi-body motion estimation (insufficient parallax), and occlusion detection (uniformly textured background regions). We adopt a robust statistical inference approach to simultaneously estimate a maximally reduced regressor, and select regions that violate the null hypothesis (co-visibility under an epipolar domain deformation) as “salient”. We show that our algorithm can perform even in the absence of camera calibration information: while the resulting motion estimates would be incorrect, the partition of the domain into salient vs. non-salient is unaffected. We demonstrate our algorithm on video footage from helicopters, airplanes, and ground vehicles. Georgios Georgiadis, Alper Ayvaci, Stefano Soatto |
CVPR | 3 |
| 2012 | Discovering discriminative action parts from mid-level video representationsabstractWe describe a mid-level approach for action recognition. From an input video, we extract salient spatio-temporal structures by forming clusters of trajectories that serve as candidates for the parts of an action. The assembly of these clusters into an action class is governed by a graphical model that incorporates appearance and motion constraints for the individual parts and pairwise constraints for the spatio-temporal dependencies among them. During training, we estimate the model parameters discriminatively. During classification, we efficiently match the model to a video using discrete optimization. We validate the model's classification ability in standard benchmark datasets and illustrate its potential to support a fine-grained analysis that not only gives a label to a video, but also identifies and localizes its constituent parts. Michalis Raptis, Iasonas Kokkinos, Stefano Soatto |
CVPR | 3 |
| 2012 | Scene-Aware Video Modeling and CompressionabstractWe describe a video compression methodology that exploits the structure of the data formation process, whereby the "source'' is the scene, and the ``channel'' includes scaling and occlusion phenomena that are critical elements of image formation. Thus our scheme involves occlusion detection, optical flow computation, texture/structure partition, and a notion of proper sampling. We show preliminary but promising results that exceed baseline compression performance, albeit at an increased computational cost. Georgios Georgiadis, Stefano Soatto |
DCC | 2 |
| 2012 | Long-Range Spatio-Temporal Modeling of Video with Application to Fire Detection
Avinash Ravichandran, Stefano Soatto |
ECCV (2) | 2 |
| 2012 | Video upscaling via spatio-temporal self-similarity
Alper Ayvaci, Hailin Jin, Zhe Lin 0001, Scott Cohen, Stefano Soatto |
ICPR | 5 |
| 2012 | Controlled Recognition Bounds for Visual Learning and ExplorationabstractWe describe the tradeoff between the performance in a visual recognition problem and the control authority that the agent can exercise on the sensing process. We focus on the problem of “visual search” of an object in an otherwise known and static scene, propose a measure of control authority, and relate it to the expected risk and its proxy (conditional entropy of the posterior density). We show this analytically, as well as empirically by simulation using the simplest known model that captures the phenomenology of image formation, including scaling and occlusions. We show that a “passive” agent given a training set can provide no guarantees on performance beyond what is afforded by the priors, and that an “omnipotent” agent, capable of infinite control authority, can achieve arbitrarily good performance (asymptotically). Vasiliy Karasev, Alessandro Chiuso, Stefano Soatto |
NIPS | 3 |
| 2012 | Fast planar object detection and tracking via edgel templatesabstractWe describe an efficient method to detect and track planar objects using a template of edge segments. Such segments are selected at multiple scales based on gradient magnitude; their positions and orientations are used to determine a canonical reference frame where the descriptor is computed based on quantized orientation. The resulting descriptors are efficiently matched using logical operations, and tracked between frames. The method yields pose estimates that are robust to scale changes, foreshortening, partial occlusions, and is suitable for use in augmented reality and human-computer interaction. Taehee Lee 0002, Stefano Soatto |
WACV | 2 |
| 2012 | Sparse Occlusion Detection with Optical Flow
Alper Ayvaci, Michalis Raptis, Stefano Soatto |
Int. J. Comput. Vis. | 3 |
| 2012 | Detachable Object Detection: Segmentation and Depth Ordering from Short-Baseline VideoabstractWe describe an approach for segmenting a moving image into regions that correspond to surfaces in the scene that are partially surrounded by the medium. It integrates both appearance and motion statistics into a cost functional that is seeded with occluded regions and minimized efficiently by solving a linear programming problem. Where a short observation time is insufficient to determine whether the object is detachable, the results of the minimization can be used to seed a more costly optimization based on a longer sequence of video data. The result is an entirely unsupervised scheme to detect and segment an arbitrary and unknown number of objects. We test our scheme to highlight the potential, as well as limitations, of our approach. Alper Ayvaci, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Learning and matching multiscale template descriptors for real-time detection, localization and trackingabstractWe describe a system to learn an object template from a video stream, and localize and track the corresponding object in live video. The template is decomposed into a number of local descriptors, thus enabling detection and tracking in spite of partial occlusion. Each local descriptor aggregates contrast invariant statistics (normalized intensity and gradient orientation) across scales, in a way that enables matching under significant scale variations. Low-level tracking during the training video sequence enables capturing object-specific variability due to the shape of the object, which is encapsulated in the descriptor. Salient locations on both the template and the target image are used as hypotheses to expedite matching. Taehee Lee 0002, Stefano Soatto |
CVPR | 2 |
| 2011 | Controlled Recognition Bounds for Scaling and Occlusion ChannelsabstractSummary form only given. In this article the simplest task (binary decision) for channels subject to scaling and occlusion phenomena, ubiquitous in remote sensing and imaging data is discussed. In order to extend Rate-Distortion theory, the tradeoff between decision performance (expected risk) and resources ("rate") is characterized. In the presence of scaling and occlusions, in order to trade off the expected risk, it is necessary to exercise control on the sensing process. Thus the natural generalization of "rate" is not the volume of the data provided by the sensor, but the amount of control authority that can be exercised on it, measured by the volume of the reachable set and the energy of the input. Stefano Soatto, Alessandro Chiuso |
DCC | 1 |
| 2011 | Edgel templates for fast planar object detection and pose estimationabstractWe describe a method to select edgels and to calculate gradient orientation-based template descriptors for edgel features. An edgel is selected within a grid block based on gradient magnitude; its position and orientation are used to determine a canonical frame where the descriptor is computed based on quantized orientation. The resulting descriptor is efficiently matched using logical operations. We demonstrate the use of the resulting edgel detection and description method for planar object detection and pose estimation. Taehee Lee 0002, Stefano Soatto |
ISMAR | 2 |
| 2011 | Multiple Instance FilteringabstractWe propose a robust filtering approach based on semi-supervised and multiple instance learning (MIL). We assume that the posterior density would be unimodal if not for the effect of outliers that we do not wish to explicitly model. Therefore, we seek for a point estimate at the outset, rather than a generic approximation of the entire posterior. Our approach can be thought of as a combination of standard finite-dimensional filtering (Extended Kalman Filter, or Unscented Filter) with multiple instance learning, whereby the initial condition comes with a putative set of inlier measurements. We show how both the state (regression) and the inlier set (classification) can be estimated iteratively and causally by processing only the current measurement. We illustrate our approach on visual tracking problems whereby the object of interest (target) moves and evolves as a result of occlusions and deformations, and partial knowledge of the target is given in the form of a bounding box (training set). Kamil Wnuk, Stefano Soatto |
NIPS | 2 |
| 2011 | Video-based descriptors for object recognition
Taehee Lee 0002, Stefano Soatto |
Image Vis. Comput. | 2 |
| 2011 | A New Geometric Metric in the Space of Curves, and Applications to Tracking Deforming Objects by Prediction and FilteringabstractWe define a novel metric on the space of closed planar curves which decomposes into three intuitive components. According to this metric, centroid translations, scale changes, and deformations are orthogonal, and the metric is also invariant with respect to reparameterizations of the curve. While earlier related Sobolev metrics for curves exhibit some general similarities to the novel metric proposed in this work, they lacked this important three-way orthogonal decomposition, which has particular relevance for tracking in computer vision. Another positive property of this new metric is that the Riemannian structure that is induced on the space of curves is a smooth Riemannian manifold, which is isometric to a classical well-known manifold. As a consequence, geodesics and gradients of energies defined on the space can be computed using fast closed-form formulas, and this has obvious benefits in numerical applications. The obtained Riemannian manifold of curves is ideal for addressing complex problems in computer vision; one such example is the tracking of highly deforming objects. Previous works have assumed that the object deformation is smooth, which is realistic for the tracking problem, but most have restricted the deformation to belong to a finite-dimensional group—such as affine motions—or to finitely parameterized models. This is too restrictive for highly deforming objects such as the contour of a beating heart. We adopt the smoothness assumption implicit in previous work, but we lift the restriction to finite-dimensional motions/deformations. We define a dynamical model in this Riemannian manifold of curves and use it to perform filtering and prediction to infer and extrapolate not just the pose (a finitely parameterized quantity) of an object but its deformation (an infinite-dimensional quantity) as well. We illustrate these ideas using a simple first-order dynamical model and show that it can be effective even on image sequences where existing methods fail. Ganesh Sundaramoorthi, Andrea Mennucci, Stefano Soatto, Anthony J. Yezzi |
SIAM J. Imaging Sci. | 3 |
| 2010 | Warping background subtractionabstractWe present a background model that differentiates between background motion and foreground objects. Unlike most models that represent the variability of pixel intensity at a particular location in the image, we model the underlying warping of pixel locations arising from background motion. The background is modeled as a set of warping layers, where at any given time, different layers may be visible due to the motion of an occluding layer. Foreground regions are thus defined as those that cannot be modeled by some composition of some warping of these background layers. We illustrate this concept by first reducing the possible warps to those where the pixels are restricted to displacements within a spatial neighborhood, and then learning the appropriate size of that spatial neighborhood. Then, we show how changes in intensity/color histograms of pixel neighborhoods can be used to discriminate foreground and background regions. We find that this approach compares favorably with the state of the art, while requiring less computation. Teresa Ko, Stefano Soatto, Deborah Estrin |
CVPR | 2 |
| 2010 | Spike train driven dynamical models for human actionsabstractWe investigate dynamical models of human motion that can support both synthesis and analysis tasks. Unlike coarser discriminative models that work well when action classes are nicely separated, we seek models that have fine-scale representational power and can therefore model subtle differences in the way an action is performed. To this end, we model an observed action as an (unknown) linear time-invariant dynamical model of relatively small order, driven by a sparse bounded input signal. Our motivating intuition is that the time-invariant dynamics will capture the unchanging physical characteristics of an actor, while the inputs used to excite the system will correspond to a causal signature of the action being performed. We show that our model has sufficient representational power to closely approximate large classes of non-stationary actions with significantly reduced complexity. We also show that temporal statistics of the inferred input sequences can be compared in order to recognize actions and detect transitions between them. Michalis Raptis, Kamil Wnuk, Stefano Soatto |
CVPR | 3 |
| 2010 | Curious snakes: A minimum latency solution to the cluttered background problem in active contoursabstractWe present a region-based active contour detection algorithm for objects that exhibit relatively homogeneous photometric characteristics (e.g. smooth color or gray levels), embedded in complex background clutter. Current methods either frame this problem in Bayesian classification terms, where precious modeling resources are expended representing the complex background away from decision boundaries, or use heuristics to limit the search to local regions around the object of interest. We propose an adaptive lookout region, whose size depends on the statistics of the data, that are estimated along with the boundary during the detection process. The result is a “curious snake” that explores the outside of the decision boundary only locally to the extent necessary to achieve a good tradeoff between missed detections and narrowest “lookout” region, drawing inspiration from the literature of minimum-latency set-point change detection and robust statistics. This development makes fully automatic detection in complex backgrounds a realistic possibility for active contours, allowing us to exploit their powerful geometric modeling capabilities compared with other approaches used for segmentation of cluttered scenes. To this end, we introduce an automatic initialization method tailored to our model that overcomes one of the primary obstacles in using active contours for fully automatic object detection. Ganesh Sundaramoorthi, Stefano Soatto, Anthony J. Yezzi |
CVPR | 2 |
| 2010 | Texture Regimes for Entropy-Based Multiscale Image Analysis
Sylvain Boltz, Frank Nielsen, Stefano Soatto |
ECCV (3) | 3 |
| 2010 | Tracklet Descriptors for Action Modeling and Video Analysis
Michalis Raptis, Stefano Soatto |
ECCV (1) | 2 |
| 2010 | Earth Mover Distance on superpixelsabstractEarth Mover Distance (EMD) is a popular distance to compute distances between Probability Density Functions (PDFs). It has been successfully applied in a wide selection of problems of image processing. This success comes from two reasons, a physical one, since it computes a physical cost to transport an element of mass between two images or two histograms, and a statistical one, since it is a cross-bin metric (as opposed to a bin-wise metric). In computer vision, these features are useful since small variation of illuminance can shift the histogram. However, histograms are not a sufficient statistic to discriminate images since they ignore all geometric correlations. In addition, transport also called flow of an histogram loose the information of geometric flow to warp one image on to an other. This paper proposes a new construction of EMD between images. This construction approximates the EMD between two images, by computing a pixel-wise transport at the complexity cost of computing an EMD between 1-D Histograms and preserves the geometrical and topological structure of the image. This construction simply relies on a segmentation of the image (also called superpixelization of the image). Results on matching on images shows the stability of the method even when the superpixelizations are highly inconsistent across images. Sylvain Boltz, Frank Nielsen, Stefano Soatto |
ICIP | 3 |
| 2010 | Feature tracking and object recognition on a hand-heldabstractWe demonstrate a visual recognition system operating on a hand-held device, with the help of an efficient and robust feature tracking and an object recognition mechanism that can be used for interactive mobile applications. In our recognition system, corner features are detected from captured video frames in a multi-scale image pyramid, and are tracked between consecutive frames efficiently. In order to perform object recognition, local descriptors are calculated on the tracked features, and quantized using a vocabulary tree. For each object, a bag-of-words model is learned from multiple views. The learned objects are recognized by computing the ranking score for the set of features in a single video frame. Our feature tracking algorithm and local descriptors are different than the Lucas-Kanade algorithm in image pyramid or the SIFT descriptor, however improving the efficiency and accuracy. For our implementation on a mobile phone, we used an iPhone 3GS with a 600MHz ARM chip CPU. The video frame is captured from a camera preview screen at a rate of 15 frames per second using the public API. The task of object recognition on a mobile phone runs at around 7 frames per second, including the feature tracking and descriptor calculation. Taehee Lee 0002, Stefano Soatto |
ISMAR | 2 |
| 2010 | Occlusion Detection and Motion Estimation with Convex OptimizationabstractWe tackle the problem of simultaneously detecting occlusions and estimating optical flow. We show that, under standard assumptions of Lambertian reflection and static illumination, the task can be posed as a convex minimization problem. Therefore, the solution, computed using efficient algorithms, is guaranteed to be globally optimal, for any number of independently moving objects, and any number of occlusion layers. We test the proposed algorithm on benchmark datasets, expanded to enable evaluation of occlusion detection performance. Alper Ayvaci, Michalis Raptis, Stefano Soatto |
NIPS | 3 |
| 2010 | Editorial for the Special Issue on Photometric Analysis for Computer Vision
Peter N. Belhumeur, Katsushi Ikeuchi, Emmanuel Prados, Stefano Soatto, Peter F. Sturm |
Int. J. Comput. Vis. | 4 |
| 2010 | Embedded Imagers: Detecting, Localizing, and Recognizing Objects and Events in Natural HabitatsabstractImaging sensors, or “imagers,” embedded in the natural environment enable remote collection of large quantities of data, thus easing the design and deployment of sensing systems in a variety of application domains. Yet, the data collected from such imagers are difficult to interpret due to a variety of “nuisance factors” in the data formation process, such as illumination, vantage point, partial occlusions, etc. These are especially severe in natural environments, where the objects of interest (e.g., plants, animals) have evolved to blend with their habitat, exhibit complex variability in shape and appearance, and perform rapid motions against dynamic backgrounds with rapid illumination changes. We describe three applications that exemplify these problems and the solutions we developed. First, we show how temporal oversampling can simplify the analysis of a slow process such as the avian nesting cycle. Then, we show how to overcome temporal undersampling in order to detect birds at a feeder station. Finally, we show how to exploit temporal consistency to reliably detect pollinators as they visit flowers in the field. Teresa Ko, Josh Hyman, Eric A. Graham, Mark H. Hansen, Stefano Soatto, Deborah Estrin |
Proc. IEEE | 5 |
| 2010 | Face verification across age progression using discriminative methodsabstractFace verification in the presence of age progression is an important problem that has not been widely addressed. In this paper, we study the problem by designing and evaluating discriminative approaches. These directly tackle verification tasks without explicit age modeling, which is a hard problem by itself. First, we find that the gradient orientation, after discarding magnitude information, provides a simple but effective representation for this problem. This representation is further improved when hierarchical information is used, which results in the use of the gradient orientation pyramid (GOP). When combined with a support vector machine GOP demonstrates excellent performance in all our experiments, in comparison with seven different approaches including two commercial systems. Our experiments are conducted on the FGnet dataset and two large passport datasets, one of them being the largest ever reported for recognition tasks. Second, taking advantage of these datasets, we empirically study how age gaps and related issues (including image quality, spectacles, and facial hair) affect recognition algorithms. We found surprisingly that the added difficulty of verification produced by age gaps becomes saturated after the gap is larger than four years, for gaps of up to ten years. In addition, we find that image quality and eyewear present more of a challenge than facial hair. Haibin Ling, Stefano Soatto, Narayanan Ramanathan, David Jacobs 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | Heartbeat of a nest: Using imagers as biological sensorsabstractWe present a scalable end-to-end system for vision-based monitoring of natural environments, and illustrate its use for the analysis of avian nesting cycles. Our system enables automated analysis of thousands of images, where manual processing would be infeasible. We automate the analysis of raw imaging data using statistics that are tailored to the task of interest. These “features” are a representation to be fed to classifiers that exploit spatial and temporal consistencies. Our testbed can detect the presence or absence of a bird with an accuracy of 82%, count eggs with an accuracy of 84%, and detect the inception of the nesting stage within a day. Our results demonstrate the challenges and potential benefits of using imagers as biological sensors. An exploration of system performance under varying image resolution and frame rate suggest that an in situ adaptive vision system is technically feasible. Teresa Ko, Shaun Ahmadian, John Hicks, Mohammad H. Rahimi, Deborah Estrin, Stefano Soatto, Sharon Coe, Michael Hamilton 0001 |
ACM Trans. Sens. Networks | 6 |
| 2009 | On the set of images modulo viewpoint and contrast changesabstractWe consider regions of images that exhibit smooth statistics, and pose the question of characterizing the "essence" of these regions that matters for recognition. Ideally, this would be a statistic (a function of the image) that does not depend on viewpoint and illumination, and yet is sufficient for the task. In this manuscript, we show that such statistics exist. That is, one can compute deterministic functions of the image that contain all the "information" present in the original image, except for the effects of viewpoint and illumination. We also show that such statistics are supported on a "thin" (zero-measure) subset of the image domain, and thus the "information" in an image that is relevant for recognition is sparse. Yet, from this thin set one can reconstruct an image that is equivalent to the original up to a change of viewpoint and local illumination (contrast). Finally, we formalize the notion of "information" an image contains for the purpose of viewpoint- and illumination- invariant tasks, which we call "actionable information" following ideas of J. J. Gibson. Ganesh Sundaramoorthi, Peter Petersen, V. S. Varadarajan, Stefano Soatto |
CVPR | 4 |
| 2009 | Nonrigid registration combining global and local statisticsabstractIn this paper we exploit normalized mutual information for the nonrigid registration of multimodal images. Rather than assuming that image statistics are spatially stationary, as often done in traditional information-theoretic methods, we take into account the spatial variability through a weighted combination of global normalized mutual information and local matching statistics. Spatial relationships are incorporated into the registration criterion by adoptively adjusting the weight according to the strength of local cues. With a continuous representation of images and Parzen window estimators, we have developed closed-form expressions of the first-order variation with respect to any general, nonparametric, infinite-dimensional deformation of the image domain. To characterize the performance of the proposed approach, synthetic phantoms, simulated MRIs, and clinical data are used in a validation study. The results suggest that the augmented normalized mutual information provides substantial improvements in terms of registration accuracy and robustness. Zhao Yi, Stefano Soatto |
CVPR | 2 |
| 2009 | Class segmentation and object localization with superpixel neighborhoodsabstractWe propose a method to identify and localize object classes in images. Instead of operating at the pixel level, we advocate the use of superpixels as the basic unit of a class segmentation or pixel localization scheme. To this end, we construct a classifier on the histogram of local features found in each superpixel. We regularize this classifier by aggregating histograms in the neighborhood of each superpixel and then refine our results further by using the classifier in a conditional random field operating on the superpixel graph. Our proposed method exceeds the previously published state-of-the-art on two challenging datasets: Graz-02 and the PASCAL VOC 2007 Segmentation Challenge. Brian Fulkerson, Andrea Vedaldi, Stefano Soatto |
ICCV | 3 |
| 2009 | Actionable information in visionabstractI propose a notion of visual information as the complexity not of the raw images, but of the images after the effects of nuisance factors such as viewpoint and illumination are discounted. It is rooted in ideas of J. J. Gibson, and stands in contrast to traditional information as entropy or coding length of the data regardless of its use, and regardless of the nuisance factors affecting it. The non-invertibility of nuisances such as occlusion and quantization induces an “information gap” that can only be bridged by controlling the data acquisition process. Measuring visual information entails early vision operations, tailored to the structure of the nuisances so as to be “lossless” with respect to visual decision and control tasks (as opposed to data transmission and storage tasks implicit in traditional Information Theory). I illustrate these ideas on visual exploration, whereby a “Shannonian Explorer” guided by the entropy of the data navigates unaware of the structure of the physical space surrounding it, while a “Gibsonian Explorer” is guided by the topology of the environment, despite measuring only images of it, without performing 3D reconstruction. The operational definition of visual information suggests desirable properties that a visual representation should possess to best accomplish vision-based decision and control tasks. Stefano Soatto |
ICCV | 1 |
| 2009 | Unsupervised multiphase segmentation: A recursive approach
Kangyu Ni, Byung-Woo Hong, Stefano Soatto, Tony F. Chan |
Comput. Vis. Image Underst. | 3 |
| 2009 | Hybrid Dynamical Models of Human Motion for the Recognition of Human GaitsabstractWe propose a hybrid dynamical model of human motion and develop a classification algorithm for the purpose of analysis and recognition. We assume that some temporal statistics are extracted from the images, and use them to infer a dynamical model that explicitly represents ground contact events. Such events correspond to “switches” between symmetric sets of hidden parameters in an auto-regressive model. We propose novel algorithms to estimate switches and model parameters, and develop a distance between such models that explicitly factors out exogenous inputs that are not unique to an individual or his/her gait. We show that such a distance is more discriminative than the distance between simple linear systems for the task of gait recognition. Alessandro Bissacco, Stefano Soatto |
Int. J. Comput. Vis. | 2 |
| 2008 | The scale of a texture and its application to segmentationabstractThis paper examines the issue of scale in modeling texture for the purpose of segmentation. We propose a scale descriptor for texture and an energy minimization model to find the scale of a given texture at each location. For each pixel, we use the intensity distribution in a local patch around that pixel to determine the smallest size of the domain that can be used to generate neighboring patches. The energy functional we propose to minimize is comprised of three terms: The first is the dissimilarity measure using the Wasserstein distance or Kullback-Leibler divergence between neighboring patch distributions; the second maximizes the entropy of the local patch, and the third penalizes larger size at equal fidelity. Our experiments show the proposed scale model successfully captures the intrinsic scale of texture at each location. We also apply our scale descriptor for improving texture segmentation based on histogram matching [15]. Byung-Woo Hong, Stefano Soatto, Kangyu Ni, Tony F. Chan |
CVPR | 2 |
| 2008 | Edge descriptors for robust wide-baseline correspondenceabstractThis paper describes a method for finding wide-baseline correspondences between images at locations along gradient edges. We find edges in scale space using established methods and develop invariant descriptors for these edges based on orientation and scale histograms. Because edges are often found on occluding boundaries, we calculate and store two descriptors per edge, one on each side, for robustness to occlusions. We demonstrate the effectiveness of edge matching in the applications of wide-baseline correspondence, structure from motion from line segments, and object category recognition on the Caltech 101 dataset. Jason Meltzer, Stefano Soatto |
CVPR | 2 |
| 2008 | Joint data alignment up to (lossy) transformationsabstractJoint data alignment is often regarded as a data simplification process. This idea is powerful and general, but raises two delicate issues. First, one must make sure that the use full information about the data is preserved by the alignment process. This is especially important when data are affected by non-invertible transformations, such as those originating from continuous domain deformations in a discrete image lattice. We propose a formulation that explicitly avoids this pitfall. Second, one must choose an appropriate measure of data complexity. We show that standard concepts such as entropy might not be optimal for the task, and we propose alternative measures that reflect the regularity of the codebook space. We also propose a novel and efficient algorithm that allows joint alignment of a large number of samples (tens of thousands of image patches), and does not rely on the assumption that pixels are independent. This is done for the case where the data is postulated to live in an affine subspaces of the embedding space of the raw data. We apply our scheme to learn sparse bases for natural images that discount domain deformations and hence significantly decrease the complexity of codebooks while maintaining the same generative power. Andrea Vedaldi, Gregorio Guidi, Stefano Soatto |
CVPR | 3 |
| 2008 | Relaxed matching kernels for robust image comparisonabstractThe popular bag-of-features representation for object recognition collects signatures of local image patches and discards spatial information. Some have recently attempted to at least partially overcome this limitation, for instance by ldquospatial pyramidsrdquo and ldquoproximityrdquo kernels. We introduce the general formalism of ldquorelaxed matching kernelsrdquo (RMKs) that includes such approaches as special cases, allow us to derive useful general properties of these kernels, and to introduce new ones. As an example, we introduce a kernel based on matching graphs of features and one based on matching information-compressed features. We show that all RMKs are competitive and outperform in several cases recently published state-of-the-art results on standard datasets. However, we also show that a proper implementation of a baseline bag-of-features algorithm can be extremely competitive, and outperform the other methods in some cases. Andrea Vedaldi, Stefano Soatto |
CVPR | 2 |
| 2008 | Filtering Internet image search results towards keyword based category recognitionabstractIn this work we aim to capitalize on the availability of Internet image search engines to automatically create image training sets from user provided queries. This problem is particularly difficult due to the low precision of image search results. Unlike many existing dataset gathering approaches, we do not assume a category model based on a small subset of the noisy data or an ad-hoc validation set. Instead we use a nonparametric measure of strangeness [8] in the space of holistic image representations, and perform an iterative feature elimination algorithm to remove the most strange examples from the category. This is the equivalent of keeping only features that are found to be consistent with others in the class. We show that applying our method to image search data before training improves average recognition performance, and demonstrate that we obtain comparative precision and recall results to the current state of the art, all the while maintaining a significantly simpler approach. In the process we also extend the strangeness-based feature elimination algorithm to automatically select good threshold values and perform filtering of a single class when the background is given. Kamil Wnuk, Stefano Soatto |
CVPR | 2 |
| 2008 | Localizing Objects with Smart Dictionaries
Brian Fulkerson, Andrea Vedaldi, Stefano Soatto |
ECCV (1) | 3 |
| 2008 | Background Subtraction on Distributions
Teresa Ko, Stefano Soatto, Deborah Estrin |
ECCV (3) | 2 |
| 2008 | Relevant Feature Selection for Human Pose Estimation and Localization in Cluttered Images
Ryuzo Okada, Stefano Soatto |
ECCV (2) | 2 |
| 2008 | Quick Shift and Kernel Methods for Mode Seeking
Andrea Vedaldi, Stefano Soatto |
ECCV (4) | 2 |
| 2008 | Dynamic Shape and Appearance Modeling via Moving and Deforming Layers
Jeremy D. Jackson, Anthony J. Yezzi, Stefano Soatto |
Int. J. Comput. Vis. | 3 |
| 2008 | 3-D Reconstruction of Shaded Objects from Multiple Images Under Unknown Illumination
Hailin Jin, Daniel Cremers, Emmanuel Prados, Anthony J. Yezzi, Stefano Soatto |
Int. J. Comput. Vis. | 6 |
| 2008 | Editorial
Cordelia Schmid, Stefano Soatto, Carlo Tomasi |
Int. J. Comput. Vis. | 2 |
| 2008 | Shape from Defocus via DiffusionabstractDefocus can be modeled as a diffusion process and represented mathematically using the heat equation, where image blur corresponds to the diffusion of heat. This analogy can be extended to non-planar scenes by allowing a space-varying diffusion coefficient. The inverse problem of reconstructing 3-D structure from blurred images corresponds to an "inverse diffusion" that is notoriously ill-posed. We show how to bypass this problem by using the notion of relative blur. Given two images, within each neighborhood, the amount of diffusion necessary to transform the sharper image into the blurrier one depends on the depth of the scene. This can be used to devise a global algorithm to estimate the depth profile of the scene without recovering the deblurred image, using only forward diffusion. Paolo Favaro, Stefano Soatto, Martin Burger 0001, Stanley J. Osher |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | On the Blind Classification of Time SeriesabstractWe propose a cord distance in the space of dynamical models that takes into account their dynamics, including transients, output maps and input distributions. In data analysis applications, as opposed to control, the input is often not known and is inferred as part of the (blind) identification. So it is an integral part of the model that should be considered when comparing different time series. Previous work on kernel distances between dynamical models assumed either identical or independent inputs. We extend it to arbitrary distributions, highlighting connections with system identification, independent component analysis, and optimal transport. The increased modeling power is demonstrated empirically on gait classification from simple visual features. Alessandro Bissacco, Stefano Soatto |
CVPR | 2 |
| 2007 | Fast Human Pose Estimation using Appearance and Motion via Multi-Dimensional Boosting RegressionabstractWe address the problem of estimating human pose in video sequences, where rough location has been determined. We exploit both appearance and motion information by defining suitable features of an image and its temporal neighbors, and learning a regression map to the parameters of a model of the human body using boosting techniques. Our algorithm can be viewed as a fast initialization step for human body trackers, or as a tracker itself. We extend gradient boosting techniques to learn a multi-dimensional map from (rotated and scaled) Haar features to the entire set of joint angles representing the full body pose. We test our approach by learning a map from image patches to body joint angles from synchronized video and motion capture walking data. We show how our technique enables learning an efficient real-time pose estimator, validated on publicly available datasets. Alessandro Bissacco, Ming-Hsuan Yang 0001, Stefano Soatto |
CVPR | 3 |
| 2007 | Joint Priors for Variational Shape and Appearance ModelingabstractWe are interested in modeling the variability of different images of the same scene, or class of objects, obtained by changing the imaging conditions, for instance the viewpoint or the illumination. Understanding of such a variability is key to reconstruction of objects despite changes in their appearance (e.g. due to non-Lambertian reflection), or to recognizing classes of objects (e.g. cars), or individual objects seen from different vantage points. We propose a model that can account for changes in shape or viewpoint, appearance, and also occlusions of line of sight. We learn a prior model of each factor (shape, motion and appearance) from a collection of samples using principal component analysis, akin a generalization of "active appearance models" to dense domains affected by occlusions. The ultimate goal of this work is stereo reconstruction in 3D, but first we have developed the first stage in this approach by addressing the simpler case of 2D shape/radiance detection in single images. We illustrate our model on a collection of images of different cars and show how the learned prior can be used to improve segmentation and 3D stereo reconstruction. Jeremy D. Jackson, Anthony J. Yezzi, Stefano Soatto |
CVPR | 3 |
| 2007 | Autocalibration and Uncalibrated Reconstruction of Shape from DefocusabstractMost algorithms for reconstructing shape from defocus assume that the images are obtained with a camera that has been previously calibrated so that the aperture, focal plane, and focal length are known. In this manuscript we characterize the set of scenes that can be reconstructed from defocused images regardless of calibration parameters. In lack of knowledge about the camera or about the scene, reconstruction is possible only up to an equivalence class that is described analytically. When weak knowledge about the scene is available, however, we show how it can be exploited in order to auto-calibrate the imaging device. This includes imaging a slanted plane or generic assumptions on the restoration of the deblurred images. Yifei Lou, Paolo Favaro, Andrea L. Bertozzi, Stefano Soatto |
CVPR | 4 |
| 2007 | Moving Forward in Structure From MotionabstractIt is well-known that forward motion induces a large number of local minima in the instantaneous least-squares reprojection error. This is caused in part by singularities in the error landscape around the forward direction, and presents a challenge in using existing algorithms for structure-from-motion in autonomous navigation applications. In this paper we prove that imposing a bound on the reconstructed depth of the scene makes the least-squares re-projection error continuous. This has implications for autonomous navigation, as it suggests simple modifications for existing algorithms to minimize the effects of local minima in forward translation. Andrea Vedaldi, Gregorio Guidi, Stefano Soatto |
CVPR | 3 |
| 2007 | Proximity Distribution Kernels for Geometric Context in Category RecognitionabstractWe propose using the proximity distribution of vector- quantized local feature descriptors for object and category recognition. To this end, we introduce a novel "proximity distribution kernel" that naturally combines local geometric as well as photometric information from images. It satisfies Mercer's condition and can therefore be readily combined with a support vector machine to perform visual categorization in a way that is insensitive to photometric and geometric variations, while retaining significant discriminative power. In particular, it improves on the results obtained both with geometrically unconstrained "bags of features" approaches, as well as with over-constrained "affine procrustes." Indeed, we test this approach on several challenging data sets, including Graz-01, Graz-02, and the PASCAL challenge. We registered the average performance at 91.5% on Graz-01, 82.7% on Graz-02, and 74.5% on PASCAL. Our approach is designed to enforce and exploit geometric consistency among objects in the same category; therefore, it does not improve the performance of existing algorithms on datasets where the data is already roughly aligned and scaled. Our method has the potential to be extended to more complex geometric relationships among local features, as we illustrate in the experiments. Haibin Ling, Stefano Soatto |
ICCV | 2 |
| 2007 | A Study of Face Recognition as People AgeabstractIn this paper we study face recognition across ages within a real passport photo verification task. First, we propose using the gradient orientation pyramid for this task. Discarding the gradient magnitude and utilizing hierarchical techniques, we found that the new descriptor yields a robust and discriminative representation. With the proposed descriptor, we model face verification as a two-class problem and use a support vector machine as a classifier. The approach is applied to two passport data sets containing more than 1,800 image pairs from each person with large age differences. Although simple, our approach outperforms previously tested Bayesian technique and other descriptors, including the intensity difference and gradient with magnitude. In addition, it works as well as two commercial systems. Second, for the first time, we empirically study how age differences affect recognition performance. Our experiments show that, although the aging process adds difficulty to the recognition task, it does not surpass illumination or expression as a confounding factor. Haibin Ling, Stefano Soatto, Narayanan Ramanathan, David Jacobs 0001 |
ICCV | 2 |
| 2007 | Correspondence Transfer for the Registration of Multimodal ImagesabstractGene expression data provide information on the location where certain genes are active; in order for this to be useful, such a location must be registered to an anatomical atlas. Because gene expression maps are considerably different from each other - they display the expression of different genes - and from the anatomical atlas, this problem is currently addressed either manually by trained experts, or by neglecting all image information and only using the pre-segmented boundaries. In this manuscript we concentrate on data discrepancy measures that take into account image information when this is present in both the target and template images. We exploit such "bi-lateral" structures to drive the correspondence process in regions where the intensity information is inconsistent, analogously to a "motion in- painting" task. Although no ground truth can be established, and prior information clearly plays a key role, we show that our model achieves desirable results on subjective tests validated by expert subjects. Zhao Yi, Stefano Soatto |
ICCV | 2 |
| 2007 | Classification and Recognition of Dynamical Models: The Role of Phase, Independent Components, Kernels and Optimal TransportabstractWe address the problem of performing decision tasks, and in particular classification and recognition, in the space of dynamical models in order to compare time series of data. Motivated by the application of recognition of human motion in image sequences, we consider a class of models that include linear dynamics, both stable and marginally stable (periodic), both minimum and non-minimum phase, driven by non-Gaussian processes. This requires extending existing learning and system identification algorithms to handle periodic modes and nonminimum phase behavior, while taking into account higher-order statistics of the data. Once a model is identified, we define a kernel-based cord distance between models that includes their dynamics, their initial conditions as well as input distribution. This is made possible by a novel kernel defined between two arbitrary (non-Gaussian) distributions, which is computed by efficiently solving an optimal transport problem. We validate our choice of models, inference algorithm, and distance on the tasks of human motion synthesis (sample paths of the learned models), and recognition (nearest-neighbor classification in the computed distance). However, our work can be applied more broadly where one needs to compare historical data while taking into account periodic trends, non-minimum phase behavior, and non-Gaussian input distributions. Alessandro Bissacco, Alessandro Chiuso, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | A Variational Approach to Problems in Calibration of Multiple CamerasabstractThis paper addresses the problem of calibrating camera parameters using variational methods. One problem addressed is the severe lens distortion in low-cost cameras. For many computer vision algorithms aiming at reconstructing reliable representations of 3D scenes, the camera distortion effects will lead to inaccurate 3D reconstructions and geometrical measurements if not accounted for. A second problem is the color calibration problem caused by variations in camera responses that result in different color measurements and affects the algorithms that depend on these measurements. We also address the extrinsic camera calibration that estimates relative poses and orientations of multiple cameras in the system and the intrinsic camera calibration that estimates focal lengths and the skew parameters of the cameras. To address these calibration problems, we present multiview stereo techniques based on variational methods that utilize partial and ordinary differential equations. Our approach can also be considered as a coordinated refinement of camera calibration parameters. To reduce computational complexity of such algorithms, we utilize prior knowledge on the calibration object, making a piecewise smooth surface assumption, and evolve the pose, orientation, and scale parameters of such a 3D model object without requiring a 2D feature extraction from camera views. We derive the evolution equations for the distortion coefficients, the color calibration parameters, the extrinsic and intrinsic parameters of the cameras, and present experimental results. Gozde Unal, Anthony J. Yezzi, Stefano Soatto, Gregory Slabaugh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | Classifying Human Dynamics Without Contact ForcesabstractWe develop a classification algorithm for hybrid autoregressive models of human motion for the purpose of videobased analysis and recognition. We assume that some temporal statistics are extracted from the images, and we use them to infer a dynamical system that explicitly models contact forces. We then develop a distance between such models that explicitly factors out exogenous inputs that are not unique to an individual or her gait. We show that such a distance is more discriminative than the distance between simple linear systems, where most of the energy is devoted to modeling the dynamics of spurious nuisances such as contact forces. Alessandro Bissacco, Stefano Soatto |
CVPR (2) | 2 |
| 2006 | Shape Representation based on Integral Kernels: Application to Image Matching and SegmentationabstractThis paper presents a shape representation and a variational framework for the construction of diffeomorphisms that establish "meaningful"correspondences between images, in that they preserve the local geometry of singularities such as region boundaries. At the same time, the shape representation allows enforcing shape information locally in determining such region boundaries. Our representation is based on a kernel descriptor that characterizes local shape. This shape descriptor is robust to noise and forms a scale-space in which an appropriate scale can be chosen depending on the size of features of interest in the scene. In order to preserve local shape during the matching procedure, we introduce a novel constraint to traditional energybased approaches to estimate diffeomorphic deformations, and enforce it in a variational framework. Byung-Woo Hong, Emmanuel Prados, Stefano Soatto, Luminita A. Vese |
CVPR (1) | 3 |
| 2006 | Control Theory and Fast Marching Techniques for Brain Connectivity MappingabstractWe propose a novel, fast and robust technique for the computation of anatomical connectivity in the brain. Our approach exploits the information provided by Diffusion Tensor Magnetic Resonance Imaging (or DTI) and models the white matter by using Riemannian geometry and control theory. We show that it is possible, from a region of interest, to compute the geodesic distance to any other point and the associated optimal vector field. The latter can be used to trace shortest paths coinciding with neural fiber bundles. We also demonstrate that no explicit computation of those 3D curves is necessary to assess the degree of connectivity of the region of interest with the rest of the brain. We finally introduce a general local connectivity measure whose statistics along the optimal paths may be used to evaluate the degree of connectivity of any pair of voxels. All those quantities can be computed simultaneously in a Fast Marching framework, directly yielding the connectivity maps. Apart from being extremely fast, this method has other advantages such as the strict respect of the convoluted geometry of white matter, the fact that it is parameter-free, and its robustness to noise. We illustrate our technique by showing results on real and synthetic datasets. OurGCM(Geodesic Connectivity Mapping) algorithm is implemented in C++ and will be soon available on the web. Emmanuel Prados, Stefano Soatto, Christophe Lenglet, Jean-Philippe Pons, Nicolas Wotawa, Rachid Deriche, Olivier D. Faugeras |
CVPR (1) | 2 |
| 2006 | Local Features, All Grown UpabstractWe present a technique to adapt the domain of local features through the matching process to augment their discriminative power. We start with local affine features selected and normalized independently in training and test images, and jointly expand their domain as part of the correspondence process, akin to a (non-rigid) registration task that yields a (multi-view) segmentation of the object of interest from clutter, including the detection of occlusions. We show how our growth process can be used to validate putative affine matches, to match a given "template" (an image of an object without clutter) to a cluttered and partially occluded image, and to match two images that contain the same unknown object in different clutter under different occlusions (unsupervised object detection). Andrea Vedaldi, Stefano Soatto |
CVPR (2) | 2 |
| 2006 | Viewpoint Induced Deformation Statistics and the Design of Viewpoint Invariant Features: Singularities and Occlusions
Andrea Vedaldi, Stefano Soatto |
ECCV (2) | 2 |
| 2006 | High Performance Feature Detection on a Reconfigurable Co-ProcessorabstractIn this paper, the authors propose a new design for feature detection used for tracking, which eliminates the need of a central computer to complete computations for the feature selection algorithm. Such a system constrains performance due to the delay in which data is transferred from camera to computer for processing. Our design suggests that feature detection computation can be done on a processor within the camera helping to reduce overall computation time for detection and increase performance for overall tracking system. However, these systems are often constrained by the processing power available to the camera. But with Benedetti and Perona's approach to Tomasi and Kanade's detection algorithm, such a design is possible to implement onto a camera system which would eliminate the delay and also improve performance over a tracking system designed on software Jia Ming Mar, Alessandro Bissacco, Stefano Soatto, Soheil Ghiasi |
FCCM | 3 |
| 2006 | Detecting Humans via Their PoseabstractWe consider the problem of detecting humans and classifying their pose from a single image. Specifically, our goal is to devise a statistical model that simultaneously answers two questions: 1) is there a human in the image? and, if so, 2) what is a low-dimensional representation of her pose? We investigate models that can be learned in an unsupervised manner on unlabeled images of human poses, and provide information that can be used to match the pose of a new image to the ones present in the training set. Starting from a set of descriptors recently proposed for human detection, we apply the Latent Dirichlet Allocation framework to model the statistics of these features, and use the resulting model to answer the above questions. We show how our model can efficiently describe the space of images of humans with their pose, by providing an effective representation of poses for tasks such as classification and matching, while performing remarkably well in human/non human decision problems, thus enabling its use for human detection. We validate the model with extensive quantitative experiments and comparisons with other approaches on human detection and pose matching. Alessandro Bissacco, Ming-Hsuan Yang 0001, Stefano Soatto |
NIPS | 3 |
| 2006 | A Complexity-Distortion Approach to Joint Pattern AlignmentabstractImage Congealing (IC) is a non-parametric method for the joint alignment of a col- lection of images affected by systematic and unwanted deformations. The method attempts to undo the deformations by minimizing a measure of complexity of the image ensemble, such as the averaged per-pixel entropy. This enables alignment without an explicit model of the aligned dataset as required by other methods (e.g. transformed component analysis). While IC is simple and general, it may intro- duce degenerate solutions when the transformations allow minimizing the com- plexity of the data by collapsing them to a constant. Such solutions need to be explicitly removed by regularization. In this paper we propose an alternative formulation which solves this regulariza- tion issue on a more principled ground. We make the simple observation that alignment should simplify the data while preserving the useful information car- ried by them. Therefore we trade off fidelity and complexity of the aligned en- semble rather than minimizing the complexity alone. This eliminates the need for an explicit regularization of the transformations, and has a number of other useful properties such as noise suppression. We show the modeling and computa- tional benefits of the approach to the some of the problems on which IC has been demonstrated. Andrea Vedaldi, Stefano Soatto |
NIPS | 2 |
| 2006 | Kernel Density Estimation and Intrinsic Alignment for Shape Priors in Level Set Segmentation
Daniel Cremers, Stanley J. Osher, Stefano Soatto |
Int. J. Comput. Vis. | 3 |
| 2006 | Two-View Multibody Structure from Motion
René Vidal, Yi Ma 0001, Stefano Soatto, S. Shankar Sastry |
Int. J. Comput. Vis. | 3 |
| 2006 | Region matching with missing parts
Alessandro Duci, Anthony J. Yezzi, Sanjoy K. Mitter, Stefano Soatto |
Image Vis. Comput. | 4 |
| 2006 | Dynamic Shape and Appearance ModelsabstractWe propose a model of the joint variation of shape and appearance of portions of an image sequence. The model is conditionally linear, and can be thought of as an extension of active appearance models to exploit the temporal correlation of adjacent image frames. Inference of the model parameters can be performed efficiently using established numerical optimization techniques borrowed from finite-element analysis and system identification techniques. Gianfranco Doretto, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Integral Invariants for Shape MatchingabstractFor shapes represented as closed planar contours, we introduce a class of functionals which are invariant with respect to the Euclidean group and which are obtained by performing integral operations. While such integral invariants enjoy some of the desirable properties of their differential counterparts, such as locality of computation (which allows matching under occlusions) and uniqueness of representation (asymptotically), they do not exhibit the noise sensitivity associated with differential quantities and, therefore, do not require presmoothing of the input shape. Our formulation allows the analysis of shapes at multiple scales. Based on integral invariants, we define a notion of distance between shapes. The proposed distance measure can be computed efficiently and allows warping the shape boundaries onto each other; its computation results in optimal point correspondence as an intermediate step. Numerical results on shape matching demonstrate that this framework can match shapes despite the deformation of subparts, missing parts and noise. As a quantitative analysis, we report matching scores for shape retrieval from a database. Siddharth Manay, Daniel Cremers, Byung-Woo Hong, Anthony J. Yezzi, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2005 | Layered Active Appearance ModelsabstractActive appearance models (AAMs) provide a framework for modeling the joint shape and texture of an image. An AAM is a compact representation of both factors in a conditionally linear model. However, the standard AAM framework does not handle images which have missing features, or allow modification of certain structures in the image while leaving neighboring ones undeformed. We introduce the layered active appearance model (LAAM), which allows for missing features, occlusion, substantial spatial rearrangement of features, and which provides a more general representation that extends the applicability of the active appearance model Eagle Jones, Stefano Soatto |
ICCV | 2 |
| 2005 | KALMANSAC: Robust Filtering by ConsensusabstractWe propose an algorithm to perform causal inference of the state of a dynamical model when the measurements are corrupted by outliers. While the optimal (maximum-likelihood) solution has doubly exponential complexity due to the combinatorial explosion of possible choices of inliers, we exploit the structure of the problem to design a sampling-based algorithm that has constant complexity. We derive our algorithm from the equations of the optimal filter, which makes our approximation explicit. Our work is motivated by real-time tracking and the estimation of structure from motion (SFM). We test our algorithm for on-line outlier rejection both for tracking and for SFM. We show that our approach can tolerate a large proportion of outliers, whereas previous causal robust statistical inference methods failed with less than half as many. Our work can be thought of as the extension of random sample consensus algorithms to dynamic data, or as the implementation of pseudo-Bayesian filtering algorithms in a sampling framework. Andrea Vedaldi, Hailin Jin, Paolo Favaro, Stefano Soatto |
ICCV | 4 |
| 2005 | Features for Recognition: Viewpoint Invariance for Non-Planar ScenesabstractMost current local feature detectors/descriptors implicitly assume that the scene is (locally) planar, an assumption that is violated at surface discontinuities. We show that this restriction is, at least in theory, unnecessary, as one can construct local features that are viewpoint-invariant for generic non-planar scenes. However, we show that any such feature necessarily sacrifices shape information, in the sense of being non shape-discriminative. Finally, we show that if viewpoint is factored out as part of the matching process, rather than explicitly in the representation, then shape is discriminative indeed. We illustrate our theoretical results empirically by showing that, even for simple scenes, current affine descriptors fail where even a naive 3-D viewpoint invariant succeeds in matching Andrea Vedaldi, Stefano Soatto |
ICCV | 2 |
| 2005 | Multi-View Stereo Reconstruction of Dense Shape and Complex Appearance
Hailin Jin, Stefano Soatto, Anthony J. Yezzi |
Int. J. Comput. Vis. | 2 |
| 2005 | A Geometric Approach to Shape from DefocusabstractWe introduce a novel approach to shape from defocus, i.e., the problem of inferring the three-dimensional (3D) geometry of a scene from a collection of defocused images. Typically, in shape from defocus, the task of extracting geometry also requires deblurring the given images. A common approach to bypass this task relies on approximating the scene locally by a plane parallel to the image (the so-called equifocal assumption). We show that this approximation is indeed not necessary, as one can estimate 3D geometry while avoiding deblurring without strong assumptions on the scene. Solving the problem of shape from defocus requires modeling how light interacts with the optics before reaching the imaging surface. This interaction is described by the so-called point spread function (PSF). When the form of the PSF is known, we propose an optimal method to infer 3D geometry from defocused images that involves computing orthogonal operators which are regularized via functional singular value decomposition. When the form of the PSF is unknown, we propose a simple and efficient method that first learns a set of projection operators from blurred images and then uses these operators to estimate the 3D geometry of the scene from novel blurred images. Our experiments on both real and synthetic images show that the performance of the algorithm is relatively insensitive to the form of the PSF. Our general approach is to minimize the Euclidean norm of the difference between the estimated images and the observed images. The method is geometric in that we reduce the minimization to performing projections onto linear subspaces, by using inner product structures on both infinite and finite-dimensional Hilbert spaces. Both proposed algorithms involve only simple matrix-vector multiplications which can be implemented in real-time. Paolo Favaro, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | A Variational Approach to Scene Reconstruction and Image Segmentation from Motion-Blur Cues
Paolo Favaro, Stefano Soatto |
CVPR (1) | 2 |
| 2004 | Shedding Light on Stereoscopic Segmentation
Hailin Jin, Daniel Cremers, Anthony J. Yezzi, Stefano Soatto |
CVPR (1) | 4 |
| 2004 | Spatially Homogeneous Dynamic Textures
Gianfranco Doretto, Eagle Jones, Stefano Soatto |
ECCV (2) | 3 |
| 2004 | Scene and Motion Reconstruction from Defocused and Motion-Blurred Images via Anisotropic Diffusion
Paolo Favaro, Martin Burger 0001, Stefano Soatto |
ECCV (1) | 3 |
| 2004 | Region-Based Segmentation on Evolving Surfaces with Application to 3D Reconstruction of Shape and Piecewise Constant Radiance
Hailin Jin, Anthony J. Yezzi, Stefano Soatto |
ECCV (2) | 3 |
| 2004 | Integral Invariant Signatures
Siddharth Manay, Byung-Woo Hong, Anthony J. Yezzi, Stefano Soatto |
ECCV (4) | 4 |
| 2004 | Multiple View Feature Descriptors from Image Sequences via Kernel Principal Component Analysis
Jason Meltzer, Ming-Hsuan Yang 0001, Rakesh Gupta 0001, Stefano Soatto |
ECCV (1) | 4 |
| 2004 | Modeling and Synthesis of Facial Motion Driven by Speech
Payam Saisan, Alessandro Bissacco, Alessandro Chiuso, Stefano Soatto |
ECCV (3) | 4 |
| 2004 | Simultaneous localization and mapping using multiple view feature descriptorsabstractWe propose a vision-based SLAM algorithm incorporating feature descriptors derived from multiple views of a scene, incorporating illumination and viewpoint variations. These descriptors are extracted from video and then applied to the challenging task of wide baseline matching across significant viewpoint changes. The system incorporates a single camera on a mobile robot in an extended Kalman filter framework to develop a 3D map of the environment and determine egomotion. At the same time, the feature descriptors are generated from the video sequence, which can be used to localize the robot when it returns to a mapped location. The kidnapped robot problem is addressed by matching descriptors without any estimate of position, then determining the epipolar geometry with respect to a known position in the map. Jason Meltzer, Rakesh Gupta 0001, Ming-Hsuan Yang 0001, Stefano Soatto |
IROS | 4 |
| 2004 | Motion Competition: A Variational Approach to Piecewise Parametric Motion Segmentation
Daniel Cremers, Stefano Soatto |
Int. J. Comput. Vis. | 2 |
| 2003 | Editable Dynamic TexturesabstractWe present a simple and efficient algorithm for modifying the temporal behavior of "dynamic textures," i.e. sequences of images that exhibit some form of temporal regularity, such as flowing water, steam, smoke, flames, foliage of trees in wind. The main goal is to design algorithms for synthesizing and editing realistic sequences of images of dynamic scenes that exhibit some form of temporal stationarity. This is an image-based rendering task, and in particular we are interested in synthesizing the temporal behavior of the scene. Gianfranco Doretto, Stefano Soatto |
CVPR (2) | 2 |
| 2003 | 3D Shape from Anisotropic DiffusionabstractWe cast the problem of inferring the 3D shape of a scene from a collection of defocused images in the framework of anisotropic diffusion. We propose an algorithm that can estimate the shape of a scene by inferring the diffusion coefficient of a heat equation. The method is optimal, as we pose it as the minimization of a certain cost functional based on the input images, and fast. Furthermore, we also extend our algorithm to the case of multiple images, and derive a 3D scene segmentation algorithm that can work in the presence of pictorial camouflage. Paolo Favaro, Stanley J. Osher, Stefano Soatto, Luminita A. Vese |
CVPR (1) | 3 |
| 2003 | Seeing Beyond Occlusions (and other marvels of a finite lens aperture)abstractWe present a novel algorithm to reconstruct the geometry and photometry of a scene with occlusions from a collection of defocused images. The presence of a finite lens aperture allows us to recover portions of the scene that would be occluded in a pin-hole projection, thus "uncovering" the occlusion. We estimate the shape of each object (a surface, including the occluding boundaries), and its radiance (a positive function defined on the surface, including portions that are occluded by other objects). Paolo Favaro, Stefano Soatto |
CVPR (2) | 2 |
| 2003 | Multi-view Stereo Beyond LambertabstractWe consider the problem of estimating the shape and radiance of an object from a calibrated set of views under the assumption that the reflectance of the object is non-Lambertian. Unlike traditional stereo, we do not solve the correspondence problem by comparing image-to-image. Instead, we exploit a rank constraint on the radiance tensor field of the surface in space, and use it to define a discrepancy measure between each image and the underlying model. Our approach automatically returns an estimate of the radiance of the scene, along with its shape, represented by a dense surface. The former can be used to generate novel views that capture the non-Lambertian appearance of the scene. Hailin Jin, Stefano Soatto, Anthony J. Yezzi |
CVPR (1) | 2 |
| 2003 | Structure From Motion for Scenes Without FeaturesabstractWe describe an algorithm for reconstructing the 3D (three-dimensional) shape of the scene and the relative pose of a number of cameras from a collection of images under the assumption that the scene does not contain photometrically distinct "features". We work under the explicit assumption that the scene is made of a number of smooth surfaces that radiate constant energy isotropically in all directions, and setup a region-based cost functional that we minimize using local gradient flow techniques. Anthony J. Yezzi, Stefano Soatto |
CVPR (1) | 2 |
| 2003 | Variational Space-Time Motion SegmentationabstractWe propose a variational method for segmenting image sequences into spatiotemporal domains of homogeneous motion. To this end, we formulate the problem of motion estimation in the framework of Bayesian inference, using a prior which favors domain boundaries of minimal surface area. We derive a cost functional which depends on a surface in space-time separating a set of motion regions, as well as a set of vectors modeling the motion in each region. We propose a multiphase level set formulation of this functional, in which the surface and the motion regions are represented implicitly by a vector-valued level set function. Joint minimization of the proposed functional results in an eigenvalue problem for the motion model of each region and in a gradient descent evolution for the separating interface. Numerical results on real-world sequences demonstrate that minimization of a single cost functional generates a segmentation of space-time into multiple motion regions. Daniel Cremers, Stefano Soatto |
ICCV | 2 |
| 2003 | Dynamic Texture SegmentationabstractWe address the problem of segmenting a sequence of images of natural scenes into disjoint regions that are characterized by constant spatio-temporal statistics. We model the spatio-temporal dynamics in each region by Gauss-Markov models, and infer the model parameters as well as the boundary of the regions in a variational optimization framework. Numerical results demonstrate that - in contrast to purely texture-based segmentation schemes - our method is effective in segmenting regions that differ in their dynamics even when spatial statistics are identical. Gianfranco Doretto, Daniel Cremers, Paolo Favaro, Stefano Soatto |
ICCV | 4 |
| 2003 | Shape Representation via Harmonic EmbeddingabstractWe present a novel representation of shape for closed planar contours explicitly designed to possess a linear structure. This greatly simplifies linear operations such as averaging, principal component analysis or differentiation in the space of shapes. The representation relies upon embedding the contour on a subset of the space of harmonic functions of which the original contour is the zero level set. Alessandro Duci, Anthony J. Yezzi, Sanjoy K. Mitter, Stefano Soatto |
ICCV | 4 |
| 2003 | On Exploiting Occlusions in Multiple-view GeometryabstractOcclusions are commonplace in man-made and natural environments; they often result in photometric features where a line terminates at an occluding boundary, resembling a "T". We show that the 2-D motion of such T-junctions in multiple views carries nontrivial information on the 3-D structure of the scene and its motion relative to the camera. We show how the constraint among multiple views of T-junctions can be used to reliably detect them and differentiate them from ordinary point features. Finally, we propose an integrated algorithm to recursively and causally estimate structure and motion in the presence of T-junctions along with other point-features. Paolo Favaro, Alessandro Duci, Yi Ma 0001, Stefano Soatto |
ICCV | 4 |
| 2003 | Tales of Shape and Radiance in Multi-view StereoabstractTo what extent can three-dimensional shape and radiance be inferred from a collection of images? Can the two be estimated separately while retaining optimality? How should the optimality criterion be computed? When is it necessary to employ an explicit model of the reflectance properties of a scene? In this paper we introduce a separation principle for shape and radiance estimation that applies to Lambertian scenes and holds for any choice of norm. When the scene is not Lambertian, however, shape cannot be decoupled from radiance, and therefore matching image-to-image is not possible directly. We employ a rank constraint on the radiance tensor, which is commonly used in computer graphics, and construct a novel cost functional whose minimization leads to an estimate of both shape and radiance for nonLambertian objects, which we validate experimentally. Stefano Soatto, Anthony J. Yezzi, Hailin Jin |
ICCV | 1 |
| 2003 | Dynamic Textures
Gianfranco Doretto, Alessandro Chiuso, Ying Nian Wu, Stefano Soatto |
Int. J. Comput. Vis. | 4 |
| 2003 | Observing Shape from Defocused Images
Paolo Favaro, Andrea Mennucci, Stefano Soatto |
Int. J. Comput. Vis. | 3 |
| 2003 | Stereoscopic Segmentation
Anthony J. Yezzi, Stefano Soatto |
Int. J. Comput. Vis. | 2 |
| 2003 | Deformotion: Deforming Motion, Shape Average and the Joint Registration and Approximation of Structures in Images
Anthony J. Yezzi, Stefano Soatto |
Int. J. Comput. Vis. | 2 |
| 2003 | A semi-direct approach to structure from motion
Hailin Jin, Paolo Favaro, Stefano Soatto |
Vis. Comput. | 3 |
| 2002 | Region Matching with Missing Parts
Alessandro Duci, Anthony J. Yezzi, Sanjoy K. Mitter, Stefano Soatto |
ECCV (3) | 4 |
| 2002 | Learning Shape from Defocus
Paolo Favaro, Stefano Soatto |
ECCV (2) | 2 |
| 2002 | DEFORMOTION: Deforming Motion, Shape Average and the Joint Registration and Segmentation of Images
Stefano Soatto, Anthony J. Yezzi |
ECCV (3) | 1 |
| 2002 | Structure from Motion Causally Integrated Over TimeabstractWe describe an algorithm for reconstructing three-dimensional structure and motion causally, in real time from monocular sequences of images. We prove that the algorithm is minimal and stable, in the sense that the estimation error remains bounded with probability one throughout a sequence of arbitrary length. We discuss a scheme for handling occlusions (point features appearing and disappearing) and drift in the scale factor. These issues are crucial for the algorithm to operate in real time on real scenes. We describe in detail the implementation of the algorithm, which runs on a personal computer and has been made available to the community. We report the performance of our implementation on a few representative long sequences of real and synthetic images. The algorithm, which has been tested extensively over the course of the past few years, exhibits honest performance when the scene contains at least 20-40 points with high contrast, when the relative motion is "slow" compared to the sampling frequency of the frame grabber (30 Hz), and the lens aperture is "large enough" (typically more than 30/spl deg/ of visual field). Alessandro Chiuso, Paolo Favaro, Hailin Jin, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2001 | Recognition of Human GaitsabstractWe pose the problem of recognizing different types of human gait in the space of dynamical systems where each gait is represented Established techniques are employed to track a kinematic model of a human body in motion, and the trajectories of the parameters are used to learn a representation of a dynamical system, which defines a gait. Various types of distance between models are then computed These computations are non trivial due to the fact that, even for the case of linear systems, the space of canonical realizations is not linear. Alessandro Bissacco, Alessandro Chiuso, Yi Ma 0001, Stefano Soatto |
CVPR (2) | 4 |
| 2001 | Dynamic Texture RecognitionabstractDynamic textures are sequences of images that exhibit some form of temporal stationarity, such as waves, steam, and foliage. We pose the problem of recognizing and classifying dynamic textures in the space of dynamical systems where each dynamic texture is uniquely represented. Since the space is non-linear, a distance between models must be defined We examine three different distances in the space of autoregressive models and assess their power. Payam Saisan, Gianfranco Doretto, Ying Nian Wu, Stefano Soatto |
CVPR (2) | 4 |
| 2001 | Real-time Virtual Object InsertionabstractWe present a system to insert virtual objects into real image sequences in real time. The system consists of offthe- shelf hardware (a camera connected to a Pentium PC) and software to (a) automatically select and track region features despite changes in illumination, (b) estimate threedimensional position and orientation of surface patches relative to an inertial reference frame despite individual pointfeatures appearing and disappearing, (c) insert a texturemapped virtual object into the scene so as to make it appear to be part of the scene and moving with it. This is all done in real time. The multi-thread C++ code, which is readily interfaced with a frame grabber as well as Matlab for development, will be made available to the public at the demonstration. Paolo Favaro, Hailin Jin, Stefano Soatto |
ICCV | 3 |
| 2001 | Real-Time Feature Tracking and Outlier Rejection with Changes in Illumination
Hailin Jin, Paolo Favaro, Stefano Soatto |
ICCV | 3 |
| 2001 | Dynamic TexturesabstractDynamic textures are sequences of images of moving scenes that exhibit certain stationarity properties in time; these include sea-waves, smoke, foliage, whirlwind but also talking faces, traffic scenes etc. We present a novel characterization of dynamic textures that poses the problems of modelling, learning, recognizing and synthesizing dynamic textures on a firm analytical footing. We borrow tools from system identification to capture the "essence" of dynamic textures; we do so by learning (i.e. identifying) models that are optimal in the sense of maximum likelihood or minimum prediction error variance. For the special case of second-order stationary processes we identify the model in closed form. Once learned, a model has predictive power and can be used for extrapolating synthetic sequences to infinite length with negligible computational cost. We present experimental evidence that, within our framework, even low dimensional models can capture very complex visual phenomena. Stefano Soatto, Gianfranco Doretto, Ying Nian Wu |
ICCV | 1 |
| 2001 | Stereoscopic Segmentation
Anthony J. Yezzi, Stefano Soatto |
ICCV | 2 |
| 2000 | Real-Time 3-D Motion and Structure of Point-Features: A Front-End for Vision-Based Control and InteractionabstractWe present a system that consists of one camera connected to a personal computer that can (a) select and track a number of high-contrast point features on a sequence of images, (b) estimate their three-dimensional motion and position relative to an inertial reference frame, assuming rigidity, (c) handle occlusions that cause point-features to disappear as well as new features to appear. The system can also (d) perform partial self-calibration and (e) check for consistency of the rigidity assumption, although these features are not implemented in the current release. All of this is done automatically and in real-time (30 Hz) for 40-50 point features using commercial off-the-shelf hardware. The system is based on an algorithm presented by Chiuso et al. (2000), the properties of which have been analyzed by Chiuso and Soatto (2000). In particular, the algorithm is provably observable, provably minimal and provably stable- under suitable conditions. The core of the system, consisting of C++ code ready to interface with a frame grabber as well as Matlab code for development, is available at http://ee.wustl.edu/-soatto/research.html. We demonstrate the system by showing its use as (1) an ego-motion estimator, (2) an object tracker, and (3) an interactive input device, all without any modification of the system settings. Hailin Jin, Paolo Favaro, Stefano Soatto |
CVPR | 3 |
| 2000 | Stereoscopic Shading: Integrating Mult1Frame Shape Cues in a Variational FrameworkabstractWe address the problem of integrating multi-frame stereo and shading cues within the framework of optimization in the infinite-dimensional space of piecewise smooth surfaces. Cue integration then reduces to the determination of regions where prior assumptions on the reflectance of the surfaces can be enforced. By combining cues, our formulation allows defining a well-posed problem even when reconstruction from stereo or shading in isolation would be ill-posed. For a simplified model we prove the necessary conditions for optimality, and propose an iterative optimization algorithm, which we implement using ultra-narrowband level set methods. Hailin Jin, Stefano Soatto, Anthony J. Yezzi |
CVPR | 2 |
| 2000 | A Geometric Approach to Blind Deconvolution with Application to Shape from DefocuabstractWe propose a solution to the generic "bilinear calibration-estimation problem" when using a quadratic cost function and restricting to (locally) translation-invariant imaging models. We apply the solution to the problem of reconstructing the three-dimensional shape and radiance of a scene from a number of defocused images. Since the imaging process maps the continuum of three-dimensional space onto the discrete pixel grid, rather than discretizing the continuum we exploit the structure of maps between (finite-and infinite-dimensional) Hilbert spaces and arrive at a principled algorithm that does not involve any choice of basis or discretization. Rather, these are uniquely determined by the data, and exploited in a functional singular value decomposition in order to obtain a regularized solution. Stefano Soatto, Paolo Favaro |
CVPR | 1 |
| 2000 | 3-D Motion and Structure from 2-D Motion Causally Integrated over Time: Implementation
Alessandro Chiuso, Paolo Favaro, Hailin Jin, Stefano Soatto |
ECCV (2) | 4 |
| 2000 | Shape and Radiance Estimation from the Information-Divergence of Blurred Images
Paolo Favaro, Stefano Soatto |
ECCV (1) | 2 |
| 2000 | Optimal Structure from Motion: Local Ambiguities and Global Estimates
Alessandro Chiuso, Roger W. Brockett, Stefano Soatto |
Int. J. Comput. Vis. | 3 |
| 2000 | Euclidean Reconstruction and Reprojection Up to Subgroups
Yi Ma 0001, Stefano Soatto, Jana Kosecka, S. Shankar Sastry |
Int. J. Comput. Vis. | 2 |
| 1999 | Euclidean Reconstruction and Reprojection up to SubgroupsabstractThe necessary and sufficient conditions for being able to estimate scene structure, motion and camera calibration from a sequence of images are very rarely satisfied in practice. What exactly can be estimated in sequences of practical importance, when such conditions are not satisfied? In this paper we give a complete answer to this question. For every camera motion that fails to meet the conditions, we give explicit formulas for the ambiguities in the reconstructed scene, motion and calibration. Such a characterization is crucial both for designing robust estimation algorithms (that do not try to recover parameters that cannot be recovered), and for generating novel views of the scene by controlling the vantage point. To this end, we characterize explicitly all the vantage points that give rise to a valid Euclidean reprojection regardless of the ambiguity in the reconstruction. We also characterize vantage points that generate views that are altogether invariant to the ambiguity. All the results are presented using simple notation that involves no tensors nor complex projective geometry, and should be accessible with basic background in linear algebra. Yi Ma 0001, Stefano Soatto, Jana Kosecka, S. Shankar Sastry |
ICCV | 2 |
| 1998 | Optimal Structure from Motion: Local Ambiguities and Global EstimatesabstractWe present an analysis of SFM from the point of view of noise. This analysis results in an algorithm that is provably convergent and provably optimal with respect to a chosen norm. In particular, we cast SFM as a nonlinear optimization problem and define a bilinear projection iteration that converges to fixed points of a certain cost-function. We then show that such fixed points are "fundamental", i.e. intrinsic to the problem of SFM and not an artifact introduced by our algorithm. We classify and characterize geometrically local extrema, and we argue that they correspond to phenomena observed in visual psychophysics. Finally, we show under what conditions it is possible-given convergence to a local extremum-to "jump" to the valley containing the optimum; this leads us to suggest a representation of the scene which is invariant with respect to such local extrema. Stefano Soatto, Roger W. Brockett |
CVPR | 1 |
| 1998 | Reducing "Structure From Motion": A General Framework for Dynamic Vision Part 1: ModelingabstractThe literature on recursive estimation of structure and motion from monocular image sequences comprises a large number of apparently unrelated models and estimation techniques. We propose a framework that allows us to derive and compare all models by following the idea of dynamical system reduction. The "natural" dynamic model, derived from the rigidity constraint and the projection model, is first reduced by explicitly decoupling structure (depth) from motion. Then, implicit decoupling techniques are explored, which consist of imposing that some function of the unknown parameters is held constant. By appropriately choosing such a function, not only can we account for models seen so far in the literature, but we can also derive novel ones. Stefano Soatto, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Reducing "Structure From Motion": A General Framework for Dynamic Vision Part 2: Implementation and Experimental AssessmentabstractFor pt.1 see ibid., p.933-42 (1998). A number of methods have been proposed in the literature for estimating scene-structure and ego-motion from a sequence of images using dynamical models. Despite the fact that all methods may be derived from a "natural" dynamical model within a unified framework, from an engineering perspective there are a number of trade-offs that lead to different strategies depending upon the applications and the goals one is targeting. We want to characterize and compare the properties of each model such that the engineer may choose the one best suited to the specific application. We analyze the properties of filters derived from each dynamical model under a variety of experimental conditions, assess the accuracy of the estimates, their robustness to measurement noise, sensitivity to initial conditions and visual angle, effects of the bas-relief ambiguity and occlusions, dependence upon the number of image measurements and their sampling rate. Stefano Soatto, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Correction to: "Reducing 'Structure From Motion': A General Framework for Dynamic Vision Part 2: Implementation and Experimental Assessment"
Stefano Soatto, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | Motion from fixationabstractWe study the problem of estimating rigid motion from a sequence of monocular perspective images obtained by navigating around an object while fixating a particular feature point. We cast the problem in the framework of "epipolar geometry", and propose a filter based upon implicit dynamical model for recursively estimating motion under the fixation constraint. This allows us to compare the quality of the estimates directly against the ones obtained assuming a general rigid motion simply by changing the geometry of the parameter space, while maintaining the same structure of the recursive estimator. We also present a closed-form static solution from two views, and a recursive estimator of the relative pose between the viewer and the scene. Stefano Soatto, Pietro Perona |
CVPR | 1 |
| 1996 | Reducing "structure from motion"abstractThe literature on recursive estimation of structure and motion from monocular image sequences comprises a large number of different models and estimation techniques. We propose a framework that allows us to derive and compare all models by following the idea of dynamical system reduction. The "natural" dynamic model, derived by the rigidity constraint and the perspective projection, is first reduced by explicitly decoupling structure (depth) from motion. Then implicit decoupling techniques are explored, which consist of imposing that some function of the unknown parameters is held constant. By appropriately choosing such a function, not only can we account for all models seen so far in the literature, but we can also derive novel ones. Casting all the different models in a common framework allows us to compare their geometric properties on common experimental grounds. Stefano Soatto, Pietro Perona |
CVPR | 1 |
| 1995 | Dynamic Rigid Motion Estimation from Weak Perspectiveabstract"Weak perspective" represents a simplified projection model that approximates the imaging process when the scene is viewed under a small viewing angle and its depth relief is small relative to its distance from the viewer. We study how to generate dynamic models for estimating rigid 3D motion from weak perspective. A crucial feature in dynamic visual motion estimation is to decouple structure from motion in the estimation model. The reasons are both geometric-to achieve global observability of the model-and practical, for a structure independent motion estimator allows us to deal with occlusions and appearance of new features in a principled way. It is also possible to push the decoupling even further, and isolate the motion parameters that are affected by the so called "bas relief ambiguity" from the ones that are not. We present a novel method for reducing the order of the estimator by decoupling portions of the state space from the time evolution of the measurement constraint. We use this method to construct an estimator of full rigid motion (modulo a scaling factor) on a six dimensional state space, an approximate estimator for a four dimensional subset of the motion space, and a reduced filter with only two states. The latter two are immune to the bas relief ambiguity. We compare strengths and weaknesses of each of the schemes on real and synthetic image sequences.> Stefano Soatto, Pietro Perona |
ICCV | 1 |
| 1995 | Visual motion estimation from point features: unified viewabstractAll methods for recursive estimation of 3-D motion from sequences of perspective images of point-features may be cast within a common framework. The unifying concept is the decoupling of the states of the dynamic observer that estimates motion and structure parameters. Two techniques are possible: explicit decoupling, following the principles of the "reduced-order observer", and implicit, via stabilization (or "compensation"). While we know how to calculate explicit decoupling for a limited number of state variables combinations, for instance using the "essential constraint" of Longuet-Higgins (1981) or the "subspace constraint" of Heeger and Jepson (1992), implicit decoupling is always possible by stabilizing an appropriate smooth function of the motion parameters. We describe some of the most "natural" choices, which consist in compensating for the image-motion of a point, a line or a plane. All the models we derive are in the form of implicit dynamical systems with parameters on different manifolds. Estimating motion may be regarded as the identification of such models, which may be carried out using general methods available in the literature. Stefano Soatto, Pietro Perona |
ICIP (3) | 1 |
| 1994 | Motion Estimation on the Essential Manifold
Stefano Soatto, Ruggero Frezza, Pietro Perona |
ECCV (2) | 1 |
| 1994 | Dynamic Visual Motion Estimation from Subspace ConstraintsabstractThe problem of estimating rigid motion from projections may be characterized using a nonlinear dynamical system, composed of the rigid motion constraint and the perspective map. The time derivative of the output of such a system, which is called the "motion field" and approximated by the "optical flow", is bilinear in the motion parameters, and may be used to specify a subspace constraint on either the direction of translation or the inverse depth of the observed points. Estimating motion may then be formulated as an optimization task constrained on such a subspace. We pose the optimization problem in a system theoretic framework as the the identification of a nonlinear implicit dynamical system with parameters on a differentiable manifold, and use techniques which pertain to nonlinear estimation and identification theory to perform the optimization task in a principled manner. The application of a general method presented in by Soatto et al. (see 33rd. IEEE conf. on Decision and Control, 1994) results in a recursive and pseudo-optimal solution of the visual motion estimation problem, which has robustness properties far superior to other existing techniques we have implemented. Experiments on real and synthetic image sequences show very promising results in terms of robustness, accuracy and computational efficiency.> Stefano Soatto, Pietro Perona |
ICIP (1) | 1 |
| 1994 | Recursive Estimation of Camera Motion from Uncalibrated Image SequencesabstractWe describe a method for estimating the motion and structure of a scene from a sequence of images taken with a camera whose geometric calibration parameters are unknown. The scheme is based upon a recursive motion estimation scheme, called the "essential filter", extended according to the epipolar geometric representation presented by Faugeras, Luong, and Maybank (see Proc. of the ECCV92, vol.588 of LNCS, Springer Verlag, 1992) in order to estimate the calibration parameters as well. The motion estimates can then be fed into any "structure from motion" module that processes motion error, in order to recover the structure of the scene.> Stefano Soatto, Pietro Perona |
ICIP (3) | 1 |
| 1993 | Recursive motion and structure estimation with complete error characterizationabstractAn algorithm that performs recursive estimation of ego-motion and ambient structure from a stream of monocular perspective images of a number of feature points is presented. The algorithm is based on an extended Kalman filter (EKF) that integrates over time the instantaneous motion and structure measurements computed by a two-perspective-views step. The key features of the authors' filter are: global observability of the model, and complete online characterization of the uncertainty of the measurements provided by the two-views step. The filter is thus guaranteed to be well-behaved regardless of the particular motion undergone by the observer. Regions of motion space that do not allow recovery of structure (e.g., pure rotation) may be crossed while maintaining good estimates of structure and motion. Whenever reliable measurements are available they are exploited. The algorithm works well for arbitrary motions with minimal smoothness assumptions and no ad hoc tuning. Simulations are presented that illustrate these characteristics.> Stefano Soatto, Pietro Perona, Ruggero Frezza, Giorgio Picci |
CVPR | 1 |