EDBT 2026 Demo / reviewers in the wild / expert
Pietro Perona
dblp:p/PietroPerona
· DBLP profile ↗
222ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0002-7583-5809ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 199 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 150 · 7 first-author · 22 since 2021Systems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-Based Action Recognition Generalizes to Untrained DomainsabstractHumans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs. third-person), and contexts (real life vs movies). Current deep learning models struggle with such generalization. We propose using features generated by a Vision Diffusion Model (VDM), aggregated via a transformer, to achieve human-like action recognition across these challenging conditions. We find that generalization is enhanced by the use of a model conditioned on earlier timesteps of the diffusion process to highlight semantic information over pixel level details in the extracted features. We experimentally explore the generalization properties of our approach in classifying actions across animal species, across different viewing angles, and different recording contexts. Our model sets a new state-of-the-art across all three generalization benchmarks, bringing machine action recognition closer to human-like robustness. Project page: vision.caltech.edu/actiondiff Code: github.com/frankyaoxiao/ActionDiff Rogério Guimarães, Frank Xiao, Pietro Perona, Markus Marks |
WACV | 3 |
| 2026 | SAVeD: Learning to Denoise Low-SNR Video for Improved Downstream PerformanceabstractLow signal-to-noise ratio (SNR) videos—such as those from underwater sonar, ultrasound, and microscopy—pose significant challenges for computer vision models, particularly when paired clean imagery for denoising is unavailable. We present Spatiotemporal Augmentations and denoising in Video for Downstream Tasks (SAVeD), a novel self-supervised method that denoises low-SNR sensor videos using only raw noisy data. By leveraging distinctions between foreground and background motion and exaggerating objects with stronger motion signal, SAVeD enhances foreground object visibility and reduces background and camera noise without requiring clean video. SAVeD has a set of architectural optimizations that lead to faster throughput, training, and inference than existing deep learning methods. We also introduce a new denoising metric, FBD, which indicates foreground-background divergence for detection datasets without requiring clean imagery. Our approach achieves state-of-the-art results for classification, detection, tracking, and counting tasks and it does so with fewer training resource requirements than existing deep-learning-based denoising methods. Project page here, Code: https://github.com/suzanne-stathatos/SAVeD. Suzanne Stathatos, Michael Hobley, Pietro Perona, Markus Marks |
WACV | 3 |
| 2026 | Unsupervised Representation Learning From Sparse Transformation AnalysisabstractThere is a vast literature on representation learning based on principles such as coding efficiency, statistical independence, causality, controllability, or symmetry. In this paper we propose to learn representations from sequence data by factorizing the transformations of the latent variables into sparse components. Input data are first encoded as distributions of latent activations and subsequently transformed using a probability flow model, before being decoded to predict a future input state. The flow model is decomposed into a number of rotational (divergence-free) vector fields and a number of potential flow (curl-free) fields. Our sparsity prior encourages only a small number of these fields to be active at any instant and infers the speed with which the probability flows along these fields. Training this model is completely unsupervised using a standard variational objective and results in a new form of disentangled representations where the input is not only represented by a combination of independent factors, but also by a combination of independent transformation primitives given by the learned flow fields. When viewing the transformations as symmetries one may interpret this as learning approximately equivariant representations. Empirically we demonstrate that this model achieves state of the art in terms of both data likelihood and unsupervised approximate equivariance errors on datasets composed of sequence transformations. Yue Song 0002, T. Anderson Keller 0001, Yisong Yue, Pietro Perona, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Model Diagnosis and Correction via Linguistic and Implicit Attribute EditingabstractHow can we troubleshoot a deep visual model, i.e. understand why it makes certain mistakes and further take action to correct its behavior? We design a ${\mathbf{M}}$ odel ${\mathbf{D}}$ iagnosis and ${\mathbf{C}}$ orrection system (MDC), an automated framework that analyzes the pattern of errors, proposes candidate causes of attributes, conducts hypothesis testing via attribute editing, and ultimately generates counterfactual training samples to improve the performance of the model. Unlike previous methods, in addition to the linguistic attributes, our method also incorporates the analysis for implicit causal attributes, those cannot to be accurately described by natural language. To achieve this, we propose an image editing module capable of leveraging both implicit and linguistic attributes to generate counterfactual images depicting error patterns and further experimentally validate causality relationships. Lastly, we enrich the training set with synthetic samples depicting verified causal attributes and retrain the model, further boosting accuracy and robustness. Extensive experiments on both generalized and specialized domains demonstrate the superiority of MDC in model diagnosis and correction. Specifically, we achieve an average relative improvement of 62.01% in HTER for face security application over state-of-the-art methods. Xuanbai Chen, Tianchen Zhao, Pietro Perona, Yifan Xing |
CVPR | 5 |
| 2025 | Is CLIP Ideal? No. Can We Fix It? Yes!abstractContrastive Language-Image Pre-Training (CLIP) is a popular method for learning multimodal latent spaces with well-organized semantics. Despite its wide range of applications, CLIP's latent space is known to fail at handling complex visual-textual interactions. Recent works attempt to address its shortcomings with data-centric or algorithmic approaches. But what if the problem is more fundamental, and lies in the geometry of CLIP? Toward this end, we rigorously analyze CLIP's latent space properties, and prove that no CLIP-like joint embedding space exists which can correctly do any two of the following at the same time: 1. represent basic descriptions and image content, 2. represent attribute binding, 3. represent spatial location and relationships, 4. represent negation. Informed by this analysis, we propose Dense Cosine Similarity Maps (DCSMs) as a principled and interpretable scoring method for CLIP-like models, which solves the fundamental limitations of CLIP by retaining the semantic topology of the image patches and text tokens. This method improves upon the performance of classical CLIP-like joint encoder models on a wide array of benchmarks. We share our code and data here for reproducibility: https://github.com/Raphoo/DCSM_Ideal_CLIP Raphi Kang, Yue Song 0002, Georgia Gkioxari, Pietro Perona |
ICCV | 4 |
| 2025 | Representational Similarity via Interpretable Visual ConceptsabstractHow do two deep neural networks differ in how they arrive at a decision? Measuring the similarity of deep networks has been a long-standing open question. Most existing methods provide a single number to measure the similarity of two networks at a given layer, but give no insight into what makes them similar or dissimilar. We introduce an interpretable representational similarity method (RSVC) to compare two networks. We use RSVC to discover shared and unique visual concepts between two models. We show that some aspects of model differences can be attributed to unique concepts discovered by one model that are not well represented in the other. Finally, we conduct extensive evaluation across different vision model architectures and training protocols to demonstrate its effectiveness. Neehar Kondapaneni, Oisin Mac Aodha, Pietro Perona |
ICLR | 3 |
| 2025 | Representational Difference ExplanationsabstractWe propose a method for discovering and visualizing the differences between two learned representations, enabling more direct and interpretable model comparisons. We validate our method, which we call Representational Differences Explanations (RDX), by using it to compare models with known conceptual differences and demonstrate that it recovers meaningful distinctions where existing explainable AI (XAI) techniques fail. Applied to state-of-the-art models on challenging subsets of the ImageNet and iNaturalist datasets, RDX reveals both insightful representational differences and subtle patterns in the data. Although comparison is a cornerstone of scientific analysis, current tools in machine learning, namely post hoc XAI methods, struggle to support model comparison effectively. Our work addresses this gap by introducing an effective and explainable tool for contrasting model representations. Neehar Kondapaneni, Oisin Mac Aodha, Pietro Perona |
NeurIPS | 3 |
| 2025 | Kuramoto Orientation Diffusion ModelsabstractOrientation-rich images, such as fingerprints and textures, often exhibit coherent angular directional patterns that are challenging to model using standard generative approaches based on isotropic Euclidean diffusion. Motivated by the role of phase synchronization in biological systems, we propose a score-based generative model built on periodic domains by leveraging stochastic Kuramoto dynamics in the diffusion process. In neural and physical systems, Kuramoto models capture synchronization phenomena across coupled oscillators -- a behavior that we re-purpose here as an inductive bias for structured image generation. In our framework, the forward process performs \textit{synchronization} among phase variables through globally or locally coupled oscillator interactions and attraction to a global reference phase, gradually collapsing the data into a low-entropy von Mises distribution. The reverse process then performs \textit{desynchronization}, generating diverse patterns by reversing the dynamics with a learned score function. This approach enables structured destruction during forward diffusion and a hierarchical generation process that progressively refines global coherence into fine-scale details. We implement wrapped Gaussian transition kernels and periodicity-aware networks to account for the circular geometry. Our method achieves competitive results on general image benchmarks and significantly improves generation quality on orientation-dense datasets like fingerprints and textures. Ultimately, this work demonstrates the promise of biologically inspired synchronization dynamics as structured priors in generative modeling. Yue Song 0002, Andy Keller, Sevan Brodjian, Takeru Miyato, Yisong Yue, Pietro Perona, Max Welling |
NeurIPS | 6 |
| 2025 | A Rapid Test for Accuracy and Bias of Face Recognition TechnologyabstractMeasuring the accuracy of face recognition (FR) systems is essential for improving performance and ensuring responsible use. Accuracy is typically estimated using large annotated datasets, which are costly and difficult to obtain. We propose a novel method for 1: 1 face verification that benchmarks FR systems quickly and without manual annotation, starting from approximate labels (e.g., from web search results). Unlike previous methods for training set label cleaning, ours leverages the embedding representation of the models being evaluated, achieving high accuracy in smaller-sized test datasets. Our approach reliably estimates FR accuracy and ranking, significantly reducing the time and cost of manual labeling. We also introduce the first public benchmark of five FR cloud services, revealing demographic biases, particularly lower accuracy for Asian women. Our rapid test method can democratize FR testing, promoting scrutiny and responsible use of the technology. Our method is provided as a publicly accessible tool at https://github.com/caltechvisionlab/frt-rapid-test. Manuel Knott 0001, Ignacio Serna, Ethan Mann, Pietro Perona |
WACV | 4 |
| 2025 | Learning Keypoints for Multi-Agent Behavior Analysis using Self-SupervisionabstractThe study of social interactions and collective behaviors through multi-agent video analysis is crucial in biology. While self-supervised keypoint discovery has emerged as a promising solution to reduce the need for manual keypoint annotations, existing methods often struggle with videos containing multiple interacting agents, especially those of the same species and color. To address this, we introduce B-KinD-multi, a novel approach that leverages pre-trained video segmentation models to guide keypoint discovery in multi-agent scenarios. This eliminates the need for time-consuming manual annotations on new experimental settings and organisms. Extensive evaluations demonstrate improved keypoint regression and downstream behavioral classification in videos of flies, mice, and rats. Furthermore, our method generalizes well to other species, including ants, bees, and humans, highlighting its potential for broad applications in automated keypoint annotation for multi-agent behavior analysis. Code available under: B-KinD-Multi Daniel Khalil, Christina Liu, Pietro Perona, Jennifer J. Sun, Markus Marks |
WACV | 3 |
| 2025 | Self-Supervised Incremental Learning of Object Representations from Arbitrary Image SetsabstractComputing a comprehensive and robust visual representation of an arbitrary object or category of objects is a complex problem. The difficulty increases when one starts from a set of uncalibrated images obtained from different sources. We propose a self-supervised approach, Multi-Image Latent Embedding (MILE), which computes a single representation from such an image set. MILE operates incrementally, considering one image at a time, while processing various depictions of the class through a shared gated cross-attention mechanism. The representations are progressively refined as more available images are incorporated, without requiring additional training. Our experiments on Amazon Berkeley Objects (ABO) and iNaturalist demonstrate the effectiveness in two tasks: object or category-specific image retrieval and unsupervised context-conditioned object segmentation. Moreover, the proposed multi-image input setup opens new frontiers for the task of object retrieval. Our studies indicate that our models can capture descriptive representations that better encapsulate the intrinsic characteristics of the objects. Our code is available at https://github.com/amazon-science/mile. George Leotescu, Alin-Ionut Popa, Diana Grigore, Daniel Voinea, Pietro Perona |
WACV | 5 |
| 2025 | A Closer Look at Benchmarking Self-supervised Pre-training with Image ClassificationabstractSelf-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data's inherent structure or context by solving a pretext task. With SSL, models can learn from abundant and cheap unlabeled data, significantly reducing the cost of training models where labels are expensive or inaccessible. In Computer Vision, SSL is widely used as pre-training followed by a downstream task, such as supervised transfer, few-shot learning on smaller labeled data sets, and/or unsupervised clustering. Unfortunately, it is infeasible to evaluate SSL methods on all possible downstream tasks and objectively measure the quality of the learned representation. Instead, SSL methods are evaluated using in-domain evaluation protocols, such as fine-tuning, linear probing, and k-nearest neighbors (kNN). However, it is not well understood how well these evaluation protocols estimate the representation quality of a pre-trained model for different downstream tasks under different conditions, such as dataset, metric, and model architecture. In this work, we study how classification-based evaluation protocols for SSL correlate and how well they predict downstream performance on different dataset types. Our study includes eleven common image datasets and 26 models that were pre-trained with different SSL methods or have different model backbones. We find that in-domain linear/kNN probing protocols are, on average, the best general predictors for out-of-domain performance. We further investigate the importance of batch normalization for the various protocols and evaluate how robust correlations are for different kinds of dataset domain shifts. In addition, we challenge assumptions about the relationship between discriminative and generative self-supervised methods, finding that most of their performance differences can be explained by changes to model backbones. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-025-02402-w. Markus Marks, Manuel Knott 0001, Neehar Kondapaneni, Elijah Cole, Thijs Defraeye, Fernando Pérez-Cruz, Pietro Perona |
Int. J. Comput. Vis. | 7 |
| 2024 | Text-Image Alignment for Diffusion-Based PerceptionabstractDiffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of these generative models for visual tasks is still an open question. Specifically, it is unclear how to use the prompting interface when applying diffusion backbones to vision tasks. We find that automatically generated captions can improve text-image alignment and significantly enhance a model's cross-attention maps, leading to better perceptual performance. Our approach improves upon the current state-of-the-art (SOTA) in diffusionbased semantic segmentation on ADE20K and the current overall SOTA for depth estimation on NYUv2. Furthermore, our method generalizes to the cross-domain setting. We use model personalization and caption modifications to align our model to the target domain and find improvements over unaligned baselines. Our crossdomain object detection model, trained on Pascal VOC, achieves SOTA results on Watercolor2K. Our cross-domain segmentation method, trained on Cityscapes, achieves SOTA results on Dark Zurich-val and Nighttime Driving. Project page: vision.caltech.edu/TADP/ Code page: github.com/damaggu/TADP Neehar Kondapaneni, Markus Marks, Manuel Knott 0001, Rogério Guimarães, Pietro Perona |
CVPR | 5 |
| 2024 | A Framework for Efficient Model Evaluation Through Stratification, Sampling, and Estimation
Riccardo Fogliato, Pratik Patil, Mathew Monfort, Pietro Perona |
ECCV (88) | 4 |
| 2024 | Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations
Riccardo Fogliato, Pratik Patil, Pietro Perona |
Int. J. Comput. Vis. | 3 |
| 2023 | BKinD-3D: Self-Supervised 3D Keypoint Discovery from Multi-View VideosabstractQuantifying motion in 3D is important for studying the behavior of humans and other animals, but manual pose annotations are expensive and time-consuming to obtain. Self-supervised keypoint discovery is a promising strategy for estimating 3D poses without annotations. However, current keypoint discovery approaches commonly process single 2D views and do not operate in the 3D space. We propose a new method to perform self-supervised keypoint discovery in 3D from multi-view videos of behaving agents, without any keypoint or bounding box supervision in 2D or 3D. Our method, BKinD-3D, uses an encoder-decoder architecture with a 3D volumetric heatmap, trained to reconstruct spatiotemporal differences across multiple views, in addition to joint length constraints on a learned 3D skeleton of the subject. In this way, we discover keypoints without requiring manual supervision in videos of humans and rats, demonstrating the potential of 3D keypoint discovery for studying behavior. Jennifer J. Sun, Lili Karashchuk, Amil Dravid, Serim Ryou, Sonia Fereidooni, John C. Tuthill, Aggelos K. Katsaggelos, Bingni W. Brunton, Georgia Gkioxari, Ann Kennedy, Yisong Yue, Pietro Perona |
CVPR | 12 |
| 2023 | Benchmarking Algorithmic Bias in Face Recognition: An Experimental Approach Using Synthetic Faces and Human EvaluationabstractWe propose an experimental method for measuring bias in face recognition systems. Existing methods to measure bias depend on benchmark datasets that are collected in the wild and annotated for protected (e.g., race, gender) and unprotected (e.g., pose, lighting) attributes. Such observational datasets only permit correlational conclusions, e.g., "Algorithm A’s accuracy is different on female and male faces in dataset X.". By contrast, experimental methods manipulate attributes individually and thus permit causal conclusions, e.g., "Algorithm A’s accuracy is affected by gender and skin color."Our method is based on generating synthetic faces using a neural face generator, where each attribute of interest is modified independently while leaving all other attributes constant. Human observers crucially provide the ground truth on perceptual identity similarity between synthetic image pairs. We validate our method quantitatively by evaluating race and gender biases of three research-grade face recognition models. Our synthetic pipeline reveals that for these algorithms, accuracy is lower for Black and East Asian population subgroups. Our method can also quantify how perceptual changes in attributes affect face identity distances reported by these models. Our large synthetic dataset, consisting of 48,000 synthetic face image pairs (10,200 unique synthetic faces) and 555,000 human annotations (individual attributes and pairwise identity comparisons) is available to researchers in this important area. Pietro Perona, Guha Balakrishnan |
ICCV | 2 |
| 2023 | Spatial Implicit Neural Representations for Global-Scale Species MappingabstractEstimating the geographical range of a species from sparse observations is a challenging and important geospatial prediction problem. Given a set of locations where a species has been observed, the goal is to build a model to predict whether the species is present or absent at any location. This problem has a long history in ecology, but traditional methods struggle to take advantage of emerging large-scale crowdsourced datasets which can include tens of millions of records for hundreds of thousands of species. In this work, we use Spatial Implicit Neural Representations (SINRs) to jointly estimate the geographical range of 47k species simultaneously. We find that our approach scales gracefully, making increasingly better predictions as we increase the number of species and the amount of data per species when training. To make this problem accessible to machine learning researchers, we provide four new benchmarks that measure different aspects of species range estimation and spatial representation learning. Using these benchmarks, we demonstrate that noisy and biased crowdsourced data can be combined with implicit neural representations to approximate expert-developed range maps for many species. Elijah Cole, Grant Van Horn, Christian Lange 0004, Alexander Shepard, Patrick Leary, Pietro Perona, Scott Loarie, Oisin Mac Aodha |
ICML | 6 |
| 2023 | MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of BehaviorabstractWe introduce MABe22, a large-scale, multi-agent video and trajectory benchmark to assess the quality of learned behavior representations. This dataset is collected from a variety of biology experiments, and includes triplets of interacting mice (4.7 million frames video+pose tracking data, 10 million frames pose only), symbiotic beetle-ant interactions (10 million frames video data), and groups of interacting flies (4.4 million frames of pose tracking data). Accompanying these data, we introduce a panel of real-life downstream analysis tasks to assess the quality of learned representations by evaluating how well they preserve information about the experimental conditions (e.g. strain, time of day, optogenetic stimulation) and animal behavior. We test multiple state-of-the-art self-supervised video and trajectory representation learning methods to demonstrate the use of our benchmark, revealing that methods developed using human action datasets do not fully translate to animal datasets. We hope that our benchmark and dataset encourage a broader exploration of behavior representation learning methods across species and settings. Jennifer J. Sun, Markus Marks, Andrew Ulmer, Dipam Chakraborty, Brian Geuther, Edward Hayes, Heng Jia, Sebastian Oleszko, Zachary Partridge, Milan Peelman, Alice Robie, Catherine E. Schretter, Keith Sheppard, Param Uttarwar, Julian Morgan Wagner, Erik Werner, Joseph Parker, Pietro Perona, Yisong Yue, Kristin Branson, Ann Kennedy |
ICML | 20 |
| 2022 | Multi-Dimensional, Nuanced and Subjective - Measuring the Perception of Facial ExpressionsabstractHumans can perceive multiple expressions, each one with varying intensity, in the picture of a face. We propose a methodology for collecting and modeling multidimensional modulated expression annotations from human annotators. Our data reveals that the perception of some expressions can be quite different across observers; thus, our model is designed to represent ambiguity alongside intensity. An empirical exploration of how many dimensions are necessary to capture the perception of facial expression suggests six principal expression dimensions are sufficient. Using our method, we collected multidimensional modulated expression annotations for 1,000 images culled from the popular ExpW in-the-wild dataset. As a proof of principle of our improved measurement technique, we used these annotations to benchmark four public domain algorithms for automated facial expression prediction. De'Aira Bryant, Siqi Deng, Nashlie Sephus, Wei Xia 0009, Pietro Perona |
CVPR | 5 |
| 2022 | Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerabstractText spotting end-to-end methods have recently gained attention in the literature due to the benefits of jointly optimizing the text detection and recognition components. Existing methods usually have a distinct separation between the detection and recognition branches, requiring exact annotations for the two tasks. We introduce TextTranSpotter (TTS), a transformer-based approach for text spotting and the first text spotting framework which may be trained with both fully- and weakly-supervised settings. By learning a single latent representation per word detection, and using a novel loss function based on the Hungarian loss, our method alleviates the need for expensive localization annotations. Trained with only text transcription annotations on real data, our weakly-supervised method achieves competitive performance with previous state-of-the-art fully-supervised methods. When trained in a fully-supervised manner, TextTranSpotter shows state-of-the-art results on multiple benchmarks. Yair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar, R. Manmatha, Pietro Perona |
CVPR | 6 |
| 2022 | Self-Supervised Keypoint Discovery in Behavioral VideosabstractWe propose a method for learning the posture and structure of agents from unlabelled behavioral videos. Starting from the observation that behaving agents are generally the main sources of movement in behavioral videos, our method, Behavioral Keypoint Discovery (B-KinD), uses an encoder-decoder architecture with a geometric bottleneck to reconstruct the spatiotemporal difference between video frames. By focusing only on regions of movement, our approach works directly on input videos without requiring manual annotations. Experiments on a variety of agent types (mouse, fly, human, jellyfish, and trees) demonstrate the generality of our approach and reveal that our discovered keypoints represent semantically meaningful body parts, which achieve state-of-the-art performance on keypoint regression among self-supervised methods. Additionally, B-KinD achieve comparable performance to supervised keypoints on downstream tasks, such as behavior classification, suggesting that our method can dramatically reduce model training costs vis-a-vis supervised methods. Jennifer J. Sun, Serim Ryou, Roni Goldshmid, Brandon Weissbourd, John O. Dabiri, David J. Anderson, Ann Kennedy, Yisong Yue, Pietro Perona |
CVPR | 9 |
| 2022 | Rayleigh EigenDirections (REDs): Nonlinear GAN Latent Space Traversals for Multidimensional Features
Guha Balakrishnan, Raghudeep Gadde, Aleix Martinez, Pietro Perona |
ECCV (17) | 4 |
| 2022 | Unsupervised and Semi-supervised Bias Benchmarking in Face Recognition
Alexandra Chouldechova, Siqi Deng, Wei Xia 0009, Pietro Perona |
ECCV (13) | 5 |
| 2022 | On Label Granularity and Object Localization
Elijah Cole, Kimberly Wilber, Grant Van Horn, Marco Fornoni, Pietro Perona, Serge J. Belongie, Andrew G. Howard, Oisin Mac Aodha |
ECCV (10) | 6 |
| 2022 | The Caltech Fish Counting Dataset: A Benchmark for Multiple-Object Tracking and Counting
Justin Kay, Peter Kulits, Suzanne Stathatos, Siqi Deng, Erik Young, Sara Beery, Grant Van Horn, Pietro Perona |
ECCV (8) | 8 |
| 2022 | Visual Knowledge Tracing
Neehar Kondapaneni, Pietro Perona, Oisin Mac Aodha |
ECCV (25) | 2 |
| 2021 | Sequence-to-Sequence Contrastive Learning for Text RecognitionabstractWe propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence structure, each feature map is divided into different instances over which the contrastive loss is computed. This operation enables us to contrast in a sub-word level, where from each image we extract several positive pairs and multiple negative examples. To yield effective visual representations for text recognition, we further suggest novel augmentation heuristics, different encoder architectures and custom projection heads. Experiments on hand-written text and on scene text show that when a text decoder is trained on the learned representations, our method out-performs non-sequential contrastive methods. In addition, when the amount of supervision is reduced, SeqCLR significantly improves performance compared with supervised training, and when fine-tuned with 100% of the labels, our method achieves state-of-the-art results on standard hand-written text recognition benchmarks. Aviad Aberdam, Ron Litman, Shahar Tsiper, Oron Anschel, Ron Slossberg, Shai Mazor, R. Manmatha, Pietro Perona |
CVPR | 8 |
| 2021 | Multi-Label Learning From Single Positive LabelsabstractPredicting all applicable labels for a given image is known as multi-label classification. Compared to the standard multi-class case (where each image has only one label), it is considerably more challenging to annotate training data for multi-label classification. When the number of potential labels is large, human annotators find it difficult to mention all applicable labels for each training image. Furthermore, in some settings detection is intrinsically difficult e.g. finding small object instances in high resolution images. As a result, multi-label training data is often plagued by false negatives. We consider the hardest version of this problem, where annotators provide only one relevant label for each image. As a result, training sets will have only one positive label per image and no confirmed negatives. We explore this special case of learning from missing labels across four different multi-label image classification datasets for both linear classifiers and end-to-end fine-tuned deep networks. We extend existing multi-label losses to this setting and propose novel variants that constrain the number of expected positive labels during training. Surprisingly, we show that in some cases it is possible to approach the performance of fully labeled classifiers despite training with significantly fewer confirmed labels. Elijah Cole, Oisin Mac Aodha, Titouan Lorieul, Pietro Perona, Dan Morris 0001, Nebojsa Jojic |
CVPR | 4 |
| 2021 | Task Programming: Learning Data Efficient Behavior RepresentationsabstractSpecialized domain knowledge is often necessary to accurately annotate training sets for in-depth analysis, but can be burdensome and time-consuming to acquire from domain experts. This issue arises prominently in automated behavior analysis, in which agent movements or actions of interest are detected from video tracking data. To reduce annotation effort, we present TREBA: a method to learn annotation-sample efficient trajectory embedding for behavior analysis, based on multi-task self-supervised learning. The tasks in our method can be efficiently engineered by domain experts through a process we call "task programming", which uses programs to explicitly encode structured knowledge from domain experts. Total domain expert effort can be reduced by exchanging data annotation time for the construction of a small number of programmed tasks. We evaluate this trade-off using data from behavioral neuroscience, in which specialized domain knowledge is used to identify behaviors. We present experimental results in three datasets across two domains: mice and fruit flies. Using embeddings from TREBA, we reduce annotation burden by up to a factor of 10 without compromising accuracy compared to state-of-the-art features. Our results thus suggest that task programming and self-supervision can be an effective way to reduce annotation effort for domain experts. Jennifer J. Sun, Ann Kennedy, Eric Zhan, David J. Anderson, Yisong Yue, Pietro Perona |
CVPR | 6 |
| 2021 | Species Distribution Modeling for Machine Learning Practitioners: A ReviewabstractConservation science depends on an accurate understanding of what’s happening in a given ecosystem. How many species live there? What is the makeup of the population? How is that changing over time? Species Distribution Modeling (SDM) seeks to predict the spatial (and sometimes temporal) patterns of species occurrence, i.e. where a species is likely to be found. The last few years have seen a surge of interest in applying powerful machine learning tools to challenging problems in ecology [2, 5, 8]. Despite its considerable importance, SDM has received relatively little attention from the computer science community. Our goal in this work is to provide computer scientists with the necessary background to read the SDM literature and develop ecologically useful ML-based SDM algorithms. In particular, we introduce key SDM concepts and terminology, review standard models, discuss data availability, and highlight technical challenges and pitfalls. Sara Beery, Elijah Cole, Joseph Parker, Pietro Perona, Kevin Winner |
COMPASS | 4 |
| 2020 | Rethinking Zero-Shot Video Classification: End-to-End Training for Realistic ApplicationsabstractTrained on large datasets, deep learning (DL) can accurately classify videos into hundreds of diverse classes. However, video data is expensive to annotate. Zero-shot learning (ZSL) proposes one solution to this problem. ZSL trains a model once, and generalizes to new tasks whose classes are not present in the training dataset. We propose the first end-to-end algorithm for ZSL in video classification. Our training procedure builds on insights from recent video classification literature and uses a trainable 3D CNN to learn the visual features. This is in contrast to previous video ZSL methods, which use pretrained feature extractors. We also extend the current benchmarking paradigm: Previous techniques aim to make the test task unknown at training time but fall short of this goal. We encourage domain shift across training and test data and disallow tailoring a ZSL model to a specific test dataset. We outperform the state-of-the-art by a wide margin. Our code, evaluation procedure and model weights are available online github.com/bbrattoli/ZeroShotVideoClassification. Biagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona, Krzysztof Chalupka |
CVPR | 4 |
| 2020 | Towards Causal Benchmarking of Bias in Face Analysis Algorithms
Guha Balakrishnan, Yuanjun Xiong, Wei Xia 0009, Pietro Perona |
ECCV (18) | 4 |
| 2020 | Synthetic Examples Improve Generalization for Rare ClassesabstractThe ability to detect and classify rare occurrences in images has important applications - for example, counting rare and endangered species when studying biodiversity, or detecting infrequent traffic scenarios that pose a danger to self-driving cars. Few-shot learning is an open problem: current computer vision systems struggle to categorize objects they have seen only rarely during training, and collecting a sufficient number of training examples of rare events is often challenging and expensive, and sometimes outright impossible. We explore in depth an approach to this problem: complementing the few available training images with ad-hoc simulated data.Our testbed is animal species classification, which has a real-world long-tailed distribution. We present two natural world simulators, and analyze the effect of different axes of variation in simulation, such as pose, lighting, model, and simulation method, and we prescribe best practices for efficiently incorporating simulated data for real-world performance gain. Our experiments reveal that synthetic data can considerably reduce error rates for classes that are rare, that as the amount of simulated data is increased, accuracy on the target class improves, and that high variation of simulated data provides maximum performance gain. Sara Beery, Dan Morris 0001, James Piavis, Ashish Kapoor, Markus Meister, Neel Joshi, Pietro Perona |
WACV | 8 |
| 2019 | Task2Vec: Task Embedding for Meta-LearningabstractWe introduce a method to generate vectorial representations of visual classification tasks which can be used to reason about the nature of those tasks and their relations. Given a dataset with ground-truth labels and a loss function, we process images through a "probe network" and compute an embedding based on estimates of the Fisher information matrix associated with the probe network parameters. This provides a fixed-dimensional embedding of the task that is independent of details such as the number of classes and requires no understanding of the class label semantics. We demonstrate that this embedding is capable of predicting task similarities that match our intuition about semantic and taxonomic relations between different visual tasks. We demonstrate the practical value of this framework for the meta-task of selecting a pre-trained feature extractor for a novel task. We present a simple meta-learning framework for learning a metric on embeddings that is capable of predicting which feature extractors will perform well on which task. Selecting a feature extractor with task embedding yields performance close to the best available feature extractor, with substantially less computational effort than exhaustively training and evaluating all available models. Alessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran, Subhransu Maji, Charless C. Fowlkes, Stefano Soatto, Pietro Perona |
ICCV | 8 |
| 2019 | Presence-Only Geographical Priors for Fine-Grained Image ClassificationabstractAppearance information alone is often not sufficient to accurately differentiate between fine-grained visual categories. Human experts make use of additional cues such as where, and when, a given image was taken in order to inform their final decision. This contextual information is readily available in many online image collections but has been underutilized by existing image classifiers that focus solely on making predictions based on the image contents. We propose an efficient spatio-temporal prior, that when conditioned on a geographical location and time, estimates the probability that a given object category occurs at that location. Our prior is trained from presence-only observation data and jointly models object categories, their spatio-temporal distributions, and photographer biases. Experiments performed on multiple challenging image classification datasets show that combining our prior with the predictions from image classifiers results in a large improvement in final classification performance. Oisin Mac Aodha, Elijah Cole, Pietro Perona |
ICCV | 3 |
| 2019 | Anchor Loss: Modulating Loss Scale Based on Prediction Difficulty
Serim Ryou, Seong-Gyun Jeong, Pietro Perona |
ICCV | 3 |
| 2019 | Teaching Multiple Concepts to a Forgetful LearnerabstractHow can we help a forgetful learner learn multiple concepts within a limited time frame? While there have been extensive studies in designing optimal schedules for teaching a single concept given a learner's memory model, existing approaches for teaching multiple concepts are typically based on heuristic scheduling techniques without theoretical guarantees. In this paper, we look at the problem from the perspective of discrete optimization and introduce a novel algorithmic framework for teaching multiple concepts with strong performance guarantees. Our framework is both generic, allowing the design of teaching schedules for different memory models, and also interactive, allowing the teacher to adapt the schedule to the underlying forgetting mechanisms of the learner. Furthermore, for a well-known memory model, we are able to identify a regime of model parameters where our framework is guaranteed to achieve high performance. We perform extensive evaluations using simulations along with real user studies in two concrete applications: (i) an educational app for online vocabulary teaching; and (ii) an app for teaching novices how to recognize animal species from images. Our results demonstrate the effectiveness of our algorithm compared to popular heuristic approaches. Anette Hunziker, Yuxin Chen 0001, Oisin Mac Aodha, Manuel Gomez-Rodriguez, Andreas Krause 0001, Pietro Perona, Yisong Yue, Adish Singla |
NeurIPS | 6 |
| 2018 | Near-Optimal Machine Teaching via Explanatory Teaching SetsabstractModern applications of machine teaching for humans often involve domain-specific, non- trivial target hypothesis classes. To facilitate understanding of the target hypothesis, it is crucial for the teaching algorithm to use examples which are interpretable to the human learner. In this paper, we propose NOTES, a principled framework for constructing interpretable teaching sets, utilizing explanations to accelerate the teaching process. Our algorithm is built upon a natural stochastic model of learners and a novel submodular surrogate objective function which greedily selects interpretable teaching examples. We prove that NOTES is competitive with the optimal explanation-based teaching strategy. We further instantiate NOTES with a specific hypothesis class, which can be viewed as an interpretable approximation of any hypothesis class, allowing us to handle complex hypothesis in practice. We demonstrate the effectiveness of NOTES on several image classification tasks, for both simulated and real human learners. Our experimental results suggest that by leveraging explanations, one can significantly speed up teaching. Yuxin Chen 0001, Oisin Mac Aodha, Shihan Su, Pietro Perona, Yisong Yue |
AISTATS | 4 |
| 2018 | It's all Relative: Monocular 3D Human Pose Estimation from Weakly Supervised Data
Matteo Ruggero Ronchi, Oisin Mac Aodha, Robert Eng, Pietro Perona |
BMVC | 4 |
| 2018 | Parsing Pose of People with Interaction
Serim Ryou, Pietro Perona |
BMVC | 2 |
| 2018 | Teaching Categories to Human Learners With Visual ExplanationsabstractWe study the problem of computer-assisted teaching with explanations. Conventional approaches for machine teaching typically only provide feedback at the instance level e.g., the category or label of the instance. However, it is intuitive that clear explanations from a knowledgeable teacher can significantly improve a student's ability to learn a new concept. To address these existing limitations, we propose a teaching framework that provides interpretable explanations as feedback and models how the learner incorporates this additional information. In the case of images, we show that we can automatically generate explanations that highlight the parts of the image that are responsible for the class label. Experiments on human learners illustrate that, on average, participants achieve better test set performance on challenging categorization tasks when taught with our interpretable approach compared to existing methods. Oisin Mac Aodha, Shihan Su, Yuxin Chen 0001, Pietro Perona, Yisong Yue |
CVPR | 4 |
| 2018 | The INaturalist Species Classification and Detection DatasetabstractExisting image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imbalanced, as some species are more abundant and easier to photograph than others. To encourage further progress in challenging real world conditions we present the iNaturalist species classification and detection dataset, consisting of 859,000 images from over 5,000 different species of plants and animals. It features visually similar species, captured in a wide variety of situations, from all over the world. Images were collected with different camera types, have varying image quality, feature a large class imbalance, and have been verified by multiple citizen scientists. We discuss the collection of the dataset and present extensive baseline experiments using state-of-the-art computer vision classification and detection models. Results show that current non-ensemble based methods achieve only 67% top one classification accuracy, illustrating the difficulty of the dataset. Specifically, we observe poor results for classes with small numbers of training examples suggesting more attention is needed in low-shot learning. Grant Van Horn, Oisin Mac Aodha, Yang Song 0009, Yin Cui, Chen Sun 0002, Alexander Shepard, Hartwig Adam, Pietro Perona, Serge J. Belongie |
CVPR | 8 |
| 2018 | Lean Multiclass CrowdsourcingabstractWe introduce a method for efficiently crowdsourcing multiclass annotations in challenging, real world image datasets. Our method is designed to minimize the number of human annotations that are necessary to achieve a desired level of confidence on class labels. It is based on combining models of worker behavior with computer vision. Our method is general: it can handle a large number of classes, worker labels that come from a taxonomy rather than a flat list, and can model the dependence of labels when workers can see a history of previous annotations. Our method may be used as a drop-in replacement for the majority vote algorithms used in online crowdsourcing services that aggregate multiple human annotations into a final consolidated label. In experiments conducted on two real-life applications we find that our method can reduce the number of required annotations by as much as a factor of 5.4 and can reduce the residual annotation error by up to 90% when compared with majority voting. Furthermore, the online risk estimates of the models may be used to sort the annotated collection and minimize subsequent expert review effort. Grant Van Horn, Steve Branson, Scott Loarie, Serge J. Belongie, Pietro Perona |
CVPR | 5 |
| 2018 | Context Embedding NetworksabstractLow dimensional embeddings that capture the main variations of interest in collections of data are important for many applications. One way to construct these embeddings is to acquire estimates of similarity from the crowd. Similarity is a multi-dimensional concept that varies from individual to individual. However, existing models for learning crowd embeddings typically make simplifying assumptions such as all individuals estimate similarity using the same criteria, the list of criteria is known in advance, or that the crowd workers are not influenced by the data that they see. To overcome these limitations we introduce Context Embedding Networks (CENs). In addition to learning interpretable embeddings from images, CENs also model worker biases for different attributes along with the visual context i.e. the attributes highlighted by a set of images. Experiments on three noisy crowd annotated datasets show that modeling both worker bias and visual context results in more interpretable embeddings compared to existing approaches. Kun Ho Kim, Oisin Mac Aodha, Pietro Perona |
CVPR | 3 |
| 2018 | Recognition in Terra Incognita
Sara Beery, Grant Van Horn, Pietro Perona |
ECCV (16) | 3 |
| 2018 | Understanding the Role of Adaptivity in Machine Teaching: The Case of Version Space LearnersabstractIn real-world applications of education, an effective teacher adaptively chooses the next example to teach based on the learner’s current state. However, most existing work in algorithmic machine teaching focuses on the batch setting, where adaptivity plays no role. In this paper, we study the case of teaching consistent, version space learners in an interactive setting. At any time step, the teacher provides an example, the learner performs an update, and the teacher observes the learner’s new state. We highlight that adaptivity does not speed up the teaching process when considering existing models of version space learners, such as the “worst-case” model (the learner picks the next hypothesis randomly from the version space) and the “preference-based” model (the learner picks hypothesis according to some global preference). Inspired by human teaching, we propose a new model where the learner picks hypotheses according to some local preference defined by the current hypothesis. We show that our model exhibits several desirable properties, e.g., adaptivity plays a key role, and the learner’s transitions over hypotheses are smooth/interpretable. We develop adaptive teaching algorithms, and demonstrate our results via simulation and user studies. Yuxin Chen 0001, Adish Singla, Oisin Mac Aodha, Pietro Perona, Yisong Yue |
NeurIPS | 4 |
| 2017 | Lean Crowdsourcing: Combining Humans and Machines in an Online SystemabstractWe introduce a method to greatly reduce the amount of redundant annotations required when crowdsourcing annotations such as bounding boxes, parts, and class labels. For example, if two Mechanical Turkers happen to click on the same pixel location when annotating a part in a given image-an event that is very unlikely to occur by random chance-, it is a strong indication that the location is correct. A similar type of confidence can be obtained if a single Turker happened to agree with a computer vision estimate. We thus incrementally collect a variable number of worker annotations per image based on online estimates of confidence. This is done using a sequential estimation of risk over a probabilistic model that combines worker skill, image difficulty, and an incrementally trained computer vision model. We develop specialized models and algorithms for binary annotation, part keypoint annotation, and sets of bounding box annotations. We show that our method can reduce annotation time by a factor of 4-11 for binary filtering of websearch results, 2-4 for annotation of boxes of pedestrians in images, while in many cases also reducing annotation error. We will make an end-to-end version of our system publicly available. Steve Branson, Grant Van Horn, Pietro Perona |
CVPR | 3 |
| 2017 | Seeing into Darkness: Scotopic Visual RecognitionabstractImages are formed by counting how many photons traveling from a given set of directions hit an image sensor during a given time interval. When photons are few and far in between, the concept of image breaks down and it is best to consider directly the flow of photons. Computer vision in this regime, which we call scotopic, is radically different from the classical image-based paradigm in that visual computations (classification, control, search) have to take place while the stream of photons is captured and decisions may be taken as soon as enough information is available. The scotopic regime is important for biomedical imaging, security, astronomy and many other fields. Here we develop a framework that allows a machine to classify objects with as few photons as possible, while maintaining the error rate below an acceptable threshold. A dynamic and asymptotically optimal speed-accuracy tradeoff is a key feature of this framework. We propose and study an algorithm to optimize the tradeoff of a convolutional network directly from lowlight images and evaluate on simulated images from standard datasets. Surprisingly, scotopic systems can achieve comparable classification performance as traditional vision systems while using less than 0.1% of the photons in a conventional image. In addition, we demonstrate that our algorithms work even when the illuminance of the environment is unknown and varying. Last, we outline a spiking neural network coupled with photon-counting sensors as a power-efficient hardware realization of scotopic algorithms. Bo Chen 0019, Pietro Perona |
CVPR | 2 |
| 2017 | Benchmarking and Error Diagnosis in Multi-instance Pose EstimationabstractWe propose a new method to analyze the impact of errors in algorithms for multi-instance pose estimation and a principled benchmark that can be used to compare them. We define and characterize three classes of errors - localization, scoring, and background - study how they are influenced by instance attributes and their impact on an algorithm's performance. Our technique is applied to compare the two leading methods for human pose estimation on the COCO Dataset, measure the sensitivity of pose estimation with respect to instance size, type and number of visible keypoints, clutter due to multiple instances, and the relative score of instances. The performance of algorithms, and the types of error they make, are highly dependent on all these variables, but mostly on the number of keypoints and the clutter. The analysis and software tools we propose offer a novel and insightful approach for understanding the behavior of pose estimation algorithms and an effective method for measuring their strengths and weaknesses. Matteo Ruggero Ronchi, Pietro Perona |
ICCV | 2 |
| 2017 | Learning Recurrent Representations for Hierarchical Behavior Modeling
Eyrun Eyjolfsdottir, Kristin Branson, Yisong Yue, Pietro Perona |
ICLR (Poster) | 4 |
| 2017 | A Simple Multi-Class Boosting Framework with Theoretical Guarantees and Empirical ProficiencyabstractThere is a need for simple yet accurate white-box learning systems that train quickly and with little data. To this end, we showcase REBEL, a multi-class boosting method, and present a novel family of weak learners called localized similarities. Our framework provably minimizes the training error of any dataset at an exponential rate. We carry out experiments on a variety of synthetic and real datasets, demonstrating a consistent tendency to avoid overfitting. We evaluate our method on MNIST and standard UCI datasets against other state-of-the-art methods, showing the empirical proficiency of our method. Ron Appel, Pietro Perona |
ICML | 2 |
| 2017 | Deciding How to Decide: Dynamic Routing in Artificial Neural NetworksabstractWe propose and systematically evaluate three strategies for training dynamically-routed artificial neural networks: graphs of learned transformations through which different input signals may take different paths. Though some approaches have advantages over others, the resulting networks are often qualitatively similar. We find that, in dynamically-routed networks trained to classify images, layers and branches become specialized to process distinct categories of images. Additionally, given a fixed computational budget, dynamically-routed networks tend to perform better than comparable statically-routed networks. Mason McGill, Pietro Perona |
ICML | 2 |
| 2016 | Multi-Level Cause-Effect SystemsabstractWe present a domain-general account of causation that applies to settings in which macro-level causal relations between two systems are of interest, but the relevant causal features are poorly understood and have to be aggregated from vast arrays of micro-measurements. Our approach generalizes that of Chalupka et. al. (2015) to the setting in which the macro-level effect is not specified. We formalize the connection between micro- and macro-variables in such situations and provide a coherent framework describing causal relations at multiple levels of analysis. We present an algorithm that discovers macro-variable causes and effects from micro-level measurements obtained from an experiment. We further show how to design experiments to discover macro-variables from observational micro-variable data. Finally, we show that under specific conditions, one can identify multiple levels of causal structure. Throughout the article, we use a simulated neuroscience multi-unit recording experiment to illustrate the ideas and the algorithms. Krzysztof Chalupka, Frederick Eberhardt, Pietro Perona |
AISTATS | 3 |
| 2016 | Cataloging Public Objects Using Aerial and Street-Level Images - Urban TreesabstractEach corner of the inhabited world is imaged from multiple viewpoints with increasing frequency. Online map services like Google Maps or Here Maps provide direct access to huge amounts of densely sampled, georeferenced images from street view and aerial perspective. There is an opportunity to design computer vision systems that will help us search, catalog and monitor public infrastructure, buildings and artifacts. We explore the architecture and feasibility of such a system. The main technical challenge is combining test time information from multiple views of each geographic location (e.g., aerial and street views). We implement two modules: det2geo, which detects the set of locations of objects belonging to a given category, and geo2cat, which computes the fine-grained category of the object at a given location. We introduce a solution that adapts state-of the-art CNN-based object detectors and classifiers. We test our method on "Pasadena Urban Trees", a new dataset of 80,000 trees with geographic and species annotations, and show that combining multiple views significantly improves both tree detection and tree species classification, rivaling human performance. Jan Dirk Wegner, Steve Branson, David Hall 0002, Konrad Schindler, Pietro Perona |
CVPR | 5 |
| 2016 | Unsupervised Discovery of El Nino Using Causal Feature Learning on Microlevel Climate Data
Krzysztof Chalupka, Tobias Bischoff, Frederick Eberhardt, Pietro Perona |
UAI | 4 |
| 2016 | Visipedia circa 2015
Serge J. Belongie, Pietro Perona |
Pattern Recognit. Lett. | 2 |
| 2015 | Describing Common Human Visual Actions in ImagesabstractWhich common human actions and interactions are recognizable in monocular still images?Which involve objects and/or other people?How many is a person performing at a time?We address these questions by exploring the actions and interactions that are detectable in the images of the MS COCO dataset.We make two main contributions.First, a list of 140 common 'visual actions', obtained by analyzing the largest on-line verb lexicon currently available for English (VerbNet) and human sentences used to describe images in MS COCO.Second, a complete set of annotations for those 'visual actions', composed of subject-object and associated verb, which we call COCO-a (a for 'actions').COCO-a is larger than existing action datasets in terms of number instances of actions, and is unique because it is data-driven, rather than experimenter-biased.Other unique features are that it is exhaustive, and that all subjects and objects are localized.A statistical analysis of the accuracy of our annotations and of each action, interaction and subject-object combination is provided. Matteo Ruggero Ronchi, Pietro Perona |
BMVC | 2 |
| 2015 | Fine-grained classification of pedestrians in video: Benchmark and state of the artabstractA video dataset that is designed to study fine-grained categorisation of pedestrians is introduced. Pedestrians were recorded “in-the-wild” from a moving vehicle. Annotations include bounding boxes, tracks, 14 keypoints with occlusion information and the fine-grained categories of age (5 classes), sex (2 classes), weight (3 classes) and clothing style (4 classes). There are a total of 27,454 bounding box and pose labels across 4222 tracks. This dataset is designed to train and test algorithms for fine-grained categorisation of people; it is also useful for benchmarking tracking, detection and pose estimation of pedestrians. State-of-the-art algorithms for fine-grained classification and pose estimation were tested using the dataset and the results are reported as a useful performance baseline. David Hall 0002, Pietro Perona |
CVPR | 2 |
| 2015 | Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collectionabstractWe introduce tools and methodologies to collect high quality, large scale fine-grained computer vision datasets using citizen scientists - crowd annotators who are passionate and knowledgeable about specific domains such as birds or airplanes. We worked with citizen scientists and domain experts to collect NABirds, a new high quality dataset containing 48,562 images of North American birds with 555 categories, part annotations and bounding boxes. We find that citizen scientists are significantly more accurate than Mechanical Turkers at zero cost. We worked with bird experts to measure the quality of popular datasets like CUB-200-2011 and ImageNet and found class label error rates of at least 4%. Nevertheless, we found that learning algorithms are surprisingly robust to annotation errors and this level of training data corruption can lead to an acceptably small increase in test error if the training set has sufficient size. At the same time, we found that an expert-curated high quality test set like NABirds is necessary to accurately measure the performance of fine-grained computer vision systems. We used NABirds to train a publicly available bird recognition service deployed on the web site of the Cornell Lab of Ornithology. Grant Van Horn, Steve Branson, Ryan Farrell, Scott Haber, Jessie Barry, Panagiotis G. Ipeirotis, Pietro Perona, Serge J. Belongie |
CVPR | 7 |
| 2015 | Tropel: Crowdsourcing Detectors with Minimal TrainingabstractThis paper introduces the Tropel system which enables non-technical users to create arbitrary visual detectors without first annotating a training set. Our primary contribution is a crowd active learning pipeline that is seeded with only a single positive example and an unlabeled set of training images. We examine the crowd's ability to train visual detectors given severely limited training themselves. This paper presents a series of experiments that reveal the relationship between worker training, worker consensus and the average precision of detectors trained by crowd-in-the-loop active learning. In order to verify the efficacy of our system, we train detectors for bird species that work nearly as well as those trained on the exhaustively labeled CUB 200 dataset at significantly lower cost and with little effort from the end user. To further illustrate the usefulness of our pipeline, we demonstrate qualitative results on unlabeled datasets containing fashion images and street-level photographs of Paris. Genevieve Patterson, Grant Van Horn, Serge J. Belongie, Pietro Perona, James Hays |
HCOMP | 4 |
| 2015 | Visual Causal Feature Learning
Krzysztof Chalupka, Pietro Perona, Frederick Eberhardt |
UAI | 2 |
| 2014 | Reconstructive Sparse Code Transfer for Contour Detection and Semantic Labeling
Michael Maire, Stella X. Yu, Pietro Perona |
ACCV (4) | 3 |
| 2014 | Improved Bird Species Recognition Using Pose Normalized Deep Convolutional Nets
Steve Branson, Grant Van Horn, Pietro Perona, Serge J. Belongie |
BMVC | 3 |
| 2014 | Hierarchical Cascade of Classifiers for Efficient Poselet Evaluation
Bo Chen 0019, Pietro Perona, Lubomir D. Bourdev |
BMVC | 2 |
| 2014 | Active Annotation TranslationabstractWe introduce a general framework for quickly annotating an image dataset when previous annotations exist. The new annotations (e.g. part locations) may be quite different from the old annotations (e.g. segmentations). Human annotators may be thought of as helping translate the old annotations into the new ones. As annotators label images, our algorithm incrementally learns a translator from source to target labels as well as a computer-vision-based structured predictor. These two components are combined to form an improved prediction system which accelerates the annotators' work through a smart GUI. We show how the method can be applied to translate between a wide variety of annotation types, including bounding boxes, segmentations, 2D and 3D part-based systems, and class and attribute labels. The proposed system will be a useful tool toward exploring new types of representations beyond simple bounding boxes, object segmentations, and class labels, and toward finding new ways to exploit existing large datasets with traditional types of annotations like SUN [36], Image Net [11], and Pascal VOC [12]. Experiments on the CUB-200-2011 and H3D datasets demonstrate 1) our method accelerates collection of part annotations by a factor of 3-20 compared to manual labeling, 2) our system can be used effectively in a scheme where definitions of part, attribute, or action vocabularies are evolved interactively without relabeling the entire dataset, and 3) toward collecting pose annotations, segmentations are more useful than bounding boxes, and part-level annotations are more effective than segmentations. Steve Branson, Kristjan Eldjarn Hjorleifsson, Pietro Perona |
CVPR | 3 |
| 2014 | From Categories to Individuals in Real Time - A Unified Boosting ApproachabstractA method for online, real-time learning of individual-object detectors is presented. Starting with a pre-trained boosted category detector, an individual-object detector is trained with near-zero computational cost. The individual detector is obtained by using the same feature cascade as the category detector along with elementary manipulations of the thresholds of the weak classifiers. This is ideal for online operation on a video stream or for interactive learning. Applications addressed by this technique are reidentification and individual tracking. Experiments on four challenging pedestrian and face datasets indicate that it is indeed possible to learn identity classifiers in real-time, besides being faster-trained, our classifier has better detection rates than previous methods on two of the datasets. David Hall 0002, Pietro Perona |
CVPR | 2 |
| 2014 | Similarity Comparisons for Interactive Fine-Grained CategorizationabstractCurrent human-in-the-loop fine-grained visual categorization systems depend on a predefined vocabulary of attributes and parts, usually determined by experts. In this work, we move away from that expert-driven and attribute-centric paradigm and present a novel interactive classification system that incorporates computer vision and perceptual similarity metrics in a unified framework. At test time, users are asked to judge relative similarity between a query image and various sets of images, these general queries do not require expert-defined terminology and are applicable to other domains and basic-level categories, enabling a flexible, efficient, and scalable system for fine-grained categorization with humans in the loop. Our system outperforms existing state-of-the-art systems for relevance feedback-based image retrieval as well as interactive classification, resulting in a reduction of up to 43% in the average number of questions needed to correctly classify an image. Catherine Wah, Grant Van Horn, Steve Branson, Subhransu Maji, Pietro Perona, Serge J. Belongie |
CVPR | 5 |
| 2014 | Distance Estimation of an Unknown Person from a Portrait
Xavier Paolo Burgos-Artizzu, Matteo Ruggero Ronchi, Pietro Perona |
ECCV (1) | 3 |
| 2014 | Detecting Social Actions of Fruit Flies
Eyrun Eyjolfsdottir, Steve Branson, Xavier Paolo Burgos-Artizzu, Eric D. Hoopfer, Jonathan Schor, David J. Anderson, Pietro Perona |
ECCV (2) | 7 |
| 2014 | Online, Real-Time Tracking Using a Category-to-Individual Detector
David Hall 0002, Pietro Perona |
ECCV (1) | 2 |
| 2014 | Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, C. Lawrence Zitnick |
ECCV (5) | 5 |
| 2014 | The Ignorant Led by the Blind: A Hybrid Human-Machine Vision System for Fine-Grained Categorization
Steve Branson, Grant Van Horn, Catherine Wah, Pietro Perona, Serge J. Belongie |
Int. J. Comput. Vis. | 4 |
| 2014 | Fast Feature Pyramids for Object DetectionabstractMulti-resolution image features may be approximated via extrapolation from nearby scales, rather than being computed explicitly. This fundamental insight allows us to design object detection algorithms that are as accurate, and considerably faster, than the state-of-the-art. The computational bottleneck of many modern detectors is the computation of features at every scale of a finely-sampled image pyramid. Our key insight is that one may compute finely sampled feature pyramids at a fraction of the cost, without sacrificing performance: for a broad family of features we find that features computed at octave-spaced scale intervals are sufficient to approximate features on a finely-sampled pyramid. Extrapolation is inexpensive as compared to direct feature computation. As a result, our approximation yields considerable speedups with negligible loss in detection accuracy. We modify three diverse visual recognition systems to use fast feature pyramids and show results on both pedestrian detection (measured on the Caltech, INRIA, TUD-Brussels and ETH data sets) and general object detection (measured on the PASCAL VOC). The approach is general and is widely applicable to vision algorithms requiring fine-grained multi-scale analysis. Our approximation is valid for images with broad spectra (most natural images) and fails for images with narrow band-pass spectra (e.g., periodic textures). Piotr Dollár, Ron Appel, Serge J. Belongie, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Merging Pose Estimates Across Space and TimeabstractNumerous 'non-maximum suppression' (NMS) post-processing schemes have been proposed for merging multiple independent object detections.We propose a generalization of NMS beyond bounding boxes to merge multiple pose estimates in a single frame.The final estimates are centroids rather than medoids as in standard NMS, thus being more accurate than any of the individual candidates.Using the same mathematical framework, we extend our approach to the multi-frame setting, merging multiple independent pose estimates across space and time and outputting both the number and pose of the objects present in a scene.Our approach sidesteps many of the inherent challenges associated with full tracking (e.g.objects entering/leaving a scene, extended periods of occlusion, etc.).We show its versatility by applying it to two distinct state-of-the-art pose estimation algorithms in three domains: human bodies, faces and mice.Our approach improves both detection accuracy (by helping disambiguate correspondences) as well as pose estimation quality and is computationally efficient. Xavier Paolo Burgos-Artizzu, David Hall 0002, Pietro Perona, Piotr Dollár |
BMVC | 3 |
| 2013 | Hierarchical Scene AnnotationabstractWe present a computer-assisted annotation system, together with a labeled dataset and benchmark suite, for evaluating an algorithm's ability to recover hierarchical scene structure.We evolve segmentation groundtruth from the two-dimensional image partition into a tree model that captures both occlusion and object-part relationships among possibly overlapping regions.Our tree model extends the segmentation problem to encompass object detection, object-part containment, and figure-ground ordering.We mitigate the cost of providing richer groundtruth labeling through a new webbased annotation tool with an intuitive graphical interface for rearranging the region hierarchy.Using precomputed superpixels, our tool also guides creation of user-specified regions with pixel-perfect boundaries.Widespread adoption of this human-machine combination should make the inaccuracies of bounding box labeling a relic of the past.Evaluating the state-of-the-art in fully automatic image segmentation reveals that it produces accurate two-dimension partitions, but does not respect groundtruth object-part structure.Our dataset and benchmark is the first to quantify these inadequacies.We illuminate recovery of rich scene structure as an important new goal for segmentation. Michael Maire, Stella X. Yu, Pietro Perona |
BMVC | 3 |
| 2013 | A Lazy Man's Approach to Benchmarking: Semisupervised Classifier Evaluation and RecalibrationabstractHow many labeled examples are needed to estimate a classifier's performance on a new dataset? We study the case where data is plentiful, but labels are expensive. We show that by making a few reasonable assumptions on the structure of the data, it is possible to estimate performance curves, with confidence bounds, using a small number of ground truth labels. Our approach, which we call Semi supervised Performance Evaluation (SPE), is based on a generative model for the classifier's confidence scores. In addition to estimating the performance of classifiers on new datasets, SPE can be used to recalibrate a classifier by re-estimating the class-conditional confidence distributions. Peter Welinder, Max Welling, Pietro Perona |
CVPR | 3 |
| 2013 | Robust Face Landmark Estimation under OcclusionabstractHuman faces captured in real-world conditions present large variations in shape and occlusions due to differences in pose, expression, use of accessories such as sunglasses and hats and interactions with objects (e.g. food). Current face landmark estimation approaches struggle under such conditions since they fail to provide a principled way of handling outliers. We propose a novel method, called Robust Cascaded Pose Regression (RCPR) which reduces exposure to outliers by detecting occlusions explicitly and using robust shape-indexed features. We show that RCPR improves on previous landmark estimation methods on three popular face datasets (LFPW, LFW and HELEN). We further explore RCPR's performance by introducing a novel face dataset focused on occlusion, composed of 1,007 faces presenting a wide range of occlusion patterns. RCPR reduces failure cases by half on all four datasets, at the same time as it detects face occlusions with a 80/40% precision/recall. Xavier Paolo Burgos-Artizzu, Pietro Perona, Piotr Dollár |
ICCV | 2 |
| 2013 | Quickly Boosting Decision Trees - Pruning Underachieving Features EarlyabstractBoosted decision trees are one of the most popular and successful learning techniques used today. While exhibiting fast speeds at test time, relatively slow training makes them impractical for applications with real-time learning requirements. We propose a principled approach to overcome this drawback. We prove a bound on the error of a decision stump given its preliminary error on a subset of the training data; the bound may be used to prune unpromising features early on in the training process. We propose a fast training algorithm that exploits this bound, yielding speedups of an order of magnitude at no cost in the final performance of the classifier. Our method is not a new variant of Boosting; rather, it may be used in conjunction with existing Boosting algorithms and other sampling heuristics to achieve even greater speedups. Ron Appel, Thomas J. Fuchs, Piotr Dollár, Pietro Perona |
ICML (3) | 4 |
| 2013 | Detecting Motion through Dynamic RefractionabstractRefraction causes random dynamic distortions in atmospheric turbulence and in views across a water interface. The latter scenario is experienced by submerged animals seeking to detect prey or avoid predators, which may be airborne or on land. Man encounters this when surveying a scene by a submarine or divers while wishing to avoid the use of an attention-drawing periscope. The problem of inverting random refracted dynamic distortions is difficult, particularly when some of the objects in the field of view (FOV) are moving. On the other hand, in many cases, just those moving objects are of interest, as they reveal animal, human, or machine activity. Furthermore, detecting and tracking these objects does not necessitate handling the difficult task of complete recovery of the scene. We show that moving objects can be detected very simply, with low false-positive rates, even when the distortions are very strong and dominate the object motion. Moreover, the moving object can be detected even if it has zero mean motion. While the object and distortion motions are random and unknown, they are mutually independent. This is expressed by a simple motion feature which enables discrimination of moving object points versus the background. Marina Alterman, Yoav Y. Schechner, Pietro Perona, Joseph Shamir |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Social behavior recognition in continuous videoabstractWe present a novel method for analyzing social behavior. Continuous videos are segmented into action `bouts' by building a temporal context model that combines features from spatio-temporal energy and agent trajectories. The method is tested on an unprecedented dataset of videos of interacting pairs of mice, which was collected as part of a state-of-the-art neurophysiological study of behavior. The dataset comprises over 88 hours (8 million frames) of annotated videos. We find that our novel trajectory features, used in a discriminative framework, are more informative than widely used spatio-temporal features; furthermore, temporal context plays an important role for action recognition in continuous videos. Our approach may be seen as a baseline method on this dataset, reaching a mean recognition rate of 61.2% compared to the expert's agreement rate of about 70%. Xavier Paolo Burgos-Artizzu, Piotr Dollár, Dayu Lin, David J. Anderson, Pietro Perona |
CVPR | 5 |
| 2012 | CompactKdt: Compact signatures for accurate large scale object recognitionabstractWe present a novel algorithm, Compact Kd-Trees (CompactKdt), that achieves state-of-the-art performance in searching large scale object image collections. The algorithm uses an order of magnitude less storage and computations by making use of both the full local features (e.g. SIFT) and their compact binary signatures to build and search the K-Tree. We compare classical PCA dimensionality reduction to three methods for generating compact binary representations for the features: Spectral Hashing, Locality Sensitive Hashing, and Locality Sensitive Binary Codes. CompactKdt achieves significant performance gain over using the binary signatures alone, and comparable performance to using the full features alone. Finally, our experiments show significantly better performance than the state-of-the-art Bag of Words (BoW) methods with equivalent or less storage and computational cost. Mohamed Aly 0001, Mario E. Munich, Pietro Perona |
WACV | 3 |
| 2012 | Unsupervised Learning of Categorical Segments in Image CollectionsabstractWhich one comes first: segmentation or recognition? We propose a unified framework for carrying out the two simultaneously and without supervision. The framework combines a flexible probabilistic model, for representing the shape and appearance of each segment, with the popular “bag of visual words” model for recognition. If applied to a collection of images, our framework can simultaneously discover the segments of each image and the correspondence between such segments, without supervision. Such recurring segments may be thought of as the “parts” of corresponding objects that appear multiple times in the image collection. Thus, the model may be used for learning new categories, detecting/classifying objects, and segmenting images, without using expensive human annotation. Marco Andreetto, Lihi Zelnik-Manor, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Pedestrian Detection: An Evaluation of the State of the ArtabstractPedestrian detection is a key problem in computer vision, with several applications that have the potential to positively impact quality of life. In recent years, the number of approaches to detecting pedestrians in monocular images has grown steadily. However, multiple data sets and widely varying evaluation protocols are used, making direct comparisons difficult. To address these shortcomings, we perform an extensive evaluation of the state of the art in a unified framework. We make three primary contributions: 1) We put together a large, well-annotated, and realistic monocular pedestrian detection data set and study the statistics of the size, position, and occlusion patterns of pedestrians in urban scenes, 2) we propose a refined per-frame evaluation methodology that allows us to carry out probing and informative comparisons, including measuring performance in relation to scale and occlusion, and 3) we evaluate the performance of sixteen pretrained state-of-the-art detectors across six data sets. Our study allows us to assess the state of the art and provides a framework for gauging future efforts. Our experiments show that despite significant progress, performance still has much room for improvement. In particular, detection is disappointing at low resolutions and for partially occluded pedestrians. Piotr Dollár, Christian Wojek, Bernt Schiele, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Distributed Kd-Trees for Ultra Large Scale Object RecognitionabstractDistributed Kd-Trees is a method for building image retrieval systems that can handle hundreds of millions of images. It is based on dividing the Kd-Tree into a “root subtree” that resides on a root machine, and several “leaf subtrees”, each residing on a leaf machine. The root machine handles incoming queries and farms out feature matching to an appropriate small subset of the leaf machines. Our implementation employs the MapReduce architecture to efficiently build and distribute the Kd-Tree for millions of images. It can run on thousands of machines, and provides orders of magnitude more throughput than the state-of-the-art, with better recognition performance. We show experiments with up to 100 million images running on 2048 machines, with run time of a fraction of a second for each query image. Mohamed Aly 0001, Mario E. Munich, Pietro Perona |
BMVC | 3 |
| 2011 | Strong supervision from weak annotation: Interactive training of deformable part modelsabstractWe propose a framework for large scale learning and annotation of structured models. The system interleaves interactive labeling (where the current model is used to semi-automate the labeling of a new example) and online learning (where a newly labeled example is used to update the current model parameters). This framework is scalable to large datasets and complex image models and is shown to have excellent theoretical and practical properties in terms of train time, optimality guarantees, and bounds on the amount of annotation effort per image. We apply this framework to part-based detection, and introduce a novel algorithm for interactive labeling of deformable part models. The labeling tool updates and displays in real-time the maximum likelihood location of all parts as the user clicks and drags the location of one or more parts. We demonstrate that the system can be used to efficiently and robustly train part and pose detectors on the CUB Birds-200-a challenging dataset of birds in unconstrained pose and environment. Steve Branson, Pietro Perona, Serge J. Belongie |
ICCV | 2 |
| 2011 | Object detection and segmentation from joint embedding of parts and pixelsabstractWe present a new framework in which image segmentation, figure/ground organization, and object detection all appear as the result of solving a single grouping problem. This framework serves as a perceptual organization stage that integrates information from low-level image cues with that of high-level part detectors. Pixels and parts each appear as nodes in a graph whose edges encode both affinity and ordering relationships. We derive a generalized eigen-problem from this graph and read off an interpretation of the image from the solution eigenvectors. Combining an off-the-shelf top-down part-based person detector with our low-level cues and grouping formulation, we demonstrate improvements to object detection and segmentation. Michael Maire, Stella X. Yu, Pietro Perona |
ICCV | 3 |
| 2011 | Multiclass recognition and part localization with humans in the loopabstractWe propose a visual recognition system that is designed for fine-grained visual categorization. The system is composed of a machine and a human user. The user, who is unable to carry out the recognition task by himself, is interactively asked to provide two heterogeneous forms of information: clicking on object parts and answering binary questions. The machine intelligently selects the most informative question to pose to the user in order to identify the object's class as quickly as possible. By leveraging computer vision and analyzing the user responses, the overall amount of human effort required, measured in seconds, is minimized. We demonstrate promising results on a challenging dataset of uncropped images, achieving a significant average reduction in human effort over previous methods. Catherine Wah, Steve Branson, Pietro Perona, Serge J. Belongie |
ICCV | 3 |
| 2011 | Predicting response time and error rates in visual searchabstractA model of human visual search is proposed. It predicts both response time (RT) and error rates (RT) as a function of image parameters such as target contrast and clutter. The model is an ideal observer, in that it optimizes the Bayes ratio of tar- get present vs target absent. The ratio is computed on the firing pattern of V1/V2 neurons, modeled by Poisson distributions. The optimal mechanism for integrat- ing information over time is shown to be a ‘soft max’ of diffusions, computed over the visual field by ‘hypercolumns’ of neurons that share the same receptive field and have different response properties to image features. An approximation of the optimal Bayesian observer, based on integrating local decisions, rather than diffusions, is also derived; it is shown experimentally to produce very similar pre- dictions. A psychophyisics experiment is proposed that may discriminate between which mechanism is used in the human brain. Bo Chen 0019, Vidhya Navalpakkam, Pietro Perona |
NIPS | 3 |
| 2011 | CrowdclusteringabstractIs it possible to crowdsource categorization? Amongst the challenges: (a) each annotator has only a partial view of the data, (b) different annotators may have different clustering criteria and may produce different numbers of categories, (c) the underlying category structure may be hierarchical. We propose a Bayesian model of how annotators may approach clustering and show how one may infer clusters/categories, as well as annotator parameters, using this model. Our experiments, carried out on large collections of images, suggest that Bayesian crowdclustering works well and may be superior to single-expert annotations. Ryan Gomes, Peter Welinder, Andreas Krause 0001, Pietro Perona |
NIPS | 4 |
| 2011 | Indexing in large scale image collections: Scaling properties and benchmarkabstractIndexing quickly and accurately in a large collection of images has become an important problem with many applications. Given a query image, the goal is to retrieve matching images in the collection. We compare the structure and properties of seven different methods based on the two leading approaches: voting from matching of local descriptors vs. matching histograms of visual words, including some new methods. We derive theoretical estimates of how the memory and computational cost scale with the number of images in the database. We evaluate these properties empirically on four real-world datasets with different statistics. We discuss the pros and cons of the different methods and suggest promising directions for future research. Mohamed Aly 0001, Mario E. Munich, Pietro Perona |
WACV | 3 |
| 2011 | Measuring and Predicting Object Importance
Merrielle Spain, Pietro Perona |
Int. J. Comput. Vis. | 2 |
| 2011 | Unsupervised Organization of Image Collections: Taxonomies and BeyondabstractWe introduce a nonparametric Bayesian model, called TAX, which can organize image collections into a tree-shaped taxonomy without supervision. The model is inspired by the Nested Chinese Restaurant Process (NCRP) and associates each image with a path through the taxonomy. Similar images share initial segments of their paths and thus share some aspects of their representation. Each internal node in the taxonomy represents information that is common to multiple images. We explore the properties of the taxonomy through experiments on a large (~10(4)) image collection with a number of users trying to locate quickly a given image. We find that the main benefits are easier navigation through image collections and reduced description length. A natural question is whether a taxonomy is the optimal form of organization for natural images. Our experiments indicate that although taxonomies can organize images in a useful manner, more elaborate structures may be even better suited for this task. Evgeniy Bart, Max Welling, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | The Fastest Pedestrian Detector in the WestabstractWe demonstrate a multiscale pedestrian detector operating in near real time (~6 fps on 640x480 images) with state-of-the-art detection performance. The computational bottleneck of many modern detectors is the construction of an image pyramid, typically sampled at 8-16 scales per octave, and associated feature computations at each scale. We propose a technique to avoid constructing such a finely sampled image pyramid without sacrificing performance: our key insight is that for a broad family of features, including gradient histograms, the feature responses computed at a single scale can be used to approximate feature responses at nearby scales. The approximation is accurate within an entire scale octave. This allows us to decouple the sampling of the image pyramid from the sampling of detection scales. Overall, our approximation yields a speedup of 10-100 times over competing methods with only a minor loss in detection accuracy of about 1-2% on the Caltech Pedestrian dataset across a wide range of evaluation settings. The results are confirmed on three additional datasets (INRIA, ETH, and TUD-Brussels) where our method always scores within a few percent of the state-of-the-art while being 1-2 orders of magnitude faster. The approach is general and should be widely applicable. Piotr Dollár, Serge J. Belongie, Pietro Perona |
BMVC | 3 |
| 2010 | Cascaded pose regressionabstractWe present a fast and accurate algorithm for computing the 2D pose of objects in images called cascaded pose regression (CPR). CPR progressively refines a loosely specified initial guess, where each refinement is carried out by a different regressor. Each regressor performs simple image measurements that are dependent on the output of the previous regressors; the entire system is automatically learned from human annotated training examples. CPR is not restricted to rigid transformations: `pose' is any parameterized variation of the object's appearance such as the degrees of freedom of deformable and articulated objects. We compare CPR against both standard regression techniques and human performance (computed from redundant human annotations). Experiments on three diverse datasets (mice, faces, fish) suggest CPR is fast (2-3ms per pose estimate), accurate (approaching human performance), and easy to train from small amounts of labeled data. Piotr Dollár, Peter Welinder, Pietro Perona |
CVPR | 3 |
| 2010 | Visual Recognition with Humans in the Loop
Steve Branson, Catherine Wah, Florian Schroff, Boris Babenko, Peter Welinder, Pietro Perona, Serge J. Belongie |
ECCV (4) | 6 |
| 2010 | Discriminative Clustering by Regularized Information MaximizationabstractIs there a principled way to learn a probabilistic discriminative classifier from an unlabeled data set? We present a framework that simultaneously clusters the data and trains a discriminative classifier. We call it Regularized Information Maximization (RIM). RIM optimizes an intuitive information-theoretic objective function which balances class separation, class balance and classifier complexity. The approach can flexibly incorporate different likelihood functions, express prior assumptions about the relative size of different classes and incorporate partial labels for semi-supervised learning. In particular, we instantiate the framework to unsupervised, multi-class kernelized logistic regression. Our empirical evaluation indicates that RIM outperforms existing methods on several real data sets, and demonstrates that RIM is an effective model selection method. Ryan Gomes, Andreas Krause 0001, Pietro Perona |
NIPS | 3 |
| 2010 | The Multidimensional Wisdom of CrowdsabstractDistributing labeling tasks among hundreds or thousands of annotators is an increasingly important method for annotating large datasets. We present a method for estimating the underlying value (e.g. the class) of each image from (noisy) annotations provided by multiple annotators. Our method is based on a model of the image formation and annotation process. Each image has different characteristics that are represented in an abstract Euclidean space. Each annotator is modeled as a multidimensional entity with variables representing competence, expertise and bias. This allows the model to discover and represent groups of annotators that have different sets of skills and knowledge, as well as groups of images that differ qualitatively. We find that our model predicts ground truth labels on both synthetic and real data more accurately than state of the art methods. Experiments also show that our model, starting from a set of binary labels, may discover rich information, such as different "schools of thought" amongst the annotators, and can group together images belonging to separate categories. Peter Welinder, Steve Branson, Serge J. Belongie, Pietro Perona |
NIPS | 4 |
| 2010 | Learning Object Categories From Internet Image SearchesabstractIn this paper, we describe a simple approach to learning models of visual object categories from images gathered from Internet image search engines. The images for a given keyword are typically highly variable, with a large fraction being unrelated to the query term, and thus pose a challenging environment from which to learn. By training our models directly from Internet images, we remove the need to laboriously compile training data sets, required by most other recognition approaches-this opens up the possibility of learning object category models “on-the-fly.” We describe two simple approaches, derived from the probabilistic latent semantic analysis (pLSA) technique for text document analysis, that can be used to automatically learn object models from these data. We show two applications of the learned model: first, to rerank the images returned by the search engine, thus improving the quality of the search engine; and second, to recognize objects in other image data sets. Rob Fergus, Li Fei-Fei 0001, Pietro Perona, Andrew Zisserman |
Proc. IEEE | 3 |
| 2010 | Vision of a VisipediaabstractThe web is not perfect: while text is easily searched and organized, pictures (the vast majority of the bits that one can find online) are not. In order to see how one could improve the web and make pictures first-class citizens of the web, I explore the idea of Visipedia, a visual interface for Wikipedia that is able to answer visual queries and enables experts to contribute and organize visual knowledge. Five distinct groups of humans would interact through Visipedia: users, experts, editors, visual workers, and machine vision scientists. The latter would gradually build automata able to interpret images. I explore some of the technical challenges involved in making Visipedia happen. I argue that Visipedia will likely grow organically, combining state-of-the-art machine vision with human labor. Pietro Perona |
Proc. IEEE | 1 |
| 2009 | Integral Channel FeaturesabstractWe study the performance of ‘integral channel features’ for image classification tasks, \nfocusing in particular on pedestrian detection. The general idea behind integral channel features is that multiple registered image channels are computed using linear and \nnon-linear transformations of the input image, and then features such as local sums, histograms, and Haar features and their various generalizations are efficiently computed \nusing integral images. Such features have been used in recent literature for a variety of \ntasks – indeed, variations appear to have been invented independently multiple times. \nAlthough integral channel features have proven effective, little effort has been devoted to \nanalyzing or optimizing the features themselves. In this work we present a unified view \nof the relevant work in this area and perform a detailed experimental evaluation. We \ndemonstrate that when designed properly, integral channel features not only outperform \nother features including histogram of oriented gradient (HOG), they also (1) naturally \nintegrate heterogeneous sources of information, (2) have few parameters and are insensitive to exact parameter settings, (3) allow for more accurate spatial localization during \ndetection, and (4) result in fast detectors when coupled with cascade classifiers. Piotr Dollár, Zhuowen Tu, Pietro Perona, Serge J. Belongie |
BMVC | 3 |
| 2009 | Pedestrian detection: A benchmarkabstractPedestrian detection is a key problem in computer vision, with several applications including robotics, surveillance and automotive safety. Much of the progress of the past few years has been driven by the availability of challenging public datasets. To continue the rapid rate of innovation, we introduce the Caltech Pedestrian Dataset, which is two orders of magnitude larger than existing datasets. The dataset contains richly annotated video, recorded from a moving vehicle, with challenging images of low resolution and frequently occluded people. We propose improved evaluation metrics, demonstrating that commonly used per-window measures are flawed and can fail to predict performance on full images. We also benchmark several promising detection systems, providing an overview of state-of-the-art performance and a direct, unbiased comparison of existing methods. Finally, by analyzing common failure cases, we help identify future research directions for the field. Piotr Dollár, Christian Wojek, Bernt Schiele, Pietro Perona |
CVPR | 4 |
| 2009 | Automatic discovery of image families: Global vs. local featuresabstractGathering a large collection of images has been made quite easy by social and image sharing websites, e.g. flickr.com. However, using such collections faces the problem that they contain a large number of duplicates and highly similar images. This work tackles the problem of how to automatically organize image collections into sets of similar images, called image families hereinafter. We thoroughly compare the performance of two approaches to measure image similarity: global descriptors vs. a set of local descriptors. We assess the performance of these approaches as the problem scales up to thousands of images and hundreds of families. We present our results on a new dataset of CD/DVD game covers. Mohamed Aly 0001, Peter Welinder, Mario E. Munich, Pietro Perona |
ICIP | 4 |
| 2008 | Unsupervised learning of visual taxonomiesabstractAs more images and categories become available, organizing them becomes crucial. We present a novel statistical method for organizing a collection of images into a tree-shaped hierarchy. The method employs a non-parametric Bayesian model and is completely unsupervised. Each image is associated with a path through a tree. Similar images share initial segments of their paths and therefore have a smaller distance from each other. Each internal node in the hierarchy represents information that is common to images whose paths pass through that node, thus providing a compact image representation. Our experiments show that a disorganized collection of images will be organized into an intuitive taxonomy. Furthermore, we find that the taxonomy allows good image categorization and, in this respect, is superior to the popular LDA model. Evgeniy Bart, Ian Porteous, Pietro Perona, Max Welling |
CVPR | 3 |
| 2008 | Incremental learning of nonparametric Bayesian mixture modelsabstractClustering is a fundamental task in many vision applications. To date, most clustering algorithms work in a batch setting and training examples must be gathered in a large group before learning can begin. Here we explore incremental clustering, in which data can arrive continuously. We present a novel incremental model-based clustering algorithm based on nonparametric Bayesian methods, which we call memory bounded variational Dirichlet process (MB-VDP). The number of clusters are determined flexibly by the data and the approach can be used to automatically discover object categories. The computational requirements required to produce model updates are bounded and do not grow with the amount of data processed. The technique is well suited to very large datasets, and we show that our approach outperforms existing online alternatives for learning nonparametric Bayesian mixture models. Ryan Gomes, Max Welling, Pietro Perona |
CVPR | 3 |
| 2008 | Learning and using taxonomies for fast visual categorizationabstractThe computational complexity of current visual categorization algorithms scales linearly at best with the number of categories. The goal of classifying simultaneously Ncat= 104- 105visual categories requires sub-linear classification costs. We explore algorithms for automatically building classification trees which have, in principle, logNcat complexity. We find that a greedy algorithm that recursively splits the set of categories into the two minimally confused subsets achieves 5-20 fold speedups at a small cost in classification performance. Our approach is independent of the specific classification algorithm used. A welcome by-product of our algorithm is a very reasonable taxonomy of the Caltech-256 dataset. Gregory Griffin, Pietro Perona |
CVPR | 2 |
| 2008 | Multiple Component Learning for Object Detection
Piotr Dollár, Boris Babenko, Serge J. Belongie, Pietro Perona, Zhuowen Tu |
ECCV (2) | 4 |
| 2008 | A Probabilistic Cascade of Detectors for Individual Object Recognition
Pierre Moreels, Pietro Perona |
ECCV (3) | 2 |
| 2008 | Some Objects Are More Equal Than Others: Measuring and Predicting Importance
Merrielle Spain, Pietro Perona |
ECCV (1) | 2 |
| 2008 | Unsupervised clustering for google searches of celebrity imagesabstractHow do we identify images of the same person in photo albums? How can we find images of a particular celebrity using Web image search engines? These types of tasks require solving numerous challenging issues in computer vision including: detecting whether an image contains a face, maintaining robustness to lighting, pose, occlusion, scale, and image quality, and using appropriate distance metrics to identify and compare different faces. In this paper we present a complete system which yields good performance on challenging tasks involving face recognition including image retrieval, unsupervised clustering of faces, and increasing precision of dasiaGoogle Imagepsila searches. All tasks use highly variable real data obtained from raw image searches on the web. Alex Holub, Pierre Moreels, Pietro Perona |
FG | 3 |
| 2008 | Memory bounded inference in topic modelsabstractWhat type of algorithms and statistical techniques support learning from very large datasets over long stretches of time? We address this question through a memory bounded version of a variational EM algorithm that approximates inference in a topic model. The algorithm alternates two phases: "model building" and "model compression" in order to always satisfy a given memory constraint. The model building phase expands its internal representation (the number of topics) as more data arrives through Bayesian model selection. Compression is achieved by merging data-items in clumps and only caching their sufficient statistics. Empirically, the resulting algorithm is able to handle datasets that are orders of magnitude larger than the standard batch version. Ryan Gomes, Max Welling, Pietro Perona |
ICML | 3 |
| 2008 | Guest EditorialabstractComputational Vision and Machine Learning have become synergistic fields of research.Modern machine learning techniques have improved the state of the art in computer vision and catalized re-thinking of key problems, such as recognition and tracking.In turn, vision has broadened the scope of machine learning, offering rich new challenges and highlighting the importance of representations.This special issue contains 15 papers at the intersection of vision and learning, most firmly rooted in both areas.The topics span a range of computer vision topics, from object recognition and tracking to lower-level visual tasks such as feature detection and finding contours and segments.The machine learning methods include discriminative and generative models, as well as both parametric and non-parametric representations.Since databases are proving to be crucial for progress in both machine learning and computer vision, we William T. Freeman, Pietro Perona, Bernhard Schölkopf |
Int. J. Comput. Vis. | 2 |
| 2008 | Hybrid Generative-Discriminative Visual Categorization
Alex Holub, Max Welling, Pietro Perona |
Int. J. Comput. Vis. | 3 |
| 2007 | Fast Terrain Classification Using Variable-Length Representation for Autonomous NavigationabstractWe propose a method for learning using a set of feature representations which retrieve different amounts of information at different costs. The goal is to create a more efficient terrain classification algorithm which can be used in real-time, onboard an autonomous vehicle. Instead of building a monolithic classifier with uniformly complex representation for each class, the main idea here is to actively consider the labels or misclassification cost while constructing the classifier. For example, some terrain classes might be easily separable from the rest, so very simple representation will be sufficient to learn and detect these classes. This is taken advantage of during learning, so the algorithm automatically builds a variable-length visual representation which varies according to the complexity of the classification task. This enables fast recognition of different terrain types during testing. We also show how to select a set of feature representations so that the desired terrain classification task is accomplished with high accuracy and is at the same time efficient. The proposed approach achieves a good trade-off between recognition performance and speedup on data collected by an autonomous robot. Anelia Angelova, Larry H. Matthies, Daniel M. Helmick, Pietro Perona |
CVPR | 4 |
| 2007 | On Constructing Facial Similarity MapsabstractAutomatically determining facial similarity is a difficult and open question in computer vision. The problem is complicated both because it is unclear what facial features humans use to determine facial similarity and because facial similarity is subjective in nature: similarity judgements change from person to person. In this work we suggest a system which places facial similarity on a solid computational footing. First we describe methods for acquiring facial similarity ratings from humans in an efficient manner. Next we show how to create feature vector representations for each face by extracted patches around facial key-points. Finally we show how to use the acquired similarity ratings to learn functional mapping which project facial-feature vectors into face spaces which correspond to our notions of facial similarity. We use different collections of images to both create and validate the face spaces including: perceptual similarity data obtained from humans, morphed faces between two different individuals, and the CMU PIE collection which contains images of the same individual under different lighting conditions. We demonstrate that using our methods we can effectively create face spaces which correspond to human notions of facial similarity. Alex Holub, Yun-hsueh Liu, Pietro Perona |
CVPR | 3 |
| 2007 | Non-Parametric Probabilistic Image SegmentationabstractWe propose a simple probabilistic generative model for image segmentation. Like other probabilistic algorithms (such as EM on a mixture of Gaussians) the proposed model is principled, provides both hard and probabilistic cluster assignments, as well as the ability to naturally incorporate prior knowledge. While previous probabilistic approaches are restricted to parametric models of clusters (e.g., Gaussians) we eliminate this limitation. The suggested approach does not make heavy assumptions on the shape of the clusters and can thus handle complex structures. Our experiments show that the suggested approach outperforms previous work on a variety of image segmentation tasks. Marco Andreetto, Lihi Zelnik-Manor, Pietro Perona |
ICCV | 3 |
| 2007 | Learning slip behavior using automatic mechanical supervisionabstractWe address the problem of learning terrain traversability properties from visual input, using automatic mechanical supervision collected from sensors onboard an autonomous vehicle. We present a novel probabilistic framework in which the visual information and the mechanical supervision interact to learn particular terrain types and their properties. The proposed method is applied to learning of rover slippage from visual information in a completely automatic fashion. Our experiments show that using mechanical measurements as automatic supervision significantly improves the visual-based classification alone and approaches the results of learning with manual supervision. This work will enable the rover to drive safely on slopes, learning autonomously about different terrains and their slip characteristics. Anelia Angelova, Larry H. Matthies, Daniel M. Helmick, Pietro Perona |
ICRA | 4 |
| 2007 | Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories
Li Fei-Fei 0001, Rob Fergus, Pietro Perona |
Comput. Vis. Image Underst. | 3 |
| 2007 | Weakly Supervised Scale-Invariant Learning of Models for Visual Recognition
Rob Fergus, Pietro Perona, Andrew Zisserman |
Int. J. Comput. Vis. | 2 |
| 2007 | Evaluation of Features Detectors and Descriptors based on 3D Objects
Pierre Moreels, Pietro Perona |
Int. J. Comput. Vis. | 2 |
| 2007 | 3D Reconstruction by Shadow Carving: Theory and Practical Evaluation
Silvio Savarese, Marco Andreetto, Holly E. Rushmeier, Fausto Bernardini, Pietro Perona |
Int. J. Comput. Vis. | 5 |
| 2007 | Automatic recognition of biological particles in microscopic images
Marc'Aurelio Ranzato, P. E. Taylor, James M. House, R. C. Flagan, Yann LeCun, Pietro Perona |
Pattern Recognit. Lett. | 6 |
| 2006 | Learning to Predict Slip for Ground RobotsabstractIn this paper we predict the amount of slip an exploration rover would experience using stereo imagery by learning from previous examples of traversing similar terrain. To do that, the information of terrain appearance and geometry regarding some location is correlated to the slip measured by the rover while this location is being traversed. This relationship is learned from previous experience, so slip can be predicted later at a distance from visual information only. The advantages of the approach are: 1) learning from examples allows the system to adapt to unknown terrains rather than using fixed heuristics or predefined rules; 2) the feedback about the observed slip is received from the vehicle's own sensors which can fully automate the process; 3) learning slip from previous experience can replace complex mechanical modeling of vehicle or terrain, which is time consuming and not necessarily feasible. Predicting slip is motivated by the need to assess the risk of getting trapped before entering a particular terrain. For example, a planning algorithm can utilize slip information by taking into consideration that a slippery terrain is costly or hazardous to traverse. A generic nonlinear regression framework is proposed in which the terrain type is determined from appearance and then a nonlinear model of slip is learned for a particular terrain type. In this paper we focus only on the latter problem and provide slip learning and prediction results for terrain types, such as soil, sand, gravel, and asphalt. The slip prediction error achieved is about 15% which is comparable to the measurement errors for slip itself Anelia Angelova, Larry H. Matthies, Daniel M. Helmick, Gabe Sibley, Pietro Perona |
ICRA | 5 |
| 2006 | Graph-Based Visual SaliencyabstractA new bottom-up visual saliency model, Graph-Based Visual Saliency (GBVS), is proposed. It consists of two steps: rst forming activation maps on certain feature channels, and then normalizing them in a way which highlights conspicuity and admits combination with other maps. The model is simple, and biologically plausible insofar as it is naturally parallelized. This model powerfully predicts human xations on 749 variations of 108 natural images, achieving 98% of the ROC area of a human-based control, whereas the classical algorithms of Itti & Koch ([2], [3], [4]) achieve only 84%. Jonathan Harel, Christof Koch, Pietro Perona |
NIPS | 3 |
| 2006 | One-Shot Learning of Object CategoriesabstractLearning visual models of object categories notoriously requires hundreds or thousands of training examples. We show that it is possible to learn much information about a category from just one, or a handful, of images. The key insight is that, rather than learning from scratch, one can take advantage of knowledge coming from previously learned categories, no matter how different these categories might be. We explore a Bayesian implementation of this idea. Object categories are represented by probabilistic models. Prior knowledge is represented as a probability density function on the parameters of these models. The posterior model for an object category is obtained by updating the prior in the light of one or more observations. We test a simple implementation of our algorithm on a database of 101 diverse object categories. We compare category models learned by an implementation of our Bayesian approach to models learned from by Maximum Likelihood (ML) and Maximum A Posteriori (MAP) methods. We find that on a database of more than 100 categories, the Bayesian approach produces informative models when the number of training examples is too small for other methods to operate successfully. Li Fei-Fei 0001, Rob Fergus, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | Beyond Pairwise ClusteringabstractWe consider the problem of clustering in domains where the affinity relations are not dyadic (pairwise), but rather triadic, tetradic or higher. The problem is an instance of the hypergraph partitioning problem. We propose a two-step algorithm for solving this problem. In the first step we use a novel scheme to approximate the hypergraph using a weighted graph. In the second step a spectral partitioning algorithm is used to partition the vertices of this graph. The algorithm is capable of handling hyperedges of all orders including order two, thus incorporating information of all orders simultaneously. We present a theoretical analysis that relates our algorithm to an existing hypergraph partitioning algorithm and explain the reasons for its superior performance. We report the performance of our algorithm on a variety of computer vision problems and compare it to several existing hypergraph partitioning algorithms. Sameer Agarwal 0001, Jongwoo Lim, Lihi Zelnik-Manor, Pietro Perona, David J. Kriegman, Serge J. Belongie |
CVPR (2) | 4 |
| 2005 | Pruning Training Sets for Learning of Object CategoriesabstractTraining datasets for learning of object categories are often contaminated or imperfect. We explore an approach to automatically identify examples that are noisy or troublesome for learning and exclude them from the training set. The problem is relevant to learning in semi-supervised or unsupervised setting, as well as to learning when the training data is contaminated with wrongly labeled examples or when correctly labeled, but hard to learn examples, are present. We propose a fully automatic mechanism for noise cleaning, called 'data pruning' and demonstrate its success on learning of human faces. It is not assumed that the data or the noise can be modeled or that additional training examples are available. Our experiments show that data pruning can improve on generalization performance for algorithms with various robustness to noise. It outperforms methods with regularization properties and is superior to commonly applied aggregation methods, such as bagging. Anelia Angelova, Yaser S. Abu-Mostafa, Pietro Perona |
CVPR (1) | 3 |
| 2005 | Hybrid Models for Human Motion RecognitionabstractProbabilistic models have been previously shown to be efficient and effective for modeling and recognition of human motion. In particular we focus on methods which represent the human motion model as a triangulated graph. Previous approaches learned models based just on positions and velocities of the body parts while ignoring their appearance. Moreover, a heuristic approach was commonly used to obtain translation invariance. In this paper we suggest an improved approach for learning such models and using them for human motion recognition. The suggested approach combines multiple cues, i.e., positions, velocities and appearance into both the learning and detection phases. Furthermore, we introduce global variables in the model, which can represent global properties such as translation, scale or view-point. The model is learned in an unsupervised manner from unlabelled data. We show that the suggested hybrid probabilistic model (which combines global variables, like translation, with local variables, like relative positions and appearances of body parts), leads to: (i) faster convergence of learning phase, (it) robustness to occlusions, and, (Hi) higher recognition rate. Claudio Fanti, Lihi Zelnik-Manor, Pietro Perona |
CVPR (1) | 3 |
| 2005 | A Sparse Object Category Model for Efficient Learning and Exhaustive RecognitionabstractWe present a "parts and structure" model for object category recognition that can be learnt efficiently and in a semi-supervised manner: the model is learnt from example images containing category instances, without requiring segmentation from background clutter. The model is a sparse representation of the object, and consists of a star topology configuration of parts modeling the output of a variety of feature detectors. The optimal choice of feature types (whose repertoire includes interest points, curves and regions) is made automatically. In recognition, the model may be applied efficiently in an exhaustive manner, bypassing the need for feature detectors, to give the globally optimal match within a query image. The approach is demonstrated on a wide variety of categories, and delivers both successful classification and localization of the object within the image. Rob Fergus, Pietro Perona, Andrew Zisserman |
CVPR (1) | 2 |
| 2005 | A Discriminative Framework for Modelling Object ClassesabstractHere we explore a discriminative learning method on underlying generative models for the purpose of discriminating between object categories. Visual recognition algorithms learn models from a set of training examples. Generative models learn their representations by considering data from a single class. Generative models are popular in computer vision for many reasons, including their ability to elegantly incorporate prior knowledge and to handle correspondences between object parts and detected features. However, generative models are often inferior to discriminative models during classification tasks. We study a discriminative approach to learning object categories which maintains the representational power of generative learning, but trains the generative models in a discriminative manner. The discriminatively trained models perform better during classification tasks as a result of selecting discriminative sets of features. We conclude by proposing a multi-class object recognition system which initially trains object classes in a generative manner, identifies subsets of similar classes with high confusion, and finally trains models for these subsets in a discriminative manner to realize gains in classification performance. Alex Holub, Pietro Perona |
CVPR (1) | 2 |
| 2005 | A Bayesian Hierarchical Model for Learning Natural Scene CategoriesabstractWe propose a novel approach to learn and recognize natural scene categories. Unlike previous work, it does not require experts to annotate the training set. We represent the image of a scene by a collection of local regions, denoted as codewords obtained by unsupervised learning. Each region is represented as part of a "theme". In previous work, such themes were learnt from hand-annotations of experts, while our method learns the theme distributions as well as the codewords distribution over the themes without supervision. We report satisfactory categorization performances on a large set of 13 categories of complex scenes. Li Fei-Fei 0001, Pietro Perona |
CVPR (2) | 2 |
| 2005 | Learning Object Categories from Google's Image SearchabstractCurrent approaches to object category recognition require datasets of training images to be manually prepared, with varying degrees of supervision. We present an approach that can learn an object category from just its name, by utilizing the raw output of image search engines available on the Internet. We develop a new model, TSI-pLSA, which extends pLSA (as applied to visual words) to include spatial information in a translation and scale invariant manner. Our approach can handle the high intra-class variability and large proportion of unrelated images returned by search engines. We evaluate tire models on standard test sets, showing performance competitive with existing methods trained on hand prepared datasets Rob Fergus, Li Fei-Fei 0001, Pietro Perona, Andrew Zisserman |
ICCV | 3 |
| 2005 | Combining Generative Models and Fisher Kernels for Object RecognitionabstractLearning models for detecting and classifying object categories is a challenging problem in machine vision. While discriminative approaches to learning and classification have, in principle, superior performance, generative approaches provide many useful features, one of which is the ability to naturally establish explicit correspondence between model components and scene features - this, in turn, allows for the handling of missing data and unsupervised learning in clutter. We explore a hybrid generative/discriminative approach using 'Fisher kernels' by Jaakkola and Haussler (1999) which retains most of the desirable properties of generative methods, while increasing the classification performance through a discriminative setting. Furthermore, we demonstrate how this kernel framework can be used to combine different types of features and models into a single classifier. Our experiments, conducted on a number of popular benchmarks, show strong performance improvements over the corresponding generative approach and are competitive with the best results reported in the literature. Alex Holub, Max Welling, Pietro Perona |
ICCV | 3 |
| 2005 | Evaluation of Features Detectors and Descriptors Based on 3D ObjectsabstractWe explore the performance of a number of popular feature detectors and descriptors in matching 3D object features across viewpoints and lighting conditions. To this end we design a method, based on intersecting epipolar constraints, for providing ground truth correspondence automatically. We collect a database of 100 objects viewed from 144 calibrated viewpoints under three different lighting conditions. We find that the combination of Hessian-affine feature finder and SIFT features is most robust to viewpoint change. Harris-affine combined with SIFT and Hessian-affine combined with shape context descriptors were best respectively for lighting changes and scale changes. We also find that no detector-descriptor combination performs well with viewpoint changes of more than 25-30/spl deg/. Pierre Moreels, Pietro Perona |
ICCV | 2 |
| 2005 | Squaring the Circles in PanoramasabstractPictures taken by a rotating camera cover the viewing sphere surrounding the center of rotation. Having a set of images registered and blended on the sphere what is left to be done, in order to obtain a flat panorama, is projecting the spherical image onto a picture plane. This step is unfortunately not obvious - the surface of the sphere may not be flattened onto a page without some form of distortion. The objective of this paper is discussing the difficulties and opportunities that are connected to the projection from viewing sphere to image plane. We first explore a number of alternatives to the commonly used linear perspective projection. These are 'global' projections and do not depend on image content. We then show that multiple projections may coexist successfully in the same mosaic: these projections are chosen locally and depend on what is present in the pictures. We show that such multi-view projections can produce more compelling results than the global projections Lihi Zelnik-Manor, Gabriele Peters, Pietro Perona |
ICCV | 3 |
| 2005 | Selective visual attention enables learning and recognition of multiple objects in cluttered scenes
Dirk Bernhardt-Walther, Ueli Rutishauser, Christof Koch, Pietro Perona |
Comput. Vis. Image Underst. | 4 |
| 2005 | Local Shape from Mirror Reflections
Silvio Savarese, Min Chen 0037, Pietro Perona |
Int. J. Comput. Vis. | 3 |
| 2004 | Is Bottom-Up Attention Useful for Object Recognition?
Ueli Rutishauser, Dirk Bernhardt-Walther, Christof Koch, Pietro Perona |
CVPR (2) | 4 |
| 2004 | A Visual Category Filter for Google Images
Rob Fergus, Pietro Perona, Andrew Zisserman |
ECCV (1) | 2 |
| 2004 | Recognition by Probabilistic Hypothesis Construction
Pierre Moreels, Michael Maire, Pietro Perona |
ECCV (1) | 3 |
| 2004 | Recovering Local Shape of a Mirror Surface from Reflection of a Regular Grid
Silvio Savarese, Min Chen 0037, Pietro Perona |
ECCV (3) | 3 |
| 2004 | Sampling Methods for Unsupervised LearningabstractWe present an algorithm to overcome the local maxima problem in es- timating the parameters of mixture models. It combines existing ap- proaches from both EM and a robust fitting algorithm, RANSAC, to give a data-driven stochastic learning scheme. Minimal subsets of data points, sufficient to constrain the parameters of the model, are drawn from pro- posal densities to discover new regions of high likelihood. The proposal densities are learnt using EM and bias the sampling toward promising solutions. The algorithm is computationally efficient, as well as effective at escaping from local maxima. We compare it with alternative methods, including EM and RANSAC, on both challenging synthetic data and the computer vision problem of alpha-matting. Rob Fergus, Andrew Zisserman, Pietro Perona |
NIPS | 3 |
| 2004 | Common-Frame Model for Object RecognitionabstractA generative probabilistic model for objects in images is presented. An object consists of a constellation of features. Feature appearance and pose are modeled probabilistically. Scene images are generated by draw- ing a set of objects from a given database, with random clutter sprinkled on the remaining image surface. Occlusion is allowed. We study the case where features from the same object share a common reference frame. Moreover, parameters for shape and appearance den- sities are shared across features. This is to be contrasted with previous work on probabilistic `constellation' models where features depend on each other, and each feature and model have different pose and appear- ance statistics [1, 2]. These two differences allow us to build models containing hundreds of features, as well as to train each model from a single example. Our model may also be thought of as a probabilistic revisitation of Lowe's model [3, 4]. We propose an efficient entropy-minimization inference algorithm that constructs the best interpretation of a scene as a collection of objects and clutter. We test our ideas with experiments on two image databases. We compare with Lowe's algorithm and demonstrate better performance, in particular in presence of large amounts of background clutter. 1 Introduction There is broad agreement in the machine vision literature that objects and object categories should be represented as collections of features or parts with distinctive appearance and mutual position [1, 2, 4, 5, 6, 7, 8, 9]. A number of ideas for efficient detection algorithms (find instances of a given object category, e.g. faces) have been proposed by virtually all the cited authors, far fewer for recognition (list all objects and their pose in a given image) where matching would ideally take a logarithmic time with respect to the number of avail- able models [3, 4]. Learning of parameters characterizing features shape or appearance is still a difficult area, with most authors opting for heavy human intervention (typically segmentation and alignment of the training examples, although [1, 2, 3] train without su- pervision) and very large training sets for object categories (typically in the order of 10 3 - 104, although [10] recently demonstrated learning categories from 1-10 examples). This work is based on two complementary efforts: the deterministic recognition system proposed by Lowe [3, 4], and the probabilistic constellation models by Perona and col- laborators [1, 2]. The first line of work has three attractive characteristics: objects are represented with hundreds of features, thus increasing robustness; models are learned from a single training example; last but not least, recognition is efficient with databases of hun- dreds of objects. The drawback of Lowe's approach is that both modeling decisions and algorithms rely on heuristics, whose design and performance may be far from optimal in Figure 1: Diagram of our recognition model showing database, test image and two competing hy- potheses. To avoid a cluttered diagram, only one partial hypothesis is displayed for each hypothesis. The predicted position of models according to the hypotheses are overlaid on the test image. some circumstances. Conversely, the second line of work is based on principled proba- bilistic object models which yield principled and, in some respects, optimal algorithms for learning and recognition/detection. Unfortunately, the large number of parameters em- ployed in each model limit in practice the number of features being used and require many training examples. By recasting Lowe's model and algorithms in probabilistic terms, we hope to combine the advantages of both methods. Besides, in this paper we choose to focus on individual objects as in [3, 4] rather than on categories as in [1, 2]. In [11] we presented a model aimed at the same problem of individual object recogni- tion. A major difference with the work described here lies in the probabilistic treatment of hypotheses, which allows us here to use directly hypothesis likelihood as a guide for the search, instead of the arbitrary admissible heuristic required by A*. 2 Probabilistic framework and notations Each model object is represented as a collection of features. Features are informative parts extracted from images by an interest point operator. Each model is the set of features extracted from one training image of a given object - although this could be generalized to features from many images of the same object. Models are indexed by k and denoted by mk, while indices i and j are used respectively for features extracted from the test image and from model images: fi denotes the i - th test feature, while f k j denotes the j - th feature from the k - th model. The features extracted from model images (training set) form the database. A feature detected in a test image can be a consequence of the presence of a model object in the image, in which case it should be associated to a feature from the database. In the alternative, this feature is attributed to a clutter - or background - detection. The geometric information associated to each feature contains position information (x and y coordinates, denoted by the vector x), orientation (denoted by ) and scale (denoted by ). It is denoted by Xi = (x, i, i) for test feature fi and X k j = (xk j k j , k j ) for model feature f k j . This geometric information is measured relatively to the standard reference frame of the image in which the feature has been detected. All features extracted from the same image share the same reference frame. The appearance information associated to a feature is a descriptor characterizing the local image appearance near this feature. The measured appearance information is denoted by Ai for test feature fi and Akj for model feature f kj. In our experiments, features are detected at multiple scales at the extrema of difference-of-gaussians filtered versions of the image [4, 12]. The SIFT descriptor [4] is then used to characterize the local texture about keypoints. A partial hypothesis h explains the observations made in a fraction of the test image. It combines a model image mh and a corresponding set of pose parameters Xh. Xh encodes position, rotation, scale (this can easily be extended to affine transformations). We assume independence between partial hypotheses. This requires in particular independence be- tween models. Although reasonable, this approximation is not always true (e.g. a keyboard is likely to be detected close to a computer screen). This allows us to search in parallel for multiple objects in a test image. A hypothesis H is the combination of several partial hypotheses, such that it explains com- pletely the observations made in the test image. A special notation H 0 or h0 denotes any (partial) hypothesis that states that no model object is present in a given fraction of the test image, and that features that could have been detected there are due to clutter. Our objective is to find which model objects are present in the test scene, given the ob- servations made in the test scene and the information that is present in the database. In probabilistic terms, we look for hypotheses H for which the likelihood ration LR(H) = P (H|{fi},{fkj}) P (H0|{fi},{fkj}) > 1. This ratio characterizes how well models and poses specified by H explain the observations, as opposed to them being generated by clutter. Using Bayes rules and after simplifications, P (H|{fi}, {f k j }) P ({fi}|{f k j }, H ) P (H ) LR(H) = = (1) P (H0|{fi}, {f kj}) P ({fi}|{f k j }, H0) P (H0) where we used P ({f k j }|H ) = P ({f k j }) since the database observations do not depend on the current hypothesis. A key assumption of this work is that once the pose parameters of the objects (and thus their reference frames) are known, the geometric configuration and appearance of the test features are independent from each other. We also assume independence between features associated to models and features associated to clutter detections, as well as independence between separate clutter detections. Therefore, P ({fi}|{f k j }, H ) = i P (fi|{f k j }, H ). These assumptions of independence are also made in [13], and undelying in [4]. Assignment vectors v represent matches between features from the test scene, and model features or clutter. The dimension of each assignment vector is the number of test features ntest. Its i - th component v(i) = (k, j) denotes that the test feature fi is matched to fv(i) = f kj, j - th feature from model mk. v(i) = (0, 0) denotes the case where fi is attributed to clutter. The set VH of assignment vectors compatible with a hypothesis H are those that assign test features only to models present in H (and to clutter). In particular, the only assignment vector compatible with h0 is v0 such that i, v0(i) = (0, 0). We obtain P (fi|fv(i), mh, Xh) LR(H) = P (H) P (v|{fk j }, mh, Xh) (2) P (H0) P (f vV i|h0) H hH i|fih P (H) is a prior on hypotheses, we assume it is constant. The term P (v|{f k j }, mh, Xh) is discussed in 3.1, we now explore the other terms. P (fi|fv(i), mh, Xh) : fi and fv(i) are believed to be one and the same feature. Differences measured between them are noise due to the imaging system as well as distortions caused by viewpoint or lighting conditions changes. This noise probability p n encodes differences in appearance of the descriptors, but also in geometry, i.e. position, scale, orientation Assuming independence between appearance information and geometry information, pn(fi|f k j , mh, Xh) = pn,A(Ai|Av(i), mh, Xh) pn,X (Xi|Xv(i), mh, Xh) (3) Figure 2: Snapshots from the iterative matching process. Two competing hypotheses are displayed (top and bottom row) a) Each assignment vector contains one assignment, suggesting a transformation (red box) b) End of iterative process. The correct hypothesis is supported by numerous matches and high belief, while the wrong hypothesis has only a weak support from few matches and low belief. The error in geometry is measured by comparing the values observed in the test image, with the predicted values that would be observed if the model features were to be trans- formed according to the parameters Xh. Let's denote by Xh(xv(i)),Xh(v(i)),Xh(v(i)) those predicted values, the geometry part of the noise probability can be decomposed into pn,X (Xi|Xv(i), h) = pn,x(xi, Xh(xv(i))) pn,(i, Xh(v(i))) pn,(i, Xh(v(i))) (4) P (fi|h0) is a density on appearance and position of clutter detections, denoted by p bg(fi). We can decompose this density as well into an appearance term and a geometry term: pbg(fi) = pbg,A(Ai) pbg,X (Xi) = pbg,A(Ai) pbg,(x)(xi) pbg,(i) pbg,(i) (5) pbg,A, pbg,(x)(xi) pbg,(i), pbg,(i) are densities that characterize, for clutter detections, appearance, position, scale and rotation respectively. Out of lack of space, and since it is not the main focus of this paper, we will not go into the details of how the "foreground density" p n and the "background density" pbg are learned. The main assumption is that those densities are shared across features, instead of having one set of parameters for each feature as in [1, 2]. This results in an important decrease of the number of parameters to be learned, at a slight cost in the model expressiveness. 3 Search for the best interpretation of the test image The building block of the recognition process is a question, comparing a feature from a database model with a feature of the test image. A question selects a feature from the database, and tries to identify if and where this feature appears in the test image. 3.1 Assignment vectors compatible with hypotheses For a given hypothesis H, the set of possible assignment vectors V H is too large for explicit exploration. Indeed, each potential match can either be accepted or rejected, which creates a combinatorial explosion. Hence, we approximate the summation in (2) by its largest term. In particular, each assignment vector v and each model referenced in v implies a set of pose parameters Xv (extracted e.g. with least-squares fitting). Therefore, the term P (v|{f k j }, mh, Xh) from (2) will be significant only when Xv Xh, i.e. when the pose implied by the assignment vector agrees with the pose specified by the partial hypothesis. We consider only the assignment vectors v for which Xv Xh. P (vH|{f k j }, h) is assumed to be close to 1. Eq.(2) becomes P (fi|fv LR(H) P (H) h(i), mh, Xh) P (H0) P (f hH i|h i|f 0) (6) ih Our recognition system proceeds by asking questions sequentially and adding matches to assignment vectors. It is therefore natural to define, for a given hypothesis H and the corresponding assignment vector vH and t ntest, the belief in vH by pn(ft|fv(t), mh , Xh ) B t t 0(vH ) = 1, Bt(vH ) = Bt-1(vH) (7) pbg(ft|h0) The geometric part of the belief (cf.(3)-(5) characterizes how close the pose X v implied by the assignments is to the pose Xh specified by the hypothesis. The geometric component of the belief characterizes the quality of the appearance match for the pairs (f i, fv(i)). 3.2 Entropy-based optimization Our goal is finding quickly the hypothesis that best explains the observations, i.e. the hy- pothesis (models+poses) that has the highest likelihood ratio. We compute such hypothesis incrementally by asking questions sequentially. Each time a question is asked we update the beliefs. We stop the process and declare a detection (i.e. a given model is present in the image) as soon as the belief of a corresponding hypothesis exceeds a given confidence threshold. The speed with which we reach such a conclusion depends on choosing cleverly the next question. A greedy strategy says that the best next question is the one that takes us closest to a detection decision. We do so by considering the entropy of the vector of beliefs (the vector may be normalized to 1 so that each belief is in fact a probability): the lower the entropy the closer we are to a detection. Therefore we study the following heuristic: The most informative next question is the one that minimizes the expectation of the entropy of our beliefs. We call this strategy `minimum expected entropy' (MEE). This idea is due to Geman et al. [14]. Calculating the MEE question is, unfortunately, a complex and expensive calculation in itself. In Monte-Carlo simulations of a simplified version of our problem we notice that the MEE strategy tends to ask questions that relate to the maximum-belief hypothesis. Therefore we approximate the MEE strategy with a simple heuristic: The next question consists of attempting to match one feature of the highest-belief model; specifically, the feature with best appearance match to a feature in the test image. 3.3 Search for the best hypotheses In an initialization step, a geometric hash table [3, 6, 7] is created by discretizing the space of possible transformations Note that we add only partial hypotheses in a hypothesis one at a time, which allows us to discretize only the space of partial hypotheses (models + poses), instead of discretizing the space of combinations of partial hypotheses. Questions to be examined are created by pairing database features to the test features clos- est in terms of appearance. Note that since features encode location, orientation and scale, any single assignment between a test feature and a model feature contains enough infor- mation to characterize a similarity transformation. It is therefore natural to restrict the set of possible transformations to similarities, and to insert each candidate assignment in the corresponding geometric hash table entry. This forms a pool of candidate assignments. The set of hypotheses is initialized to the center of the hash table entries, and their belief is set to 1. The motivation for this initialization step is to examine, for each partial hypothesis, only a small number of candidate matches. A partial hypothesis corresponds to a hash table entry, we consider only the candidate assignments that fall into this same entry. Each iteration proceeds as follows. The hypothesis H that currently has the highest likeli- hood ratio is selected. If the geometric hash table entry corresponding to the current partial hypothesis h, contains candidate assignments that have not been examined yet, one of them, ( m fi, f h j ) is picked - currently, the best appearance match - and the probabilities p bg(fi) m and pn(fi|f h j , mh, Xh) are computed. As mentioned in 3.1, only the best assignment Figure 3: Results from our algorithm in various situations (viewpoint change can be seen in Fig.6). Each row shows the best hypothesis in terms of belief. a) Occlusion b) Change of scale. Figure 4: ROC curves for both experiments. The performance improvement from our probabilistic formulation is particularly significant when a low false alarm rate is desired. The threshold used is the repeatability rate defined in [15] m vector is explored: if pn(fi|f h j , mh, Xh) > pbg(fi) the match is accepted and inserted in m the hypothesis. In the alternative, fi is considered a clutter detection and f h j is a missed detection. The belief B(vH ) and the likelihood ratio LR(H) are updated using (7). After adding an assignment to a hypothesis, frame parameters X h are recomputed using least-squares optimization, based on all assignments currently associated to this hypothe- sis. This parameter estimation step provides a progressive refinement of the model pose parameters as assignments are added. Fig.2 illustrates this process. The exploration of a partial hypothesis ends when no more candidate match is available in the hash table entry. We proceed with the next best partial hypothesis. The search ends when all test scene features have been matched or assigned to clutter. 4 Experimental results 4.1 Experimental setting We tested our algorithm on two sets of images, containing respectively 49 and 161 model images, and 101 and 51 test images (sets P M - gadgets - 03 and JP - 3Dobjects - 04 available from http : //www.vision.caltech.edu/html - f iles/arc hive.html). Each model image contained a single object. Test images contained from zero (negative exam- ples) to five objects, for a total of 178 objects in the first set, and 79 objects in the second set. A large fraction of each test image consists of background. The images were taken with no precautions relatively to lighting conditions or viewing angle. The first set contains common kitchen items and objects of everyday use. The second set (Ponce Lab, UIUC) includes office pictures. The objects were always moved between model images and test images. The images of model objects used in the learning stage were downsampled to fit in a 500 500 pixels box, the test images were downsampled to 800 800 pixels. With these settings, the number of features generated by the features detector was of the order of 1000 per training image and 2000-4000 per test image. Figure 5: Behavior induced by clutter detections. A ground truth model was created by cutting a rectangle from the test image and adding noise. The recognition process is therefore expected to find a perfect match. The two rows show the best and second best model found by each algorithm (estimated frame position shown by the red box, features that found a match are shown in yellow). Pierre Moreels, Pietro Perona |
NIPS | 2 |
| 2004 | Self-Tuning Spectral ClusteringabstractWe study a number of open issues in spectral clustering: (i) Selecting the appropriate scale of analysis, (ii) Handling multi-scale data, (iii) Cluster- ing with irregular background clutter, and, (iv) Finding automatically the number of groups. We first propose that a ‘local’ scale should be used to compute the affinity between each pair of points. This local scaling leads to better clustering especially when the data includes multiple scales and when the clusters are placed within a cluttered background. We further suggest exploiting the structure of the eigenvectors to infer automatically the number of groups. This leads to a new algorithm in which the final randomly initialized k-means stage is eliminated. Lihi Zelnik-Manor, Pietro Perona |
NIPS | 2 |
| 2003 | Object Class Recognition by Unsupervised Scale-Invariant LearningabstractWe present a method to learn and recognize object class models from unlabeled and unsegmented cluttered scenes in a scale invariant manner. Objects are modeled as flexible constellations of parts. A probabilistic representation is used for all aspects of the object: shape, appearance, occlusion and relative scale. An entropy-based feature detector is used to select regions and their scale within the image. In learning the parameters of the scale-invariant object model are estimated. This is done using expectation-maximization in a maximum-likelihood setting. In recognition, this model is used in a Bayesian manner to classify images. The flexible nature of the model is demonstrated by excellent results over a range of datasets including geometrically constrained classes (e.g. faces, cars) and flexible objects (such as animals). Rob Fergus, Pietro Perona, Andrew Zisserman |
CVPR (2) | 2 |
| 2003 | A Bayesian Approach to Unsupervised One-Shot Learning of Object CategoriesabstractLearning visual models of object categories notoriously requires thousands of training examples; this is due to the diversity and richness of object appearance which requires models containing hundreds of parameters. We present a method for learning object categories from just a few images (1 /spl sim/ 5). It is based on incorporating "generic" knowledge which may be obtained from previously learnt models of unrelated categories. We operate in a variational Bayesian framework: object categories are represented by probabilistic models, and "prior" knowledge is represented as a probability density function on the parameters of these models. The "posterior" model for an object category is obtained by updating the prior in the light of one or more observations. Our ideas are demonstrated on four diverse categories (human faces, airplanes, motorcycles, spotted cats). Initially three categories are learnt from hundreds of training examples, and a "prior" is estimated from these. Then the model of the fourth category is learnt from 1 to 5 training examples, and is used for detecting new exemplars a set of test images. Li Fei-Fei 0001, Rob Fergus, Pietro Perona |
ICCV | 3 |
| 2003 | An Improved Scheme for Detection and Labelling in Johansson DisplaysabstractConsider a number of moving points, where each point is attached to a joint of the human body and projected onto an image plane. Johannson showed that humans can effortlessly detect and recog- nize the presence of other humans from such displays. This is true even when some of the body points are missing (e.g. because of occlusion) and unrelated clutter points are added to the display. We are interested in replicating this ability in a machine. To this end, we present a labelling and detection scheme in a probabilistic framework. Our method is based on representing the joint prob- ability density of positions and velocities of body points with a graphical model, and using Loopy Belief Propagation to calculate a likely interpretation of the scene. Furthermore, we introduce a global variable representing the body’s centroid. Experiments on one motion-captured sequence suggest that our scheme improves on the accuracy of a previous approach based on triangulated graph- ical models, especially when very few parts are visible. The im- provement is due both to the more general graph structure we use and, more significantly, to the introduction of the centroid variable. Claudio Fanti, Marzia Polito, Pietro Perona |
NIPS | 3 |
| 2003 | Mutual Boosting for Contextual InferenceabstractMutual Boosting is a method aimed at incorporating contextual information to augment object detection. When multiple detectors of objects and parts are trained in parallel using AdaBoost [1], object detectors might use the remaining intermediate detectors to enrich the weak learner set. This method generalizes the efficient features suggested by Viola and Jones thus enabling information inference between parts and objects in a compositional hierarchy. In our experiments eye-, nose-, mouth- and face detectors are trained using the Mutual Boosting framework. Results show the method outperforms applications overlooking contextual information. We suggest that achieving contextual integration is a step toward human-like detection capabilities. Michael Fink 0002, Pietro Perona |
NIPS | 2 |
| 2003 | Visual Identification by Signature TrackingabstractWe propose a new camera-based biometric: visual signature identification. We discuss the importance of the parameterization of the signatures in order to achieve good classification results, independently of variations in the position of the camera with respect to the writing surface. We show that affine arc-length parameterization performs better than conventional time and Euclidean arc-length ones. We find that the system verification performance is better than 4 percent error on skilled forgeries and 1 percent error on random forgeries, and that its recognition performance is better than 1 percent error rate, comparable to the best camera-based biometrics. Mario E. Munich, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Unsupervised Learning of Human MotionabstractAn unsupervised learning algorithm that can obtain a probabilistic model of an object composed of a collection of parts (a moving human body in our examples) automatically from unlabeled training data is presented. The training data include both useful "foreground" features as well as features that arise from irrelevant background clutter - the correspondence between parts and detected features is unknown. The joint probability density function of the parts is represented by a mixture of decomposable triangulated graphs which allow for fast detection. To learn the model structure as well as model parameters, an EM-like algorithm is developed where the labeling of the data (part assignments) is treated as hidden variables. The unsupervised learning technique is not limited to decomposable triangulated graphs. The efficiency and effectiveness of our algorithm is demonstrated by applying it to generate models of human motion automatically from unlabeled image sequences, and testing the learned models on a variety of sequences. Yang Song 0023, Luis Goncalves, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | Local Analysis for 3D Reconstruction of Specular Surfaces - Part II
Silvio Savarese, Pietro Perona |
ECCV (2) | 2 |
| 2002 | A Digital Antennal Lobe for Pattern Equalization: Analysis and DesignabstractRe-mapping patterns in order to equalize their distribution may greatly simplify both the structure and the training of classifiers. Here, the properties of one such map obtained by running a few steps of discrete-time dynamical system are explored. The system is called 'Digital Antennal Lobe' (DAL) because it is inspired by recent studies of the antennallobe, a structure in the olfactory sys(cid:173) tem of the grasshopper. The pattern-spreading properties of the DAL as well as its average behavior as a function of its (few) de(cid:173) sign parameters are analyzed by extending previous results of Van Vreeswijk and Sompolinsky. Furthermore, a technique for adapting the parameters of the initial design in order to obtain opportune noise-rejection behavior is suggested. Our results are demonstrated with a number of simulations. Alex Holub, Gilles Laurent 0001, Pietro Perona |
NIPS | 3 |
| 2002 | Special Issue on Partial Differential Equations in Image Processing, Computer Vision, and Computer Graphics
Olivier D. Faugeras, Pietro Perona, Guillermo Sapiro |
J. Vis. Commun. Image Represent. | 2 |
| 2002 | Visual Input for Pen-Based ComputersabstractThe design and implementation of a camera-based, human-computer interface for acquisition of handwriting is presented. The camera focuses on a standard sheet of paper and images a common pen; the trajectory of the tip of the pen is tracked and the contact with the paper is detected. The recovered trajectory is shown to have sufficient spatio-temporal resolution and accuracy to enable handwritten character recognition. More than 100 subjects have used the system and have provided a large and heterogeneous set of examples showing that the system is both convenient and accurate. Mario E. Munich, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Local Analysis for 3D Reconstruction of Specular SurfacesabstractWe explore the geometry linking the shape of a curved mirror surface to the distortions it produces on a scene it reflects. Our analysis is local and differential. We assume a simple calibrated scene composed of lines passing through a point. We demonstrate that local information about the geometry of the surface may be recovered up to the second order from either the orientation and curvature of the images of two intersecting lines, or from the orientation of the images of three or more intersecting lines. An explicit solution for calculating shape and position of spherical mirror surfaces is given. Silvio Savarese, Pietro Perona |
CVPR (2) | 2 |
| 2001 | Learning Probabilistic Structure for Human Motion DetectionabstractDecomposable triangulated graphs have been shown to be efficient and effective for modeling the probabilistic spatio-temporal structure of brief stretches of human motion. In previous work such model structure was handcrafted by expert human observers and labeled data were needed for parameter learning. We present a method to build automatically the structure of the decomposable triangulated graph from unlabeled data. It is based on maximum-likelihood. Taking the labeling of the data as hidden variables, a variant of the EM algorithm can be applied. A greedy algorithm is developed to search for the optimal structure of the decomposable model based on the (conditional) differential entropy of variables. Our algorithm is demonstrated by learning models of human motion completely automatically from unlabeled real image sequences with clutter and occlusion. Experiments on both motion captured data and grayscale image sequences show that the resulting models perform better than the hand-constructed models. Yang Song 0023, Luis Goncalves, Pietro Perona |
CVPR (2) | 3 |
| 2001 | Shadow Carving
Silvio Savarese, Holly E. Rushmeier, Fausto Bernardini, Pietro Perona |
ICCV | 4 |
| 2001 | Grouping and dimensionality reduction by locally linear embeddingabstract(LLE) Locally Linear Embedding is an elegant nonlinear dimensionality-reduction technique recently introduced by Roweis and Saul [2]. It fails when the data is divided into separate groups. We study a variant of LLE that can simultaneously group the data and calculate local embedding of each group. An estimate for the upper bound on the intrinsic dimension of the data set is obtained automatically. Marzia Polito, Pietro Perona |
NIPS | 2 |
| 2001 | Unsupervised Learning of Human Motion ModelsabstractThis paper presents an unsupervised learning algorithm that can derive the probabilistic dependence structure of parts of an object (a moving hu- man body in our examples) automatically from unlabeled data. The dis- tinguished part of this work is that it is based on unlabeled data, i.e., the training features include both useful foreground parts and background clutter and the correspondence between the parts and detected features are unknown. We use decomposable triangulated graphs to depict the probabilistic independence of parts, but the unsupervised technique is not limited to this type of graph. In the new approach, labeling of the data (part assignments) is taken as hidden variables and the EM algo- rithm is applied. A greedy algorithm is developed to select parts and to search for the optimal structure based on the differential entropy of these variables. The success of our algorithm is demonstrated by applying it to generate models of human motion automatically from unlabeled real image sequences. Yang Song 0023, Luis Goncalves, Pietro Perona |
NIPS | 3 |
| 2001 | Monocular Perception of Biological Motion in Johansson Displays
Yang Song 0023, Luis Goncalves, Enrico Di Bernardo, Pietro Perona |
Comput. Vis. Image Underst. | 4 |
| 2000 | Towards Detection of Human MotionabstractDetecting humans in images is a useful application of computer vision. Loose and textured clothing, occlusion and scene clutter make it a difficult problem because bottom-up segmentation and grouping do not always work. We address the problem of detecting humans from their motion pattern in monocular image sequences; extraneous motions and occlusion may be present. We assume that we may not rely on segmentation or grouping and that the vision front-end is limited to observing the motion of key points and textured patches in between pairs of frames. We do not assume that we are able to track features for more than two frames. Our method is based on learning an approximate probabilistic model of the joint position and velocity of different body features. Detection is performed by hypothesis testing on the maximum a posteriori estimate of the pose and motion of the body. Our experiments on a dozen of walking sequences indicate that our algorithm is accurate and efficient. Yang Song 0023, Xiaolin Feng, Pietro Perona |
CVPR | 3 |
| 2000 | Towards Automatic Discovery of Object CategoriesabstractWe propose a method to learn heterogeneous models of object classes for visual recognition. The training images contain a preponderance of clutter and learning is unsupervised. Our models represent objects as probabilistic constellations of rigid parts (features). The variability within a class is represented by a join probability density function on the shape of the constellation and the appearance of the parts. Our method automatically identifies distinctive features in the training set. The set of model parameters is then learned using expectation maximization. When trained on different, unlabeled and unsegmented views of a class of objects, each component of the mixture model can adapt to represent a subset of the views. Similarly, different component models can also "specialize" on sub-classes of an object class. Experiments on images of human heads, leaves from different species of trees, and motor-cars demonstrate that the method works well over a wide variety of objects. Max Welling, Pietro Perona |
CVPR | 3 |
| 2000 | Monocuolar Perception of Biological Motion - Clutter and Partial Occlusion
Yang Song 0023, Luis Goncalves, Pietro Perona |
ECCV (2) | 3 |
| 2000 | Unsupervised Learning of Models for Recognition
Max Welling, Pietro Perona |
ECCV (1) | 3 |
| 2000 | Viewpoint-Invariant Learning and Detection of Human HeadsabstractWe present a method to learn models of human heads for the purpose of detection from different viewing angles. We focus on a model where objects are represented as constellations of rigid features (parts). Variability is represented by a joint probability density function (PDF) on the shape of the constellation. In the first stage, the method automatically identifies distinctive features in the training set using an interest operator followed by vector quantization. The set of model parameters, including the shape PDF, is then learned using expectation maximization. Experiments show good generalization performance to novel viewpoints and unseen faces. Performance is above 90% correct with less than 1 s computation time per image. Wolfgang Einhäuser, Max Welling, Pietro Perona |
FG | 4 |
| 2000 | An Intelligent Vision-Only Operator Interface for Dexterous RobotsabstractThe development of a vision-only intelligent interface for dexterous robots is described, capable of tracking operator motions: evaluating trajectory feasibility, and providing visual feedback to the operator. This interface analyses the operator's arm motion for safety, and then converts it into trajectory commands for a mechanical arm. Potentially dangerous situations, such as singularity and self collision, are displayed to the operator as graphical icons superimposed over the robot video images. The paper summarizes the main features of the current implementation, presents the approach developed for singularity identification and describes the performance of the demonstration prototype. Paolo Fiorini, Gene Chalfant, Yuichi Tsumaki, Enrico Di Bernardo, Pietro Perona |
ICRA | 5 |
| 2000 | Bayesian reasoning on qualitative descriptions from images and speech
Gudrun Socher, Gerhard Sagerer, Pietro Perona |
Image Vis. Comput. | 3 |
| 2000 | Special Issue on the Second International Conference on Scale Space Theory in Computer Vision: Guest Editors' Comments
Olivier D. Faugeras, Mads Nielsen, Pietro Perona, Bart M. ter Haar Romeny, Guillermo Sapiro |
J. Vis. Commun. Image Represent. | 3 |
| 2000 | Call for Papers: Special Issue on Partial Differential Equations in Image Processing, Computer Vision, and Computer Graphics
Olivier D. Faugeras, Pietro Perona, Guillermo Sapiro |
J. Vis. Commun. Image Represent. | 2 |
| 1999 | What Do Planar Shadows Tell About Scene Geometry?abstractA method for reconstructing 3D scene geometry from a set of projected shadows is presented. It is composed of two stages. First, the scene geometry is retrieved up to three scalar unknowns using only the information contained in the observed shadow edges on the image plane. Then, the three remaining unknowns are computed making use of the known depths at three points. This technique improves upon previous results in that it does not require the presence of a reference plane in the background. A mathematical analysis is presented using dual-space geometry, a formalism that provides adequate tools to carry out all the derivations in a compact and intuitive manner. A linear algorithm based on singular value decomposition (SVD) is presented leading to a closed form solution for reconstruction. Jean-Yves Bouguet, Pietro Perona |
CVPR | 3 |
| 1999 | Visual Signature Verification Using Affine Arc-LengthabstractSignatures can be acquired with a camera-based system with enough resolution to perform verification. This paper presents the performance of a visual-acquisition signature verification system, emphasizing on the importance of the parameterisation of the signature in order to achieve good classification results. A technique to overcome the lack of examples in order to estimate the generalization error of the algorithm is also described. Mario E. Munich, Pietro Perona |
CVPR | 2 |
| 1999 | Continuous Dynamic Time Warping for Translation-Invariant Curve Alignment with Applications to Signature VerificationabstractThe problem of establishing correspondence and measuring the similarity of a pair of planar curves arises in many applications in computer vision and pattern recognition. This paper presents a new method for comparing planar curves and for performing matching at sub-sampling resolution. The analysis of the algorithm as well as its structural properties are described. The performance of the new technique applied to the problem of signature verification is shown and compared with the performance of the well-known Dynamic Time Warping algorithm. Mario E. Munich, Pietro Perona |
ICCV | 2 |
| 1999 | Monocular Perception of Biological Motion - Detection and LabelingabstractComputer perception of biological motion is key to developing convenient and powerful human-computer interfaces. Successful body tracking algorithms have been developed; however initialization is done by hand. We propose a method for detecting a moving human body and for labeling its parts automatically. It is based on maximizing the joint probability density function (PDF) of the position and velocity of the body parts. The PDF is estimated from training data. Dynamic programming is used for calculating efficiently the best global labeling on an approximation of the PDF. The computational cost is on the order of N/sup 4/ where N is the number of features detected. We explore the performance of our method with experiments carried on a variety of periodic and non-periodic body motions viewed monocularly for a total of approximately 30,000 frames. Point-markers were strapped to the joints of the subject for facilitating image analysis. We find an average of 2.3% labeling error; the experiments also suggest a high degree of viewpoint-invariance. Yang Song 0023, Luis Goncalves, Enrico Di Bernardo, Pietro Perona |
ICCV | 4 |
| 1999 | 3D Photography Using Shadows in Dual-Space Geometry
Jean-Yves Bouguet, Pietro Perona |
Int. J. Comput. Vis. | 2 |
| 1998 | Real-Time 2-D Feature Detection on a Reconfigurable ComputerabstractWe have designed and implemented a system for real-time detection of 2-D features on a reconfigurable computer based on Field Programmable Gate Arrays (FPGA's). We envision this device as the front-end of a system able to track image features in real-time control applications like autonomous vehicle navigation. The algorithm employed to select good features is inspired by Tomasi and Kanade's method. Compared to the original method, the algorithm that we have devised does not require any floating point or transcendental operations, and can be implemented either in hardware or in software. Moreover, it maps efficiently into a highly pipelined architecture, well suited to implementation in FPGA technology. We have implemented the algorithm on a low-cost reconfigurable computer and have observed reliable operation on an image stream generated by a standard NTSC video camera at 30 Hz. Arrigo Benedetti, Pietro Perona |
CVPR | 2 |
| 1998 | Using Hierarchical Shape Models to Spot Keywords in Cursive Handwriting DataabstractDifferent instances of a handwritten word consist of the same basic features (humps, cusps, crossings, etc.) arranged in a deformable spatial pattern. Thus, keywords in cursive text can be detected by looking for the appropriate features in the "correct" spatial configuration. A keyword can be modeled hierarchically as a set of word fragments, each of which consists of lower-level features. To allow flexibility, the spatial configuration of keypoints within a fragment is modeled using a Dryden-Mardia (DM) probability density over the shape of the configuration. In a writer-dependent test on a transcription of the Declaration of Independence (/spl sim/1300 words, /spl sim/7500 characters), the method detected all eleven instances of the keyword "government" with only four false positives. Michael C. Burl, Pietro Perona |
CVPR | 2 |
| 1998 | Scene Segmentation from 3D MotionabstractWe propose an EM approach combined with the modified separation matrix scheme to perform 3D motion segmentation of the image sequence which contains multiple moving objects. We observe that, given the detected features and their 2D optical flow, in most cases the objects or their flow are separated very well from each other in space. The separation matrix method modified by using normalized cuts achieves expected grouping results for these cases. However, when the objects are overlapped spatially but undergoing independent motions, such as Ullman's co-axial transparent cylinder demonstration, there will be no proper affinity to perform segmentation by that way. Pure underlying 3D motion becomes the only cite to segment the scene. We exploit the EM algorithm to deal with such difficult cases. The scheme is tested on vast number of synthetic image sequences. Results with real image sequence are also given. Xiaolin Feng, Pietro Perona |
CVPR | 2 |
| 1998 | Probablistic Affine Invariants for RecognitionabstractUnder a weak perspective camera model, the image plane coordinates in different views of a planar object are related by an affine transformation. Because of this property, researchers have attempted to use affine invariants for recognition. However, there are two problems with this approach: (1) objects or object classes with inherent variability cannot be adequately treated using invariants; and (2) in practice the calculated affine invariants can be quite sensitive to errors in the image plane measurements. In this paper we use probability distributions to address both of these difficulties. Under the assumption that the feature positions of a planar object can be modeled using a jointly Gaussian density, we have derived the joint density over the corresponding set of affine coordinates. Even when the assumptions of a planar object and a weak perspective camera model do not strictly hold, the results are useful because deviations from the ideal can be treated as deformability in the underlying object model. Thomas K. Leung, Michael C. Burl, Pietro Perona |
CVPR | 3 |
| 1998 | A Probabilistic Approach to Object Recognition Using Local Photometry and Global Geometry
Michael C. Burl, Pietro Perona |
ECCV (2) | 3 |
| 1998 | Camera-Based ID Verification by Signature Tracking
Mario E. Munich, Pietro Perona |
ECCV (1) | 2 |
| 1998 | A Factorization Approach to Grouping
Pietro Perona, William T. Freeman |
ECCV (1) | 1 |
| 1998 | Reach Out and Touch Space (Motion Learning)
Luis Goncalves, Enrico Di Bernardo, Pietro Perona |
FG | 3 |
| 1998 | 3D Photography on Your DeskabstractA simple and inexpensive approach for extracting the three-dimensional shape of objects is presented. It is based on 'weak structured lighting'; it differs from other conventional structured lighting approaches in that it requires very little hardware besides the camera: a desk-lamp, a pencil and a checker-board. The camera faces the object, which is illuminated by the desk-lamp. The user moves a pencil in front of the light source casting a moving shadow on the object. The 3D shape of the object is extracted from the spatial and temporal location of the observed shadow. Experimental results are presented on three different scenes demonstrating that the error in reconstructing the surface is less than 1%. Jean-Yves Bouguet, Pietro Perona |
ICCV | 2 |
| 1998 | Learning to Recognize Volcanoes on Venus
Michael C. Burl, Lars Asker, Padhraic Smyth, Usama M. Fayyad, Pietro Perona, Larry Crumpler, Jayne Aubele |
Mach. Learn. | 5 |
| 1998 | Reducing "Structure From Motion": A General Framework for Dynamic Vision Part 1: ModelingabstractThe literature on recursive estimation of structure and motion from monocular image sequences comprises a large number of apparently unrelated models and estimation techniques. We propose a framework that allows us to derive and compare all models by following the idea of dynamical system reduction. The "natural" dynamic model, derived from the rigidity constraint and the projection model, is first reduced by explicitly decoupling structure (depth) from motion. Then, implicit decoupling techniques are explored, which consist of imposing that some function of the unknown parameters is held constant. By appropriately choosing such a function, not only can we account for models seen so far in the literature, but we can also derive novel ones. Stefano Soatto, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Reducing "Structure From Motion": A General Framework for Dynamic Vision Part 2: Implementation and Experimental AssessmentabstractFor pt.1 see ibid., p.933-42 (1998). A number of methods have been proposed in the literature for estimating scene-structure and ego-motion from a sequence of images using dynamical models. Despite the fact that all methods may be derived from a "natural" dynamical model within a unified framework, from an engineering perspective there are a number of trade-offs that lead to different strategies depending upon the applications and the goals one is targeting. We want to characterize and compare the properties of each model such that the engineer may choose the one best suited to the specific application. We analyze the properties of filters derived from each dynamical model under a variety of experimental conditions, assess the accuracy of the estimates, their robustness to measurement noise, sensitivity to initial conditions and visual angle, effects of the bas-relief ambiguity and occlusions, dependence upon the number of image measurements and their sampling rate. Stefano Soatto, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Correction to: "Reducing 'Structure From Motion': A General Framework for Dynamic Vision Part 2: Implementation and Experimental Assessment"
Stefano Soatto, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Orientation diffusionsabstractDiffusions are useful for image processing and computer vision because they provide a convenient way of smoothing noisy data, analyzing images at multiple scales, and enhancing discontinuities. A number of diffusions of image brightness have been defined and studied so far; they may be applied to scalar and vector-valued quantities that are naturally associated with intervals of either the real line, or other flat manifolds. Some quantities of interest in computer vision, and other areas of engineering that deal with images, are defined on curved manifolds;typical examples are orientation and hue that are defined on the circle. Generalizing brightness diffusions to orientation is not straightforward, especially in the case where a discrete implementation is sought. An example of what may go wrong is presented.A method is proposed to define diffusions of orientation-like quantities. First a definition in the continuum is discussed, then a discrete orientation diffusion is proposed. The behavior of such diffusions is explored both analytically and experimentally. It is shown how such orientation diffusions contain a nonlinearity that is reminiscent of edge-process and anisotropic diffusion. A number of open questions are proposed at the end. Pietro Perona |
IEEE Trans. Image Process. | 1 |
| 1997 | Orientation diffusionsabstractDiffusions provide a convenient way of smoothing noisy brightness images, of analyzing images at multiple scales, and of enhancing discontinuities. Some quantities of interest in computer vision are defined on curved manifolds; typical examples are orientation and hue that are defined on the circle. Generalizing diffusions to orientation is not straightforward, especially in the case where a discrete implementation is sought. An example of what may go wrong is presented. A method is proposed to define diffusions of orientation-like quantities. First a definition in the continuum is discussed, then a discrete orientation diffusion is proposed. The behavior of such diffusions is explored both analytically and experimentally. It is shown how such orientation diffusions contain a nonlinearity that is reminiscent of edge-process and anisotropic diffusion. A number of open questions are proposed at the end. Pietro Perona |
CVPR | 1 |
| 1996 | Recognition of Planar Object ClassesabstractWe present a new framework for recognizing planar object classes, which is based on local feature detectors and a probabilistic model of the spatial arrangement of the features. The allowed object deformations are represented through shape statistics, which are learned from examples. Instances of an object in an image are detected by finding the appropriate features in the correct spatial configuration. The algorithm is robust with respect to partial occlusion, detector false alarms, and missed features. A 94% success rate was achieved for the problem of locating quasi-frontal views of faces in cluttered scenes. Michael C. Burl, Pietro Perona |
CVPR | 2 |
| 1996 | Motion from fixationabstractWe study the problem of estimating rigid motion from a sequence of monocular perspective images obtained by navigating around an object while fixating a particular feature point. We cast the problem in the framework of "epipolar geometry", and propose a filter based upon implicit dynamical model for recursively estimating motion under the fixation constraint. This allows us to compare the quality of the estimates directly against the ones obtained assuming a general rigid motion simply by changing the geometry of the parameter space, while maintaining the same structure of the recursive estimator. We also present a closed-form static solution from two views, and a recursive estimator of the relative pose between the viewer and the scene. Stefano Soatto, Pietro Perona |
CVPR | 2 |
| 1996 | Reducing "structure from motion"abstractThe literature on recursive estimation of structure and motion from monocular image sequences comprises a large number of different models and estimation techniques. We propose a framework that allows us to derive and compare all models by following the idea of dynamical system reduction. The "natural" dynamic model, derived by the rigidity constraint and the perspective projection, is first reduced by explicitly decoupling structure (depth) from motion. Then implicit decoupling techniques are explored, which consist of imposing that some function of the unknown parameters is held constant. By appropriately choosing such a function, not only can we account for all models seen so far in the literature, but we can also derive novel ones. Casting all the different models in a common framework allows us to compare their geometric properties on common experimental grounds. Stefano Soatto, Pietro Perona |
CVPR | 2 |
| 1996 | Visual input for pen-based computersabstractHandwriting may be captured using a video camera, rather than the customary pressure-sensitive tablet. This paper presents a simple system based on correlation and recursive prediction methods that can track the tip of the pen in real time with sufficient spatio-temporal resolution and accuracy to enable handwritten character recognition. The system is tested on a large and heterogeneous set of examples and its performance is compared to that of three human operators and a commercial high-resolution pressure-sensitive tablet. Mario E. Munich, Pietro Perona |
ICIP (2) | 2 |
| 1996 | Monocular tracking of the human arm in 3D: real-time implementation and experimentsabstractWe have developed a system capable of tracking a human arm in 3D and in real time. The system is based on a previously developed algorithm for 3D tracking which requires only a monocular view and no special markers on the body. In this paper we describe our real-time system and the insights gained from real-time experimentation. Enrico Di Bernardo, Luis Goncalves, Pietro Perona |
ICPR | 3 |
| 1996 | Visual input for pen-based computersabstractHandwriting may be captured using a video camera, rather than the customary pressure-sensitive tablet. This paper presents a simple system based on correlation and recursive prediction methods that can track the tip of the pen in real time with sufficient spatio-temporal resolution and accuracy to enable handwritten character recognition. The system is tested on a large and heterogeneous set of examples and its performance is compared to that of three human operators and a commercial high-resolution pressure-sensitive tablet. Mario E. Munich, Pietro Perona |
ICPR | 2 |
| 1996 | Scale-Space Properties of Quadratic Feature DetectorsabstractWe consider the scale-space properties of quadratic feature detectors and, in particular, investigate whether, like linear detectors, they permit a scale selection scheme with the "causality property", which guarantees that features are never created as the scale is coarsened. We concentrate on the design of one dimensional detectors with two constituent filters, with the scale selection implemented as convolution and a scaling function. We consider two special cases of interest: the constituent filter pairs related by the Hilbert transform, and by the first spatial derivative. We show that, under reasonable assumptions, Hilbert-pair quadratic detectors cannot have the causality property. In the case of derivative-pair detectors, we describe a family of scaling functions related to fractional derivatives of the Gaussian that are necessary and sufficient for causality. In addition, we report experiments that show the effects of these properties in practice. We thus demonstrate that at least one class of quadratic feature detectors has the same desirable scaling property as the more familiar detectors based on linear filtering. Paul Kube, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1995 | Monocular Tracking of the Human Arm in 3DabstractWe address the problem of estimating the position and motion of a human arm in 3D without any constraints on its behavior and without the use of special markers. We model the arm as two truncated right-circular cones connected with spherical joints. We propose to use a recursive estimator for arm position, and to provide the estimator with error signals obtained by comparing the projected estimated arm position with that of the actual arm in the image. The system is demonstrated and tested on a real image sequence.> Enrico Di Bernardo, Luis Goncalves, Enrico Ursella, Pietro Perona |
ICCV | 4 |
| 1995 | Visual Navigation Using a Single CameraabstractWe assess the usefulness of monocular recursive motion estimation techniques for vehicle navigation in the absence of a model for the environment. For this purpose we extend a recently proposed recursive motion estimator, the Essential filter, to handle scale estimation. We examine experimentally the accuracy with which the motion and position of the vehicle may be computed on an 8000 frame indoors sequence. The issues of sampling time frequency and number of necessary features in the environment are addressed systematically.> Jean-Yves Bouguet, Pietro Perona |
ICCV | 2 |
| 1995 | Finding Faces in Cluttered Scenes Using Labeled Random Graph MatchingabstractAn algorithm for locating quasi-frontal views of human faces in cluttered scenes is presented. The algorithm works by coupling a set of local feature detectors with a statistical model of the mutual distances between facial features it is invariant with respect to translation, rotation (in the plane), and scale and can handle partial occlusions of the face. On a challenging database with complicated and varied backgrounds, the algorithm achieved a correct localization rate of 95% in images where the face appeared quasi-frontally.> Thomas K. Leung, Michael C. Burl, Pietro Perona |
ICCV | 3 |
| 1995 | Dynamic Rigid Motion Estimation from Weak Perspectiveabstract"Weak perspective" represents a simplified projection model that approximates the imaging process when the scene is viewed under a small viewing angle and its depth relief is small relative to its distance from the viewer. We study how to generate dynamic models for estimating rigid 3D motion from weak perspective. A crucial feature in dynamic visual motion estimation is to decouple structure from motion in the estimation model. The reasons are both geometric-to achieve global observability of the model-and practical, for a structure independent motion estimator allows us to deal with occlusions and appearance of new features in a principled way. It is also possible to push the decoupling even further, and isolate the motion parameters that are affected by the so called "bas relief ambiguity" from the ones that are not. We present a novel method for reducing the order of the estimator by decoupling portions of the state space from the time evolution of the measurement constraint. We use this method to construct an estimator of full rigid motion (modulo a scaling factor) on a six dimensional state space, an approximate estimator for a four dimensional subset of the motion space, and a reduced filter with only two states. The latter two are immune to the bas relief ambiguity. We compare strengths and weaknesses of each of the schemes on real and synthetic image sequences.> Stefano Soatto, Pietro Perona |
ICCV | 2 |
| 1995 | Pyramidal implementation of deformable kernelsabstractIn computer vision and increasingly, in rendering and image processing, it is useful to filter images with continuous rotated and scaled families of filters. For practical implementations, one can think of using a discrete family of filters, and then to interpolate from their outputs to produce the desired filtered version of the image. We propose a multirate implementation of deformable kernels, capable to further reduce the computational weight. The "basis" filters are applied to the different levels of a pyramidal decomposition. The new system is not shift-invariant-it suffers from "aliasing". We introduce a new quadratic error criterion which keeps into account the inherent system aliasing. By using hypermatrix and Kronecker algebra, we are able to cast the global optimization task into a multilinear problem. An iterative procedure ("pseudo-SVD") is used to minimize the overall quadratic approximation error. Roberto Manduchi, Pietro Perona |
ICIP | 2 |
| 1995 | Visual motion estimation from point features: unified viewabstractAll methods for recursive estimation of 3-D motion from sequences of perspective images of point-features may be cast within a common framework. The unifying concept is the decoupling of the states of the dynamic observer that estimates motion and structure parameters. Two techniques are possible: explicit decoupling, following the principles of the "reduced-order observer", and implicit, via stabilization (or "compensation"). While we know how to calculate explicit decoupling for a limited number of state variables combinations, for instance using the "essential constraint" of Longuet-Higgins (1981) or the "subspace constraint" of Heeger and Jepson (1992), implicit decoupling is always possible by stabilizing an appropriate smooth function of the motion parameters. We describe some of the most "natural" choices, which consist in compensating for the image-motion of a point, a line or a plane. All the models we derive are in the form of implicit dynamical systems with parameters on different manifolds. Estimating motion may be regarded as the identification of such models, which may be carried out using general methods available in the literature. Stefano Soatto, Pietro Perona |
ICIP (3) | 2 |
| 1995 | Deformable Kernels for Early VisionabstractEarly vision algorithms often have a first stage of linear-filtering that 'extracts' from the image information at multiple scales of resolution and multiple orientations. A common difficulty in the design and implementation of such schemes is that one feels compelled to discretize coarsely the space of scales and orientations in order to reduce computation and storage costs. A technique is presented that allows: 1) computing the best approximation of a given family using linear combinations of a small number of 'basis' functions; and 2) describing all finite-dimensional families, i.e., the families of filters for which a finite dimensional representation is possible with no error. The technique is based on singular value decomposition and may be applied to generating filters in arbitrary dimensions and subject to arbitrary deformations. The relevant functional analysis results are reviewed and precise conditions for the decomposition to be feasible are stated. Experimental results are presented that demonstrate the applicability of the technique to generating multiorientation multi-scale 2D edge-detection kernels. The implementation issues are also discussed.> Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1994 | Automating the hunt for volcanoes on VenusabstractOur long-term goal is to develop a trainable tool for locating patterns of interest in large image databases. Toward this goal we have developed a prototype system, based on classical filtering and statistical pattern recognition techniques, for automatically locating volcanoes in the Magellan SAR database of Venus. Training for the specific volcano-detection task is obtained by synthesizing feature templates (via normalization and principal components analysis) from a small number of examples provided by experts. Candidate regions identified by a focus of attention (FOA) algorithm are classified based on correlations with the feature templates. Preliminary tests show performance comparable to trained human observers.> Michael C. Burl, Usama M. Fayyad, Pietro Perona, Padhraic Smyth |
CVPR | 3 |
| 1994 | Overcomplete steerable pyramid filters and rotation invarianceabstractA given (overcomplete) discrete oriented pyramid may be converted into a steerable pyramid by interpolation. We present a technique for deriving the optimal interpolation functions (otherwise called 'steering coefficients'). The proposed scheme is demonstrated on a computationally efficient oriented pyramid, which is a variation on the Burt and Adelson (1983) pyramid. We apply the generated steerable pyramid to orientation-invariant texture analysis in order to demonstrate its excellent rotational isotropy. High classification rates and precise rotation identification are demonstrated.> Hayit Greenspan, Serge J. Belongie, Rodney M. Goodman, Pietro Perona, Subrata Rakshit, Charles H. Anderson |
CVPR | 4 |
| 1994 | X-Y separable pyramid steerable scalable kernelsabstractA new method for generating X-Y separable, steerable, scalable approximations of filter kernels is proposed which is based on a generalization of the singular value decomposition (SVD) to three dimensions. This "pseudo-SVD" improves upon a previous scheme due to Perona (1992) in that it reduces convolution time and storage requirements. An adaptation of the pseudo-SVD is proposed to generate steerable and scalable kernels which are suitable for use with a Laplacian pyramid. The properties of this method are illustrated experimentally in generating steerable and scalable approximations to an early vision edge-detection kernel.> Douglas Shy, Pietro Perona |
CVPR | 2 |
| 1994 | Scale-space Properties of Quadratic Edge Detectors
Paul Kube, Pietro Perona |
ECCV (1) | 2 |
| 1994 | Motion Estimation on the Essential Manifold
Stefano Soatto, Ruggero Frezza, Pietro Perona |
ECCV (2) | 3 |
| 1994 | Automated Analysis of Radar Imagery of Venus: Handling Lack of Ground TruthabstractLack of verifiable ground truth is a common problem in remote sensing image analysis. For example, consider the synthetic aperture radar (SAR) image data of Venus obtained by the Magellan spacecraft. Planetary scientists are interested in automatically cataloging the locations of all the small volcanoes in this data set; however, the problem is very difficult and cannot be performed with perfect reliability even by human experts. Thus, training and evaluating the performance of an automatic algorithm on this data set must be handled carefully. We discuss the use of weighted free-response receiver-operating characteristics (wFROCs) for evaluating detection performance when the "ground truth" is subjective. In particular, we evaluate the relative detection performance of humans and automatic algorithms. Our experimental results indicate that proper assessment of the uncertainty in "ground truth" is essential in applications of this nature.> Michael C. Burl, Usama M. Fayyad, Pietro Perona, Padhraic Smyth |
ICIP (3) | 3 |
| 1994 | Diffusion Networks for On-Chip Image Contrast NormalizationabstractA new method for normalizing and quantizing images is presented. The method is based on calculating a local reference frame for the image gray levels. The levels of the reference frame are calculated using biased diffusions that are linked to the original image. The method is conceived to be integrated with sensing elements on the image plane of a camera. Its mathematical properties are analyzed and its performance is experimentally demonstrated. A circuital implementation has been designed, constructed and tested; it consists of a 20 nodes 1-D non-linear resistive grid. Experimental results are shown.> Pietro Perona, Marco Tartagni |
ICIP (1) | 1 |
| 1994 | Dynamic Visual Motion Estimation from Subspace ConstraintsabstractThe problem of estimating rigid motion from projections may be characterized using a nonlinear dynamical system, composed of the rigid motion constraint and the perspective map. The time derivative of the output of such a system, which is called the "motion field" and approximated by the "optical flow", is bilinear in the motion parameters, and may be used to specify a subspace constraint on either the direction of translation or the inverse depth of the observed points. Estimating motion may then be formulated as an optimization task constrained on such a subspace. We pose the optimization problem in a system theoretic framework as the the identification of a nonlinear implicit dynamical system with parameters on a differentiable manifold, and use techniques which pertain to nonlinear estimation and identification theory to perform the optimization task in a principled manner. The application of a general method presented in by Soatto et al. (see 33rd. IEEE conf. on Decision and Control, 1994) results in a recursive and pseudo-optimal solution of the visual motion estimation problem, which has robustness properties far superior to other existing techniques we have implemented. Experiments on real and synthetic image sequences show very promising results in terms of robustness, accuracy and computational efficiency.> Stefano Soatto, Pietro Perona |
ICIP (1) | 2 |
| 1994 | Recursive Estimation of Camera Motion from Uncalibrated Image SequencesabstractWe describe a method for estimating the motion and structure of a scene from a sequence of images taken with a camera whose geometric calibration parameters are unknown. The scheme is based upon a recursive motion estimation scheme, called the "essential filter", extended according to the epipolar geometric representation presented by Faugeras, Luong, and Maybank (see Proc. of the ECCV92, vol.588 of LNCS, Springer Verlag, 1992) in order to estimate the calibration parameters as well. The motion estimates can then be fed into any "structure from motion" module that processes motion error, in order to recover the structure of the scene.> Stefano Soatto, Pietro Perona |
ICIP (3) | 2 |
| 1994 | Rotation invariant texture recognition using a steerable pyramidabstractA rotation-invariant texture recognition system is presented. A steerable oriented pyramid is used to extract representative features for the input textures. The steerability of the filter set allows a shift to an invariant representation via a DFT-encoding step. Supervised classification follows. State-of-the-art recognition results are presented on a 30 texture database with a comparison across the performance of the k-NN, backpropagation and rule-based classifiers. In addition, high accuracy estimation of the input rotation angle is demonstrated. Hayit Greenspan, Serge J. Belongie, Rodney M. Goodman, Pietro Perona |
ICPR (2) | 4 |
| 1994 | Inferring Ground Truth from Subjective Labelling of Venus ImagesabstractIn remote sensing applications "ground-truth" data is often used as the basis for training pattern recognition algorithms to gener(cid:173) ate thematic maps or to detect objects of interest. In practical situations, experts may visually examine the images and provide a subjective noisy estimate of the truth. Calibrating the reliability and bias of expert labellers is a non-trivial problem. In this paper we discuss some of our recent work on this topic in the context of detecting small volcanoes in Magellan SAR images of Venus. Empirical results (using the Expectation-Maximization procedure) suggest that accounting for subjective noise can be quite signifi(cid:173) cant in terms of quantifying both human and algorithm detection performance. Padhraic Smyth, Usama M. Fayyad, Michael C. Burl, Pietro Perona, Pierre Baldi |
NIPS | 4 |
| 1993 | Recursive motion and structure estimation with complete error characterizationabstractAn algorithm that performs recursive estimation of ego-motion and ambient structure from a stream of monocular perspective images of a number of feature points is presented. The algorithm is based on an extended Kalman filter (EKF) that integrates over time the instantaneous motion and structure measurements computed by a two-perspective-views step. The key features of the authors' filter are: global observability of the model, and complete online characterization of the uncertainty of the measurements provided by the two-views step. The filter is thus guaranteed to be well-behaved regardless of the particular motion undergone by the observer. Regions of motion space that do not allow recovery of structure (e.g., pure rotation) may be crossed while maintaining good estimates of structure and motion. Whenever reliable measurements are available they are exploited. The algorithm works well for arbitrary motions with minimal smoothness assumptions and no ad hoc tuning. Simulations are presented that illustrate these characteristics.> Stefano Soatto, Pietro Perona, Ruggero Frezza, Giorgio Picci |
CVPR | 2 |
| 1992 | Boundary Detection in Piecewise Homogeneous Textured Images
Stefano Casadei, Sanjoy K. Mitter, Pietro Perona |
ECCV | 3 |
| 1992 | Steerable-Scalable Kernels for Edge Detection and Junction Analysis
Pietro Perona |
ECCV | 1 |
| 1992 | Steerable-scalable kernels for edge detection and junction analysis
Pietro Perona |
Image Vis. Comput. | 1 |
| 1991 | Deformable kernels for early visionabstractA technique is presented that allows (1) computing the best approximation of a given family using linear combinations of a small number of basis functions; and (2) describing all finite-dimensional families, i.e. the families of filters for which a finite-dimensional representation is possible with no error. The technique is general and can be applied to generating filters in arbitrary dimensions. Experimental results that demonstrate the applicability of the technique to generating multi-orientation multiscale 2-D edge-detection kernels are presented. The implementation issues are also discussed.> Pietro Perona |
CVPR | 1 |
| 1990 | Detecting and localizing edges composed of steps, peaks and roofsabstractThe projection of depth or orientation discontinuities in a physical scene results in image intensity edges which are not ideal step edges but are more typically a combination of step, peak and roof profiles. Most edge detection schemes ignore the composite nature of these edges, resulting in systematic errors in detection and localization. The problem of detecting and localizing these edges is addressed, along with the problem of false responses in smoothly shaded regions with constant gradient of the image brightness. A class of nonlinear filters, known as quadratic filters, is appropriate for this task, while linear filters are not. Performance criteria are derived for characterizing the SNR, localization and multiple responses of these filters in a manner analogous to Canny's criteria for linear filters. A two-dimensional version of the approach is developed which has the property of being able to represent multiple edges at the same location and determine the orientation of each to any desired precision. This permits junctions to be localized without rounding. Experimental results are presented.> Pietro Perona, Jitendra Malik |
ICCV | 1 |
| 1990 | Scale-Space and Edge Detection Using Anisotropic DiffusionabstractA new definition of scale-space is suggested, and a class of algorithms used to realize a diffusion process is introduced. The diffusion coefficient is chosen to vary spatially in such a way as to encourage intraregion smoothing rather than interregion smoothing. It is shown that the 'no new maxima should be generated at coarse scales' property of conventional scale space is preserved. As the region boundaries in the approach remain sharp, a high-quality edge detector which successfully exploits global information is obtained. Experimental results are shown on a number of images. Parallel hardware implementations are made feasible because the algorithm involves elementary, local operations replicated over the image.> Pietro Perona, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1989 | A computational model of texture segmentationabstractAn algorithm for finding texture boundaries in images is developed on the basis of a computational model of human texture perception. The model consists of three stages: (1) the image is convolved with a bank of even-symmetric linear filters followed by half-wave rectification to give a set of responses; (2) inhibition, localized in space, within and among the neural response profiles results in the suppression of weak responses when there are strong responses at the same or nearby locations; and (3) texture boundaries are detected using peaks in the gradients of the inhibited response profiles. The model is precisely specified, equally applicable to grey-scale and binary textures, and is motivated by detailed comparison with psychophysics and physiology. It makes predictions about the degree of discriminability of different texture pairs which match very well with experimental measurements of discriminability in human observers. From a machine-vision point of view, the scheme is a high-quality texture-edge detector which works equally on images of artificial and natural scenes. The algorithm makes the use of simple local and parallel operations, which makes it potentially real-time.> Jitendra Malik, Pietro Perona |
CVPR | 2 |